Unlocking Decades of News Media Assets: Intelligent Media Asset Management for a Media Group
An LLM-powered media asset management platform unifies decades of news content and diverse media assets in one storage and processing foundation. AI and semantic search answer asset queries in seconds, while NLP and multimodal models auto-generate news drafts and match images — freeing professionals from repetitive busywork.
A media group with decades of content assets runs a news archive that spans text, images, audio and video — covering everything from the print era to the mobile internet age. For this organization, "storing the content" was never the problem. The real problem is "using it": when an editor needs a set of photos from twenty years ago to illustrate a story, can they actually find them? This project answered that question with an LLM-powered intelligent media asset management platform.
Background and pain points
Content asset management in media has long been plagued by structural problems:
- Huge assets, inefficient retrieval. Decades of news content and diverse media data come in different formats with incomplete tags. Traditional retrieval relies on manually entered keywords — editors routinely spend hours hunting for material, asking around and hoping for luck.
- Long creation pipelines, heavy repetition. Writing, sourcing images, matching images and reviewing are largely manual. One story pulls in multiple systems and multiple colleagues, and whether the material can be found decides how fast the story ships.
- Data assets asleep, commercial value unclaimed. A large share of the archive has never been reused. Its potential licensing value, reuse value and knowledge value all sit in a state of "nobody knows how to use it."
- Talent consumed by busywork. Journalists spend their time on mechanical tasks — searching, organizing, comparing — leaving little room for creativity and strategic judgment.
In one sentence: content isn't scarce; what's missing is a mechanism that turns "sleeping assets" into "productivity on demand."
What we did
Built on a unified storage foundation with AI at its core, the platform forms a complete chain from "storing" to "using":
- Unified storage and efficient processing. Decades of news content and multiple media data types are brought into one platform with a single cataloging and management system — content scattered across departments and media finally has "one master index."
- AI and intelligent retrieval. With natural language processing and multimodal understanding, the platform supports semantic — not just keyword — search. An editor describes the material in one sentence, and the system locates relevant content across text, images and video, answering asset queries in seconds.
- Auto-generated news drafts with matched images. NLP and multimodal models enter the creation loop: the platform drafts news stories from journalistic elements and automatically matches semantically relevant images, compressing the serial workflow of "write — find image — attach image" into a single generation step.
- Unlocking the commercial value of data assets. Beyond search and creation, the platform structures content features so that "what's usable, how it can be used, and who it's licensed to" becomes queryable and operable — providing the data foundation for content redevelopment and monetization.
Results
The value of the platform shows up as a double shift in content productivity and in how talent spends its time:
- Retrieval moves from "digging through the archive" to "instant hits." Semantic search replaces manual keywords, dramatically cutting the time editors spend finding material. Decades of accumulated content become genuinely available on demand for the first time.
- Creation and review get significantly faster. Auto-generated drafts and auto-matched images speed up story turnaround. Editors no longer start from a blank page, and review runs more smoothly on stable first drafts. The whole production cadence accelerates.
- Data assets get unlocked. The archive moves from dormant storage to searchable, operable assets. Potential commercial value becomes visible — and quantified — for the first time, opening new paths to monetization.
- Professionals return to high-value work. Mechanical search, organization and repetitive writing are taken over by the system. Journalists' time flows back to story planning, fact-checking and deep reporting — creative and strategic work.
For this media group, the platform is not just a productivity tool; it's a restructuring of how work gets done — give the mechanical labor to machines, and give people back to the content.
Lessons learned: from project to OntiCards
This project confirmed a rule we keep rediscovering: in content-intensive industries, AI's value is not in "generating" but in making generation grounded. Auto-written articles without accurate access to content assets produce text that merely looks like news. Only when a model can pinpoint material and understand semantic relationships does generation have quality.
That experience directly shaped how OntiCards works. Semantic search on top of unified storage maps to the entity and relationship modeling in data cards — the "which asset illustrates which story, what cites what" relationships are made explicit, and AI organizes content accordingly. Auto-generation and image matching map to what the Q&A engine and terminology bank enforce for consistent business definitions. And making content assets operable maps to the full loop — assets that are queryable, governable and reusable — under the data card system.
If your organization is also sitting on a mountain of dormant content, look at it from three angles: Is there a unified description system for content? Is retrieval semantic or only keyword-based? Which steps in the creation pipeline could be automated? The clearer the answers, the shorter the path to unlocking your assets.