Hollywood studios have decades of film footage. Broadcasters sit on vast news archives. Podcast networks have accumulated thousands of hours of conversations. YouTube creators have years of videos that attract little attention after publication.
All this archived content is valuable to AI companies seeking to train their models on how the real world works, helping developers build everything from robotics to autonomous vehicle systems and customer service avatars.
Shutterstock, known primarily for its stock photography business, is expanding into the business of licensing content for AI training, acting as a broker between AI developers and copyright owners ranging from media companies and film studios to podcast producers and digital creators.
“We’ve always enabled creative people to monetize their work in a rights-first environment,” Mitch Rotter, director of strategic content partnerships, data and AI, at Shutterstock, told The AI Innovator. “That is exactly what we’re doing with training material.”
Need more clues? Ask the Sherlock chatbot in the lower right corner to summarize this story, explain technical concepts or answer other questions.
Shutterstock is part of a growing class of companies seeking to supply one of AI’s most important inputs: licensed training data. Getty Images has expanded its own AI licensing efforts, while firms such as Scale AI, Appen and Lionbridge specialize in preparing and annotating datasets for model developers. Several startups also help creators monetize their work for AI training.
For Shutterstock, its clients want content that is “100%” created by humans, Rotter said. Creators still retain ownership; AI companies will get licenses to use the content for a limited time and only for training, not generation. Shutterstock brokers deals between content owners and AI companies.
Rotter didn’t disclose how much AI companies will pay for the content, only saying that “it’s a material part of what we’re doing.”
Real-world content desired
Early generations of large language models were trained largely on data scraped from the internet, a practice that triggered lawsuits from publishers, authors, artists and other copyright holders. As legal scrutiny has intensified, many AI companies have begun seeking licensed sources of content instead.
That has created a new market for what amounts to the raw materials of artificial intelligence.
Shutterstock began working with AI model developers roughly five years ago when companies sought large volumes of licensed images and videos for training purposes, according to Rotter. Over time, those requests became more sophisticated. Developers wanted additional types of content, better metadata and more specialized datasets.
Today, the company’s Data Partner Network includes not only images and short-form video but also television programming, films, podcasts, scientific datasets, architectural content, 3D models and conversational audio.
The market’s demands often differ from what media executives expect.
AI companies are frequently less interested in a brand’s prestige than in the characteristics of the content itself. A news archive may be valuable not because it covers politics or finance but because it contains examples of well-structured writing. A TV archive may be useful because it has thousands of hours of conversations, human movement or objects interacting in physical environments.
In many cases, the metadata describing content can be more valuable than the content itself. “The metadata is really the map,” Rotter said. “The raw material itself is not as valuable.”
Metadata can include transcripts, annotations, scene descriptions, camera specifications and object-recognition tags that help AI developers identify precisely the material they need.
A sports archive, for example, may help train computer vision systems because it contains continuous examples of moving objects, changing perspectives and human motion. Conversational audio can help developers build more natural customer-service agents. Content in underrepresented languages can help expand models beyond English-speaking markets.
A lifeline for journalism?
For media companies facing pressure on traditional business models, the prospect of monetizing archives arrives at an opportune moment. Digital advertising remains volatile. Subscription growth has slowed across much of the industry. Local news organizations continue to struggle financially.
Rotter argues that many organizations are overlooking assets they already own. “A lot of archival stuff is just gathering dust,” he said.
That includes not only traditional media companies but also YouTube creators, podcasters and production studios that have accumulated years of content. Much of that material generates little ongoing revenue even though it may contain characteristics valuable to AI developers.
Whether AI training ultimately becomes a major revenue source for publishers remains unclear.
Large media companies such as publishers and news organizations have already begun signing direct licensing agreements with AI firms. But Shutterstock is betting there is also demand for an intermediary that can package, prepare and distribute content at scale while handling rights management, metadata processing and commercial negotiations.
Rotter reassures that “we don’t do anything in the synthetic space. Our partners come to us because they want non-synthetic content.”
For years, publishers worried that artificial intelligence would diminish the value of their content. Instead, AI ironically may be creating a new market for some of their oldest assets.
“This is a very real opportunity,” Rotter said. “It’s happening now.”










