AI companies are hungry for data, and this is creating a new economy.
According to a Reuters report, veteran companies in the internet industry that have long archived massive amounts of data appear to be generating significant revenue by licensing their old data archives to AI technology companies for use as training data for AI models.
According to industry insiders, major tech companies at the forefront of AI development—such as OpenAI, Microsoft, Google, and Meta—are seeking large-scale licensing deals for images, videos, and other content to train their AI models. These AI companies are reportedly paying anywhere from 5 cents to $1 per photo, and $1 or more per video, depending on the buyer and the type of material.
The Reuters article also provides specific pricing details: Defined.ai, which licenses data to Microsoft, Google, Meta, and Apple, estimates prices at $1 to $2 per image, $2 to $4 for short videos, and $100 to $300 per hour for feature-length content. Nude images require special handling and cost $5 to $7 each. Text is priced at $0.001 per word.
Ted Leonard, CEO of stock photo service Photobucket, says the company is currently negotiating with multiple tech companies to license its 13 billion photos and videos at 5 cents to $1 per photo and $1 or more per video. However, even such a massive archive doesn't seem to satisfy AI companies' demand. According to Leonard, one company told him it needed more than 1 billion videos.
Similarly, stock photo service Shutterstock has signed usage agreements with Amazon, Google, Meta, and Apple covering hundreds of millions of images, videos, and music files. The initial contracts ranged from $25 million to $50 million, and have since been expanded further. According to Reuters, Shutterstock has already signed a contract with OpenAI as well.
Spanish platform Freepik has licensed the majority of its 200-million-image archive to two major tech companies at 2 to 4 cents per image. According to CEO Joaquin Cuenca Abela, five more similar deals are in the pipeline.
Photobucket's Leonard believes the company is on solid legal ground. He points out that the terms of service updated in October grant the company "unlimited rights" to sell uploaded content for AI system training purposes. He views licensing as an alternative to advertising, helping to sustain the company's ability to continue offering free accounts.
However, in February, the U.S. Federal Trade Commission (FTC) warned companies against retroactively changing their terms of service regarding AI usage.
For example, if a company adopts more permissive data practices—such as beginning to share consumers' data with third parties or use that data for AI training—and only notifies consumers of this change through a surreptitious, retroactive amendment to its terms of service or privacy policy, this could be unfair or deceptive.
The agency is investigating the training data agreement Reddit made with Google. Reddit's already human-evaluated, high-quality data is considered one of the most valuable data sources for AI companies.
Source
