Discussion over the provenance of training data for large language models is still ongoing, but AI companies continue to form partnerships with news organizations. The latest such move sees it announced that ChatGPT developer OpenAI has entered into a strategic partnership with the UK's economic newspaper "Financial Times" (FT).

Through this partnership, FT will provide article data to OpenAI, enabling OpenAI to use FT's data for training large language models. FT will also use OpenAI's technology to develop new AI tools for its readers.

In addition, FT staff will begin using ChatGPT Enterprise. Even so, FT states that it remains "committed to human journalism." John Ridding, CEO of FT, said, "It's right, of course, that AI platforms pay publishers for the use of their material," noting that the agreement "demonstrates that OpenAI values FT journalism and wants to understand how AI systems make use of content." Neither company has disclosed the financial terms of the agreement.

"Aside from the benefits for the FT, there are broader implications for the industry. Of course, it's right that AI platforms pay publishers for the use of content. OpenAI understands the importance of transparency, attribution, and compensation. At the same time, it's clearly in users' interest for these products to include trusted sources. What's simply not possible is to turn back the clock," Ridding added.

FT stated that OpenAI understands the importance of maintaining transparency, giving credit, and compensating for content.

Large language models, exemplified by GPT-4, have their performance heavily influenced by the quality of the data used to train them. Up until now, AI companies have scraped as much as they can from the public internet without creators' consent, and they are constantly searching for new data sources to keep the outputs generated by these models up to date. Training AI models on news is one such method, but some publishers are wary of providing their content to AI companies for free. For example, the New York Times and BBC have banned OpenAI from scraping their websites.

As a result, OpenAI has been entering into financial agreements with major publishers to continue training its models. Last year, the company partnered with German publisher Axel Springer, training its models on the latest articles from the US's "Politico" and "Business Insider," as well as Germany's "Bild" and "Die Welt." The company has also signed agreements with the Associated Press, France's Le Monde, and Spain's Prisa Media. However, it's said that the content licenses OpenAI offers to publications range from $1 million to $5 million, considerably less than what other companies such as Apple are offering.


Source