Court documents unsealed on September 17, 2026, reveal that a Microsoft executive viewed AI's harvesting of content as potentially "the biggest theft of labor in human history." The remarks emerged in the lawsuit brought by The New York Times (NYT) and other news organizations against OpenAI and Microsoft, cited by plaintiffs from internal documents and testimony. Microsoft has stated that the remarks reflect the personal views of an individual employee, not the company's official position. What has newly come to light is an internal concern that the mechanism by which AI learns from articles and answers readers directly could simultaneously threaten the revenue of content creators and AI's own source of information.
Whose Words, About What "Theft"?
The remarks came from Brent Hecht, Microsoft's Director of Applied Science. The plaintiffs' publicly filed brief opens by quoting his description of the practice as possibly "the biggest theft of labor in human history." Further, on page 11 of the brief, an internal comment is cited suggesting that many people around the world would view large models absorbing people's work as theft on an unprecedented scale.
The "labor" referenced here isn't solely about jobs that AI might eliminate in the future. At issue is the relationship between the work people put into creating content—articles and the like—and those who use the results of that work for training. Hecht's concern connects to the recognition that most creators never intended such use and were not compensated for it.
However, this does not mean Microsoft has legally admitted to copyright infringement. In response to inquiries from Ars Technica, the company's public relations team characterized Hecht's statement as a personal opinion, not a legal analysis or corporate position. Microsoft maintains that its own AI products do not replace news sites and qualify as fair use, given their different purpose and character compared to the original copyrighted works.
Fair use is a legal framework in the U.S. that allows copyrighted material to be used without permission under certain conditions. What has been made public here is a brief in which plaintiffs argue against that defense—the court's ruling is a separate matter. Nor can the entirety of the underlying internal documents or testimony be read from this brief alone; the remarks are being encountered through context selected by the plaintiffs.
Learning From Articles vs. Substituting for Articles
OpenAI co-founder Greg Brockman also internally assessed the model's ability to complete news article text. Page 53 of the plaintiffs' brief cites a remark stating that when given a passage from an NYT article, the model appeared to complete the sentence accurately. This was an internal evaluation of a specific use case, not a guarantee that the current version of ChatGPT can always reproduce any article in full.
Even so, article reproduction has direct business implications for news organizations. If readers can get what they need from an AI's answer instead of opening the article, their incentive to return to the outlet that invested in reporting and writing weakens. The plaintiffs link the model's capacity to handle news content with its potential to displace the market for news organizations.
The plaintiffs' brief also states that Nick Turley, who leads ChatGPT at OpenAI, internally described the situation as an "existential threat" to publishers. The view was that as the AI product improves, its substitutability increases as well. The ability to handle articles skillfully is a strength for the AI product, but from the perspective of those selling articles, it also means the AI product is growing to capture the same demand.
It's also important to distinguish between reproduction of learned text and the mechanism by which external information is referenced when generating a response. The former is a matter of text emerging from what the model learned during training; the latter refers to a usage pattern where information obtained through search, for example, is fed to the model to construct an answer. The plaintiffs' brief discusses these separately, and the mere fact that a given article appeared in an output does not by itself determine through which pathway it was used.
This distinction also matters when measuring harm. The degree of textual overlap indicates the extent of reproduction, while whether readers move on to the original outlet is a matter of usage behavior. The plaintiffs use both types of evidence to build their argument that AI displaces the demand for reading the original article.
What Does the 87–93% Click-Through Drop Actually Measure?
Page 68 of the plaintiffs' brief cites Microsoft data showing that click-through rates to NYT's sites dropped by 87–93% when comparing Bing Chat to conventional Bing Web Search. What's being compared is the rate at which users move from an answer or search results to NYT's site. This is a different figure from claiming that NYT's overall traffic dropped by the same percentage.
Converting the 87–93% decline in NYT-directed click-through rate cited in the plaintiffs' brief into an index with conventional Bing search set at 100 yields a range of 7 to 13.
| Comparison | Relative Index of NYT Click-Through Rate |
|---|---|
| Bing Web Search | 100 (baseline) |
| Bing Chat | 7–13 |
Source: Plaintiffs' brief, page 68 (PDF page 76), SF1536, published September 17, 2026. The index was calculated as "100 × (1 − decline rate)." A 93% decline yields 100 × (1 − 0.93) = 7, while an 87% decline yields 100 × (1 − 0.87) = 13. The value of 100 does not represent an actual click-through rate of 100%—it is a baseline set to make the conventional search figure easier to compare against.
What this conversion shows is that, in the comparison cited by the plaintiffs, the rate at which users move from an AI answer to the original site is smaller than with conventional search. However, this particular passage does not reveal the absolute click-through rate, the measurement period, or the sample size. This is data presented by the plaintiffs, and it should be distinguished from a figure that a court has determined reflects the actual magnitude of the effect.
Accordingly, one cannot use this index to calculate a percentage decline in subscription revenue, nor can it be directly applied to all current Copilot users. Without data on the number of times search results were viewed or the rate at which visits convert into subscriptions, it's not possible to translate this into a dollar impact on the business. The difference in referral traffic shown by the numbers and any economic damage estimated from it need to be read as separate matters.
When News Revenue Suffers, AI's Own Information Supply Also Thins
Microsoft's internal documents recognized an uncomfortable dynamic for the company itself: that its end product threatens the economic foundation of a supplier it depends on. Page 2 of the plaintiffs' brief cites language suggesting that the AI content strategy has set off a vicious cycle that simultaneously degrades both model performance and the web as a whole. This should be read not as a measured outcome of future model quality decline, but as an internal warning about the sustainability of the business.
For news organizations to keep reporting new facts, they need revenue to fund that work. If AI's use of articles to generate convenient answers reduces readers' return to the original outlets, and that weakens revenue, then the supply of fresh information that AI itself depends on could also thin out. The internal documents introduced by the plaintiffs point to a possible link between losses suffered by content creators and long-term harm to AI companies themselves.
Still, the existence of this internal concern is not the same determination as whether individual uses constitute copyright infringement. Microsoft is contesting the very premise that its products substitute for news sites, and OpenAI had not immediately responded to Ars's inquiry at the time of publication. What remains to be seen is how far the plaintiffs' figures and statements will be credited against the companies' counterarguments.
What the court must ultimately determine is which specific uses of AI substitute for the demand to read original articles, and what data under what conditions supports that determination. This disclosure adds both forceful internal statements and concrete referral-traffic data to that ongoing dispute. The scope of compensation required for the use of news content will ultimately be defined not by the intensity of the language used, but by the court's judgment on the actual realities of that usage.
