Build1 publisher3 min readPublished
OpenAI's own executives call ChatGPT substitutive of the journalism it trained on
Discovery quotes unsealed on Thursday have Nick Turley, Greg Brockman and Satya Nadella describing the products as substitutes, while OpenAI and Microsoft have told a Manhattan court that their training does not replace journalism.
The Engineer · Build desk
What happened
- Filings unsealed on Thursday in the Manhattan copyright case show OpenAI and Microsoft executives privately describing their AI products as substitutes for journalism, according to the news outlets suing them.
- Nadella agreed under oath that conversing with chatbots has substituted for needing to go to the underlying source of the information.
- The Trump administration backed the AI companies on September 1, telling Judge Sidney Stein that training is "extraordinarily" transformative, and the defendants filed their own fair-use arguments on September 4.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint The defense was pleaded on two grounds, transformation and non-substitution, and the unsealed testimony puts the defendants' own witnesses against the second one, leaving transformation to carry more weight before Stein.
- contradiction Microsoft treats Nadella's sentence as commentary on how people find information; the newspapers treat it as an admission of market harm, and how Stein reads it decides what the sentence costs.
- exposure The paywall-hack exchange puts conduct during data collection into the record. That attack does not depend on what the model outputs.
- precedent With dozens of copyright suits pending on the same transformation question, a substitution finding here becomes the citation the next set of plaintiffs reaches for.
The defense as pleaded has two parts. OpenAI and Microsoft told the Manhattan federal court that training on millions of newspaper articles transforms the copyrighted material into new content, and that the result does not compete with or replace the news organizations' journalism [2]. The unsealed quotes go at the second part. The Trump administration's brief goes to the first: it told US District Judge Sidney Stein on September 1 that AI training is "extraordinarily" transformative [14].
Turley was describing what the product does to publishers. The outlets quote OpenAI's head of ChatGPT saying they face an "existential threat" from products that are "largely substitutive" and "will get more and more substitutive as they get better" [4]. Brockman was describing capability at the specific task. The filing quotes him writing that large language models are "particularly good at predicting text of news articles," "excellent at news," and "very good at any news task" [8].
Microsoft's answer is about characterisation. "Satya's testimony and Microsoft's position in this case are perfectly consistent," a spokesperson said, adding that "He spoke to broad principles and changes underway in how people find and consume information" [6]. The same filing quotes Microsoft's director of applied science, Brent Hecht, saying "millions of people" would consider AI companies "'hoovering up' all their work" to be "an astonishing theft of unprecedented proportions" [10]. Microsoft said those comments "reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views" [11].
One quote sits on a different axis from output substitution. Brockman responded "ah nice" when he was told that OpenAI employees had discovered a "hack" to get around the New York Times' paywall, according to the outlets' filing [9]. That goes to how the material was obtained.
Steven Lieberman represents eight newspapers in the case, the New York Daily News and seven others [12][18]. "The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong," he said in a statement [12]. The suit was first filed in 2023, and the plaintiff publications include Ziff Davis titles, the Intercept and the Center for Investigative Reporting [13]. OpenAI did not immediately respond to requests for comment on the filing, and a New York Times spokesperson declined to comment [17].
For anyone building on frontier models, the question is what would have to be true for this record to reach a supply contract. The report covers two defendants and one defense; contract terms are not disclosed [19]. Dozens of other copyright complaints are pending against tech companies, and according to the report those cases will likely turn on the same transformation question [16]. The event that would move provenance from a legal argument into a procurement clause is a holding that substitution defeats fair use for news training data. Stein has the parties' arguments in filings dated September 4 [15].
What to watch
- Whether Stein's ruling treats transformation and market substitution as separable questions, or lets one carry the other.
- Whether plaintiffs in the dozens of other pending complaints move to unseal comparable discovery from their defendants.
- Whether OpenAI issues a substantive response to the filing after not immediately responding to requests for comment.