Security2 publishers2 min readPublished
Unsealed NYT v. OpenAI filing puts Microsoft's "astonishing theft" memo into the public docket
Court documents unsealed Thursday quote Microsoft and OpenAI employees describing the training corpus as theft and the products as substitutive. Microsoft says those words are not the company's position.
The Watch · Security desk

What happened
- A filing in the newspapers' copyright case against Microsoft and OpenAI was unredacted Thursday in Manhattan Federal Court, exposing material the two companies had insisted be kept confidential.
- Microsoft's Director of Applied Science, Brent Hecht, wrote that millions of people will soon consider large models "hoovering up" all their work an astonishing theft of unprecedented proportions.
- An internal Microsoft document says the company's AI content strategy started a "doom loop" in which an end product threatens the economic foundations of its essential suppliers.
- A Microsoft spokeswoman said Hecht's comments are one employee's individual perspective, are not a legal analysis, and that Copilot is not a substitute for publishers' journalism.
Compiled by The WatchSomething wrong?How this is made
Why it matters
- exposure The Authors Guild and the individual writers in the same consolidated litigation inherit these quotations without running their own unsealing fight.
- decision A buyer weighing provenance risk on Copilot or ChatGPT can put the vendor's internal wording next to the vendor's filed position and make the two reconcile before signing.
- precedent Redaction designations in AI copyright cases are now worth contesting, because contesting them here produced the defendants' own sentences.
- constraint Any transformative-use argument now has to be made in front of a judge reading OpenAI's own policy director on labor substitution.
A publisher arguing that its archive sits inside a training corpus normally has to infer the contents from model output. The newspapers' motion, filed earlier this month, quotes the defendants' own employees and asks the court to rule on acquisition-, training- and distribution-based infringement [3][4].
Both sides now argue substitution. Jack Clark, OpenAI's policy director, wrote that the company was "creating systems that substitute for the labor of the people that define the 'culture' of society" [11]. An OpenAI software engineer testified that "no matter how prominently we show the links, users won't click" [12]. According to 404 Media's account of the filing, Microsoft executives including Satya Nadella testified under oath that clicks to the Times and other news sites fell by more than 90 percent on Bing [9][10].
A Microsoft policy document quoted in the motion extends the point to employment, saying generative AI could significantly disrupt the employment of the very people who generated the data the foundation model was trained on, and that large language models are a product that destroys its own supply chain [13]. 404 Media also reports documents showing OpenAI created "a hack to get around nytimes paywall," and that cofounder Greg Brockman replied "ah, nice" [14].
Microsoft's answers are in the record too. On the doom loop document, a company spokeswoman said it concerned "broad principles and changes underway in how people find and consume information," not the copyright questions before the court [16]. In a declaration filed with Microsoft's own summary judgment motion on Sept. 4, and refiled with fewer redactions Thursday, another Microsoft executive framed Brent Hecht as someone on the payroll to play a contrarian [17][18].
A separate unredacted passage alleges the companies copied "billions of web pages" for "horse trading"-like deals, selling content to each other instead of licensing it from the news organizations. Microsoft did not address a question about that allegation [21].
Steven Lieberman, the attorney leading the Daily News' legal effort, said the disclosures prove the companies are knowingly committing theft instead of operating inside fair use as they claim [19]. "Well, now the cat is out of the bag," Lieberman said. "Finally, the world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior" [20].
No judge has ruled on either motion. Neither account reports an enterprise customer being sued over training-data provenance. For a company negotiating indemnity language on Copilot or ChatGPT, the difficult sentences are now the vendors' own, in a public docket, where Jason Kint, the CEO of the trade group Digital Content Next, found them [23].
What to watch
- Whether the judge rules on the newspapers' motion, Microsoft's Sept. 4 motion, or sends the substitution question to trial.
- Whether the Authors Guild and author plaintiffs cite the same unredacted exhibits in their own filings.
- Whether OpenAI answers publicly on the paywall-hack documents and the "billions of web pages" trading allegation.