Build1 publisher2 min readPublished
Unsealed NYT v. OpenAI brief quotes Microsoft and OpenAI insiders on scraping and paywalls
New York Times lawyers quote a 2023 Microsoft memo calling AI scraping 'the largest theft of labor in human history' in briefs a court unsealed September 17. For teams training on unlicensed data, the case puts dataset custody and internal email in the same record.
The Engineer · Build desk

What happened
- The Times sued OpenAI and Microsoft for copyright infringement in December 2023, with the Daily News and the Center for Investigative Reporting as co-plaintiffs.
- In January 2024 Brent Hecht, the memo's author, presented slides finding Copilot's answer engine cut click-throughs to nytimes.com by as much as 93% compared with Bing search.
- According to the brief, OpenAI president Greg Brockman answered a researcher's "hack to get around nytimes paywall" with "ah nice."
- Every quote comes from the Times' brief, the underlying exhibits remain sealed, and OpenAI and Microsoft did not comment.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure A company that takes delivery of a partner's training set opens itself to the argument that it knew what was in it, which is the case the brief makes against Microsoft.
- precedent Internal notes describing a model as substitutive go straight to a copyright plaintiff's competition claim, so later training-data plaintiffs can be expected to ask for the same kind of material.
- constraint Any judgement of what these memos prove has to wait for the sealed exhibits, because the public record is one side's selected excerpts.
- decision Teams training on unlicensed data now have to review dataset manifests and the threads discussing them together, since this brief pairs per-domain counts with the emails about how data was obtained.
Summary judgment is the stage where each side asks the judge to decide some questions without a trial, arguing that the facts are not really in dispute [12]. The surest way to show a fact is undisputed is to quote the other side saying it. So the Times' brief is built from OpenAI and Microsoft emails, slides and depositions [12].
For anyone running a training pipeline, the most useful passage is about custody. Satya Nadella said in a deposition this year that "anything that is paywalled should be licensed by anyone who wants to use it" [5]. Had he known OpenAI scraped paywalled content, he said, he would have invoked Microsoft's right to require OpenAI to retrain its models [6]. The same filing says "OpenAI delivered the entire GPT-3 training dataset to Microsoft" [7]. Data also went the other way, through projects called "Project Taxi" and "Project Mango", including material from the Bing index [7]. The brief uses this to argue that Microsoft held the dataset whose contents Nadella says he did not know [15].
Brent Hecht's memo predates the Times' December 2023 complaint by about 11 months [1]. At the 93% figure Hecht reported in January 2024, Copilot's click-through rate to nytimes.com was 7% of Bing's [2]. He called the pattern a "doom loop" that would "hurt the performance of our models and the entire web at the same time" [4]. "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers," he wrote [14].
Few plaintiffs get their theory of the case drafted by the defendant's head of product. Nick Turley, head of ChatGPT, wrote internally that ChatGPT is "largely substitutive" and "will get more and more substitutive as they get better" [10]. Substitution is the Times' core argument: that the companies copied the papers' journalism to train models that now compete with it [13].
In my view, each quote carries weight because it sits next to a dataset fact. Hecht's memo on its own is one employee's opinion about an industry [2]. Nadella's licensing standard becomes a problem when set against a GPT-3 dataset Microsoft held [5][7]. Brockman's "ah nice" becomes one when set against more than 91,692 copies of the plaintiffs' articles in OpenAI's mid-training datasets [8][9]. For a team training on data it did not license, I'd review the manifests before the email policy. A domain filter on the Common Crawl-derived set in this case would have taken its count of over 2 million nytimes.com documents to zero [9].
What to watch
- Unsealing of the exhibits behind the Times' quotes, which would show the Hecht memo and the Brockman exchange in their original context.
- The judge's ruling on which questions in the summary-judgment motions can be decided without a trial.
- Whether OpenAI and Microsoft contest the 91,692-copy and 2 million-document counts when they respond publicly.