Skip to content

Product1 publisher3 min readPublished

Unsealed brief quotes Microsoft's applied science director calling AI training a 'theft of labor'

Redactions came off a 92-page plaintiffs' brief in the consolidated OpenAI copyright case on Thursday. Inside are internal Microsoft and OpenAI documents on substitution, memorised training data and collapsing publisher click-throughs.

The Product Desk · Product desk

Photograph accompanying Unsealed brief quotes Microsoft's applied science director calling AI training a 'theft of labor'
Photo: thenextweb.com

What happened

  • Redactions came off the News Plaintiffs' 92-page combined summary judgment brief in the consolidated OpenAI copyright litigation before Judge Sidney H. Stein on Thursday, after covering most of its damaging passages.
  • It quotes Brent Hecht, Microsoft's Director of Applied Science, describing large models taking people's work as "an astonishing theft of unprecedented proportions" and, in a second document, perhaps the "largest theft of labor in human history".
  • Microsoft's own data in the brief shows click-through rates falling 87% to 93% for the Times's websites, 83% to 91% for the Daily News group's, and 51% to 94% for Ziff Davis's, measuring Copilot against traditional Bing search.
  • The brief says OpenAI obtained the New York Times Annotated Corpus, over 1.8 million articles from 1987 to 2007, under a licence limited to non-commercial research.
  • The plaintiffs are The New York Times, the Daily News group, the Center for Investigative Reporting and Ziff Davis.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • exposure The damaging passages were written by employees of the two defendants, so a buyer assessing a model's training provenance can now quote the vendor's own staff instead of a plaintiff's characterisation of them.
  • constraint A fair use defence is harder to run in front of a jury once the defendant's own applied science director has written that a win for his employer would mock the doctrine.
  • decision Publishers weighing a licensing deal have an internal Microsoft baseline for the referral traffic they are being asked to trade away, which changes what a reasonable licence fee looks like.
  • contradiction The line being passed around came from a research scientist, not from a board member or the chief executive, so reading it as Microsoft's corporate position stretches it past what the filing supports.

The one sentence in the filing that describes what users do belongs to an OpenAI software engineer the brief does not name: "No matter how prominently we show the links, users won't click." [12] Nick Turley, who runs ChatGPT, wrote that publishers face an "existential threat", and said the products are "largely substitutive, period" [13]. Satya Nadella agreed under oath that chatbots had substituted for going to the underlying source [14]. Internal OpenAI documents call ChatGPT "the modern newsstand" [15].

One Microsoft document in the brief says the company's AI content strategy "has started a 'doom loop'" that will damage both its own models and the entire web [8]. Another puts it in eight words: "LLMs are a product that destroys its supply chain." [11] Brent Hecht also wrote that a ruling for his own employer would arguably "make a complete mockery of the idea of 'fair use'" [10].

Two numbers in the brief count the same sort of thing on different terms. The plaintiffs say OpenAI's mid-training datasets hold more than 91,692 copies of their works [18]. Separately, they say WebText2 held at least 6,552 items from the Times, 18,609 from the Daily News group and 66,780 from Ziff Davis [19]. Those three add to 91,941 [1], which is already above the floor given for every mid-training copy [3]. Ziff Davis is about 73% of that WebText2 total [2], and Ziff Davis publishes CNET, ZDNET, PCMag and Mashable [7].

On what the training data held, OpenAI's VP of Research gave the plaintiffs a usable sentence: "We train our networks to memorize the training data. That's their objective." [16] By November 2019 the company was worrying internally about regenerating copyrighted works [17]. Around 2017, Greg Brockman wrote that he was "deeply motivated by the gazillions" he hoped to make from commercialising the technology [21]. Nick Ryder later told Brockman about "a hack to get around nytimes paywall", and Brockman replied "ah nice." [22]

For a team picking a model this month, the useful distinction is between what the vendor sells and what its own staff wrote down. The sales material is about citation and referral traffic. The internal record is about substitution, including a Microsoft warning of a "real risk" that the technology could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained" [23]. Two procurement questions fall straight out of the filing: which of these plaintiffs' titles are in the training data, and under what licence. The document is the plaintiffs' own combined summary judgment brief before Judge Sidney H. Stein, so its selection of internal material is adversarial by design [5].

What to watch

  • How Judge Stein treats these internal documents when he rules on summary judgment.
  • Whether Microsoft or OpenAI disputes the context of the quoted documents or moves to re-seal them.
  • Whether other publishers put their own Copilot and ChatGPT referral figures on the record.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories