Skip to content

Product2 publishers3 min readPublished

Microsoft measured a 93 percent click-through drop to nytimes.com from its own answer engine

A newly unredacted brief in the New York Times' case quotes a Microsoft director calling AI scraping the largest theft of labor in human history. The underlying exhibits are still sealed, so the quotes arrive without their original context.

The Product Desk · Product desk

What happened

  • Newly unredacted filings in the Times' case quote a January 2023 memo by Microsoft's Director of Applied Science, Brent Hecht, calling AI scraping the largest theft of labor in human history.
  • Microsoft's own data shows its Copilot answer engine cut click-through rates for the New York Times' domain by as much as 93 percent compared with traditional Bing search.
  • The documents say OpenAI's mid-training datasets alone hold more than 91,692 copies of works published by the Times, the Daily News and the Center for Investigative Reporting.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Anyone shipping an answer surface over third-party content now has a vendor-measured number for what it does to the source's traffic. The referral-traffic paragraph in the deck has to be backed by your own outbound click logs.
  • exposure Turley's description of chatbots as largely substitutive sits in a public filing, where any plaintiff can reuse it against OpenAI and against other products built on the same models.
  • constraint If paywalled material needs a license to be used for grounding, retrieval pipelines that fetch logged-out or bypassed pages need legal sign-off before launch.
  • contradiction The admissions cut at the market-harm pillar of fair use, and the rulings so far have gone the AI companies' way. A product team reading only the filings will overprice the litigation risk and underprice the licensing bill.

A reader asks Copilot a question, reads the answer, and never opens the tab. Microsoft's own instrumentation counted the result: for every 100 clicks the traditional Bing results page sent to the Times' domain, the answer engine sent as few as 7 [11].

Brent Hecht, Microsoft's Director of Applied Science, wrote in a January 2024 internal presentation that the decline was a "doom loop" that would "hurt the performance of our models and the entire web at the same time" [4]. The same document, as quoted in the Times' brief, said: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain'" [5].

Every team shipping a summary or answer surface over someone else's content tells itself the same thing: the answer filters the user, so the clicks that survive are worth more. Microsoft counted the clicks. Nick Turley, OpenAI's Head of ChatGPT, wrote in internal communication that products like the chatbot are "largely substitutive" and "will get more and more substitutive as they get better" [6].

The volume behind that is now on the docket. One Common Crawl-derived dataset held more than 2 million documents from nytimes.com alone, roughly 22 times the number of copies logged across the three publishers in OpenAI's mid-training data [8][12]. The Times also alleges the companies bypassed paywalls undetected and deliberately stripped copyright notices out of training data [9].

TechCrunch reported that much of the new material comes from the Times' own brief rather than the exhibits, which remain sealed, and that the quotes appear without their original context [10]. Judges have so far been receptive to the argument that training is fair use, and earlier this month the Trump administration filed a brief defending OpenAI's unlicensed use of copyrighted work for training [13][14]. Near-term exposure for anyone building on these models is licensing cost and settlement leverage.

Does your surface reduce outbound clicks to the source? Measure that in clicks you can see in your own logs. Then ask whether your model vendor will state in writing that paywalled material in the training and grounding path was licensed. Substitutive plus unwarranted is the pairing that ends up in someone's discovery request. Most retrieval-augmented features shipped this year would fall into it.

Satya Nadella testified in a deposition earlier this year that "anything that is paywalled should be licensed by anyone who wants to use it...for grounding or training" [15], and said that had he been made aware OpenAI trained on paywalled information, he would have invoked Microsoft's right to require OpenAI to retrain its models [16].

What to watch

  • Whether the court unseals the underlying exhibits, which would show the quoted lines with their original context.
  • Whether any ruling weighs Microsoft's own click-through data as evidence of market harm under the fair use test.
  • Whether Microsoft acts on the retraining right Nadella described in his deposition.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories