Product1 publisher3 min readPublished
The DOJ's fair use brief tells you who can afford to lose the argument
Washington has filed on the side of fair use in the Times case against OpenAI, and the passage that matters to anyone shipping on scraped text is its estimate of who could still afford to train if licensing won.
The Product Desk · Product desk

What happened
- The Justice Department has filed in the New York Times' copyright case against OpenAI, arguing that training large language models on copyrighted work is fair use and citing precedent for it.
- Much of the filing's opening rests on national security, the claim that restricting what American companies may train on hands China the lead in AI.
- A later passage argues that an adverse ruling would leave only the largest technology companies able to pay licensing fees, with the money flowing to legacy publishers on the strength of archive volume.
- Techdirt, a long-standing fair use advocate, says the DOJ's analysis is correct and that the Times' copyright theories would expose ordinary reporting, including the Times' own, to liability.
- In separate litigation against Anthropic, Judge William Alsup found that AI training was easily fair use.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- constraint A filing is not a holding, so nobody rolling out a model on Monday gains a line they can hand a customer's counsel; the defensible claim is still about their own corpus.
- cost If licensing wins, the DOJ's own framing sets the price band: an expense the best-capitalised firms absorb and smaller trainers cannot, which makes incumbents the quiet beneficiaries of a publisher victory.
- exposure The blast radius reaches teams that never trained a model, because the same doctrine underwrites search indexing, book scanning, reverse engineering and research data mining.
- contradiction The brief's warmest endorsement arrives with a warning that this DOJ has spent its presumption of regularity, so its news value to a board outruns its persuasive value in front of the judge.
By mid-morning somebody on your team will have pasted the headline into a channel and asked whether the training-data row can come off the risk register. The clearest answer comes from the filing's supporters. Techdirt, which says the DOJ has the law right, also says this DOJ burnt through the presumption of regularity with a run of bad briefs in bad cases, so judges will stay skeptical even when it files something reasonable [2][9]. A brief is an argument, and the only fair use holding anywhere in this record belongs to one district judge in a different case [8][15].
Separate the thing being pitched from the thing being done. The geopolitical opening is the pitch, and Techdirt, which wants the DOJ to win, calls that argument weak and unnecessary [4]. The substance is about money. The filing describes licensing entry barriers as functioning primarily as large subsidies for old mainstream media companies, and says it is not in the public interest for the largest technology companies to hold an oligopoly on model training [6].
Read the competition passage as a cost estimate rather than a doctrine [5]. By its own logic, a plaintiff victory is a line item for firms with the deepest balance sheets and a shutdown notice for the open-weight teams underneath them [14]. The DOJ's own illustration of the value at stake is domestic and small: a writer with no budget generating an image for an article instead of commissioning a photographer or buying a licence [7].
Here is what teams tell themselves: the direction of travel is settled enough to ship on. Here is what the record supports: one district court holding [8] and one filed argument a court is free to discount [9]. Neither of those is a sentence you can paste into a vendor questionnaire and stand behind.
The forcing function is a 2x2. One axis is whether you can produce a corpus inventory, source by source, that a customer's counsel could read without follow-up questions. The other is whether you could absorb per-source licensing at your current revenue. Teams holding both the inventory and the money priced this some time ago, and the filing changes nothing for them. With the inventory but not the money, the risk is existential and legible, which is enough to start swapping sources for licensed, first-party or synthetic data while the case runs. Money without an inventory is a discovery problem, and it gets paid twice, once for the licences and once to reconstruct what was ingested. The fourth cell, no inventory and no capacity to license, is where the brief reads as relief, and it is also where an adverse ruling lands hardest, because the doctrine in dispute also carries search engines, book scanning, reverse engineering and text and data mining for research [11].
The one part of this a rollout owner can work on is the inventory, and it is also the only part that does not depend on how a judge rules [15].
What to watch
- Whether the court in the Times case engages the DOJ's competition reasoning or ignores it when it reaches fair use.
- Whether other courts follow Alsup's holding that AI training is easily fair use, or split from it.
- Whether licensing markets start quoting per-corpus prices only the largest buyers can clear, the outcome the DOJ brief predicts.