Build1 distinct publisher3 min readPublished
The 20-page filing concedes an output can infringe on its own while asking that training be judged apart from it. If the judge agrees, the risk builders carry moves into the serving path, where it can be tested.
The Engineer · Build desk

product
Trump administration files in support of OpenAI's fair-use defence in the Times case1 distinct publisher
invest
OpenAI's 14 exits land on the two seats a $1T listing has to defend1 distinct publisher
invest
WikiHow's suit against OpenAI asks for a remedy that outlives the damages award1 distinct publisher
build
The AI-training bans live on the big infrastructure blogs, not the small publications1 distinct publisher
Compiled by The EngineerSomething wrong?How this is made
The April 2025 opinion in this litigation already split the pipeline into acquisition, training and output, and said each stage can raise its own copyright question [11]. Those stages are not abstractions to anyone who has shipped a model. Acquisition is crawlers, purchase contracts, and whatever a vendor told you about a corpus. Training is a job that runs once and leaves weights behind. Output is the serving path, where requests arrive and tokens leave.
The government's brief describes training mechanically: copy the works, convert the data into numerical representations, learn linguistic patterns from those representations [5]. That description is doing the legal work. It puts the resulting artifact on the statistics side of the line, which is how the filing gets to "highly transformative" and to tasks the original articles never served [3].
For that framing to survive contact with a plaintiff, one thing has to be true: the weights must not hand the article back. The brief does not pretend otherwise. It concedes that an output can be its own copyright problem when a model reconstructs and distributes protected material, and argues only that such outputs should not decide the training question [6]. The Times has alleged that reconstruction since it sued in December 2023 [8].
This is the allocation a builder should want regardless of who wins. You cannot un-train a model, and a dataset audit run after the fact produces a document, not a control. A reproduction test is a control. Write an eval that prompts for known passages, measure what comes back, keep the traces, gate releases on the number. If publishers have to prove infringement at the output level [7], the evidence sits in a system you already own and can instrument.
Then count the stages: the court named three [11], and the filing argues about only two of them [17]. This brief does not address acquisition at all, and the transformativeness of training says nothing about how the bytes arrived. The government itself says fair use stays fact-dependent and that it is not asking the court to excuse every use [15].
The licensing section is the loudest part. Pay per work, the brief argues, and only the largest technology companies can afford to train, while publishers with the deepest archives collect the largest payments [12]; it calls that a route to oligopoly in model development and a subsidy for established media, and warns that U.S. developers would trail foreign competitors under different IP rules [13][14]. An anti-oligopoly argument filed in support of OpenAI [9] is a shape worth noticing.
None of it binds the judge, who is free to ignore the statement of interest [10]. Two things have to be true before your exposure actually shifts: the court has to adopt the separation, and your serving path has to be able to tell you when it reproduces someone else's text. Building that detection capability is within your control, and it is the cheaper of the two.
Ranked by verification strength, evidence, and original report placement.
The Justice Department filed a 20-page statement of interest on September 1st in the consolidated OpenAI copyright litigation in the Southern District of New York.
TechCrunch reported the filing on Wednesday; runtimewire.com lists TechCrunch as its primary source.
The government argues that copying written works to train an LLM is highly transformative, because the resulting model learns statistical relationships among words and can perform tasks different from those served by the original articles.
The Justice Department urged the court to analyze model training separately from any output that reproduces protected expression.
The filing says training an LLM requires copying works, converting the data into numerical representations, and using those representations to learn linguistic patterns.
The filing concedes that an output may present a separate copyright problem when a model reconstructs and distributes protected material, while arguing that potentially infringing outputs should not determine whether the earlier training process qualifies as fair use.
Distinct publishers with included, body-backed reporting in this cluster.
runtimewire.com
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One retelling, no filing in hand
The specifics are precise — 20 pages, September 1, the Southern District of New York — and precise details are checkable, which keeps this from scoring lower. But every one of them arrives through runtimewire.com's summary of TechCrunch's summary. No passage of the brief is quoted, no docket entry is cited, and the two documents the story leans on for context, the April 2025 opinion and the Copyright Office report, are described rather than shown.
Nothing deployed to observe
A statement of interest is not something anyone installs. This reporting contains no release, no change in data sourcing, no licensing deal signed or abandoned, and no lab adjusting its output filters in response — so there is no uptake here to score, only a legal argument awaiting a judge.
Headline outruns the docket
'Backs OpenAI' does more work than the underlying event supports: a non-binding filing in a case the judge has already sliced into stages, in which the government explicitly declines to bless every use and concedes outputs may infringe. Pulling the other way, the story flags its own limits twice and the framing above it is properly conditional. Hence a modest overstatement, not a serious one.
Advocacy relayed as description
Almost every substantive proposition here is an argument authored by a party seeking an outcome. The oligopoly warning and the foreign-competitor point are industrial policy wearing fair-use clothing; OpenAI collects the benefit if the court agrees, The Times absorbs the loss, and publishers with large archives are cast as rent-seekers by the same brief that would relieve their counterparty. The one counterweight on the page is the Copyright Office's narrower reading, which the reporting does surface.
Firm event, soft particulars
Two things hold up well: a filing exists, and it asks the court to treat training and output as separate questions. Beyond that the ground softens — how far the brief actually pushes past the Copyright Office, and whether it addresses acquisition at all, we know only through paraphrase. Enough to act on directionally; not enough to quote in a memo.