Skip to content

Topic

Training Data Provenance

Determining the origin and authorship of web text used to build model training and retrieval corpora, and the contamination risk from synthetic pages.

Current stories

build3 publishers

Naming Amodei and Mann turns Anthropic's data-acquisition approval chain into discovery

Sony and Warner want evidence about which founder approved which corpus, and that request is aimed at internal records rather than model weights. The complaint is effectively a test case for any lab whose pipeline cannot answer that question per corpus.

Perspective Coverage

3 publishers
Builder
Builder 37%
Operator
Operator 32%
Investor
Investor 31%

Reality

Evidence62
Adoption
Insufficient
Hype gap+18
Incentives72
Confidence58
build4 publishers

A 30-day retention clause routes Nvidia's sensitive work to its own Nemotron models

Anthropic's June decision to keep Fable usage logs for 30 days has turned zero data retention into a procurement gate at Nvidia, Booz Allen and Palantir. The metadata channel stays open under the clause they are demanding.

Perspective Coverage

4 publishers
Builder
Builder 26%
Operator
Operator 35%
Investor
Investor 39%

Reality

Evidence58
Adoption62
Hype gap+18
Incentives74
Confidence55
build1 publisher

The EU AI Act pushes training-data provenance down into the crawler

The AI Act's general-purpose AI obligations became enforceable in August, and the parts of them that touch a data pipeline depend on records only the collector can write while it is fetching. Penalties scale with turnover.

Publishers:dev.to

Reality

Evidence32
Adoption
Insufficient
Hype gap+22
Incentives70
Confidence44
product6 publishers

Sony and Warner name Anthropic's co-founders personally in their lyric suit

The publishers want $150,000 for every alleged infringement and say Claude will recite lyrics verbatim, which moves output filtering into the column of legal controls rather than polish for anyone shipping a text feature.

Perspective Coverage

6 publishers
Builder
Builder 32%
Operator
Operator 31%
Investor
Investor 37%

Reality

Evidence62
Adoption
Insufficient
Hype gap+30
Incentives78
Confidence58
leadership3 publishers

OpenAI asks Judge Stein to measure copying by output rather than by corpus

Three summary judgment briefs in the consolidated news publisher case describe the same conduct using denominators that never meet, and the one the court adopts will set what licensed text is worth to every buyer.

Publishers:abcnews.comppc.landtheguardian.com

Perspective Coverage

3 publishers
Builder
Builder 28%
Operator
Operator 34%
Investor
Investor 38%

Reality

Evidence58
Adoption25
Hype gap+12
Incentives84
Confidence62

Earlier coverage

  1. Sony and Warner attach a $150,000 per-work price to Anthropic's training corpus

    Product · August 30, 2026 · 1 publisher

  2. WikiHow's suit against OpenAI asks for a remedy that outlives the damages award

    Invest · August 24, 2026 · 1 publisher

  3. Alsup fined the download, not the training: the copyright risk in your stack is provenance

    Product · August 23, 2026 · 1 publisher