Skip to content

Build1 publisher2 min readPublished

Account tier decides whether customer content becomes training data

An anonymous page accused platforms of selling user content to AI labs without naming one, while the published policies at OpenAI and Anthropic already split training defaults by the account a customer signed.

The Engineer · Build desk

Illustration accompanying Account tier decides whether customer content becomes training data

What happened

  • An anonymous site called fuck-off.ai circulated on September 19th accusing software companies of widening their terms so user content becomes training material, with leaving as the customer's remedy.
  • Its central scenario is a frontier lab writing a check big enough to end a platform's internal debate over licensing user data. The page does not say which lab, which platform, how much, under what contract, or when.
  • The Federal Trade Commission warned in February 2024 that adopting broader data practices retroactively through quiet terms changes, AI training among them, could be unfair or deceptive.
  • Anthropic's July 8th, 2026 privacy update lets consumers control whether their conversations improve its models and excludes Team, Enterprise and developer-platform accounts from those consumer terms.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure Because the FTC has ordered companies to delete models built on unlawfully obtained data, a terms rewrite that reaches backwards puts trained weights in scope, not just the disclosure text.
  • capability A team that needs clean provenance can buy it, with permission and intended use fixed before transfer and a 5% buyer-side fee as the comparison price against training on its own users.
  • contradiction The page argues as though one industry default applies everywhere, while the policies differ by product and customer class, so its claim cannot be tested against any particular platform.

Read the vendor policies as configuration and the question gets specific. OpenAI's data-use policy, updated March 13th, 2026, says content from individual services may be used for training unless the user opts out, and business services sit under different data commitments [7]. The default follows the contract the customer signed. According to runtimewire, a privacy promise can work as a product feature for enterprise buyers while consumer data stays an input elsewhere in the same company, and customers still experience one brand [10].

The regulatory limit is narrower than the accusation, and harder to engineer around. A separate FTC warning told AI companies to honor commitments made in privacy policies, marketing materials and product marketplaces [5]. The agency has previously required companies to delete models and algorithms built with unlawfully obtained data [6]. That remedy applies to the trained artifact.

Founders at Mozilla Data Collective, Kled, Verb and Protege are building marketplaces on the opposite premise: if the data is worth selling to AI developers, the people producing it keep control and get paid [11]. Mozilla's version prices the permission first. Verified providers set their own licensing prices and receive the license fee, and Mozilla charges buyers a separate 5% platform fee, with permission, price and intended use stated before a dataset changes hands [12]. On a $200,000 license the buyer pays $210,000 and the provider keeps $200,000 [17]. Mozilla says E.M. Lewis-Jong, who led the crowdsourced speech corpus Common Voice across hundreds of languages, started the collective in 2025 after looking for a platform that gave data-producing communities control over licenses and value exchange [13].

Kled is running the consumer version. It announced a $5.5M seed on March 10th, taking stated total financing to $10M, with Wischoff VC, Aglae, K5 Global, Parable VC and Cox Exponential in the round [14]. That leaves $4.5M raised before this one [16]. Contributors upload material for licensed datasets and can be paid when that data is purchased, and Kled's upload, earnings and customer figures are company-reported [15].

The anonymous essay rests on weaker evidence than the marketplaces do. fuck-off.ai says comparable content deals reach tens or hundreds of millions of dollars, and it supplies no example to price against [3]. Its description of platform behaviour is also too broad to establish that every product runs the same defaults and exceptions, which vary by product and customer class [9]. The usable part of the record is published: the two vendors' training defaults by account class [7][8], the FTC's February 2024 position on retroactive terms changes [4], and a live marketplace that attaches a price to permission [12].

What to watch

  • Whether any platform discloses an actual licensing deal with a named lab, a payment and a date.
  • Whether the FTC brings an action over a terms change that moved already-collected user content into training.
  • Whether Mozilla Data Collective or Kled publish audited volumes and payouts instead of company-reported figures.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories