Skip to content

Build2 publishers2 min readPublished

Cloudflare's new AI training opt-out is backed by network-level enforcement Apple, Google and Microsoft have agreed to honor

Fewer than 1% of Cloudflare sites block search crawlers while 17% block training. A single mixed-use crawler used to collapse those two into one decision, and the new setting splits it for three named operators.

The Engineer · Build desk

Illustration accompanying Cloudflare's new AI training opt-out is backed by network-level enforcement Apple, Google and Microsoft have agreed to honor

What happened

  • Cloudflare has launched a Disallow AI Training setting that keeps a site indexed for search while refusing the same crawler permission to train a model on its content.
  • The tradeoff it targets comes from mixed-use crawlers, a single crawler doing both search indexing and training, where refusing one use has meant refusing the other.
  • Apple, Google and Microsoft either honour the setting today or have committed to honour it within a specified time frame.
  • Cloudflare classifies bots by behaviour, and three behaviours are separately controllable: search, training, and user-directed agents such as chat fetch bots and browser-use agents.
  • Cloudflare says the Block setting now means something different, because Block and "Block on pages with ads" previously did not apply to mixed-use crawlers.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Every site that used to infer its training answer from its search answer now has a separate field to fill in, and leaving it untouched is a position too.
  • constraint Honouring the preference is the operator's job and the evidence arrives after the crawl, so an owner who needs an enforced guarantee is still stuck with the blunt tool that costs indexing.
  • exposure Settings chosen when mixed-use crawlers sat outside Block may now reach crawlers their owners never intended to touch, and the change lands without anyone re-picking a value.
  • precedent Cloudflare is making itself the one place a summaries preference gets registered. Cloudflare then negotiates with the operators in place of the individual publisher.

Turn the setting on and Cloudflare publishes a `Disallow:` directive in your robots.txt, which is where the name comes from [8]. Cloudflare's post also argues that a robots.txt directive alone cannot solve the problem: anyone can publish one, but it cannot identify who is crawling, determine why they are crawling, or stop a crawler that ignores it [9]. The rest of the design is network-side. Publish the preference, identify the crawler, classify the purpose, block the ones that ignore it, then report what each operator actually does on Radar [10].

This setting exists to avoid blocking [12], and Cloudflare's stated reason is that blocking removes a crawler without changing how crawlers behave [11]. So for the three named operators, compliance rests on a commitment plus an audit trail: the Accountable designation requires URL-level visibility into which pages were made available for training, metrics showing how content appeared in search, an operator-level opt-out for AI summaries, and assurance that opting out of training will not affect traditional search results [13]. Cloudflare says it has been talking to operators directly since July, and that almost all of them agreed site owners should have control and transparency into how their content is used [18]. An operator's agreement cannot be verified from outside, but the per-URL training reports can, and I would want to see one before loosening a config that currently blocks.

The demand side offers one comparison: fewer than 1% of Cloudflare sites block search bots, and 17% enable some mechanism to block training [4][5]. Training refusers therefore outnumber search refusers by more than seventeen to one [6]. Both figures are shares of Cloudflare's own sites, counted as sites. Cloudflare did not publish a traffic-weighted version. For the 17% to say something about the advertising- and subscription-funded web the post points to [19], it would have to hold when the count is weighted by traffic or by pages of content. The other 83% of sites have no training block at all [7].

Summaries are the harder half. According to Cloudflare, a site-wide yes or no is too blunt, because how much of your content appears in a summary matters as much as whether it appears at all [16]. The goal it states for early next year is a single control on Cloudflare for how much of a site's content is included, set once instead of with each operator separately [17].

What to watch

  • Whether the single Cloudflare control for how much content appears in AI summaries ships by early next year, and what proportion it defaults to.
  • What Radar shows about each Accountable operator once URL-level reports on pages made available for training start arriving.
  • Whether a fourth operator earns the Accountable designation, or one of the three loses it for missing a time-bound commitment.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories