Skip to content

Build2 publishers3 min readPublished

Cloudflare's Bot Preference Sync makes robots.txt an output of the edge, not a policy of its own

Bot Preference Sync writes Cloudflare's AI bot settings into your robots.txt from the Free tier up. The file becomes a rendering of dashboard state, and Cloudflare's bot database decides what it says.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Cloudflare's Bot Preference Sync makes robots.txt an output of the edge, not a policy of its own
Generated illustration

What happened

  • Cloudflare announced Bot Preference Sync, which mirrors a customer's AI bot configuration into robots.txt, is available from Free tier to Enterprise, and can be toggled off.
  • It builds on the July 1, 2026 controls that let site owners set separate rules for Search, Agent and Training traffic.
  • On sites with an existing file, the generated block is inserted at the top and the owner's existing Disallow lines are preserved.
  • Mixed-use crawlers that meet Cloudflare's transparency conditions can keep search access even when the customer has selected Disallow Training.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Site owners now choose where their AI policy is authored, and for these categories the dashboard wins: anything typed into the file by hand is downstream of a setting elsewhere.
  • constraint The sync only reconciles Cloudflare's category-level controls, so any rule outside those categories still has to be kept in step by hand.
  • exposure A site's published instructions can change because Cloudflare reclassified a bot, with Radar's directory as the place to find out why.
  • precedent Exemption from a customer's own block is now priced in disclosure, which turns one vendor's verification bar into a condition of crawler access.

Cloudflare's stated reason for building this is behavioural rather than tidy-desk. When a site's published preference and its enforced rule disagree, some crawlers treat the disagreement as grounds to disregard the preference or to try bypassing the block, the company writes in its announcement [15]. Accept that and robots.txt changes job. It stops being the policy and becomes corroboration for the policy, which lives in the rule that drops the request.

The two layers drifted in the first place because bot identity is not stable. One user agent can build a search index, fetch a page for an AI agent, and collect material for model training, and Cloudflare's July taxonomy treats those as three separate behaviours even when a single operator performs all three [9]. Cloudflare's July 1 argument was that these mixed-use crawlers disadvantage site owners precisely because they make the wanted use inseparable from the unwanted one [10]. A hand-maintained file cannot keep up with that. A generated one can, at the cost of depending on whoever maintains the generator.

Then there is what "Disallow Training" actually means at the edge. Selecting it makes Cloudflare publish a no-training preference while keeping separate enforcement against crawlers that fail its transparency requirements [4]. Those requirements number four, and only one of them describes crawler behaviour: respecting a no-training preference by any mechanism. The other three are disclosure duties owed to the site owner, namely an opt-out from AI summaries, URL-level visibility into which pages were made available for training along with search metrics, and public evidence that declining training does not cost the site its ordinary search results [6][19].

None of that reaches a crawler unwilling to say who it is. Cloudflare's verification framework depends on bots identifying themselves and honouring the relevant robots.txt preference, and syncing a file cannot compel an unidentified crawler [16]. The company's own reference point for that limit, per the runtimewire account, is the accusation it made against Perplexity last year [17]. So this is housekeeping for the cooperative half of the traffic, with the enforcement layer left to handle the rest.

One wrinkle in the reporting: runtimewire describes the feature as rolling out in an upcoming launch [12], while Cloudflare's own post announces it as available across tiers and switchable at any time [20]. Existing users of the older managed robots.txt feature will be asked to review their settings before moving across [14]. New customers get it on by default, which means the published file on a new Free tier site will be Cloudflare's output unless somebody deliberately turns it off [18].

That default is the part worth thinking about, because Cloudflare's own examples pull in opposite directions: an e-commerce store that wants everything crawled and trained on so its sofas surface in a chatbot answer, and an ad-funded publisher that wants to stay in the search index while keeping articles out of the training set [11]. The file will now say whichever of those the owner clicked. That is an improvement, provided the click was deliberate.

What to watch

  • Whether any mixed-use crawler operator actually publishes URL-level training visibility and evidence of no search penalty, and which bots land on Cloudflare's tracked list first.
  • What happens to a customer's published directives when BotBase reclassifies a bot and the generated section changes without the owner touching anything.
  • Whether default-on for new customers draws objections from crawler operators, publishers or regulators about a CDN authoring site policy.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories