Build2 distinct publishers3 min readUpdated
Bot Preference Sync writes Cloudflare's AI bot settings into your robots.txt from the Free tier up. The file becomes a rendering of dashboard state, and Cloudflare's bot database decides what it says.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Cloudflare's stated reason for building this is behavioural rather than tidy-desk. When a site's published preference and its enforced rule disagree, some crawlers treat the disagreement as grounds to disregard the preference or to try bypassing the block, the company writes in its announcement [9]. Accept that and robots.txt changes job. It stops being the policy and becomes corroboration for the policy, which lives in the rule that drops the request.
The two layers drifted in the first place because bot identity is not stable. One user agent can build a search index, fetch a page for an AI agent, and collect material for model training, and Cloudflare's July taxonomy treats those as three separate behaviours even when a single operator performs all three [15]. Cloudflare's July 1 argument was that these mixed-use crawlers disadvantage site owners precisely because they make the wanted use inseparable from the unwanted one [16]. A hand-maintained file cannot keep up with that. A generated one can, at the cost of depending on whoever maintains the generator.
Then there is what "Disallow Training" actually means at the edge. Selecting it makes Cloudflare publish a no-training preference while keeping separate enforcement against crawlers that fail its transparency requirements [7]. Those requirements number four, and only one of them describes crawler behaviour: respecting a no-training preference by any mechanism. The other three are disclosure duties owed to the site owner, namely an opt-out from AI summaries, URL-level visibility into which pages were made available for training along with search metrics, and public evidence that declining training does not cost the site its ordinary search results [10][18].
None of that reaches a crawler unwilling to say who it is. Cloudflare's verification framework depends on bots identifying themselves and honouring the relevant robots.txt preference, and syncing a file cannot compel an unidentified crawler [12]. The company's own reference point for that limit, per the runtimewire account, is the accusation it made against Perplexity last year [13]. So this is housekeeping for the cooperative half of the traffic, with the enforcement layer left to handle the rest.
One wrinkle in the reporting: runtimewire describes the feature as rolling out in an upcoming launch [19], while Cloudflare's own post announces it as available across tiers and switchable at any time [1]. Existing users of the older managed robots.txt feature will be asked to review their settings before moving across [3]. New customers get it on by default, which means the published file on a new Free tier site will be Cloudflare's output unless somebody deliberately turns it off [20].
That default is the part worth thinking about, because Cloudflare's own examples pull in opposite directions: an e-commerce store that wants everything crawled and trained on so its sofas surface in a chatbot answer, and an ad-funded publisher that wants to stay in the search index while keeping articles out of the training set [17]. The file will now say whichever of those the owner clicked. That is an improvement, provided the click was deliberate.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Cloudflare announced Bot Preference Sync, available to all customers from the Free tier to Enterprise, which updates a site's robots.txt to reflect what the customer has set in its AI bot configuration and can be turned on or off at any time.
Cloudflare says Bot Preference Sync will be switched on by default for new customers.
Existing users of Cloudflare's older managed robots.txt feature will be asked to review their settings before moving across.
If a customer has an existing robots.txt, Cloudflare says its generated section will be added at the top while preserving the site's existing Disallow directives.
On July 1, 2026, Cloudflare launched options to manage Search, Agent and Training traffic separately.
Search and Agent traffic can be allowed, blocked on pages serving ads, or blocked everywhere.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Mechanics well documented by the vendor, unverified in the field
The feature's behaviour is described in first-party detail (tiers, prepend-and-preserve, three Search/Agent options, Disallow-Training semantics, four verification requirements) and independently restated with additional operational facts by one secondary publisher. What is absent is any third-party test of crawler compliance, of edge enforcement effectiveness, or of BotBase classification accuracy, so the evidence base is descriptive rather than measured.
Broad availability, no uptake data
Adoption evidence is limited to distribution facts: a launch announcement, the July 1 category controls it builds on, availability from Free tier to Enterprise, and default-on for new customers with a review step for legacy managed-robots.txt users. There are no customer counts, zone counts, traffic figures, or disclosures of how many crawler operators meet the transparency conditions, so real uptake is unmeasured and the score reflects distribution potential only.
Mildly overstated by the vendor framing, tempered in coverage
Cloudflare's 'say it once' and 'transparency is the price of admission' framing implies more leverage than the mechanism has: the sync only publishes preferences and enforces against crawlers Cloudflare can identify, and category-level sync ignores custom rules and per-operator arrangements. The secondary coverage explicitly walks this back - a synced file is not an access-control system, and the Perplexity episode shows undeclared crawlers are unaffected - so the gap is modest rather than large.
Vendor-authored launch plus derivative coverage
The primary source is Cloudflare announcing its own product and, in the same post, defining the transparency criteria that determine which third-party crawlers keep access - a strong commercial and positioning interest, reinforced by Cloudflare owning the classification database (BotBase) and the public scoreboard (Radar). The secondary publisher names the Cloudflare blog as its primary source and frames the story as Cloudflare claiming the policy layer, so it adds interpretation but not independent verification.
Facts firm, consequences unmeasured
Two sources agree closely on the product mechanics and the secondary adds verifiable operational detail, so the factual core is solid. Confidence is held down by single-source items (default-on, migration review, PM attribution), a tense discrepancy between 'announced' and 'upcoming launch', the truncation of both source bodies, and the total absence of adoption or compliance measurement.
build
Cloudflare's one-click AI block names GPTBot, not the bot that decides if ChatGPT cites you1 distinct publisher
build
Route leak prevention moves into the protocol, and two Tier-1s are stripping the signal1 distinct publisher
build
The AI-training bans live on the big infrastructure blogs, not the small publications1 distinct publisher
build
OpenAI's cheap tier becomes a routing problem: Terra $2/$12, Luna $0.20/$1.20, seats untouched2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 21, 2026
1 article · August 21, 2026