Build1 distinct publisher3 min readUpdated
The managed robots.txt file lists eight crawlers. Three of them are the training half of a pair whose search half is left allowed, and Perplexity is not in the file at all.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Cloudflare's managed robots.txt feature writes a block for eight named user agents, and the list runs along the seam between training crawlers and search crawlers [1][2]. The practical result is that teams who flip the switch to stop AI answer engines from using their content are instead removing themselves from training corpora while leaving the citation-facing bots untouched [9].
The file is published in full on Cloudflare's own docs page: Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent, each with Disallow: / [1]. Above them sits a wildcard block that allows everything and carries the line "Content-signal: search=yes, ai-train=no, use=reference" [3]. That is not ambiguous drafting. Search is explicitly permitted; training is explicitly refused.
Read against Cloudflare's crawler reference table, the pattern repeats three times. GPTBot is blocked and OAI-SearchBot is not; ClaudeBot is blocked and Claude-SearchBot is not; Applebot-Extended is blocked and plain Applebot is not [4]. For the first two pairs the table labels the blocked agent an AI Crawler and the unblocked one AI Search [5]. The table has no Applebot-Extended row at all, so Cloudflare's own reference cannot tell you which side of the line it sits on [6]. Perplexity is not named anywhere in the file [7].
The confusion has a documented cost. An r/SEO post from April, with 53 points and 40 comments, claimed Cloudflare had cut the author's site off from ChatGPT, from Perplexity and from Google's AI Overviews [8]. None of those three is what the eight agents govern [9]. What they do govern is real: GPTBot controls inclusion in OpenAI's training data, and Google-Extended controls grounding in Gemini Apps [10]. OpenAI's documentation is one line on the relationship between its bots: "Each setting is independent of the others" [11].
Google-Extended is the entry that does the most damage, because the name reads as if it governs AI answers in Search. Google's crawler documentation says otherwise: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search" [12]. It also notes that Google-Extended "doesn't have a separate HTTP request user agent string" and that crawling is done with existing Google user agent strings, which disposes of the common advice to grep your access logs for it [13].
The control that does exist for AI Overviews is the snippet. Google's AI features page states that robots.txt directives for Googlebot are the control for how sites are crawled for Search, and points owners to nosnippet, data-nosnippet, max-snippet or noindex to limit what is shown from their pages [14]. Opting out of the AI answer means opting out of the snippet.
On the recurring claim that Cloudflare turned this on for existing domains without owners acting: the source searched the docs, the changelog, four announcement posts and the July 2025 press release and found no such case, while flagging that as a failure to find rather than proof of absence [15]. The adjacent change that gets mistaken for it is the 1 July 2025 press release line that "every new domain will now be asked if they want to allow AI crawlers", a question at onboarding, for new domains, governing the traffic block rather than the file [16].
One default does appear without owner action, and it is not a block. Free plan domains with no robots.txt of their own and no managed file will serve Cloudflare's Content Signals Policy, which Cloudflare says "does not express any specific preferences about your content" [17].
What to watch: whether Cloudflare adds the search-side agents to the managed list or renames the setting, whether OAI-SearchBot and Claude-SearchBot pick up their own switches, and whether the crawler reference table ever gets an Applebot-Extended row [4][6].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Cloudflare prints its managed robots.txt file in full on its managed robots.txt documentation page; it names eight user agents with Disallow: / -- Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent.
The author's assessment: the block runs along the training and search seam, and that looks deliberate.
The managed file's wildcard section reads: User-Agent: * / Content-signal: search=yes, ai-train=no, use=reference / Allow: /.
In the managed block, GPTBot is included and OAI-SearchBot is not; ClaudeBot is included and Claude-SearchBot is not; Applebot-Extended is included and plain Applebot is not.
For the GPTBot/OAI-SearchBot and ClaudeBot/Claude-SearchBot pairs, Cloudflare's crawler reference table calls the blocked agent an AI Crawler and the one left alone AI Search.
Cloudflare's crawler reference table has no Applebot-Extended row, so it cannot say what Cloudflare classifies that agent as.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Primary documents quoted verbatim, one publisher
The core factual load is carried by primary artifacts reproduced in full or quoted word for word: Cloudflare's managed robots.txt block, Cloudflare's crawler reference table, Cloudflare's Free-plan and edge-prepend documentation, OpenAI's independent-settings line, and Google's crawler and AI-features pages. Those are checkable and directly on point. Deductions: everything reaches the reader through a single publisher with no independent corroboration, one link in the chain (Applebot-Extended's classification) is explicitly unverifiable because the reference table has no such row, and the auto-enable question rests on a search that found nothing, which the author correctly declines to treat as proof.
Mechanism and defaults documented, usage unquantified
There is real deployment-shaped evidence: Cloudflare asks every new domain at onboarding whether to allow AI crawlers, Free-plan domains without their own robots.txt are served the Content Signals Policy by default, and the managed file is prepended at the edge in production. But nothing in the cluster quantifies how many domains have managed robots.txt switched on, how many crawlers honour it, or what traffic changes followed. The only end-user signal is one Reddit report whose configuration was never verified. Availability and defaults are established; actual uptake is not.
Popular and vendor framing overstate the file's reach
Positive: the surrounding claims run ahead of the artifact. The widely upvoted complaint attributes loss of ChatGPT citation, Perplexity presence and Google AI Overviews to the managed file, yet OAI-SearchBot and Perplexity are absent from it and Google states Google-Extended affects neither Search inclusion nor ranking. Cloudflare's own 'instructs known AI crawlers to stay away' framing is likewise broader than a file that leaves the AI-search halves of three vendor pairs allowed. The gap is moderate rather than extreme because the file does impose real costs Cloudflare can legitimately claim -- training inclusion and Gemini Apps grounding -- and the author's own strongest interpretive claim, that the training/search seam is deliberate, is presented as opinion.
Vendor marketing framing plus light author self-reference
Documented incentives are visible on two sides but neither is dominant. Cloudflare markets the switch in benefit language ('instructs known AI crawlers to stay away') and prompts every new domain about AI crawlers, which gives it reason to present the control as broader than the eight agents warrant; the analysis here is largely a check on that framing. On the publishing side, the author points readers to an earlier post of their own instead of restating the OpenAI bot breakdown, a mild traffic incentive, and the piece's appeal rests on debunking a popular grievance. Nothing in the cluster shows sponsorship, affiliate relationships, or a commercial stake in Cloudflare or any AI vendor.
Well-sourced on mechanics, open on causation
Confidence is solid for the mechanical claims -- the file's contents, the training/search split, and the vendor scope statements are quoted from primary documentation and would be cheap to re-verify. It is materially lower for the questions readers care most about: whether Cloudflare ever enabled the managed file without owner action (an unresolved failed search), what actually caused the r/SEO complainant's losses (unverified), and how Cloudflare classifies Applebot-Extended (no table row). Single-publisher sourcing and one explicitly subjective inference about deliberateness keep this in the middle band.
build
ChatGPT-User outfetched Googlebot for 34 days on one small site. Read your logs.1 distinct publisher
build
The AI-training bans live on the big infrastructure blogs, not the small publications1 distinct publisher
build
ChatGPT tops Google's paid-click share at 4.75%, and growth teams should reprice the auction1 distinct publisher
invest
Nearly half of ChatGPT's advisor citations were advisors' own websites1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 18, 2026