Build1 publisher3 min readPublished
Cloudflare's one-click AI block names GPTBot, not the bot that decides if ChatGPT cites you
The managed robots.txt file lists eight crawlers. Three of them are the training half of a pair whose search half is left allowed, and Perplexity is not in the file at all.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Cloudflare prints its managed robots.txt file in full on its managed robots.txt documentation page; it names eight user agents with Disallow: / -- Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent.
- The author's assessment: the block runs along the training and search seam, and that looks deliberate.
- The managed file's wildcard section reads: User-Agent: * / Content-signal: search=yes, ai-train=no, use=reference / Allow: /.
- In the managed block, GPTBot is included and OAI-SearchBot is not; ClaudeBot is included and Claude-SearchBot is not; Applebot-Extended is included and plain Applebot is not.
- For the GPTBot/OAI-SearchBot and ClaudeBot/Claude-SearchBot pairs, Cloudflare's crawler reference table calls the blocked agent an AI Crawler and the one left alone AI Search.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Cloudflare's managed robots.txt feature writes a block for eight named user agents, and the list runs along the seam between training crawlers and search crawlers [1][2]. The practical result is that teams who flip the switch to stop AI answer engines from using their content are instead removing themselves from training corpora while leaving the citation-facing bots untouched [9].
The file is published in full on Cloudflare's own docs page: Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent, each with Disallow: / [1]. Above them sits a wildcard block that allows everything and carries the line "Content-signal: search=yes, ai-train=no, use=reference" [3]. That is not ambiguous drafting. Search is explicitly permitted; training is explicitly refused.
Read against Cloudflare's crawler reference table, the pattern repeats three times. GPTBot is blocked and OAI-SearchBot is not; ClaudeBot is blocked and Claude-SearchBot is not; Applebot-Extended is blocked and plain Applebot is not [4]. For the first two pairs the table labels the blocked agent an AI Crawler and the unblocked one AI Search [5]. The table has no Applebot-Extended row at all, so Cloudflare's own reference cannot tell you which side of the line it sits on [6]. Perplexity is not named anywhere in the file [7].
The confusion has a documented cost. An r/SEO post from April, with 53 points and 40 comments, claimed Cloudflare had cut the author's site off from ChatGPT, from Perplexity and from Google's AI Overviews [8]. None of those three is what the eight agents govern [9]. What they do govern is real: GPTBot controls inclusion in OpenAI's training data, and Google-Extended controls grounding in Gemini Apps [10]. OpenAI's documentation is one line on the relationship between its bots: "Each setting is independent of the others" [11].
Google-Extended is the entry that does the most damage, because the name reads as if it governs AI answers in Search. Google's crawler documentation says otherwise: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search" [12]. It also notes that Google-Extended "doesn't have a separate HTTP request user agent string" and that crawling is done with existing Google user agent strings, which disposes of the common advice to grep your access logs for it [13].
The control that does exist for AI Overviews is the snippet. Google's AI features page states that robots.txt directives for Googlebot are the control for how sites are crawled for Search, and points owners to nosnippet, data-nosnippet, max-snippet or noindex to limit what is shown from their pages [14]. Opting out of the AI answer means opting out of the snippet.
On the recurring claim that Cloudflare turned this on for existing domains without owners acting: the source searched the docs, the changelog, four announcement posts and the July 2025 press release and found no such case, while flagging that as a failure to find rather than proof of absence [15]. The adjacent change that gets mistaken for it is the 1 July 2025 press release line that "every new domain will now be asked if they want to allow AI crawlers", a question at onboarding, for new domains, governing the traffic block rather than the file [16].
One default does appear without owner action, and it is not a block. Free plan domains with no robots.txt of their own and no managed file will serve Cloudflare's Content Signals Policy, which Cloudflare says "does not express any specific preferences about your content" [17].
What to watch: whether Cloudflare adds the search-side agents to the managed list or renames the setting, whether OAI-SearchBot and Claude-SearchBot pick up their own switches, and whether the crawler reference table ever gets an Applebot-Extended row [4][6].