Build1 publisher3 min readPublished
Every AI-crawler opt-out in a 185-site European robots.txt census blocks the whole site
Pennyforge's 1 October scan of 185 European sites found 35 AI-crawler opt-outs in robots.txt, every one sitewide. That gives the EU consultation on opt-out signals, closing 3 November, a baseline where three in four sites show crawlers no AI rule at all.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Another 39 sites, including zalando.de and sciencedirect.com, answer requests for robots.txt with HTTP 403 from CDN or bot walls.
- CCBot is the most-blocked agent on the panel, disallowed by 26 sites, and not one site explicitly allows it.
- Only nine sites carry an explicit Allow for any AI agent, mostly consumer-electronics retailers such as mediamarkt.nl and saturn.de.
- MistralAI-SearchBot, the newest search bot on the list, has no allow, disallow or mention on any of the 185 sites.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A TDM signal that depends on path-level or purpose-level rules would ask publishers for a practice that none of the 185 sites shows today.
- exposure Under the convention the census cites, the 39 sites behind 403 walls have no policy a polite crawler can read, whatever their CDN is set up to block.
- cost Per-agent lists cost an edit for every new bot, and the five sites blocking 10 of 12 agents had not made that edit for Mistral's search bot.
- capability At $0 in data and a seven-minute pass, the method can be rerun on a team's own site list and dated before the consultation closes.
In robots.txt, the only way to state purpose is the agent name [18]. The census describes Europe's de facto opt-out as a fragmented, voluntary robots.txt with per-user-agent rules [18]. According to the census, the convention, RFC 9309 of 2022, was written for a web with one kind of crawler [6]. Pennyforge counts at least three jobs now: bulk collection for training, search-index crawling for answer engines, and on-demand fetches for one user and one page [5]. A site that wants to refuse the first and admit the other two has to name each bot and give it a rule [18].
Few sites have written those rules. The scan checked each robots.txt against a fixed list of 12 AI user-agents [4]. Subtracting the 77 silent files from the 121 served leaves 44 sites that mention any of the 12 [1]. Thirty-five of those 44 block at least one agent from the whole site, about 80% [2]. Pennyforge put the electronics retailers' Allow lines down to "the visible-in-AI-search incentive, not a content policy" [15].
After CCBot and GPTBot, the full-disallow counts run ClaudeBot 18, Google-Extended and Bytespider 17 each, and Meta-ExternalAgent 16 [13]. The census reads CCBot's lead as concern about Common Crawl's role as a training-data conduit [13]. The post does not give counts for OAI-SearchBot, ChatGPT-User, PerplexityBot or Perplexity-User [13]. Without them, the number of sites that block training while admitting search cannot be read from the summary. Pennyforge cites a per-site file with the full 12-agent table [19].
Five sites block 10 of the 12 agents: amazon.de, faz.net, derstandard.at, rtbf.be and lalibre.be [11]. No site names MistralAI-SearchBot [16], so it is one of the two agents each of those five leaves open, by omission [3]. "Policy is written for yesterday's crawlers," Pennyforge wrote [16]. The single government entry is gov.uk, blocking Meta-ExternalAgent alone [12].
The best method decision in the post concerns /.well-known/ai-crawler. Twenty sites answered HTTP 200 for that path [17]. All 20 were single-page apps returning their HTML shell, the same page they would hand back for any path at all [17]. A checker that trusted status codes would have reported 20 adopters [4]. Pennyforge classified by content type and found none [17].
These rates describe this panel. Retail shops make up 89 of the 185 sites, about 48% [5]. The census says opt-outs cluster by sector, with publishers blocking while governments and shops stay mostly quiet [12]. For the 18.9% opt-out rate [10] to describe another list of sites, that list needs a similar sector mix; by the census's own sector reading, a media-heavy panel would push it up [12]. It is also a single seven-minute pass [7]. The Commission's consultation asks what a text-and-data-mining opt-out signal should be, and it closes 33 days after that pass [7] [6].
What to watch
- A second dated Pennyforge pass on the same 185-site panel and method, showing whether any site adds MistralAI-SearchBot or a path-level rule.
- Release of per-site counts for OAI-SearchBot, ChatGPT-User, PerplexityBot and Perplexity-User; those numbers would show directly how many sites split training from search.
- Whether the Commission's consultation outcome specifies a signal beyond per-agent robots.txt rules, such as a file at a fixed path with a defined format.