Skip to content

Build1 publisher3 min readPublished

Every AI-crawler opt-out in a 185-site European robots.txt census blocks the whole site

Pennyforge's 1 October scan of 185 European sites found 35 AI-crawler opt-outs in robots.txt, every one sitewide. That gives the EU consultation on opt-out signals, closing 3 November, a baseline where three in four sites show crawlers no AI rule at all.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Every AI-crawler opt-out in a 185-site European robots.txt census blocks the whole site
Generated illustration

What happened

  • Another 39 sites, including zalando.de and sciencedirect.com, answer requests for robots.txt with HTTP 403 from CDN or bot walls.
  • CCBot is the most-blocked agent on the panel, disallowed by 26 sites, and not one site explicitly allows it.
  • Only nine sites carry an explicit Allow for any AI agent, mostly consumer-electronics retailers such as mediamarkt.nl and saturn.de.
  • MistralAI-SearchBot, the newest search bot on the list, has no allow, disallow or mention on any of the 185 sites.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A TDM signal that depends on path-level or purpose-level rules would ask publishers for a practice that none of the 185 sites shows today.
  • exposure Under the convention the census cites, the 39 sites behind 403 walls have no policy a polite crawler can read, whatever their CDN is set up to block.
  • cost Per-agent lists cost an edit for every new bot, and the five sites blocking 10 of 12 agents had not made that edit for Mistral's search bot.
  • capability At $0 in data and a seven-minute pass, the method can be rerun on a team's own site list and dated before the consultation closes.

In robots.txt, the only way to state purpose is the agent name [18]. The census describes Europe's de facto opt-out as a fragmented, voluntary robots.txt with per-user-agent rules [18]. According to the census, the convention, RFC 9309 of 2022, was written for a web with one kind of crawler [6]. Pennyforge counts at least three jobs now: bulk collection for training, search-index crawling for answer engines, and on-demand fetches for one user and one page [5]. A site that wants to refuse the first and admit the other two has to name each bot and give it a rule [18].

Few sites have written those rules. The scan checked each robots.txt against a fixed list of 12 AI user-agents [4]. Subtracting the 77 silent files from the 121 served leaves 44 sites that mention any of the 12 [1]. Thirty-five of those 44 block at least one agent from the whole site, about 80% [2]. Pennyforge put the electronics retailers' Allow lines down to "the visible-in-AI-search incentive, not a content policy" [15].

After CCBot and GPTBot, the full-disallow counts run ClaudeBot 18, Google-Extended and Bytespider 17 each, and Meta-ExternalAgent 16 [13]. The census reads CCBot's lead as concern about Common Crawl's role as a training-data conduit [13]. The post does not give counts for OAI-SearchBot, ChatGPT-User, PerplexityBot or Perplexity-User [13]. Without them, the number of sites that block training while admitting search cannot be read from the summary. Pennyforge cites a per-site file with the full 12-agent table [19].

Five sites block 10 of the 12 agents: amazon.de, faz.net, derstandard.at, rtbf.be and lalibre.be [11]. No site names MistralAI-SearchBot [16], so it is one of the two agents each of those five leaves open, by omission [3]. "Policy is written for yesterday's crawlers," Pennyforge wrote [16]. The single government entry is gov.uk, blocking Meta-ExternalAgent alone [12].

The best method decision in the post concerns /.well-known/ai-crawler. Twenty sites answered HTTP 200 for that path [17]. All 20 were single-page apps returning their HTML shell, the same page they would hand back for any path at all [17]. A checker that trusted status codes would have reported 20 adopters [4]. Pennyforge classified by content type and found none [17].

These rates describe this panel. Retail shops make up 89 of the 185 sites, about 48% [5]. The census says opt-outs cluster by sector, with publishers blocking while governments and shops stay mostly quiet [12]. For the 18.9% opt-out rate [10] to describe another list of sites, that list needs a similar sector mix; by the census's own sector reading, a media-heavy panel would push it up [12]. It is also a single seven-minute pass [7]. The Commission's consultation asks what a text-and-data-mining opt-out signal should be, and it closes 33 days after that pass [7] [6].

What to watch

  • A second dated Pennyforge pass on the same 185-site panel and method, showing whether any site adds MistralAI-SearchBot or a path-level rule.
  • Release of per-site counts for OAI-SearchBot, ChatGPT-User, PerplexityBot and Perplexity-User; those numbers would show directly how many sites split training from search.
  • Whether the Commission's consultation outcome specifies a signal beyond per-agent robots.txt rules, such as a file at a fixed path with a defined format.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories