Skip to content

Build1 publisher3 min readPublished

Cloudflare's new-domain defaults block AI agents on pages that carry ads

Cloudflare blocks AI agent and training traffic by default on ad-carrying pages of domains added since September 15, 2026. Search crawlers stay allowed, so a browsing assistant can be refused a page that a search indexer reads freely.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Cloudflare split AI traffic into Search, Agent and Training categories on July 1, 2026, with Agent covering software that acts for a person, usually in real time.
  • Site owners can allow, block, or block only on ad-carrying pages for each category, and can override the new defaults.
  • A Disallow AI Training option lets owners refuse training while still letting accountable mixed-use crawlers index their sites for search.
  • A blocked agent may get a 403, a Cloudflare challenge, a rate-limit response or another error, and cannot always tell whether the block was deliberate.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure A browsing assistant on a new domain loses ad-supported pages that a search indexer can still fetch, so the same URL can succeed for one product and fail for another.
  • constraint Because the default block is set per page, one successful fetch tells an agent little about a domain, and refusals have to be handled page by page.
  • decision Agent teams have to choose what a challenge or 403 triggers; the post argues it should end the attempt and send the agent to an API, feed or sitemap.

On a domain onboarded after September 15, the default policy takes two inputs. One is the category a request falls in. The other is whether the page displays ads [7]. By default no category is blocked across a whole site, and two of the three are blocked only where ads appear [16]. An agent that fetches an ad-free page on a new domain gets through. The same agent fetching an ad-supported article on that domain is blocked, while a search crawler gets both pages [17].

The Agent row matters most for anyone shipping an assistant that browses. Asked for the latest information on a product, an assistant may open several pages, read them and return what is relevant, and the post notes that this is different from a traditional search crawler [20]. According to the post, Cloudflare's Agent definition is very close to how many AI agents work today [18].

The post does not explain how Cloudflare decides which category a request belongs to, or how it determines that a page displays ads. The categories are defined by what the traffic is for [2][3][4]. So whatever signal Cloudflare uses to infer purpose decides which row a request lands in. A team that wants its traffic classified correctly needs that signal documented before it can do anything about it. Crawlers that do several jobs make this harder, and for that case the post points to Cloudflare's documentation on mixed-purpose crawlers [8].

The failure handling contains one piece of good engineering. Spicrawl, the fetching tool the post uses as its example, reports an upstream bot challenge as ERR::UPSTREAM::CHALLENGE. It passes the target site's own HTTP status back in an X-Target-Status header [11]. A retry loop tends to merge two questions: did my fetcher fail, or did the site refuse me? This design keeps them apart. The post makes the same point: a scraper's HTTP response is not always the same as the target site's actual response [12]. I think any fetch layer in front of an agent should expose the origin status the same way.

The post rules out retrying until the block disappears [13]. "If a website has clearly decided not to allow your automated traffic, treat that as a decision," the post's author wrote [14]. The alternatives it lists are reading robots.txt, terms of service and API documentation, and using an API, feed, sitemap or other official source where one exists [15].

Existing sites keep their settings [6]. The new defaults therefore reach more Cloudflare sites only as new domains onboard or existing owners opt in. The post's forecast is conditional: if more site owners use the new defaults, some AI agents may encounter more blocked pages [19].

What to watch

  • Cloudflare documentation on how a request is assigned to Search, Agent or Training, and whether agent operators can declare their category.
  • Whether existing Cloudflare customers switch to the September defaults by hand; each one that does extends the ad-page agent block to an established site.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories