FoxyInvoice prerenders every public Angular route to static HTML at build time so crawlers that skip JavaScript get the full page text. A build step and a one-line drift check take the place of the always-on SSR server its team declined to run.
Reality
- Evidence40
- Adoption15
- Hype gap+20
- Incentives70
- Confidence40
DataDome says malicious automated traffic grew 124% in the year to June. Its June test of 21,491 popular sites found that claiming a trusted crawler's name is usually enough to get through the door.
Reality
- Evidence55
- Adoption60
- Hype gap+20
- Incentives82
- Confidence50
A site logged twenty requests for its unreachable /.env file from a client naming itself ClaudeBot. GreyNoise counted 824 addresses scanning under forged AI crawler names, and saw no request for robots.txt in the whole campaign.
Reality
- Evidence60
- Adoption45
- Hype gap−12
- Incentives65
- Confidence62
Google-Extended and Applebot-Extended are robots.txt tokens that never issue a request, so the nine agents OpenAI, Anthropic and Perplexity document are the only AI traffic a WordPress site can confirm for itself.
Reality
- Evidence74
- Adoption48
- Hype gap−8
- Incentives68
- Confidence60
Before starting a column about getting cited by LLMs, a developer curled medium.com/robots.txt and found GPTBot, ClaudeBot and six other agents disallowed. Common Crawl and Google-Extended are still allowed.
Reality
- Evidence62
- Adoption40
- Hype gap−8
- Incentives68
- Confidence55
A dev.to post sets out two ways assistant retrieval fails while ordinary search keeps working, one in robots.txt rules written per user-agent and one in copy that appears only after JavaScript runs. It offers no measurement.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
Debugging an agent endpoint has to start on the wire, because a client that reports a parse error is describing the last thing that touched the bytes rather than the machine that swapped them out.
Reality
- Evidence64
- Adoption20
- Hype gap+8
- Incentives32
- Confidence62
One site logged 13,491 crawler requests in thirty days and could prove the origin of roughly 2,200 of them, because half the traffic came from vendors that publish no addresses to check against.
Reality
- Evidence48
- Adoption34
- Hype gap+28
- Incentives55
- Confidence52
A 41-day test switched off sitemaps and breadcrumbs so client-side navigation was the only way in, and the crawlers behind ChatGPT and Claude never arrived. Googlebot did, then needed 41 days to notice the fix.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+22
- Incentives62
- Confidence41
In a 30,180-request test across roughly 1,000 pages, Googlebot followed JavaScript-injected links. GPTBot, ClaudeBot and PerplexityBot did not follow them at all.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+26
- Incentives62
- Confidence41
The managed robots.txt file lists eight crawlers. Three of them are the training half of a pair whose search half is left allowed, and Perplexity is not in the file at all.
Reality
- Evidence72
- Adoption24
- Hype gap+34
- Incentives46
- Confidence61
A developer read the current Terms of Service for every feed in his news digest. The clauses that would break a summarization pipeline clustered in VC-data sites and well-lawyered corporate blogs.
Reality
- Evidence61
- Adoption21
- Hype gap+14
- Incentives58
- Confidence56
A developer's 34-day Caddy log audit found 40.3 assistant fetches a day against Googlebot's 35.8. None of that traffic appears in Search Console.
Reality
- Evidence44
- Adoption24
- Hype gap+28
- Incentives58
- Confidence52