Skip to content

Build1 publisher3 min readPublished

ChatGPT-User outfetched Googlebot for 34 days on one small site. Read your logs.

A developer's 34-day Caddy log audit found 40.3 assistant fetches a day against Googlebot's 35.8. None of that traffic appears in Search Console.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying ChatGPT-User outfetched Googlebot for 34 days on one small site. Read your logs.
Generated illustration

What happened

  • In mid-July the author's Google clicks in his home market, Israel, dropped by almost half, and buyer-intent queries that had brought steady leads disappeared from Search Console.
  • The author stopped looking at Google Search Console dashboards and began reading raw server logs, where he found AI systems reading his site constantly.
  • The dataset is 34 days of complete Caddy access logs from a small business site with about 70 real human visitors a day.
  • All reported counts are HTTP 200 responses only, over 34 days.
  • ChatGPT-User made 40.3 fetches a day; Googlebot made 35.8 a day.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer published 34 days of HTTP 200 counts pulled from the Caddy access logs of a small business site, and in that window OpenAI's live-retrieval agent ChatGPT-User fetched pages 40.3 times a day against Googlebot's 35.8 [5][4][3]. The consequence is not the scoreboard between two user agents; it is that the site's most frequent readers left no trace in the dashboard the owner had been staring at, and he only found them by giving up on Search Console and parsing raw logs [2][1].

The trigger was commercial. In mid-July his Google clicks in his home market of Israel fell by almost half, and buyer-intent queries vanished from Search Console [1]. The logs told a different story: the site takes about 70 real human visitors a day [3], and across the four buckets he reported, roughly 298 machine fetches a day, about 4.3 automated requests for every human [20][21].

The breakdown matters more than the total. He grouped by intent rather than bot name: 1,907 live user retrievals, about 56 a day, meaning a person asked an assistant something and the assistant pulled a page [9][6]; 6,096 AI-search indexing requests from OAI-SearchBot, Perplexity, bingbot and Applebot, about 177 a day [10]; 897 training crawls from GPTBot, ClaudeBot and Amazonbot [11]; and 1,233 requests from classic Google [12]. That works out to 7.2 non-Google machine fetches for every Google one [22]. Two caveats on that framing. The indexing bucket is mostly not new: bingbot alone accounts for about 88 percent of it [23], since Bing crawls the site 158 times a day against Google's 36, a 4.4 to 1 ratio [7]. And the live-retrieval bucket is effectively two vendors, with ChatGPT-User and Claude-User covering about 99 percent of it [24].

Whether Bing's crawl rate is worth caring about depends on a second-hand claim: the author cites Seer Interactive finding that 87 percent of SearchGPT citations match Bing's top organic results, versus 56 percent for Google [8]. If that holds, the crawler nobody optimises for is upstream of the assistant doing the fetching.

Per-page data is where this gets operationally useful. ChatGPT-User's most-fetched URLs were the homepage at 115 fetches, a WhatsApp automation guide at 108, and a WhatsApp bot pricing guide third [14]. Claude-User concentrated hard: 233 of its 519 fetches, 45 percent, went to a single spam-detection post [15]. GPTBot, the training crawler, spent its budget on /signin and /forgot-password, including twelve hits on the login page [16]. Perplexity-User logged zero live retrievals in 34 days even though PerplexityBot made 103 indexing requests [17].

The method is unglamorous and that is the point. The analyzer is about 60 lines of Python over Caddy's JSON logs, matching user-agent regexes, with one ordering rule: the *-User patterns must be tested before the generic bot patterns because the loop stops at the first match [18]. The filter that makes the numbers honest is counting only HTTP 200s. His first pass counted every request and was inflated by security scanners spoofing OpenAI user agents, including one claiming to be GPTBot while probing /wp-admin on a site that does not run WordPress, all landing on 404s [19].

This is one site in one niche, so treat the ratios as a method rather than a benchmark. What to watch on your own logs: the live-retrieval count as a weekly series, which the author now tracks as a leading indicator for citations [29]; whether your pricing and comparison pages are in the assistant fetch list, since that is a sales conversation happening without you; and whether training crawlers are burning fetches on auth pages [16].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories