Product1 distinct publisher2 min readUpdated
TollBit's report has unauthorised crawlers wearing Google's name, rotating IPs and renting home broadband. Access control now buys attenuation, not exclusion.
The Product Desk · Product desk
Compiled by The Product DeskSomething wrong?How this is made
A blocklist reads a request for signals: which user agent it claims, which address range it came from, how fast it arrives. Route a scrape through a large network of consumer devices and every one of those signals reads as somebody at home on a laptop [3]. At that point the publisher is not deciding whether to admit a crawler. It is guessing which of its readers is one, and paying to guess on every request.
The arithmetic favours the crawler. A retry after a block costs the scraper another rotated address [3]. Each evaluation costs the site engineering time it has already spent. People Inc. has deep expertise here and still has not sealed it [4]. Impersonating Google is the neatest version of the problem, because Googlebot is the one identity most publishers will not turn away, so the borrowed name inherits the benefit of the doubt the publisher extends for commercial reasons.
Which is why the number worth pulling out of the licensing market is not the headline count. Rob Kelly's tally has 94 publicly announced deals, with about four in ten carrying training rights [6]. That is roughly 38 deals with training rights and 56 without [9]. A clear majority of the signed market is now buying live access to serve answers rather than a corpus to train on, which is the same direction Fast Company argues the whole fight should move: from what gets scraped to how it gets used [5].
The trouble is what that side pays. Pew put clicks on links inside a Google AI summary at about 1% of visits, against 15% on a results page without one [8], a reduction of roughly 93% [10]. Digiday found brands unable to connect AI presence to business outcomes, and being the authoritative source in an answer is not monetizable in itself [7]. So the asset publishers are being asked to license is placement, and placement returns almost nothing in referral terms. Any usage-based regime has to be priced per retrieval, because no click is arriving to pay for it.
One more thing is worth naming about where the evidence comes from. TollBit sells payment rails between publishers and AI crawlers [1]. Its report is therefore an accurate description of the population its own product cannot reach: an operator willing to rent home broadband to look like a reader is not queueing up to install a meter. Whatever replaces the blocklist has to be something a publisher can do to its content after the fetch, without the fetcher's cooperation. Nothing in the material describes that mechanism yet.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The report shows evidence that "bad" bots sometimes try to masquerade as legitimate bots like Google's, often rotate IP address if a first scrape is blocked, and that some industrial-scale scraping companies use huge networks of devices in people's homes to make traffic look like it comes from real people.
Rob Kelly of the Media and the Machine Substack tallied 94 publicly announced deals and found only about four in 10 now include training rights, with the market shifting from "buy content to build better models" toward "license content to deliver better answers."
Pew Research clocked clicks on links inside a Google AI summary at about 1% of visits, compared with 15% on a results page without one.
TollBit builds payment rails between publishers and AI crawlers, and publishes a State of the Bots report.
People Inc.'s Chief Innovation Officer, Jonathan Roberts, outlined the company's approach in TollBit's report: aggressively block unauthorised bots while allowing access to legitimate crawlers, which typically means some kind of licensing agreement.
Roberts concedes that unauthorised crawlers have become harder and harder to identify and block, and that even People Inc. has not been entirely successful at blocking it all.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One column relaying others' data
The factual core is attributed and specific — a named executive's concession, three named evasion techniques, a 94-deal tally, and Pew's 1%/15% split — but all of it reaches the reader secondhand through a single opinion column, with no primary report excerpt, no methodology, and no independent corroboration of the crawler-detection findings.
Deals scaling, enforcement not
There is real, countable adoption on the licensing side — 94 announced deals, with the mix moving toward retrieval rights — and blocking is clearly deployed at scale by at least one large publisher. But the column states plainly that no monetization or enforcement mechanism has become an industry standard, and the proposed answer-layer licence check has no known implementation, so operative adoption stops well short of the story's prescription.
Mostly deflationary, prescription runs ahead
The framing is corrective rather than promotional: it deflates both the training-data payday and the idea that blocking can hold. The modest overstatement comes from building a strategic prescription — police the answer, verify licences at citation time — on top of a single vendor's unverified detection data and an unimplemented mechanism, while the empirical parts (Pew figures, deal tally) are proportionate to their sources.
Interested parties throughout the chain
Nearly every voice has a commercial position in the outcome: TollBit sells payment rails between publishers and crawlers and authored the report underpinning the evasion findings; People Inc.'s executive is describing a licensing-dependent strategy in that vendor's publication; ProRata's business model is cited as proof that attribution works; and the column itself carries a promotion for the author's own newsletter. None of these interests is disclosed as a caveat in the piece.
Directionally credible, thinly sourced
The direction of travel — blocking degrades, training rights lose value, retrieval displaces referrals — is coherent and partly backed by numbers from named third parties. Confidence is held down by having one publisher, no primary documents, no adversary or platform response, and no quantification of the leakage the thesis depends on.
product
OpenAI shipped a teen ChatGPT. The over-65 cohort doubled to 23% and got nothing.1 distinct publisher
product
White House lets vetted firms hack back and leaves liability blank for 60 days1 distinct publisher
product
Two years of jobs data invert the AI risk rankings: clerks shrink, managers grow1 distinct publisher
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 21, 2026