Build1 distinct publisher3 min readPublished
Three groups received counts and cluster descriptions from roughly 250,000 conversations each. The raw chats stayed on Anthropic's servers, and so did the classifier that produced the numbers.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The redaction figures are more useful than their size suggests. Anthropic removed or altered 1.9% of Stanford's clusters, and those clusters accounted for 4.28% of Stanford's conversation sample [8][9]. Divide one by the other and the cut clusters were roughly 2.3 times the average size [1]. Oxford's equivalent ratio is close to even at 1.16, METR's is 1.64 [2]. Across the three samples the withheld material covers about 27,700 conversations, near 3.7% of the pilot's 750,000 [3]. Anthropic says most removals concerned descriptions that could reveal how users bypassed safeguards [10]. That is a defensible reason to cut, and it is also the category an outside evaluator would most want to read.
The interface determines what is answerable at all. Researchers write questions called facets, Anthropic runs them, Claude classifies and clusters the conversations, and the outside group receives category counts, percentages and cluster prose [5][6]. Nobody outside the lab ran an unrestricted search or saw a raw conversation [5]. So every number in METR's Claude Code work is the output of a classifier owned by the company whose product is under study, applied to data the evaluator cannot inspect. There is no move available inside this arrangement that lets METR pull two hundred sessions by hand and check whether a facet was scored the way it was meant.
Anthropic's contractual review covered privacy, information that could help users evade safeguards, confidential company information, and research accuracy, with the outside groups keeping control of their questions and conclusions [7]. Three of those four are unremarkable. Research accuracy is the one that matters, because it seats the subject of the study in the loop on whether findings about it are correct.
The sampling frame sets a second limit. The Claude.ai samples came from Free, Pro and Max users, with Team, Enterprise and API traffic excluded, while METR's Claude Code sample was consumer users who had opted into data use for model improvement [4]. Stanford's headline result, that 56% of conversations with an actionable task involved consequential or higher work and 12% reached its high-stakes tier, concentrated in legal and financial guidance [13], is therefore a statement about individual subscribers. It says nothing measured about the deployments that carry contracts and administrators.
Imperial College London's red team could not reidentify users from the released material, but it did link one cluster to a widely used open-source project from distinctive language associated with the tool [11]. Anthropic's stated remedy is higher minimum-user thresholds and less distinctive cluster descriptions [12]. Both raise the floor on what an outsider can resolve, which is the trade being made: privacy margin bought with descriptive detail.
None of this reads as bad faith. It is a question of what aggregate counts can bear. They can describe a distribution, including the finding that 72% of conversations were human-led and friction appeared in 49.7% [14][15]. They cannot support an audit, because auditing means going to look at the specific thing you doubt, and that is the one operation the pipeline does not offer.
Ranked by verification strength, evidence, and original report placement.
Claude classified the conversations, grouped similar answers and generated descriptions for the resulting clusters. The external groups received category counts, percentages and cluster descriptions, and Anthropic published those aggregate outputs as a CC BY 4.0 dataset on Hugging Face.
Stanford's SALT Lab paper examined 249,834 Claude.ai conversations. Among conversations involving an actionable task, 56% concerned work classified as consequential or higher, meaning it affected other people or would be difficult to reverse, and 12% reached the researchers' high-stakes category, with legal and financial guidance among the areas where consequential delegation was concentrated.
Anthropic opened a controlled route for outside researchers to study how people use Claude, giving three research groups aggregated findings from real conversations while keeping the underlying chats on Anthropic's servers.
Anthropic described the pilot in a five-post thread on X and an accompanying research note on August 26th. Stanford University's Social and Language Technologies Lab, Oxford University's Human Information Processing Lab and the independent evaluation group METR designed their own studies using Anthropic Insights, the privacy-preserving analysis system previously called Clio.
Each group worked from a separate sample of roughly 250,000 conversations collected during April and May 2026. Stanford and Oxford studied Claude.ai conversations, METR examined Claude Code sessions, putting the pilot's combined scope at about 750,000 conversations rather than one shared dataset.
The Claude.ai samples came from Free, Pro and Max users, with Team, Enterprise and API traffic excluded. The Claude Code sample covered consumer users who had opted into allowing their data to be used for model improvement, and its collection window crossed a model release so METR could compare activity before and after.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific and quantified, but almost entirely lab-supplied
The cluster carries unusually concrete numbers — sample sizes, per-group redaction rates by cluster and by sample share, and a completed Stanford paper on 249,834 conversations — plus a published CC BY 4.0 aggregate dataset that third parties can inspect. But nearly every figure originates in Anthropic's own thread, research note and technical appendix, only one of the three studies is complete, and the source itself notes that classification errors are hard to detect because researchers cannot read the underlying chats. One publisher, no independent replication.
Real pilot at scale, one of three studies delivered
Adoption is concrete but bounded: three named external groups actually ran studies, roughly 750,000 conversations were processed, and an aggregate dataset shipped publicly. Against that, this is explicitly a pilot on consumer tiers only, Oxford's study is underway, METR has not published, and there is no evidence of other labs adopting a comparable access route.
Transparency framing runs slightly ahead of verifiability
The substantive facts are modest and well documented, and the source is deliberately caveat-forward rather than promotional. The gap is mild and comes from the framing of the arrangement as external research access when Anthropic writes no research questions but does control the sampling, the classifier, the redaction and the release — and when the reported usage percentages rest on Clau95de judging Claude, a method whose failure mode already forced removal of an Oxford facet.
Publisher of the evidence also reviews and redacts it
Anthropic benefits reputationally from demonstrating third-party scrutiny while retaining contractual review rights that include research accuracy, and it alone reports how much was removed. Removals are justified by legitimate safeguard-evasion and privacy concerns, but those same categories are unfalsifiable from outside. The sampling choices — consumer tiers in, Enterprise and API out — also keep commercially sensitive traffic out of scope.
Internally consistent single-source account
The account is detailed, arithmetically coherent, and grounded in a published dataset and one completed paper, which supports moderate confidence in the described mechanics. Confidence is capped by the absence of any second publisher, the reliance on lab self-reporting for the governance numbers, the model-as-judge basis for the usage percentages, and two studies still outstanding.
build
The $559M-versus-$12.3B quarter matters more than the $65B run rate4 distinct publishers
product
Thomson Reuters spent $40M to make a $450K training run worth doing2 distinct publishers
leadership
Anthropic's own telemetry: 93% of permission prompts approved. Budget for blast radius, not reviewers1 distinct publisher
invest
The retail route into Anthropic is mostly a fee on everything that is not Anthropic1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026