Leadership1 distinct publisher3 min readPublished
A Stanford lab found that over half of roughly 250,000 conversations involved handing AI work that affects other people or is hard to undo. Someone else's tool measured it.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
The mechanically interesting part is which tool did the work. Anthropic Insights, previously named Clio, is not an instrument built for outsiders; it is the privacy-preserving system Anthropic's own teams use to analyse patterns across millions of Claude conversations [2]. The pilot therefore did not establish that independent measurement of AI usage is possible in principle. It established that it is a query away on the vendor's own analysis stack, for whoever the vendor admits [1].
Then the arithmetic. Each team worked with roughly 250,000 Claude.ai or Claude Code conversations from April and May 2026 [4]. The Stanford SALT Lab reports that over half of conversations involved a person delegating a consequential task, defined as work that affects others or is hard to undo, which cuts against prior research suggesting people keep that sort of work for themselves [6]. Applied to the sample, that is upwards of 125,000 conversations in two months [7]. Separately, in nearly three-quarters of conversations the human set direction while Claude assisted, and usually adapted the output rather than using it verbatim [8]. The complement is on the order of 62,000 conversations where that description did not hold [9].
What the published results do not let anyone do is cross those two findings. The cell a board actually wants is consequential work delegated without the human directing it, and the aggregate release does not produce that cell. The next honest step for anyone in legal or finance is to note that professional guidance was where consequential delegation peaked, legal and financial questions in particular [c6b], and that people vary in how much they understand what Claude produced even when they were the ones steering [c8b].
Anthropic's stated reason for opening the door is that real-world interaction data sits inside a handful of labs, leaving outsiders to choose between the labs' own analyses of real usage and their own analyses of public datasets that skew casual [10]. The pilot buys a third option on the lab's terms. Those terms: review rights confined to user privacy, information that would help people violate usage policies, Anthropic's confidential information, and research accuracy [11], with researchers otherwise free to publish findings that are inconvenient for Anthropic [12]. Three of those four categories are narrow. "Research accuracy" is the one with room in it, and it is the category a lab would reach for if a finding were both awkward and contestable. Nothing in the account says it was used that way. The point is that it exists and is where pressure would land.
For operators the governance consequence is simple to state and awkward to answer. Anthropic says this is the first time external researchers have run public independent studies on an AI company's own usage data [5], the aggregate data from each project is being released [13], and Anthropic has run a further privacy audit of what was shared and is inviting expressions of interest for future work [14]. Category-level facts about how employees delegate consequential work are now produced by parties whose access is granted by the model vendor, not by the employer. The material covers Claude.ai and Claude Code conversations [4]; it does not say what happens to enterprise-tier traffic, which is the question worth putting to a vendor in writing.
Ranked by verification strength, evidence, and original report placement.
Anthropic's contractual review rights over the studies were limited to user privacy, information that could help people violate its usage policies, Anthropic's confidential information, and research accuracy.
Anthropic says it otherwise had no say in the content of the findings and that the researchers are free to publish their results even if they are inconvenient for Anthropic.
Anthropic ran a pilot earlier in the year giving external researchers access to aggregate, real-world Claude usage data; three research groups designed their own studies, Anthropic ran the data collection on their behalf, and the groups conducted their own independent analysis.
Anthropic Insights, formerly named Clio, is the privacy-preserving tool Anthropic's own teams use to analyse usage patterns across millions of Claude conversations.
The three partners were the Social and Language Technologies (SALT) Lab at Stanford University, the Human Information Processing Lab at the University of Oxford, and METR, a non-profit that evaluates frontier AI models.
Each group developed its own research questions and used Anthropic Insights to conduct privacy-preserving analysis of roughly 250,000 Claude.ai or Claude Code conversations from April-May 2026.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Quantified but single-sourced and partly unpublished
The headline findings carry concrete denominators - roughly 250,000 conversations from April-May 2026, over half consequential, nearly three-quarters human-directed - which is more than a typical vendor post offers. But everything in the cluster comes from Anthropic's own write-up: the Oxford lab's report is not yet public, the METR productivity analysis is cut off mid-sentence in the supplied text, and no sampling methodology or classifier definition for 'consequential task' is disclosed. Only the SALT Lab findings are linked for inspection.
Three institutions, one pilot cohort, data now public
Real but small-scale uptake: three named research institutions completed studies, aggregate datasets from all three projects were released publicly, and an expression-of-interest form was opened to widen access. That is concrete adoption of the access mechanism, not of a broadly used programme - the cohort is a single pilot, no downstream users of the released data are documented, and no follow-on access grants are disclosed.
Framing runs ahead of what is currently inspectable
Modest overstatement. The 'we believe this is the first time' novelty framing and the assurance that Anthropic 'otherwise had no say' in findings are self-asserted and unverifiable, and the independence story is undercut slightly by a retained 'research accuracy' review right that only Anthropic defines. Against that, the core delegation and human-direction numbers are specific and the aggregate data was actually released, so the gap is narrow rather than promotional - the substantive finding is arguably underplayed relative to the process framing.
Subject publishes on its own transparency and its own tool
Strong incentive alignment to disclose favourably: Anthropic is simultaneously the data holder, the operator of the analysis tool being showcased, the party whose openness is being praised, and the holder of contractual review rights over the studies - including a 'research accuracy' category. It also benefits reputationally and in policy debates from being seen as the lab that opened its usage data, and it selected which early results to summarise while two writeups remain unpublished.
Moderate: specific numbers, one interested source
Confidence is limited by structure rather than vagueness. The factual spine - partners, sample size, window, disclosed shares, data release - is stated precisely and is unlikely to be misreported by the publisher. But the cluster contains a single interested publisher, the two derived volume estimates depend entirely on that publisher's shares, and interpretation of the consequential-delegation finding cannot be checked until the underlying classifier and the remaining writeups are public.
build
Anthropic's Insights pilot sets the ceiling on what outside evaluators can verify1 distinct publisher
science
Half of 249,834 Claude conversations carried consequential work, and the hard ones ran twice as long1 distinct publisher
build
The $559M-versus-$12.3B quarter matters more than the $65B run rate4 distinct publishers
invest
The retail route into Anthropic is mostly a fee on everything that is not Anthropic1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.