Product1 distinct publisher3 min readPublished
A Stanford-affiliated project pooled seven consented chat datasets and applied Anthropic's own filter. Forty-eight percent of conversations were discarded, and the discards were the sensitive ones.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
A research project called the AI Observatory, co-led by Stanford Trustworthy AI Research (STAIR) Lab PhD candidate Anka Reuel, has published an independent analysis of real chatbot conversations drawn from seven existing datasets collected with user consent, covering models including Claude and Gemini [3]. It matters because vendor-published usage reports are currently the main evidence base for how people actually use these systems, and, as Reuel puts it, "There is no independent source to corroborate it" [1][2].
The most concrete result is a replication test. The Anthropic Economic Index is one of the most widely cited sources of AI usage data, and by design it covers work- and productivity-related uses of Claude, filtering out conversations that are not [7]. When the Observatory team ran Anthropic's method over its own corpus, 48% of conversations were filtered out [8]. The discarded material was not neutral residue. Non-work conversations that got filtered were more likely to involve health and relationships (44.2%, against 31.2% in Anthropic's analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%) [9][10][11][12]. That is roughly five times the harassment and hate rate and about seven times the sexual content rate [1][2].
OpenAI's own 2025 report points the same direction from the inside: it found that only 30% of consumer use was work-related [13], which leaves 70% outside the frame that productivity-indexed reporting is built to see [4]. Anthropic has published separate posts on companionship use and on attempts to generate CSAM [14]. David Widder of UT-Austin's School of Information, who is not involved in the project, said the value of the Observatory is the bird's eye view rather than findings sectioned off into separate reports [15].
The time series is more interesting than the headline gap. Across datasets spanning 2023 to 2025 [16], conversations in WildChat, one of the largest and most detailed sets in the study, got longer and more elaborate, measured by prompt tokens, response tokens, and conversation turns [17]. Small talk rose significantly, which the researchers read as increasing companionship use, while assistants disclosed that they were chatbots less often [18]. Exchanges labelled sensitive, including sexual harassment and hate speech, declined, which the team suggests may reflect more effective safeguards [19].
Usage also splits sharply by product, in topics, interaction style, conversation structure, and both the likelihood and the type of sensitive use [20]. Grok and Gemini drew more information retrieval; Grok was especially popular for news and politics and was also where misinformation concentrated, consistent with other research, and xAI did not respond to a request for comment [21]. Anthropic's models skewed to coding, Gemini to social and roleplay, ChatGPT to homework [22]. Even model versions diverged: conversations were shorter on GPT-3.5 and longer and more iterative on GPT-4o [23]. The Observatory's stated purpose is to give researchers and policymakers an independent read [4], on the argument that consequential decisions about benefits and risks are being made on very limited data [5].
Watch whether Anthropic or OpenAI publishes the filtered-out share of their own corpora, which is the single number that would make their indexes auditable. Watch also whether the declining sensitive-use trend holds once companionship-heavy products are measured separately, and whether any regulator cites consented third-party datasets rather than vendor blog posts.
Ranked by verification strength, evidence, and original report placement.
Reuel is co-lead of a new research project called the AI Observatory, a public platform that aggregated and analyzed real AI conversations with popular models like Claude and Gemini, collected with users' consent through seven existing datasets.
The AI Observatory's stated intent is to provide independent sources of information for researchers and policymakers to assess how people are using generative AI.
The AI Observatory found that AI use differs significantly across models and has changed over time, and its research shows many more sensitive behaviors than are captured in reports from major AI companies, which the researchers say focus more on work than on personal use.
AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but AI researchers say they only release the data they want the public to see.
Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab, says of vendor usage reports: "There is no independent source to corroborate it."
When the AI Observatory team applied Anthropic's methods to their dataset, they found that nearly half of the conversations, or 48%, would have been filtered out.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific numbers, one outlet, no reproducible artifact
The cluster provides named researchers, a described corpus (24,521 conversations, 85,633 turns, 5,000 users, 52 models, 2023-2025), a stated method (reapplying Anthropic's filter), and precise comparative percentages, plus an uninvolved outside expert and an on-record Anthropic statement. It is nonetheless a single publisher with no link to the study or platform, no publication or peer-review status, no verification that the reimplemented filter matches Anthropic's pipeline, and self-acknowledged underrepresentation of sensitive uses in donated data.
Just published; no uptake signals yet
Adoption evidence is limited to the platform's publication and one uninvolved academic commenting favorably. The cluster shows no citations, policymaker use, researcher access numbers, or vendor engagement with the Observatory's findings, while the incumbent vendor indices it challenges are described as widely cited and built on 1M-1.5M conversation datasets.
Framing outruns corpus scale
The 'independent check' framing and the vanishing-half headline claim more authority than a 24,521-conversation donated corpus can carry against 1M and 1.5M conversation vendor datasets, especially since the researchers concede sensitive uses are likely underrepresented and since the filter comparison was reimplemented rather than audited. The overstatement is moderate rather than severe: the specific figures are reported plainly, the caveats appear in the article, and Anthropic's scoping of its index to work uses is not in dispute.
Challenger researchers, non-responsive vendors, institutional overlap
Every actor here has a stake in the measurement narrative. The Observatory team gains standing and funding relevance by establishing vendor reports as inadequate; Anthropic and OpenAI control which usage data is published and how it is scoped; Anthropic responded on the record while OpenAI and xAI did not, leaving one side unrebutted. The publisher is MIT Technology Review and one co-lead is described as an MIT Media Lab graduate, an affiliation overlap the article does not disclose.
Moderate-low: one outlet, plausible mechanism, unverified artifact
The underlying mechanism — a work-scoped filter removing personal and sensitive traffic — is internally coherent and partly corroborated by OpenAI's own 30% figure, and the named sources are checkable. Confidence is capped by single-publisher coverage, absent study access, no independent replication of the filter reimplementation, and interpretive claims about companionship and safeguards that the article itself hedges.
invest
Scalable Capital puts ChatGPT, Claude and Grok inside the European order ticket2 distinct publishers
product
A school agenda shipped with "Vitoiis" and a planet named Marc, and no one read it first1 distinct publisher
product
Incogni ranks 13 AI assistants by privacy risk: bigger is worse, except ChatGPT1 distinct publisher
product
OpenAI and Anthropic publish eight and thirteen hours of downtime in the same 90 days1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · August 18, 2026