Published · 5d agoProduct3 min read
Half the conversations vanish: an independent check on AI usage reports
A Stanford-affiliated project pooled seven consented chat datasets and applied Anthropic's own filter. Forty-eight percent of conversations were discarded, and the discards were the sensitive ones.
Not a builder's beat, but builders have a standing stake in it.See today for builders

What happened
- AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but AI researchers say they only release the data they want the public to see.
- Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab, says of vendor usage reports: "There is no independent source to corroborate it."
- Reuel is co-lead of a new research project called the AI Observatory, a public platform that aggregated and analyzed real AI conversations with popular models like Claude and Gemini, collected with users' consent through seven existing datasets.
- The AI Observatory's stated intent is to provide independent sources of information for researchers and policymakers to assess how people are using generative AI.
- Reuel says stakeholders are currently making highly consequential decisions about AI's benefits and risks based on very limited data.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
A research project called the AI Observatory, co-led by Stanford Trustworthy AI Research (STAIR) Lab PhD candidate Anka Reuel, has published an independent analysis of real chatbot conversations drawn from seven existing datasets collected with user consent, covering models including Claude and Gemini [3]. It matters because vendor-published usage reports are currently the main evidence base for how people actually use these systems, and, as Reuel puts it, "There is no independent source to corroborate it" [1][2].
The most concrete result is a replication test. The Anthropic Economic Index is one of the most widely cited sources of AI usage data, and by design it covers work- and productivity-related uses of Claude, filtering out conversations that are not [7]. When the Observatory team ran Anthropic's method over its own corpus, 48% of conversations were filtered out [8]. The discarded material was not neutral residue. Non-work conversations that got filtered were more likely to involve health and relationships (44.2%, against 31.2% in Anthropic's analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%) [9][10][11][12]. That is roughly five times the harassment and hate rate and about seven times the sexual content rate [1][2].
OpenAI's own 2025 report points the same direction from the inside: it found that only 30% of consumer use was work-related [13], which leaves 70% outside the frame that productivity-indexed reporting is built to see [4]. Anthropic has published separate posts on companionship use and on attempts to generate CSAM [14]. David Widder of UT-Austin's School of Information, who is not involved in the project, said the value of the Observatory is the bird's eye view rather than findings sectioned off into separate reports [15].
The time series is more interesting than the headline gap. Across datasets spanning 2023 to 2025 [16], conversations in WildChat, one of the largest and most detailed sets in the study, got longer and more elaborate, measured by prompt tokens, response tokens, and conversation turns [17]. Small talk rose significantly, which the researchers read as increasing companionship use, while assistants disclosed that they were chatbots less often [18]. Exchanges labelled sensitive, including sexual harassment and hate speech, declined, which the team suggests may reflect more effective safeguards [19].
Usage also splits sharply by product, in topics, interaction style, conversation structure, and both the likelihood and the type of sensitive use [20]. Grok and Gemini drew more information retrieval; Grok was especially popular for news and politics and was also where misinformation concentrated, consistent with other research, and xAI did not respond to a request for comment [21]. Anthropic's models skewed to coding, Gemini to social and roleplay, ChatGPT to homework [22]. Even model versions diverged: conversations were shorter on GPT-3.5 and longer and more iterative on GPT-4o [23]. The Observatory's stated purpose is to give researchers and policymakers an independent read [4], on the argument that consequential decisions about benefits and risks are being made on very limited data [5].
Watch whether Anthropic or OpenAI publishes the filtered-out share of their own corpora, which is the single number that would make their indexes auditable. Watch also whether the declining sensitive-use trend holds once companionship-heavy products are measured separately, and whether any regulator cites consented third-party datasets rather than vendor blog posts.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but AI researchers say they only release the data they want the public to see.
ReportedView cited source - [2]
Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab, says of vendor usage reports: "There is no independent source to corroborate it."
- [3]
Reuel is co-lead of a new research project called the AI Observatory, a public platform that aggregated and analyzed real AI conversations with popular models like Claude and Gemini, collected with users' consent through seven existing datasets.
ReportedView cited source - [4]
The AI Observatory's stated intent is to provide independent sources of information for researchers and policymakers to assess how people are using generative AI.
ReportedView cited source - [5]
Reuel says stakeholders are currently making highly consequential decisions about AI's benefits and risks based on very limited data.
- [6]
The AI Observatory found that AI use differs significantly across models and has changed over time, and its research shows many more sensitive behaviors than are captured in reports from major AI companies, which the researchers say focus more on work than on personal use.
ReportedView cited source
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- technologyreview.comEileen Guo5d agoWe still don’t know how people are really using AI
- technologyreview.comThomas Macaulay5d agoThe Download: how people really use AI, and Flock’s design choices
Additional citations
- Anka Reuel, Stanford STAIR Lab, quoted by MIT Technology Review
- Anka Reuel
- David Widder, UT-Austin School of Information



