AI observatory finds half of AI chat usage unrelated to work

Companies like Anthropic and OpenAI publish regular usage reports, but the data they release are heavily filtered, presenting only a partial picture. Researchers warn there is no independent source to verify these numbers, and decisions based on the limited information have far-reaching implications.
The project aiming to close the gap
Anka Reuel, a PhD student in computer science at Stanford’s STAIR lab, co-leads the AI Observatory. The public platform aggregates and analyzes real conversations with popular models such as Claude and Gemini, collected with user consent from seven existing datasets. Its goal is to give researchers and policymakers an independent source on how people actually use generative AI.
What companies hide behind filtering
Anthropic’s economic metric, one of the most frequently cited sources, focuses on work and productivity uses and filters out conversations that are not work-related. When the Observatory applied the same methodology to its own dataset, 48% of the conversations were discarded. The filtered conversations contained markedly higher rates of sensitive topics: health and relationships (44.2% versus 31.2%), adult or illegal content (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). OpenAI’s ChatGPT report for 2025 showed a similar picture, with only 30% of consumer usage tied to work.
Changes over time: more companies, less self-disclosure
The WildChat dataset, the largest and most detailed in the study, showed that conversations grew longer and more complex between 2023 and 2025, with more tokens in prompts, more tokens in responses, and more dialogue rounds. At the same time the volume of “small talk” jumped, suggesting increased use for companionship; conversely, the tendency of assistants to disclose they are chatbots declined. Researchers also identified a drop in conversations classified as “sensitive use,” harmful or restricted content, which may indicate improved safety mechanisms.
Not all models equal: Grok and Gemini lead in information retrieval
Usage differed markedly across models in topics, interaction style, conversation structure, and the type and volume of sensitive cases. Grok and Gemini were used more frequently for information retrieval. Grok stood out especially for news and politics, but also aggregated a higher concentration of misinformation, a finding consistent with earlier studies. According to David Widder, professor at the School of Information, University of Texas at Austin, who studies human-AI interaction and is not involved in the project, the holistic view the Observatory provides—rather than separate reports for each topic—helps researchers understand the full spectrum of usage more consistently.