Executive Summary
Almost everything the public knows about how AI actually gets used comes from reports that vendors wrote about their own logs, using samples they selected themselves. The AI Observatory, released in August 2026, is the first large attempt to break that dependence. Researchers at a dozen or so institutions, MIT and Stanford among them, annotated seven public conversation corpora with verified consent grounds under a single taxonomy, then rebuilt the occupational classification gate that Anthropic had published under CC-BY and MIT licences and ran it back over their own data.
The result was half. 47.9% of conversations were classed as non-occupational and dropped out before any task distribution was counted, and the discarded side was much denser in health, relationships, harassment and sexual content than the side that survived. That number should not be read as a finding about Anthropic's data. The paper itself refuses to present it as an estimate of Claude traffic. It is a finding about what one line of preprocessing does.
The uncomfortable part is that this test was only possible because of Anthropic. The prompt and the taxonomy were published, so someone outside could run the same rule again. Anthropic has since changed its measurement method three times. The problem this report is about is not concealment. It is the absence of reproducibility.
Four Numbers This Report Rests On
Four figures carry the argument. The first two show what a single rule removes from a statistic, and how differently that same rule behaves depending on which sample it meets. The second two show that holding the sample fixed is not enough, because the window of time also moves the conclusion. One of those came from the independent corpus, the other from a vendor's own data.
47.9%
Excluded by the work filter
Of 22,956 conversations pooled from the six sources the gate could run on
34.2–61.9%
Same rule, spread across sources
AI Archive lowest, LMSYS-Chat-1M highest
35% → 50%
Personal use, weekday to weekend
The first thing Anthropic saw after switching to continuous sampling
+1,049.5%
Growth in prompt length
Inside WildChat alone, April 2023 to July 2025
47.9% of What, Exactly
Start with what the study is. The AI Observatory: A Public Measure of Real-World AI Use is a preprint with three co-first authors: Shayne Longpre of MIT, Anka Reuel of Stanford and Dayeon Ki of the University of Maryland. The affiliation list runs through MIT, Stanford, Northeastern, Johns Hopkins, Berkeley, Carnegie Mellon, Brown, NYU, Waterloo and Maryland, along with EleutherAI and Cohere. What this group did was not build a model or release a benchmark. They gathered conversations that were already public and annotated all of them against one yardstick. Reuel's account of why is short. Of the usage statistics vendors publish, she says, "There is no independent source to corroborate it."
The haul is seven sources, 23,158 conversations and 85,633 turns. There are widely used public datasets such as WildChat and ShareGPT, Grok conversations crawled from public share links, and NIO data captured through browser instrumentation. Onto that they fixed a single taxonomy of 145 attributes, spanning four levels — prompt, response, turn and conversation — and grouped into nine families that run from language and media format through topic and sensitive use. Ten trained authors refined the scheme across five pilot rounds, and GPT-4.1 performed the actual annotation. How closely that automated labelling tracked human judgement was measured separately: two annotators independently labelled a validation set of 594 turns across 120 conversations, a third adjudicated disagreements, and the pipeline's output was compared against that consensus. Annotating all seven sources cost $5,680.
1.1Borrowing Someone Else's Rule and Running It on Your Own Corpus
All of that is corpus construction. What made the study news is the experiment that followed. The 27 March 2025 edition of Anthropic's Economic Index works in two stages. A binary filter keeps only work-related conversations, and whatever survives is mapped by the Clio pipeline onto O*NET occupational task clusters. Anthropic published the system prompt and the taxonomy for that process. The researchers rebuilt it exactly and ran it back over their own corpus.
The rebuilt pipeline has five stages. ① The is_occupational_task binary filter decides whether a conversation is work-related; only what passes goes on to ② top-level cluster, ③ mid-level cluster and ④ base-level cluster, before ⑤ multi-label assignment of occupational skills. Conversations judged non-occupational stop at stage ①. They never appear in the task distribution at all.
The gate could be run on six of the seven sources. NIO's data use agreement permits only pre-agreed aggregate annotations to leave the enclave, which made executing the pipeline impossible in the first place. That left 22,956 conversations pooled from six sources as the comparison set, and close to half of them fell at the first hurdle.
Source: AI Observatory §3.2 and Appendix D.14. NIO was excluded because its data use agreement did not allow the gate to be applied.
1.2The Line the Paper Draws Around Itself
At exactly this point the paper narrows its own claim with two sentences. The first: "We do not claim it as an estimate for Claude.ai traffic or for representative AI use overall." 47.9% is not an estimate of what actually happens on Claude. The second: "the gate is not a neutral preprocessing step." That one is the real argument. A line that passes quietly under the name of preprocessing is not neutral.
The paper goes further and records, in a footnote, a fact that weakens its own result. Later Anthropic editions do not apply this occupational filter, and still arrive at broadly similar usage distributions. The rule the researchers reproduced is already an old edition. Why the experiment still matters is a question section 6 returns to.
1.3When a Number Travels, Its Definition Stays Behind
Following 47.9% from the paper into press coverage and onward into summaries demonstrates this report's thesis one more time. MIT Technology Review's write-up described the attributes of the filtered conversations and named the comparison group as "Anthropic's analysis." In the paper, the figures that follow are not Anthropic's data at all; they are the conversations that passed the gate inside the same corpus. The share of sexual content is 15.7% in the paper and 16.7% in the article; the harassment baseline is 5.6% in the paper and 5.66% in the article. The same piece cites OpenAI's work-related share as "30%," where the original paper's table gives 27% as of June 2025.
None of this is a complaint about reporters. OpenAI's own paper puts 73% in a table (June 2025) and writes "about 70%" in the prose of its conclusion (as of July 2025). At every link in the citation chain, rounding and reference dates drift a little. Numbers are light and definitions are heavy, so only the numbers make the trip. That is why this report treats 47.9% as a question of what it is a percentage of, rather than what the percentage is.
What Was on the Discarded Side
Losing half the data is not yet a problem in itself. If you set out to produce work statistics, removing conversations that are not about work is the ordinary thing to do. The problem is that what left was not random. Holding the sources and the annotations fixed, the researchers split the corpus with nothing but the gate, then compared the conversations it removed (CLIO-No) against the ones it kept (CLIO-Yes) on the same yardstick. Two groups, same data, same annotations, one rule between them.
Four attributes diverged most sharply. The contrast below is not the Observatory against Anthropic; it is the removed conversations against the retained ones inside the same corpus.
The differences were tested at the parent-category level with Fisher's exact test. When you are checking well over a hundred attributes, a few will look significant by chance alone, so the false discovery rate was controlled with the Benjamini–Hochberg procedure and only the survivors were counted. The four below cleared that bar.
Source: AI Observatory §3.2, redrawn from the right panel of Figure 4. Only attributes that survived FDR correction are shown.
Health and relationships run at 44.2% on the removed side and 31.2% on the retained side, so that category is widespread in both groups. The other three behave differently. Harassment and hate stand at 27.5% against 5.6%, nearly a fivefold gap, and sexual content at 15.7% against 2.4%, more than sixfold. The conversations that safety policy and risk assessment most need to see are the ones a work frame cuts away first. The paper is careful to call this contrast a property of the gate as applied to its own sources, but even inside that qualification the implication is plain.
2.1What a Complete-Looking Taxonomy Leaves Out
One sentence the paper attaches to this result is the heart of the section. A framework, it writes, "can look comprehensive while still systematically filtering out uses that are socially important, policy-relevant, or disproportionately tied to harm." The trouble is not the omission itself but its invisibility. The O*NET occupational task taxonomy looks complete from the inside. What is missing cannot be seen from within it.
David Widder of the University of Texas at Austin, who was not involved in the research, frames the issue as one of placement. Having a bird's-eye-view analysis, he says, helps researchers understand different uses more consistently than leaving that information "sectioned off into a separate report." His question is blunter still: "When we want to ask, for example: is Anthropic's general-purpose AI system … used mostly for good or mostly for bad … we don't have a way of answering that question because that information is proprietary."
A work filter makes a statistic clean. But that cleanliness is defined by what was thrown away, not by what remains. Reading only what a report counted tells you nothing about which world it is describing. You learn that only when you can also see what it discarded before counting began.
Why One Sample Never Generalises
47.9% is an average pooled from six sources. Broken out by source, the spread is much wider. The same gate discarded 34.2% in AI Archive and 61.9% in LMSYS-Chat-1M. Not one line of the rule changed. Only the data it met changed. The paper concedes that with opt-in sources it cannot identify which property produced the spread, and argues that the size of the spread is itself the point.
Lining up the seven sources shows why. They are not the same kind of data. How they were collected, when, and from whom all differ.
| Source | Annotated conversations | Turns | Period | Collection method |
|---|---|---|---|---|
| WildChat | 14,648 | 50,152 | 2023-04 to 2025-07 | Free API proxy service, explicit consent in the terms of use |
| AI Archive | 2,076 | 15,955 | 2023–2025 | Archive of conversations users chose to share publicly |
| LMSYS-Chat-1M | 2,000 | 3,900 | 2023-04 to 08 | Visitors to the Vicuna demo and arena |
| Chatbot Arena | 2,000 | 2,526 | 2023–2024 | Head-to-head model preference votes |
| Grok | 1,234 | 3,750 | 2025 snapshot | Crawl of share links users posted publicly |
| ShareGPT | 1,000 | 8,438 | 2023 | Shared exports via a browser extension |
| NIO | 200 | 912 | 2024-06 to 2025-04 | Browser instrumentation, IRB-equivalent review plus a data use agreement |
Source: AI Observatory, Table 1 (preprint edition). WildChat's 14,648 annotated conversations are a stratified sample drawn from roughly 4.8 million, not a census.
The structural differences the table does not show are just as large. Grok conversations average 1,322.9 tokens per response, while the other sources sit between 51.8 and 181.0. That is not a matter of a few percent. ShareGPT averages 4.42 turns per conversation and AI Archive 4.34, so people go back and forth repeatedly, whereas Arena conversations, collected for preference voting, end after a turn or two. The collection surface sets the shape of the conversation before anything else does.
Shape is not the only thing that differs. What people use the tool for splits by source too. Information seeking accounts for 67.1% of Grok conversations but only 26.2% of WildChat. ShareGPT is concentrated in content generation at 63.6%, and AI Archive in information analysis at 53.9%. Which source you pick answers the question "what do people mainly use AI for?" before you have started.
For anyone trying to measure risk, there is a harder passage. Academic-policy issues such as homework outsourcing appear at 23.1% in WildChat and 40.4% in AI Archive. Misinformation concentrates in Grok at 15.3%, against 4.2–10.2% elsewhere. The rarer the category — sexual content, harassment, privacy, copyright — the wider the relative gap between sources. The paper's conclusion is matter-of-fact. Measure the size of a specific risk from any single source and the value will be wrong.
3.1On the Same Platform, Swapping the Model Changes the Conversation
If the differences appeared only between sources, this would end as the familiar observation that datasets differ. The researchers held source and period fixed and varied only the model. Within Chatbot Arena, Gemini's responses run 81.0% longer than other models and Grok's 78.0% longer, while Claude's are 59.8% shorter. In NIO, Gemini conversations are more than twice the length of ChatGPT ones. Same platform, same period, same user pool.
Swapping the model inside the same product changes the texture of the conversation as well. The researchers report that conversations from the era when ChatGPT ran on GPT-3.5 were shorter, and that after the move to GPT-4o they became longer and more multi-turn. Concentrations by developer are clear too. People reached for Claude on coding, Gemini on social and role-play, and ChatGPT on homework. Vendor reports generally do not count these differences among their own models separately.
The conclusion the paper draws here is that "platform interface and user composition shape how AI is used as much as the model itself." From that follows a warning: evaluations that measure only model behaviour on benchmarks can systematically misestimate the risks that actually arise in deployment.
3.2Why Three Companies' Numbers Do Not Contradict Each Other
The point sharpens when vendor reports are laid side by side. In July 2025 Microsoft published a study mapping 200,000 Bing Copilot conversations onto O*NET work activities, using the same classification axis as Anthropic but a different product and sample. Separately, a study of 41,735 Grok interactions on X found roles that private one-to-one studies never see. On a public timeline, Grok gets summoned as arbiter of truth, as advocate and as adversary. The researchers inductively derived ten roles, and the most common one is still information provision; the public setting simply layered dispute resolution on top of it. The operational figures that study reports are of a kind private-conversation research cannot produce either. Grok answered only 62% of the requests that summoned it, and half of the answers it did give had not passed twenty views after 48 hours. Change the deployment surface and the same model acquires a different purpose.
There are also cases where the taxonomy manufactures the conclusion. In OpenAI's usage analysis, relational and emotional conversations account for just 1.9% of all messages. Yet an external estimate cited in the same paper ranks "therapy and companionship" as the most common use of generative AI. Coding tells a similar story. Computer programming is 4.2% of consumer ChatGPT conversations, while coding's share of Claude work conversations is incomparably higher. Neither side is lying. The denominators, the taxonomies and the user bases are different.
How the categories are cut changes the picture too. In the same OpenAI analysis, three categories — practical guidance, seeking information and writing — account for nearly 80% of all conversations. Make the boxes big and most use falls into a handful of them; make them fine-grained and the same use scatters across many, so none looks large. There is little to gain, then, from setting two companies' tables against each other and asking which is right. Shayne Longpre, who co-led the research, puts it plainly: "No single company report tells the whole story."
The same pattern of sample composition producing the conclusion appeared in our report on the geography of open-weight model use. That piece watched it happen in a corpus of research papers. This one watches it happen in vendors' reports about themselves.
A Different Window Makes It a Different Dataset
That changing the sample changes the conclusion is familiar enough. The study goes one step further. Hold the sample fixed, change the window of time, and the conclusion still changes. The researchers picked WildChat, the only one of the seven sources with multi-year timestamps, and measured the shift from April 2023 to July 2025. The dataset's name stayed the same. What was inside it did not.
Six things moved significantly across those two years and three months. The top three concern how much bigger conversations got. The bottom three concern how people and models talk to each other.
Source: AI Observatory §3.3. All changes shown survived FDR correction. Percentages are changes in means; pp denotes changes in prevalence.
Prompt length grew more than tenfold. Responses doubled, and turns per conversation rose. Phatic expressions, the small social pleasantries that open and close a message, increased by 21.5 percentage points, while self-disclosure of the "I am an AI assistant" kind fell by 12.0. The share of responses containing formatted lists rose 25.8 points. User habits and model register moved together. What people talked about moved too: adult and illicit topics fell by 11.2 points and academic-policy uses such as homework outsourcing by 13.9, while the rate of pasting code into prompts rose 15.2 points. In the paper's phrase, "a static WildChat sample therefore mixes multiple temporal regimes." Put 2023 conversations and 2025 conversations in one bucket and average them, and the average describes neither year.
4.1The Total Grew While Individuals Narrowed
Here a counterintuitive finding appears. Isolate returning users and compare their first month against their last, and individuals do not use the tool more. Turns per conversation compress, the number of distinct functions they use falls from four to three, and sensitive-use flags drop from two to one. Prompt length and topic distribution show no significant change. In other words, the corpus swelled not because individual use deepened but because the user population changed. The paper's words are "user-composition shifts," not "within-user intensification."
What a shift in composition means becomes visible when users are grouped. The researchers clustered returning users by topic distribution into 24 personas, then consolidated those into six higher-level groups. In the fantasy, fiction and entertainment group, role-play runs 10.4 percentage points above baseline, sexual content 29.5 points above, and jailbreak attempts 20.0 points above. The technically oriented learner group shows responses containing code 27.4 points more often, and privacy-implicating turns 14.8 points more often. These are people using the same service, and the grain of their use runs in separate directions.
Source: AI Observatory §3.4, Figures 7-8, reinterpreted. Individual figures for the other four groups (Reflective Learners, Culturally Engaged Communicators, STEM & Engineering Education, Business/Tech/Communication) are not cited in the body and are omitted here.
This gives another reason to doubt that what the gate in section 2 removed was scattered small talk. Sensitive use is not sprinkled evenly across the corpus; it clumps in particular user groups. Filter such use once on the criterion of work, and what drops out is not a handful of conversations but a cohort of users. That said, these clusters were observed within a single source while the gate experiment pooled six, so the two cannot be joined into a causal claim.
Note that this analysis was possible only in WildChat, because it was the only source with both stable user identifiers and enough temporal history. The persona clustering was confined to one source for the same reason. What bounded the analysis was not the research design but the presence or absence of metadata. Section 5 picks that up.
4.2The Same Thing Showed Up in a Vendor's Own Data
If the timing problem belonged only to independent corpora, this section would end here. But on 26 June 2026 Anthropic published its sixth Economic Index report. It is titled "Cadences," and half of it is an account of changing the measurement method. Announcing a move to continuous telemetry sampled a little each day, it wrote: "in contrast to the seven-day samples each previous Economic Index report drew on." Read the other way round, every report until then had been a one-week snapshot.
The first thing continuous observation revealed was a day-of-week effect. Personal-purpose conversations sit around 35% on weekdays and jump to just under 50% at weekends.
Source: Anthropic Economic Index, sixth report, "Cadences" (26 June 2026). The day-of-week effect first observed after the shift to continuous telemetry.
Which days a seven-day window happens to cover, in other words, can move the headline number by more than ten percentage points. All five earlier reports were photographs taken through that window. Something similar happened with the fifth report. The March 2026 edition reported academic-purpose conversations falling from 19% to 12% and personal use rising from 35% to 42%, and Anthropic itself attributed much of the movement to winter school holidays and a rise in new sign-ups in February. Usage patterns had not changed so much as the calendar had turned.
What matters is that the Observatory's critique and the vendor's own revision arrived at the same conclusion. In the independent corpus, prompt length changed more than tenfold across two years and three months. In the vendor's own data, weekday and weekend split the personal-use share by close to fifteen points. The argument is not about who is right, but that the rules and the windows keep moving. If they do, then the date of collection and the rule used must travel with the number.
What Actually Blocks an Independent Corpus
Ask why independent measurement took this long and the answer usually starts with money. Yet the total annotation cost of this study was $5,680. WildChat was the largest line item at $3,904 and the rest ran to a few hundred dollars each, with another $182 for automated annotation of the validation set. That is less than a mid-sized organisation's monthly cloud bill. Cost was not what had been holding independent measurement back.
What actually stopped the pipeline was rights. NIO's data use agreement allows only pre-agreed aggregate annotations to leave the enclave, so reproducing the gate there was impossible and six of seven sources became the comparison set. Chatbot Arena licenses prompts under CC-BY 4.0 but responses under CC-BY-NC 4.0, so the non-commercial restriction propagated into derived annotations on the response side. LMSYS had to be handled in line with its deletion-request provisions. Licences attach not only to the original data but to whatever is layered on top of it.
The form of consent differed by source as well. WildChat obtained explicit permission through the terms of use of a free API proxy service, while AI Archive rests on the fact that a user pressed "share publicly." The Grok data was crawled from conversations users posted themselves. That last item has a shadow over it. In August 2025 thousands of shared Grok conversations turned up indexed by Google search, an incident the paper cites in its own bibliography. It is a case where the scope of consented disclosure and the scope of actual disclosure came apart. An open licence does not mean the user pictured that situation. The way data licences propagate into derivatives, and what that costs, is something we examined separately in our report on the price of auditing open dataset licences.
5.1What You Fail to Record Now Cannot Be Recovered Later
If practitioners take one thing from this report, it is probably this. The temporal analysis in section 4 was possible in only one of seven sources not because of research capacity but because of missing records. There were no timestamps, or no model identifiers, or no stable identifier to stitch the same person's sessions together. The persona analysis was confined for the same reason. Anything not recorded at collection time cannot be reconstructed afterwards by any method.
Any organisation accumulating internal AI usage logs would do well to check the fields below now. They are the paper's appendix tables on licensing and consent, translated into operational terms.
| Field to record | Analysis you lose without it |
|---|---|
| Timestamp | Adjustment for time. Averages that blend several eras get read as current values |
| Model identifier and version | Separating the effect of a model swap from a change in user behaviour |
| Stable user identifier | Telling deepening individual use apart from a shift in user composition |
| Language | Checking bias by region and language community |
| Consent basis | Determining after the fact what may be re-analysed or shared externally |
| Redistribution rights | Setting distribution terms for annotations and derivatives, since restrictions propagate |
| Annotation model and version | Reproducing labels. The same prompt yields different labels once the model changes |
Adapted from the per-source consent and licensing conditions in Appendices C and E of the AI Observatory, recast for internal log design.
The problem of tracing the lineage of training data is one we have already covered in our report on provenance in reinforcement learning reward data. What is new in this case is that the same problem repeats identically in usage data rather than training data. It is not only what goes into a model that needs a lineage. The records left behind by using one need the same precision.
So How Should These Numbers Be Read
Reading this far and concluding "so trust the Observatory instead of vendor statistics" would invert the argument exactly. The Observatory is not a new authority. It is a second observation point that can be compared against the first, and it has plenty of flaws of its own. The paper lists them itself, so start there.
First, annotation accuracy is not complete. Against human consensus, the median parent-category F1 across the nine families is 0.856, and at the fine-grained prompt-function level it falls to 0.648. Every comparison in the body of the paper is therefore made at parent-category level only. Second, the annotation is not deterministic. In the paper's words, "Neither temperature 0 nor a fixed seed guarantees determinism through the API." The least stable family happens to be sensitive use, which means the figures in section 2 of this report should be read as approximations. Third, the validation set was built by the same people who designed the taxonomy. There was no blind external annotation, and the paper says of this that it is evidence the pipeline was faithful to their annotation scheme, not evidence that the scheme itself is valid.
Coverage is narrow as well. Because every source is opt-in or based on public sharing, conversations people would rather not disclose are structurally underrepresented. Private deployments, on-device assistants and services widely used across the Global South are largely out of scope. The corpus keeps growing, too: the paper's PDF says 23,158 conversations, while the project website and the NeurIPS 2026 camera-ready say 24,521, with almost all of the increase coming from a second crawl of the Grok source in April 2026. Usefully, the 47.9% this report is about is identical in both editions.
6.1The Target Has Already Moved Twice
There is also no reason to hide the fact that the rule the Observatory reproduced is an old edition. Anthropic has changed its measurement method several times since, and disclosed each change.
| Edition | Date | Measurement method |
|---|---|---|
| 1st | 2025-03-27 | Occupational gate → Clio → O*NET task clusters. The edition the Observatory reproduced |
| 4th, "Economic primitives" | 2026-01-15 | Five economic primitives in place of the filter. Claude use split three ways: 46% work, 19% study, 35% personal. Sample of 1M consumer plus 1M API conversations |
| 5th, "Learning curves" | 2026-03-24 | Study 19%→12%, personal 35%→42%, explained by Anthropic itself as winter holidays plus a rise in new sign-ups |
| 6th, "Cadences" | 2026-06-26 | Seven-day samples → continuous telemetry, alongside its own survey (linked sample of roughly 9,700 respondents) |
Source: the text of each Anthropic Economic Index report. Anthropic states that the sixth report's survey is "not representative of the general population"; about 30% of respondents work in computing and mathematics.
In fifteen months the filter became a set of primitives, the snapshot became continuous observation, and a survey was added to the logs. The target the Observatory aimed at has already shifted twice. That does not weaken this report. Rules changing this often is precisely the argument for being able to replay them from outside. The situation this piece objects to is exactly the one where values measured on a yardstick that keeps changing get quoted side by side as though they were a time series.
6.2Publication Is What Made the Test Possible
And here is the turn in the story. The reason the researchers reproduced Anthropic's rule of all rules is not that Anthropic is unusually untrustworthy. It is that Anthropic's was the only rule that could be reproduced. The company released the data behind its March 2025 edition under CC-BY 4.0 and the code under an MIT licence. From that moment the number stopped being marketing material and became something testable. Statistics from a company that never published its rules cannot even be criticised this way. The possibility of criticism is a by-product of openness.
An Anthropic representative's reply in the coverage of the Observatory belongs here too. The company's published research, they said, reflects its research teams' specific questions and interests, and it is important to support external independent research. In the same article, OpenAI did not respond to requests for comment. What a company says matters less than what can be replayed from outside, and that shows up here as well.
So the conclusion cannot be to wait for better-behaved vendors. Vendor logs are unmatched in scale and coverage, and a 23,000-conversation independent corpus is no substitute for them. What is needed is not a replacement but a public measurement layer where the same rule can be run against different data: vendors publishing their classification prompts and taxonomies, independent researchers applying those to their own corpora, and a structure in which the gap between the two results can be interrogated. What the Observatory did was the first execution of exactly that structure.
The next time you meet a sentence of the form "X% of AI users do Y," four questions are worth asking. What is the denominator? Consumers only, or enterprise API traffic as well. What was discarded before counting? Which conversations never made it into the tally. When was it collected? How many days, and which days of the week. And can the same rule be run against different data? A number that cannot answer the fourth question has not yet been verified by anyone.
Editor's Note
A word on why Pebblous spent so long with this study. The premise of our work is that diagnosing data requires first knowing how the data was gathered. This research extends that premise from training data to usage data, at scale, for the first time. 47.9% turned out to be a property not of a model but of one line of preprocessing. It is also the cleanest demonstration we have seen of the claim that the selection rule is the conclusion.
Understand data quality as a checklist of consistency tests and this case stays invisible. Run the same gate over six sources and the rate spreads from 34.2% to 61.9%: hold the rule fixed, change the sample, and the conclusion changes. Run it inside WildChat alone and prompt length shifts more than tenfold with time, while in Anthropic's own data weekday and weekend split personal use by close to fifteen points: hold the sample fixed, change the window, and the conclusion changes again. Documenting those two axes, composition and timing, and making them reproducible is much closer to what we mean by data quality.
In practice it gets more concrete. Every organisation building an internal AI usage dashboard arrives at the same fork: which conversations count as work? This report puts a measured figure on the error that choice introduces. The table in section 5 can be used as it stands as a checklist of what to record now so that logs remain analysable later. The reason the Observatory could run its temporal analysis in only one of seven sources was the absence of exactly those fields.
Finally, the lesson of this episode is not that Anthropic behaved badly. It is that Anthropic published its rules, and that made verification possible. The unit of trust is not good intentions but a reproducible pipeline. That, we think, is why an independent layer for data diagnosis and provenance needs to exist.
Pebblous Data Communication Team
21 August 2026
References
Every figure in this report was checked directly against a primary source. For the AI Observatory that means the body and appendices of the paper itself; for the Anthropic Economic Index and OpenAI's usage analysis, the reports and papers each organisation published. What follows groups those primary sources, the academic literature the original paper cites as the basis of its method, and adjacent Pebblous pieces on the same theme.
Primary Sources
- 1.Longpre, S., Reuel, A., Ki, D., et al. (2026). The AI Observatory: A Public Measure of Real-World AI Use. Preprint. Project platform: ai-observatory.org
- 2.Anthropic. (2026-01-15). Anthropic Economic Index Report: Economic Primitives.
- 3.Anthropic. (2026-03-24). Anthropic Economic Index Report: Learning Curves.
- 4.Anthropic. (2026-06-26). Anthropic Economic Index Report: Cadences. The move from seven-day samples to continuous telemetry, and the day-of-week effect.
Measurement Methodology (Academic)
- 5.Tamkin, A., McCain, M., Handa, K., et al. (2024). Clio: Privacy-Preserving Insights into Real-World AI Use. arXiv:2412.13678.
- 6.Handa, K., Tamkin, A., McCain, M., et al. (2025). Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations. arXiv:2503.04761. The original edition of the gate the Observatory reproduced.
- 7.Chatterji, A., Cunningham, T., Deming, D., Hitzig, Z., Ong, C., Shan, C., & Wadman, K. (2025). How People Use ChatGPT. NBER Working Paper No. 34255.
- 8.Tomlinson, K., Jaffe, S., Wang, W., Counts, S., & Suri, S. (2025). Working with AI: Measuring the Applicability of Generative AI to Occupations. arXiv:2507.07935. O*NET mapping of 200,000 Bing Copilot conversations.
- 9.Mei, Y., Wolfe, R., Weber, I., & Saveski, M. (2026). Grok in the Wild: Characterizing the Roles and Uses of Large Language Models on Social Media. arXiv:2602.11286.
- 10.Benjamini, Y., & Hochberg, Y. (1995). Controlling the False Discovery Rate. Journal of the Royal Statistical Society B, 57(1), 289–300.
Corpora and Data Governance (Academic)
- 11.Zhao, W., Ren, X., Hessel, J., Cardie, C., Choi, Y., & Deng, Y. (2024). WildChat: 1M ChatGPT Interaction Logs in the Wild. ICLR 2024.
- 12.Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2024). LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset. ICLR 2024.
- 13.Chiang, W.-L., Zheng, L., Sheng, Y., et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv:2403.04132.
- 14.Longpre, S., Mahari, R., Chen, A., et al. (2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv:2310.16787.
- 15.Mireshghallah, N., Antoniak, M., More, Y., Choi, Y., & Farnadi, G. (2024). Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild. arXiv:2407.11438.
- 16.National Center for O*NET Development. O*NET OnLine. CC BY 4.0. The occupational task taxonomy used in the Economic Index reproduction.
Press
- 17.Guo, E. (2026-08-18). We Still Don't Know How People Are Really Using AI. MIT Technology Review.
- 18.Bellan, R. (2025-08-20). Thousands of Grok Chats Are Now Searchable on Google. TechCrunch. The consent-scope mismatch cited in the paper's bibliography.
Adjacent Pebblous Reports
- 19.Pebblous. Geography, Not Openness, Decided Which AI Models Science Runs On. Another case of sample composition producing the conclusion.
- 20.Pebblous. The Real Invoice for Free Data. How licence restrictions propagate into derivatives.
- 21.Pebblous. Nobody Can Trace, Atom by Atom, the Data That Reward-Verified RL Learns From. The same problem on the training-data side.
- 22.Pebblous. The Tools Are Named, the Sources Are Sealed.