Executive Summary
This article rechecks a follow-up study by the marketing analytics firm Graphite, which ran its original measurements again on Claude Opus 5.5, against the primary sources and the raw data the firm published. The most quoted line says em dash use fell by 99%, from 2.92 to 0.015 per 1,000 words. The number the same study reports a few lines later is 2,548: the count of "AI tells," words and phrases and frames that the study flags, and only 4% below the previous version. One looks like a cliff and the other looks almost flat, but the two numbers were never counting the same thing. The first is the frequency of a single feature. The second is the size of a set fixed by a threshold.
Open the raw data and one more number appears that the study text never prints. In the same corpus, human writing carries 3.14 em dashes per 1,000 words. That puts Opus 5.5 at roughly 1/205 of the human rate. All four current models sit below the human line without exception. In the same table, GPT-4.1 and GPT-5 sit above twice it. When readers were learning to treat the em dash as an AI marker, the folk rule was actually correct, and today it points the other way. Everything up to here is written in the study and its raw data.
The reading we draw is this. Detection compares two distributions. In this study the machine-side distribution moves wholesale with every model version. The human-side distribution is pinned to 10,000 web articles published before ChatGPT, and there is no way to refresh it. The web after that date already mixes both kinds of writing, so any attempt to redraw the baseline contaminates the baseline. What deserves more attention is that the freeze is not peculiar to this one study. Open the evaluations that show detectors holding up across model generations, and the human writing they compare against is, every time, from before 2022.
If one side moves, the other side sits still, and the still side cannot be fixed, then the labels "human-written" and "AI-written" produced between them acquire a shelf life as a matter of structure. And for data quality, a signal flipping sign is worse than a signal dying. A dead signal shows up as falling accuracy. An inverted signal keeps producing wrong labels quietly and with high confidence. This article is not a guide to spotting AI text. It is about how to store the labels that have already been applied.
1/205
Opus 5.5 em dash rate, against human writing
0.015 per 1,000 words. In the same corpus people use 3.14
2.27x
GPT-4.1 em dash rate, against human writing
When the folk rule formed, the models above the human line. Same corpus
2,548
AI tells counted in Opus 5.5
4% below the previous version. The size of a set fixed by a threshold
6% lower
Mean strength of 11 well-known tells, against Opus 5
Mean strength by Cohen's d. Against Opus 4 it is 53% lower
When the folk rule formed, GPT used twice the human rate
Graphite is a content marketing growth agency. Its September 2026 study, "AI Tells," set 10,000 web articles published before ChatGPT against 90,000 AI articles written to the same topics, then counted which words and phrases ran disproportionately high on either side. The follow-up published on 1 October 2026 added Claude Opus 5.5 and GPT-6 Astra and applied the same yardstick again. This is not a peer-reviewed paper. It is research a company posted on its own blog. The method and the limitations are stated in the text, though, and the per-feature numbers were released as raw data under CC BY 4.0. Every figure in this article was read out of those files directly.
The most widely repeated line concerns the em dash, the long '—' that readers have been trained to watch for. The follow-up reports that Opus 5.5 went from 2.92 to 0.015 of them per 1,000 words. What the text never states is the comparison point. In the same measurement, human writing uses 3.137847 per 1,000 words. That value is sitting in the raw data. Set it as the reference, line up all eleven rows in one table, and a different picture appears than the quoted sentence suggests.
| Model | Em dashes / 1,000 words | vs. human (3.14) |
|---|---|---|
| GPT-4.1 | 7.131936 | 2.27x |
| GPT-5 | 6.921365 | 2.21x |
| Human (published before 2022-11-30) | 3.137847 | reference |
| Claude Opus 5 | 2.924242 | 0.93x |
| Gemini 2.5 Pro | 1.430118 | 1/2.2 |
| Claude Opus 4 | 0.661113 | 1/4.7 |
| GPT-6 Astra | 0.368904 | 1/8.5 |
| GPT-5.6 Sol | 0.305514 | 1/10.3 |
| Claude Opus 4.6 | 0.202808 | 1/15.5 |
| Claude Opus 5.5 | 0.015347 | 1/204.5 |
| Gemini 3.1 Pro | 0.013723 | 1/228.7 |
Source: the layer1__punct_emdash_per1k row of ai_tells_features.csv.gz, published by Graphite. Based on 9,974 aligned topics. The same file carries a second em dash feature under a different definition, but this row is the one that matches the 2.92 printed in the study text. Split the table at the human row and two models stand above it, eight below.
The top two rows are where this article starts. The models in circulation while readers were learning to treat the em dash as an AI marker were GPT-4.1 and GPT-5, and both really did use it at more than twice the human rate. The folk rule was true when people learned it. All four models in current use now sit below the human line: Claude Opus 5.5, Gemini 3.1 Pro, GPT-6 Astra and GPT-5.6 Sol. The signal has not faded. It has reversed direction. Graphite itself raised the possibility of overcorrection one version earlier, though the models it named then were GPT and Gemini, while Claude Opus 5 was sitting at roughly the human rate. One version later, Claude dropped below both of them.
1.1Four conditions hold this number up
Lift the table out on its own and it turns false quickly. Four conditions travel with it. First, the human writing it compares against is web articles published before 30 November 2022. It is not writing produced by people in 2026. Second, the AI writing is SEO-style web articles generated from a single prompt. Third, the model versions are fixed. Change the names in the table and the values change with them.
The fourth condition is the one that goes missing most often. These are values measured in this corpus. Readers mostly learned the em dash heuristic from chat windows, and this study did not measure conversations. Graphite attaches the same caveat in its limitations section: tells may differ in other contexts, such as conversations with AI models. So the table cannot be restated as "chatbots have stopped using em dashes." How large that gap gets shows up numerically in the next subsection.
1.2Three teams measured the same model and the baselines differed fivefold
The collapse in Opus 5.5's em dash use has been reproduced independently. All three measurements agree on direction. The absolute value for the previous version, Opus 5, varies sharply by who did the measuring. The three measurements below asked the models to write different kinds of text.
| Measured by | What the model was asked to write | Opus 5 | Opus 5.5 |
|---|---|---|---|
| Graphite | Single fixed prompt, SEO-style web articles | 2.92 | 0.015 |
| Arena | Text Arena high-reasoning production responses | 15.2 | 0.8 |
| Arize | 20 research briefs, 5 genres x 5 domains | 12.9 | 0.05 |
| (Human baseline, Graphite) | Web articles published before 2022-11-30 | 3.14 | · |
All units are em dashes per 1,000 words. The Arena figures come from a tally reported by BleepingComputer on 26 September 2026, and the Arize figures from that firm's own reproduction the same month. The three measurements do not contradict each other. They differ in what the models were asked to write.
In Graphite's corpus, Opus 5 used em dashes less often than people did. The same model used them four to five times the human rate in Arena's and Arize's corpora. The reversal in the Section 1 table still holds, then, but what reversed is not "the model across all contexts." It is the model as measured in this corpus. The independent reproductions strengthen the article rather than weakening it. Same model, same feature, and a fivefold swing depending on what it was compared against: that is the shortest available illustration of the argument this article reaches in Section 6.
The Arena measurement adds one more thing. It counted semicolons alongside em dashes, and over the same interval the semicolon rate fell from 6.10 to 1.64. The visible punctuation signals were suppressed together, which supports the overcorrection reading Graphite introduced. It should not be read as across-the-board improvement, though. In the same tally, mean response length rose from 453 to 481 words. Arize's closing sentence is neutral and accurate: "So the claim holds up halfway. 'Fixed' is too strong. 'Much better' is fair."
The 99% and the 4% count different things
Two numbers out of the same study circulate side by side. One is 99% and the other is 4%. One says almost everything is gone; the other says almost nothing has changed. Put them on the same line and the study looks self-contradictory. The two numbers were never counting the same object.
The 99% is a drop in the frequency of one feature. Count how often a single trait, the em dash, appears per 1,000 words and divide by the previous version: 0.015347 over 2.924242 gives a 99.5% decrease. The 4% is a change in the size of a set. After length normalization, Graphite collects the words and phrases and frames that are used at least twice the human rate and that clear a minimum article count, and calls that collection the model's tells. The minimums are 500 articles for single words, 250 for two-word phrases, 100 for three-word phrases, and 120 for frames. The list assembled that way went from 2,666 for Opus 5 to 2,548 for Opus 5.5. Move the threshold slightly and the number moves with it.
So the two numbers cannot be divided into a ratio. Ninety-nine over four points at nothing: the numerator and the denominator live in different worlds. The figure below lays those two worlds over a shared horizontal axis. The top series is the em dash rate across four Claude versions; the bottom series is the tell count for the same four. The dashed line across the top is the human baseline of 3.14.
The two vertical axes use different units, so the heights should not be compared against each other. What to read is the shape of the slope. The top series is a cliff: it climbs near the human baseline at Opus 5 and drops to the floor at the next version. The bottom series is a gentle incline across all four.
2.1A third number sits between the two
There is one more number to read between the 99% and the 4%. Like the original study, the follow-up tracks a separate set of 11 well-known tells: the em dash, words such as delve, and other habits readers had already adopted as AI markers. Measured as mean Cohen's d, the average strength of those eleven is 6% lower in Opus 5.5 than in Opus 5, and 53% lower than in Opus 4. The itemized list of the eleven has not been published.
That 6% is one half of the sharpest contrast in this article. A single tell lost 99% of its frequency, while the whole bundle it belongs to fell by six percent across the same version step. The most famous signal dropped dramatically and the rest barely moved. If work was actually done to erase well-known tells, it seems to have aimed at the most conspicuous item rather than at the bundle.
2.2GPT is moving the other way
The sentence "AI writing is converging on human writing" turns false the moment its subject is "AI," because the model families point in opposite directions. Claude has closed the distance to human writing; GPT has widened it. Tallying the raw data directly, the GPT-side tell total rose 47.8%, from 2,491 to 3,681. Over the same span the Claude side fell from 2,666 to 2,548. Union the published tell lists across all nine models and the count is 12,877; restrict it to the three current-generation models and it is roughly 7,000.
Any claim about resemblance between model families has to carry the study's own reservation with it. Every model received the same prompt, so part of the resemblance may come from the task rather than from the models. Strip that caveat and shorten the finding to "all the models resemble each other," and the study is being made to say something it did not say.
2.3The reversal is not confined to the em dash
If the em dash alone had flipped, it could be set aside as a coincidence. So we put the same question to the vocabulary side. Taking eight words commonly named as LLM vocabulary, including the two the original study cites as examples of well-known tells, we summed their rates straight out of the raw data: delve, tapestry, nuance, crucial, landscape, foster, stakeholder, realm. This sum is not a metric Graphite publishes. It is a tally this article built from the open raw data.
| Model | Eight words summed / 1,000 words | Cohen's d vs. human |
|---|---|---|
| Human | 0.439665 | reference |
| GPT-4.1 / Opus 4 / Gemini 2.5 Pro | around 2.62 | +1.04 to +1.13 (well above human) |
| Claude Opus 5 | 0.443966 | +0.004 (essentially level) |
| GPT-5.6 Sol / GPT-6 Astra / Claude Opus 5.5 | around 0.31 | −0.111 / −0.115 / −0.124 (below human) |
Per-1,000-word rates for the eight words, summed from the same raw data and set against the human value. This is not a metric Graphite published. It is this article's tally.
The result has the same shape as the em dash. A widely known signal dropped below the human rate within one model generation of becoming widely known, and the vocabulary side did exactly what the punctuation side did. So the reversal in Section 1 is not an accident that befell one mark. A rule people apply by eye ages faster the better known it becomes, and the way it ages is not by fading. It is by turning around.
'This matters' moved into the vacancy
A tell count falling from 2,666 to 2,548 says only that the list got slightly shorter. It does not say that 118 entries dropped out and the remainder stayed put. This is where the follow-up study gets interesting. Other items moved into the vacated slots, and the newcomers have enough family resemblance that Graphite gave them category names. The most prominent category is the one it calls "Flagging Importance": expressions that tell the reader what they are reading is important. The second is "Contrastive Phrasing," the frame that dismisses one thing in order to raise another, of which the highest-scoring member, "is more than a _ it," reaches 98 times the human rate.
| Category | Phrase | vs. human |
|---|---|---|
| Flagging Importance telling the reader it is important | this matters | 116x |
| why _ matters | 92x | |
| just as important | 13x | |
| matters most | 12x | |
| Contrastive Phrasing setting one thing against another | is more than a _ it | 98x |
| rather than simply | 32x | |
| not only about | 13x | |
| instead it | 8x | |
| Others | looking ahead the | 40x |
| adds another layer | 27x | |
| what comes next | 24x | |
| dependable | 23x | |
| in practice | 7x |
Usage frequency as a multiple of the human rate. Some coverage reported dependable as Opus 5.5's top tell, which the primary source does not say. That word is one of four examples under the "Evaluative Adjectives" category and sits at 23x. The highest multiple for this model is this matters at 116x. Worth adding: in the original study, dependable was a tell for Astra, at 59x.
Phrases like these are what filled the space the em dash vacated. The em dash is a single character, easy to count and easy to strike. "This matters" is a two-word frame, and removing it means changing how the model closes a sentence. The signal moved down to a less conspicuous layer, which also means it got harder for a person to filter by eye.
3.1More than half the list turns over between versions
Druck's remark holds up numerically. Counting how much the tell lists of two consecutive versions overlap in the open raw data, the shared fraction runs between 18.7% and 45.5% depending on the version step. In every step it falls well short of half. The remainder appears on one side only. That is the turnover between two consecutive versions from the same company.
At the family level the picture shifts a little. Sixty-five percent of tells are unique to a single model family, and Astra and Sol, both from the same company, share 45%. Something like a fingerprint survives in how a company builds its models, then. But even that fingerprint is more than half redrawn at each version change.
What changes is clearest at the level of individual phrases. The phrases Opus 5.5 uses more than its predecessor are "can help you" (8x), "is especially helpful" (12x) and "makes it easier" (6x). Going the other way, Opus 5 used "is genuinely" (26x) and "matters enormously" (15x) more. One version step inside one company, and the lists diverge this far.
3.2Different companies, different habits
Set Opus 5.5 against GPT-6 Astra and the two lean opposite ways. Opus 5.5 leans superlative: "one of the best _ about" runs at 79 times Astra's rate, and "the most popular" at 45 times. Astra leans toward reservation: "does not establish" runs at 275 times Opus 5.5's rate, and "may provide" at 18 times. One tilts toward elevating something, the other toward drawing a line around it.
Secondary coverage dropped an order of magnitude here too. One article reported Astra's corrective framing at more than 100 times the human rate, where the primary source puts the composite measure at 12x. What exceeds 100x are the individual phrases inside that category: "not simply" at 157x and "the _ is not simply" at 576x. A category's composite and the individual phrases within it differ by an order of magnitude, so they need to be written separately.
The study also measures which parts of speech produce the gap between the two models. Nouns, verbs, adjectives and adverbs together account for about 79% of it. That 79% is not a share of the gap between humans and AI. It is a part-of-speech breakdown of the gap between the two models. Mixing the two figures produces an entirely different claim.
The style score's yardstick, written by Anthropic
The follow-up study does more than count words. One of its measures scores overall style. Graphite calls it "mannered prose" and defines it as prose that substitutes metaphor and flourish where a direct statement would do. That definition is not Graphite's own.
The results: Opus 5 scores 16.75 and Opus 5.5 scores 10.57, a 37% drop across one version. Human writing measured on the same scale scores 6.65. Opus 5.5 therefore still sits at about 1.6 times the human score, and Astra at 1.2 times. Nothing fell to the floor here the way the em dash did, but the decline is real.
4.1Three layers folded into one place
The quotation above folds three distinct layers together. First, the company that builds the model named a stylistic flaw in its own models and shipped an instruction for suppressing it in public documentation. Second, an outside study borrowed that definition verbatim and used it to score models from the same company. Third, the scorer is the previous version, Claude Opus 5, and the pool being scored includes its successor, Opus 5.5.
There is a fourth layer. The Anthropic document carrying that definition is a section of the prompting guide for Claude Fable 5.1, and Fable is not among the models this study measures. A style prescription written for one model line became the yardstick for scoring another. People at Graphite say they use the instruction routinely, telling current models to avoid mannered prose, and report that it works reasonably well.
Pebblous original diagram (reinterpreting Section 4.1). Each layer, taken alone, is a fact. The problem appears only once all four sit in one sentence — the definition is Anthropic's, the scorer is an Anthropic model, and the subject being scored is the next version in the same line.
4.2That does not make the number untrustworthy
Jumping to "Claude scored Claude, so the number is worthless" would again put words in the study's mouth. Graphite names the possibility of evaluator bias in its own limitations section, and cross-validated by separately scoring about 2,000 articles with a phrase-extraction method. The two methods correlate at Spearman 0.75. The scoring was run blind. A measure whose basis is disclosed to that degree does not deserve to be thrown out wholesale.
The sentence that can be written here is about the provenance of the yardstick, not about reliability. The study's other measures count words, so anyone counting gets the same answer. The mannered prose score alone is a value assigned by a model from the same family as one of the subjects, using a definition written by the company that builds that family. That one cell is not independent of the rest. A second thread runs alongside it: the study's pipeline passes every input article through a GPT-4.1 summarization step before features are extracted. Either way, part of the yardstick remains the same kind of object as the thing being measured.
The grade of the study itself deserves one note. This is not a peer-reviewed paper but research a content marketing company posted on its own blog. Two of its three authors hold doctorates with published backgrounds in machine learning and natural language processing, the method and limitations are described in the text, and the per-feature numbers were released as raw data. As the company's CEO put it, anyone else can open the same files and look for different patterns. That release is why this article could pull figures in Sections 1 and 2 that the study text never prints.
4.3OpenAI is the one that named the em dash in public
So did Anthropic remove the em dash on purpose? The public documents do not confirm it. No expression pointing at the em dash appears anywhere in the Opus 5.5 announcement, the release notes, or the prompting documentation. What the announcement says about style is that the model puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules the user gives it. A tester is quoted saying "it writes the way I do." The em dash is in none of it.
OpenAI makes the contrast. In November 2025 Sam Altman named the em dash directly, announcing that telling ChatGPT in custom instructions to avoid it would finally be obeyed. That too is a statement about following user instructions, not a declaration that the default behaviour changed. And yet the larger measured decline, in Graphite's numbers, belongs to the company that never named it. Stated intent and measured change failing to line up is as far as the confirmed facts go. The two figures have different denominators, so they cannot be turned into a claim about which company erased the habit more thoroughly.
Even the closest available answer to the question of intent exists only in the form of a guess. The Graphite text says the decline may reflect deliberate efforts by AI companies to remove these tells. The person who ran the measurement phrases it the same way. Asked about it, he said his guess is that the labs have targeted evals and are trying to remove certain things, and marked it as a guess.
4.4The primary source does not dismiss detectors
Some write-ups of this study ran under headlines announcing that your AI detector is already obsolete, with body text arguing that traditional detection methods lose their usefulness before they can establish themselves. Those write-ups do not quote the last line of Graphite's limitations section, which says the opposite.
This article stays on the same side of that line. Graphite has never dismissed detectors, and the argument here does not rest on Graphite alone, which is why Section 5 brings in independent academic literature. That a sentence the primary source wrote in plain contradiction got dropped so a headline could be built is itself a map of where the careful footing lies in this subject.
The human baseline, frozen at November 2022
Everything so far has been about the machine side. This section looks at the other one. Calling a piece of writing AI-flavoured requires a reference for what human writing usually looks like. In this study that reference is 10,000 Common Crawl articles published before 30 November 2022, the date ChatGPT was released. The reason for cutting there is plain. The web after that date contains AI-written text, so using it as the human reference contaminates the reference. A companion Graphite study reports that about half of the articles published online in the first quarter of 2026 are AI-generated. How much AI text has accumulated on the web, and how sharply the share splits by domain, is covered in an earlier article of ours.
The people who made the cut know what it costs. Their limitations section says so.
5.1One side moves and the other cannot be fixed
Two asymmetries sit here. The machine-side distribution moves wholesale at every version change: as Section 3 showed, consecutive versions share between 18.7% and 45.5% of their tells. The human-side distribution is pinned at November 2022, and there is no way to refresh it. Redrawing it would mean collecting writing produced by people in 2026, and the way to confirm that people wrote it is the very tool being built. The reasoning closes on itself.
Pebblous original diagram (reinterpreting Section 5.1). The horizontal axis is time. The human baseline is pinned and does not move; the AI-side line rises and falls with every version. The distance between them is the shelf life of the label.
The human side is not standing perfectly still either. The person who ran the measurement passed along an anecdote: people are changing their writing habits so as not to sound like AI. He labelled it anecdotal himself, so it cannot carry weight as evidence. The direction is clear enough, though. The distance between a baseline drawn in 2022 and the writing people produce now widens on its own, and widens faster if there is a movement to avoid sounding like a machine.
If one side moves, the other sits still, and the still side cannot be fixed, then verdicts issued between them acquire a shelf life through no one's fault. It is a product of the structure. The question here is how long that shelf life runs.
5.2Detectors are weakest against models that just shipped
Graphite counted words. It did not measure detector performance. From here, then, the academic literature takes over. Three numbers are drawn from two papers published in 2026. One shows how much of a new model's output a commercial detector misses right after that model ships. One runs the other direction, measuring how confident a machine verdict on human writing tends to be. The third asks how far the same estimator's output spreads when applied to human writing from different years.
| What it shows | Figure | Conditions |
|---|---|---|
| The machine side moves | AI recall of 34.0% for the commercial detector Pangram 66% of AI-written text judged human-written |
Immediately after GPT 5.4 shipped, on arXiv abstracts. The same detector is near-perfect on older models (arXiv 2606.25152) |
| The direction of the errors | 60.4% of human-written text judged "machine" at p≥0.95 99.4% of predictions land in the extreme-confidence bands |
MAGE benchmark, leave-one-domain-out training, fine-tuned RoBERTa. The paper's own phrasing: "its confidence is often misplaced" (arXiv 2607.03680) |
| The human side ages | The same estimator puts 2.3% of 2010 writing and 15.1% of 2020 writing at LLM-generated | A prevalence estimator, not a binary classifier. Sentence level. A simulation setting that assumes LLMs became available after 2010. The true value for both years is near zero (arXiv 2606.25152) |
2020 is two years before ChatGPT was released. An estimate of 15.1% LLM-generated for that year's writing means the estimator reads drifting human style as machine-side. The control is the 2.3% the same estimator returns for 2010 writing.
Supporting figures point the same way. Faced with a generator held out of training, the accuracy of a RoBERTa-based detector falls from 99.26% to 60.30%. One benchmark holds the domain fixed, swaps only the generator, and watches accuracy drop from above 95% to below 60%. Tie down every other condition, change nothing but which model did the writing, and the result swings this far.
Something should be kept separate here. What this article covers is drift produced by model generations turning over, not cases where a person edits text deliberately to evade detection. The latter has an adversary; the former does not. We took those tools apart in an earlier teardown, where evasion on Korean sentences is worked through in detail. Why detector false positives are statistically hard to reduce is covered on a different axis in another article. The failure axis here is not the structure of false positives. It is the aging of the reference distribution.
5.3We also looked hard for evidence that detectors hold up
Collect only the figures above and the easy next step is "AI detectors are finished." So we went looking for evidence on the other side. It exists. Three pieces of it.
- In a July 2026 evaluation by Epoch AI, Pangram missed none of the 297 pieces generated by three contemporary frontier models from default prompts. The same evaluation produced no false positives across 495 human-written pieces.
- In a peer-reviewed paper published by researchers at Vrije Universiteit Brussel in June 2026, Pangram caught 97.5% of 40 master's theses generated entirely by GPT-4o deep research. Turnitin scored all 40 of the same batch at "0 to 20% AI."
- An NBER working paper reports that Pangram's detection power held up even under a constraint capping the false positive rate at 0.5%.
So this article cannot argue that detectors are useless. That is not the primary source's claim, and it conflicts with the academic literature besides. And the people building detectors never leaned on punctuation heuristics to begin with. GPTZero moved to deep learning in autumn 2023 and stopped using perplexity and burstiness. A visible signal like the em dash flipping sign does not immediately disable a neural detector. The two layers are different layers.
5.4Line up the human samples those studies used
Going out to look for a counterexample turned up something else. Put the three evaluations above in one table, asking when the writing they treated as "human-written" was produced, and a single shape appears. The primary source of this article is on the same list.
| Study | Human-side sample | When it was written |
|---|---|---|
| Epoch AI (2026-07) | 495 pieces from 99 authors | Before 2022. Only authors with enough pre-2022 output were selected, and dates were verified against the Wayback Machine |
| Vrije Universiteit Brussel (2026-06, peer-reviewed) | 40 master's theses | Written before 2019 |
| Pangram 4 (vendor documentation) | 2 million-item false positive benchmark | Pre-2022 commercially licensed documents |
| Graphite (primary source of this article) | 10,000 Common Crawl articles | Before 30 November 2022 |
The four are independent of one another and were built for different purposes: an academic evaluation, a peer-reviewed paper, vendor technical documentation, and a marketing firm's research. At the point of choosing writing that is certainly human, all four made the same choice.
The working definition of "certainly human-written" in this field has effectively become "published before 2022." The Pangram documentation does not even explain the cutoff. That reads as a sign the choice has become obvious enough to need no explanation.
So the honest answer runs like this. Evidence that detectors hold up does exist. But the low false positive rates that evidence reports are all false positive rates against pre-2022 human writing. We found no study reporting a false positive rate against writing produced by people in 2026. Not because researchers were lazy, but because there is no clean contemporary human corpus from which to produce that number. The positive evidence and the argument of this article do not collide. They are two faces of one structure. Detectors work well, measured against a baseline frozen in 2022.
How far one product's numbers scatter across conditions is now readable too. Pangram missed none of the default-prompt output from three frontier models, and the same product judged 66% of GPT 5.4's AI-written text as human right after that model shipped. A 0% miss rate and a 66% miss rate belong to the same product. Detectors are strong against models they have caught up with and weak against models that just shipped. That is what the shelf life consists of.
A report pointing the other way belongs here as well. One German-language authorship verification benchmark reports finding no evidence that the AI era, 2023 through 2025, produced a systematic distribution shift in authorship verification. The same paper also reports that temporal drift is the strongest single degradation factor: in the review genre, a ten-year gap costs 0.21 in F1. The authors' own stated limitation is clear. Their corpus carries no annotation for AI-assisted writing, so the design cannot measure how much AI is mixed in. The task is authorship verification, not AI detection, and the language is German. Writing the state of the literature honestly: temporal drift is established, and no study has yet isolated an AI-era effect independently.
5.5Why the em dash broke first
One question remains. Among thousands of tells, why did the most famous one break first and break hardest? A 2026 paper that dissected detectors with explainable AI methods offers a general rule: within a domain, the feature with the highest discriminative power is also the feature most vulnerable when the domain shifts. Discriminative power and fragility come from the same cause.
Apply that rule to the em dash and the story fits. It was the sharpest signal, which is why it became the best known, and being best known is why it became the first target for correction. One line needs drawing, though. No paper has measured the em dash's own contribution to detection. The rule above is general, and that the em dash broke first is Graphite's observation. The two can be read together, but it cannot be written as "a paper measured the em dash." And the cross-domain drop in that paper is itself asymmetric by direction, so citing one direction alone overstates it.
An inverted signal is worse than a dead one
A signal can age in two ways. It can die, or it can invert. For data quality these are entirely different accidents. A dead signal lowers accuracy, and lower accuracy registers on a dashboard. Someone says something looks wrong. An inverted signal does not lower accuracy. It returns the opposite answer. And the answer arrives with high confidence.
The figures in Section 5 show the shape of that confidence. The detector that judged 60.4% of human-written text to be machine-written at probability above 0.95 placed 99.4% of all its predictions in the extreme-confidence bands. In the paper's phrasing, its confidence is often misplaced. A hesitant wrong answer gets caught. A decisive wrong answer passes straight through. An inverted signal produces wrong labels quietly and steadily.
6.1One rule that is running right now
None of this is abstract. Plenty of organizations have a line in their content policy or editorial guidelines saying heavy em dash use marks a draft as AI-written. Any organization still running that rule is now running a rule that is wrong in exactly the opposite direction. Under this study's conditions, current models all use the em dash less often than people do, so the rule now tilts toward flagging human writing as AI. Not one character of the rule changed. What the rule does changed.
The same question returns wherever training data is bought and sold. A guarantee that text is human-written is already priced in. The next step is asking what that guarantee was based on: which classifier, run when, compared against which reference corpus. Opus 5.5 at least was an announced replacement carrying a version number, so the date the behaviour changed is knowable. Where a model changes without a version bump, even the start date of the drift is unknown.
6.2None of the four specifications has a field for the verdict
So where should those conditions be recorded? Several specifications already exist for recording origin and history, which seems like the natural place. We opened four: C2PA 2.4, the international content provenance specification; Article 50 of the EU AI Act and its code of practice; Hugging Face dataset cards together with Croissant metadata; and the Data Provenance Standards 1.0.0 from the Data & Trust Alliance. The gap has the same shape in all four.
| Specification | Field for AI status | Verdict date | Verdict criteria | Classifier version | Reference corpus date |
|---|---|---|---|---|---|
| C2PA 2.4 | digitalSourceType | none | none | none | none |
| EU AI Act Article 50 | Machine-readable marking + fully AI-generated / AI-assisted | none | none | none | none |
| HF dataset card / Croissant | Free-text description | none | in prose | none | none |
| D&TA Data Provenance Standards 1.0.0 | Method (machine-generated, etc.) | none | none | none | none |
C2PA's when is the time the content was created, and softwareAgent.version is the version of the creating tool. Neither carries a value about a verdict. The only place a date survives under EU Article 50 is the record of human involvement claimed for the editorial exemption. All three date fields in the D&TA standard point at the data itself.
The reasons for the gap are reasonable ones. C2PA is a declaration specification, not a detection specification. Its implementation guidance explicitly says not to display "AI detected" to users. The design records only what the creator discloses. Article 50 likewise runs in the direction of providers marking their own outputs. All four were designed to record who made something, and none of them to record who judged it and how.
The difficulty is that declared material is a tiny minority. Everything else comes down to somebody's verdict. And the repository keeps the verdict while the conditions behind it vanish. The sharpest instance is in the D&TA standard. Of its 22 fields, three are devoted to when the data was created, and not one to when, or by what yardstick, the human-or-AI verdict attached to that data was issued. For completeness, C2PA 2.4 is on the international standards fast track as ISO/DIS 22144, which means this gap is about to set as a standard.
What belongs in the repository, then, is not the verdict but the conditions the verdict stood on: the date it was issued, the criteria used, the classifier version, and the date of the reference corpus. Why those four have to travel with the label was covered once already in an earlier article that asked the same question about medical imaging labels. How far EU Article 50 extends the marking obligation is written up separately. So this article adds a confirmation: no current specification has a field for those four.
6.3Four things that can be done now
This is no place to close on pessimism. The 2026 literature on drift offers four mitigations with measured evidence behind them, ordered here from the most immediately applicable.
- Watch distribution shift on the input side first. Without remeasuring detector performance, tracking how far the linguistic feature distribution of incoming text has moved gives advance warning of falling label reliability. A correlation of 0.416 is reported between the magnitude of shift in past-tense verb distribution and cross-model generalization accuracy. Since it does not require waiting for a new model to appear, this is the most operational of the four.
- A small recalibration recovers a good deal. Ten labelled samples per class raised the detection rate at the 1% false positive operating point from 0.588 to 0.690. That is not full retraining. It also falls short of in-distribution performance. A mitigation, not a fix.
- Mix several generations of generators into training. The greater the diversity of generators used in training, the better the generalization to generators never seen before. The answer is mixing several generations in, not adding one more recent model.
- Look at approaches that measure properties of the generation process instead of vocabulary. Some research uses signals that do not depend on any particular model's lexical taste, such as the variance in unpredictability across human writing, or the damping of that variance in the later part of a generated passage. These wear down more slowly in principle as new models arrive. We could not verify the conditions behind the reported figures, so this one is recorded as a direction only.
None of the four is a way to detect better. All four are ways to manage how fast a verdict goes stale, and that difference is the conclusion of this article. A label is not an observed fact. It is a verdict issued by setting two distributions against each other, and a verdict moves when its comparison point moves. A pipeline that stores only labels is a warehouse stacking food with the expiry dates rubbed off.
Why Pebblous Cares
Sections 1 through 6 record what we confirmed in Graphite's two studies, their published raw data, and the academic literature branching out from there. This section is the part those documents do not say. It is about why Pebblous spent so long looking at this research.
7.1That label is the thing we work on
The work at Pebblous is asking about the condition of data before it enters training. DataClinic diagnoses the state of a dataset, AI-Ready Data covers the cleanup that precedes training, and data lineage records what came from where. Among those, this research touches the provenance label. Most training data curation today runs on a two-slot label: human-written and AI-written. Filtering synthetic data, weighting human text, honouring the provenance clauses in a contract, all of it leans on those two slots. What this research shows is that the signal used to apply that label changed sign within a single model version.
7.2The label stays, its meaning shifts
Data quality discussions usually cover missing values, duplicates and range violations. What surfaces here is a different kind of defect. No one touched the cell marked "human-written," and the baseline it was filled against has aged since. It is not a missing value, not a duplicate, not out of range, so existing quality checks will not catch it. And the training runs that weighted on the strength of that label have already happened.
Honesty requires noting that this is not somebody else's problem. Two months ago Pebblous covered a study estimating the share of AI-assisted writing across 1.19 million biomedical papers from the excess use of LLM-preferred words. The premise of that method is that when an LLM touches a text, marker vocabulary rises above its ordinary human level. As Section 2 showed, current-generation models use that very family of words less often than people do. The two measurements differ in corpus, word list and prompt conditions, so we cannot say the estimate is wrong. The sentence that can be written is conditional: counting by excess marker vocabulary presupposes a model generation that overused that vocabulary, and if the current generation uses it less often than people do, the same yardstick applied unchanged can err on the side of undercounting.
7.3One more item on the diagnostic list
Diagnosing a dataset has so far meant asking about the state of its values. One line gets added. When, with what, and against which reference was the provenance label on this dataset applied? As Section 6 showed, no current specification has a field for those four, so recording them today means writing them by hand into the free-text section of a dataset card. Put the other way around, the position of defining that field as a specification is vacant. A question of the same kind came up once before, on the side of measuring AI usage rates with an independent corpus.
7.4The problem that starts after marking
Provenance marking has so far been debated as a question of whether to mark at all: whether to embed a watermark, whether to attach C2PA, how to comply with Article 50 of the EU AI Act. The question this research poses arises only once the marking is done. How long does an applied label stay true? Pebblous is positioned to put that question into Korean first. And our own publishing pipeline runs a yardstick that measures AI style by frequency, which means we can speak to the way such a yardstick wears out from our own records. We already went through that self-application record in an earlier article, so it gets one line here.
The verbatim quotations and figures in this article were checked directly against the text of Graphite's two studies and the raw data files they released under CC BY 4.0. Every value in the Section 1 and Section 2 tables was reconfirmed by reopening those files. No figure was taken from secondary coverage; only quotations were, with the source named. Sections 1 through 6 are what we confirmed, and Section 7 is the part those documents do not say, so please read them separately. Thank you for reading this far.
References
The sources differ in grade, so they are listed in groups. The first group is the research this article examines and the raw data its authors released; every figure in the body was checked against it. The second group is specifications and company announcements. The third group passed through news coverage, and every use of it names the attribution inside the sentence. The fourth group is the academic literature behind Sections 5 and 6.
The research examined and its open raw data (primary)
- 1.Graphite. AI Tells: Opus 5.5 Update. 1 October 2026. graphite.io · Source of the em dash drop from 2.92 to 0.015, the 2,548 tell count, the 6% decline in the mean strength of 11 well-known tells, the mannered prose move from 16.75 to 10.57, the per-category phrase multiples, and the phrases swapped between versions.
- 2.Graphite. AI Tells. 16 September 2026. graphite.io · Source of the methodology (10,000 human articles, the 2022-11-30 cutoff, the threshold conditions, the GPT-4.1 summarization step), the verbatim limitations section, the "may have overcorrected" line about GPT and Gemini, and the 12,877-tell union across all models.
- 3.Graphite. AI Tells open raw data (CC BY 4.0). ai_tells_features.csv.gz · ai_tells_tells_vs_human.csv.gz · Source of the 11-model em dash figures in Section 1, the human baseline of 3.137847, the eight-word sum in Section 2, and the 18.7 to 45.5% shared fractions in Section 3. Requested citation: Graphite, "AI Tells" (2026).
- 4.Graphite. More Articles Are Now Created by AI Than Humans. graphite.io · The tally putting about half of online articles in Q1 2026 at AI-generated.
Specifications and company announcements (primary)
- 5.Anthropic. Claude Fable 5.1 prompting documentation, "Writing density" section · The definition of mannered prose lives here, and Graphite carried it verbatim into its scoring prompt.
- 6.Anthropic. Claude Opus 5.5 announcement. 22 September 2026 · The style description in the Communication section. The confirmation that no expression pointing at the em dash appears in the announcement, the release notes or the prompting documentation was made here.
- 7.Sam Altman (OpenAI). Post on X. 14 November 2025 · The announcement that turning off em dashes via custom instructions now works properly.
- 8.C2PA Technical Specification 2.4 and Guidance for AI and ML 2.4. spec.c2pa.org · The meanings of digitalSourceType, when and softwareAgent.version, and the implementation guidance against displaying "AI detected" to users. ISO/DIS 22144 fast-track status.
- 9.EU AI Act Article 50 · Code of practice on transparency for AI-generated content (AI Office, 10 June 2026) and Commission guidelines (20 July 2026).
- 10.Data & Trust Alliance. Data Provenance Standards v1.0.0 · 22 fields. The confirmation that all three date fields point at the data itself was made from the field definitions.
- 11.Hugging Face dataset card specification and Croissant (including the Responsible AI extension) · Confirmation that rai:dataCollection and similar fields are free text.
Coverage and independent reproductions (used with attribution named)
- 12.Russell Brandom, TechCrunch. 1 October 2026. techcrunch.com · ⚠️ The only source for the Greg Druck quotations, so only the quotations were taken from it. No figure was. The piece carries several figures that conflict with the primary source: 12,877 written as "more than 13,000," dependable (23x) given as the top tell, a composite measure of 12x written as "more than 100x," and a model name that does not appear in Graphite's work.
- 13.VentureBeat. 30 September to 1 October 2026 · Quotations from Druck and Ethan Smith. The figures match the primary source.
- 14.Mayank Parmar, BleepingComputer. 26 September 2026 · The Text Arena production-response tally (Opus 5 at 15.2, Opus 5.5 at 0.8, semicolons from 6.10 to 1.64, mean response length from 453 to 481 words).
- 15.Jim Bennett, Arize AI. September 2026 · Independent reproduction across 20 research briefs (Opus 5 at 12.9, Opus 5.5 at 0.05). Source of the closing line, "'Fixed' is too strong. 'Much better' is fair."
Academic literature (Sections 5 and 6)
- 16.Ren, Raghavan & Garg. Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift. arXiv:2606.25152, 23 June 2026 · Pangram's 34.0% recall on GPT 5.4, and the prevalence estimate rising from 2.3% for 2010 writing to 15.1% for 2020 writing.
- 17.Shen, Wang, Zou & Bu. Rethinking AI-Generated Text Detection: A Strong Baseline and the Distribution-Shift Problem That Remains. arXiv:2607.03680, 4 July 2026 · 60.4% of human writing judged machine at p≥0.95, 99.4% of predictions in the extreme-confidence bands, and recalibration from 0.588 to 0.690 with ten samples per class.
- 18.Pudasaini, Miralles-Pechuán, Lillis & Llorens Salvador. Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy. arXiv:2603.23146 v2, 22 April 2026 · The general rule that highly discriminative features are the most vulnerable to domain shift. The caveat that the cross-domain drop is asymmetric by direction also comes from here.
- 19.Xia, Stańczak & Roth. Explaining Generalization of AI-Generated Text Detectors Through Linguistic Analysis. arXiv:2601.07974 v2, EACL 2026 · Predicting generalization performance in advance from input feature distribution shift, correlation 0.416.
- 20.Pu et al. Breaking the Generator Barrier. arXiv:2604.13692, 15 April 2026 · Training generator diversity and generalization to unseen generators.
- 21.Kiefer et al. When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era. arXiv:2608.17979, 18 August 2026 · Temporal drift at −0.21 F1, with no AI-era effect detected. ⚠️ The task is authorship verification and the language is German, so it should not be read as detector performance.
- 22.Wang et al. M4GT-Bench. arXiv:2402.11175 v2, 2024 · From 99.26% to 60.30% on unseen generators. / Dugan et al. RAID. arXiv:2405.07940 v2, 2024 · From above 95% to below 60% when the domain is fixed and the generator swapped.
- 23.Basani & Chen. Diversity Boosts AI-Generated Text Detection. arXiv:2509.18880 v3, TMLR 2026 · ⚠️ We could not verify the conditions behind the figures, so it is cited as a direction only. / Sun, Bao, Cui & Zhang. When AI Settles Down. arXiv:2601.04833, 8 January 2026 · Damping of variance in the later part of a generated passage.
- 24.Epoch AI. AI detectors rarely flag human writing. 15 July 2026 · 0% miss rate across 297 pieces from three frontier models, 0% false positives across 495 human pieces, and the selection criterion putting the sample before 2022. / Van Vlasselaer, Van Droogenbroeck & Spruyt. International Journal for Educational Integrity 22(16), DOI 10.1007/s40979-026-00226-w, 29 June 2026 · Peer-reviewed, 40 master's theses written before 2019. / Jabarian & Imas. Artificial Writing and Automated Detection. NBER WP 34223, 2025 · Detection power under a false positive rate ceiling.
Related articles on the Pebblous blog
- 25.A Better Model Will Not Fix AI-Written Korean (Sections 5.2 and 7.4) · The Mathematical Limits of Detecting AI-Written Text (Section 5.2) · AI-Written Pages Piled Up on .com at Ten Times the .edu Rate (Section 5) · 89% of New Biomedical Papers Carry LLM Vocabulary (Section 7.2) · No Provider Could Show That the Evaluated Model Is the One Answering You (Section 6.1) · A Chest X-Ray Model Latched Onto Chest Drains at Block 13 (Section 6.2) · On August 2, Machine-Made Text Gets a Tag (Section 6.2) · A Work Filter Cut Half the Conversations Out of AI Use Statistics (Section 7.3) · Reading AI Watermarks Like Wastewater Testing · A watermark is a signal that gets planted; this article is about a signal that gets left behind.