Executive Summary
Nine in ten open-access biomedical papers published in December 2025 use the words LLMs favor more often than human authors used to. That figure comes from researchers at the University of Tübingen, who read through the full text of 1,194,287 papers in PubMed Central. The average across all of 2025 was 77%, and in 2023 it was 19%. For anyone who trains or retrieves on the literature, "written by a human" is no longer the default assumption.
But the traces are not spread evenly inside a paper. Pick one paragraph at random and Discussion measures 68% while Methods measures 32%, a gap of more than two to one. Methods rises to 54% when the whole section is measured at once, so it cannot be called a clean zone either.
Because the method watches only shifts in word frequency, it cannot separate a sentence that was polished from a passage that was generated from scratch. The 89% is not the share of papers written by AI. It is the share of papers where an LLM left a statistical trace.
Key Numbers
The first two cards say how many papers carry a trace. The last two say where the traces cluster.
Source: Holzwarth, González-Márquez, Kobak, arXiv:2608.10715 (2026-08-11)
89%
Papers published in December 2025
77% across all of 2025, 19% in 2023
1.19M
Open-access papers analyzed
Full texts from 2017 to 2025, 379 marker words
68% vs 32%
Discussion vs Methods paragraphs
Gap holds after matching text length
37% vs 72%
English-speaking vs other countries
Korea highest at 85%, the UK lowest at 28%
What the 89% Counts
The only raw material behind this estimate is how often certain words appear. The researchers took 379 words whose usage jumped after ChatGPT was released and used them as markers. A few of them stand out, such as delve, but many are style-carrying function words like these or potential. The point is not to catch a handful of buzzwords. It is to watch the total volume of the words that rise together whenever a sentence gets smoothed out.
That list was not built for this study. The same lab had already compiled it in earlier work by scanning PubMed abstracts for words whose usage rose sharply in 2024, and it was carried over to the full-text analysis as it stood. The only added step was checking that nearly all of those words rose again in the 2025 abstracts in this dataset. So the 89% rests on a word list chosen by an earlier paper. If a sentence was polished in a way that list does not cover, this yardstick will not register it.
What the approach needs is a baseline from the years when only humans wrote. The researchers extrapolated the pre-2023 trend to estimate how often these words would have appeared in 2025 without LLMs, then converted the amount by which the observed counts exceed that baseline into a share of papers. The material is 1,194,287 open-access papers in PubMed Central, published between 2017 and 2025.
Those 1.19 million papers are not everything in PMC. Starting from the open-access snapshot of 23 January 2026, which holds 7,125,722 papers, the researchers removed non-English papers and papers without a publication date, then kept only those with all four of Introduction, Methods, Results and Discussion, where each section and the abstract hold at least 250 characters of English. That leaves 1.60 million, of which 1.19 million were published in 2017 or later. The 89% comes out of that subset, so short reports without section headings and papers in journals that are not open access were never in the calculation.
The yearly figures rose three years running, and the same papers measured on their abstracts alone come out far lower.
| Period | Full text | Abstract only |
|---|---|---|
| 2023 | 19% | 9% |
| 2024 | 52% | 31% |
| 2025 (annual average) | 77% | 53% |
| December 2025 | 89% | 68% |
Estimates based on full texts and on abstracts alone. Source: arXiv:2608.10715 Table 1
The gap between the two columns comes from length. The longer the text, the more chances a marker word has to show up at least once. So the 89% is closer to the share of papers judged to have an LLM's touch somewhere inside them. A paper with a few edited sentences and a paper with a fully generated passage do not separate under this method.
Last Year the Number Was 12–16%
Readers who have seen numbers on this topic before will remember much lower ones. The same lab published a study in Science Advances in July 2025 that scanned 15.1 million PubMed abstracts and reported that at least 13.5% of 2024 abstracts had passed through an LLM. The press translated that into one paper in seven.
The two numbers are not in conflict. The earlier figure was a lower bound. If you count only the excess in word frequency, papers that used an LLM but happened not to use any of those words never enter the count. This study proposes a formula that fills the gap by dividing the excess by the probability that a human would not have used the word, which recovers the full share.
A single word shows how large the correction is. In December 2025 abstracts, these appeared in 50% of them, and the estimate for humans writing alone was 33%. Counting the frequency difference alone gives 17%. Correcting it gives 25%.
Most of the distance between 12–16% and 89%, then, comes from the yardstick rather than from any change in what people do. Measured again with the new method, 2024 comes out at 31% for abstracts and 52% for full texts. Set that beside the 13.5% the earlier study found in the same year's abstracts, and the correction alone more than doubles the value.
Whether the correction formula works was checked on a synthetic corpus whose answers were known in advance. The researchers planted 500 marker words across 100,000 texts, varied the fraction touched by an LLM from 0 to 1, and asked whether the formula recovered it. The gap between estimate and truth stayed under 0.02 across the whole range. The nature of that check deserves to be stated plainly, though. It shows the formula solving correctly in a world built the way the formula assumes, and it is not a validation against real papers labeled as human-written or LLM-edited.
The new method also stands on assumptions. Since the baseline is an extrapolation of the pre-2023 trend, the excess is overstated if human style shifted for some other reason. A force pushes the other way as well. As more authors deliberately avoid words that sound like an LLM, real usage runs higher than the estimate. The researchers wrote down both directions.
Discussion Runs Twice as High as Methods in the Same Paper
A Discussion paragraph is twice as likely to involve an LLM as a Methods paragraph. The authors wrote that sentence into their own abstract. Pick one paragraph at random and Discussion comes out at 68%, Methods at 32%.
Measure whole sections and every figure rises, because a long section has more room to hold marker words and length leaks into the result. Putting the researchers' length-matched numbers, taken from 255-word crops of each section, next to the numbers for the sections as they are shows the difference.
| Section | Per paragraph (255 words) | Whole section |
|---|---|---|
| Discussion | 68% | 78% |
| Abstract | 67% | 68% |
| Introduction | 59% | 63% |
| Results | 46% | 58% |
| Methods | 32% | 54% |
December 2025. Source: arXiv:2608.10715 Table 1
Methods sits lowest because of what it contains. Reagent concentrations, instrument model numbers, sample sizes and statistical procedures can only be written by whoever ran the experiment, and they leave little room for smoothing a sentence. Discussion is where results get widened into meaning and limitations get recorded, so generic prose passes with less notice. The abstract reaches 67% even per paragraph because an abstract is a summary, and summarizing is the easiest job to hand to an LLM.
Reading Methods as a clean zone would still be a mistake. Measured as a whole section it comes out at 54%. The distance between 32% per paragraph and 54% per section says less about intensity than about the chance that at least one paragraph inside a long section was touched. The low end is only the low end. It is not zero.
The Unit of Judgment Drops from the Paper to the Section
Everything above is what the paper reports. The paper goes only as far as producing the observations a policy discussion would need, and moving those numbers over to the corpus side is left to whoever reads them.
Scientific literature is the source material that scientific AI trains on and searches through. If contamination were uniform, the decision would be one of two things: filter it out, or use it as it is. Once the figure doubles from section to section, the unit of judgment drops from the paper to the section. A single trust score attached to a whole paper covers 32% in Methods and 68% in Discussion with the same value.
When retrieval-augmented generation pulls evidence passages, there is little reason to weigh Methods and Discussion equally. What reproduction needs sits mostly in Methods, and that is where the figure is lowest. In the other direction, for anyone training on style as well as content, the more Discussion paragraphs go in, the larger the share of the corpus that feeds a model sentences resembling its own output. How that feedback shows up in performance is outside what this paper measures, and for now it can only be stated as a concern.
What practice needs is not a yes-or-no answer on contamination but a map of it. Section, publication year and author country, the axes along which the figure differs, mostly exist as metadata already. Keep those axes on the record when a corpus is built and the weights can be redrawn later. Leave them out and there is no way to recompute. A dataset assembled by pushing the literature in wholesale carries a 2019 Methods section and a 2025 Discussion section at the same weight.
You can also hold the same yardstick up to your own corpus. The code for this analysis is on GitHub, the marker word list is tabulated in the earlier study's repository, and the PMC open-access data is there for anyone to download. What has to travel with the code is the baseline, not the code itself. The pre-LLM word frequencies have to be rebuilt inside that corpus, and if the corpus is thin on material from before 2023, there is nothing to build a baseline on and the method does not stand up at all.
The literature is not the only place where the assumption of human-written data is coming loose. How the same problem showed up in survey responses is covered in You Asked People. Increasingly, AI Is Answering., and an attempt to screen a paper corpus with a classifier is in An AI Detector Screened 2.6 Million Cancer Papers and Flagged 260,000 Fakes. For what has happened on the review side rather than in the papers themselves, see our analysis of 76,139 ICLR 2026 peer reviews.
Editor's Note: When Pebblous talks about AI-Ready Data, axes like these are among the values we think have to travel with the data. Checking whether a value is correct does not settle whether the data can be used for training. Where the text came from, and from which section and which year, has to stay on the record for that judgment to be made again later.
Korea's 85% Reads Two Ways
The axis with the widest spread is the author's country. For 2025 the figures are 85% for Korea, 82% for China and 80% for Taiwan. English-speaking countries as a group come out at 37%, with the UK at 28%. Non-English-speaking countries as a group reach 72%.
All 2025 papers. Source: arXiv:2608.10715 Figure 3, Table S3
Add the time axis and the direction shows. Before 2023 the marker words appeared far less often in papers from non-English-speaking countries. By 2025 those frequencies had nearly caught up with the English-speaking ones. The researchers describe this as linguistic convergence.
The same convergence pulls in two directions. One is that the language barrier has come down. A researcher who used to spend a great deal of time writing in English can put that time into research instead, and fewer papers lose points in review over their English prose. The other is that the range of writing narrows. The researchers list hallucinated citations, the homogenization of intellectual diversity, and the weakening of critical review when editing is handed off as risks that come with it.
The risk the paper presses hardest on is the practice of not disclosing that an LLM was used. Guidance on how to use these tools already exists in several places, yet many authors do not follow it and do not report their use, and if that continues, what erodes is not individual findings but society's trust in science as a procedure. Which new guidance should be written is not this paper's job. What the authors say in closing is roughly that a policy discussion needs numbers first.
For corpus builders this passage arrives as an awkward problem. That the figures differ by country means author country can serve as a weighting axis, but using that axis as a filter starts by screening out the papers of researchers who crossed the language barrier with a tool. Word frequency tells you how a text was written. It does not tell you whether the text is right. What the 89% asks for is not less faith in the literature but the habit of writing down which parts of it you trust, and on what grounds.
References
Academic Papers
- 1.Holzwarth, L., González-Márquez, R., & Kobak, D. (2026). "Most biomedical publications show signs of LLM-assisted writing." arXiv:2608.10715, 2026-08-11.
- 2.Kobak, D., González-Márquez, R., Horvát, EÁ., & Lause, J. (2025). "Delving into LLM-assisted writing in biomedical publications through excess vocabulary." Science Advances, 11(27), eadt3813.
Data and Code
- 3.Kobak Lab. "kobaklab/llm-usage-in-pmc." Estimation and validation code for this paper.
- 4.Berens Lab. "berenslab/llm-excess-vocab." Excess vocabulary list and analysis code from the earlier study (2025), and the source of the 379 marker words used here.
- 5.National Library of Medicine. "PubMed Central Open Access Subset." The source corpus for the analysis, as of the 23 January 2026 snapshot.