Executive Summary

Nine researchers in physics and complex systems rebuilt the spread of LLMs as an epidemic model. People fall into three compartments: those who do not use LLMs or use them lightly, those who use them while still reading, writing and reasoning for themselves, and those for whom the model has taken over the cognitive work. Adding a single term to the rates that move people between compartments changed the result. The model produced two thresholds rather than one. The value that lets dependence take hold and the value that lets a population come back are different numbers, and in the interval between them the population can sit in either state under identical conditions. The paper names its own headline result as this: collective change can be abrupt even while individual adoption stays gradual. In its own words, prevention can be considerably easier than reversal.

The paper collected no data of its own. It leans instead on other people's experiments, and the strongest of them is a field trial with roughly a thousand high school students. Students given a tool that mimics a standard chatbot interface scored 48% higher while they had it. Once it was taken away they scored 17% lower than classmates who had never had access at all. Students given a tutor whose prompts were designed to safeguard learning barely showed that drop. When the paper argues that the question is not whether you use these systems but how you couple to them, this trial is the evidence it rests on.

Holding the model up against the world is where it runs into a wall. Only two of its five parameters can be approximated even loosely from public statistics; no statistical agency publishes the return rate, the recovery rate or collective reinforcement. Everything that is measured sits on the prevention side, and everything unmeasured sits on the reversal side. Corporate AI logs record how much people use the tools, never which compartment a given person is in. Testing this paper would take a kind of data that does not yet exist.

+48% → −17%

Grades with the tool, then without it

Field trial with roughly a thousand high school students. Measured by a study this paper cites, not by this paper

3 of 5

Parameters with no public statistics

Return rate, recovery rate, collective reinforcement. All three sit on the reversal side

0.40 ↔ 0.50

Coming-back threshold and crossing-over threshold

Dimensionless values from the illustrative parameters the paper chose. The span between them is the hard-to-reverse zone

28.3% vs 54%

US generative AI adoption in the same year

Change the survey definition and the figure nearly doubles. Adoption is not transmission pressure itself

1

Three Compartments and Four Arrows

The paper is called "Large-Language Models as a Cognitive Virus." It went up on arXiv on 3 September 2026, filed under physics and society. Nine authors signed it, with Barcelona at the centre: the Complex Systems Lab at Universitat Pompeu Fabra, ICREA, the Institut de Biologia Evolutiva, and researchers on the neuroengineering side, joined by the University of Padua, the Allen Discovery Center at Tufts and the Wyss Institute at Harvard. Three of the nine hold appointments at the Santa Fe Institute. The body runs 12 pages with 137 references. No journal version exists yet.

The word in the title deserves an early note. Calling something a virus is not the same as arguing that LLMs are harmful, and the paper blocks that reading itself in section 1.

"The viral analogy does not imply that LLM-human interactions are intrinsically parasitic. Biological viruses range from pathogens to mutualists and evolutionary partners."

The borrowing is a calculation framework, not a verdict. Epidemiology sorts a population into compartments by state and writes the rates of movement between them as differential equations. This paper does the same with three compartments, and unlike the epidemic models most people have seen, there is no immune compartment. The subject is movement between states people can enter and leave repeatedly, not a disease you catch once and never catch again.

The paper also lists the precedents for aiming this machinery at human behaviour. Technology diffusion has been written in epidemic equations before, and compartment models have been built for drug and social media addiction. Section 2 lays out that lineage. This work adds one more entry to it.

So what plays the part of the virus here? Even with that word in the title, the paper states that the model alone cannot settle the question.

"the fact that the state variables describe hosts does not imply that a host state is the pathogen. The equations model the epidemiology of coupling and therefore do not, by themselves, determine the identity of the viral analogue."

Something close to an answer arrives at the end of section 5 rather than in the main argument. There is a wider ecosystem in which culturally transmitted usage practices, human coupling states, institutions and technical infrastructure are entangled, and the long-lived technological lineage embedded in it is what the cognitive virus corresponds to. The equations cover one layer of that ecosystem: movement among the human states. That sentence narrows in advance what the model is entitled to claim.

1.1The Line Between Compartments Is Not Volume of Use

The first compartment holds people who do not use LLMs, or use them only lightly. The second holds people who use them while keeping reading, writing, reasoning and verification in their own hands and reaching for other information sources too. The third holds people for whom LLM-mediated work has strongly displaced the work they used to do themselves. The paper's own labels for the three are uncoupled, autonomous coupled and persistently dependent.

Note that what separates the second from the third is not hours or session counts. People in both compartments use LLMs. Someone may prompt ten times a day and check the grounds for every answer afterwards; someone else may prompt once a day and let that one answer close the question. The model does not sort those two by frequency. The compartment follows from which work is still being done by the person.

The paper puts that criterion in two words: scaffolding and substitution. Scaffolding lowers the immediate load while leaving, or building, the capacity to do the task alone later. Substitution removes the need to do the cognitive work at all. The same model falls on the first side when it is used as something to rebuild an argument against, interrogate and check, and on the second when it finishes the synthesis, the judgment and the writing on the person's behalf. Moving from the second compartment to the third is the minimal drawing of that shift.

Cognitive offloading is therefore not the thing being criticised. The paper describes it as a normal and often beneficial component of human cognition, and its headline example is reading and writing. Literacy recruits and reorganizes neural circuits that were already there, and symbols written down outside the head extend memory and open inferences that were not available before. The dividing line is whether external support builds internal capacity or takes its place.

There is already a vocabulary for the distinction. Following an essay one of the co-authors wrote a decade ago, the paper separates complementary cognitive artifacts, which strengthen capacities beyond their immediate use, from competitive ones, which improve performance while potentially weakening the underlying skill. Then it states the test: not whether the tool is used, but which cognitive capacities remain when the tool is withdrawn. That test appears here for the first time and comes back with numbers attached in section 4.

1.2The Four Arrows Between Compartments

Three compartments leave four paths between them. The path from uncoupled to autonomous coupling runs through exposure: social practice, institutional demand, features a platform pushes at you. The paper calls the strength of that push the transmission pressure. There is a path back from autonomous coupling to uncoupled as well, the return rate, which will come up more often in this piece than any other parameter. The path from coupling into dependence is habit hardening. Coming back the other way, from dependence to autonomous coupling, runs on training, verification practices and what the paper calls deliberate cognitive friction.

Stopping at four paths is itself a choice. No arrow jumps straight from uncoupled to dependent, because dependence is taken to develop mainly through regular use. That leaves out the newcomer who hands everything over on the first try. The paper flags the simplification and says what it buys: with that case removed, the simplifications isolate the interaction between contagion-like technological adoption and collective protection of cognitive autonomy. A model takes its character from what it erases as much as from what it keeps.

The substitution to avoid here is reading transmission pressure as an adoption rate. The paper states in its own text that the incidence term should be interpreted as an effective host-side social or institutional transmission pressure. The share of people currently using LLMs is a consequence of that pressure, not the pressure itself. Section 6 returns to why that distinction keeps collapsing in practice, with the adoption statistics laid side by side.

Uncoupled No use, or only light use Autonomous coupled Uses LLMs, but keeps reading, writing and verification Persistently dependent Cognitive work has been substituted Transmission pressure Return rate Dependency onset Recovery rate Social practice, institutional demand, platform-pushed features Training, verification practices, deliberate cognitive friction One thing is still missing from this picture. Section 2 adds one more arrow.

Our redrawing of the state transition structure in the paper's Figure 1(a). The state names and the meaning of each arrow follow that figure's caption; the layout and the English labels were added here.

1.3Good for a Person and Good for a Population Are Different Things

There is a reason to go to the trouble of building a model like this. What happens to one person and what happens to a whole population need not point the same way. One experiment the paper cites in section 1 shows the gap clearly. Writers who received story ideas from an LLM produced stories judged more creative, better written and more enjoyable, and the effect was larger for writers who were less creative to begin with. But the stories written with LLM help resembled each other more. Individuals improved while the range of stories the group produced narrowed.

The Pebblous blog has covered which tasks to hand to AI, sorted along cognitive dimensions. That piece dealt with how one person divides up work; this paper asks what shape the result takes once those individual divisions accumulate across a population. The two questions sit on different levels.

2

Only One Term Is New

Everything so far is textbook structure. Anyone who has studied epidemic models can draw compartments, draw arrows and write rates. All this paper adds is a single term hung on the return path. That one term is also what changes the result.

Think of the return path as splitting in two. One share is people quitting on their own: the tool turns out to be a nuisance, the output is not good enough, the job changes. The other share depends on the people around you. Working under your own steam gets easier when plenty of people nearby are doing the same. If several colleagues ask for your evidence in a meeting, coming prepared with evidence becomes the ordinary thing to do; if nobody asks, the person who prepares it starts to look eccentric. The paper writes this second share as proportional to the square of the autonomous fraction.

"The nonlinear term κU²C introduces a cooperative (Allee-like) mechanism. The underlying assumption is that LLM-independent cognitive practice is socially reinforced and cultural expectations that reward independent reasoning become more effective when autonomous individuals are common."

The squaring is the part that matters. Triple the autonomous fraction from 0.2 to 0.6 and the squared quantity goes from 0.04 to 0.36, a factor of nine. The restoring force grows far more steeply than the headcount does. The same holds in reverse: once autonomous people start to thin out, the restoring force collapses faster than their numbers fall.

The coupled fraction multiplying the end of the term carries meaning too. This force acts only on people who already use LLMs. It is not a mechanism for making non-users use them even less; it is a mechanism for pulling users back toward work they do with their own hands.

2.1A Name Borrowed From Ecology

Feedback of this shape already has a name: the Allee effect in ecology. A population large enough to find mates easily and withstand predators together holds up well, but drop below a certain line and those advantages vanish at once, leaving even the survivors struggling. The paper lifts the mechanism directly and attaches it to cognitive autonomy, citing a conservation ecology textbook as the source.

It also cites an experimental precedent for the same property in human convention: once a committed minority in a group passed a certain size, the convention the whole population had been following flipped to a different one. That literature is what sits under the paper whenever it uses the word tipping.

2.2Written Out, It Is Three Lines

The three lines below are that same argument written in symbols. Each line is the people entering a compartment minus the people leaving it, and the three fractions sum to 1. Reading them takes no differential calculus, only attention to the signs: a term with a minus in front is a share flowing out of that compartment, a term with a plus is a share flowing in.

dU/dt = −λUC + ρC + κU²C

dC/dt = λUC − (μ+ρ)C + σD − κU²C

dD/dt = μC − σD

U is the uncoupled fraction, C the autonomous coupled fraction and D the persistently dependent fraction; the three sum to 1. λ is transmission pressure, ρ the return rate, μ dependency onset, σ the recovery rate, and κ the strength of collective reinforcement. Transcribed from the paper's equations (1) through (3).

κU²C appears with a plus on the first line and the same term with a minus on the second. Coupled people are moving back to uncoupled. Take that term away and what remains is an ordinary one-threshold transmission model. Leave it in and there are two thresholds.

3

Going Over and Coming Back Had Different Thresholds

A threshold is the value past which the population settles somewhere else. In a model you find it by looking for the equilibria, the states where the fractions stop changing over time, and checking whether each one is stable or unstable. The equilibria here fall out as roots of a quadratic, and a quadratic has two answers. That is the root of the split into two thresholds.

The first threshold is the return rate plus collective reinforcement. Below that value, in a population where nobody is coupled yet, coupling fails to spread and dies out. The second threshold is twice the square root of collective reinforcement times the return rate. In a population where coupling has already spread, transmission pressure has to fall below this second value before the population returns to its original state. The two values are not equal. When collective reinforcement exceeds the return rate, the second threshold drops below the first and a gap opens between them.

To draw its figures the paper fixes one set of parameters: return rate 0.10, collective reinforcement 0.40, dependency onset 0.20, recovery rate 0.10. These were not measured anywhere; they are illustrative values chosen to produce the plots, a point this piece returns to later. With that set in the equations, the two thresholds come out as numbers.

Quantity Expression Value at the illustrative parameters
Crossing-over threshold return rate + collective reinforcement 0.50
Coming-back threshold 2 × √(collective reinforcement × return rate) 0.40
Width between the two thresholds (√collective reinforcement − √return rate)² 0.10
Point where the two valleys are equally deep solved numerically ≈ 0.420
Condition for bistability collective reinforcement > return rate met

Every value is transcribed from a number printed in the paper's text or figure captions. These quantities are dimensionless and should not be read as percentages.

3.1Two Endings Are Possible Under the Same Conditions

The paper calls the interval where transmission pressure lies between 0.40 and 0.50 bistable. The hard-to-reverse zone this piece names, the lock-in interval in the table of section 5 and the hysteretic regime in the passage quoted there are all names for that same stretch. Inside it, a population that stays wholly uncoupled is stable, and so is a population where most people have coupled. Which one you are in is settled not by present conditions but by how you got there. A population climbing from below stays uncoupled until it reaches 0.50; a population descending from above stays coupled until it reaches 0.40. Two societies under identical pressure can therefore sit in different places.

Transmission pressure rises → 1 0 Uncoupled fraction 0.40 0.50 Hard-to-reverse zone All uncoupled (stable) Unstable boundary Mostly coupled (stable) Climbing, it drops here Descending, it must reach here before it can rise

Our redrawing of the bifurcation structure in the paper's Figure 2(a), from the same equations and the same illustrative parameters. Solid lines are stable equilibria; the dashed line is the unstable boundary separating the two stable states. We computed the curve coordinates ourselves by putting the illustrative parameters into the paper's equation (7).

3.2Two Valleys and One Marble

The same structure comes through a second picture too: a landscape with a marble rolling across it. The image belongs to the paper, not to us.

"The dynamics can be pictured as a marble moving over the landscape until it settles in a valley. … Nevertheless, the marble may remain trapped in the high-autonomy valley until this valley disappears at λTC."

At low transmission pressure the landscape has a single valley, on the high-autonomy side. Past 0.40 a second valley appears opposite it. As pressure rises further that valley deepens, and around 0.420 the two are equally deep. From there the offloading side is the better place to be, yet the marble has not moved, because the wall of its own valley is still standing. The wall vanishes entirely at 0.50, and the marble rolls all the way to the far floor in one go.

At low pressure There is only one valley on the high-autonomy side At about 0.420 The two valleys are equally deep The marble is still held by the wall Past 0.50 The autonomy valley disappears The marble rolls to the other side

Our redrawing of the potential landscape in the paper's Figure 3(c), split into three scenes. The horizontal axis is average cognitive competence, with high autonomy on the right. This is a conceptual sketch of how many valleys there are and how deep they are relative to each other rather than an exact rendering of the curve; the choice of scenes follows the paper's own description.

3.3Which Is Why Prevention and Reversal Cost Different Amounts

Two thresholds carry a simple practical implication. The force it takes to stop a population crossing over is not the force it takes to bring it back afterwards. Section 5 of the paper settles the asymmetry in one sentence: "Prevention can therefore be considerably easier than reversal."

The same structure is stated from another angle in the significance paragraph that sits ahead of the abstract, where collective change is allowed to be abrupt even while each person's adoption stays gradual.

"These results suggest that abrupt collective changes can emerge even when individual adoption is gradual."

Section 5 explains in words how that can happen, through two social processes that lock together. In the first, use spreads as people learn from other people, as workplaces and schools adopt the tools, and as whatever works gets copied. In the second, an environment where reading, writing, reasoning and checking are repeated and rewarded holds autonomous thinking up. Wider offloading weakens the second environment, and a weaker environment makes offloading the easier and more attractive choice. That feedback is why a small rise in pressure produces a large response from the population. The squared term in section 2 is this feedback doing its work inside the equations.

The authors themselves mark where this model stops answering. A sentence in section 3 states that this minimal model does not determine how fast the transition unfolds in real time. Saying the population shifts abruptly past a threshold is a claim about the structure between states, not a prediction that the shift takes months rather than years.

4

Where the Numbers on Lost Competence Came From

Everything counted so far has been headcount: what share of people sits in which compartment. Section 4 takes one step further, assigning a competence score to each compartment and averaging across the population. How that average moves the moment a threshold is crossed is the content of the section, and it is also the part of this piece that needs the most careful reading.

The first thing to pin down is what competence means here. The paper's definition is a narrow one.

"Here Γ should be narrowly interpreted as cognitive competence (CC): the capacity available to the human when external support is removed, rather than the total capability of the coupled human-AI system."

Performance with the tool in hand and capacity left over once the tool is gone are routinely run together, and every number in the rest of this section means something else entirely without that distinction.

4.1The Number 0.425 Comes Out of an Assumption

For its figures the paper sets competence at 1 in the uncoupled state, 0.5 under autonomous coupling and 0.1 under persistent dependence. Where those three values came from is the point. They were not measured, and the paper says so itself in the very next sentence.

"This ordering is an illustrative modelling assumption, not a general claim about LLM use. In scaffolded or augmentative regimes, coupling could leave subsequent CC unchanged or even increase it."

With those values in place, the average behaves like this. At the moment the crossing-over threshold is reached, average competence falls from 1 to about 0.425; on the way back, with pressure lowered again, it returns to 1 only from a position that has climbed to about 0.617. Both numbers are printed in section 3 of the paper. The drop going down is larger than the climb coming back.

How those values arise fits on one sheet of paper. The average competence of coupled people comes from the share that dependency onset and the recovery rate fix between them, multiplied by the assumed values for coupling and dependence respectively, which gives 0.233 at the illustrative parameters. On top of that sits however much of the uncoupled population is left. At the transition point 0.25 of it remains; at the return point 0.5 does. Twice the remaining fraction means twice the contribution from that side.

Reporting this as a 57.5% cut in cognitive ability would therefore be wrong. The scale itself is an assumption, and what it measures is not the combined performance of a person plus an AI. That is why this piece kept the value out of the summary cards. Showing the shape of the picture the model produces is as far as this number goes.

4.2The Genuinely Measured Numbers Sit in a Different Paper

Has anyone actually measured the capacity that remains once the tool is withdrawn? One of the works this paper cites measured exactly that: a 2025 PNAS field experiment in high school mathematics. Nearly a thousand students were given a generative AI tutor, in two versions. One mimicked a standard ChatGPT interface, which the authors call GPT Base; the other carried prompts designed to safeguard learning, which they call GPT Tutor.

Condition Grades while the tool was available Grades after it was taken away
GPT Base (standard chatbot interface) +48% −17%
GPT Tutor (prompts safeguarding learning) +127% negative effects largely mitigated

The second column is the improvement relative to students who never had access, and the −17% in the third column is measured against those same students. The figures are as printed in that paper's abstract; they were not measured by the cognitive virus paper this report is about.

Grades rose sharply while the tool was in hand. Once it was withdrawn, that group did worse than classmates who had never touched it. The mechanism the authors report is plain: without safeguards, students used the model as a "crutch" during practice sessions, so the ability to solve problems alone never developed. With the same underlying tool, the version designed to safeguard learning barely produced that drop.

The cognitive virus paper compresses this into a single citation line in its section 5: unrestricted access to generative AI improves performance while the tool is present but lowers subsequent unaided performance, and pedagogically constrained AI substantially mitigates that effect. The experiment puts numbers on what the model's assumed values state in words. Joining the two calls for care, though. The experiment's figures are individual grades and the model's parameters are population fractions. Experimental results cannot be converted into parameter estimates for the model.

4.3Two Numbers This Piece Decided Not to Use

The citation list also holds a 2025 EEG study that measured brain connectivity while participants wrote essays and then re-measured months later with the conditions swapped. What the published abstract and the researchers' official project site do confirm is that connectivity scaled down systematically as external support increased, and that the LLM group fell behind at quoting sentences they themselves had written minutes earlier.

Two other numbers travel widely with press coverage of that study and appear neither in the abstract nor on the official site. We ran into both during research, failed to find any basis for them when checking against the sources, and left them out. The researchers have posted a request to journalists on the project site.

"Please do not use the words like 'stupid', 'dumb', 'brain rot', 'harm', 'damage', 'passivity', 'trimming' and so on. It does a huge disservice to this work, as we did not use this vocabulary in the paper, especially if you are a journalist reporting on it." The same page adds: "Lots of media and people used LLMs to summarize the paper. It adds to the noise." This report came out of a pipeline that used LLMs as a tool, so that second remark applies to us as squarely as to anyone. Which is why every figure in the body was checked against the source document itself, and anything that failed to check out was left out.

5

Immunization Is Not Quitting AI

A model that talks about abrupt transitions invites prescriptions that drift toward bans. This paper goes the other way. It uses the word immunization and then states twice, in sections 4 and 5, that immunization is not about blocking contact.

"At the population level, immunization does not mean preventing contact with AI. It means preserving the practices and institutions that keep human cognition active: unaided problem solving, verification, critical discussion, periods of deliberate disengagement, maintenance of non-AI skills, and educational designs in which the model supports rather than completes the cognitive task."

The paper adds that immunization therefore need not mean lowering total usage. Immunization is shaping the coupling so that heavy use travels together with verification, active reasoning, autonomous alternatives and recovery.

5.1Only Four of the Six Interventions Move the Thresholds

Section 4 of the paper collects the interventions in a six-row table, each row naming the parameter it moves and the consequence. We added one column: whether the intervention changes the tipping landscape itself or leaves the landscape alone and only changes the burden carried in the coupled state. That split restates in table form what the same section spells out as two distinct levels.

Intervention What it moves Do the thresholds move?
Reduce propagation of substitutive or dependency-producing coupling Lowers transmission pressure Yes
Preserve autonomous alternatives and routes back to unaided cognition Raises the return rate Yes. Raises both thresholds and shrinks the lock-in interval
Strengthen collective autonomy Raises collective reinforcement Yes, though the lock-in interval can widen
Favor recovery over collective lock-in Raises the ratio of return rate to collective reinforcement Yes. When the two thresholds merge the transition becomes continuous
Prevent progression to dependency Lowers dependency onset No
Promote recovery from dependency Raises the recovery rate No

The left two columns transcribe the six rows of the paper's Table 1. The right column is our one-cell summary of the two levels that section 4 of the same paper sets out.

5.2A Privilege That Belongs Only to the Return Path

The second row is the most powerful lever in this model. Raise the return rate and both thresholds rise together while the span between them shrinks. Once the return rate reaches or exceeds collective reinforcement, bistability disappears altogether. The landscape never splits into two valleys, and instead of a sharp fall the system moves gradually with pressure. This is the only way to make the hard-to-reverse zone go away.

The third row has to be read the other way around. Strengthening a culture of independent thought does raise the threshold against invasion, which is plainly good while most of the population is still uncoupled. But the same value also widens the span between the two thresholds. Once the autonomous fraction has already been badly depleted, a strong culture works in the direction of making the return harder. That is why the paper's prescription is not to push collective reinforcement as high as it will go.

"The relevant intervention is consequently not simply to maximize κ, but to reinforce collective autonomy together with sufficiently strong return processes ρ so that protection does not come at the cost of a broad hysteretic regime."

A map makes the same condition visible. Put collective reinforcement on the horizontal axis and transmission pressure on the vertical and the plane divides into three. Below lies a region where autonomy prevails; above, a region where only the offloading state is stable; between them, a band where both states are possible. The band exists only when collective reinforcement exceeds the return rate. Along the line where the two are equal the thresholds merge, and that line is the boundary between smooth continuous adoption and a discontinuous transition whose outcome depends on the path taken. Knowing which region we stand in requires knowing those two values, and those are precisely the two sitting at the centre of the measurement gap in section 6.

collective reinforcement κ grows → transmission pressure λ meet at κ=ρ 0.40 0.50 0.40 autonomy prevails offloading only both states possible

We redrew the paper's Figure 3(b) (κ, λ) phase map from the same two boundary equations (λ_SN = 2√(κρ), λ_TC = ρ+κ). ρ is fixed at the paper's illustrative value of 0.10, the same one used throughout this report. Above the bold orange curve only the offloading state is stable; below the gray dashed curve autonomy prevails; the orange band between them is the hysteretic lock-in zone. Reading off κ=0.40 reproduces the 0.40 and 0.50 already given in section 3's table.

5.3The Levers That Leave the Thresholds Where They Are

The last two rows are a different kind of thing. Lowering dependency onset or raising the recovery rate shrinks the dependent share among coupled people and lifts average competence. At the illustrative parameters two of every three coupled people are in the dependent state, and raising the recovery rate brings that share down. A useful change. The positions of the two thresholds, however, do not budge. The paper writes this property out in equations and explains that these interventions can therefore lower the dependent share without reducing total usage at all.

On recovery the paper's line shifts a little. Once dependence has already hardened, educational design alone will not handle it. Recent work on problematic LLM use points to mechanisms such as loss of control, emotion regulation, cognitive bias and habitual reliance, so behavioural self-regulation and, in severe cases, psychological intervention such as cognitive behavioural therapy may be needed. The immunization metaphor runs all the way from cultural and institutional prevention to recovery at the level of the individual. It also means the recovery-rate lever may not turn on school or company rules alone.

So the levers fall into two layers. Transmission pressure, the return rate and collective reinforcement change the tipping landscape itself; dependency onset and the recovery rate leave the landscape alone and change only the cognitive cost of being in the coupled state. When the two layers get mixed in conversation, a single verification mandate comes to be mistaken for having prevented a tipping point. Section 6 takes real cases and asks which layer they land in.

5.4The Same Activity Reads Into Either Layer

In section 4 the examples start to overlap. The ones given for raising the return rate are protected unaided tasks, periods of deliberate disengagement and the maintenance of non-LLM skills. A few paragraphs later the examples given for lowering dependency onset and raising recovery are metacognitive training, verification requirements and periodic unaided practice. But a protected unaided task and periodic unaided practice are the same activity on the ground. A rule that drafts get written without AI on Friday mornings can be filed under either name.

Inside the model the two names have different fates. One moves the thresholds and the other does not. Yet the paper offers no operational test for deciding which side a given practice belongs on. Since the authors describe their work as a minimal model, calling this an error would be unfair. For anyone trying to carry the model into a real workplace, though, the gap stays open. Finding out which layer a rule actually operated in means observing how people's states moved.

Observation of that kind is not impossible. A 2026 creativity experiment cited by this paper managed it once. University students, 196 of them, were split into three groups: a human-only group working alone, a general-AI group using ChatGPT freely, and a regulated-AI group asked to generate their own ideas first and only then collaborate with ChatGPT. On the first task the general-AI group produced the best results. Then the tool was taken away from everyone and a new task was set, and the order inverted. The general-AI group's creativity fell to the level of the human-only group, and the regulated-AI group, which had gained nothing at all on the first task, outperformed both.

Process analysis shows why. The general-AI group most often dictated to ChatGPT to generate the solutions directly, while the think-first group used the tool to improve ideas they had produced themselves. The condition that performed best in the moment and the condition that left the most behind once the tool was gone were not the same condition.

What this experiment measured is individual subsequent performance, not a population-level return rate. The direction is still clear. At least one controlled study has established what changes when a route back is designed in. Confirmations like this, though, have barely travelled outside the lab.

6

Where This Model Stops

This paper collected no user data. No survey, no logs, no experiment. It writes down equations, shows what shape those equations produce, and then argues from other people's experiments that the shape is plausible. The authors state the status of their own result: not that such runaway dynamics must occur, but that they can occur under plausible forms of social learning and cooperative reinforcement, and that this possibility is the central result.

Skip that sentence and the whole paper means something else. The model does not say which threshold we are currently near, and it does not hold the material needed to say so. This section looks at why that material is missing.

6.1The Limits the Authors Wrote Down Themselves

The closing part of section 5 lists three limits. Treating the population as a mean leaves out the heterogeneity of real networks. Cognitive autonomy is cut into three sharp compartments rather than left as a continuous quantity. And the feedback by which changed cognition alters subsequent adoption is not yet in the model. With the one already disclosed in section 3, the count is four: this minimal model does not determine how fast the transition unfolds in real time.

To those we add what our own check against the source turned up. The paper points to supplementary material in three places. The full derivation of average competence, the reduction to one dimension, and the fact that the potential is a quartic with two wells are all said to live there. But the arXiv submission carries no supplementary material. The public PDF runs 12 pages, the last of them ending in references, with no ancillary files listed and no later version. Half the core derivation is missing from the public record, so this piece cites nothing from the supplement either.

6.2The Levers the World Is Actually Turning

Outside the model something else stands out. Plenty of organisations are already doing something about AI dependence, and most of them are turning levers that leave the thresholds where they are. The table below maps real-world responses gathered during research onto the model's levers. None of these organisations designed their response with reference to this model; the mapping is a hypothetical reading of how each one lines up.

Real-world response Mapped onto the model Do the thresholds move?
University exams returning to handwritten answer books Access blocked, but only during the exam Everyday pressure is untouched
Oral examinations revived Verification that reveals state. No category for it in the model No
AI tutors designed to safeguard learning Lowers dependency onset No
Internal policies mandating human review and verification Lowers dependency onset No
Organisations that make AI usage rates a performance metric Raises transmission pressure Yes, upward
The AI-free day recommended in consulting circles Raises the return rate Yes, though the effect has never been measured

The real-world cases were gathered from public reporting and policy documents up to 12 September 2026. The right two columns are our reading against the paper's parameters, not a classification the paper published.

Read the table downward and one line remains. Everyone is locking doors or fitting devices that check after the fact, and almost nobody is building a door you can walk back in through. The only case that aims squarely at the return rate turned out to be the recommendation from consulting circles, and even there we found no record of anyone measuring what changed in the organisations that adopted it. On the opposite side sit organisations raising pressure by treating usage itself as an achievement.

Fairness requires a caveat on the third row. The place where a learning-safeguarding tool was confirmed to protect post-withdrawal competence is the high school trial in section 4, and that confirmation came from a condition the researchers designed and deployed themselves. Products of the same class are on the market, but we found no record of anyone taking the tool away from students who used those products and measuring what was left. Even on the levers that do not move the thresholds, the properly measured effects amount to a handful.

The smartphone precedent makes this gap less surprising. Policies banning phones in schools outright are already widespread across many countries, yet studies testing their causal effect on outcomes such as test scores can be counted on one hand. One large district reported a significant increase; a national analysis using a different method found an effect close to zero. If the evidence is this thin for a device that can be shut out completely, then a technology embedded in daily work, where shutting it out is barely possible, is at an earlier stage still.

6.3Three of the Five Have No Statistics at All

To hold the model up against reality, the parameters need numbers in them. We went looking for all five, one at a time, and the table records what came back.

Parameter Public statistics? Nearest available thing
Transmission pressure Proxies in abundance National generative AI adoption surveys
Dependency onset Only indirectly Self-report surveys of knowledge workers
Return rate None One diary study of an enforced four-day break. Ten participants
Recovery rate None Clinical intervention literature on problematic use
Collective reinforcement Effectively none General organisational culture instruments

Results of a search across international, governmental and academic statistics published up to 12 September 2026. Items in the right column are material pointing in a similar direction, not estimates of the parameters.

Start with the row marked abundant. Adoption rates exist in every survey. They just do not agree with one another. Measure the same United States at roughly the same time and one survey returns 28.3% while another returns 54%, because having heard of the tools, having tried them once, having used them last month and using them weekly are all packed into the single word adoption. In international comparison the figures scatter again: an OECD average of 36.8%, 32.7% across the EU, 44% for one specific service in the United States. In Korea, a government survey published in March 2026 found 44.5% of the public reporting experience with generative AI, up 11.2 percentage points on the previous year.

The scatter itself is data backing the caution stated in section 1. Adoption is a consequence of transmission pressure, and it doubles depending on what you decide to count as adoption. No one of those numbers can be laid on the same ruler as the model's thresholds. A separate Pebblous piece takes up what a measurement captures and what it drops: AI use statistics remeasured on an independent corpus.

Dependency onset is in slightly better shape, on weak foundations. A 2025 study cited by this paper collected 936 real work examples from 319 knowledge workers and reported that higher confidence in AI goes with a significant reduction in actually enacting critical thinking. Several studies point the same way. All of them rest on self-report, though, and self-report itself wobbles. In one survey of students, roughly 60% said they used AI while roughly 90% believed their peers did. The 30-point gap is explained as social desirability bias.

For the remaining three there is no seat at all. No statistical agency publishes an AI abandonment rate or a recovery rate. Usage is measured everywhere in the world and state transitions are measured nowhere. The closest record to a return rate is a diary study of ten participants cut off from use for four days, which is a different phenomenon from returning of one's own accord and cannot be generalised from that sample. On recovery, the literature that at least points in a similar direction is the clinical intervention work the paper itself cites in section 5. That material covers how to help an individual, not what share of a population comes back. Hence the note that the right column holds no estimates.

6.4The Gap Has the Same Shape as the Model's Asymmetry

The two tables laid on top of each other show which side the measurement leans to. Everything that is measured is a prevention-side lever: how far it has spread, how dependent people report feeling. Everything unmeasured is a reversal-side lever: how many come back, how many recover, how much of the force propping up autonomy is left. The model says reversal is harder than prevention, and the world's data says the reversal side is not visible at all. Those two sentences having the same shape is the clearest observation this report got out of its research.

The same gap repeats in corporate metrics. Most companies now use AI in at least one function, far fewer track any meaningful indicator, and what they do track are diffusion measures such as volume and breadth of use. No dashboard carries a label separating who is autonomously coupled from who is dependent.

In education the diagnosis has at least been written down in policy language. A 2026 OECD report stated that students using general-purpose chatbots produce better work day to day, but that the advantage disappears or reverses when access is removed in exams, while tools designed with pedagogical purpose sustain the improvement. The finding has the same structure as this paper's. The evidence behind that report is quite likely the same high school trial this paper cites, though, so we do not call it independent confirmation. Either way it stops at diagnosis and never descends to a prescription for building return routes into institutions.

Competence eroding at the individual level has already begun to be documented in other fields. The Pebblous blog covered one such record in the deskilling of endoscopists. This paper works one level up from that, asking what landscape forms when individual skill decline accumulates across a population. The data that would join the two levels is exactly the data that does not exist.

7

Why Pebblous Is Watching

Pebblous diagnoses data and issues quality reports on it. A cognitive model looks a long way from that work, and there is one reason we read this paper to the end: the place where the paper gets stuck has the same shape as the place we meet every day.

7.1The Data That Exists Does Not Measure What You Want to Know

Five parameters carry the conclusions of twelve pages, and not one of them has ever been measured. Transmission pressure looks approximable by an adoption rate, except that the paper itself defines it as a pressure, and actual adoption rates double within the same country on a change of definition alone. At the return rate and the recovery rate there is not even a proxy. The AI logs organisations keep measure usage, not state. With no label separating who is autonomously coupled from who has crossed into dependence, the transition the model predicts has no way of being observed.

Saying there is no data and saying the data that exists does not measure what you want to know are two different statements. The second is the kind of gap we work with constantly. Logs accumulate, in a form that cannot answer the question. Closing that gap starts not with collecting more data but with deciding what to record as an event.

7.2Irreversibility Happens at Both Ends of the Pipeline

One piece already on this blog covers a model trained on junk data that never returns to baseline even after retraining on clean data. That one deals with irreversibility on the model side. This paper says irreversibility of the same shape can appear on the user-population side. Its name is hysteresis, and its cause is that the way back is not the way in.

Side by side, the two pieces make one story. On the side where training data goes in, and on the side where the people using that model are, the cost of reversing damage exceeds the cost of preventing it. If data quality is handled purely as after-the-fact correction, that asymmetry goes missing.

7.3What to Record Next to the Adoption Rate

The metrics companies watch are adoption, usage and hours saved. This model says that identical adoption rates lead to different destinations when the return rates differ. And as sections 5 and 6 showed, the interventions companies actually reach for today cluster almost entirely in the layer that leaves the thresholds alone. Of the three levers that do move the thresholds, the one companies turn deliberately is transmission pressure, and they turn it upward.

The practical suggestion therefore comes down to one line. How open the way back is has to be measured, and measuring it requires a record. How many pieces of work were finished without AI, what share of drafts a person wrote first, how output held up with the tool withdrawn. These three items, absent from every dashboard today, are the closest observable stand-ins for a return rate. All three can begin as one extra label on records that already exist, with no new system.

7.4The Instrument Has Already Been Built

The oral examinations catch the eye again here. The table in section 6 filed them under verification, though an oral exam is really a procedure for finding out what remains once the tool is gone. That is exactly the paper's definition of cognitive competence. Schools revived the format to catch cheating, and what they ended up building is a state classifier. Nobody keeps the results as data. This paragraph is our reading rather than anything the paper says.

Pebblous works on the side of making data that AI can use. This paper puts the question from the opposite side: on the side of the people using AI, what should be recorded as data? Turning adoption logs into longitudinal records carrying state labels, writing exit and re-entry down as events, narrowing the distance between self-report and behavioural record. All of it is a data quality problem.

This piece connects to what Pebblous does not because the paper proves our product is necessary, but because the place the paper gets stuck is the place we have to answer every day. The equations, threshold values and verbatim quotations in the body were checked directly against the full text of the public arXiv version, and the figures from the cited empirical work were confirmed in each original abstract. Anything that failed to check out was left out, and the table mapping real-world cases onto the model's parameters is marked, in place, as our reading. Please read the assessment of the research and our own positioning as separate things. Thank you for reading this far.

R

References

The evidence behind this piece comes in three strands. The model's equations, threshold values and verbatim quotations were checked against a downloaded full text of the public arXiv version of reference 1. The empirical figures were confirmed in each original paper's abstract, and the body notes in each case that they were not measured by the paper this report is about. Some statistics and policy documents could not be opened at source and rest on search summaries; those entries are marked as such.

The Backbone of This Report (Full Text Checked)

  • 1.R. Solé, G. Ruffini, F. Castaldo, M. Tuccio, L. F. Seoane, M. de Domenico, S. F. Elena, D. C. Krakauer, M. Levin. "Large-Language Models as a Cognitive Virus." arXiv: 2609.03344, v1 submitted 3 September 2026, 12 pages and 3 figures, physics.soc-ph. Equations (1) through (20), the threshold values 0.50, 0.40 and approximately 0.420, the six interventions in Table 1, the average cognitive competence values 0.425 and 0.617, and all ten verbatim sentences quoted in the body were confirmed against the full text of this version. Reference markers inside the quotations were removed for readability. No journal version exists; the licence is CC BY-NC-SA 4.0. The supplementary material the body points to in three places is not included in the public submission.

Empirical Work the Paper Cites (Original Abstracts Checked)

  • 2.H. Bastani, O. Bastani, A. Sungu, H. Ge, Ö. Kabakcı, R. Mariman. "Generative AI without guardrails can harm learning: Evidence from high school mathematics." PNAS 122, e2422633122 (2025). doi:10.1073/pnas.2422633122. Field experiment with nearly a thousand high school students. The 48% and 127% improvements, the 17% reduction after access was removed, and the word "crutch" are all printed in the abstract.
  • 3.S. S. H. Wong, S. X. Qiu. "Think First, ChatGPT Later: Guiding Human–AI Collaboration for Learning Gains in Independent Human Creativity." Educational Psychology Review 38, 45 (2026). doi:10.1007/s10648-026-10118-7. The 196 university students, the three groups and the reversal on the second task are described in the abstract.
  • 4.H.-P. Lee, A. Sarkar, L. Tankelevitch, I. Drosos, S. Rintel, R. Banks, N. Wilson. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pp. 1–22 (2025). 319 knowledge workers, 936 real work examples. Reports a negative relationship between confidence in AI and the enactment of critical thinking.
  • 5.N. Kosmyna et al. arXiv:2506.08872 (2025), with the project site brainonllm.com. A preprint, not yet peer reviewed. The request quoted in the body comes from the FAQ page of the project site. Two figures that circulated widely in press coverage could not be confirmed in the abstract or on that site and were left out of the body.
  • 6.A. R. Doshi, O. P. Hauser. "Generative AI enhances individual creativity but reduces the collective diversity of novel content." Science Advances 10, eadn5290 (2024). doi:10.1126/sciadv.adn5290.
  • 7.D. Centola, J. Becker, D. Brackbill, A. Baronchelli. Science 360, 1116 (2018). The experiment the paper cites as a precedent for tipping in social convention.
  • 8.F. Courchamp, L. Berec, J. Gascoigne. Allee Effects in Ecology and Conservation (Oxford University Press, 2008). The ecological source the paper gives for the collective reinforcement term.
  • 9.Four works the body borrows concepts from. E. F. Risko, S. J. Gilbert. Trends in Cognitive Sciences 20, 676 (2016), for the definition of cognitive offloading. · D. C. Krakauer, Nautilus issue 40 (2016), for the distinction between complementary and competitive cognitive artifacts; he is a co-author of the paper. · S. Dehaene, L. Cohen, J. Morais, R. Kolinsky. Nature Reviews Neuroscience 16, 234 (2015), cited in the reading-and-writing passage in 1.1 as the basis for literacy reorganising preexisting neural circuits. · H.-Y. Liao, C.-H. Ko, C.-F. Yen. Biomedical Journal 49, 100998 (2026), the problematic-LLM-use literature behind the clinical intervention remark in 5.3. Bibliographic details for these four were transcribed from reference 1's own bibliography; we did not open each original.

Statistics and Policy Documents

  • 10.Stanford HAI, AI Index 2026 (published April 2026, 2025 data). Global adoption 53%, United States 28.3%, organisations 88%. The contrasting 54% figure for the United States is the Federal Reserve Bank of St. Louis's own tracker value for August 2025.
  • 11.OECD member-country individual surveys (2025 data) averaging 36.8%; Eurostat, ages 16–74, 32.7% (2025, the first year with a generative AI question); Pew Research Center survey of US adults, 44% for ChatGPT on its own and 49% for chatbots as a whole (February 2026, 5,119 respondents). None of these three source pages could be opened directly and all rest on search summaries. Check each institution's original before citing.
  • 12.OECD, Digital Education Outlook 2026 (January 2026). The source for the statement that the advantage disappears or reverses when AI access is removed. Its evidence plausibly overlaps with reference 2, so the body does not call it independent confirmation.
  • 13.Korean Ministry of Science and ICT, 2025 Survey on Internet Usage (published 31 March 2026, 50,750 household members). Generative AI experience rate 44.5%, up 11.2 percentage points year on year.
  • 14.Diary study of enforced disconnection by researchers at KAIST, arXiv:2603.26099 (2026). Ten participants, four days. A qualitative study of imposed disconnection rather than voluntary abandonment.
  • 15.A 2026 CHI-family study reporting the 30-point gap between students' self-reported AI use and their perception of peer use, interpreted as social desirability bias. Reached via search summary; the original bibliographic record could not be confirmed.
  • 16.Two strands of NBER working papers on the effects of school smartphone policies: one reporting a significant increase in test scores in a large district, and a 2026 national analysis reporting an effect close to zero for the locked-pouch approach. Both are stated in the body.
  • 17.The AI-free day is a prescription that came out of consulting circles. We could not confirm the original report title and publication date at a primary source, so the body does not attribute it to any named firm's report. The state of corporate AI metric tracking likewise comes via consulting aggregation.

Adjacent Pebblous Pieces