Executive Summary
A researcher at the Oxford Internet Institute collected every public version of the safety frameworks published by the twelve developers that put one out after the 2024 AI Seoul Summit, then put a number on how legible their revisions are. The yardstick is called the silent revision rate, and the thing it measures is not the document. It measures the account the developer published about its own revision. Across the eight version pairs that came with an account, the study traces 383 material changes, and for two-thirds of them a reader working from the account alone could not tell what had changed or which way. The changes the account never touches at all are a smaller share, close to half. The band between those two figures belongs to changes where the account named the kind of commitment involved and stopped there.
The established finding is a single one. Silence tracks the direction of the change. Among changes that came with an account, those that loosened a commitment were silent at 0.75 and those that tightened one at 0.50. Seven of the eight pairs show the same slope inside the pair. On the question of form, the signal that survives is enumeration. Developers who wrote longer accounts did not go silent any less often. Developers who broke their revisions into discrete items did.
California and the European Union already attach duties to framework revision. Both duties point at why the framework changed rather than at what changed in it. A justification explains a reason, an enumeration states a change, and only the second one lets anyone check the revision later. Korean companies face the same drafting question right now, since the AI Framework Act takes effect on 22 January 2026 and safety and reliability documentation is being written this quarter. Publishing a document and leaving its revisions open to audit were separate jobs from the start.
Three numbers carry this measurement
The first number measures the distance between what a developer said and what it did to its own commitments. The second shows that the distance depends on which way the change went. The third shows that closing the distance depends on enumeration, not on word count. All three sit on different denominators, so the denominator travels with the number.
Source: Zhu, L. Y. (2026), Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers, arXiv:2609.08789v1 [cs.CY], 8 September 2026. A single-author preprint, under workshop review.
67%
Material changes a reader cannot identify from the developer's account
257 of 383 material changes across the eight version pairs that published an account. The share the account never mentions at all is 53%
0.75 : 0.50
Silence rate for loosening changes against tightening ones
135 of 181 on the loosened side (weakened, removed, relocated), 25 of 50 on the tightened side. Odds ratio 2.93, Fisher's exact test p=0.002
−0.58 : −0.17
Correlation with silence: item count against word count
Spearman rank correlation across the eight pairs. Breaking the account into items lowered silence. Writing more words did not
What the silent revision rate measures
Frontier AI developers promised at the 2024 AI Seoul Summit to publish safety frameworks. Twelve of them went on to publish one, and those twelve are the entire corpus of this study. So the documents exist. What happened after that is the subject here. The documents kept being revised, and the revisions did not become public in the way the documents did. Louis Yiven Zhu of the Oxford Internet Institute posted Silent Revision to arXiv on 8 September 2026, a single-author preprint and the first attempt to put a number on that gap. Its opening sentence states the problem in one line. “A standard that can be revised without anyone noticing is not a standard that anyone can be held to.”
The target of this measurement comes first. The silent revision rate does not count how much a document changed. It counts the share of material changes to a framework's commitments that the developer's own published account does not identify. The numerator therefore describes a property of the developer's statement about its document rather than a property of the document. The same document paired with a more detailed account scores lower.
1.1What counts as material
The unit is neither a sentence nor a clause. It is a commitment, which the paper defines as “a statement in which the provider commits itself to a practice at any strength from must to may”. A policy document written in the present tense is using the standing-commitment register, so those statements count at the strongest rung. Counted this way, the study compared 710 commitment instances across version boundaries in the twelve traced pairs.
Whether a change to one of those commitments is material turns on six dimensions: scope of application, trigger and threshold, actor, strength of obligation, extent of disclosure, and the consequence a result obligates. Movement along any one of them makes the change material. The obligation-strength dimension follows the legal convention that separates mandatory from permissive language, and one xAI version pair shows how far that ladder can drop on a single word.
| Earlier version | Later version |
|---|---|
| “If xAI learned of an imminent threat of a significantly harmful event, including loss of control, we would take steps to stop or prevent that event, including potentially the following steps:” | “Should it happen that xAI learns of an imminent threat of a significantly harmful event, including loss of control, we may take steps such as the following to stop or prevent that event:” |
XAI-2-035, Appendix F. One modal verb moves from would to may, the obligation-strength dimension moves with it, and the outcome is coded as a weakening. xAI's four pairs published no revision account, so no disclosure code attaches to this change.
Which way a change moved is recorded separately, in six outcome codes: kept, strengthened, weakened, removed, relocated, added. When a single commitment moves both ways at once, the codebook sends it to weakened by rule and flags it as mixed. Google DeepMind's pre-deployment review of the safety case is one of those. The set of deployments that must pass through the review widened from general availability to external deployments generally, while the body that performs the review went from a named corporate governance body to an unnamed governance function. The first move tightens and the second loosens, and the rule sends the case to the loosening side. The weakened share reported later therefore contains cases of this shape.
1.2Announcement comes in grades
Every material change carries one of four codes describing how the developer's account handled it. If a reader of the account can tell that this commitment moved in this direction, the change is announced. If the account names the commitment or its class without giving the direction or the content, it is partly announced. If an account exists and does not point at the change even at the level of class, it is silent. If the developer published no account at all, the code is no account, and those changes stay out of the denominator, because there is no account to measure them against.
▲ The four disclosure codes and the boundary between the two standards. Rebuilt by Pebblous from Table 1 of arXiv:2609.08789v1.
That distinction also explains why the strict 0.67 and the lenient 0.53 are both reported. Rather than pick one, the paper publishes both and says it does so to expose the judgement inside the measurement. Put partly announced on the silent side and the rate is 0.67; put it on the announced side and it is 0.53. If those 14 points are hard to picture, the clearest case in the corpus belongs to Microsoft.
The earlier version promised deeper capability assessments at least once every six months. The later version promises them when there are material changes to the deployed model's risk profile. A fixed cadence became an event trigger. The changelog did handle the change: “Adjusting the cadence by which we repeat deeper capability assessment to align with emerging industry standards.” It says the cadence was adjusted and does not say in which direction. Why partly announced is hard to count as announced sits entirely in that one case.
MSFT-1-015, Appendix F. The adjudication note reads “Fixed six-month minimum becomes event-triggered; change log names the cadence without direction”.
Google DeepMind's revision has the same shape. In the capability threshold for machine learning R&D acceleration, “substantially accelerating (e.g. 2x) from 2020–2024 rates” became “substantially accelerating from historical rates”. The quantitative anchor of 2x went out, and the preceding “can or has been used” lost its first half, so only actual use remains. The announcement reported that the capability level definitions had been “sharpened”. Which definition was sharpened, and which way, it did not report.
Two-thirds of what, exactly
Numbers like this travel by shedding their denominator. This one narrows three times. Twelve developers are in the corpus, but Amazon, Cohere, G42, NVIDIA and Magic carry a single labelled version each, so there is no pair to compare. Magic re-posted its file under the same version number, which leaves no label to separate a before from an after. The seven developers with two or more versions yield nineteen consecutive pairs, of which twelve were actually traced and eight produced a rate. The 67% comes from those eight.
| Layer | Count | What drops out |
|---|---|---|
| Developers in the corpus | 12 | Five carry a single labelled version, so there is nothing to compare |
| Consecutive version pairs | 19 | Anthropic 8, xAI 4, DeepMind 3, OpenAI / Meta / Microsoft / Naver 1 each |
| Pairs traced | 12 | Five redlined pairs are 0 by construction and are not traced |
| Pairs with a rate | 8 | xAI's four pairs published no account, so no rate is defined |
Rebuilt by Pebblous from Appendices J and N. Two untraced pairs, one at Anthropic and one at DeepMind, also drop out on the way from nineteen to twelve.
With a sample of eight, the intervals belong in the same sentence as the point estimates. The 95% confidence interval on the strict 0.67 runs from 0.62 to 0.72, and the lenient 0.53 runs from 0.48 to 0.58. No single unusual pair is dragging the mean either. All eight pairs with an account clear 0.55 on the strict standard, and grouping the rated pairs by developer leaves the values between 0.61 at OpenAI and 0.81 at Microsoft.
The zero that sits against redlined pairs deserves care. That zero is a definition rather than a result. A revision published as a whole redline exposes every textual change, so nothing in it can be silent, and the paper concedes in §7 that the construction is partly circular. A redline shows how the text differs and says nothing about which differences are material or which way they point. Beyond that, the developer picks the form of the notice. Anthropic redlined most of its revisions and did not redline its largest one.
The value falls when the same changes are regrouped under a different unit. Counting document sections rather than commitments gives 0.49 across 99 units, and a majority rule that calls a section silent only when most of its changes are silent gives 0.66. The paper reads the two as different experiences: the commitment-level rate is the value a reader checking one commitment meets, and the section-level rate is the value a reader working through the whole document meets. Either way of counting leaves more than half of the changes unexplained.
To count version pairs at all you need the versions in hand first, and that part was not easy. The manifest runs to 52 rows, 43 of them frameworks and 9 of them companion documents. Thirty-five files came from the developer directly, fourteen from the Internet Archive and one from a portal that had closed off access. Three versions known to exist were never recovered. Each file is stored as published, with a plain-text extraction, a hash recomputed from disk and a record of where it came from. The cost anyone pays before they can compare two revisions is written out in that manifest.
2.1Nine documents that changed without a version bump
Assembling the corpus turned up a finding of its own. Nine files sit at the same address under the same version label with different text, re-uploads of a document that had already been published. The paper keeps them as separate rows instead of folding them together as copies of one document, since a reader who downloaded before and a reader who downloaded after end up holding different documents under the same name. Eight of the nine were judged non-material, covering typesetting, links and corrected wording. One Meta file carried two material changes.
- In the sentence that sets the scope of the whole framework, most advanced and match or were deleted. Models that match frontier capability now fall outside the framework.
- The trigger for withholding release moved from a catastrophic outcome to a threat scenario.
The version number stayed where it was and no new identifier appeared. This sits one layer away from a problem Pebblous covered earlier in No provider could show that the evaluated model is the one answering you, and the shape is the same. There, a served model changed while its name stayed put. Here, the document that governs that model changes while its name stays put. In both layers the only fixed point available to an outside observer is the name, and the name is not holding still.
2.2A language model did the first-pass coding
§7 lists the study's limitations in order of how much they matter, and the coding procedure comes first. The first pass over 710 commitments across version boundaries was not performed by a person. It was performed by an agentic language model system from the Claude family, working from a frozen codebook as its only instruction. Each row carries verbatim quotations from both versions, a rationale and a confidence grade, and each material change carries either a quotation pulled from the revision account or a record that none was found. The quotation fields were validated against the corpus text in 2,130 checks, and what that validation guarantees stops at the fidelity of the quotations.
The author then adjudicated 244 of the 710 rows by hand. Those adjudications changed 4 outcome codes and 54 disclosure codes, and 46 of the 54 moved from silent to partly announced. They moved away from the side that props up the strict 67% and toward the side that lowers it. The paper adds that the direction of that movement matters, since the adjudicator designed the codebook and holds the hypothesis. Inter-coder reliability was not computed at the time of submission, and that fact sits at the top of the limitations list.
[Interpretation] An AI system doing the first-pass coding of revisions to AI safety documents reads easily as a weakness of the study. It is also exactly what the paper argues about. Because the record states that adjudication moved 46 cases against the author's own conclusion, a reader can weigh what kind of number this is. The claim that auditability comes from writing down what changed and which way it went is being practised by the paper's own methods section.
Silence tracks direction
Among changes that came with an account, 135 of the 181 on the loosening side were silent, and 25 of the 50 on the tightening side were. The odds ratio is 2.93, and Fisher's exact test gives p=0.002. Silence is not evenly distributed, and the slope runs with the direction of the change: the paper's established label reaches that far and no further.
Those 181 loosened changes are not only reductions in the strength of a commitment. The figure aggregates three outcomes: commitments weakened, commitments removed outright, and commitments relocated to a companion document. Inside that aggregate, removals were silent most often at 35 of 42, and all four relocations were silent. Add the 181 to the 152 additions and the 50 tightened changes and you get the 383 material changes on the eight pairs with an account. Removals and relocations are a share of the first bucket rather than a fourth one beside it, so lining up four values and summing them inflates the denominator.
▲ Silent revision rate by direction of change, strict standard. The odds ratio between the loosened and tightened sides is 2.93, with Fisher's exact test at p=0.002.
The overall volume tilts the same way. Of the 299 traced material changes left once additions are set aside, 229 weaken, remove or relocate a commitment, a share of 77%. The number flips if the denominator goes missing here. With the 174 additions folded back in for a total of 473, the same 229 becomes 48%. The 77% in the abstract stands on the 299 that exclude additions. No single developer is producing the tilt either. Recomputed with any one provider left out, the share stays between 0.70 and 0.79. Narrowing the bucket does not move it much either: leave out the relocations to a companion document and it is 0.76, and leave out the cases that moved both ways and landed on weakened by rule and it is 0.73.
The same slope shows up inside individual pairs. In seven of the eight pairs with an account, loosening changes were silent more often than tightening ones, by margins running from 0.03 to 0.67. The exception is Anthropic's revision from version 1.0 to 2.0, where the tightening changes were silent by 0.13 more.
3.1What sat in the places that went silent
Rates alone make it hard to feel how specific this finding is. Appendix F prints eighteen cases with the original and revised text side by side, and three of them follow. The first two come from Anthropic's revision from version 1.0 to 2.0, and the account points at none of the three.
| The commitment that went away | What took its place |
|---|---|
| “we commit to pause the scaling and/or delay the deployment of new models whenever our scaling ability outstrips our ability to comply with the safety procedures” | “we will act promptly to reduce interim risk to acceptable levels … The CEO and Responsible Scaling Officer may approve the use of interim measures that provide the same level of assurance as the relevant ASL-3 Standard” |
| “We will manage our plans and finances to support a pause in model training if one proves necessary” | “We will set expectations with internal stakeholders about the potential for such pauses.” |
| “We will also continue to enable external research and government access for model releases” (OpenAI, twelve-item changelog) | Sentence removed. None of the twelve items points at the removal. |
ANT-1-001, ANT-1-052 and OAI-1-040, Appendix F, quoted verbatim and abridged where marked. The appendix flags the first and the third as headline examples, the third of them for the third-party involvement category.
Third-party involvement is the smallest of the nine categories: 21 material changes, 18 of them with an account, silent at 0.67. The commitment that went missing inside that small category is the one keeping external researchers and government able to reach the model. In a piece about auditability, that single case weighs differently from the others, because a commitment to keep the channel for outside verification open was itself closed, and nobody outside was told.
Some commitments look gone and are not. In Anthropic's revision from version 2.2 to 3.0, the promise to give staff a pathway for raising concerns about the policy and about the risk levels of models dropped out of the framework, and it survives in a companion document, the noncompliance policy. The paper codes a case like that as relocated rather than removed. Doing so requires collecting the companion documents alongside the frameworks, which is why nine companions sit in the corpus. All four relocations are silent. Nothing anywhere records that the commitment moved, so a reader has no way to tell a deletion from a sideways move.
3.2Governance clauses came out with the highest silent revision rate
If silence tracks direction, the next question is whether it tracks subject matter. The paper sorts commitments into nine categories. Six of them are evidentiary, covering how a developer will evidence risk or capability, and governance, security and mitigation stand outside that group as three further strata. Here is how the values split.
| Category | Material | With account | Strict | Lenient |
|---|---|---|---|---|
| Scope of evaluation | 29 | 27 | 0.52 | 0.33 |
| Trigger and threshold | 68 | 59 | 0.66 | 0.42 |
| Method | 64 | 56 | 0.73 | 0.64 |
| Third-party involvement | 21 | 18 | 0.67 | 0.61 |
| Disclosure | 46 | 41 | 0.68 | 0.46 |
| Consequence a result obligates | 59 | 56 | 0.57 | 0.43 |
| Governance | 84 | 64 | 0.77 | 0.73 |
| Security | 35 | 30 | 0.63 | 0.47 |
| Mitigation | 67 | 32 | 0.72 | 0.56 |
Table 6, Appendix H. The material column counts changes across all twelve traced pairs, and the account column counts those falling on pairs that published one, which is the denominator of the rates. The nine denominators sum to 383 and the material column sums to 473. The first six rows are the evidentiary categories and the last three are the further strata.
[Fact] The appendix states that no test is applied across categories. Nothing in this table supports saying that one category is significantly more silent than another. The intervals are wide as well. The strict interval for third-party involvement runs from 0.44 to 0.84. Pooled by group, the six evidentiary categories come to 0.65 strict and 0.48 lenient, and the remaining three strata to 0.72 and 0.63.
[Interpretation] Reading down the columns, the governance row catches the eye twice. Governance carries the most material changes, 84 of them, and also the highest strict rate at 0.77. Its lenient rate of 0.73 leaves almost no room for the partly announced band. The provisions that set out who decides changed more often than any other kind and were explained least often. Anthropic's financial pause-readiness commitment and its staff reporting pathway both came from this category. The opposite end of the column is scope of evaluation, whose strict 0.52 is the lowest of the nine.
3.3The headline change was announced
At the point where these numbers invite a reading about concealment, the paper drives a nail the other way. The account Anthropic published with its revision from version 2.2 to 3.0 ran to 14,274 words and explained at length why it was withdrawing unilateral pause commitments. That largest change was coded as announced. The silence in that same revision fell on everything else. The commitment to delete model weights where security safeguards cannot be met disappeared, and the post addressed the class of hard unilateral commitments it was restructuring without pointing at this commitment. Of that revision's 86 material changes, 51 are silent even on the lenient reading.
[Fact] §6 states that the study does not infer intent. The sentence reads “we do not infer intent, since the people who write changelogs may simply champion the new instruments”. Silence is a name for what an outsider can see, not a name for a motive.
[Interpretation] The result survives without the motive. The more natural it feels to describe the instrument you have just added, the more easily the instrument you have just withdrawn passes without description. Organisational research offers two accounts of this shape. Structural secrecy, where the division of labour leaves no department able to see the whole, predicts uneven silence and stops there. Decoupling, where the formal structure and the actual activity come apart, predicts that unflattering changes will be less visible. Silence that tracks direction fits the second prediction.
Not length, enumeration
The number in this study that looks most quotable comes from here. The three pairs whose developers published a narrative announcement run at 0.74, and the five that published an itemised changelog run at 0.63. The lesson seems to write itself: stop summarising in prose and break the revision into items. The paper declines to present the contrast as an established result. On the permutation test that respects the nesting of changes within pairs it gives p=0.071, and the difference in per-pair means gives p=0.089. A test that treats each change as an independent observation gives p=0.041, and the paper rejects that test itself as overstating precision. On the lenient standard the difference vanishes altogether, at an odds ratio of 1.19 and p=0.46.
§6 sorts the study's own results into two grades. The direction result is established and the form result is suggestive. It adds that any effect of form, should one exist, lives inside partial announcement. Developers who write itemised changelogs go silent as well. They shift from not pointing at a change at all toward naming its class and leaving it there.
On form, the clear signal is a pair of correlation coefficients. The Spearman correlation between the word count of a revision account and the silent revision rate is −0.17, near enough to nothing. The correlation with the number of discrete items the developer published is −0.58. The contrast is visible at a glance once the eight pairs are laid out by length.
| Version pair | Form | Words | Items | Material | Silent rate |
|---|---|---|---|---|---|
| Anthropic 2.2 → 3.0 | Narrative | 14,274 | 3 | 86 | 0.73 |
| Anthropic 1.0 → 2.0 | Itemised | 4,710 | 11 | 76 | 0.57 |
| Naver 2024 → 2.0 | Narrative | 2,139 | 3 | 22 | 0.77 |
| OpenAI Beta → v2 | Itemised | 1,971 | 12 | 41 | 0.61 |
| Meta 1.1 → 2 | Itemised | 1,353 | 7 | 80 | 0.69 |
| Google DeepMind 2.0 → 3.0 | Narrative | 863 | 3 | 30 | 0.73 |
| Google DeepMind 3.0 → 3.1 | Itemised | 358 | 6 | 27 | 0.56 |
| Microsoft v1 → 2026 | Itemised | 291 | 6 | 21 | 0.81 |
All eight pairs from Table 3. The rate is the strict one, and the word and item columns describe the account the developer published for that revision. The caption to Table 1 records a limitation on the Naver pair. Its 2024 version exists publicly only as an English summary page, so that pair over-counts additions.
With the longest account set next to the shortest, the length hypothesis collapses. Anthropic's 14,274-word narrative covered 86 material changes and scored 0.73. Microsoft's 291-word changelog covered 21 and scored 0.81, the highest in the corpus. Microsoft managed more silence with roughly one forty-ninth of the words. At the other end sits OpenAI, whose 1,971 words carried twelve items over 41 changes and produced the lowest rate in the corpus at 0.61. The paper's sentence: “what lowers silence is enumeration at the level of the change”.
4.1Enumeration does not prevent loosening, it makes it visible
The distinction is easiest to see in Meta's revision. The earlier version said a model reaching the high risk threshold would not be released. The later version says it will be deployed once sufficient mitigations are defined, implemented and validated. That is a plain weakening. And the changelog recorded it in one accurate line: “High threshold measure changed from ‘Do not release’ to ‘Deploy with mitigations.’” That single line shows what an enumeration duty produces. It does not stop a commitment from loosening. It leaves the loosening checkable from outside.
The other half of the same contrast turns up at OpenAI. Its twelve-item changelog stated in item four that persuasion was moving out of the framework, and the same changelog said nothing about the narrowing of a commitment to run evaluations continually. In the later version that commitment applies only to deployments with a plausible chance of crossing a threshold, and a footnote takes models distilled, fine-tuned or quantized from a model already found not to cross a high capability threshold out of the ordinary scope of additional measures. An enumeration makes visible as much as it enumerates.
META-1-032 and OAI-1-021, Appendix F. The appendix flags the first as a headline example, and the adjudication note on the second records that it is “absent from the twelve-item changelog”.
The law asks why and does not ask what
Framework revision is not unregulated. California's Transparency in Frontier Artificial Intelligence Act requires large developers to publish their framework, and to publish the modified framework together with a justification within thirty days of a material modification. The European Union's General-Purpose AI Code of Practice requires signatories to update their safety and security framework at least once a year and to give each update to the AI Office within five business days. Neither regime leaves revision alone.
The paper's aim is the artefact those duties point at. California asks for a reason. Europe asks for the updated document. Neither asks anyone to enumerate what changed, commitment by commitment, and in which direction. California's statute does not define material modification either. The abstract puts it plainly: “The statutory remedy therefore exists and specifies the wrong artefact.” The sentence that follows is the spine of this whole report. “A justification explains why a framework changed, an enumeration states what changed, and only the latter makes revision auditable.”
Split around 1 January 2026, when TFAIA took effect, the three pairs closed before that date run at 0.61 and the five closed after it run at 0.71, with the weakening share holding at 0.78 and 0.76. [Fact] The paper states that three pairs a side carry no causal claim, and goes no further than saying the justification duty arrived alongside no reduction in silence. A sentence claiming the regulation failed cannot come out of this sample.
5.1An enumeration duty already exists in another field's statute
§7 concedes that the study has no comparison class. Nothing in it says whether two-thirds is high next to other regulated standards. So the enumeration duty arrives looking like a new proposal. The thing itself is already written into law elsewhere. In the United States, the regulation governing changes to an approved drug application, 21 CFR 314.70, puts it at (a)(6):
“A supplement or annual report must include a list of all changes contained in the supplement or annual report.”
21 CFR 314.70(a)(6)
The same regulation also sorts changes by direction. A labelling change that adds or strengthens a contraindication, warning, precaution or adverse reaction can be put into effect as soon as the supplement is received ((c)(6)(iii)(A)), while a change with substantial potential to have an adverse effect on safety or effectiveness goes through a prior approval supplement ((b)(1)). Changes that tighten toward safety move quickly, and changes running the other way carry the heavier path.
[Interpretation] We take that contrast to be the most usable thing in this report. The asymmetry this study found by measurement is that loosening changes more often stop being visible. Drug regulation built that asymmetry into its design in advance and wrote direction into the rule, while the two frontier AI regimes attach one duty to material modification in general and draw no line by direction. It also means nobody has to invent an enumeration duty. The work looks more like carrying over a regulatory design that has run in another field for decades.
On version retention, ISO 9001:2015 is the counterpart. Its documented-information clauses require obsolete versions to be identified and retained, or marked as withdrawn, so that a superseded version is not used by mistake. Neither the California statute nor the European Code carries an equivalent duty to keep earlier versions. Software practice points the same way, having long ago separated the commit message that records why from the changelog that lists what changed by category. Safety frameworks have not made that separation yet.
What happens without a retention duty has a specimen inside the corpus. Anthropic's Frontier Compliance Framework explains the changes across four versions in an itemised changelog, and one of its sections follows the California language almost word for word in promising a changelog with a justification within thirty days of a material update. That is drafting one step ahead of the statute. At the time the corpus was assembled, though, the three earlier versions had come down from the hosting portal. You can read what the document says it changed. You cannot check it against what the earlier text said.
5.2The Naver case in Appendix F was an announced change
Naver's AI Safety Framework is the one Korean entry in the corpus. The version published at the 2024 AI Seoul Summit and the 2.0 version published on 8 July 2026 at the Seoul Forum on AI Safety Science form a pair. That pair runs at 0.77 strict and 0.45 lenient. A limitation travels with those figures. As the caption to Table 1 records, Naver's 2024 version exists publicly only as an English summary page, so the pair over-counts additions.
The more interesting part is that the change Appendix F prints for Naver falls on the announced side. The earlier version carried a quantitative trigger: evaluate large language models quarterly, and evaluate sooner than the three-month mark when performance appears to have increased six times over. Version 2.0 drops that trigger, and the press release named the replacement, “replaces a single performance-based criterion with separate criteria for context, use case and impact”. Using Naver as an example of a Korean company going silent gets the case backwards.
As far as we could confirm, no Korean company besides Naver has published a safety framework of this kind. That is not proof of absence. It does mean that for most Korean companies the issue is not yet a question about revision. It is a design question about a document being written for the first time. The AI Framework Act takes effect on 22 January 2026, and safety and reliability documentation is being drafted at many companies at once right now. Building a place for revision history into that document from the start and bolting version control on afterwards carry different costs. Pebblous covered the Korean regulatory side in The AI Safety Dossier Is Really a Data Lineage Record, and the question of who verifies such a document in Illinois Becomes the First State to Put Frontier AI Under Annual Outside Audit.
Industry-wide transparency scores move in the same direction. Stanford CRFM's Foundation Model Transparency Index scored an average of 40.69 in its December 2025 edition, down from 58 in 2024, with training data, training compute, and downstream use and impact named as the most opaque areas. It measures something other than what this paper measures, so the two cannot be measured with the same yardstick. Both measurements still point the same way on the visibility of change after the fact.
Why Pebblous cares
Pebblous diagnoses data and issues quality reports on it. The revision history of a policy document sounds at first like somebody else's field. We read this paper to the end because the question it sets has the same shape as the question we put to a dataset.
6.1A clean snapshot and an auditable path are different properties
Diagnosing a dataset means looking at movement between states, not at the current state. Which records arrived when, what dropped out, which labels were applied again, and where each of those changes was written down. Lineage sits among the conditions for AI-Ready Data for that reason. A snapshot being clean and the path to that snapshot being auditable are two different properties, and without the second the grounds for trusting the first grow thin. This paper applies the same distinction to policy documents and puts a number on it. Publishing a document and leaving its revisions open to comparison are separate jobs, and 383 changes are the demonstration.
6.2Loosening gets written down less on the data side too
The asymmetry where loosened commitments went silent far more often than tightened ones is a familiar shape in dataset management. Adding data lands in the release notes. Relaxing an exclusion rule usually does not. A decision to loosen a filter or lower a quality threshold passes quietly, and the mark it leaves in a model's internal representations surfaces several generations later. The paper's observation that the trigger-and-threshold category keeps losing its quantitative anchors describes the same drift. A factor of 2x becomes historical rates, a ceiling on the dishonesty rate for a named benchmark becomes a criterion incorporating a margin of security, and the definition of frontier capability leaves the document altogether. Thresholds drifting in a policy document and thresholds drifting in a data pipeline are hard to tell apart. The related problem of labels going stale when the rule behind them changes is one Pebblous took up separately in Provenance Tracking Cut Rule-Revision Relabeling to 14.7%.
6.3Four design rules for teams drafting a framework now
The founder of the Midas Project, a nonprofit watchdog that has tracked unannounced edits to AI policy documents, once described the work this way: “If every AI company had a change log, this work would be unnecessary… That would be the ultimate transparency.” For a team drafting a safety framework or a set of safety and reliability documents this quarter, the study reads less as a warning and more as a set of design rules.
- Treat publishing the document and keeping its revisions auditable as two separate pieces of work. The second is hard to add later.
- Rather than writing a longer account, enumerate change by change. The number that correlated with silence was the item count, not the word count.
- Record the direction of each item, whether it tightened or loosened. Changes that loosen are the ones most likely to fall out of the record.
- Keep the previous version when you publish a new one. With no original to compare against, a changelog is not evidence either.
Those four carry over to a dataset release policy without modification. Writing down which rule changed, when, and in which direction, and keeping the earlier version alongside it, is the same object a data catalogue calls lineage.
6.4Records, not declarations
Pebblous has written once about there being no way to confirm that the model being served is the model in the documentation, and once about the safety and reliability documentation the AI Framework Act asks for amounting to a data lineage record. This piece is the third, joining those two along the axis of version control. All three share one view. The real question in verification is not what an organisation declared it would do. It is whether what it did survives as a record. There is now one more yardstick for the distance between published and auditable, and putting that yardstick in the hands of the people writing these documents is what this report is for.
The figures and verbatim quotations in this article were checked directly against the full text and appendices of the public arXiv version. The paper's body and its appendices disagree in two places, and we followed the appendices in both. Statements about developer documents and statutes stay inside what the paper collected and cited. Results the paper marks as established and results it marks as suggestive are kept apart in the body, and we would ask that the distinction travel with any quotation. Thank you for reading this far.
References
The measurements and verbatim quotations in the body were checked against the downloaded full text of reference 1, in both its body and its appendices. Statements about developer documents and statutory language do not go beyond what that paper cited. Three sections of the US Code of Federal Regulations, the Foundation Model Transparency Index, and the wording of Naver's 2.0 announcement were checked separately in the originals, and those items are marked. Prior work carried over from the paper's own citations is marked as well.
Primary source (full text checked)
- 1.Zhu, L. Y. (2026). Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers. arXiv:2609.08789v1 [cs.CY], submitted 8 September 2026. Single author, Oxford Internet Institute; the Comments field states that the paper is under review at a NeurIPS 2026 workshop (AISciK). Paper and data CC BY 4.0, code MIT. The 383, 257 and 203 counts, the rates by direction, the eight pairs of Table 3 and the verbatim comparisons of Appendix F were all confirmed in this version.
- 2.Corpus and code address given by the same paper: github.com/louisyzhu/frontier-safety-framework-corpus. This report did not open the repository to check its contents.
Prior measurement and theory
- 3.Wan, A. et al. (2025). The 2025 Foundation Model Transparency Index. arXiv:2512.10169, Stanford CRFM. Average 40.69 (down from 58 in 2024), highest IBM at 95, lowest xAI and Midjourney at 14 each. Checked by Pebblous in the original.
- 4.Amos, R. et al. (2021). Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset. WWW ’21, arXiv:2008.09159. 1,071,488 English privacy policies from over 130,000 websites across more than twenty years. The closest precedent for the corpus method, and what this paper adds to it is a comparison axis, the developer's own account of the revision.
- 5.Stelling et al. (2025), a snapshot assessment scoring twelve providers against 65 criteria. Liang et al. (2024), field completeness across 32,111 model cards. The citations and figures for both items are carried over from reference 1, and we did not open the originals.
- 6.Ananny, M. & Crawford, K. (2018), on the transparency ideal confusing seeing with knowing, which is reference 1's starting point. The organisational theory carried into the body comes from Vaughan (1996) on structural secrecy, Meyer & Rowan (1977) and Brunsson (1989) on decoupling, and Power (1997) on auditability, all within the range reference 1 cites.
Statutes and standards
- 7.California Transparency in Frontier Artificial Intelligence Act (TFAIA, SB 53): framework publication, publication of the modified framework and a justification within thirty days of a material modification, no definition of material modification, effective 1 January 2026. EU General-Purpose AI Code of Practice: update at least annually, each update to the AI Office within five business days. Both items are described only within the range cited by reference 1.
- 8.21 CFR 314.70: (a)(6) the duty to include a list of all changes in a supplement or annual report, (b)(1) the changes requiring a prior approval supplement, (c)(6)(iii)(A) the changes-being-effected-on-receipt path for adding or strengthening a contraindication, warning, precaution or adverse reaction. Pebblous checked the text of all three provisions directly.
- 9.ISO 9001:2015 §7.5.2 and §7.5.3 (identification of documents and control of obsolete versions), ISO/IEC 5259-3:2024 (data quality management system requirements). Pebblous covered the latter in The AI Safety Dossier Is Really a Data Lineage Record.
- 10.Korea's Framework Act on the Development of Artificial Intelligence and Establishment of a Foundation of Trust (AI Framework Act), effective 22 January 2026. A different law on a different date from California's TFAIA, which took effect on 1 January 2026.
Industry and watchdog work
- 11.The Midas Project, AI Safety Watchtower, which continuously tracks the safety and security policy documents of sixteen companies. The body quotes founder Tyler Johnston, and cites the scoring from the companion Seoul Commitment Tracker, where no company scored above B− and six received a failing grade.
- 12.Naver AI Safety Framework (ASF) 2.0, published 8 July 2026 at the Seoul Forum on AI Safety Science (SFASS). Pebblous confirmed the correspondence between the Korean wording of the announcement and the English quoted in the body. It follows the version published at the 2024 AI Seoul Summit.
- 13.Keep a Changelog (keepachangelog.com), a changelog specification that takes enumeration by category as its norm. The body cites it only as a qualitative contrast and does not cite adoption statistics.