Executive Summary
This article reads the AI standards body that OpenAI, Anthropic and Google are trying to build against the thing that actually lets the original, the Financial Industry Regulatory Authority, examine a broker-dealer. The three companies have been meeting as a working group since at least July; that was reported in September, and OpenAI confirmed it two days later. Most commentary treats the story as either a backroom cartel or a step forward for self-regulation. We skip that frame and ask one question instead. In finance the ledger the referee reads came before the referee. Does the blueprint now being drawn have anything in that place?
We obtained all five public design documents in full and put the same question to each. None of them answers it. The most statute-like of the five says that examiners will audit developer practice from data selection through post-deployment monitoring, and then never says what evidence of that practice the developer has to create, in what form, or for how long. Another limits in advance the documents an auditor may see, naming model cards as the example. In a third, the only sentence that commands anyone to retain anything is a prohibition aimed at evaluators. The finding has to be stated narrowly to stay true. The idea is not missing from the field: in Europe these same companies comply with a technical documentation list set by law, and the item drops out only in the blueprint they are drawing in the United States.
The original side of the analogy is more complicated than that. Building that layer cost finance about ten years and an annual outlay approaching four and a half times its own estimate, and in March 2026, four years into full industry operation, the SEC approved deleting records more than three years old. So the argument here is not that AI should build the same thing. It is that the most expensive and most contested layer is being left unresolved while the body gets built first. For developers in Korea the last two sections are where this lands. The AI Framework Act already uses the same 1026 threshold the American design papers use, and the clause that tells anyone to keep records is not on that track.
~10 years
Time finance took to assemble the records its referee reads
Consolidated Audit Trail rule adopted July 2012, industry-wide compliance completed December 2022
~4.5x
How far that record layer's annual budget ran past the original estimate
$55 million estimated in 2016 against an actual approaching $250 million; cumulative cost near $1 billion
0
Record-keeping clauses found across the five design documents
All five checked against the full original text. Annex XI of the EU AI Act, held as a control, has them
67%
Material safety framework changes the same companies' own accounts left unverifiable
Oxford study: 257 of 383 material changes across eight version pairs. 95% CI 0.62 to 0.72
A July proposal, a September confirmation
On 14 July 2026 Demis Hassabis of Google DeepMind posted a personal essay titled A Framework for Frontier AI and the Dawning of a New Age on X and on a personal newsletter. In an Axios interview the same day, Hassabis said the hope was to have the body running within the year. The skeleton of the proposal is FINRA: an industry-funded private body that examines broker-dealers under the supervision of the Securities and Exchange Commission, with that structure carried over to frontier AI as it stands.
The essay sets out five design elements. An industry-funded self-regulatory organisation is established under the oversight of a federal agency. Its board includes “independent leading technical experts and open-source representatives”. Developers voluntarily submit models for safety testing before release, and once the regime proves effective and robust the submission becomes mandatory. Testing covers high-risk areas including cyber-offensive capability and biological risk, and agentic tests look for attempts to circumvent safeguards or signs of deception. Coverage extends to every model designated frontier-class, “no matter their country of origin or whether they are open or closed”. On money, the essay puts the burden on industry: the work needs top technical talent and compute, so funding “would need to be substantial and likely mostly come from industry”.
That one clause is everything the essay says about the board. The provision that independent directors must outnumber industry directors belongs to another design paper, discussed below, and not to this essay. The 1026 FLOP compute threshold so often quoted alongside it is not here either. The essay leaves frontier-class designation to thresholds on a benchmark suite that the standards body sets and refreshes periodically rather than to a compute figure. It also says who writes those benchmarks. The first round is “developed in consultation with Frontier Labs”, and only once the body has its own technical capacity does it build private test items of its own to guard against overfitting. The refresh cycle starts at roughly quarterly, and benchmarks that go stale or saturate are retired and replaced.
The wording of the pre-submission clause matters. The essay says developers would “voluntarily share models with the Standards Body for review up to 30 days before release”. That reads as a ceiling of thirty days rather than a floor of thirty days. Section 4 sets the thirty days beside the review windows evaluators have actually received.
Two months later the proposal turned out to be moving inside meeting rooms. On 13 September, Leo Schwartz of The Information reported that a working group from OpenAI, Anthropic and Google had been meeting regularly since at least July. Attendance is at the executive level below chief executive, and the subject is a body that would set safety benchmarks and rules. OpenAI confirmed the meetings publicly two days later, on 15 September. Chris Lehane, the company's global policy chief, said at a Washington briefing that the three had been in discussion for weeks, and the same briefing confirmed that the starting point was the July proposal from Hassabis.
This is a story with no charter and no press release yet. Separating what is a published document from what rests on anonymous sourcing keeps the rest of the article steady. The table below is that separation.
| Grade of evidence | Content | Source |
|---|---|---|
| Officially confirmed | OpenAI acknowledges discussing safety matters with Anthropic and Google. The idea started with the July proposal from Hassabis | Lehane briefing, 2026-09-15 |
| Published proposal | Submission up to 30 days before release, a board including independent technical experts and open-source representatives, industry funding, cyber, bio and deception testing, frontier-class designation by benchmark threshold | Hassabis essay, 2026-07-14 |
| Published proposal | Google's FARO concept, Anthropic's framework, OpenAI's Frontier Governance Framework | Public documents from each company, 2026-05 to 06 |
| Anonymous sourcing | Working group meetings since July, an account of Amodei leading them, a report that Altman told an internal all-hands the industry should build this without government help | The Information, 2026-09-13 |
| Anonymous sourcing | Reports that Treasury Secretary Bessent developed a FINRA-style independent regulator concept and that Chief of Staff Wiles is reviewing it, and that a draft White House executive order stalled on internal disagreement | Secondhand from Bloomberg and The Information reporting |
Table 1. What is confirmed in this story and what is merely reported. The body below states the first three rows flatly and cites the last two only as things that have been reported.
As of the publication date, 17 September 2026, this body has no name and no charter. Membership, funding, the supervising agency and the composition of a first governing council: not one of them is fixed in a public document. So this is not an article about what has been decided. It is an article about what the blueprint contains and what it leaves empty, before anything is decided.
What FINRA leans on when it examines
Borrowing an analogy means looking at how the original runs. At the end of 2025 FINRA oversaw 3,184 member broker-dealers and 639,723 registered representatives. Staff numbered about 4,200 as of 31 December 2024, and net revenue in 2024 was $1.6849 billion, 60% of it regulatory revenue. Fines are accounted for separately under a policy of keeping them out of capital planning. The examination count appears only as a formula, more than 2,000 a year, and no official yearly table could be confirmed.
On size alone this is a large institution. Size is not why it can put a finger on a member firm's violation. It can do that because what has to be sitting there when an examiner walks in was settled before the body existed. SEC Rules 17a-3 and 17a-4 are that settlement. Rule 17a-4 requires blotters, general ledgers and the rest of the basic books to be preserved for at least six years, with the first two years in an easily accessible place; order tickets and many trade-related records for at least three years; and customer account records for a further six years after the account closes. Retention period, retention format and retrieval speed are all written into the text of the rule.
What we are looking for here is the part on format. Before the 2022 amendment the default was preservation on non-rewriteable media with an index; the amendment added an audit-trail alternative. A firm choosing the alternative needs an electronic recordkeeping system that provides four things. Every modification and deletion made to a record. The date and time of the creation, modification or deletion. The identity of the individual who did it. And the fourth is better quoted than summarised: “will permit re-creation of the original record if it is modified or deleted”. The information needed to bring back the original record after it has been changed or removed. The records and their audit trails must also be downloadable and transferable immediately in both a human-readable form and a machine-processable form.
[Interpretation] The requirement is not that a change history exist. It is that the change history alone be enough to restore the prior state. Which is another way of saying that what changed has to be knowable without relying on the account given by whoever changed it. The measurement in Section 4 captured precisely that absence of recoverability.
Firm-level books are one thing; seeing across a whole market is another. After the flash crash of 2010 the SEC decided to pull order flow from the entire market into one place, the Consolidated Audit Trail. Rule 613 was adopted on 11 July 2012. Approving the plan of operation took another four years and four months, Phase 1 reporting began in June 2020, and industry-wide compliance including customer and account information was completed in December 2022. Roughly a decade passed between adopting the rule and getting the whole industry reporting. At full operation the system takes in an estimated 58 billion data points a day.
Cost grew along with the calendar. The cleanest statement of it is a sentence SEC Commissioner Hester Peirce wrote on 17 April 2026: “the estimated annual budget of $55 million in 2016 expanded to, until very recently, an actual annual budget of almost $250 million”. On cumulative cost, then-Commissioner Mark Uyeda wrote in September 2023 that “The cumulative costs for CAT may soon be approaching a billion dollars”. The same statement quotes a comment letter putting 2023 operating cost at about 5.2 times the original plan estimate, and that quotation carries an ellipsis inside it. That is why the multiple itself stays out of this article and only the year-on-year figures from Peirce are used.
2.1Four years after it was built, finance is winding it back
Read only this far and the story sounds like a case for copying finance. The institution this article holds up as a control has been moving in the opposite direction for four years. In September 2025 SEC Chairman Paul Atkins called the operating cost of CAT implausibly inflated and cut the annual run rate through a conditional exemptive order. On 27 March 2026 an amendment to the plan of operation was approved. The SEC press release describes saving $50 million to $70 million a year against the 2025 budget, and permitting the plan processor to “delete certain CAT data, including all CAT data older than three years”, which means deleting every CAT record more than three years old. Intermediate lifecycle linkages no longer have to be generated unless regulators specifically ask.
Three weeks later, on 17 April, the SEC issued a concept release and opened comment on rethinking the structure of market surveillance itself. The question Commissioner Peirce asked that day cuts against this article's own control: “Could the Commission or the self-regulatory organizations continue to conduct effective market surveillance if the CAT were eliminated?” The statement goes on to say that “the very concept of a massive surveillance database may simply be incompatible with the principles of civil liberty”.
▲ CAT timeline — about a decade from rule adoption (2012) to full industry compliance (2022), then scaled back in 2025-2026 over cost and privacy. Pebblous reconstruction from SEC originals (references 10, 11, 12).
The lesson from finance is therefore not that AI should build a CAT. Record infrastructure is needed before the referee, and it is also far more expensive, far slower to build, and liable to be wound back on cost and privacy grounds once finished. Finance spent a decade and close to four and a half times its own annual estimate on that layer, then agreed to delete anything older than three years four years after the last firms came online. Starting without the layer in the blueprint means standing up the body while the most expensive and most contested part stays unresolved.
The analogy stops here. Order records in finance and training or evaluation records in AI are different in kind. An order is a discrete event fixed at the moment it occurs, which makes it easy to load into a standard format, while the lineage of training data and the logs of an evaluation run require agreement on what counts as one unit before anything can be recorded. The privacy question has a different shape as well. What CAT ran into was the trading history of individual investors; what sits in that place on the AI side is trade secrets and model weights. Cost and privacy drove the wind-back, and that much carries over as a preview of the objections a training-data lineage duty will meet. Pebblous covered where FINRA has actually drawn the line for AI agents in Robinhood Just Let AI Trade Your Stocks and Swipe Your Card.
Five design documents in one column
Five design documents for this body are public: the July essay from Hassabis, a design paper by Mark Thomas, a senior fellow at the Institute for Progress, published in Lawfare, and the federal governance frameworks the three companies each issued around June. We put a single question to all five. Does it say what the developer must create, in what format, and for how many years, so that the examining side has something to read?
The scope of the verdict first. All five were obtained in full and swept end to end for wording covering retention, records, documentation and training data. For the Hassabis essay the full text on the author's own newsletter, identical to the X post, was the copy used. Not one of the five contains a sentence telling a developer to create and preserve a record.
Two passages in the essay come closest. One lists best practice among frontier labs and includes “publishing model cards with technical details”. The other asks for human-readable output tokens so that a model's reasoning can be understood. Both ask for something to be shown, not for something to be kept. The remaining best practices have the same character: hardening internal cybersecurity, vetting key personnel, allocating enough resources to safety and security research. No retention period, no retention target and no submission format appears anywhere in the document.
| Document | What it tells developers to do | Clause requiring records to be created and kept |
|---|---|---|
| Hassabis essay 2026-07-14 |
Submit models up to 30 days before release, cyber, bio and deception testing, frontier-class designation by quarterly-refreshed benchmarks, watermarking and readable reasoning tokens recommended | None. Publishing model cards is recommended, which is not a retention duty |
| IFP and Lawfare design paper 2026-07-30 |
Mandatory membership above 1026 FLOP, examination as the main channel of enforcement, audits of developer practice from data selection through post-deployment monitoring | None |
| Google FARO concept 2026-06 |
Publish and comply with an in-house framework, one procedural audit a year, attest compliance to FARO before releasing a new model | None. Auditor document access limited to model-card-type material |
| Anthropic framework 2026-06 |
Six-monthly risk report and system cards, notification of a critical safety incident within 15 days, unmodified access guaranteed to independent evaluators | None. The only sentence commanding retention is a prohibition aimed at evaluators |
| OpenAI Frontier Governance Framework 2026-05-28 |
Produce safety and security model reports, decide every six months whether an update is needed, publish framework changes in a changelog within 30 days | None (no retention period clause) |
| Control: EU AI Act Article 53 and Annex XI | Draw up technical documentation, keep it up to date, and provide it on request to the AI Office and national authorities | Provenance and curation methodology of training data, evaluation protocols and results, records of red teaming and model adaptation, training compute and time |
Table 2. The five design documents against the EU AI Act. The third column counts only clauses about records the developer has to create and hold. Clauses about what to publish and what to report do not go in it.
▲ Five design documents vs. the EU AI Act on record-retention clauses. Pebblous reconstruction of Table 2 (checked against full original texts).
3.1The most detailed design leaves the same place empty
Of the five, the Lawfare paper reads most like a statute. Membership is pinned to “any entity maintaining an AI model trained on 10^26 or more floating point operations (FLOPs)”, the supervising agency is to be carved out of the Center for AI Standards and Innovation inside the Commerce Department, and the board is to seat two or three more independent directors than industry directors. On enforcement it says this: “AI enforcement would primarily be an examination function. Enforcement staff—separate from the technical staff who design tests—would audit developer practices across the full model life cycle, including data selection, training, and post-deployment monitoring.”
[Interpretation] The scope of the audit is written that broadly, and no sentence follows telling anyone what evidence of that practice has to survive so it can be checked later. Auditing data selection requires a record of which candidate pool was used and what was dropped under which rule. Auditing training requires a record of which version was trained on which data. That is exactly the work Rule 17a-4 does in finance. The design paper takes the structure of FINRA and leaves behind the layer that makes FINRA work.
Google's FARO concept empties the same place by a different route. This document does put an audit in. It is an annual audit the document itself calls “procedural”, checking whether the developer followed the framework it published. And the document limits the scope of what an auditor may see: “Auditors should have access to predefined, standardized, and focused sets of documents (such as model cards) to promote consistency and prevent the risk of loss or disclosure of sensitive IP.” A predefined standard set of documents, model cards for example, with the stated reason being the risk of leaking sensitive intellectual property. The final audit report also goes to FARO confidentially.
The line drawn in the 2022 paper by Raji and colleagues, which classified audit regimes across several industries, runs like this. Where the auditee chooses and pays the auditor, sets the scope and the access, and faces limited consequences afterwards, the audit “may be functionally indistinguishable from first-party audit”. The second and third of those conditions apply to the annual procedural audit in the FARO concept as written. The same paper supplies a case showing this is not hypothetical. A voluntary audit of the hiring-tool company HireVue was assessed as “highly limited”, because what it actually examined was the company's documentation of one candidate assessment tool, with no independent evaluation of the data or the model itself.
The paper also answers the intellectual property objection: “Protecting proprietary information is not a proper response, as all audit systems provide some form of privileged access to auditors, and disclosure does not need to be direct nor absolute.” Disclosure need be neither direct nor absolute, and the working example the paper gives is the face recognition vendor test at the National Institute of Standards and Technology, where companies protect their models by having them run through a dedicated programming interface rather than handing them over. [Interpretation] Narrowing which documents can be seen and mediating how they are seen are two different answers to the same worry. The FARO document picks the first.
There is history in the choice of model cards, too. A 2020 paper on internal audit frameworks from the same first author puts model cards and datasheets at the entry condition of an audit. The auditor arrives at the first stage with a checklist, confirms that every document the development cycle should have produced is present, and only starts the audit once that holds. The FARO document uses the same artefacts as the far edge of what an audit can reach rather than the line it starts from.
In Anthropic's framework a single word sums up the arrangement. Nowhere in the document does retain appear as a developer obligation. The one place the word is used as a command is a prohibition directed at evaluators: “Evaluators should be bound by obligations not to copy, retain, or disclose confidential information”. Evaluators must not copy, retain or disclose. How much a developer must keep and for how long goes unwritten, while the party forbidden to keep anything is specified.
[Fact] All three company documents write their disclosure duties out in detail. Anthropic commits to a six-monthly risk report and system cards; OpenAI commits to safety and security model reports and to publishing a changelog within 30 days of any framework change. OpenAI's own wording is that “changes and justifications for material updates documented in a changelog and published within 30 days of the update”, which is a voluntary answer to the very problem measured in The developers' own accounts do not say what two-thirds of the framework changes were, an earlier Pebblous report. Disclosure is genuinely increasing. A developer discloses the conclusion it has assembled, and the material for retracing how that conclusion was reached stays inside the company.
In the Hassabis essay the same problem hangs on the time axis. As Section 1 noted, the test items in that proposal refresh roughly quarterly, and saturated benchmarks are retired and replaced. [Interpretation] Under a yardstick that changes every quarter, comparing one version against another requires last quarter's items, the conditions they ran under and the results to be held somewhere. Without that, each quarter's verdict means something only within that quarter, and no one can retrace which way a model moved relative to its previous version. Refreshing is the device that keeps the yardstick from going stale, and with no retention the same device erases comparability every quarter. Auditing is not the only thing that breaks when the record axis is missing.
3.2How far the absence finding can be stretched
Stretch this into a claim that AI has no notion of a record-keeping duty and it becomes false on the spot. Article 53 of the EU AI Act requires providers of general-purpose AI models to “draw up and keep up-to-date the technical documentation of the model, including its training and testing process and the results of its evaluation”, and Annex XI enumerates the minimum contents. Section 1 of the annex covers training methodology and techniques, key design choices and their rationale, the “type and provenance of data and curation methodologies (e.g. cleaning, filtering, etc.)”, measures to detect unsuitable data sources, and the floating point operations and time used for training. Models with systemic risk pick up Section 2 on top: evaluation strategies and results, evaluation criteria and metrics, and records of internal and external “adversarial testing (e.g. red teaming), model adaptations, including alignment and fine-tuning”.
So the finding in this report has to be narrowed to stay accurate, and narrowing it sharpens it. AI does not lack the idea of a record-keeping duty. The same companies that already keep that list in Europe left the corresponding axis out of the body they are designing in the United States.
3.3The prior literature has no such column either
Reading the gap in the five documents as negligence by three companies catches only half of it. The 2022 paper by Raji and colleagues quoted above is the prior work that deals most directly with designing a third-party audit ecosystem for this field. It lays audit regimes from finance, environment, medicine, aviation and food out in one table and splits institutional design into five axes: identification and scoping of audit targets, auditor independence, auditor powers and access, auditor professionalisation and standards of conduct, and post-audit consequences. The observation from that survey closest to this article reads: “All audit systems have some mechanism of access to otherwise confidential information: e.g., access to enter a facility to inspect the premises and records or the reporting or access to financial statements.” The paper goes on to call the lack of access to data and algorithmic systems the greatest weakness of the current AI audit ecosystem.
[Interpretation] None of the five axes covers what the audited party has to create. Access is defined as the power to see records that already exist, and the example the paper reaches for in explaining that power is records and documents in a financial audit. In finance those records exist by rule first, so settling access is enough. In AI that premise does not hold. The gap this article found in five design documents is not peculiar to them. The literature that examined the original of the analogy most carefully also designs the right to look and stops there.
[Fact] The 2020 paper from the same first author, mentioned in 3.1, works the opposite side. Starting from audit practice in aviation and medical devices, it proposes a procedure in which every stage of the development cycle leaves documentation, and calls the result a “transparency trail” left in the development process. The medical device parallel is the concrete one. US federal regulation requires device manufacturers to document design inputs and outputs, reviews, verification, validation, transfer and changes, and to hold them in a design history file, and the paper says of that industry that “Audit document trails are as important as the drug products and devices themselves”.
[Interpretation] The material existed in 2020. That paper put it in the audit procedure inside a company, and two years later, when the same author organised the design of an external audit ecosystem into five axes, the axis did not come across. The duty to keep ended up on the inside as good practice and the right to look ended up on the outside as institutional design. The body now under discussion inherits the second.
This paper cannot be pulled in on either side of the current argument. Its limitations section says plainly that it deliberately declined to say whether reform should come through legislation, regulation, enforcement or self-regulation. It does not address FINRA directly, taking the Public Company Accounting Oversight Board as its main comparator instead. The same limitations section carries one more warning. The effects of the five axes may be entangled, and auditor access may only mean something where the auditor is independent.
What happens when only the explanation is left
The result of a missing record duty does not have to be imagined. A measurement covering these same companies came out four days ago, and Pebblous wrote it up in a report dated 13 September 2026. A researcher at the Oxford Internet Institute collected every version of the safety frameworks published by twelve developers since the 2024 AI Seoul Summit and measured whether the revisions could be identified from the developers' own accounts.
Across eight version pairs that came with an account, 383 material changes were traced, and for 257 of them the developer's notice alone does not reveal what changed or how. That is 0.67 on the strict criterion, with a 95% confidence interval of 0.62 to 0.72. What this measured matters: not the documents but the statements developers made about their own revisions. The documents were all public. Even with the documents public, what had changed could not be reconstructed.
Of the nine categories, the one closest to this question is third-party involvement. Taking the original report's own wording: third-party involvement is the smallest of the nine categories, 18 of its 21 material changes came with an account, and the silence rate is 0.67. And a commitment to give external researchers and governments access to models went missing inside that small category. A promise to keep the channel for outside verification open was closed without the outside hearing about it.
Two cautions belong together here. First, the 0.67 in this paragraph is the value over 21 changes in the third-party involvement category, and the 0.67 in the paragraph above is the value over all 383. The same number came up by coincidence and the denominators differ. Third-party involvement has the smallest sample, and its confidence interval runs from 0.44 to 0.84. Second, the paper's appendix states that no significance tests were run between categories. Nothing in that table supports saying one category goes silent significantly more than another.
▲ Two different 0.67s — the all-383 value and the third-party-involvement-21 value have different denominators. Reconstructed from the Oxford Internet Institute study (cited via Pebblous's report of 13 September 2026).
4.1A 30-day window against the time third parties actually got
Records are not the only problem. How the 30-day pre-submission window at the centre of the Hassabis proposal meets real evaluation practice is also checkable in public material. METR is one of the few organisations that publishes the access conditions it received alongside its frontier model evaluations. For GPT-4.5 the access METR received started about a week before release, and METR noted that had a risk surfaced after release, the additional exposure would have been that one week. Conditions improved for GPT-5: checkpoint access from 10 July gave roughly three weeks before release, and an assurance checklist, a final checkpoint and a version exposing the reasoning process were opened in stages.
Three intervals side by side
- Pre-submission window in the Hassabis proposal: up to 30 days before release, voluntary
- Access METR actually received: about 1 week for GPT-4.5, about 3 weeks for GPT-5
- Standard in METR's own evaluation protocol: a small team on a budget of a few million dollars should be able to finish in about a month
From METR's public reports and protocol documents. The three measure different things, so read them for scale rather than subtracting one from another.
Thirty days is not a generous window. It is slightly longer than what has actually been granted so far. The wording of the essay also points to a ceiling rather than a floor. And something has to be settled before the length of the window: what the evaluator receives during it. Getting final weights only and getting the composition of the training data and the evaluation history of previous versions are not the same thirty days. The design documents state the length of the window and say nothing about what crosses it, which puts them in the same place as the gap in Section 3.
METR has written down what failed to cross. The GPT-5 report contains a table called the assurance checklist. The left column holds assumptions METR considered necessary for its risk verdict, and the right column holds the answers OpenAI gave for each. There are three assumptions. The model was not trained to hide relevant information. There is no reason its sabotage capability would jump sharply above existing models. METR's results do not conflict with results and evidence held by OpenAI researchers. The answers in the right column include “Relevant capabilities were not artificially suppressed in the model” and “There are no elicitation techniques or internal tools known to drastically improve performance”.
[Interpretation] The table does not show that pre-deployment evaluation is sloppy. METR is in fact a rare case of an evaluator publishing a table of what it is relying on. The table shows what the conclusion of a pre-deployment evaluation stands on. The three assumptions holding up the verdict were confirmed by the developer in words rather than observed by the evaluator. Had the training process been recorded in a standard format and submitted, a history would sit in those cells. Assertions sit there instead.
METR has also passed judgement on the 30-day unit itself. In the GPT-4.5 write-up the organisation says pre-deployment evaluations “do not address risks from internal deployment”, and that because public deployment is reversible, a short pre-deployment evaluation does very little to reduce a model's lifetime risk even when perfectly accurate. The alternative it names as promising is “external review of the AI developer’s internal results”: the developer and the third party agree in advance which evaluations run under which elicitation method, the developer hands over the results, the third party reviews them alongside its own, and where commercial leakage is a concern the work happens at the developer's facility on the developer's hardware. [Interpretation] This approach widens what crosses the window instead of widening the window. That requires the developer to hold internal evaluation results and their run conditions in a shape agreed beforehand. It is the clearest illustration of where the record axis and the access axis interlock.
A proposal that goes one step further on access is already on the table. We Must Pace the Frontier, an essay Dario Amodei of Anthropic published on 12 September 2026, proposes giving external evaluators resident status: a desk in the office, a badge, a company laptop, and authority comparable to an internal risk assessment team. The essay draws the comparison with bank examiners stationed alongside bank staff, and attaches the condition that the evaluator must be able to keep public judgement. [Interpretation] This proposal opens the access axis, not the record axis. A resident evaluator can sit at that desk and still see only what the company is showing right now, unless the per-version lineage of the training data was built in the first place. Neither axis substitutes for the other. The warning in 3.3 also catches here. Auditor access may only mean something where the auditor is independent, and the desk, the badge and the laptop come from the company being evaluated.
Three positions and antitrust
Three companies sitting in one room do not necessarily want the same thing. What each has written in public splits on the shape of the body they want. The strength of the evidence differs as well.
The Hassabis position is the one set out above: an industry-funded self-regulatory organisation under federal oversight, starting voluntary and moving to mandatory. A signed essay makes it the most solid of the three. Amodei puts something different on the table in his essay of 12 September. Anthropic would take resident evaluators first, then coordination among companies in democratic countries, then an international agreement including China. The second step comes with a condition: “For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations.”
The Altman position comes in two layers. In public the tone is favourable to cooperation. On 15 September Altman said “I think it’s great for our industry to say we want to come together and coordinate”, adding that accidents are unavoidable and transparent reporting of them is needed. On what the body should look like, an internal all-hands reportedly heard that the large labs should build it themselves without government support. That second layer rests on anonymous sourcing and does not carry the same weight as the other two.
5.1One company asks for the waiver, another says it isn't needed
Two days after Amodei asked for a waiver, a company sitting in the same room took the opposite position in public. Chris Lehane of OpenAI said at the Washington briefing on 15 September that the three companies did not see any need for an antitrust exemption to discuss safety. The grounds were existing precedent and practice: competitors in aviation have shared safety-sensitive information for years, and the Obama administration established that information sharing to counter cybersecurity threats is permissible. The line Lehane left was “It’s better to try to work together to prioritize safety”.
[Interpretation] The two look like a head-on collision, and they may not be talking about the same object. Jointly setting safety standards and jointly slowing the pace of releases read differently under competition law. The first is standard-setting cooperation, a category handled for decades. The second edges toward competitors agreeing to restrict output. The waiver Amodei's essay asks for is attached to the second step, the coordination of pace, and what Lehane said was unnecessary is safety discussion in general.
Pace coordination is not confined to the September essay from Amodei. The July essay from Hassabis, explaining the approach of the standards body, says the regime could be tightened if the seriousness of the situation demands it, and gives as an example of tightening “coordinating a slowdown in development among the Frontier Labs”. [Interpretation] The distinction in the previous paragraph still holds: the general safety discussion Lehane called exempt from any waiver and the pace coordination Amodei asked a waiver for are different things. The change is that pace coordination already sits in the document this whole effort started from, written there as a function of the body. Antitrust here is not an issue that might arise later. It is a line already in the blueprint.
The regulators answered quickly and coldly. On the same 15 September, Treasury Secretary Scott Bessent rejected requests for liability protection for AI companies at a House hearing, and Federal Trade Commission Chairman Andrew Ferguson said everyone should be deeply suspicious of requests for antitrust exemptions. Ferguson added that the companies are effectively asking for a barrier to entry and that the barrier would defend their own incumbency, with the caveat that this was a personal view and that federal AI policy is set by the president.
Opposition came from inside the industry too. Aidan Gomez, co-founder of Cohere, criticised the dominant American labs for forming a “cartel”, noting that the issue is not whether rules exist but “who writes them, who gets to participate”. That word is Gomez's and not this article's verdict. The question it points at does overlap with this report's own. Who decides the material the rules will stand on belongs to the same ground as who writes the rules.
▲ Two branches of the antitrust question — coordinating on safety standards is not the same as coordinating release pace. Reconstructed from statements by Lehane, Amodei, Bessent, Ferguson and Gomez (references 6, 7, 8).
A separate route is moving in Congress at the same time. At the same briefing Lehane said OpenAI supports the external safety evaluator provisions of the FRONTIER Act. Discussing a self-regulatory body while backing a federal mandate is not a contradiction, because the two instruments aim at different places. The federal draft that would turn auditors into a licensed profession is covered separately in A House Draft Wants to Make AI Audits a Licensed Profession.
These three companies coming together is not itself new. In July 2023 Anthropic, Google, Microsoft and OpenAI launched the Frontier Model Forum, and Amazon and Meta joined in May 2024. The forum has issued technical reports on risk taxonomies, capability evaluation and mitigations, and in March 2025 members signed an agreement to share vulnerability and threat information. A new body is under discussion because of what the forum does not do. It has no external audit function, and members largely report on themselves. The body now being discussed departs from the forum in one respect, an intent to hold the power to judge, and that intent brings with it, immediately, the problem that the basis for a judgement has to survive outside the member firm.
Korea is on the next track over
The threshold under discussion in the United States is already law in Korea. Article 24(1) of the Enforcement Decree of the Framework Act on the Development of Artificial Intelligence and the Establishment of a Foundation for Trust defines the presidential-decree standard referenced by Article 32 of the Act as a system meeting all three requirements, the first of which is that cumulative compute used in training is at least 10 to the 26th floating point operations. The other two are that the system is built and operated with state-of-the-art technology, and that its risk level may have a broad and serious effect on human life, physical safety and fundamental rights.
[Fact] The 1026 FLOP figure is the value used by US Executive Order 14110, revoked in January 2025, the value the Lawfare design paper revived as a membership test, and the value Google's document names as a provisional standard. It sits ten times above the 1025 the EU AI Act uses as its systemic-risk threshold for general-purpose models. The Hassabis essay carries no compute figure at all, which is worth recording separately; that text leaves frontier-class designation to benchmarks the body would set. American design papers fill that slot with compute, Korean law fills it with compute plus two further requirements, and the original proposal left it empty and handed it to the body. The Act took effect on 22 January 2026, so Korea has pulled by statute the same trigger that was revoked once in the United States and then came back in a design paper. [Interpretation] Paragraph 2 of the same article says the specific method for calculating cumulative compute will be determined and published by the Minister of Science and ICT. The trigger is set and the definition of the measurement that pulls it is still empty. That is the same shape this article found in the American design papers.
6.1The duty to create and keep documents falls to the other track
Korea's AI Framework Act runs its duties along two tracks. One is the frontier track, caught by a compute threshold. The other is the high-impact track, caught by domain, such as medicine or hiring. So which of the two carries the clause about records?
| Provision | What it catches | What it requires |
|---|---|---|
| Article 32 safety assurance duty |
Cumulative compute at or above 1026 FLOP (frontier track) | Identify, assess and mitigate risk across the lifecycle, monitor safety incidents and build a risk management system, submit the results to the Minister. Specific methods and submission items delegated to ministerial notice |
| Article 34(1) duties of high-impact AI operators |
Domain-based high-impact AI (high-impact track) | Subparagraph 2: establish and implement a plan for explaining matters including an overview of the training data. Subparagraph 5: create and retain documents that allow the content of safety and reliability measures to be confirmed |
Table 3. Based on the statutory text at the Korea Law Information Center (Act No. 20676, in force 22 January 2026). Article 32 has no subparagraph corresponding to the creation and retention of documents.
Between the duties borne by a company training a frontier model and those borne by a company applying AI in medicine, hiring or education, the one that says to keep records is the latter. Article 32 carries a submission duty, and the contents of that submission have been passed to a ministerial notice. [Interpretation] The axis missing from the American design papers has not vanished from Korean law; it sits on the track next door. The two gaps are not the same shape, and they overlap in this: in the norms that govern models above the threshold, records come later.
The other half of this is when enforcement actually begins. The high-impact AI impact assessment in Article 35 is an endeavour provision, and the Ministry of Science and ICT has announced a grace period of at least one year during which fact-finding investigations and administrative fines are suspended. Substantive enforcement slides to early 2027. For domestic developers that period is not a reprieve but a design window. The notice under Article 32(3) has not yet filled in what must be submitted, which means deciding now whether the artefacts being kept today will line up with what that notice eventually asks for.
6.2Five artefacts worth keeping now
While the notice stays empty, the most concrete ready-made list to work from is Annex XI of the EU AI Act. Its items are not hypothetical demands, since companies putting models on the European market already comply with them, and they are also the list most likely to be consulted if the body under discussion in the United States later adds a records clause. For a team in Korea drafting safety and reliability documentation right now, five items are worth making room for inside the document.
- Provenance and curation method of the training data. Not only where it came from but what was filtered out under which rule. Annex XI gives cleaning and filtering as its examples.
- Basis for the cumulative compute figure. Not a single number but how that number was counted. Once the notice fixes a calculation method, recounting retroactively becomes a live possibility.
- Evaluation protocols and results. Not the conclusion that something passed but what was run under which conditions. Keeping it reproducible is what lets a third party see the same thing.
- Records of red teaming and model adaptation, alignment and fine-tuning included. This is what proves the version being deployed is the version that was evaluated.
- Per-version change history, recorded change by change: what changed, when, and in which direction. The absence of exactly this item is what the Oxford study measured.
These five are a design choice rather than a compliance cost. Building a place for history inside the document from the start costs differently from bolting version control on after the writing is done. Pebblous set out the Korean institutional side in The AI Safety Dossier Is Really a Data Lineage Record, and the route where the state fixes the qualifications of the verifier in Illinois Becomes the First State to Put Frontier AI Under Annual Outside Audit and California Decides Who Can Audit AI, but Auditors Decide the Test. The California case shows the same kind of gap. The qualifications of the auditor were settled and the choice of yardstick was left to the auditor.
Why Pebblous cares
Pebblous diagnoses data and issues quality reports on it. Which body three American tech companies decide to build looks at first glance like somebody else's neighbourhood, and we followed it to the end because the place where the discussion tripped is the place we handle every day.
7.1We diagnose other people's data from outside too
DataClinic works from outside as well, on data and model outputs that belong to someone else, and leaves the diagnosis in a form a third party can check again. Do that work and the first thing you hit is not the yardstick but the material. A diagnosis request arrives, we open it, and often the data in its current state is there while the path that led to that state is not. Without knowing which record arrived when, what was dropped, and which label was reapplied, all we can issue is an observation rather than a verdict. A FINRA-style standards body will hit the same problem in the same order. Settling who the referee is does not help if what the referee is meant to read was never prepared, and the verdict settles into an observation.
7.2Evaluation means something only if its subject holds still
Pre-deployment evaluation as an instrument stands on the premise that the thing evaluated is fixed. Holding that premise requires proving three things. Which data the model was trained on. Which version passed the evaluation. Whether the model being served now is that version. All three are data lineage questions, and none of them is answered by examining model weights. Where the path from the composition of the training data to the model's internal representations goes unrecorded, an evaluation result cannot be more than a snapshot of that version at that moment. We put the third question to providers in Korea and wrote up the result in No provider could show that the evaluated model is the one answering you.
7.3The layer both outcomes need
Most commentary on this story divides over self-regulation against statutory regulation. One layer is needed whichever way that fight goes, and few pieces point at it. An auditable data trail is that layer. If the self-regulatory body wins, the body needs something to read. If statutory regulation wins, the regulator needs something to read. If both lose, nobody has anything to read, and even then whoever is buying a model will still ask what they are supposed to look at before buying. That is why Pebblous keeps returning to this layer. Who becomes the referee matters less than what gets read to reach a verdict, and the thing in that place is the artefact AI-Ready Data has been defining. Only Two Countries Can Verify the World's Most Powerful AI, which measured how many countries can run the verification at all, sits on the same axis.
7.4What has to be settled before the body
This report wants to leave one sentence behind. Standing up a referee and assembling the ledger the referee reads are separate jobs, and in finance, the original, the ledger came first and it was the part that took a decade and serious money. If the body now being designed starts with that layer unresolved, adding the layer later costs more than adding it now. For companies in Korea the same point arrives in a plainer form. Building something you can submit is cheaper than waiting for the notice to tell you what to submit.
This article covers a story still taking shape. All five design documents were checked against the full original text, and the absence finding is bounded by those five plus the EU AI Act held as a control. Among the differences between the three companies, the Altman position rests on anonymous sourcing, which is why the table grades the evidence. Once the body has a charter, the third column of Table 2 may well get filled. This article will happily go out of date then, and if the table is what confirms the date has passed, it will have done its job. Thank you for reading this far.
References
The spine of this article is an absence finding, so the scope of the evidence is recorded item by item. All five design documents (the Hassabis essay, the Lawfare design paper, and the Google, Anthropic and OpenAI frameworks) were obtained in full and swept end to end for wording covering retention, records, documentation and training data. The retention periods and audit-trail requirements in SEC Rule 17a-4 come from the text of the Code of Federal Regulations; the SEC statements and press release, the EU AI Act articles and the Korean AI Framework Act provisions were each checked in the original. Both prior papers were checked against the arXiv full text. The Oxford figures were not recalculated and are carried over as published in an earlier Pebblous report.
Design documents (checked in full)
- 1.Thomas, M. (2026). Designing a FINRA for Frontier AI. Lawfare, 30 July 2026. Senior fellow at the Institute for Progress. Membership at 1026 FLOP, a CAISI spin-out, board composition, examination-led enforcement, a levy on compute providers. No record retention or submission clause was found.
- 2.Google (2026). A Pragmatic Approach to AI Governance in America. June 2026, under the name of Kent Walker. The FARO concept, an annual procedural audit, pre-release attestation, a limit on the documents auditors may access, and 1026 FLOP named as a provisional standard.
- 3.Anthropic (2026). Anthropic’s Advanced AI Framework. June 2026. Covered-entity criteria of 1025 FLOP plus a revenue test, six-monthly risk reports, system cards, notification of critical safety incidents within 15 days, and the independent evaluator provisions. The sentence quoted in the body prohibiting evaluators from retaining information is here.
- 4.OpenAI (2026). OpenAI’s Frontier Governance Framework. Published 28 May 2026. The document states that it addresses California's TFAIA and the EU general-purpose AI code of practice. Safety and security model reports, a six-monthly decision on whether to update, and publication of framework changes in a changelog within 30 days. No retention period clause was found.
- 5.Hassabis, D. (2026). A Framework for Frontier AI and the Dawning of a New Age. 14 July 2026. The same text as the X post appears on the author's own newsletter, and that full text was the copy checked. FINRA is named explicitly as the model, and the board composition, industry funding, voluntary submission up to 30 days before release, quarterly benchmark refresh and coordination of pace among frontier labs are all in this text. There is no compute threshold and no record retention clause. The Axios interview published the same day was checked alongside it.
Reporting on the events
- 6.Schwartz, L. (2026). The Information, 13 September 2026. Report that a working group from the three companies had been meeting regularly since at least July. Anonymously sourced, and graded as such in Table 1.
- 7.AFP (2026). OpenAI, Anthropic and Google are working to create an AI standards body. 15 September 2026. Carries Lehane's official confirmation and the remarks from Altman, Gomez and Amodei.
- 8.Amodei, D. (2026). We Must Pace the Frontier. 12 September 2026. The original of the resident evaluator proposal and the antitrust waiver request. Both verbatim quotations in the body come from here.
- 9.The other side of the antitrust question (Bessent's answers at the House hearing, the remarks by FTC Chairman Ferguson) and Lehane's support for the FRONTIER Act were cross-checked across several Washington reports dated 15 September 2026. This part relies on individual outlet summaries, and the primary transcripts could not be confirmed.
The finance control (SEC originals)
- 10.SEC (2026). SEC Approves Amendment to NMS Plan to Further Reduce Costs of the Consolidated Audit Trail. 27 March 2026. Savings of $50 million to $70 million a year and permission to delete data older than three years. Source of the verbatim quotation in the body.
- 11.Peirce, H. M. (2026). Statement on the Costs, Risks, and Privacy Concerns of the Consolidated Audit Trail. 17 April 2026. The $55 million against $250 million comparison, the question about eliminating CAT and the sentence on civil liberty are all in this statement.
- 12.Uyeda, M. T. (2023). Statement on CAT Funding. 6 September 2023. The sentence about cumulative cost approaching a billion dollars is the commissioner's own; the figure of about 5.2 times belongs to a comment letter quoted in the statement and that quotation contains an ellipsis. That is why this article does not use the multiple as a figure in the body.
- 13.17 CFR § 240.17a-4, Records to be preserved by certain exchange members, brokers and dealers. The retention periods in the body come from paragraphs (a), (b) and (c); the audit-trail requirements and the quoted sentence from paragraph (f)(2). Read alongside SEC Rule 17a-3, SEC Rule 613 (Consolidated Audit Trail) and the 2022 amending final rule (34-96034). The FINRA scale figures come from the 2026 Industry Snapshot and the 2024 annual financial report.
Statutes and scholarship
- 14.EU AI Act Article 53 and Annex XI. The duty to draw up and keep technical documentation up to date, and its minimum contents. The control column in Table 2 and the checklist in Section 6 come from here.
- 15.Framework Act on the Development of Artificial Intelligence and the Establishment of a Foundation for Trust (Act No. 20676), Articles 32, 34 and 35, and Article 24 of its Enforcement Decree (Presidential Decree No. 36053). Checked against the statutory text at the Korea Law Information Center. The Act took effect on 22 January 2026.
- 16.Raji, I. D., Xu, P., Honigsberg, C. & Ho, D. (2022). Outsider Oversight: Designing a Third Party Audit Ecosystem for AI Governance. AIES ’22, arXiv:2206.04737. The five-axis split of institutional design opens Section 3; the sentence on audit access carried into the body and the finding on indistinguishability from first-party audit are in the Table 1 summary in Section 3; the intellectual property counter-argument and the NIST example are in Section 4.3.2; the HireVue case is in Section 4.3. The statement that the paper deliberately declined to choose between legislation, regulation, enforcement and self-regulation, and the warning about interaction between the axes, are from the limitations section in Section 5. The paper's main comparator is the PCAOB, and it does not address FINRA directly.
- 17.Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D. & Barnes, P. (2020). Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing. FAT* ’20, arXiv:2001.00973. The phrase transparency trail is in Section 4.1, the sentences on the medical device design history file and audit document trails are in Section 3.2, and the passage placing model cards and datasheets at the entry condition of an audit is in Section 4.4.
- 18.METR public evaluation reports (GPT-4.5, GPT-5) and the example evaluation protocol. Source of the access periods and the budget standard. The assurance checklist table and its three assumptions are in the GPT-5 report; the limits of pre-deployment evaluation and the proposal for external review of internal results are in the GPT-4.5 write-up.
Related Pebblous articles
- 19.The developers' own accounts do not say what two-thirds of the framework changes were · No provider could show that the evaluated model is the one answering you · The AI Safety Dossier Is Really a Data Lineage Record
- 20.Illinois Becomes the First State to Put Frontier AI Under Annual Outside Audit · A House Draft Wants to Make AI Audits a Licensed Profession · California Decides Who Can Audit AI, but Auditors Decide the Test · Robinhood Just Let AI Trade Your Stocks and Swipe Your Card · Only Two Countries Can Verify the World's Most Powerful AI