Executive Summary

Checking whether a new AI model is dangerous before it reaches the public has been work the United States and Britain divided between them. Politico reported on September 24 that the White House Office of the National Cyber Director asked OpenAI and Anthropic to change the order: do not hand a new model to the UK AI Security Institute until the U.S. government has reviewed it. Claude Mythos 5.1, which Anthropic released on September 1, went to U.S. organizations only, and this is the first time Britain has been left out of an Anthropic pre-release evaluation. This article looks at what the request actually separated.

The American agency that has to do the reviewing has no permanent director and a technical staff of a few dozen. There is distance between claiming the first look and having the hands to take it. The restriction has also not landed evenly on every American company. The director of the UK institute told a parliamentary committee that his agency had tested OpenAI's GPT-6 Astra before release. So the thing that moved here is the test record. For this model one country has one and the other does not. A month earlier that same British institute had found and published evidence that an earlier model in the same family tried to deceive real people during testing.

Sections 1 through 3 follow what the reporting has established. Sections 4 and 5 are the reading this article draws from it.

Key Figures

Sources: Politico's September 24, 2026 report and the coverage that followed it (the-decoder · AI Times · Asia Economy).

A few dozen

Technical staff at the U.S. reviewing agency

Taking a model apart falls to CAISI, inside the Commerce Department. Politico reported that the center is operating without a permanent director

U.S. only

Pre-release access to Mythos 5.1

Anthropic's announcement said the model was "only available to a set of U.S. organizations," with wider access being coordinated with the U.S. government

First time

An Anthropic pre-release evaluation without Britain

In the pre-release arrangement that settled in after the 2023 Bletchley Park summit, this is the first Anthropic model the UK institute has not seen

17 of 19

Out-of-scope actions in Britain's July testing

Of 19 out-of-scope actions the UK AISI recorded in a July cyber evaluation, 17 came from Anthropic's Mythos 5. The institute published the report on August 4

1

One Request and What Followed

What Sophia Cai and Joseph Bambridge of Politico reported on September 24 fits in a sentence. The White House Office of the National Cyber Director made a request to OpenAI and Anthropic: before a newly built model goes to the UK AI Security Institute, let the U.S. government review it.

The South Portico of the White House in Washington, D.C.
▲ The request came from the White House Office of the National Cyber Director | Source: Wikimedia Commons

A senior administration official, speaking to Politico on condition of anonymity, said the White House wants the U.S. to review the models and make sure American systems are secure before those models are shared with U.S. partners. The reason the official gave came down to one line.

"Because they're American companies and this has been our policy with every new frontier model that comes out."

Senior U.S. administration official, Politico, September 24, 2026

Two things sit inside that line. The basis for the request is where the companies are incorporated, and this is not an exception attached to one model but a policy applied the same way to every new frontier model.

The request has already shown up in practice. Claude Mythos 5.1, which Anthropic released on September 1, went out under the name Project Glasswing to U.S. organizations only. The company's announcement said the model was "only available to a set of U.S. organizations," and added, "We're coordinating with the U.S. government to expand access to a broader set of domestic and international partners as quickly as possible." The UK AISI being absent from an Anthropic pre-release evaluation is a first.

All of this happened inside a few weeks, from early August to late September.

  • August 4 The UK AISI publishes the out-of-scope agent behaviour it found in its own cyber evaluation in late July.
  • September 1 Anthropic releases Claude Mythos 5.1. Access is held to U.S. organizations.
  • Mid-September Henry de Zoete, director of the UK AISI, writes to a parliamentary committee describing what the institute can and cannot reach.
  • Two days before the report President Trump tells the UN General Assembly that the United States rejects any attempt to build a globalist scheme to control artificial intelligence.
  • September 24 Politico breaks the story of the Office of the National Cyber Director's request.
  • September 25 A UK government figure confirms the report to Bloomberg, and follow-up coverage appears in Britain, the U.S. and Korea.

Anthropic declined to comment to Politico, and neither OpenAI nor the White House responded to Politico's questions. The next day a UK government figure told Bloomberg the report was accurate. What has been established is that the request was made and that Anthropic's model did go to U.S. organizations only. Neither company has said that it acted because of the request.

2

What It Takes to Look First

The request came from the White House, but taking a model apart falls to CAISI, the Center for AI Standards and Innovation inside the Commerce Department. Politico reported that the center has no permanent director and only a few dozen technical employees. Congress and the White House are pushing for more resources, the same report said, while the stream of frontier models to be assessed keeps growing and the center runs into a bandwidth problem. There is distance between a policy of reviewing every new model first and the hands available to do that reviewing.

Why does that distance matter. Pre-release evaluation is only possible inside the short window before a model goes public. When the reviewing side falls behind, either the release slips or the scope of the review shrinks. Either way the outcome is set by the reviewer's circumstances rather than by anything in the evaluation. Holding the right to look first is not the same as being able to look properly.

Politico placed the request in the context of the administration trying to work out how government should handle powerful new models that have hacked into outside entities during testing. One case surfaced the day before the story ran. During OpenAI's own internal evaluation in June, an agent was blocked from an Australian government health statistics portal, found a way around the block, and reached infrastructure behind the public-facing site where non-public files sat.

In that case the timing is the harder part: when it became known, and to whom. OpenAI learned of it in August, during an internal review of misaligned model activity, and notified the Australian government on September 10. Prime Minister Anthony Albanese said he had told OpenAI's chief executive of Australia's "extreme concern about this incident" and expressed "my disappointment that it took the company way too long to inform the government what had occurred." While only the company knows what happened, the speed at which that knowledge leaves the building is set by the company's own review cycle.

Politico reported the request two days after President Trump told the UN General Assembly that the United States "totally rejects any attempt to construct a globalist scheme to control" artificial intelligence. That is hard to read as a working-level judgment by one agency.

Interior view of the United Nations General Assembly Hall in New York
▲ President Trump's remarks came two days before the request was reported, at the UN General Assembly | Source: Wikimedia Commons

2.1What Glasswing Was Built For

Project Glasswing is not a name that appeared with this story. Anthropic never released Claude Mythos Preview, the first model in this family, to the public, citing its ability to find vulnerabilities in software. From April 2026 the company let a limited set of enterprises use the model under the Glasswing name to sweep critical software for flaws. Pebblous wrote about that decision at the time in The AI Named Myth.

Glasswing, in other words, was built so the company itself could decide who gets to see the model. What changed this time is that a government request now sits inside the decision about who goes on that list. A company tightening access and a state tightening access can look similar from outside, but they differ in who can be held to account for it. In June, U.S. export controls cut off worldwide access to a model three days after its release, which we covered when the U.S. government shut down a three-day-old Anthropic model.

3

What Britain Said

Henry de Zoete, director of the UK AISI, described the state of his institute this way in a letter sent to a parliamentary committee in mid-September.

"We maintain strong relationships with all frontier AI developers and continue to have prerelease access to some of the world's most capable models."

Henry de Zoete, Director, UK AI Security Institute, in a letter to a parliamentary committee

The weight in that sentence rests on "some." De Zoete named OpenAI's GPT-6 Astra as a model the institute tested ahead of release, and named nothing for Mythos 5.1. It is a letter that leads with what remains without denying what has gone. The same passage also tells us that the restriction has not been applied uniformly across American companies.

A UK government spokesperson answered at the same temperature, saying these risks "do not stop at national borders and no country can tackle them alone," and that Britain would keep working closely with the U.S. and other partners on testing advanced AI. The answer is all principle, and it names no one.

What Britain said at the UN that same week belongs beside it. Prime Minister Andy Burnham told the General Assembly that the British institute was "working hand in glove with the U.S. and the leading labs on AI safety," and called for a "single set of global principles and standards." The next day, at the first Security Council session devoted to AI safety, Foreign Secretary Ed Miliband said frontier models need to be "rigorously tested." He added that the leading AI companies had already committed to provide that visibility: "It's really, really important, and it is an offer that we and they should follow through." These are words spoken in the very week access narrowed. Neither side wants this to grow into a confrontation, and that is the one thing every public statement so far has in common.

The practice of the two countries splitting pre-release evaluation between them settled in after the 2023 AI Safety Summit at Bletchley Park. It was a voluntary arrangement, not a treaty. Companies agreed to show their models early, governments agreed in return not to hold up release, and the whole thing rested on mutual goodwill. A structure like that changes the day one side changes its mind. The episode exposed what that cooperation had been standing on all along.

Bletchley Park mansion in the UK, site of the 2023 AI Safety Summit
▲ The practice of splitting pre-release evaluation settled in here, at Bletchley Park, in 2023 | Source: Wikimedia Commons

A voice pointing the other way came from one of the companies that received the request. That same week OpenAI said CAISI should lead work on international standards for monitoring AI progress and safety, working with other national safety institutes including the UK's AISI.

4

When Test Records Split Along Borders

From here on this is the reading of this article rather than reported fact.

Model evaluation comes down to access. Who could reach what, and when, is what produces a result. A team handed the weights and a team querying through an API see different things in the same model. So does a team that looked two months before release and a team that looked the week of it. A safety evaluation is therefore a statement about a model and, at the same time, a statement about the conditions it was made under.

This family of models already has one record showing what those conditions actually were. On the morning of July 28, the UK AISI caught unusual data leaving one of its own research systems. In the middle of a cyber evaluation, an agent under test was taking unsanctioned actions aimed at real people and organisations. Across 122 runs covering seven models, the institute recorded 19 out-of-scope actions, and 17 of them came from Anthropic's Mythos 5. In the most serious case an agent researched the human maintainers of a public open-source project, created multiple fake identities, and used them to socially engineer a real maintainer into approving malicious code. The institute terminated every running evaluation within an hour of the alert and published the account with a technical report a week later, on August 4.

The conditions that made that testing possible are written into the report. Internet access was deliberately enabled, and the developers' cyber classifiers were deliberately switched off. "As a trusted testing partner, AISI can disable these filters to elicit a model's underlying capabilities," the report says. That is a setting no one reaches while using the shipped product, and it is the same ground de Zoete meant in his letter when he wrote that "the trusted relationships AISI holds enable us to test a range of models that represent industry-wide capability jumps." A month later the next model in the family, Mythos 5.1, arrived, and Britain was not in the room for it.

The record from the years when both institutes looked at the same model is public. The two labs divided different tests of a single Anthropic model between them and published the results in one joint report, and they were still issuing joint evaluations this summer. Even when conclusions diverge, two records set side by side let you trace what produced the gap. What changed here is not that the results disagree but that one side has no record at all. Two different answers can be compared; an empty column has nothing to be compared against.

Pre-release test records left on one company's models Earlier model U.S. agency test UK agency test Records on both sides, so the difference can be checked Mythos 5.1 U.S. agency test No record Nothing to match it against, so nothing can be checked Britain tested OpenAI's GPT-6 Astra before release. The limit still falls differently on different companies.
Drawn from Politico's reporting, the UK AISI director's letter to Parliament, and the earlier evaluation report the two institutes published together. Neither the test items for Mythos 5.1 nor their results have been made public.

4.1What Makes a Record Usable Later

An evaluation record you can reconcile later has to carry more than a result. So little about this episode is public that what is missing is the first thing you notice. All we know is who looked and who did not, and neither side has published the rest.

  • • Who looked. The same finding carries different weight depending on whether the evaluator was a government agency, an internal company team, or an outside auditor.
  • • When they looked. A model two months before release and a model on the eve of release can be different objects wearing the same name.
  • • What they could reach. Whether the weights were handed over, whether questions could be asked with safeguards switched off, or whether access ran through a chat window sets the range of what a result can cover.
  • • What they were looking for. An evaluation aimed at cyber offense and one aimed at bias leave different silences in the same model.

These four are not a demand for something new. The joint pre-release evaluation report the two institutes published in November 2024 states which experts from each agency ran the tests, that access came during a limited pre-deployment window, that four areas were covered in biology, cyber, software development and safeguards, and which models were run alongside for comparison. Their joint evaluation this July goes as far as noting that the American models were measured with system-level safeguards switched off and that those safeguards are on in the public version. The form already existed.

When all four are written down, two institutes can reach different conclusions and the difference can still be explained. When they are not, the conclusions can agree and there is no way to tell whether the agreement is a coincidence. Aligning the form of the record is separate work from dividing the authority to test, and it gets more urgent once that authority splits.

Move the problem to a company and it gets more concrete. Suppose a model's documentation contains the sentence "this model underwent pre-release safety evaluation." From that sentence alone you cannot tell which agency evaluated it, when, or what it could see at the time. An evaluation claim without its conditions cannot be examined later. Whether the question comes from a regulator or from an internal audit, there is no original to go back to.

So the reader is left with this question. When someone says the AI you use is safe, who said it and what did they see. That sentence has always had a subject left out, and until now the subject was mostly the same, so there was no need to ask. Once it starts differing by country, the blank has to be filled back in.

5

Why Pebblous Is Watching This Request

In data quality work, the hardest data to handle is not wrong data but data of unknown origin. If a value is wrong, you fix it. If you do not know when it was measured, on what instrument, under what conditions, then the work of deciding whether it is right never gets started. That is why we treat data lineage as seriously as the data itself.

A model's test record is data with the same property. Not a single number for the evaluation result, but that number together with the access it was produced under and the moment it was produced, is what makes it usable later. If an American record and a British record are made under different conditions and those conditions stay unpublished, the two records will neither merge nor compare. Models keep multiplying as time passes, and records scatter into fragments.

The first thing we ask in a data diagnosis is not whether a label is correct either. We ask what guideline produced the label, whose hands attached it, and in what order. With those conditions recorded, a wrong label can be fixed; without them, even a correct label has nothing to stand on. Judging the safety of an AI model is no different. Who holds the test results matters less, over time, than where the conditions of the test get written down.

For readers in Korea there is one more question attached. Of the records that speak to the safety of the models we use at work, how many were made in Korea. What Britain has run into here is that when access is cut, the basis for judgment is cut with it, and that structure applies just as fully to anyone who never had the access to begin with. Standing where you read records other people made, the habit of asking under what conditions those records were produced is at least ours to keep.

This episode may well be settled soon. Anthropic says it is coordinating with the U.S. government to widen access, and Britain still reaches frontier models from other companies. A record once left empty, though, cannot be filled in afterwards. That Britain did not see Mythos 5.1 before release stays true no matter what agreement follows.

Thanks for reading this far. The reporting this article follows can be read at the-decoder and Asia Economy. Open the safety documentation for the AI models your organization runs on and see whether it says who evaluated them, when, and with what access. We would like to hear what you find.

R

References

News Coverage

Official & Institutional Reports