Executive Summary

On September 23, 2026, the UN Security Council held its first meeting with the safety risk of AI as the single item on the agenda. France, holding the September presidency, convened it during the General Assembly's high-level week, and Jean-Noël Barrot, France's Minister for Europe and Foreign Affairs, chaired. The chief executives of OpenAI, Anthropic and Hugging Face spoke in turn, alongside the co-chair of the UN's Independent International Scientific Panel on AI. There was no vote, and neither a resolution nor a presidential statement came out of it. This article looks at what the verification and the notification asked for there actually require.

The four requests converge on one point. For states to check each other's commitments, there has to be a record to check, kept in a shared form. Five of those are already running, and they do not line up. The EU AI Act gives providers of high-risk systems 15 days to report a serious incident to a national supervisory authority, and in the same law tells providers of general-purpose models to notify the AI Office without undue delay, with no date attached. California works in 24 hours and 15 days, and the report goes to the state Office of Emergency Services. Where international agreement jams is not the principle. It is here.

Sections 1 through 4 follow the record of the meeting and the text of each rule. Section 5, which reads this question as a matter of the evidence an organization can produce, is this article's interpretation rather than anything said in the chamber.

Key Numbers

Sources: record of the Security Council's 10228th meeting (auto-transcribed, not an official UN record), EU AI Act Article 73, California SB 53, OECD AI Papers No. 34.

Zero

Resolutions or statements adopted

A high-level briefing is a format that carries no voting procedure at all

2·10·15 days

Reporting deadlines under EU AI Act Article 73

Two days for a widespread infringement, ten for a death, fifteen otherwise. They go to Member State market surveillance authorities

24 hours

Shortest deadline in California SB 53

For an imminent risk of death or serious physical injury. That one goes to a law enforcement or public safety agency with jurisdiction; other critical safety incidents go to the Office of Emergency Services within 15 days

29

Fields in the OECD common reporting framework

Narrowed down from 88 candidates, and 7 of them are mandatory. It fixes neither a deadline nor a recipient

1

A First Meeting on AI Safety, and No Document Left Behind

This was not the first time AI had reached the Council's agenda. The Council and its members had taken up the technology's bearing on peace and security in six formal and informal meetings before this one. What made the 10228th meeting on September 23 a first lies elsewhere. Never before had a meeting narrowed its focus to the safety risk created by increasingly capable AI systems themselves.

The format was a high-level briefing. The item was listed as artificial intelligence and international security, under maintenance of international peace and security, and Jean-Noël Barrot, Minister for Europe and Foreign Affairs of France, presided for the September presidency. Four people took the floor: Yoshua Bengio, co-chair of the UN's Independent International Scientific Panel on AI; Sam Altman, chief executive of OpenAI; Dario Amodei, chief executive of Anthropic; and Clément Delangue, chief executive of Hugging Face. Amodei joined by video.

No document survived the meeting. Neither a resolution nor a presidential statement was adopted. A briefing in the Council is not where decisions get made; it is where members put their positions on the record, and no voting procedure attaches to it. The only thing produced that afternoon, then, was a transcript of what was said. That connects to everything below, because a record is precisely what this meeting was asking for.

Interior of the UN Security Council chamber — the horseshoe table and briefer's seat
▲ The Security Council chamber at UN Headquarters in New York — the September 23 briefing was held in this same room | Source: Wikimedia Commons (James D. Forrester, CC BY 4.0)
2

Every Request Came Down to Verify and Notify

Amodei brought three ideas to the Council. The first was a global agreement that starts narrow, built from items everyone can support, such as a ban on using AI to make biological weapons. The second is the passage that has been quoted most often since. He asked the Council to build "evaluation and verification systems that keep pace with AI development so that states can have visibility into frontier model capability and can verify each other's commitments." The third was common global standards for testing models for loss-of-control and misuse risks, together with "a notification system for AI incidents that are significant to global security."

He also put his own company's practice on the table. Anthropic has committed to embed external evaluators inside the company with employee-like access, and he compared them to a food inspector. An inspector who receives finished goods outside the plant and one who stands inside the process are doing different jobs. The two arrangements also demand different amounts of documentation.

Altman's request ran the same way. He proposed a mechanism in which national and international frontier AI standards complement each other, and set out what those standards would do: measure capabilities, assess risks, determine whether safeguards are sufficient, and preserve meaningful human oversight as systems become more autonomous. He also said why they have to be common ones. "We need common standards so countries can compare evidence, verify compliance, and have a shared language and understanding about what is happening." On incidents he was more specific still, calling for "accurate and speedy incident reporting, classification and reporting protocols so the world can learn from failures before they become catastrophes."

Bengio borrowed his model of verification from other industries. Frontier AI should be licensed like medicine, aviation and nuclear energy, he argued, so that safe development is what gets rewarded. What he asked of developers came in two parts: demonstrate to independent experts that a system is safe to train and safe to deploy, and monitor and report all security incidents. He then named three ways to fix the situation, the first being genuine scientific independence from the companies — and he pinned to it "a common definition of AI incidents and a shared reporting mechanism." That was the one statement of the day that went straight at the specification of the record to be kept. The scientific panel Bengio co-chairs had issued its first thematic brief two days earlier, on September 21, and it treated a single agent incident as evidence of loss of control. That incident could be reconstructed at all because a record of it existed.

Of the four, Delangue named the most concrete form. Stronger standards for monitoring and incident disclosure are what the global community needs, he said, and the way he proposed to get there was "mandatory sharing of full agent traces." Everything an agent did — which tools it called, what it reached, in what order — leaves the building intact. His grounds were not only what his own company had lived through. He said we now know that "similar incidents had been happening months earlier in secret at a handful of frontier labs without monitoring." Set the four requests side by side and the common element shows. All four asked for a record to be kept, and Bengio went as far as to say a common definition is needed. What none of them said is which fields that definition should contain.

Four requests — converging on two demands Amodei Altman Bengio Delangue Evaluation & verification visibility into frontier capability Incident notification what to disclose, and when Both require a record kept in the same form — continued in Section 4
▲ Pebblous original diagram — the four Council statements sorted into two demands

Verification is not an act of trusting the other side's good faith. It is an act of reading what the other side produced. To read it, both sides have to write things down in the same fields and the same units. If the inventory of training data, the procedure an evaluation followed, and when and how an incident was found are each kept in a different form, agreement between states stays a sentence and nothing more.

3

The Split Was Over Who Gets to Be the Verifier

Where three companies and one scientist pointed the same way, the Council members did not. Michael Kratsios, director of the White House Office of Science and Technology Policy, speaking for the United States, carried over what the president had said at the General Assembly the day before: "the United States totally rejects any attempt to construct a globalist scheme of control of superintelligence." Technology sovereignty, he added, is realized in technology adoption. He did not, however, push back on verification as such. In the same statement he said the United States has engaged frontier labs on testing and evaluation of new model capabilities and has established cybersecurity clearinghouses for coordination on vulnerability mitigation. What is being refused is not verification but a picture in which an international body is the verifier.

The Security Council chamber's circular seating, seen head-on
▲ The Security Council chamber, seats arranged in a circle around the table — that day, member states spoke in different directions | Source: Wikimedia Commons (James D. Forrester, CC BY 4.0)

Ed Miliband, the UK Foreign Secretary, stood on the other side. The core of his statement was that "governments must have the visibility to understand and assess what AI companies are doing," and he added that the leading companies have already committed to provide that visibility, which makes it an offer both sides should follow through on. China called for respect for digital sovereignty, saying every country has the right to choose AI technologies, products and services for itself. It also warned that splitting into blocs raises the risk of fragmented technical standards and incompatibility between systems. Russia took issue with the step before all that. It questioned whether so generic a subject falls within the Council's mandate in the first place, and said it belongs in more specialized multilateral fora. The example it reached for is a pointed one. The Group of Governmental Experts on lethal autonomous weapons in Geneva adopted a consensus report three weeks earlier, and that report contains a working definition of such systems. Definitions, the argument goes, are already being written in another room.

France stepped away from the chair to give its own position and cut the work into four challenges: evaluating models independently before they are made available and throughout their life cycle; being transparent about how models function and also about how they malfunction, particularly in the event of an incident; establishing internal and external oversight from training through deployment; and holding AI companies legally accountable when an incident occurs, including during the development phase. "We are not starting from scratch," Barrot said, pointing to the international network of AI safety institutes and its work on defining common evaluation methods.

Smaller states brought a different axis. Somalia said it supports mandatory incident reporting, layered technical safeguards and independent audits, and argued that African states need the institutional and technical capacity to evaluate AI systems independently rather than relying on external determination. The Democratic Republic of the Congo bundled risk assessment prior to deployment, incident reporting and traceability together as the companies' share, and said developing countries must not be reduced to mere recipients of standards they did not adequately participate in shaping. Liberia pressed this axis hardest of all. "Norms without verification are insufficient," its representative said, calling for credible auditing frameworks, robust red teaming standards and independent technical evaluation so that corporate and state commitments translate into measurable safety.

So what the day agreed on stops at the sentence saying verification is needed. It split over who the verifier would be, and the question underneath was never touched. France said disclosure has to happen when an incident occurs; Bengio said a common definition of incidents has to stand. Which fields that definition should carry, and what goes into them, never reached the agenda. That question is not waiting on political agreement. Several answers to it already exist.

4

Five Rulebooks Already Run on Different Clocks

The incident notification system asked for at the Council sounds like a proposal to build something that does not exist. In fact several are already running, and the trouble is that they do not mesh. For the same incident, the deadline differs by rule, the recipient differs, and so does what counts as an incident in the first place.

Start with the ones that bind, where the difference is sharpest. Article 73 of the EU AI Act requires providers of high-risk AI systems to report a serious incident to the market surveillance authorities of the Member States, and sets three clocks running. The baseline is immediately after a causal link is established and no later than 15 days after becoming aware; a widespread infringement cuts that to two days, and a death to ten. Yet Article 55 of the same law tells providers of general-purpose models with systemic risk to keep track of, document and report information about serious incidents "without undue delay" to the AI Office and, "as appropriate," to national competent authorities. There is no date. Inside a single law, both the recipient and the clock diverge.

Rule Deadline Goes to Force
EU AI Act Article 73
High-risk systems
15 days / 2 days widespread infringement / 10 days death Member State market surveillance authorities Legal obligation
EU AI Act Article 55
GPAI models with systemic risk
Without undue delay AI Office, and national authorities as appropriate Legal obligation
California SB 53
Frontier developers
15 days / 24 hours for imminent risk of death or serious injury Office of Emergency Services
the 24-hour report goes to an authority with jurisdiction
Legal obligation
OECD common reporting framework
Incidents and near misses
None Not specified Voluntary
NIST AI RMF 1.0
Risk management framework
No reporting clause Not specified Voluntary

Taken from the text of each rule. California SB 53 is the Transparency in Frontier Artificial Intelligence Act, signed in September 2025, and it defines a critical safety incident to include loss of control and the deliberate subversion of safeguards.

Laying the clocks over one another on a single line makes the gaps easier to see. Below is how much time the same incident is allowed before it has to be reported, depending on which rule it falls under.

One incident, different clocks — from discovery to report Discovery 24 hours SB 53 imminent risk of death 2 days EU Article 73 widespread infringement 10 days EU Article 73 · death 15 days EU Art. 73 baseline SB 53 baseline No deadline EU Article 55 OECD · NIST 0 → days elapsed
▲ Pebblous original diagram — deadlines taken from EU AI Act Article 73, Article 55 and the text of California SB 53. Horizontal spacing follows the actual number of days

The deadlines are not the only thing that differs. What gets counted as an incident differs too. The critical safety incidents SB 53 names include unauthorized modification of model weights, harm from the materialization of a catastrophic risk, loss of control, and the deliberate subversion of safeguards. That last definition carries a clause excluding the evaluation context, which is why the incident OpenAI had this summer inside its own evaluation environment fell outside the statute's reporting scope. The same event runs into a different threshold under EU Article 73, which asks first whether harm reached a person. More rules can be added and a band still stays uncounted.

Only one of the three binding rules writes into its own text what a report has to contain. SB 53 requires the reporting mechanism the Office of Emergency Services establishes to include four fields: the date of the incident, the reasons it qualifies as a critical safety incident, "a short and plain statement" describing it, and whether it was associated with internal use of a frontier model. The EU AI Act sets deadlines and hands the form to Commission guidance. So at the level of statutory text, one side has four fields, the other still has a blank, and beside them the OECD waits with a twenty-nine-field proposal. SB 53 also exempts submitted reports from the California Public Records Act. The report is kept, but no one outside reads it.

4.1An Attempt at a Common Vocabulary Already Exists

Work on aligning the forms has been done. The OECD's AI governance working party and its expert group on AI incidents published a common reporting framework in 2025. They first drew 88 criteria characterizing AI incidents out of the OECD's AI system classification framework, the AI Incidents Database, product recall portals and their own incident monitor, then filtered by relevance, frequency and expert opinion down to 29. Those 29 group into eight dimensions, seven of which are mandatory.

What this framework filled in, and what it left open, are both clear. It filled in the vocabulary. For the first time there is a baseline, across countries and sectors, for which fields an incident record should carry. What it left open is the obligation. Who must file, when, and to whom is left to national law and policy, and actual collection is being trialled through voluntary submissions to the OECD's AI Incidents Monitor. That monitor has so far gathered incidents mostly by scraping media coverage. An incident that was never reported is absent from the count from the start.

Set against other industries, the distance shows. Aviation has spent decades refining its investigating bodies and its reporting forms, and vaccine adverse events, vehicle defects and security vulnerabilities each have a mandatory reporting system of their own. Bengio invoking medicine, aviation and nuclear energy was a bid to borrow that maturity. The full agent traces Delangue asked for, by contrast, are not defined as a field by any of the five rulebooks above. One request made at the Council already has five specifications; the other has none.

5

Why Pebblous Is Watching This Meeting

From here on this is our reading. The Council's meeting was addressed to the handful of companies building frontier models and to member state governments, and it says nothing about the data practices of an ordinary company. Even so, the question that jammed in that room has the same shape as ours, only at a different scale. When we say we are using AI safely, what evidence can we put in front of someone else?

Verification between states is hard for more reasons than mutual distrust. There is no agreed format for the record that could stand in for trust. Brought down to the level of an organization, the problem arrives much sooner. The moment a customer's security review, an auditor's diligence or a regulator's inquiry lands, what is demanded is not a policy document but values. When something was trained and on what, by what procedure it was evaluated, when an incident became known: you have to be able to produce all of it on the spot.

The deadlines make this especially plain. Every deadline above is counted from the moment of becoming aware. Which makes where your organization records the moment it became aware of an incident the starting point of compliance itself. If that moment lives only in someone's memory or a chat thread, the number 15 is not counting anything.

Roughly how far your organization's evidence currently reaches shows up in the four questions below. This is not a checklist the Council handed out; it is this article carrying the question over into our own work.

  • Is what your organization calls an incident written down anywhere? Every rule defines it differently, so without your own definition standing first, nothing can be fitted to any form.
  • Does the moment an incident became known survive as a value? If it does, who records it and where is it kept?
  • Do the procedure and results of an evaluation of a model or a service stay readable after the person who ran it has moved on?
  • Can a single event be filed in two or more forms at once? For an organization answering to both the EU and a US state law, that is already the situation.

This is why, when we talk about AI-Ready Data, we ask about the structure of the record before the volume collected. For data to be used a second time later, what was done, when and under what conditions has to travel with that data. Trying to fit records to an international standard after it arrives leaves the stretch already passed beyond recovery. Events keep happening while the standard is late.

Thank you for reading this far. Every statement quoted here can be checked by anyone in the record of the Security Council's 10228th meeting. What records does your own organization point to when it says it is using AI safely? If you have ever come up short on evidence in an audit or a due-diligence review, we would be glad to hear that story.

R

References

Official Meeting Records

Statutes and Regulations

Standards and Reports

News Coverage