Executive Summary
Twenty jurisdictions that host no frontier model developer have started writing documents that govern general-purpose AI. What those documents cover is remarkably similar. Assess systemic risk, evaluate and verify the model, prohibit certain things, and report when something goes wrong. Those four. Go down to the clause level and count what legal force those sentences actually carry, though, and the convergence of form turns out not to be a convergence of force. In the paper's own phrasing, these countries converge in form and diverge in force.
The break comes in the same place every time. They built the power to observe and left the power to act mostly blank. Evaluation and verification are not a neglected corner. They rank near the top by provision count, and sixteen of the twenty have something in place. What is missing is legal force. Eleven jurisdictions have stood up an AI safety institute or an equivalent evaluation body, and not one of them, the European AI Office included, carries a duty to monitor prohibition breaches. Capability sits in one institution and consequence in another, which makes this a problem of mandate design rather than of technology.
So where do the evaluations and safety documents that companies produce actually go? If no regulator receives them, or the one that does cannot ask a follow-up question, then the real recipient is not a regulator at all. It is customer due diligence, procurement review, audit, and the partner contract. Change the recipient and you change the design brief for the document. A regulatory gap does not remove the demand for evidence; it moves the address the evidence is sent to.
Four numbers hold this report up. The first two say how light the documents carrying these rules are. The second two say whether anything is waiting to receive what those rules turn up. All four have different denominators, and stripping the denominator makes every one of them false.
85 of 382 positive provisions
sit in binding law in force
binding instruments that define
general-purpose AI at all
named safety or evaluation institutes
holding enforcement power (of 144 actors)
operative incident reporting provisions
that set both a deadline and a consequence
The underlying source is a mapping study posted to arXiv on 18 August 2026 by five researchers from the Arcadia Impact AI Governance Taskforce and The Future Society. It is a working paper without peer review, and the data is a snapshot as of 30 July 2026. The instrument, actor and definition tables are published under CC BY 4.0, so some of the figures below were checked directly against the original CSV files; where that was not possible, this report says so.
Editor's note. Where the evidence you produce ends up, and who actually reads it, is a question Pebblous cannot stay out of; documenting the quality of data and models is what we do. This piece is not a summary of regulatory developments. It places 587 provisions on a four-link accountability chain and counts, with denominators attached, where the chain breaks. We have not pulled the conclusion our way. The observation stops at the point where the recipient of the evidence moves away from the regulator; what to prepare from there is left to the reader's own setting.
Twenty countries filled in the same four boxes
The twenty jurisdictions in this study share one trait. None of them hosts a frontier developer, which means none of them can reach the point where a model's capabilities are decided. Australia, Brazil, Canada, Chile, France, Germany, India, Israel, Japan, Kenya, Nigeria, Peru, Singapore, South Africa, South Korea, Switzerland, Taiwan, the United Arab Emirates and the United Kingdom, with the European Union taking the twentieth slot as a supranational jurisdiction. Each of them has to govern models built elsewhere, by companies they cannot license or tax, by talking to whoever they can actually reach. Deployers, purchasers, platforms, and supervised financial institutions are who they can reach.
The researchers read the primary legal texts of these twenty jurisdictions across four domains. Systemic risk assessment, evaluation and verification, prohibitions and the monitoring and detection of breaches, and serious incident reporting. Those four boxes are not an arbitrary bundle of topics but the order of an accountability chain. Measure the risk, verify what the measurement produced, draw the line that must not be crossed, and raise the alarm when it is. When an earlier link is empty, the later ones lose their footing. That order is also the spine of this report.
1.1Absence was recorded as data too
The reading produced four separate populations: 101 instruments, 587 provision records, 144 governance actors, and 52 definitions of general-purpose AI. Those four numbers are four different denominators. Attach an instrument count to a discussion of provisions, or read an actor ratio as if it were a provision ratio, and the sentence goes wrong end to end. That is why every figure below arrives with its denominator attached.
The 587 figure carries a methodological decision inside it. Positive provisions, meaning clauses that actually require or recommend something, number 382. The remaining 205 are confirmed absences. Each one records that the team followed a stated search procedure through statute books, regulatory guidance and strategy documents, ran expert verification, and then wrote down that this jurisdiction has nothing in this domain. That is treated as a finding about the country rather than a hole in the research. It separates silence from decision, and it is how a comparative governance study preserves its negative results.
Set this against the existing AI policy trackers and the value of the design becomes clear. The OECD AI Policy Observatory and the IAPP-family trackers count somewhere between 900 and more than 1,000 policy initiatives across 80-plus jurisdictions. Far broader. But those are jurisdiction-level inventories, so they do not record who an instrument binds, to what obligation, or with how much legal force. Nine hundred initiatives and 101 instruments carrying 85 binding provisions are not competing counts. They are counts at different resolutions.
1.2Form converges, force diverges
On the topic side the convergence is unmistakable. Fifteen of the twenty engage with at least three of the four domains, and eleven have positive provisions in all four. Every jurisdiction except South Africa has written something into at least one box. Move to the force side and the picture splits. Of the 382 positive provisions, the ones sitting in binding law that is in force number 85, or 22%. The rest are guidance, strategy, and bills not yet passed. And even those 85 are not spread evenly.
Engaging with a domain is also not the same as covering it. Only five jurisdictions have two or more provisions in each of the four domains: the European Union, the United Kingdom, Australia, Japan and Brazil. Kenya and Korea fill all four boxes but hold a single incident reporting provision each; India and Singapore fill all four but hold a single risk assessment provision each. In the paper's own count, nominal coverage is eleven jurisdictions and substantive coverage is five. The United Arab Emirates is the sharpest case of a chain broken at both ends. It holds 32 positive provisions, nearly all of them massed in evaluation and prohibitions, with two in risk assessment and none at all in incident reporting.
Here is how the binding provisions are distributed. The European Union and Brazil take more than half of the 85 between them, the United Kingdom and Korea follow, and ten jurisdictions hold none at all.
Binding provisions by jurisdiction (denominator: 85 binding of 382 positive provisions)
The paper names only the top four and the next group (Taiwan, Australia, Singapore, Peru, Canada, the Dubai International Financial Centre, India). The 21 and the 0 in the last two rows are the residual after subtracting the top four from 85. The DIFC is a subnational unit and is not part of the count of twenty jurisdictions. Source: the paper's clause-level tallies.
The jurisdictions with no hard-law instrument whatsoever number seven: Japan, Israel, Switzerland, Kenya, Chile, Nigeria and South Africa. France and Germany join them. Both rank at the very top of the composite index and hold no domestic hard-law AI instrument at all, with their entire binding layer arriving through the direct effect of EU law. The tenth slot belongs to the United Arab Emirates, whose only binding instrument sits in the subnational Dubai International Financial Centre. Seven and ten are different counts and should not be mixed.
Brazil needs a caveat as well. Eighteen reads like the second-strongest regulator after the EU, but a large share of that sits inside a bill that is not yet law. PL 2338/2023 passed the Senate unanimously in December 2024, moved to the Chamber of Deputies, and is now waiting on a rapporteur's opinion in special committee. As of August 2026 it has still not been enacted, so the paper's description of a bill pending in the legislature for more than three years still holds. That said, Brazil also holds three binding instruments already in force, so reading it as a pure preview would be inaccurate too.
The United Kingdom carries a caveat pointing the other way, and the paper treats it as the sharpest case in the sample. The UK has sixteen instruments and 44 positive provisions, the most of anyone outside the European Union, and twelve of those provisions are binding. But all twelve sit inside two sectoral statutes, the Online Safety Act 2023 and the Crime and Policing Act 2026, and not one of them places a duty on a provider of a general-purpose AI model. Even the single functioning reporting provision asks for a reporting procedure to exist rather than for anything to be reported. The binding provisions are real; they are aimed elsewhere. Australia shows the same shape at a different scale, with thirteen instruments and 45 positive provisions, again the largest outside the EU, and binding provisions that are entirely sectoral.
Which is why, in this sample, activity and force lie on different axes. India generates 20 positive provisions from nine instruments and touches all four domains, yet holds one hard-law provision in force. Peru sits near the bottom of the composite score and holds binding criminal provisions. Israel has three instruments and eight provisions; Singapore has thirteen instruments and 27 provisions, and Singapore's generative AI framework recommends notification once a materiality threshold is crossed without ever setting the threshold, while Israel's anti-money-laundering and counter-terrorist-financing directives bolt onto a statutory regime that already runs. What makes the difference is not the volume of documents but whether the document is wired to a machine that works.
1.3The blueprints split six ways, and one came back blank
The twenty did not converge on a single regulatory model. The paper distills six recurring cross-tier architectures. The Korean pattern is an anchor: one horizontal framework act with subordinate legislation hung beneath it. The French and German pattern inherits its entire binding layer through the direct application of supranational law. The UK and Australian pattern disperses across many subordinate instruments with no horizontal statute. Kenya's pattern is declarations and strategy only; the Taiwanese and Emirati pattern binds through sectoral hard law; and the Chilean and Nigerian pattern consists of proposals.
And one of the twenty fits none of these boxes. South Africa is the only jurisdiction with no in-scope provisions at all during the study period. The team went through statute books, regulatory guidance and strategy documents and ran expert verification, and no instrument passed the scope test. Its one candidate, the draft National AI Policy, was withdrawn before adoption in April 2026. The reason for that withdrawal is not in the body of the paper. This report returns to it in the final section.
They never defined what they were regulating
You would think writing a rule requires writing down what the rule is about. Not in this sample. Of the 101 instruments, 42 define general-purpose AI or something equivalent. The other 59, meaning 58% of the total, layer rules on top of no definition at all. That much can be counted directly from the published instrument data.
The imbalance shows up next. Split by legal tier, and definitions cluster in the lighter documents. Of the 20 binding instruments, five carry a definition. Among soft-law and non-binding instruments and strategy documents, close to half do. The precision is absent where it would bite and present where it produces nothing.
Fifteen of the twenty binding instruments regulate general-purpose AI without writing down what general-purpose AI is. Put the other way round, the documents with the most carefully drafted definitions are the documents where getting the definition wrong costs nothing. The relationship looks less like coincidence than like a consequence of drafting practice. In a document with no consequences, a broad definition is free.
2.1The term international forums use most is the one domestic law uses least
Count which term each of the 52 definitions actually defines and the ordering is inverted. General-purpose AI, the umbrella category that international conferences and treaty discussions reach for most often, appears eight times. Generative AI, one of its subcategories, appears twenty-seven times and dominates. This distribution is also verifiable in the published data.
Terms taken as the object of a definition (denominator: 52 definitions across 17 jurisdictions)
Tallied directly from the definitions table in the published dataset.
Open the definitions and you can guess why. Most of them describe what a model produces rather than what it can do. It generates text, images, code. That is the easier sentence to write. Definitions that carry generality itself, the property separating general-purpose from narrow AI, into operative language are rare. Twenty-six of the 52, exactly half, describe a model; four describe a system. The paper adds one more line here. Outside Korea, not a single jurisdiction in this sample adopted the EU AI Act's general-purpose AI model classification by name. What did get adopted was the high-risk system classification.
2.2Jurisdictions with something to align to, and jurisdictions without
How scope gets drawn matters far more to a practitioner. Of the 17 jurisdictions that hold definitions, 15 set scope by asking whether the description fits your model. Only the European Union and Korea set scope by properties of the model itself, and only those two attach both a quantitative trigger and a power to designate a model into scope. Korea adds a use-context requirement on top. Singapore diverges in yet another direction and scopes by output modality alone, so a general-purpose model that generates no content sits outside while a small generative tool sits inside.
What that difference does in compliance work is straightforward. Where scope runs on a threshold or a designation, there is something to align to. A number, an authority, a register. Check whether your model crosses the number; if it does not, document why. Where scope runs on description matching, there is nothing to align to. Every market is a fresh judgment call, and the judgment leaves no record.
Definitions with a quantitative trigger number three across the entire sample. Two in the EU, one in Korea. And there is a common error in setting those two numbers side by side.
"A model trained between those two figures is therefore a general-purpose AI model with systemic risk in the European Union and falls outside Korea's class altogether."
The EU AI Act presumes systemic risk once cumulative training compute exceeds 1025 floating-point operations, and Article 52 leaves a provider room to rebut that presumption with evidence. Korea's Enforcement Decree sets 1026, one order of magnitude higher. But the Korean figure is not a presumption. It is one of three cumulative conditions that must all be satisfied: the compute threshold, application of state-of-the-art technology, and a risk of broad and serious impact on human life, physical safety and fundamental rights. All three have to hold before a model falls into the class. This triple-cumulative structure is confirmed twice over, once in the Korean definition record in the published dataset and once in the body of the paper. So the tidy "EU at 1025, Korea at 1026" contrast, ten times apart, erases the difference in kind between the two regimes.
On what a threshold is designed to filter, we have written before about a US bill that used revenue as its threshold, and the EU timeline with its actual scope of application is covered in our piece on the August 2026 AI Act deadline.
They built the bodies that measure and left off the power to act
The first two links in the chain have to be read together. Measuring a risk and verifying what the measurement produced sit in separate provisions, but they break in the same place. The measuring gets done, and what it produces goes nowhere.
Start with the first link. Sixteen jurisdictions hold at least one risk assessment provision, so the function itself is widely recognised. Only four have put any of it into binding law: Korea with five provisions, Brazil with three, and the European Union and Australia with one each. Of those ten, two reach the developer. In every other jurisdiction the function is acknowledged without being required.
Who bears the duty points the same way. Forty-six percent of risk assessment provisions fall on public bodies themselves rather than on regulated entities, and the United Kingdom alone accounts for a third of those with ten. On timing, fewer than one in three provisions reaches a model before it is released, and of the nineteen pre-deployment provisions, fourteen land on providers. Where the state does the assessing, it is invariably watching from outside after release. Not one provision borne by an authority is tagged pre-deployment, and many attach to no point in the chain at all: risk registers, coordination centres, monitoring mandates handed to bodies that will become regulators later.
But the place this link truly breaks is elsewhere. Provisions requiring the results of an assessment to be shared with a named institution number fewer than a quarter, and 65% do not require sharing at all. The paper's summary is short: most assessments go nowhere, and the analysis stays inside the organisation that produced it. Where the assessor is the authority itself, the omission means little, since the maker and the recipient are the same body. The problem is the majority of provisions that fall on providers or deployers, where there is simply no route for a finding to reach anyone who could act on it. The same defect that returns in incident reporting has already occurred once, at the first link in the chain.
The second link, evaluation and verification, asks whether what is claimed about a model actually gets tested, and by whom. One thing to establish up front. This is not a neglected domain. With 136 positive provisions it ranks near the top of the four, and sixteen of the twenty jurisdictions hold at least one. The gap is not a product of anyone being lazy. What is short here is legal force, not attention.
Filter those 136 by legal force and the number collapses fast. Six sit in binding law. Of those six, four actually impose an evaluation mechanism on someone, while the other two merely establish an evaluation body or procedure without saying what is to be evaluated. And of those four, exactly two reach the entity that built the model: Article 55(1)(a) of the EU AI Act and Article 30 of Korea's 2025 Basic Act.
How 136 evaluation and verification provisions narrow
nineteen out of every twenty carry no binding force
the other 2 build evaluation architecture only
EU AI Act Art. 55(1)(a) · Korea Basic Act Art. 30
Source: the paper's clause-level tallies (§5.4). The remaining two are the European Commission's ex ante conformity assessment and an audit requirement under a Brazilian judicial resolution.
3.1The one being tested is the one doing the testing
Move to who performs the evaluation and one default emerges. Of the 136 provisions, 90 record self-assessment, and 30 recognise an independent third party; the two overlap where an instrument offers a choice. The builder checks itself and so does the adopter. In eleven of the seventeen jurisdictions holding evaluation provisions, self-assessment is the majority model.
More interesting is where independence comes from when it appears, because it usually does not come from AI law. The strongest evaluation duty in the United Arab Emirates sits in a 2026 central bank directive addressed to licensed financial institutions, requiring an annual cybersecurity review of third-party AI suppliers to be performed by an external expert. That is ordinary vendor risk management, applied to AI because AI entered the vendor stack. The most explicit third-party certification architecture in the sample, the Dubai International Financial Centre's 2026 recognition and certification regime, was likewise built on accreditation powers the data protection authority already held. Bills in Kenya and Nigeria have regulators accredit AI auditors, again an assurance model borrowed from financial and standards regulation.
3.2Capacity gets built before obligations do
Sort the provisions again by who bears them and a substantial share turns out to be the state building its own capacity. Thirty place an evaluation duty on a state body, and 23 create evaluation architecture without imposing a mechanism on anyone at all. Safety institutes, sandboxes, standardisation programmes, evaluation toolkits, ministerial reporting cycles. Roughly one provision in six in this domain is a provision that builds an institution.
The United Kingdom, Japan, Canada, Nigeria, Singapore and the European Union are all doing this, and in the same order. Capacity first, obligations later. The European Commission committed to funding EU evaluation capacity including cybersecurity through 2027, and Canada's 2026 strategy promised to expand its safety institute. Nigeria's bill requires the minister to publish an annual report assessing compliance and sectoral impact, and Singapore proposes a baseline set of mandatory safety tests and an accreditation mechanism to come. As the paper puts it, none of these creates an obligation for providers.
3.3"Run the test" is written down; "what counts as passing" is not
The instruments name sixteen distinct evaluation mechanisms. Safety testing, capability testing, performance testing, red-teaming, controllability testing and human uplift testing generate evidence about what a model does. Audit, conformity assessment, certification, verification and accreditation check claims already made, or the party that made them, or the evaluator. Security testing looks at integrity. These measure different things.
The trouble comes next. Most provisions name the mechanism and leave thresholds, benchmarks and failure conditions blank. There is no statement of what passing means. The distribution also tilts toward mechanisms that generate evidence about the model. Safety testing leads with 21 provisions but spreads thin, and human uplift testing, the only mechanism that asks what a model adds to what a user could already do, appears twice in the entire sample. Controllability testing, which asks whether a system can be stopped, overridden or prevented from resisting shutdown, appears four times and only in two jurisdictions.
The paper calls this pattern convergence in silence. The AI middle powers converged on requiring tests and did not converge on what passing a test means. In most cases they did not address the question at all. Anyone who has watched a quality requirement without a metric collapse under audit will recognise the shape.
3.4Nowhere is there power to act on what gets found
This is the core of the report. Eleven jurisdictions have established an AI safety institute or an equivalent evaluation body. Not one of them, the European AI Office included, carries a duty to monitor prohibition breaches. And the paper goes a step further.
"The institutional capacity exists, but the legal connection to it does not. This disconnection is structural: no named safety or evaluation institute in the corpus holds enforcement powers, and of the 31 AI-specific governance actors, only four do, and one of which exists solely in a Bill."
The paper's conclusion states the same fact from the other side: eleven jurisdictions created bodies with an evaluation mandate, and the European AI Office is the only one with power to act on what its body finds, with the rest established as advisory. The two statements do not conflict. The Office is classified separately as an enforcement-holding regulator rather than as a safety institute, and its powers do not sit in the same place as the duties it bears as an evaluation body. Saying that capability and consequence live in different institutions is what you get when you lay the two sentences on top of each other. When quoting either, it is safer to say which one you are using.
Translated into actor counts: of the 144 governance actors mapped, 31 were created specifically for AI, and four of those hold enforcement powers. The European AI Office, the UAE's AI and Advanced Technology Council, the UAE Federal AI and Data Authority established in June 2026, and Kenya's Office of the AI Commissioner. The last one needs a footnote. It exists only inside a bill still sitting in committee. Kenya's AI Bill was published in February 2026 and had its first Senate reading in April; because it touches county government it must also go through the National Assembly, and as of August 2026 it has not been enacted. The status the paper recorded at its snapshot is unchanged four months later.
The paper finishes the arithmetic itself. Outside the direct effect of EU law, the purpose-built AI actors with power to enforce anything are two Emirati bodies and one Kenyan office that does not yet exist. The table below reproduces the nine actors classified as dedicated AI regulators in the published dataset. Those nine are a different population from the paper's 31: the AI-specific column in the public CSV is empty for all 144 rows, so the count of 31 cannot be reproduced, while the nine can be recovered directly from the actor-type column. The two calculations happen to land on four either way.
| Dedicated AI regulator | Jurisdiction | Enforcement power |
|---|---|---|
| Federal AI and Data Authority | UAE | Yes |
| AI and Advanced Technology Council | UAE | Yes |
| European AI Office | EU, France, Germany | Yes |
| Office of the AI Commissioner | Kenya | Yes — enabling law not enacted |
| AI Safety Institute | Australia | Advisory only |
| AI Safety Institute | Canada | Advisory only |
| AI Safety Institute | Germany | Advisory only |
| AI Security Institute | United Kingdom | Advisory only |
| IndiaAI Safety Institute | India | Advisory only |
The five internationally recognised AI safety institutes all sit in the lower block. Australia, Canada, Germany, the United Kingdom, India. These are bodies constituted to measure and to advise, and not because they lack competence. As the paper judges it, attaching a consequence to what they find is a question of mandate rather than of the documents they produce.
Prohibitions show the same shape. Only six jurisdictions name a monitor independent of the regulated party (the European Union, Taiwan, Japan, the United Kingdom, Singapore and the United Arab Emirates), and the monitor named is always an institution that already existed for another purpose. Courts, police, market surveillance authorities, procurement ministries, certification bodies. And never once an AI institute.
3.5Three days before the snapshot, the one exception got stronger
One spot on this map is already out of date, and it happens to be the middle of the argument. The paper's reference list cites the EU Digital Omnibus as a proposal with no number, linking to a Council document, with an access date of 31 July 2026. By then the regulation had already been published in the Official Journal on 24 July as Regulation (EU) 2026/1744 and had entered into force three days later on 27 July. Six days before the manuscript went out.
Reading this as an error in the paper would overstate it. The body of the paper codes Article 75(1a), which that regulation inserted, as binding and in force, and counts it as one of the three provisions that carry both a deadline and a consequence. The substantive judgment was right; only the bibliographic status lagged. Set alongside the authors' own line that this map dates within months, it is less a flaw than an illustration of the subject.
What matters is what went into that regulation. The European AI Office was expressly granted powers to request information, to conduct remote and on-site investigations, to make commitments binding, to find infringements and to impose fines and periodic penalty payments. Where an on-site investigation is obstructed, the competent authority of the member state concerned assists and, where appropriate, requests the support of police or an equivalent enforcement body. Its exclusive supervisory perimeter widened as well, so that systems built on a general-purpose AI model now fall under the Office not only when the same provider builds them but when providers within the same corporate group do.
So the change immediately after the snapshot did not narrow the gap. It widened it. With eleven jurisdictions holding evaluation bodies and none of them gaining power to act, the one that had power from the start acquired investigation, on-site inspection, binding commitments, fines and periodic penalties on top. The rest remain advisory. This is not a counterexample to the paper's conclusion but an acceleration of it.
3.6So evaluation runs outside the law
Read this far and you might conclude that pre-deployment evaluation of frontier models simply is not happening. It is, and the paper is precise about it. The absence lies in the law rather than in the coordinated governance landscape. A third-party evaluation ecosystem has formed over the past three years, and organisations such as METR and Apollo Research conduct pre-deployment evaluations of frontier models under voluntary access agreements with developers.
The terms of those agreements are not set by the evaluators. They report that access is inconsistent and that the time allowed is measured in days rather than weeks. International discussions ask for three things: evaluation before deployment, a third party independent of the developer, and an international body with audit and certification authority. None of the three appears in the provisions of this sample. Evaluation exists, but what requires it is contract rather than law. That distinction returns as a practical question in the final section.
The problem of the evaluated model drifting from the served one is covered in our piece on silent model updates, and Singapore's AI Verify, the representative case of publishing a methodology and nothing more, is covered in our report on Singapore's AI trust infrastructure.
When something goes wrong, who gets told
The remaining two links are prohibitions and incident reporting. They look like separate domains and share a single defect. Something is prohibited without anyone being told to notice a breach, and notification is required when something goes wrong without saying who must be told, or by when.
4.1Prohibitions target what people do, not what models can do
There are 127 prohibition provisions. Split them by what is prohibited and the direction is clear. Ninety target a use, 24 target a capability, and 13 target an outcome. What people do with a system comes up nearly four times as often as what a system is able to do. More than half the use-based prohibitions belong to the European Union, Brazil and Taiwan, while capability-based prohibitions cluster in the United Arab Emirates (10 of its 19), Australia (6) and India (3).
There is a reason for the skew. The areas where prohibitions concentrate, manipulation, child safety, human rights and the administration of justice, are areas that already have a body of law banning the conduct, a settled answer to who is bound, and an enforcement route that runs. Covering the AI version needs an amendment rather than a new regulatory regime. Capability-based prohibitions have no such precedent inside domestic AI law, which is why they show up in only a handful of jurisdictions.
The problem is when enforcement can happen. A capability can be tested before deployment, but a use or an outcome can only be enforced after the wrongful act occurs or the harm materialises. International red-line discussions are largely framed around capability categories and assume a structure in which developers demonstrate the absence of specific dangerous capabilities before deployment. The provisions in this sample fit that structure poorly.
Detection makes it sharper still. Of the 127, 56% say nothing about who would notice a breach. Among those that say something, more than half have the regulated party monitoring itself, and 30% name no institution as bearing the monitoring duty. Provisions naming a monitor other than the regulated party come to 14%.
And the strongest prohibitions in form are the least likely to be caught. Of 39 absolute prohibitions, 27 specify no detection mechanism, a noticeably higher rate than for conditional prohibitions or provisions restricting an activity. The reason lies in where absolute prohibitions sit. They sit in criminal law and in ethics charters, and those documents leave the surfacing of a breach to complaints, routine policing and reputational pressure. They impose no continuing duty on any party to monitor for violations. Conditional prohibitions, by contrast, sit inside administrative regimes where supervisory machinery already runs, and they can point straight at that machinery.
The paper's conclusion, transposed directly: the strictest measures are the hardest to enforce. Absolute prohibitions, the strongest in form and the most directly connected to international red-line agreements, are the least likely to have a breach detected. Worth adding that no prohibition provision in this corpus defines the capability threshold at which it would bite.
The grammar of prohibition splits in two as well, which matters in practice. Some jurisdictions simply prohibit; others permit subject to conditions. Peru, Chile, France, Germany and Singapore rely entirely on absolute prohibitions. An international agreement written in one grammar does not implement cleanly where the other is the operative form. With thresholds undefined and the sentence structure itself divided, a red-line agreement descending to the clause layer hits two walls rather than one.
4.2Incident reporting is where absence dominates
Serious incident reporting is the only one of the four domains where absence outweighs presence. Of the 105 provision records tagged to this domain, the operative ones, meaning those that impose or recommend a reporting duty in some form, number 27, and the remaining 78 are confirmed absences. The distribution of those 27 does not track overall governance activity either. India holds five reporting provisions out of 20 positive provisions in total, the most of anyone outside the European Union. Australia and the United Kingdom, with 45 and 44 positive provisions, hold two each in this domain, and the United Arab Emirates holds 32 positive provisions and none.
Eight jurisdictions have no self-standing reporting duty. Add France and Germany and it becomes ten, since neither has a domestic duty and their obligations arrive only through the direct effect of EU law, so the count is eight or ten depending on how you count. The paper uses both figures depending on context.
Whether an operative reporting provision actually completes the chain is judged on four elements: recipient, trigger, deadline, and consequence for failure. The OECD's proposed common reporting framework sets out 29 criteria for what a report should contain, and in this sample the provisions specifying both a deadline and a consequence number three of the 27.
| The 3 provisions that complete the chain | Deadline | Consequence for failure |
|---|---|---|
| EU AI Act Article 73 | Tiered schedule | Administrative fines |
| EU AI Act Article 75(1a), inserted by the 2026 Digital Omnibus | Inherits the Art. 73 schedule | Administrative fines |
| Brazil, CNJ Resolution 615/2025 (generative AI in the judiciary) | 72 hours | Corrective measures |
Even inside the EU AI Act, Article 55(1)(c) fails to complete the chain. It requires providers of general-purpose AI models with systemic risk to report serious incidents to the European AI Office and is backed by the same fines, but its deadline reads "without undue delay", so there is no clock. The paper adds a comparison here: mature reporting regimes outside AI take these elements together as a matter of course. The GDPR carries a 72-hour deadline, a designated supervisory authority, and penalties for failure to notify, all at once. A deadline without a recipient, or a recipient without a consequence, is the unusual arrangement.
4.3A borrowed channel borrows its triggers too
How the reporting channel was arranged splits too. Some built a new AI-specific channel; others routed AI incidents into a regime already running. The European Union builds. All five of its reporting provisions create AI-specific channels. Japan's two procurement provisions, Brazil's judicial resolution, Kenya's 2025 strategy commitment and Singapore's generative AI framework are on the same side. India, Israel and Korea borrow. Three of India's five ride on cyber incident response arrangements or intermediary liability rules, and two of Israel's three ride on anti-money-laundering and data protection regulation.
Of the eight borrowed routes, seven sit in guidance. The donor domains are four: financial crime, cyber incidents, data protection and health breach notification, and online safety channels. A structural problem rides along with them. A borrowed channel keeps the trigger it was built with. A duty attached to fraud reporting fires when fraud occurs, not when a model fails. A general-purpose AI incident that produced no data breach, no service outage and no suspicious transaction can pass through without switching on any duty at all.
One more limit runs through the whole sample. Every reporting duty asks a single actor to report events within its own span of control. Failures that only become visible across multiple models or multiple deployments therefore have no recipient anywhere in this corpus. The detection asymmetry identified in the same team's earlier incident response work compounds this: the developer holds both the visibility into a safety problem and the discretion over whether to escalate it. Actual incident counts are not falling. The AI Incident Database tally cited by Stanford HAI records 362 incidents in 2025 against 233 in 2024.
The broader structure of naming risks while never naming the party responsible is something we looked at in our piece on the accountability gap in AI risk taxonomies.
Governance was built on borrowed institutions
It is tempting to read every gap so far as a lack of preparation, and the paper explicitly warns against that reading. What is visible at the clause layer may be the trace of a different design rather than the trace of a failure. Time to put the weight on the other side of the scale.
Start with the actor numbers. Of the 144 governance actors mapped, 31 were created for AI. Eighty-four predate AI governance and later received a formal AI mandate, and the remaining 29 touch AI functionally without any formal mandate. Together that is 113 actors, four in five, that existed before the mandate they now exercise. By type, ministries lead with 53, followed by sectoral regulators at 33, advisory bodies at 21 and data protection authorities at 15. Dedicated AI regulators number nine and standards bodies two.
Look at enforcement power across all 144 and roughly half can enforce something. Forty-eight percent hold enforcement powers, 37% hold none, and 14% hold limited powers. The problem is not the total but the placement. The institutions with power and the institutions with an AI mandate are in different places. That is the result of loading general-purpose AI oversight onto ministries and existing sectoral supervisors instead of creating new institutions, and it is the instrument-layer pattern repeating one layer up. The obligations were built on borrowed regimes, and so were the bodies that carry them.
The consequence is that obligations go to whoever those institutions could already reach. Forty-six percent of systemic risk assessment provisions fall on public bodies, mostly on the side that evaluates systems it did not build. In evaluation and verification, 43 provisions point at deployers and 40 at providers. Outside the direct effect of EU law and Korea's 2025 Basic Act, almost nothing reaches a model before it is deployed.
5.1Sectoral routes actually work
There is one piece of countervailing evidence. Of the 101 instruments, 25 are sectoral, a quarter of the total. Those 25 carry 97 of the 382 positive provisions, again a quarter, which is to say they pull their weight. And within them, five financial instruments stand out. Of the 55 soft-law provisions in this sample with an identifiable enforcement route, 15 sit in four of them.
The reason is not the form of the document. It is the supervisory relationship. Financial supervisors already hold authority over the same firms, and that sector already runs an incident reporting route with a named recipient and a fixed clock. General-purpose AI risk does not arrive there as a new kind of harm but as a new route to a familiar one, which means the existing supervisory apparatus can be pointed at it directly.
The most striking case is not financial but judicial. Brazil's National Council of Justice, a constitutional body governing the courts, produced the only sectoral instrument in this sample covering all four domains, and one of the three complete reporting chains from the previous section lives there. Report within 72 hours, with corrective measures attached.
What gets inherited is not only the strengths, though. A borrowed route keeps the trigger it was built around, and it tends to produce guidance rather than statute. None of the five financial instruments is a statute. The trade buys speed and real compliance pressure at the cost of permanence. And a financial supervisor reaches licensed institutions and nobody else. The coverage of a sectoral route ends precisely at the boundary of its sector.
The definitional problem from section 2 is inherited along the same route. Sectoral instruments largely do not state the category of system they apply to, because coverage flows not from a statement about technology but from who is already supervised. The object is inferred rather than written. Of the five financial instruments in this sample, exactly one defines the category of system it covers: the UAE central bank's 2026 directive.
You can also measure how far these routes reach up toward the model layer. Of the 97 sectoral provisions, 25 fall on providers, and of those, four take the general-purpose AI model itself as the object of evaluation. Three in Singapore, one in the United Kingdom. Sectoral routes do work, but where they work is mostly after a model has already entered someone's business.
5.2Tier states legal form, not behavior
One proposition in this report is likelier than any other to be misquoted. The paper puts considerable effort into counting the legal tier of each instrument and then tells you not to predict behavior from that tier. Two cases at opposite ends of the sample show why.
Japan holds five instruments and no hard law at all. Four soft-law instruments and one strategic scaffold. Yet the coding does not read like advice. Every positive provision in the 2026 government procurement guidelines was coded as a requirement or enforcement provision, in a document that creates no legally enforceable obligation. A Japanese interviewee explained why to the researchers. Guidance issued by a ministry is understood by those subject to it as an instruction to comply, and the compliance it produces runs through sourcing, bidding and procurement rather than through enforcement.
"…though tier is a reliable statement of an instrument's legal form, it is an unreliable proxy for the behavior it produces: a jurisdiction with no hard law is not necessarily a jurisdiction without compulsion."
Add a case the paper does not use and the proposition holds in the other direction too. Korea, which has the most complete binding chain on the page of anyone in the sample, has not yet started the enforcement clock on it. In the legislative notice for the AI Basic Act's Enforcement Decree, the Ministry of Science and ICT stated that it would run a grace period on administrative fines for at least one year so that the regime can settle and firms can prepare. During the grace period, punitive measures such as fact-finding investigations and fines are in principle minimised and replaced by support-desk consultation. One thing not to misread: what is deferred is the sanction, not the obligation. The obligations themselves arise from the date of entry into force.
Set Japan and Korea side by side and the paper's proposition is confirmed in both directions. A country with no hard law where the rules bite, and a country with hard law whose enforcement clock has not started. Neither is visible from the text of the provisions alone. It is worth citing as the moment a study that counts provisions writes down the limits of its own method.
Korea built the upstream and borrowed the downstream
In this sample Korea occupies an unusual position. In the paper's words it is the only jurisdiction to legislate binding upstream obligations under its own power. Its own legislature wrote a definition, attached a compute threshold, and imposed risk assessment and evaluation duties on whatever crosses that threshold. That is the 2025 Basic Act and the 2026 Enforcement Decree, mapped as 16 provisions across five instruments.
The published data shows the same picture. Of the five binding instruments that define general-purpose AI, two are Korean. The other three are one in Australia, one in Brazil, and one shared by the EU, France and Germany. Korea is the only jurisdiction in the sample holding two such instruments. This cross-tabulation is not in the body of the paper and comes straight from the public CSV.
It also shows how far a ranking can diverge from actual legislation. Korea places fourth on the composite score the study used to pick its twenty, while France and Germany, near the very top of that same score, hold no domestic hard-law AI instrument at all. The paper says outright that the composite barely predicts what any country actually built.
What the provisions name as a risk says something about Korea too. Across the whole sample, the risk most often named in risk assessment provisions is cross-cutting risk with no specified domain, at 29 provisions, followed by fundamental rights and cyber at 21 each, democracy and societal manipulation at 15, loss of control and CBRN uplift at 11 each, and concentration of power at five. Catastrophic domains appear only where a jurisdiction regulates the model layer upstream or where an authority watches from outside. Yet Korea's Basic Act, the one instrument in the sample that created binding upstream duties under its own power, names no catastrophic domain at all. It is coded to cross-cutting risk, critical infrastructure, public health, fundamental rights and financial stability. A chain built to reach upstream, with the thing it is meant to catch written in a different vocabulary.
Grouping by region makes the position clearer still. Japan, Singapore and India lean on evaluations that the government performs or publishes itself, and hold neither binding evaluation duties on providers nor model evaluation duties. Taiwan skips horizontal legislation and adapts sectoral hard law to general-purpose AI, while Australia combines the largest body of soft-law evaluation material outside the EU with binding provisions that are entirely sectoral. In the paper's phrasing, Korea is the regional exception. International forums treat Asia-Pacific as a bloc, and in practice it is not converging on a common model: the distance between Korea's framework-act approach and the voluntary instruments of Japan and Singapore is as wide as the distance between either of them and the European Union.
There is a grouping in the other direction as well. Kenya, Nigeria, Brazil and Chile have all drafted comprehensive horizontal bills that combine mandatory duties with concrete governance institutions. Outside the European Union, these four are the only documents that would impose binding evaluation, risk assessment and incident reporting duties on providers all at once, and none of the four has been enacted. Three of them sit in the bottom half of the study's composite score. Does regulatory ambition move inversely to the presence of a domestic AI industry? The paper records this as a question its mapping raises but cannot answer, and leaves it as future comparative work.
6.1The fourth link is empty
Then comes the gap. Neither the Basic Act nor its Enforcement Decree creates a serious incident reporting channel. The only Korean reporting provision the paper coded sits outside AI legislation altogether. It is the 2026 financial sector AI guideline, which requires a financial institution to report immediately to the supervisory authority when an AI incident that could escalate into systemic risk has occurred or threatens to. The report rides the pre-existing financial incident reporting system rather than any AI-specific channel. It is a duty written by a financial supervisor for its own licensees, and it reaches one sector. The guideline is also voluntary, so nothing attaches to non-compliance.
Apply the four-element test from the previous section and Korea has a recipient and a trigger and lacks a deadline and a consequence. The paper attaches one caveat: a subordinate notice under the Basic Act, applying to high-performance AI operators above the compute threshold, was out for comment at the time of writing and could fill part of the gap. That the most finely drafted upstream regime in the sample borrows another regime downstream means the pattern running through every section above applies to Korea without exception.
The clock has not started either. By the government's own account no domestic model currently meets the compute threshold, and the grace period on administrative fines runs for at least a year. In the country that legislated the chain most carefully, no domestic company is currently caught by it. This does not need to be read as criticism. It is the structural consequence of a country without a frontier developer trying to regulate the model layer, which makes it the story of all twenty jurisdictions in the sample. Our separate treatment of Korea's high-impact AI regime is in the Basic Act governance report.
6.2So who receives the documents we file
From here it stops being regulatory news and becomes a practical question. What this study counts, in the end, is what gets produced and who it goes to, and the finding is that the "who" is mostly empty, or occupied by someone without the authority to ask a follow-up question. That produces four consequences for anyone doing data work.
One, the recipient of evidence moves. When a regulator holds no power to read a document and ask about it, demand for the document does not disappear. The address changes. Customer due diligence, procurement review, audit and partner contracts take the place. The METR and Apollo Research pre-deployment evaluations from the previous section are exactly this structure. They are not required by law; they run on voluntary access agreements with developers, which is why the evaluators cannot set the terms of access or the time allowed. Procurement producing compliance in Japan is the same mechanism. When the recipient is not a regulator, the design brief for the document is not a regulatory form but the other side's review process.
Procurement deserves more than a passing glance. An analysis the paper cites divides procurement conditions into two kinds. Conditions that specify what evidence a supplier must produce are neutral as to where a model was built, and Japan's 2026 procurement guidelines are the example. Conditions that favour domestic suppliers are classified as a first step toward protectionism, and no instrument in this sample takes that form. The most advanced models are built in two jurisdictions, but they are purchased and deployed in all twenty. If several sizeable economies align on the evidence they demand from systems entering their markets, the incentives shift regardless of where a developer sits. That is the leverage available to a country with no regulatory authority.
There is also a route to borrowing consequences. A jurisdiction with no means of attaching consequences at home can run its own serious incident database aligned to the harm categories used by the 2025 General-Purpose AI Code of Practice and share it with the European AI Office. The Office has powers to demand information from providers, conduct its own model evaluations, and require mitigation or withdrawal from the market. In the same passage, though, the paper attaches two caveats. Bilateral arrangements of this kind reproduce the asymmetry in which the better-resourced side sets the agenda and the methods, and the Office itself is under-resourced relative to its mandate, which could leave implementation of the Code of Practice led by providers rather than by the Office. Both points come by way of cited literature and should be read that way.
Two, "run the test" arrives without a metric. Sixteen evaluation mechanisms are named, and most provisions leave thresholds, benchmarks and failure conditions blank. A quality requirement with no metric does not survive an audit. The definitional problem compounds it: because most definitions scope by what a model outputs rather than what it can do, regulation framed around outputs places no requirement at all on the provenance and composition of training data. If the passing criterion never arrives, defining it yourself and documenting the basis is the cheaper path later.
Three, how many compliance packages are you going to build? In the 15 jurisdictions that scope by description matching, every market is a fresh judgment call. In the European Union and Korea, which scope by threshold and designation, a single number decides, and the two numbers disagree and mean different things. One work the paper cites notes that when implementing guidance is missing, firms build interim compliance packages on their own interpretation, and those harden into jurisdiction-specific systems that raise the cost of consolidation later. That observation comes through the paper rather than from the original, and should be read accordingly. The practical implication is clear enough: settling on a common evidence schema before things harden is the cheaper move.
Four, an argument about sequence. This is the paper's own third contribution. International debate concentrates on what to prohibit, while the more pressing question for AI middle powers may be what they can verify. And the distance between an evaluation being performed and an evaluation having consequences is the shortest span between the governance these countries already have and governance that actually binds a duty-holder. This sample supports the claim with data. These twenty built more evaluation capacity than law to use it with.
6.3Why South Africa came back empty
The thread left hanging in section 1 ends here. South Africa is the only jurisdiction in this sample with no in-scope provisions at all, and the one attempt to fill that gap was the draft National AI Policy. It was withdrawn before adoption in April 2026. The reason is not in the body of the paper. It is one line in the appendix published alongside the dataset. According to that appendix on Zenodo, the draft was withdrawn after fabricated citations were found in its reference list.
That line has to be used with its source flagged, because the appendix itself notes that some rows warrant re-verification before publication. Even with the caveat attached, something remains. When a document meant to fill a regulatory gap collapses because of unverified output, the gap is not filled and a second one opens. The question this report has followed from the beginning comes round here. Had someone held the authority to read the document and ask a follow-up question, those citations would have been caught before publication.
The same question points back at the study itself. The paper discloses that 116 of the 587 provisions were extracted manually and 373 with model assistance, and that 31 of the 52 definitions were also model-assisted. Arabic, German, Hebrew and Korean were verified through round-trip back-translation between two different models because the team had no speakers of those languages, with material discrepancies resolved by expert review. Provision text is preserved in the original language so that every coding decision can be checked against the source, each record carries a confidence flag, and low-confidence records required two-person verification. The codebook itself started as a model-assisted draft and went through fourteen major revisions. That a dataset mapping AI governance had to carry its own provenance and verification design is worth noticing in itself.
6.4How far to trust this report
Five caveats belong with any citation. One, the source is a working paper without peer review. Two, the data is a snapshot as of 30 July 2026, and the authors wrote that it dates within months. Three, the twenty are a purposive rather than a random sample, so nothing here generalises worldwide. Four, anything outside the four governance domains is out of scope. Copyright, labour, competition and procurement in general are not on this map.
The fifth matters most. The authors disclose the provenance of their own classification. The six dimensions used to sort definitions and the criteria for the four domains derive from the EU AI Act, the OECD and G7 texts, which means the finding that the European Union covers the field most completely contains both the fact that the EU does cover more and the origin of the framework doing the measuring. That reads less like an admission weakening the conclusion than like a statement of how far the conclusion can be pushed.
A line on the authors as well. This paper is not the output of a standing institute but of the Arcadia Impact AI Governance Taskforce, a research fellowship cohort. One co-author is affiliated with The Future Society, a co-organiser of the international AI red lines campaign. What makes that interesting is that the paper's third contribution is an argument about sequence, that "what can we verify" is more urgent than "what should we prohibit". It is a passage relativising the priority of the authors' own campaign, so the disclosure reads as context that raises confidence rather than as a weakness.
Finally, what this report could not confirm. The current status of the Nigerian and Chilean bills as of August 2026, how the UAE Federal AI and Data Authority operates in practice, the headcount and budget of the European AI Office, and the volume of safety and trust filings actually submitted in Korea. The last of those may simply have no public statistics yet, given how recently the regime came into force. Absence is recorded here as absence.
References
Primary sources
- 1.Schwab, J., Naidoo, N., Barazzutti, F., Lee, S., Machado, C., "Mapping General-Purpose AI Governance Across Twenty AI Middle-Power Jurisdictions," 2026-08-18 (working paper, not peer reviewed, data snapshot 2026-07-30). arXiv:2608.19278 — source of every provision, instrument, actor and definition figure in this report
- 2.Same authors, accompanying dataset. Zenodo, v1.0.0, CC BY 4.0. DOI: 10.5281/zenodo.21978946 — tables of 101 instruments, 144 actors and 52 definitions, plus the out-of-scope instrument appendix. The 587 clause-level records are not included in this version
Adjacent academic literature
- 3.Gomez, F., Ball, M., Harre, M., Preston, L., Schwab, J., Machado, C., "Designing Escalation Criteria for International AI Incident Response," 2026-04-25. arXiv:2604.23183 — source of the detection asymmetry concept. Two authors overlap with the primary source and the first author co-designed its codebook, so this should be read as the same team's earlier work rather than as independent prior research
- 4.Heim, L., Koessler, L., "Training Compute Thresholds: Features and Functions in AI Regulation," 2024-08-07. arXiv:2405.10799 — argues that training compute is currently the best available first-pass filter for selecting what regulatory oversight should cover. The claim that thresholds trigger nothing is not from this paper; it is the primary source's finding about its own sample
- 5.Okolo, C. T., Raji, M., "The Global Majority in International AI Governance," 2026-01-23. arXiv:2601.17191 — source of the observation that sectoral routes trade permanence for speed and compliance pressure
Law, policy and statistics
- 6.Regulation (EU) 2026/1744 of 8 July 2026 (Digital Omnibus on AI), OJ L, 2026/1744, published 2026-07-24, in force 2026-07-27. EUR-Lex — source for the AI Office's investigation, on-site inspection, binding commitment, fine and periodic penalty powers and for the widened exclusive supervisory perimeter (text checked 2026-08-24)
- 7.Regulation (EU) 2024/1689 (EU AI Act), OJ L, 2024/1689, 2024-07-12 — Article 52 (rebuttable presumption of systemic risk), Article 55(1)(a) and (c), Article 73
- 8.Korea, Framework Act on the Development of Artificial Intelligence and Establishment of a Foundation of Trust (2025) and its Enforcement Decree (in force 2026-01-22), Article 24(1) — the three cumulative conditions of compute, state-of-the-art technology and serious risk
- 9.Korea, Ministry of Science and ICT, legislative notice of the AI Basic Act Enforcement Decree (2025-11-12) — grace period on administrative fines of at least one year and the integrated support centre. Confirmed via Electronic Times reporting of 2025-11-12 (article, in Korean)
- 10.Korea, Financial Sector AI Guidelines (2026) — the only Korean serious incident reporting provision coded by the paper
- 11.Brazil, PL 2338/2023 — passed the Senate 2024-12-10, pending in a Chamber of Deputies special committee (checked 2026-08)
- 12.Kenya, Artificial Intelligence Bill 2026 (Senate Bills No.4) — first Senate reading 2026-04-02, referred to the ICT standing committee, not enacted as of August 2026
- 13.CNJ Resolução nº 615/2025 (governance of generative AI in the Brazilian judiciary) — 72-hour reporting of harmful events, with corrective measures
- 14.OECD, "Towards a Common Reporting Framework for AI Incidents," 2025 — proposes 29 criteria for what an incident report should contain
- 15.Stanford HAI, "AI Index Report 2026" — AI Incident Database tally of 362 incidents in 2025 and 233 in 2024 (cited via the primary source)
Related Pebblous reports
- 16.Pebblous, high-impact AI governance under Korea's AI Basic Act and what the safety-and-trust filings actually contain. report/korea-ai-basic-act-highimpact-governance
- 17.Pebblous, when the evaluated model and the served model drift apart. report/silent-model-updates-disclosure-gap
- 18.Pebblous, what the EU AI Act's August 2026 deadline actually covers. report/eu-ai-act-august-2026-deadline-reality
- 19.Pebblous, Singapore's AI Verify as the representative case of publishing a methodology and nothing more. report/singapore-ai-trust-infrastructure-2026
- 20.Pebblous, taxonomies that name risks without naming the party responsible. blog/ai-risk-taxonomy-accountability-gap
- 21.Pebblous, a US bill that used revenue as its regulatory threshold. blog/great-american-ai-act-revenue-threshold