Executive Summary

This report reads a single preprint that classified, end to end, how UK listed companies have written about AI risk in 9,821 annual reports. It went up on 1 October 2026, the UK AI Security Institute's societal resilience team funded the work, and the classification code and result data are open alongside it. Two figures from it will be quoted most often. Reports naming AI as a risk rose from 2.8% to 41.2%, and the reports describing an AI incident that actually happened can be counted on one hand.

Something else did not rise. Among the reports that name AI as a risk, the share that goes on to say something substantive has not moved from around 10% since 2023. The gap opened because more reports started writing, not because the writing thinned out year by year, and most of that middle ground is not copy-and-paste boilerplate either. Sentences that name where the risk sits and stop short of what the company will do about it account for the bulk of AI risk disclosure. The incident count does not transfer cleanly either. The author put the scorecard for that label in the same paper, showing it got nothing right in the validation sample, then reread every flagged passage in the corpus by hand and kept four. Between seven and four sits a verification step.

That is what the study reports; what follows is this report's reading of it. The author asks you not to read the gap as concealment. What an annual report tells you is not how often incidents happen but where incidents get written down. The question then moves from corporate honesty to form design. Which field on which form records the failure your company had with AI, and who is allowed to read that field? When it sits empty, there is nowhere for remediation, insurance or regulation to stand.

41.2%

Reports naming AI as a risk

2025 · of 1,561 reports. In 2020 it was 2.8%

10.4%

Of those, graded substantive

2025 · of the 643 reports naming a risk. Flat near 10% since 2023

78.5%

Middle grade, naming only where

2025 · of the 643 reports naming a risk. The boilerplate grade is 11.0%

4 reports

Describing an actual incident

2020–2026 · of 9,821 reports. What remained after the author reread the 7 the classifier flagged

1

Before the numbers came a counting rule

The study is a preprint posted to arXiv on 1 October 2026. It has not been peer reviewed, the author lists an independent affiliation, and the acknowledgements credit the UK AI Security Institute (AISI) societal resilience team for funding and for advice on the research direction. The classification code is under Apache-2.0 and the result data under CC BY 4.0. What kind of body the UK AI Security Institute is came up before, in our piece on its review of a US model.

Before reading any figure from the body, it is worth establishing what this study counted, because every percentage that follows stands on the rules fixed here. There are three of them: the gate that decides which passages become candidates, the classification that decides what label a passage gets, and the aggregation that decides how those labels add up to a verdict on a single report.

1.1Some words do not open the gate

A keyword gate comes first. A passage reaches the LLM classifier only if it contains at least one term from a fixed list. The list holds core terms such as artificial intelligence, machine learning, large language model and generative AI; techniques that clearly belong to AI, such as neural network, computer vision and natural language processing; named AI products and vendors; and applied terms such as robotic process automation, predictive analytics and chatbot. The author also pins down what does not get through.

Generic terms such as “data analytics,” “digital transformation,” and “automation” are not triggers on their own. Source: arXiv:2610.02281v1, §3.2

That choice carries a price, and the author states it directly in the limitations section: the gate creates deliberate omissions. A company that avoids the word AI and writes around it as "intelligent systems" or by the name of its own platform never gets picked up. So the figures here should be read as the volume of discussion explicitly labelled AI. The volume of AI-related discussion overall is larger than that.

Noise in the other direction is documented too. Short acronyms collide with unrelated usage. ML in a water utility's report is megalitres, and the LLM in a director's biography is a master of laws. Some terms are matched case-sensitively so that a capitalised AI does not trigger inside Shanghai or Chairman. Passages that clear the gate come to 24,189, and each one carries two paragraphs on either side of the triggering sentence along with it as context.

1.2Six labels, and two of the definitions carry this report

The stage-one classifier assigns each passage one or more of six labels: adoption, risk, vendor, harm, general/ambiguous and none. Two of those definitions are what the rest of this report leans on, and every figure that follows comes out of these two sentences.

Label The author's definition What it takes to qualify
Risk A downside or exposure attributed to AI The word risk appearing near the word AI is not enough; the downside has to be attributed to AI rather than merely mentioned near it
Harm A past, specific AI-caused incident A passage about something that could happen does not qualify. It has to have happened

Stage-one label definitions from §3.3 of the paper. The adoption label carries its own extra condition: first-person company description indicating actual deployment.

At stage two, risk passages are sorted into ten categories and scored 1 to 3 for how explicit the attribution is. A 3 means a causal expression ties AI to the risk inside the same sentence, and the author calls only a signal of 3 explicit. The same stage also assigns a specificity grade, which section 3 takes up separately.

Keyword gate 1+ core term → 24,189 passages clear the gate Stage one — one or more of six labels Adoption Risk Vendor Harm General None Multiple labels per passage allowed · trigger sentence carries 2 passages of context each side Stage two — risk-labeled passages only · 10 categories — competitive, cybersecurity, information integrity, workforce impact, etc. · Attribution strength 1–3 — a 3 ties AI and risk causally in the same sentence ("explicit") · Specificity grade — boilerplate / mid-tier / substantive (covered in section 3)

Pebblous reconstructed the paper's three-stage classification procedure (§3.2–3.3) as a diagram. Only passages that clear the keyword gate get a stage-one label, and only those labeled risk go on to receive a category, attribution strength, and specificity grade at stage two. Every percentage that follows is an output of one of these three stages.

1.3The denominator is filings, not companies

The number most likely to be carried over wrong is 9,821. That is a count of reports, meaning filings. The companies behind them number 1,362, drawn from a population of 1,469 UK-incorporated listed companies as the subset for which machine-readable reports could be obtained. For 2025 alone it is 1,561 documents from 884 companies. Every "% of reports" in the paper is weighted by document.

What mixing the two units does shows up inside a single section of the paper. For 2025 the document basis gives 41.2% and the company basis 41.6%, near enough identical, but in the partial 2026 data they separate into 66.4% and 72.0%. Two rulers on the same phenomenon disagree by 5.6 percentage points. Why checking the denominator before quoting a figure is a habit worth having got a long treatment in our piece on Microsoft's AI diffusion metric.

The 2026 figures are partial data, 339 reports gathered through April 2026. The author includes them to show the direction of travel and says not to treat them as final. This report uses the 2026 numbers only when talking about direction.

2

The breadth grew, the depth stayed

UK listed companies filed 1,007 annual reports in 2020, and 28 of them named AI as a risk. In 2025, 643 out of 1,561 did the same. As a share that is 2.8% to 41.2%, roughly fifteenfold, and roughly fifteenfold is the phrase the author uses in the conclusion. The raw counts go from 28 to 643, about twenty-three times, but part of that is the corpus growing, which makes it the wrong multiple to quote as growth.

The column printed next to it in that table is what this section is about. The author recorded, year by year, the share of reports graded substantive. Against all reports as the denominator it runs from 0.3% in 2020 to 4.3% in 2025. The distance between the two columns widened from 2.5 to 36.9 percentage points, and the author writes that this gap grew every year.

Year Reports AI named as risk Substantive Gap
20201,0072.8%0.3%2.5pp
20211,3284.2%0.6%3.6pp
20221,8534.8%0.9%3.9pp
20231,9059.8%1.2%8.6pp
20241,82830.4%3.3%27.1pp
20251,56141.2%4.3%36.9pp

From Table 4 of the paper. Both percentages use that year's total report count as the denominator. The 2026 figures are partial data, 339 reports gathered through April, and are left out of this table.

2.1Three curves from the same data, three different shapes

The author ran three different filters over the same data and laid the resulting curves on top of each other. Count every risk mention with no condition attached and 2025 gives 41.2%. Keep only the ones with an attribution signal of 3, meaning AI and the risk are tied causally inside one sentence, and it drops to 30.3%. Keep only the ones graded substantive and it is 4.3%. Put those three on one chart and they part ways.

0% 15% 30% 45% 2020 2021 2022 2023 2024 2025 41.2% 30.3% 4.3% Reports naming AI as a risk Explicit attribution (2025) Substantive only

Redrawn from Figure 6 and Table 4 of the paper. Both lines use that year's total report count as the denominator. The explicit-attribution filter appears as a single point because the paper gives only the 2025 value in the body text.

The reading the author puts in the figure caption runs like this. Filtering to explicit attribution lowers the level but leaves the growth trajectory largely intact, and filtering to substantive flattens the curve almost completely. The sentence that follows holds the study's conclusion in one line.

… restricting to substantive signals flattens the curve almost entirely (4.3% in 2025): most of the growth is in disclosure lacking specific mechanisms, controls, or targets. Source: arXiv:2610.02281v1, Figure 6 caption

2.2The gap did not open because the writing got thinner

The denominator has to change once more here. The 4.3% takes all 1,561 reports as its base. Take only the 643 that name a risk and the same 67 reports become 10.4%. How that 10.4% moved over time is the quietest and most important result in the study.

Moderate disclosure dominates in every year, and since 2023 the composition has been broadly stable (about 10% of risk-mentioning reports substantive) even as volume has grown; the higher substantive shares in 2020–2022 (up to 18%) rest on very few reports. Source: arXiv:2610.02281v1, §4.6

That one paragraph blocks two misreadings at once. First, there is no trend in this data of disclosure getting flimsier year by year. The substantive share has held near 10% since 2023. The 36.9-point gap opened because more reports started writing. Second, using the 18% from 2022 as evidence that things used to be better gets it wrong. Only 89 reports named a risk that year, and the share rests on 16 of them.

The other side of 2025 belongs in the record too. Of the 1,561 reports that year, 918 did not name AI as a risk at all. Four in ten wrote something, which is also to say six in ten wrote nothing.

3

Most risk disclosure names the spot and stops

"Not specific" usually means one of two things: nothing was said, or boilerplate that would fit any report was pasted in. What this study counted is neither. Sort the 643 reports that named AI as a risk in 2025 by specificity grade and the boilerplate level comes to 11.0%, the substantive level to 10.4%, and the remaining 78.5% bunches in the middle grade.

What a middle-grade sentence actually looks like is something the author illustrates directly while explaining the ruler. Read the three together and it becomes clear where the line is drawn.

Grade The author's example sentence What is there and what is missing
Boilerplate “AI is a rapidly evolving technology that may impact our business” Drop it into any report and nothing looks out of place
Middle “AI regulation may affect our compliance obligations” Names where the risk sits; no mechanism, no mitigation
Substantive “We allocated £5M to bring three high-risk AI systems into conformity with the EU AI Act by Q3 2025” A mechanism, plus controls, named systems and a measurable target

The three examples §3.3 of the paper gives for risk disclosure. The condition for the substantive grade is that the passage "describes a specific mechanism and provides concrete controls, named systems, or measurable targets."

The quickest way to check whether that ruler is lax is to look at something it actually graded substantive. One of the 2025 passages left in the open data belongs to Bloomsbury Publishing. The company wrote its estimated impact from AI-related risk in three bands, under £1m, £1m to £10m and £10m to £26m, and added that stakeholders including shareholders had been engaged on the issue throughout 2024/25. A passage earns the substantive grade when the monetary bands and the engaged parties are written down together.

3.1Hold the same ruler up and each disclosure type has a different shape

The specificity grade was not applied to risk disclosure alone. The author ran the same ruler over adoption and vendor disclosure, and stacking the three bands shows risk disclosure as the least specific of them. The share graded substantive is 10.4% for risk, 33.8% for adoption and 64.5% for vendor disclosure.

Substantive Middle Boilerplate Risk 10.4 78.5 11.0 Adoption 33.8 55.5 10.8 Vendor 64.5 24.7 10.7 Unit: % of the 2025 reports carrying each disclosure type

Redrawn from Table 7 of the paper. The long left end of the vendor band is not a sign of better disclosure, and the author pins that down on the spot: a vendor mention by definition names a provider, which already satisfies the named-system condition in the grading criteria. It is not directly comparable with the other types.

That disclosure does not mirror usage shows up when you set this against other data. The Office for National Statistics found, as of June 2026, that 29% of UK businesses were using at least one AI technology, rising to 35% among those with ten or more employees and 49% among those with 250 or more. Over the same period, the share of UK listed company annual reports that name an AI vendor even once is 19.2%. You cannot subtract one from the other and call the difference a gap, because the populations differ: the ONS surveys businesses of all sizes excluding finance and insurance, and this study looks at reports filed by listed companies. One direction is sayable. Disclosure is a subset of usage, and a subset of the cases a company judged important enough to write down. In the same release the ONS reads its own numbers as showing that AI has so far had a relatively limited transformative impact. Nothing in this data closes off the possibility that the writing is shallow because the usage is shallow.

The boilerplate grade sits between 10.7% and 11.0% across all three, near enough identical. What separates the types is not how much boilerplate they carry but how long the middle band runs, and it runs longest in risk disclosure. Companies write about what they are using in relatively specific terms, and about what it could break by pointing at the spot.

3.2The categories did broaden, but the category ranking is not a measure to trust

The author writes that risk disclosure broadened from a 2020 focus on competitive concerns out to cybersecurity, information integrity and workforce impact. That directional statement is supported by the study. What the same paper blocks is comparing category figures to the decimal or ranking them, because human and machine agreement on the ten risk categories is 0.33. Set against 0.73 for the stage-one type classification, that is less than half.

Prior work also keeps the growth in volume from being read as something peculiar to AI. A study tracking US 10-K narrative disclosure from 1996 onward reported that as reports got longer, redundancy and boilerplate rose while readability and specificity fell, and the UK study cites exactly that passage as the starting point for asking whether AI disclosure follows the same pattern. One more possibility sits on top of it. The UK study cites the argument that as generative AI increasingly writes narrative disclosure, the authorship and reliability of the analysed text changes in kind. No study has yet measured what share is written that way.

That documents get anchored to their own prior year as they grow longer was also measured before AI. A study tracking the MD&A section of US 10-Ks across 28,142 firm-years from 1997 to 2006 scored year-over-year change as a cosine distance. As of 2006 a filing sat 0.10 away from the same company's document a year earlier, and 0.64 to 0.75 away from the documents of other companies in the same industry. Corporate narrative disclosure is far closer to its own last year than to anyone else's writing. Prior work cited there found that cautionary language, once laid down, is not removed even after litigation risk falls. Read the UK study's 11.0% boilerplate without that baseline and it tilts toward an impression that AI disclosure is uniquely formulaic.

Taking sentences companies wrote about themselves as your data puts a known problem in the scale, and that problem showed up in the same shape in a survey that asked about data readiness by self-report. What makes this study different is that instead of asking for self-reports, it held one ruler up to documents already filed.

4

Verification stands between seven and four

One of the six stage-one labels is harm. It attaches to a passage describing a specific, AI-caused incident that already occurred. The automated classification raised that label on seven of the 9,821 reports. By passage it is nine, across the whole span from 2020 to 2026. This is the figure from the study most likely to be quoted, and the one that most needs care in the carrying.

The author put the reason for that care into the paper itself, by printing the harm label's scorecard from the validation stage exactly as it came out.

The validation set contains no true harm instance: all 40 harm-tagged passages were false positives. Source: arXiv:2610.02281v1, Table 2 caption

Of the 474 human-labelled passages, zero qualified as harm, and the classifier put a harm label on 40 of them. Precision 0.00. So the author went back and reread by hand all nine flagged passages in the corpus, and judged that five passages across four reports do describe an AI-caused incident that occurred. What sits between seven and four is not a trace of overstatement being walked back but a record of verification working. Wanting to inspect and finding no record to inspect came up in our piece on the AI audit trail at a financial regulator, and this case runs the other way. The record existed, the label on it was wrong, and a person counted again.

4.1The harm label fails in a different way from the rest

There is a conclusion that comes too easily here, which is that the classifier as a whole cannot be trusted. The author's account blocks it. Recall is consistently high, between 0.91 and 0.97, across every label that could be scored. The harm label had nothing to get right, so it carries no recall figure at all. For the labels where precision is low, the author gives a separate cause.

… manual inspection of low-agreement multi-label categories (especially risk) suggests that much of the gap reflects higher LLM label completeness: the classifier tends to assign more, and frequently valid, labels per passage than the annotator initially did. Source: arXiv:2610.02281v1, §3.4

The harm label is the one that explanation does not reach. There is no gap here created by an annotator under-labelling. There were zero correct labels to assign, and the classifier assigned 40. Two different kinds of failure sit mixed together in one table, and a single averaged metric hides the difference between them.

This is not a problem specific to this classifier. It is in the nature of metrics for rare events. Work on how precision and recall shift with prevalence showed a case where the same classifier at the same accuracy gives precision 0.60 on balanced data and 0.33 on imbalanced data. Prevalence of harm in this corpus is seven in 9,821, about 0.07%. At that rarity, precision 0.00 does not point at an incapable model. It is how the metric behaves in front of a rare event. A replication study covering 27 tasks that used LLMs for annotation reached the same place when it wrote that researchers using generative AI for automated annotation must always validate.

4.2The number this report leans on hardest cannot be recounted from outside

The study is unusually open. The code is Apache-2.0, the result data CC BY 4.0, and even the report-level aggregation rule is written into the body: a report takes the most frequent grade among its passages, and ties round up. Apply that rule to the open data and 2025's 67, 505 and 71 come out without a cell out of place. Corpus size, the year-by-year trend, the market split, the year-over-year change by sector, the specificity distribution and the vendor tallies can all be recounted from outside.

One cell sits where it cannot be recounted. What is published is the 474 rows of human labels in the validation set, and that label distribution matches the support column of Table 2 exactly. What is not published is the record of what the classifier assigned to those 474 passages. Searching the repository and the release assets turns up no prediction log. So precision 0.00 and the 40 false positives exist only as the author's narration. The denominator verifies completely and the numerator cannot be recomputed from outside.

The author adds one more caveat on top. A single annotator produced the reference labels, and wherever the human labels and the classifier output disagreed, every such case was reread and settled toward whichever side was judged correct. With no independent second annotator there is also no human-to-human agreement baseline. The author's phrase for it is a curated benchmark rather than an unambiguous ground truth.

This is not a fact that counts against the author. It says that being reproducible and being fully recountable are different layers. And the argument of this report applies to the paper itself. What is not left on the record does not get verified. Why opening the intermediate artefacts of a verification run alongside the results is a quality question becomes visible right here.

4.3"AI harm" does not come in one shape

The nine passages the automated label flagged are in the open data verbatim. The author's adjudication notes are not published, so which combination of those nine survived as the four cannot be settled from outside. What can be read directly is what a specific incident description with a monetary figure actually looks like, and the two examples are quite different from each other.

Company (2025 report) The event as written Where AI sits
GEAR4MUSIC An outsourced AI-driven marketing system ran into problems, disrupting the first half of the financial year, and a misallocation of costs led to marketing overspend A tool the company was using malfunctioned
TRIFAST Two deepfake messaging frauds impersonating executives were reported as internal control incidents. The loss is a £400,000 provision plus an estimated £200,000 judged unrecoverable An outside attacker used AI as a weapon against the company. The company is the victim

Taken from the 2025 harm-tagged passages in the open data (CC BY 4.0). Both companies wrote this in their own public reports, and they are cited here only as real examples of specific incident description. Exactly which combination the author's manual rereading left as the four cannot be confirmed, because the adjudication notes are not published.

Can those two go in the same field? Anyone who has designed an incident form knows how early this question splits. One is a failure of a tool we chose, and the other is the instrument of an attack aimed at us. Accountability, remediation and insurance all run down different tracks from there. Where the definition of a reportable event starts counting an incident came up with a different case in our piece on whether something that happens during model evaluation counts as an incident.

5

Some write less, and the gap comes with conditions

The variation hidden behind the overall average is substantial. For 2025, 57.8% of reports from Main Market companies named AI as a risk, against 6.4% of reports from AIM companies. That is 51.4 percentage points, a factor of nine. Take the FTSE 100 on its own and it is 68.6%, more than tenfold. The author singles out this gap in the conclusion as well.

Four conditions come attached to that comparison, and it is more accurate to keep all four in the same paragraph as the figure.

  • Of the 1,561 reports from 2025, 687 drop out of this comparison: 676 with no market segment assigned and 11 from Aquis. What goes into the comparison is 703 Main Market reports and 171 AIM reports.
  • Recompute with a corrected mapping that pushes most of the unassigned into the Main Market and it becomes 46.9% across 1,287 reports. The direction of the gap survives; the width narrows.
  • Coverage of the data source, as the paper states it, is about 95% for the FTSE 350 and about 28% for AIM. The two populations were not swept at the same density.
  • The validation set was drawn from large Main Market reports published in 2023 and 2024. As the author writes directly, that set does not independently verify classification performance on AIM reports.

With every condition attached, the direction still holds. Reading that direction as "AIM companies are under no obligation to write it" gets it wrong, though. The author sets out the legal position in a footnote. A UK-incorporated AIM company is also required by section 414C of the Companies Act to describe principal risks and uncertainties in its strategic report. What does not apply is the Listing Rules, the Disclosure Guidance and Transparency Rules, and the Corporate Governance Code. AIM's own Rule 26 asks a company to state on its website which governance code it follows; it does not set what goes in the annual report.

5.1Two sectors sit low side by side for reasons that are not the same

Split the data across the twelve critical national infrastructure sectors and the lowest are energy at 20.0% and data infrastructure at 10.0%. Next to each other at the bottom of the table, the two are easy to bundle into one story, and the author adds a sentence immediately after to stop that.

The two differ, however: Energy is low on every signal … whereas Data Infrastructure reports adoption at a mid-table 55% … yet rarely discusses AI as a risk. Source: arXiv:2610.02281v1, §4.5

Put the other signals for the two sectors next to each other and the statement turns into numbers. In energy, 30.7% of reports mention AI and 21% report adoption. In data infrastructure, AI mentions run to 60.0% and adoption to 55%.

Sector 2025 reports AI mention Adoption Risk mention
Energy14030.7%21%20.0%
Data infrastructure2060.0%55%10.0%

The two rows this section needs, from Table 6 and §4.5 of the paper. For a reference point, the sector with the most reports is finance, at 657 reports with 74.7% mentioning AI and 47.0% naming a risk.

Three caveats come with the 10.0% for data infrastructure, and quoting it without them is not on. First, it rests on 20 reports from 2025, and the author writes that 20 reports make the estimate imprecise. Second, sectors with fewer than 30 reports carry a note to treat them as indicative only. Nuclear at 3, water at 17, defence at 22, government at 23, chemicals at 24 and telecoms at 29 all fall below that line. Telecoms' 26.7-point year-over-year rise also rests on 29 reports. Third, not one of the 15 companies used in validation is a data infrastructure business, and the author writes that down too.

One fact survives all three caveats. A sector that reports adoption at 55% reports risk at 10%. It is not that they write little because they use little; the item they use and the item they write down are out of alignment. The precision of the numbers cannot be trusted, and the misalignment itself is still visible inside 20 reports.

5.2The rule the paper describes has since changed

One footnote about AIM needs its date checked before it is read. On 5 August 2026 the London Stock Exchange replaced the governance provision in AIM Rule 26. The old version worked on comply-or-explain: state which code you apply, how you comply with it, and where you depart from it, why. The new version states that it neither requires nor expects comply-or-explain against that code, and instead mandates disclosure of five items: board composition, the role of each director, remuneration and performance, the risk and control framework, and investor relations. The item on each director's role includes the effective management of risk.

None of which makes the paper wrong. The footnote is written against the old version, and the old version is what applied across most of the 2020 to 2026 span this study covers. What is worth marking is the timing. The institutional condition behind the gap this paper measures moved two months before the paper went up, and it moved in the direction of making risk and control a disclosure item. AIM Rule 26 governs website disclosure, though, so it is not the same field as what goes in the annual report. Which way next year's data moves is answerable by running the same pipeline once more.

6

The rules allow vagueness, and no form has an incident field

Before reading the 78.5% middle ground as the product of companies writing without care, the right order is to check which rules those sentences were written under.

6.1What the rules require stops at "describe"

Writing AI risk into the strategic report of a UK listed company happens under four conditions at once. Lay the four on top of each other and you can see where the middle ground comes from.

Condition What it says Source
What the law requires Describe the principal risks and uncertainties facing the company Companies Act 2006, s.414C(2)(b)
The specificity ask Entity-specific description is recommended, and the guidance states of itself that it is non-mandatory best practice FRC Guidance on the Strategic Report
Audit scope The auditor's opinion is confined to the financial statements and offers no assurance on the strategic report Same guidance
Liability scope Liability to third parties is excluded, and liability to the company requires knowledge or recklessness Companies Act 2006, s.463

The first three were checked against the legislation and the FRC guidance itself; the last against the substance of the provision. The Corporate Governance Code provision requiring a declaration on the effectiveness of internal controls applies to companies in the listing categories to which the Code applies, for financial years beginning on or after 1 January 2026, and AIM is out of scope.

Read the four conditions in sequence and one sentence comes out. When the paper says 41.2% named AI as a risk and only around 10% of those were substantive, the rest are not sitting somewhere the rules were broken. They are sitting where the rules ask nothing. The more specifically a sentence is written, the more verifiable it becomes and the more it turns into material for a dispute, and there is weak incentive to make that choice inside a document that is not audited and whose liability turns on knowledge or recklessness.

The legal side pulls both ways too. The puffery defence developed in US case law protects optimistic statements that cannot be objectively verified, at the materiality stage. In the other direction, the safe harbour in the US Private Securities Litigation Reform Act requires meaningful cautionary statements, so earning the protection means writing more specifically, not less. A device that protects vagueness and a device that demands specificity sit at different stages of the same process. It is possible to read the middle ground as the narrow path between them.

All of that is a mechanism offered, not a causal link established. We found no empirical study testing the relationship between the specificity of AI risk disclosure and litigation risk directly. The empirical work showing that the length and wording of risk-factor disclosure changes when litigation risk changes covers general US risk factors and does not isolate AI. There is no basis today for writing the sentence "UK companies wrote vaguely for legal reasons."

6.2The author puts four readings side by side

The author stops at the same place. Calling the gap between volume and quality the most policy-relevant observation in the study, the author offers four explanations for it, not mutually exclusive and with no ranking between them. Declining to pick one is itself the paper's posture.

  • Limits of the medium — the annual report as a document may structurally favour cautious, general language about an emerging risk. If so, low specificity is a ceiling in the medium rather than evidence of a management failure.
  • Immature management frameworks — AI risk management is new at most companies. A company may recognise the risk and still lack the process, the expertise or the internal data to write it down specifically.
  • Disclosure as signalling — the writing may exist to tell regulators and investors that the company is aware. That fits the finding that the general/ambiguous type is the most common and has the largest absolute growth since 2020.
  • Rational vagueness — spelling out AI risks and mitigations can invite legal exposure or an adverse investor reaction. A company may rationally choose vagueness.

Which of the four it is decides what there is to do. If it is a limit of the medium, a different instrument is needed. If the frameworks are immature, it improves with time. If it is strategic vagueness, it takes an intervention that forces specificity. Those three are the branches the author separates in the conclusion as well, and the author adds there that which one it is remains a question this study can raise but not yet answer.

6.3The UK has no desk that takes an AI incident

The author names three places to look instead of annual reports: regulatory enforcement records, incident databases and litigation records. Checking the state of those three in the UK as of October 2026 makes the shape of the first route clear.

Desk What it takes Status
Information Commissioner's Office (ICO)Personal data breaches, reportable within 72 hours of awarenessIn force
Financial Conduct Authority (FCA)Operational incidents generally. Not AI-specificEffective March 2027
Bank of EnglandOperational incidents at financial market infrastructuresEffective March 2027
MHRAAdverse events involving medical devices, AI includedIn force
Cyber Security and Resilience BillCyber incidentsIn passage

Checked against each body's published documents. That this table has no row for AI incidents is the point of this section.

Personal data leaks out, you go to the ICO. A financial institution's systems go down, you go to the FCA. A medical device malfunctions, you go to the MHRA. Nowhere takes the bare fact that an AI got something wrong and a person was harmed by it. The Centre for Long-Term Resilience, in written evidence submitted to Parliament in 2024, noted that the Department for Science, Innovation and Technology has no centralised, up-to-date picture of incidents involving AI systems and recommended establishing a central incident reporting system. As of 2026 that recommendation has not been implemented.

One self-assessment from the regulatory side captures the character of this gap well. In its 2026 AI Airlock Phase 2 report, the MHRA wrote that current post-market surveillance assumes performance change arrives as a discrete, detectable event. AI performance degradation and automation bias, though, are detected only after harm has occurred. The shape of incident the form assumes and the shape of the incident that happens are out of alignment.

The second route the author names, incident databases, is not a reporting desk either. The AI Incident Database is a curated archive built from voluntary submissions and media collection, and by incident number it stands at roughly 1,700 as of October 2026. The OECD AI Incidents Monitor is still in beta and collects automatically from news, with no submission function at all. The cumulative totals of the two archives are not numbers you can place side by side, because their collection units and criteria differ. They share one property: neither receives anything by obligation. Our piece on a UN briefing covered the path by which an incident record becomes evidence, and a US state case where layoff notices started asking whether AI was the cause is the other face of the same gap.

6.4The EU laid the evidence down before the duty to report

Set the EU AI Act alongside it as a contrast and the difference is plain. What the EU built is not one reporting duty but the chain of evidence that makes reporting possible. Four articles run in sequence.

  • Article 12 — high-risk AI systems must be capable of automatically recording events over the system's lifetime.
  • Article 19 — those logs must be kept for a period appropriate to the purpose, at least six months.
  • Article 72 — providers must actively and systematically collect, document and analyse the relevant data.
  • Article 73 — serious incidents, once established, are reported to the authorities: 10 days for a death, 2 days for a widespread infringement, otherwise 15 days after the causal link is established.

The UK has none of the first three steps. Which is why four incidents written into annual reports does not read as a problem of corporate attitude. Reporting becomes possible only after a record is made to exist, made to survive, and gathered to be read, and those earlier steps are not laid down as institutions. A structure where the provisions exist while the authority and resources to enforce them sit elsewhere was set out with other countries' cases in our piece on the enforcement gap in middle-power AI regulation. Here the problem sits one step earlier than enforcement: there is no record.

6.5The US study points the same way with a different ruler

The US has its closest counterpart study. Sweeping more than 30,000 10-K filings, it reported that 4% of filers mentioned AI in the risk factors item in 2020 and 43% did in 2024. The UK paper calls it the most comparable study and notes that the two results agree in direction. And that is as far as it goes.

Axis UK study US study
DenominatorAll reports filed that yearAll companies filing that year
What is measuredAn LLM judges whether AI is described as a riskWhether any of 25 keyword terms appears in the item
Latest year2025 (41.2%)2024 (43%)
Specificity metricFigures graded on three levelsNone

Compiled by setting the methods sections of the two papers against each other. The denominators and units effectively overlap; what is measured, and the year, do not.

Place 43% and 41.2% side by side and write "the two countries look similar" and the two numbers appear to have counted the same thing, when one counts the appearance of a keyword and the other judges the character of a description. Move to the specificity axis and there is no place to compare at all. The US study states in its own limitations that it did not use a systematic method to quantify how much of AI risk disclosure is boilerplate. No US number exists to set against the UK's 10.4% and 78.5%.

None of this needs reading as a weakness in this report. It is section 1 getting demonstrated once more. Two teams looked at the same phenomenon with different rulers, and the two numbers do not meet. That is why the counting rule has to be checked before the figure is quoted.

6.6Korea's business report has no matching field at all

Running the same study in Korea would start with deciding where to take the denominator from, because the Korean business report (사업보고서) has no standalone item corresponding one-to-one with "principal risks and uncertainties" in the UK strategic report. A standalone section called Risk Factors exists in the securities registration statement, not in the statutory contents of the business report. AI risk scatters into the risk management narrative inside Business Overview, or into the Management's Discussion and Analysis. As far as we could confirm, the corporate disclosure form standards do not designate AI risk as a separate item either.

The shape of the incident desk matches the UK's. The AI Framework Act, in force since January 2026, requires operators above a certain size to maintain a risk management system and to submit the results of that implementation periodically, but it is not built around notifying the authorities when an individual incident occurs. The 72-hour notification and report under Article 34 of the Personal Information Protection Act is close to the only route that works in practice. So in Korea too, the institutional path by which an AI incident gets recorded opens when that incident came with a personal data leak. How responsibility for data incidents climbed to the board is set out in our piece on the amendment to that Act.

Why This Matters to Pebblous

A study that left its diagnosis on the record

Pebblous works on asking what state data is in before it enters training. DataClinic diagnoses incoming datasets, and AI-Ready Data covers the cleanup that happens before training. What makes this study unusual is that it left the diagnosis itself on the record. It classified 9,821 reports, measured performance against 474 passages, and when the most important label came back at precision 0.00, it printed that number in a table and then went back over every flagged passage in the corpus by hand. Anyone who has run an automated classification pipeline knows this order. Measure, see where it breaks, and have a person look again at the broken part.

The contrast from section 4 turns into a working lesson here. Most of the low-precision labels were low because the annotator had under-labelled, and the harm label alone was different: there was nothing to get right. A rare label fails in another way, and a single averaged metric erases that difference. That is the reason to look at per-label performance separately, and to write human rereading into the procedure for sparse labels.

A filled field and a field that tells you what to do are different questions

Data quality discussions usually cover missing values, duplicates and ranges. The defect this study exposes has a different grain. The 78.5% of risk disclosure is sentences that pass the schema completely. The word AI is there, the word risk is there, and the two are tied causally in one sentence. Not empty, correctly formatted, on topic. And with nothing in it about what will be done, or how.

The line the author draws is clear. A sentence saying AI regulation may affect compliance obligations is middle; a sentence saying how much is allocated to compliance, for which systems, by when, is substantive. Whether a field is filled and whether that field tells you what to do are different questions, and this is where the second one becomes as important as the first in designing a quality metric.

Two things you can try today

One is to hold the same ruler up to your own organisation's documents. Pull the passages where AI appears, from a business report or an internal risk register, and set them against the author's three grades. Boilerplate, or does it name where the risk sits, or does it say what will be done and how? The paper put the full production prompt in an appendix and published the code and data, so that ruler can be borrowed as is.

The other is a question. Which field on which form records the failure we had with AI? If not the annual report, then where? As section 6 showed, neither the UK nor Korea has that field set aside, and the desk opens only when personal data leaked. The EU, by contrast, laid down the steps of making, keeping and gathering the record before the duty to report. Record first and report second is not an unfamiliar order to anyone who has designed a data pipeline.

If you do decide to create the field, the two cases from section 4 become the design question immediately. Can an AI tool the company was using malfunctioning and an outside attacker using AI as a weapon against the company go in the same field? It is the first fork in building an incident taxonomy, and if the split is not made there, nobody will later be able to say what the aggregate number counted.

Before the ban list comes the record list

Discussion of AI management has dwelt heavily on what to prohibit. What this study shows is the step before that. What gets left on the record. Without a record there is nowhere for regulation, insurance or remediation to stand, and with a record that stops at naming where the risk sits, there is nothing to be done with the reading.

What Pebblous can do is put that distinction into Korean first, from the data side. That a filled field and a field that tells you what to do are different things; that the instrument for measuring the difference is already published; and that the rarer the event, the more an averaged metric lies. All three are sentences that transfer into a data quality specification.

Every verbatim quotation in this report was checked directly against the full text of the arXiv preprint, the data in the public repository, the UK legislation, the FRC guidance and the EU AI Act provisions. We could not open the Official Journal text for the EU provisions and confirmed them in a republished version instead, so the application dates and amending regulation numbers are not used in the body; only the substance of the articles and their deadlines. Which four reports the author's manual rereading kept, and what labels the classifier assigned to the 474 validation passages, are not published, and we did not fill those gaps with other numbers. Sections 1 through 5 are what the study and the primary documents report; the later part of section 6 and this section are the part those documents do not cover, so please read them separately. Thank you for reading this far.

References

The subject of this report

  • 1.The AI Risk Observatory: What Can We Learn from AI Disclosures in Annual Reports About Societal Resilience?. arXiv:2610.02281v1 [cs.AI], 1 October 2026. ⚠️ A v1 preprint, not peer reviewed. The author lists an independent affiliation, and the acknowledgements record funding and directional advice from the UK AI Security Institute (AISI) societal resilience team. Every verbatim quotation in this report was checked against this full text.
  • 2.AI-Risk-Observatory — code and data repository (release dataset-v1.1). Code Apache-2.0, data CC BY 4.0. The table figures in sections 2, 3, 4 and 5 were recounted against this data. The classifier's predicted labels are not included.

UK rules — legislation and guidance

  • 3.Companies Act 2006, s.414C — the duty to describe principal risks and uncertainties in the strategic report. It applies to UK-incorporated AIM companies too.
  • 4.Companies Act 2006, s.463 · Financial Services and Markets Act 2000, s.90A — limits on liability for statements in the strategic report. Confirmed from the substance of the provisions.
  • 5.Financial Reporting Council, Guidance on the Strategic Report (February 2026 edition) — §1.5 and §1.17 recommending entity-specific description while stating that the guidance is non-mandatory best practice, §2.13 confining audit scope to the financial statements, and the scope table at §9.1.
  • 6.Financial Reporting Council, UK Corporate Governance Code 2024, Provision 29 — applies to the listing categories to which the Code applies, for financial years beginning on or after 1 January 2026. AIM is out of scope.
  • 7.London Stock Exchange, AIM Rules for Companies Rule 26 — both the January 2026 edition (the version the paper's footnote is written against) and the edition effective 5 August 2026 were checked. The new edition drops comply-or-explain and mandates disclosure of five items including the risk and control framework.

Routes by which incidents get recorded

  • 8.Regulation (EU) 2024/1689 (EU AI Act), Articles 12 · 19 · 72 · 73 — automatic logging, retention for at least six months, active and systematic collection and analysis, and serious incident reporting deadlines. ⚠️ We could not reach the Official Journal text and confirmed these in a republished version, so the application dates and amending regulation numbers are not used in the body.
  • 9.Centre for Long-Term Resilience, written evidence submitted to the UK Parliament — the observation that DSIT has no centralised picture of AI incidents, and the recommendation to establish a central incident reporting system.
  • 10.MHRA, AI Airlock Phase 2 Programme Report (2026) — a regulator's own assessment that current post-market surveillance is structurally unable to detect AI performance change. FCA PS26/2 (published 18 March 2026, effective 18 March 2027) — operational incident reporting, not AI-specific.
  • 11.AI Incident Database (Responsible AI Collaborative) · OECD AI Incidents Monitor (beta) — neither is a reporting desk. Their cumulative totals use different collection units and are not set against each other here.
  • 12.Personal Information Protection Act of Korea, Article 34 (with Enforcement Decree Articles 39 and 40) · Framework Act on the Development of Artificial Intelligence and Establishment of a Foundation of Trust (in force 22 January 2026) — the two Korean routes. We could not obtain the full text of the Framework Act, so only its structure is described, with no verbatim quotation.

Comparators and prior work

  • 13.Uberti-Bona Marin, L. G., Rijsbosch, B., Spanakis, G., Kollnig, K. (2026). Are Companies Taking AI Risks Seriously? A Systematic Analysis of Companies' AI Risk Disclosures in SEC 10-K forms. arXiv:2508.19313 / ECML PKDD 2025 Workshops, pp. 86–106, Springer. DOI 10.1007/978-3-032-19096-3_6 — the US comparator in section 6.5.
  • 14.Chiu, I. H.-Y. (2025). Using generative artificial intelligence in corporate narrative reporting. Cambridge Forum on AI: Law and Governance 1:e43. DOI 10.1017/cfl.2025.10038 — on generative AI writing disclosure. It does not measure a share.
  • 15.Dyer, T., Lang, M., Stice-Lawrence, L. (2017). The evolution of 10-K textual disclosure. Journal of Accounting and Economics 64(2–3): 221–245 — the pre-AI baseline. ⚠️ We could not access the full text and used only the directional statement. Brown, S. V., Tucker, J. W. (2011). Large-Sample Evidence on Firms' Year-over-Year MD&A Modifications. Journal of Accounting Research 49(2): 309–346 — primary evidence that narrative disclosure anchors on copying the prior year's document. The 2006 figures of 0.10 against 0.64–0.75 in the body come from p.320, and the citation to Nelson & Pritchard (2007) on cautionary language not being removed from p.314. The difference score is one minus the cosine similarity with the prior year's document.
  • 16.Saito, T., Rehmsmeier, M. (2015). PLoS ONE 10(3): e0118432 — precision collapses mathematically at low prevalence. Pangakis, N., Wolken, S., Fasching, N. (2023). Automated Annotation with Generative AI Requires Validation. arXiv:2306.00176 — 9 of 27 tasks fell below 0.5 on precision or recall.
  • 17.Donelson, D. C. et al. (2024). The effect of securities litigation risk on firm value and disclosure. Contemporary Accounting Research 41(3): 1785–1818 — when litigation risk changes, the wording of risk-factor disclosure changes. ⚠️ We confirmed the abstract only, and it does not isolate AI disclosure.
  • 18.ONS, Artificial intelligence in UK businesses: 2023 to 2026 (20 July 2026, BICS Wave 159) — actual AI usage rates among UK businesses. The populations differ, so these are neither subtracted from nor placed alongside the disclosure shares in the body. The reading that the transformative impact has so far been relatively limited appears in that release too.

Related Pebblous pieces