Executive Summary
This article takes one thread out of the Microsoft Digital Defense Report published on 1 October 2026 — the parts about time — and traces it back to primary sources. The report sets two clocks side by side. One runs from the moment a vulnerability is discovered in the wild to the moment it becomes a weapon. The other runs from exposure to the moment an enterprise finishes fixing a critical vulnerability facing the internet. The first has dropped below a day. The second stretches across one to two months. Where that second figure is recorded, Microsoft's own guidance sits beside it. It is 72 hours.
The most widely quoted sentence in the report describes frontier models chaining 32 steps without human direction and taking over an emulated enterprise network. That evaluation was not run by Microsoft. It was run by the UK AI Security Institute, and the original write-up — which Microsoft lists in its own reference section — records the same result as two runs out of ten and three runs out of ten. The original adds a caveat: these results cannot say whether the same models would succeed against a well-defended target. Neither the number of attempts nor the caveat made it into the report.
Everything above is fact. Here is the interpretation. The firmest sentence in this report is not in the passage about attackers; it closes the section on patching. Microsoft writes that the question has moved from whether a patch exists to whether an organization can identify its exposed assets, and in the section on its own red team it writes that an inventory is not enough, that a graph of reachability is required, and that most environments cannot answer that question. What separates one organization from another, then, is not model capability but the state of the records a company keeps about its own systems. The report does not say attackers have won. Its phrasing is that the balance will eventually be restored while the near term favors the attacker, and it also notes that most observed campaigns are still human-directed.
Under 24 hours
Discovery to weaponization
Median measured from discovery in the wild, not from public disclosure
30–60 days
Enterprise fix for exposed critical CVEs
The same report recommends 72 hours; the US federal deadline is 7–14 days
2–3 runs in 10
Completion rate on the 32-step chain
The denominator the original evaluation records. Mythos Preview 3, GPT-5.5 2
1 victim
Documented target of autonomous ransomware
Sysdig's published analysis of JADEPUFFER records a single victim organization
165 Trillion Signals Is Not a Count of Attacks
The Digital Defense Report is Microsoft's annual publication. This year's edition appeared on 1 October 2026 and, unless noted otherwise, covers July 2025 through June 2026. Year-over-year comparisons use July 2024 through June 2025. That window is the first thing to check when reading the report, because the same events can trend up or down depending on which twelve months you cut.
Second comes the passage where the report states the scale of its own observation. Near the front sits a sentence about processing more than 165 trillion security signals per day, and bundled with it are 4.7 million malicious files blocked daily, 5.2 billion emails scanned daily, 31 million identity risk detections daily, 35,000 security engineers, and 15,000 partners. Those figures describe what Microsoft handles and analyzes in a day, not how often the world was attacked in a day. Read 165 trillion as a count of attacks and the whole report turns into a different document.
Third is the shape of that observation network. Nearly every number in the report comes from Microsoft products and the customer estate running them. Whatever happens outside that estate does not show up in this window, and Microsoft sells security products. The point of noting this is not to score a hit but to sort the sentences: which ones are observation, which are citation, and which are forecast. Two passages that this article leans on turn out to have sources outside the report, which made it possible to check them separately.
1.1Two Counting Methods, Two Different Answers
The report demonstrates once, inside its own pages, how much the answer depends on what you claim to be counting. Asked which ransomware family was most active in 2026, it publishes two different rankings two pages apart. Count public leak-site postings and one family leads; count Microsoft's own detection telemetry and a different one does.
| Counting method | First | Second |
|---|---|---|
| Public leak-site postings | Qilin 16.4% (+386% year over year) | Akira 8.3% |
| Defender detection telemetry | Akira 22% (+150% year over year) | Qilin 14% (+111%) |
Both tallies sit on neighboring pages of the same report. Microsoft itself notes that leak-site data can give a distorted picture, since ransomware operators choose which victims to post and can inflate what they post.
Publishing both figures is conscientious handling. The scene also previews the rest of this article. When the boundary of what is being counted goes blurry, the question "what is the biggest threat" has no stable answer. As we will see, the bottleneck the report identifies on the defensive side is the same kind of problem. When the boundary of what exists inside your own organization goes blurry, the question "what do we fix first" has no stable answer either.
One outside caution about reading the numbers is worth carrying along. An analysis published days after release argued that figures such as 88% or 1.3 billion, whose survey methods were never disclosed, should be treated as directional indicators rather than precise measurements. This article uses them for direction only.
A Model Finished 32 Steps, Two or Three Times in Ten
The fast clock is gathered on page 14. The sentence saying the median time from discovery in the wild to weaponization has fallen below 24 hours, Microsoft's own estimate that CVEs tracked during 2026 will reach 72,000, and the story of the 32-step attack chain all sit on that page. The three are different kinds of statement. The first is an observation, the second a forecast, the third a citation.
2.124 Hours Is Not Measured from the Disclosure Date
That median comes with a starting point attached. What Microsoft wrote is "discovered in the wild," not "published as a CVE." Miss that distinction and the figure collides immediately with other reports. Verizon's 2026 breach investigations report was covered as finding a five-day median from disclosure to first observed exploitation, and the two are not in conflict; they are different rulers. On page 14 Microsoft notes that pre-disclosure windows, in which exploitation begins more than 30 days before publication, have become routine, and cites a remote code execution flaw in a Cisco security product rated 10.0 that was used in ransomware roughly 30 days before it was disclosed.
CVE volume forecasts have to be sorted by source as well. The 72,000 figure is Microsoft's own estimate; for the same year, the Forum of Incident Response and Security Teams put the number at 59,427 and one industry analyst at 70,135. Three separate forecasts, three different methods. Since 48,185 CVEs were actually published in 2025, arithmetic along the lines of "2026 will be double" matches none of them.
2.2Where the Most-Quoted Sentence Actually Comes From
Page 14 carries the passage the press repeated most often, and it reads as follows.
Microsoft neither designed nor ran that evaluation. The UK AI Security Institute did, and Microsoft is the party citing the result in its own report. That citation is not hidden: both the institute's evaluation write-up and its paper appear in the report's reference list. So the originals could be opened, and opening them turned up a number that the report's body text does not carry.
The AI Security Institute's evaluation post, dated 30 April 2026, puts it this way. GPT-5.5 completed the full 32-step range in two attempts out of ten, making it the second model to do so. Mythos Preview, the first model to clear that range, managed it three times out of ten. The same post estimates that a human expert would need roughly 20 hours to finish the same chain, notes that the range contains none of the active defenders, defensive tooling or alerting penalties a real environment normally has, and then adds that these results cannot say whether GPT-5.5 would succeed against a well-defended target.
| Item | Microsoft report, p.14 | AI Security Institute original |
|---|---|---|
| Result | Full domain compromise | Full domain compromise |
| Conditions | Emulated enterprise network, no defenders | No active defenders, detection or alerting penalties |
| Attempts | Not stated | 10 |
| Completions | Not stated | Mythos Preview 3 · GPT-5.5 2 |
| Human expert time | Not stated | About 20 hours |
| Generalization caveat | Not stated | Cannot speak to well-defended targets |
The two documents do not contradict each other. Microsoft recorded the result; the original evaluation recorded the denominator alongside it. For human expert time, the same team's paper says 14 hours while the evaluation post says about 20 hours. The later figure is an estimate for the chain as a whole.
Restoring the denominator does not weaken the claim; it sharpens it. Two or three times in ten means the capability is not yet reliable even on a practice range with nobody defending it, and it also means successes are now appearing where there had been none. The trend the report is pointing at survives intact. What is nowhere in the original evaluation is the sentence "AI takes over enterprise networks without humans, every time."
2.3A Second Range from the Same Team Barely Moved
The paper Microsoft lists in its references is a multi-step attack evaluation the same team published in March 2026. It covers two ranges: a 32-step enterprise network attack called "The Last Ones" and a seven-step industrial control system attack called "Cooling Tower." Both were built by the security firms SpecterOps and Hack The Box, with evaluation design and interpretation resting with the authors. Neither range has active defenders or detection tooling, a limitation the paper states outright.
Seven model generations were measured, from GPT-4o in August 2024 to Opus 4.6 in February 2026. At a 10M-token budget, mean steps completed on the enterprise range rose from 1.7 to 9.8; at 100M tokens, the figure rose from 11.0 for Opus 4.5 to 15.6 for Opus 4.6. The best single run reached 22 of 32 steps, well past the previous generation's best of 13, and the paper converts that into roughly 6 hours of the 14 hours it estimates for a human expert. So no model in this paper finished the range, and completion arrived in the next generation at two or three times in ten. Put the two panels side by side and what appears is not a capability that materialized overnight but a curve climbing across several generations.
The paper also explains why the curve climbs. It names two mutually reinforcing trends: performance for a given model scales log-linearly with inference-time compute, with no observed plateau up to 100M tokens, and each new generation travels further on the same token budget. Taking only the first, increasing from 10M to 100M tokens yields gains of up to 59%. Then the paper attaches a sentence about who can do that. Scaling inference-time compute, it writes, requires no specific technical sophistication; unlike custom scaffolding, expert prompting or tailored tooling, actors of all skill levels can increase token budgets, which makes this a particularly accessible path to improved performance. The cost is printed too. One 100M-token attempt with Opus 4.6 runs about $80 at standard API pricing as of publication.
In the next breath the same paper draws a line. All the models it measured still fall far short of end-to-end completion, and performance drops sharply after attack phases that require specialist knowledge in reverse engineering, cryptography and malware development. Nor was the run quick. That 22-step attempt took approximately 10 hours of wall-clock time on the authors' infrastructure, and the steps it cleared amount to about 6 hours of human expert effort, so laying the two values against each other puts the machine on the slower side. Since the paper states that the 14-hour human figure comes from summing per-step effort estimates rather than from timed human trials, this comparison should be read for direction only. What changed on this range is not speed but the fact that nobody has to sit alongside.
The industrial control range tells a different story. Even the most recent models averaged 1.2 to 1.4 of seven steps, with a maximum of three. The paper describes this as models beginning to complete steps reliably for the first time. That is neither "every model failed" nor "every model finished." Microsoft's line about open-weight models trailing closed models by seven months in attack orchestration belongs in the same reading, though those seven months are Microsoft's own characterization rather than the paper's.
2.4One Documented Victim for the First Autonomous Ransomware
One case from outside the laboratory appears in the report too. In early July 2026 came the first documented automated ransomware extortion incident, which the Sysdig Threat Research Team named JADEPUFFER. What Microsoft wrote is not that it caught the activity, but that it observed AI-orchestrated intrusions sharing elements with it and that the volume is low. Then it adds one sentence: initial access consistently began at internet-facing services with known unpatched vulnerabilities.
Open Sysdig's original and that one sentence acquires a face. The way in was an instance of the open-source tool Langflow left exposed to the internet, still carrying a missing-authentication vulnerability that already had a number (CVE-2025-3248). Sysdig gives four grounds for calling the operation autonomous: payload comments in which the operator writes down its own reasoning, a recovery sequence that deleted, diagnosed, rebuilt and resubmitted in fifteen lines within 31 seconds of a failed login, actions that only make sense if free-text responses from the target were read, and more than 600 distinct payloads delivered in a short window.
In the same post Sysdig draws its own line. It had no visibility into JADEPUFFER's system prompt or agent configuration, so its data cannot determine how far a human was involved. The documented victim is a single operational database holding 1,342 configuration entries. Yet the day after the report came out, one outlet ran a plural headline declaring that autonomous ransomware had hacked real organizations. Where two primary sources wrote "volume is low" and "one victim organization," the facts had inflated that much within a day.
Remediation Is Slow Because the Tests Are Missing
The slow clock lives on pages 47 and 48, and here too the conditions come off first. The 30–60 day figure is not an average across all vulnerabilities but a median range for enterprise remediation of internet-exposed critical CVEs, and the original phrasing is closer to "can take" than to "takes." In the same paragraph the report supplies a comparison group: US federal agencies are required to fix anything listed as a known exploited vulnerability within 7 to 14 days.
And in the recommendation box just below sits Microsoft's own standard: fix new vulnerabilities affecting internet-facing or identity systems within 72 hours. The numerical axis of this article is therefore not 24 hours against 30–60 days but 72 hours against 30–60 days. The first pair measures different things with different statistics and cannot be turned into a ratio. The second pair is a recommendation and a measurement, written by the same company about the same population on those same two pages, and those can be laid against each other directly.
| Value | What it measures | Population | Kind |
|---|---|---|---|
| Under 24 hours | Discovery in the wild to weaponization | Observed vulnerabilities | Median (observed) |
| 72 hours | Deadline to fix a new vulnerability | Internet-facing and identity systems | Recommendation |
| 7–14 days | Deadline for listed vulnerabilities | US federal agencies | Regulation |
| 30–60 days | Actual enterprise remediation | Internet-exposed critical CVEs | Median range (observed) |
These four values were not measured with the same ruler. Dividing the first by the last to say remediation is "so many times slower" would set two different axes against each other. The pair that can be compared is the second and the fourth.
Different as these four are, all of them are denominated in time, so a single scale can hold them. In the figure below the vertical axis carries no meaning; only the horizontal axis is drawn to scale, from day 0 to day 60. Two things are worth looking at: the point where the recommended deadline ends does not meet the point where the measured band begins, and the weaponization bar covers a single day, which on this axis shrinks almost to a dot.
3.1Microsoft Does Not Blame the Delay on Negligence
An explanation for the slowness appears on page 12. After stating that remediation is inherently much slower than discovery, the report gives a reason: many systems lack robust unit and integration testing, which prevents them from shipping code changes quickly. The lag comes not from a shortage of security tooling but from the speed of the pipeline that fixes and ships software. That is a question of engineering maturity rather than of the security organization, and it is why the rest of this article turns toward data.
The same paragraph turns into a forecast. The world will likely live through a multi-year period in which the count of known-but-unfixed vulnerabilities rises sharply, and well-prepared attackers can stockpile the zero-days they find along the way. Elsewhere the report restates the same forecast as a multi-year window in which known-but-unfixed vulnerabilities spike and then recede — a signal not to read this as a one-year story. The forecast from section 2 lands in the same place. If CVEs tracked in 2026 do reach Microsoft's estimated 72,000, the inflow grows while the outflow pipeline holds its speed. The backlog accumulates by exactly that difference.
Tallies from other institutions point the same direction. Verizon's 2026 report was covered as finding that the median remediation time for known exploited vulnerabilities grew from 32 days to 43 year over year, while the share fully remediated fell from 38% to 26%. Those measure something different and cannot sit on the same line as Microsoft's figures, but they independently support the conclusion that remediation runs in tens of days and is getting worse.
An Inventory Is Not Enough, a Graph Is Essential
So far this article has followed what the report counted. From here on it looks at the sentences the report writes as its own conclusions. The patching section does not end on a number. Microsoft writes that the problem has moved.
The first line of the recommendation box is blunter. It opens with "know what to patch" and goes on to say an organization should maintain a complete, current inventory covering third-party systems, remote management tools and legacy devices so that exposed assets do not fall out of view. That recommendation is not an instruction to buy a patching tool. It is an instruction to reconcile the books.
4.1A Sentence That Surfaced While Describing Microsoft's Own Red Team
Elsewhere sits the sentence that goes furthest. It appears where Microsoft describes how its own red team operates, and the context helps. That section opens by saying measurement is what separates a mature program from a merely active one, then describes tracking remediation rates, repeat findings and time to fix, and weighing just as heavily whether defenders detect and contain the red team's activity. Next comes a remark that the real blast radius when a single developer is compromised runs far wider than any individual permission implies. Then this.
What matters is that this is not an abstract recommendation. It arrives while a red team is explaining its own measurement practice and reaches the point of saying what it takes to size up one developer's blast radius. Attached to it, without parentheses, is a clause: most environments cannot answer that question. For a sentence written by a company that sells security products, it leads with the customer's condition rather than the product on offer.
The difference between an inventory and a graph fits in one line. An inventory answers "what do we have." A graph answers "what can be reached from here." If the AI bill of materials that clears out shadow AI, which Pebblous covered earlier, was the first question, this report has raised the one that follows it. The two should not be mistaken for one another, though. The graph the red team describes maps reachability across an organization's network and permissions, which is a different scale from a software dependency graph. Carrying research findings about code-level reachability analysis cutting false positives by 70–95% over into claims at the organizational asset level would be its own kind of overreach.
4.2The Point of Entry in Section 2 Was Exactly This Problem
Two sentences from section 2 meet here. Microsoft wrote that initial access in AI-orchestrated intrusions consistently began at internet-facing services with known unpatched vulnerabilities, and the entry point in the incident Sysdig documented was a Langflow instance left open to the internet. Same fact, different resolution. Whether or not the attacking side used an autonomous agent, the door they came through was a door that was not on the books.
On how completely organizations keep those records, several survey figures exist outside Microsoft. In one 2026 industrial security survey, 21% of organizations said they had a complete inventory of their operational technology assets, while 88% in the same survey rated their own OT security program as mature. That 88% comes from a different survey than Microsoft's 88% in section 5, so the two are better kept apart. A separate tally found the share reporting full visibility into technology assets falling from 47% to 43% year over year. All of these reached us through secondary coverage of surveys, so this article does not build on them. Their direction does point where the report's clause points.
Agents Grow the Ledger Before They Grow the Attack Surface
The report's agent section runs across pages 22 and 23. Microsoft observed that 88% of enterprises are already experimenting with agents and that 82% of leaders plan broader deployment within the next 12 to 18 months. Attached to those is a projection of roughly 1.3 billion agents in operation by 2028, a figure cited as an "industry forecast" rather than a Microsoft observation. Worth noting that the three numbers differ in the standing of their sources.
Among the figures in that section, the clearest are the ones about prompts. Microsoft sorted prompts arriving at generative AI agents between February and May 2026 by intent: malicious use and policy violation 45%, sensitive data leakage 17%, system prompt leakage 14%, goal hijacking and downstream tool abuse 12%, system intelligence disclosure 11%, resource exhaustion and financial harm 1%. The body attaches a caveat that a single attempt can fall into several risk categories.
What stands out is the character of three of those six buckets. Sensitive data leakage, system prompt leakage and system intelligence disclosure are not attacks that break an agent to plant something inside it; all three are attempts to pull out what the agent can already reach. Attackers are after the agent's permission scope, and when that scope is not written down anywhere, the size of a loss cannot be estimated even after the fact.
The sentence that closes the section points the same way. Every agent needs its own identity rather than credentials borrowed from a human employee, which makes securing AI agents in 2026 a matter of executing on inventory, identity and governance for every agent, with defenders able to attribute, revoke and enforce least privilege at scale.
5.1All Three Risks the Report Names Are Record-Keeping Problems
For risks that arise when an organization adopts agents, the report lists three. Read them together and they are not three different attack techniques but three faces of the same ledger problem.
- Ungoverned proliferation — agent identities multiply faster than an organization can track them. Without central visibility, the report says, agent sprawl creates a risk larger than shadow IT.
- Excessive privilege — agents accumulate access over time and rarely have it revoked. Broad access persists in stale form long after the original purpose has changed.
- No accountable owner — when no person is responsible for a given agent's access, there is nobody to call once something goes wrong.
None of the three has anything to do with model capability. A smarter model does not resolve them, and neither does a safer one. One condition resolves them: when an agent is added, a line gets added to the ledger with it. Pebblous covered the same ground through Korean cases in 200 agents run the company and none of them have an employee ID and the AI agents enterprises cannot turn off, or even see.
Can Your Company Answer Within a Day?
Microsoft proposes changing the metric. The same sentence appears twice, once in the executive summary and once in the body: reporting should move from "number of patches deployed" to reduced exposure, increased detection coverage and shortened mitigation time. Pages 87 and 88 carry a section titled "from asset visibility to active defense," which lays out four stages climbing from asset visibility to continuous vulnerability identification, from there to threat detection and active defense, and from there to AI-accelerated insight.
Yet no published figure measured with those new metrics appears in this report, nor in the outside material surveyed while preparing this article. The practice of counting patches is to be abandoned, and the replacement ruler has no markings on it. That gap leaves readers with a clear job. Since nobody is measuring it for you, each organization has to time itself.
6.1Two Things to Measure First
Turn the report's sentences into questions and two remain. One is about reachability, the other about how fast you can answer. Both can be timed by a person with a stopwatch before anyone buys a tool.
- Reachability — can you draw what a single developer and the automation acting for them reach directly, what the pipelines they run can trigger, and what sits two hops out through everything those pipelines connect to? Those three are exactly the scope the red team says it measures.
- Time to answer — ask "how many of our assets are exposed to the internet today, and what version is each one running," and count the hours until an answer comes back. Given that the report's recommended deadline is 72 hours, the budget available for identification is far shorter than that.
Neither question can be answered by a security organization alone. Asset lists, owners and dependency relationships are usually scattered across infrastructure, data and product teams, and none of those records was created for security purposes. That is why this report's framing of the defensive task — reconciling records rather than buying more tools — lands heavily in practice. How exceptions accumulate until exposure remains is something Pebblous treated separately in nine in ten detection rule exclusions are still there three years later.
6.2What Cuts Against This Reading
Gathering the counter-evidence is the honest thing to do. First, the strongest case — completing all 32 steps — happened two or three times in ten, on a range with no defenders. Second, on the industrial control range measured by the same team, even the most recent models averaged 1.2 to 1.4 of seven steps. Third, frequently quoted figures such as 88% or 1.3 billion come with no disclosed survey method. Fourth, the report's observation window is Microsoft products and the customer estate, and this company sells security products. Fifth, as the plural headline about autonomous ransomware hacking real organizations showed within a day, facts inflate easily between primary and secondary sources on this topic.
Microsoft draws its own lines too. The balance between attacker and defender will eventually be restored, and the original wording is that attackers are reaching the gains first in the near term. It also writes that even now, with frontier systems demonstrating end-to-end autonomy, most observed campaigns remain human-directed. This article does not say attackers have won because the primary sources do not say it.
How long "near term" lasts, though, the report measures differently elsewhere. It writes that the spike in known-but-unfixed vulnerabilities will span several years, and the reason it gives is the absence of testing infrastructure. The claim that the attacker's advantage is temporary and the claim that the temporary stretch may run for years are standing side by side.
Why This Matters to Pebblous
Sections 1 through 6 report what was verified in the full report and in the originals that report cites. This section is the part those documents do not cover, and it is about why Pebblous spent so long with this report.
7.1The Same Task Under Two Names
Every Pebblous product stands at the stage before data gets used. DataClinic diagnoses the condition of a dataset, AI-Ready Data handles the cleanup that precedes training, and data lineage records where each thing came from. The bottleneck this report names on the defensive side is fundamentally the same task. How accurately and how currently has an organization recorded its own assets and their dependencies? Security calls this exposed-asset identification; the data side calls it lineage. The scales differ, as section 4 already noted. The graph the red team describes maps network and permission reachability, and that is not the same object as a code dependency graph.
7.2How Data Quality Converts into Exposure Time
The causal chain this report builds is simple. Unrecorded assets mean exposure goes unidentified; unidentified exposure means nothing can be prioritized; without priorities remediation runs late; late remediation leaves a window open for attackers. Contrary to the assumption that data quality problems show up only as model performance, in security they show up as exposure time. The accuracy and currency of asset records are measurable data quality indicators in their own right, and this report gives a primary-source account of how those indicators convert into security outcomes. Now that agents are running at 88% of enterprises, the set of things that need recording has grown as well.
7.3An Answer for Partner Conversations
Organizations building a security budget usually count detection and blocking tools first. This report says there is a record-keeping layer ahead of them in line. Asked what an investment in data lineage and asset records returns, one can now answer with a primary source: not model reliability, but reduced exposure time. The distance between a 72-hour recommendation and a 30–60 day measurement is the size of what that investment addresses. And the gap in section 6 — new metrics prescribed with no published figures measured by them — also means the seat is open for whoever builds that ruler first.
7.4A Layer Entered the Conversation
AI security discourse has so far been written mostly in the language of models. Which model is more dangerous, which guardrail to apply. By writing that the difference comes from an organization's records rather than from its models, this report has admitted the data record layer into that discourse as a first-class item. What Pebblous can do is put that arrival into Korean first, and this article is one step of it.
The verbatim quotations in the body were checked directly against the full PDF of the Microsoft Digital Defense Report 2026, the UK AI Security Institute's evaluation post and paper, and Sysdig's original analysis. Page numbers follow the print edition, and the multi-column layout allows for an error of about one page either way. Survey and tally figures that reached us through secondary coverage are marked as such in the sentences where they appear and carry no weight in the argument. Sections 1 through 6 are what was verified; section 7 is the part those documents do not cover, so please read them separately. Thank you for reading this far.
References
Sources are grouped because they differ in standing. The first group is the report this article examines, and every verbatim quotation and page number in the body was checked against it. The second group holds the primary evaluation documents that report lists in its own references; the denominator and the caveat in section 2 came from there. The third group reached us through press coverage, and each time one of its figures is used in the body, the attribution is written into the sentence.
Subject report (checked against the primary text)
- 1.Microsoft. Microsoft Digital Defense Report 2026. 1 October 2026. microsoft.com — pages cited: p.4–5 (observation capacity block), p.7–8 (metric shift in the executive summary), p.12 (structural reason remediation is slow, multi-year forecast), p.14 (weaponization median, 72K CVE forecast, 32-step chain, open-weight seven-month gap, JADEPUFFER), p.22–23 (agent adoption rates, prompt intent distribution, three risks), p.47–48 (30–60 days, KEV 7–14 days, 72-hour recommendation, pre-disclosure window, conclusion of the patching section), p.75–76 (two ransomware tallies), p.87–88 (from asset visibility to active defense).
- 2.Microsoft. Microsoft Digital Defense Report 2026 — Executive Summary. 1 October 2026 — a separate summary PDF. Some figures use different statistics from the main text, and this article follows the main text.
- 3.Microsoft Security Blog. Insights from the 2026 Microsoft Digital Defense Report. 1 October 2026. microsoft.com
Primary evaluation documents the report cites
- 4.AI Security Institute (UK). Our evaluation of OpenAI's GPT-5.5 cyber capabilities. 30 April 2026. aisi.gov.uk — reference 4 in the Microsoft report. The completion counts of two in ten and three in ten, the estimate of roughly 20 hours for a human expert, and the caveat about well-defended targets all come from here.
- 5.AI Security Institute (UK). Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios. arXiv:2603.11214v3, March 2026. arxiv.org — reference 6 in the Microsoft report. The builders of the two ranges (SpecterOps, Hack The Box), the absence of defenders and detection, mean completed steps of 1.7→9.8 and 11.0→15.6, the best single run of 22 of 32 against a previous best of 13, gains of up to 59% when the inference budget rises from 10M to 100M tokens with no observed plateau, the statement that scaling inference-time compute requires no specific technical sophistication, the roughly $80 cost of a single 100M-token attempt, the approximately 10 hours of wall-clock time for the best run, the 14-hour human expert estimate (summed per-step estimates rather than a timed baseline), the sharp drop after phases requiring specialist knowledge, and the industrial control range averaging 1.2–1.4 steps (maximum 3) all come from here. ⚠️ This paper does not contain the two-or-three-in-ten completion counts; that figure comes from reference 4.
- 6.Sysdig Threat Research Team. JADEPUFFER: Agentic ransomware for automated database extortion. 1 July 2026. sysdig.com — the naming and documentation, the initial access vector, the four grounds for autonomy, the self-imposed caveat about having no visibility into the system prompt, and the single documented victim organization all come from here.
Press and tallies (used with attribution stated)
- 7.Help Net Security. AI cybersecurity threats: Microsoft report. 2 October 2026. helpnetsecurity.com / BleepingComputer. Microsoft says threat actors are ahead in the early AI race. bleepingcomputer.com
- 8.Tech Times. Microsoft 2026 Security Report: Autonomous Ransomware Has Hacked Real Organizations. 2 October 2026 — cited in sections 2 and 6 as an instance of secondary coverage overstating the case. The two primary sources write "volume is low" and "one victim organization."
- 9.Verizon 2026 Data Breach Investigations Report (43-day median remediation, 26% fully remediated, five days from disclosure to exploitation), FIRST 2026 Vulnerability Forecast (59,427), Jerry Gamblin's annual forecast (70,135), actual CVEs published in 2025 (48,185), Honeywell 2026 OT Cybersecurity Benchmark Report (21% complete inventory / 88% self-rated maturity), Flexera 2025 State of ITAM Report (43% full visibility) — ⚠️ for all of these the original report PDF could not be opened directly, and the figures came via coverage and aggregation. They are used in the body only as context for comparison.
- 10.Studies reporting reduced false positives from code-level reachability analysis (empirical work on SBOM-based vulnerability management and similar) — mentioned in section 4 only to mark the difference in scale from claims at the organizational asset level.
Related Pebblous articles
- 11.The AI Bill of Materials (AI BOM) That Clears Out Shadow AI (section 4) — the story of building the inventory. 200 Agents Run the Company. None of Them Have an Employee ID. (section 5) · The AI Agents Enterprises Can't Turn Off, or Even See (section 5) · Nine in Ten Detection Rule Exclusions Are Still There Three Years Later (section 6) · Hundreds of AI Agents Breached 440 PaperCut Servers at 395 Organizations — an earlier piece on where claims of autonomous attack reach their limit.