Executive Summary

This report rereads Mods, which Anthropic shipped for Claude Code on 1 October 2026, against the official documentation and the public source of the built-in guard. A mod is a JavaScript or TypeScript function that ships inside a plugin and runs inside the Claude Code process. It can catch a tool call, rewrite a prompt, redraw the screen and add commands of its own. It is on by default. Most of the launch coverage wrote that much and stopped. Powerful, and no sandbox. But the absence of a sandbox is not a Claude Code fact. VS Code says in its own documentation that the extension host has the same permissions as VS Code itself. Stopping there says nothing yet.

The documentation records something more specific. First, the deny rules you wrote hold in some places and not in others. Claude Code evaluates permission rules deny-first, but a mod that handles the final verdict on a tool call answers after that verdict, and its answer can replace it. Nothing in the engine reverses that answer. A built-in guard seated ahead of the mods does, and the guard loads only when the machine has managed settings or the user is signed in on a Team or Enterprise plan. Second, the side that blocks defaults to letting the call through. If a hook throws or runs past its time limit, Claude Code skips it and the command it was holding runs. Third, Mods is not an isolated box but a chokepoint you can intercept. There is exactly one road to the outside, every call along that road raises another event, and you can print what a mod reaches for before you install it. Two doors lead out of that chokepoint.

That is what the documentation and the public source state; what follows is this report's reading of them. The three are less three separate defects than one shape: the place where a rule is written and the scope a rule applies to are not the same. Block a file from being read and the plugin code still reads it. The rule is not wrong; the rule covers a different subject. This shape was given a name fifty years ago.

45 / 80

Events and API methods you can hook

Per the v2.1.289 docs. 45 named events, 80 methods, and every one of those method calls is an event too

15

Events the built-in guard hooks

Counted directly in the guard's source. Files, network, processes and tool execution pass straight through

10s / 1s

Time limits for a hook and its failure handler

Per the v2.1.289 docs. Ten seconds for one hook, one second for the handler that catches its failure

0 / 3

Official examples with a failure handler

Counted by reading all three published example mods in full. Not even the one that holds dangerous commands has one

1

Functions that moved inside the process

Mods is not a separate new product. It is one part of plugins. Put a hooks module inside a plugin folder and Claude Code calls that module's functions from inside its own process. It runs from v2.1.287 in the terminal and from v2.1.286 in the desktop app, and it is on by default. The hooks you used to write in settings files are still alive. The documentation nails that down: "Nothing about them is deprecated."

What a function receives is an event. There are 45 named events, and the API a module uses to reach outside has 21 namespaces and 80 methods. Both numbers come from recounting the tables in the v2.1.289 documentation, and the 80 excludes five import helpers for stored state and five return fields on the usage query. One more condition attaches here. Each of those method calls is itself an event. The public type declaration file is on an older build (v2.1.277), where the same count comes to 20 namespaces and 78 methods, so mixing the two baselines in one sentence gets the number wrong.

A hook has three moves available to it on an event. Observe and pass it on, rewrite it and pass it on, or answer for itself without calling what comes next. The third is the axis of this whole report. A hook that receives a tool call and answers "deny" without calling next stops the command; answer "approve" and the call goes through before the permission prompt appears.

1.1The hooks module has no files and no network

The module's own reach is narrow. The documentation says it outright.

"The hooks module itself has no Node.js APIs, no timer globals such as setTimeout, and no network or file access of its own." Source: mods API documentation.

Reading a file, starting a program or using the network all have to go through that API. This design is what makes section 2 possible. One road means you can put a checkpoint on it. What the module does through that API, though, happens with the permissions of the user who ran it. What is narrow is the module's vocabulary, not the module's authority.

1.2How wide the same frame stretches

Several of Claude Code's own features have already moved onto this frame. The /plugin list shows six, and four of those have public source. Open the published ones and you can see how far the same frame stretches. diff, which shows you what changed, touches 10 of the 21 namespaces and calls 23 API methods. The built-in guard that serves as the security default calls two: read settings, write a log. Both are the same kind of code in the same seat, and their reach differs by that much.

1.3Hooks run on six surfaces of seven, and drawing shows up on two

"Supported in both the terminal and the desktop app" carries half the fact. The surfaces table in the official documentation shows that the range where hooks run and the range where a mod's drawing reaches the screen are not the same. Hooks run on the unattended surfaces too. That fact is what section 6 builds on.

Where Claude Code runs Do hooks run? Can you see what a mod draws?
claude in a terminal (including an editor's built-in terminal and the JetBrains plugin)YesYes
Desktop app, Code tab (excluding WSL sessions)YesYes (except terminal-only elements)
WSL sessions in the desktop appNo (plugins do not work at all)No
The VS Code extension's chat panelYesNo
claude -p and the Agent SDKYesNo
Remote control from claude.ai or mobileYes (in the session on your machine)In the terminal on your machine
Cloud sessionsYes (for plugins whose settings carry over)No

Transcribed from the surfaces table in the Mods overview documentation. Hooks run on six of the seven surfaces, and drawing appears on two. Which is to say hooks also run where nobody is looking at a screen.

2

There is no isolation, but there is a chokepoint

The place the explainer pieces all stopped is "no sandbox." It is not a false statement, but it describes only part of the design. What Anthropic put into Mods is a chokepoint. As section 1 showed, the hooks module has no files and no network, and there is exactly one road to the outside. Every call made along that road raises another event.

"Every one of these calls is itself an event, named for its namespace and method without the $. … A mod earlier in the chain can observe, rewrite, or refuse your call, which is how an organization restricts what mods reach." Source: mods API documentation.

On top of that, a mod ahead of you can alter the API object a later mod receives before handing it over. And because the chokepoint only works if the code can be read statically, a mod that reaches the API in an unreadable way is refused at load. That is an unusual device in an extension ecosystem.

2.1No isolation is the default for this category

So where does the sentence "there is no sandbox" belong? Open the official documentation of the products you would compare it against and the answer appears. VS Code writes this about itself: "The extension host has the same permissions as VS Code itself." A developer tool's extensions running with host permissions is the default for this category, which makes that sentence a statement about the category rather than a critique of the product. That is why this report has to move to the next box.

2.2Two doors lead out of the chokepoint

The first door is child processes. The official documentation states that even with sandboxing on, processes a mod starts run outside it. Network policy works the same way. Switch web requests off for the organisation and the requests a mod sends through the API are refused, but a program the mod starts reaches the network with the user's own access.

The second door is the boundary of what permission rules cover. The plugin security documentation settles it in one sentence: Claude Code's permission rules and sandboxing cover the tool calls Claude makes, not the code a plugin runs itself. The organisation admin documentation states the same fact with an example. Deny Read(.env) and a mod can still read that file with the file-read method, or start a program that does. And the very rule the permissions documentation recommends as the proper way to protect secrets is that deny rule. The two documents look like they undercut each other, but there is no contradiction. Two actors touch the same file by two paths, and the rule hangs on only one of them.

What rules and sandboxing cover = the tool calls Claude makes Read(./.env) — blocked by a deny rule Bash(…) — blocked by a deny rule Outside that scope = what mod code does itself $.fs.read $.process.run $.http.fetch with the user's permissions blocked read the same ./.env file

Pebblous drew the two sentences from the plugin security documentation and the organisation admin documentation as a diagram. The left path and the right path reach the same file, and the deny rule hangs on the left one only. The point of the figure is that the rule is not wrong; the rule covers a different subject.

2.3This design has a name, and its lineage already knows where it stops

Anthropic's documentation never gives this structure a name. The one used here is borrowed from elsewhere. Handing every authority over as one object and letting calls on that object be intercepted again is not new. The object-capability literature calls it attenuation by interposition. By the definition Mark S. Miller set out in his 2006 doctoral thesis, pass a reference along unchanged and the receiver holds undiminished authority; to attenuate it you interpose an object in the middle and control how it forwards. The move where one mod alters the API object it hands to the next overlaps with that definition.

Give it a name and the limit its lineage already knows comes along too. The JavaScript implementations of this model (SES and Endo, and LavaMoat for dependency isolation) all stop at a common place. A compartment is a language-level boundary, not an operating-system boundary. It constrains the reference graph and not process authority. An independent audit report even documents a case where editing an environment variable skipped the isolation load entirely. So the first door above, child processes running outside, is not a hole that was punched through. That authority was never inside the mediated scope to begin with. This contrast is not a story about other implementations managing it while only Claude Code does not. They all stop in the same place.

2.4What the chokepoint changes is what gets drawn on screen

One summary that spread widely after launch has to be corrected here before we go on: that Mods automatically strips sensitive data. It is wrong twice over. No such filter is a built-in Claude Code feature, and the example mod cited as evidence does not delete anything either. That example "visually masks" email addresses until you hover over them. It does not remove them, and hovering brings them back. What a mod seated at the chokepoint changes is what gets drawn on screen, not the value Claude has already read. Miss that distinction and you end up treating a screen cover as a protection.

3

Where deny rules hold, a condition is attached

What this section measures is not how people write rules but how far the engine applies them. What happened when people were asked to write permission rules up front was covered separately in the experiment on pre-written permission rules; here we only look at how far a well-written rule carries. The conclusion first. On a machine that has neither managed settings nor a session signed in on a company plan, a mod you installed can approve a tool call that your own deny rule refused. Strip the conditional clause and it becomes an exaggeration. Keep it and it is a design fact.

3.1The order of the rules, and the seat that answers after them

Claude Code's permission rules have a stated precedence. The permissions documentation puts it this way: "Rules are evaluated in order: deny, then ask, then allow. The first match in that order determines the outcome." The same section nails down the corollary: "An allow rule can't carve an exception out of a deny rule." Deny is the strongest.

The paragraph after that records the new seat. A mod that handles the final verdict event on a tool call answers after the rules and the settings hooks have all decided, and "its answer can replace theirs." The documentation splits how far that replacement goes into four bullets, and the last one is the centre of this report.

"Deny rules: on a machine with managed settings, or when you're signed in with a Team or Enterprise plan, deny rules hold over the mod by default, and your organization can change that. Anywhere else, the mod can approve a call that a deny rule refuses." Source: Permissions documentation, "Extend permissions with hooks."

3.2What keeps deny first is another mod

Why the condition is there becomes clear in the organisation admin documentation. What restores deny-first is not the engine. Another mod does it. Claude Code loads a built-in guard called sec-default@builtin ahead of every mod a user installs, and the user cannot switch it off. It has two load conditions, though: the machine has managed settings, or the user is signed in on a Team or Enterprise plan. Users who authenticate with an API key or through a cloud provider get the guard only on a machine with managed settings. The same document adds: "Where the guard doesn't load, neither option applies."

The order falls out like this. The mod seated first in the chain is outermost. It sees the event before anyone else and receives the result after everyone else, and it decides whether the ones behind it run at all. The hooks written in settings files also have a seat in this line, and which seat decides how much force they carry. A pre-tool-use hook declared in managed settings runs ahead of the first mod, and a block there is final. The same hook in your own settings file runs at the very end of the chain, so if a mod ahead of it answers for itself without calling next, it never runs at all. The asymmetry from section 3.1 shows up once more, in the same shape, in settings hooks.

There is also a box that comes after the user. An organisation can list its own mods in managed settings and seat them either ahead of a user's mods or behind them. The ones seated behind run after every mod the user installed, and if a mod in an earlier box answers for itself without calling next, the later box never sees that call. The fourth box in the figure below is that seat. What happens when a policy mod sitting there fails is covered separately in section 4.4.

The line one tool call passes through Pre-tool-use hook in managed settings a block here is final Built-in guard seated only with managed settings or a company plan the mods you installed can answer after the rules decide the mod seated behind yours the organisation's policy mod Claude Code itself the remaining settings hooks run here In any box, if a hook throws or runs past its time limit Claude Code skips that box and moves on (section 4) When the dashed box is empty, the answer from the third box replaces the answer from the rules

Pebblous drew the chain order section of the events documentation and the guard's load conditions from the organisation admin documentation as one line. The argument of this report is that only the guard box is dashed. The machine where that box is empty is the ordinary single-user install on a personal plan.

So why not install the guard yourself? Its source is public. Its own README closes that path: "It is a plugin folder like any other, but its one move that matters, next.to, is refused outside a managed tier, so loading it with --plugin-dir seats a plugin that can only pass." You can read it. You cannot put it to work.

3.3What the guard watches and what it waves through

Then does deploying managed settings make you safe? That depends on what the guard actually watches. Open the registration file in the public source, count every on( call, and you get 15 events hooked; what the guard itself calls through the API is two things, read settings and write a log. Those 15 go to keeping what the organisation configured (managed CLAUDE.md and rules, settings, the tool list and descriptions, the final verdict, and whether other mods load) out of the user tier's reach. The README's closing line covers the rest, and that list includes file access, network requests, process execution and tool calls themselves. All pass through. That is what the README means where it narrows its own scope: "seated outermost, this plugin keeps exactly those out of the user tier's reach and adds no policy of its own. Everything else passes through untouched." The guard protects the controls an organisation already had. It does not protect the user's files.

3.4The two switches an organisation holds are strict in opposite directions

The guard has two options that only managed settings can switch on: one locks loading to mods the organisation deployed, the other releases deny rules so installed mods may override them. The second is locked by default. Read the published policy code and the two are written strict in opposite directions. The locking one locks on any value that is not empty, and the releasing one releases only on the value exactly true. Write the string 'true' or the number 1 and it stays locked. It is built so that a typo never moves in the direction that breaks protection, and a test name states that principle in words: "a value mistyped still locks, as a managed lock reads."

On a machine where the guard is alive and a deny rule actually holds, the user sees a single line. The wording in the README and the return string in the source do not differ by one character.

<plugin> tried to lift a deny rule in your settings from a <tool> call (<rule>); the deny rule holds over the plugins you install (allowModsToOverrideDenyRules)

Confirmed by comparing the notice string in the guard's README against the return string in its source. The machine where this line never appears is the machine where the dotted box in section 3.2 is empty.

3.5Measured against a fifty-year-old yardstick, one box stays empty

The comparison below is not in the official documentation. The yardstick is one this report held up. The first document to write down what a mechanism that mediates access must have is James P. Anderson's 1972 report, and the three requirements in it are still quoted today. It must be tamper proof, it must always be invoked, and it must be small enough to analyse and test. Hold that yardstick against the built-in guard and two boxes fill while one stays empty.

The three requirements, 1972 Built-in guard Basis
Tamper proofMetThe user cannot switch it off, and seated outermost it decides whether the ones behind it run
Small enough to analyse and testMetThe registration file is 3,643 bytes, each decision has its own file and unit tests, and the source is public
Always invokedNot metIt loads only with managed settings or a session signed in on a company plan

The verbatim of the three requirements comes from the US Department of Defense evaluation criteria, which quotes the Anderson report directly; the two right-hand columns were confirmed in the official documentation and the public source. Neither the Anderson report nor the guard's README points at the other.

That the empty box happens to be "always invoked" is hard to look past. Reading it as a design failure would be wrong, though. As its README narrows its own scope, that artefact was not built to be a mediator. One sentence is as far as this goes. Measure the only thing seated there against a mediator's yardstick and one box stays empty.

4

When the thing that blocks fails, the call goes through

The most tempting line in the Mods pitch is that you can build your own safety device. Catch a dangerous command in a hook, ask a person, then decide whether to proceed. That line has a back side, and the back side is in the same documentation. A safety device you build yourself lets the call through when it fails instead of blocking it.

4.1Fail, and it gets skipped

The events documentation splits failure handling into two lines. When a hook with no failure handler throws, exceeds its time limit, or returns a malformed result, this happens.

"It failed before calling next: Claude Code skips it, and the next handler runs in its place." / "It failed after next resolved: that result stands, and nothing runs a second time." Source: mods events documentation.

What is left behind is written down too. One line. A log line naming the mod, the event and the reason goes somewhere, and the command that was being held runs. The hold example in the same documentation states the outcome directly in a comment: "Claude Code skips a hook that times out, so the held command would run." Which means the command goes before the person has answered.

4.2Closing it takes a handler you attach yourself, and the price is one second

How to make it close is on the same page. Attach a failure handler to the registered hook so it answers in that hook's place. The example the documentation ships looks like this.

on('tool.call', { tool: 'Bash' }, guard).catch(async ($, e, next) => { // next.error.kind is 'throw' or 'timeout' — it tells you how guard failed return { deny: 'The command guard failed, so this command was not run: ' + next.error.kind } })

The official example from the events documentation, reproduced as published. While guard works, this handler never runs once. The documentation's explanation closes like this: "Without the handler, Claude Code would skip guard and run the command."

The price is time. A hook gets ten seconds of its own, and the handler that catches its failure and answers in its place gets one. Try to gather the material you need for that judgement inside the one second and you walk into the same trap again.

The other three rows of the table sit on the same scale. A hook that rewrites a prompt gets 50 milliseconds, and the hooks that run as a session ends get 1.5 seconds for all of them together. A process a mod starts, by contrast, gets 30 seconds by default and up to ten minutes if you ask for it. As the first row of the table notes, that wait does not count against the hook's own ten seconds, so a hook's judgement in its own seat is cut off at ten seconds and one second while the work a hook hands to an outside program runs for up to ten minutes.

Limit Value
A hook's own execution time on one event (waits on next and on API calls are excluded)10 seconds
Execution time for a prompt-editing hook50 milliseconds
Execution time for a failure handler1 second
Budget for all session-end hooks combined1.5 seconds
Timeout for a process a mod starts30 seconds default, 10 minutes max

The five rows this section needs, taken from the limits table in the v2.1.289 documentation. The gap between ten seconds and one second is the point of this section.

4.3None of the three published official examples attaches a failure handler

So did Anthropic's own published examples avoid this trap? We opened the source of all three mods in the official example repository in full and counted whether their registration chains attach a failure handler. One holds dangerous commands, one shows token usage, one replays a session. None of the three attaches one.

Official example mod What it does Events hooked Failure handler
blast-radiusHolds dangerous shell commands and asks a person2 kindsNone
token-weatherShows token usage above the prompt3 kindsNone
replay-theaterReplays a session7 kindsNone

The full source in the official example repository was read on 6 October 2026 and every registration call counted. With a sample of three we do not infer intent. What can be written down is the observation and no further.

The example that holds dangerous commands is the one that catches the eye. Its code does catch exceptions from its own logic internally and turn them into a denial. But that is not a failure handler registered with the engine. The failures the engine defines, throwing and timing out, are not covered by that internal handling. And the example's own README lowers its rank for you: "This is a safety net, not a permission system. Use permission rules to be sure." The README also lists several command shapes it does not catch.

4.4A policy mod the organisation installs stands in the same position

This default does not apply only to devices individuals build. The organisation admin documentation, while explaining how to use your own policy mods, attaches the same warning. If the hook that vets whether other mods load throws or exceeds its time limit, Claude Code skips it, so "the check fails open and the mod it was checking loads." And if the worker thread that runs installed mods dies three times, every non-built-in mod goes down. The organisation's policy mod goes down with them.

Only the built-in guard runs the other way, and even that is conditional. Read the public source and there are exactly two places inside the guard with a failure handler attached, and the two behave differently. The final-verdict side falls to an unchecked denial when it fails, and a test name records the outcome in words: "when the rules cannot be evaluated the call is refused." The side that vets whether other mods load looks at when the failure happened. If it failed after already letting something pass, that pass stands; if it failed before deciding, it refuses. So rather than "the guard always fails closed," the accurate sentence is that the guard closes when it fails before granting anything and keeps the pass when it fails after granting one. The case where it cannot read the policy at all folds toward blocking.

4.5A sentence from 1975 wrote down this failure shape, name and all

The quotation here comes not from the official documentation but from a 1975 paper. Setting it beside the failure above is ours. The second of the eight design principles J. H. Saltzer and M. D. Schroeder set out in 1975 is fail-safe defaults. The closing sentence of that item describes the failure shape above almost exactly.

"a design or implementation mistake in a mechanism that explicitly excludes access tends to fail by allowing access, a failure which may go unnoticed in normal use." Source: Saltzer & Schroeder (1975), the fail-safe defaults item.

A hook that holds dangerous commands is by definition a mechanism that explicitly excludes access. Its failure mode is allowing, and since all that remains is one log line, the rest of the sentence about going unnoticed holds too. Saltzer and Schroeder never saw Mods, and the Mods documentation never cites them. The classic is a yardstick, not an answer key. Treat it as something already solved in 1975 and you lose sight of where this default could actually be changed today. That place sits inside a two-line failure handler.

5

The two lines you can print before you install

Everything so far has been about what happens after a mod loads. What you can do before it loads is in the documentation too. Run the validation command against the plugin folder you are about to install and what that code receives and what it calls come back as two lines. One line for the events it takes, one for the API methods it calls.

❯ ./register.js hooks: session.start, tool.call, ui.render{component=Pane} ❯ ./register.js calls: $.fs.read, $.http.fetch, $.store.set, $.ui.open

The sample output published in the organisation admin documentation, reproduced as printed. A module that reads or writes environment variables gets another line naming those variables, and a module that uses stored state gets a line naming its keys. Two lines is the baseline and the count grows with what the mod uses.

5.1What the two lines tell you and what they do not

This device carries one rare property. If a mod reaches the API in a way that cannot be read statically, Claude Code refuses to load it. It does not merely print a list; it keeps out the code that cannot be listed. The documentation also puts the lines worth looking at into a table: reading and writing files, running programs, network requests, reading environment variables and settings, submitting prompts.

What it does not tell you is just as clear. The line says what a mod calls, not what it does. A single line for the process-execution method in particular amounts to "anything." As section 2 showed, a process started that way runs outside the chokepoint, and neither permission rules nor network policy reach it. Review also carries a timestamp. With marketplace auto-update on, the file you read can change on disk later, and that problem was covered separately in the piece on plugins that keep their name and change their contents.

5.2Which boxes the other tools in this category filled

Put this capability declaration alongside the extension systems of other developer tools and the position of Mods becomes visible. Every cell in the table below was confirmed in each product's official documentation or marketplace policy documentation, and it does not rank them. "Documented as absent" and "the documentation does not address the question" are different cells, so they are written separately.

Tool Isolation Review Signing and publisher checks Capability declaration Can it override a permission verdict?
Claude Code Mods None (documented) Not in the documentation Not in the documentation Yes — prints two lines, and refuses to load what it cannot read Conditionally, yes
VS Code extensions None (documented) Yes — automated checks at publish time Yes — every extension signed None Not in the documentation
JetBrains plugins Not separately documented Yes — a person reviews every new submission and update Yes — signed by both author and marketplace None Not in the documentation
Chrome extensions Not confirmed in this research Yes — every new submission and update Yes Yes — manifest permission declarations Not in the documentation
Gemini CLI extensions Yes — an OS-level sandbox option None (documented) None A manifest exists, but it is not a capability declaration Not in the documentation

Filled from each product's primary documentation. Rows for Cursor and OpenAI Codex were left out because they could only be confirmed through secondary sources. The isolation cell for Chrome was also left blank for want of a primary confirmation.

Three sentences come out of the table. First, "no isolation" is not a feature peculiar to Mods. The VS Code documentation from section 2.1 says the same thing about its own product. Second, two rows fill the capability declaration box — Mods and Chrome — and among the four developer tools only Mods does. Review and signing go unfilled in two rows as well: Mods, whose documentation addresses neither, and Gemini CLI, whose documentation says it has neither. Third, only one row carries a mark in the last column, and not because the others cannot be overridden. Claude Code is the only one whose documentation answers that question explicitly, and its answer is a conditional yes.

5.3Even the tightest review stopped at the same place

So is the answer to fill the two empty boxes? The product that fills them most tightly in that table is JetBrains. A person reviews every new submission and update, and both author and marketplace sign. Under that regime, 15 malicious AI plugins were found and removed in June 2026. What they took was the AI provider API keys developers had configured, and JetBrains wrote publicly about why its own checks missed them.

"Historically, our Plugin Verifier tool was architected as a compatibility and API-usage checker rather than a dedicated data-flow or anti-malware scanner. Because the core APIs used by the plugins appeared normal in isolation, individual hardcoded endpoints and custom TLS configurations were not flagged during initial ingestion." Source: JetBrains Platform Blog, June 2026.

This case does one job in this report. Review and signing each guarantee a different thing, and the scope of that guarantee is set by what the tool was built to look at. What the review that passed guaranteed was compatibility and API usage, not what the code does. So reading it as "review is useless" gets it wrong. JetBrains writes in the same post that it is rolling out a new layer of checks. How long the plugins sat on the marketplace and how many installs they had are not in that official post, so this report does not state them either.

There is evidence pointing the same way from the research side. The early study measuring the effect of permission declarations in browser extensions (Felt et al., 2011) started from the finding that nearly every extension asks for at least one dangerous permission at install time, and argued that when warnings appear constantly and rarely precede a bad outcome, users get less information out of an install-time warning. A capability list informs you only while the items on it are rare. Run that logic over the two lines from section 5.1 and a question stands up. If the process-execution method ends up in most mods worth installing, that line stops being a warning and becomes background. Mods is five days old, so there is no material for measuring the actual distribution. This paragraph is therefore not a prediction. It records the road the same device travelled in another ecosystem.

5.4Five settings an organisation can use, and what each does not cover

There are five handles an organisation can hold right now. All five are in the official documentation, and this report invents no new control. Each is only useful alongside the box it does not cover.

The five do not sit on one layer. Locking loading to the organisation's own deployments and releasing deny rules for override are the two guard options from section 3.4, so they mean something only on a machine where the guard loads. The other three work without the guard, but in exchange they cannot block one chosen mod. They stop everything installed at once, or they decide where plugins may come from, and that is as far as they go.

Setting or option What it does What it does not cover
allowManagedModsOnlyLoads only the mods the organisation deployed, plus built-in modsBuilt-in mods. Settings hooks, the status line and /goal keep running
allowModsToOverrideDenyRulesReleases deny rules so installed mods may override them (locked by default)Machines where the guard does not load. There, neither of the two options applies
disableAllHooksStops the mods a user installed in every sessionBuilt-in mods. Settings hooks and the status line stop with them, which is a steep price
--safe-modeRuns one session without any installed modsBuilt-in mods. Your other user settings stop as well
Marketplace restriction settingsDecide which sources a plugin may be installed fromBehaviour after loading. A mod is a plugin, so only the installable range is set

Taken from the Mods overview documentation and the organisation admin documentation. Five lines further down, the same document attaches a note that covers all of them: "None of these controls sandboxes a mod. A mod you allow runs as the user, with the user's access to files, processes, and the network."

6

The layer that writes the record runs on the same permissions

What we have looked at so far is tool calls and the verdicts on them. The event roster has other lines too, covering what Claude reads and what is left behind for people. This section is not about a missing record. The situation where someone wants to inspect and there is no ledger to open was covered in a separate piece on that subject; here we look at the other side. The record does get written. The layer that writes it is replaceable.

6.1Four places that can be changed

We list only what appears in both the official documentation and the public type declarations. There are four. First, an event builds the attribution text attached to commits and pull requests. The sentence that marks a piece of work as Claude's is set by that hook's return value. Second, the system prompt is assembled from named sections, a hook attaches to each one, and returning an empty value drops that section. Third, a hook attaches to messages arriving from another session or agent, and the documentation spells out the order: "A message that's held for your approval reaches the hook first, so a mod can read a message you haven't approved yet." Fourth, there is a call that injects a prompt as though the user had typed it. Normally a sentence saying a mod sent it goes in front, and an option to send it as the user's own words drops that sentence.

Two more features in the official documentation belong here, with a condition attached. One event rewrites each line destined for the transcript before it is stored, and another can block analytics records. These two need the condition read with them. As checked on 6 October 2026, the type declaration file in the public repository is on build v2.1.277, and neither event exists in it. The documentation is written against v2.1.289. The documentation even warns about that gap itself: "The copy on GitHub can be older than the Claude Code version you have installed." So we treat these two as documented features and do not claim to have confirmed them in source.

6.2Which boxes the guard covers and which it does not

Check the seats in section 6.1 against the guard's list of 15 from section 3.3 and the boxes split. The attribution text and the system prompt sections are among the things the guard keeps out of the user tier's reach. The design intent is that the instructions an organisation manages reach the model without being edited by a user's mod. The session-family events, on the other hand, are named on the guard's pass-through list, and the analytics side is not among the 15 the guard hooks at all, so it goes straight by. Which means that even on a machine where the guard is on, the seat where session and agent messages land first and the analytics side are both open.

Hold the guard's 15 against the seats in 6.1 What the guard keeps from the user tier Commit / PR attribution text System prompt sections on a machine where the guard loads What the guard lets through as-is Session / agent messages Analytics records open even with the guard on Left is the guard's block list; right is its pass-through list, or events it never hooks

Pebblous drew this by holding the guard's list of 15 from section 3.3 against the seats sections 6.1 and 6.2 name. The left two are kept out of the user tier's reach; the right two — session/agent messages and analytics records — stay open even on a machine where the guard is on.

In an agent pipeline, the thing serving as the audit record is the transcript, and the layer that writes that transcript runs on the same permissions as the thing it records. That changes what it takes to trust it. What was written comes second; what was loaded at the time of writing comes first. The question an earlier piece asked about where permission approval and the audit trail live, across a vendor boundary, moves inside a single process here.

6.3A new interception layer appears mid-session

The easiest way to build a mod is to ask Claude. The flow the documentation describes runs like this. A person asks in plain language, and Claude writes a module into a folder named for the session id. When the first file is saved, Claude Code asks once: enable hot reloading for this session? Say yes and the mods in that folder load when the turn ends, and after that "reload at the end of each turn that changes them." That answer lasts for the session, including after you resume it.

Approval happens once per session, and changes after it land per turn. Writing the files has its own gate, though. In the default permission mode and in the mode that auto-approves edits, that folder is a protected path, so each file prompts. What happens in other permission modes was not confirmed against primary documentation in this research, so we do not state it. This is the place where the rhetoric about an agent reprogramming itself becomes literally true, and the range over which it is true is that session and that folder.

6.4So did Mods create a new risk?

The evidence pointing the other way has to carry equal weight for this report to do its job. Start with the fact that the comparison baseline is nothing at all. Settings hooks were running shell commands with the user's permissions long before Mods, and there was no command to print what they hooked and what they called. VS Code extensions have no capability declaration either. Within this category, the thing that increased is information, and Mods is the side that increased it. It even comes with a device that refuses to load code it cannot read statically. So what Mods enlarged may be the shape and visibility of the risk more than its size.

The case against is not weak either. As section 5.3 showed, a capability list informs you only while its items are rare, and the ecosystem with full review and signing stopped there too. On the agent side, a separate literature targets the approval machinery itself. A synthesis paper posted in January 2026 gathers 78 studies going back to 2021, classifies 42 prompt injection techniques against coding assistants, and reports that adaptive strategies exceed an 85% success rate even against current defences. One path included there overlaps exactly with the shape of this report. If an injected instruction switches on tool auto-approval in the editor's settings file, every tool call after it runs without human confirmation. That puts the switch that turns approval off inside what the agent can write to, and this report adds one line to that list. The layer that answers permission verdicts on your behalf sits inside what gets loaded.

So what this report recommends is not stacking on more devices. It is knowing how far the devices you have reach. The argument that you have to see something before you can stop it was covered in an earlier piece, and what gets added here is the condition that even seeing depends on what is loaded. Three lines finish the check. Does my machine meet the conditions under which the guard loads? Does the call list of the mod I am about to install include process execution? Does the safety device I built have a failure handler?

Why This Matters to Pebblous

The runtime we work in became programmable

We should declare our interest first. Pebblous runs several agent pipelines on Claude Code in production, and the pipeline that produced this report is one of them. So this is not news about somebody else's tool. As the table in section 1.3 records, hooks run on the unattended surfaces too. For an organisation like ours that drives pipelines with claude -p, that means a mod can attach in the places where nobody is watching the screen. Turn the question DataClinic asks of an incoming dataset, and the question AI-Ready Data asks before training, back on the pipeline itself, and it reads like this. What code was loaded when this output was produced?

The quality of a record comes from what was loaded

Data quality work usually deals with missing values, duplicates and ranges. The defect this report points at sits one level above that. As section 6.2 shows, in an agent pipeline the layer that writes the record and the thing being recorded run on the same permissions. That changes what it takes to trust the sentence "the transcript says so." The quality of that record is set by a condition that comes before its contents: what was loaded at the moment it was written. Anyone who has designed data lineage knows the shape. The trustworthiness of lineage comes from how far the thing writing it sits from the thing it describes. Measure the transcript against the requirements for audit evidence that an earlier piece drew from banking, and the independence box is empty.

Three things you can check today

There are three things an organisation reading this can do now. One, check whether deny rules hold on your machine. Does it have managed settings, and are you signed in on a company plan? If neither, the dotted box in section 3.2 is empty on your machine. The procedure is already in the official documentation: run /status in a session and read the settings-sources line for a managed settings entry. Two, run the validation command against a mod before you install it and read the list of calls. If process execution is on that list, that one line means "anything." Three, if you built a safety device as a mod, check whether it has a failure handler. Without one, the device lets the call through when it fails instead of blocking it. All three are methods the official documentation already describes. This report invents no new control; it records how far the existing ones reach.

An item with no box on the review sheet yet

One note on the Korean market, where we work. The AI tool adoption review criteria that are publicly documented there are built around axes like the provenance of the model, the chance of contaminated training data, exposure to prompt injection and how personal data is handled. There is no box asking what is currently loaded into the developer tool. We also found no public case of a Korean company deploying Claude Code with managed settings. Where there is nothing, we say there is nothing.

AI agent governance has so far stayed on the question of what to forbid. This case points one step ahead of that. The place where a prohibition is written and the scope it applies to can differ. What Pebblous can do is put that distinction into words from the data side. That having a rule and having a rule that reaches are different things, and that the procedure for checking the boundary is already in public documentation. Both are sentences that transfer into a data quality specification.

Every verbatim quotation here was checked directly against fourteen pages of Claude Code's official documentation, the public source, README and tests of the built-in guard, three modules in the official example repository, and the official documentation of the products used for comparison. Mods is five days old, so adoption or usage figures do not exist and none are used. The "turning point" verdict that secondary coverage spread had no attributable speaker, so it was dropped; the claim of automatic sensitive-data deletion is corrected in section 2.4. We describe no attack method and no bypass procedure, and the only code quoted is the examples the official documentation already publishes. Sections 1 through 5 and sections 6.1 to 6.3 are what the official documentation and public source state; sections 2.3, 3.5, 4.5 and this section are the part those documents do not cover, so please read them separately. Thank you for reading this far.

References

The subject — official documentation and public source

  • 1.Anthropic, Claude Code docs — Mods. overview · admin · events · api · reference · create. Append .md to any of these URLs and the full markdown comes back as is. The 45 events, 21 namespaces, 80 methods and the limits table were recounted directly from the v2.1.289 tables on the reference page.
  • 2.Anthropic, Permissions — the deny, ask, allow precedence and the four bullets in "Extend permissions with hooks." The verbatim in section 3.1 comes from here.
  • 3.Anthropic, Plugin security — the boundary of what permission rules and sandboxing cover. The verbatim in section 2.2.
  • 4.Anthropic, Deploy managed settings — the per-OS deployment paths and the procedure for confirming they apply by reading the settings-sources line in /status.
  • 5.Built-in guard sec-default source — README, hooks/register.ts (3,643 bytes), hooks/policy/, tests/. The 15 events it hooks and the 2 methods it calls were counted in the registration file, and the notice string users see was checked against the return string in the source. The repository has no test covering the case where the guard does not load.
  • 6.mods/types/claude-code.d.ts — as checked on 6 October 2026, the version string on its first line reads v2.1.277. That is 12 patches behind the documented baseline (v2.1.289), and the two events qualified in section 6.1 are absent from it. Counted against this version, the totals come to 20 namespaces and 78 methods.
  • 7.Official example mod repository — the source of blast-radius, token-weather and replay-theater was read in full and the registration calls and failure handlers counted. The repository describes itself as shared as is, without support.

Design principles — the classics and the lineage

  • 8.Saltzer, J. H., & Schroeder, M. D. (1975). The Protection of Information in Computer Systems. Proceedings of the IEEE 63(9), 1278–1308. Full text — the verbatim from the fail-safe defaults item in section 4.5 was confirmed in the full text on the author's own site.
  • 9.Anderson, J. P. (1972). Computer Security Technology Planning Study, ESD-TR-73-51. Original PDF. ⚠️ The original is a scanned image and machine extraction failed, so the verbatim of the three requirements in section 3.5 was taken from the US DoD evaluation criteria, Part II §6.1, which quotes it directly.
  • 10.Miller, M. S. (2006). Robust Composition: Towards a Unified Approach to Access Control and Concurrency Control. PhD thesis, Johns Hopkins University. Full PDF — the definition of attenuation by interposition used in section 2.3. ⚠️ We confirmed the substance of the definition and did not quote at page level.
  • 11.Endo / SES · LavaMoat documentation · MetaMask, LavaMoat and the Ledger Software Supply Chain Attack · LeastAuthority, MetaMask Plugin System + LavaMoat audit report — the "language-level boundary, not an OS boundary" point in section 2.3 and the case where an environment variable skipped the isolation load.
  • 12.Felt, A. P., Greenberg, A., Chin, E., Hanna, S., & Wagner, D. (2011). The Effectiveness of Application Permissions. USENIX WebApps '11. PDF — the warning-fatigue argument in section 5.3. ⚠️ Machine extraction of the PDF failed and we could not secure a verbatim, so the body carries the direction of the finding without figures. Kariryaa et al. (SOUPS 2021) and Reeder et al. (CHI 2018) belong to the same line, but we could not confirm their sample figures at first hand and did not use them.
  • 13.Maloyan, N., & Namiot, D. (2026). Prompt Injection Attacks on Agentic Coding Assistants. arXiv:2601.17548 (submitted 2026-01-24) — the figures in section 6.4 were confirmed in the abstract: 78 studies synthesised, 42 attack techniques classified, and adaptive strategies exceeding 85% success. The abstract does not say which tools were tested or how many, so we do not state it; the path that switches on auto-approval settings reached us through a secondary summary, so it is described without figures.

Official documentation of the products compared

Related Pebblous blog posts