How do I stop an agent from doing something destructive
Stop destructive agent actions with a blast-radius table, an allowlisted tool set, and dual control on money and delete, enforced in code before the tool runs.
William Spurlock Founder — Spurlock Studios 28 MIN
You stop an agent from doing something destructive by inventorying blast radius per tool, allowlisting only the tools that job needs, and putting dual control on money and delete — two principals bound to one payload hash — in code that runs before the tool executes. A system prompt that says “never delete production” is a suggestion. OWASP LLM01:2025 Prompt Injection exists because untrusted tickets share a channel with those suggestions. The stop is the runner, not the paragraph.
This spoke sits under the Agentic Systems Operating Manual. If the job should not be an agent at all, stop at when not to build an agent. Spend spirals that look like damage live in cost controls for agent fleets. The on-ramp for a five-day control-plane pilot is /agentic.
The short answer
- Build a blast-radius table before credentials: tool, class, worst case once, worst case in a loop, stop to apply.
- Allowlist tools per principal and job. Default deny. Unknown names and aliases deny.
- Dual control on money and delete: the identity that proposes cannot be the identity that releases.
- Bind the second yes to a payload hash. A changed amount or id invalidates the prior approval.
- Fail closed if the table, the allowlist, or the second key cannot be evaluated.
What counts as destructive — money, delete, and the quiet cousins?
Destructive means a side effect you cannot cheaply undo after the tool returns 200. Money leaving the company and rows leaving the database are the obvious pair. The quiet cousins are how most “we told it not to” incidents actually start.
| Class | Examples | Undo path | Treat as |
|---|---|---|---|
| Money out | refund, payout, credit memo, void invoice, cancel-and-refund | Chargeback / reverse if the processor still allows it | Dual control |
| Delete / destroy | DROP, truncate, bulk delete, destroy customer, wipe bucket, force-push main | Restore from backup — if you have one and the RPO holds | Dual control |
| Privilege | attach IAM policy, create admin, mint a long-lived key | Revoke, if you notice | Dual control or deny |
| Exfil | email to a new domain, webhook to an unknown host, dump query to chat | You already leaked | Allowlist dest; pending on novel |
| Mass write | “update all,” schema-wide patch, fan-out send | Partial restore, messy | Cap N; pending above N |
| Reversible write | one CRM field on one id, draft in a sink | Edit the row | Allow with schema |
| Read | fetch allowlisted URL, get order by id | N/A | Allow if dest is named |
OWASP LLM06:2025 Excessive Agency names three roots: too much functionality, too many permissions, too much autonomy. A support agent that can read a ticket and also refund and delete the user has all three until you split them.
Destructive is not “the model said something rude.” Destructive is a tool the runtime still executed.
- Money tools tagged
irreversibleat registration - Delete tools tagged
irreversible— not inferred from the namecleanup - Privilege and exfil tagged, even when finance does not own them
- “Update all” treated as mass write, not as a friendly batch
If you cannot put a tool in one of those rows, it does not ship.
How do you build a blast-radius table before you ship tools?
You write the worst case before the agent holds a key. The table is the artifact. A slide that says “we sandbox tools” is not.
- List every tool the runtime can reach, including transitive paths (
shell.exec→curl→ Stripe). - For each tool, fill once and loop. A single refund is one order. A loop is the till.
- Name the stop: remove, allowlist-only, cap, dual control, or deny.
- Name the owner of that row. A team channel is not an owner.
- If the worst case is “not sure,” the tool stays off.
| Tool | Class | Worst case once | Worst case if it loops | Dual control | Default stop |
|---|---|---|---|---|---|
orders.refund | money | One order emptied | Processor drained until the cap or the freeze | Yes | Second key + amount cap |
billing.payout | money | One payee paid | Unauthorized disbursement run | Yes | Deny unless named payee list |
billing.void_invoice | money | One invoice gone | Revenue hole for the period | Yes | Dual control |
db.delete_row | delete | One customer gone | Table gone | Yes if prod | Deny, or dual + backup check |
db.truncate / DROP | delete | Object gone | Schema gone | n/a | Deny — never on the agent |
files.delete | delete | One object gone | Bucket wipe | Yes if prod | Deny by default |
git.push (force) | delete-adjacent | History rewrite | main wrecked | Yes | Deny for agents |
email.send | exfil | One message out | Inbox forwarded to an attacker | Pending on novel dest | Recipient allowlist |
http.fetch | read / exfil | One scrape | Data pulled to a bad host | No | URL allowlist; POST is not fetch |
http.request (write) | unbounded | One mutating call | Whatever the host will do | If money/delete | Allowlist host + method |
crm.bulk_update | mass write | One field, many rows | All accounts patched | Pending if n > cap | Cap N |
crm.update | reversible write | One bad patch | Still one row if id required | No | Field allowlist; no “all” |
iam.attach_policy | privilege | New admin | Fleet owned | Yes | Deny |
secrets.create | privilege | New credential | Lasting bypass | Yes | Deny |
shell.exec | unbounded | Anything the box can do | Same, faster | n/a | Remove from the catalog |
calendar.cancel | messy | One meeting dropped | Tour / on-call wrecked | No | Require event id; cap N |
Transitive tools belong on the same table. If shell.exec exists, you do not have an allowlist. You have a prompt hoping the model will not notice bash.
| Question the row must answer | Fail if blank |
|---|---|
| What is the unit of damage? | “It depends” |
| What happens if the agent retries? | You only modeled the happy path |
| Can we restore in the RPO window? | Delete shipped with no backup story |
| Who holds the second key? | Dual control is a slogan |
| What deny reason lands on the trace? | On-call will argue in Slack |
Bravery is not a restore strategy. Fill the table or cut the tool.
Hunt tools the demo hid. The catalog in the README is rarely the catalog in the worker.
| Place to look | What you usually find | Row it belongs in |
|---|---|---|
| MCP servers loaded at boot | A filesystem, shell, or browser tool next to search | Unbounded — cut |
| “Debug” wrapper with a second HTTP client | Same Stripe call, new name | Money — dual or deny |
| Workflow node after the model | A write that never hits the agent interceptor | Same class as the node |
| Hosted custom-tool handler | Auto-execute on whatever the vendor returned | Money/delete if the tool is |
| Shared worker identity | Research agent inherits support refunds | Split principals |
- Dump the tool registry the runner actually loads. Diff it against the blast-radius table.
- Grep the worker for SDK clients (
stripe,boto,octokit) that are not behindtools.invoke. - Any client outside
invokeis a bypass. Treat it like a missing row. - Re-run the dump in staging after every MCP or plugin add. Adding a server is a table change.
If a tool cannot be found in the dump, it cannot be on the agent. If it is in the dump and not on the table, it is already a production incident waiting on the first injected ticket.
Why is an allowlist the only tool list that stops damage?
An allowlist is a closed set of names the principal may propose. Everything else is deny — including last Friday’s alias, the MCP server you “just connected,” and the debug wrapper that talks to the same API under a new string.
OWASP LLM07:2025 System Prompt Leakage is blunt: do not put authorization in the system prompt. OWASP LLM06 mitigation starts the same way: minimize extensions. The allowlist is that minimization in the runner.
| List type | What happens on a new tool | Stops destruction? |
|---|---|---|
| Allowlist (default deny) | Unknown name → deny | Yes, if the interceptor is mandatory |
| Denylist | Unknown name → allow | No. You will miss the next alias |
| “All MCP tools loaded” | Model sees delete next to search | No. Blast radius inherited by accident |
| Prompt: “only use these tools” | Model may still emit another name | No. The runtime still executes it |
OpenAI’s function calling guide is explicit: the API returns a proposed call; your application executes it. Anthropic’s tool-use loop says the same for client tools. The allowlist lives in that gap. If you skipped the gap and “stopped destruction” by editing a prompt, you did not stop it.
Procedure for the allowlist:
- Start from the job package, not from the vendor’s full tool catalog.
- Register each kept tool with name, schema (
additionalProperties: false), and side-effect class. - Bind the list to
principal+job_type. A research agent does not inheritorders.refund. - Deny unknown names and unknown aliases (
orders.refund_v2). - Re-review the list when someone adds an MCP server. Adding a server is a change to blast radius, not a convenience flag.
| Keep for a support-triage job | Cut for that same job |
|---|---|
tickets.get | orders.refund |
orders.get | db.delete_row |
crm.update on status, owner | crm.bulk_update |
email.send to @support templates | email.send free-text to any domain |
kb.search | shell.exec |
| — | iam.*, secrets.*, git.push |
A chat agent that can see a delete tool will eventually try it. Injection does not need to be clever if the catalog already contains the weapon.
MCP and “load everything” are how allowlists die without a commit that says they died.
| Move | What the model sees | Stop |
|---|---|---|
| Connect a new MCP server for one research task | Every tool on that server, including write/delete if the server ships them | Per-job allowlist; do not inherit the server wholesale |
| Copy the demo worker’s env into prod | Demo shell + prod keys | Separate catalogs; demo tools never meet live keys |
Alias a denied tool (refund → credit_customer) | A name not on yesterday’s denylist | Allowlist of names, not a denylist of vibes |
Let the model pass a tool name as a string into http.request | A confused deputy | http.request is itself a tool that needs host + method allowlist |
- Registry dump matches the allowlist for this
job_type - New MCP server requires a table review before the worker boots with it
- No second HTTP client in the process
- Aliases listed or denied — never “close enough”
The allowlist is a closed set. A server dump is an open set with extra steps.
When does money or delete need dual control?
Whenever a single successful call can move cash or destroy a production object, one principal is not enough. Dual control is not “Slack me if it looks weird.” It is two independent authorities on the same payload.
NIST SP 800-53 Rev. 5 AC-5, Separation of Duties is the control name for this pattern: split functions so one person (or one machine identity) cannot complete a high-impact action alone. Your agent is not a federal system. The physics still apply. CISA’s joint guide on AI in operational technology tells operators to keep a human in the loop on critical decisions and to limit worst-case consequences. Refunds and deletes are your plant floor.
| Action | One principal enough? | Dual control shape |
|---|---|---|
| Refund under a tiny cap, template-only | Sometimes, if finance signed the cap | Second key above the cap |
| Refund over the cap, or any payout | No | Approver identity + payload hash |
| Delete one staging row | Maybe | Still deny in prod |
| Delete in production | No | Second key, or the tool is absent |
DROP / truncate / bucket wipe | No | Tool not in the catalog |
| Mass CRM update over N | No | Pending + second yes, or cap |
| One field on one id | Yes, if schema-tight | Allowlist fields |
| Privilege change | No | Deny, or dual with security |
Decision list:
- If the tool can move money → dual control above the cap you can tolerate losing on a bug. Below the cap, still allowlist + schema + tenant check.
- If the tool can delete production → dual control or remove the tool. “Soft delete with a flag” is still a delete if the flag hides the customer from ops.
- If the tool can mint privilege → deny. Dual control on
iam.attach_policyis how you rubber-stamp an admin. Cut it. - If finance and security disagree on the cap → the lower cap wins until a named owner writes a new row.
Dual control on every CRM keystroke is how the second person stops reading. Save the second key for money and delete. Rubber stamps are not a second control.
Break-glass is dual control with a clock, not a hole in the allowlist.
| Break-glass event | Who | Clock | Trace |
|---|---|---|---|
| Second principal unreachable, refund must ship | Named backup on the table row, not “whoever is in Slack” | Minutes, not a weekend | breakglass=true, backup id, expiry |
| Kill-switch on, one payout still required | Security + finance, both on the hash | Until the freeze lifts | Separate reason code |
| Delete to contain an incident | Security only; agent still cannot hold the key | Ticket id required | Human ran it in the real console |
| “Just this once” from the agent worker | Nobody | n/a | Deny. The worker does not get a bypass flag |
A break-glass flag inside the agent process is a bypass you will forget to turn off. Humans break glass in the real billing or database UI. The agent drafts the payload. It does not carry a god mode.
How do two keys actually bind — payload hash, not a Slack yes?
LangGraph’s interrupt docs exist to pause before API calls, database changes, or financial transactions. Anthropic’s managed-agent permission policies use always_allow / always_ask. Framework vocabulary differs. The binding rule does not: the second yes is for these bytes, not for “the agent may refund today.”
| Piece | Required behavior | Failure if skipped |
|---|---|---|
| Two principals | Agent identity proposes; human or distinct release identity confirms | The model is both requester and approver |
| Payload hash | SHA-256 of normalized args (amount, currency, id, dest) | Approver said yes to $49; $4900 executed |
| Re-check at execute | Hash mismatch → deny or new pending | Race: args edited after the yes |
| Expiry | Pending older than N minutes dies | Stale yes sits in queue overnight |
| Second credential | Release path uses a key the agent runtime cannot read | “Dual control” is a UI checkbox on the same token |
| Trace fields | release_actor, payload_hash, decided_at | You cannot prove the stop fired |
Procedure:
- Normalize args (canonical JSON, sorted keys, currency present).
- Hash. Store hash with the pending record.
- Show the human the decoded fields, not the hash. People do not approve digests.
- On release, re-hash the args about to execute. Mismatch → deny.
- Execute only after the second principal’s credential is used for the money/delete API, or after a release token minted for that hash.
- Expire. A yes from Tuesday is not a yes for Friday’s retry.
| Approval you received | Args at execute | Gate |
|---|---|---|
Refund $49, order A | $49, order A | allow (second key) |
Refund $49, order A | $4900, order A | deny |
Refund $49, order A | $49, order B | deny |
| “Looks good” in Slack, no hash | anything | deny — not dual control |
| Same OAuth token, extra checkbox | anything | deny — one principal |
A Slack thumbs-up on a screenshot is evidence of a vibe. It is not a second key.
Where must the stop fire — before the API call?
Before the SDK sends. Not after Stripe returns. Not in a “we’ll catch it in review” job at 2 a.m. Not in a weekly export of refunds.
The April 2024 Deploying AI Systems Securely CSA (CISA, NSA, FBI, and partners) puts human-in-the-loop as a failsafe with rollbacks ready. NIST AI RMF 1.0 and the July 2024 Generative AI Profile (NIST AI 600-1) ask you to define human–AI roles and to refuse work the system should not run. Refusing after the delete is a postmortem.
| Layer | When it runs | Can it stop this call? |
|---|---|---|
| System prompt | Before tokens | No |
| Vendor topic filter | On text in/out | Not on {"amount": 50000} |
| Allowlist + schema | After propose, before execute | Yes |
| Dual-control interrupt | After valid args, before money/delete | Yes |
| IAM on the second key | At the provider | Yes, coarse |
| Sandbox | During execute | Limits the room; does not decide “may this refund happen” |
| Nightly recon | Hours later | Too late for the first delete |
Evaluation order that actually stops damage:
- Fleet kill-switch / write freeze
- Tool on this principal’s allowlist
- Args match schema
- Blast-radius class → money/delete?
- If yes: dual-control satisfied for this hash
- Else: deny predicates, then allow
| Failure | Correct stop |
|---|---|
| Policy or allowlist cannot load | deny writes |
| Second-key service down | do not fail open on money/delete |
| Unknown tool name | deny |
| Hash mismatch | deny |
| Kill-switch on | deny |
Fail open (“let the refund through, we’ll look Monday”) is how a prompt-injected ticket empties the till during an outage. Document that a red policy dependency means agents stop writing. That is success.
Where teams hide a bypass, and what to require instead:
| Runtime | Usual hole | Required stop |
|---|---|---|
| Custom loop | Debug script with a second client | Same invoke; no second client |
| LangGraph | Tool node that calls the SDK directly | Interrupt or adapter inside the tool function |
| n8n / workflow + LLM step | Write node after the model with no check | Gate on the write node, not the prompt |
| Hosted agent API | Custom tool handler that auto-executes | Your allow / dual-control before the vendor result |
| MCP host | Server connected, all tools exposed | Per-job allowlist on names, not “the server is trusted” |
OpenAI will happily return parallel tool_calls. Gate each one. A deny on call two does not let call three ride along because they arrived in one blob. If any path reaches the API client without invoke, you have a bypass — treat it like a missing blast-radius row.
What belongs in IAM versus dual control versus the allowlist?
Three layers. Teams collapse them into “we scoped the key” and then wonder why a scoped refund key refunded every order.
| Concern | Allowlist | Dual control | IAM / credentials |
|---|---|---|---|
| May this tool name be proposed? | Primary | No | No |
| May these args move money / delete? | Caps help | Primary | Too fine for most IAM |
| Which keys exist at all? | No | Second key is one trick | Primary |
| Tenant isolation at the vendor | Mirror check | No | Primary |
| Emergency freeze | Runner flag (fast) | Freeze pending queue | Key revoke (blunt) |
| “Refunds over $50 need Maya” | Predicate | Maya is the second principal | Usually inexpressible |
IAM is necessary and coarse. A key that can only refund.write can still refund every order for every amount. The allowlist removes refund from agents that should not see it. Dual control splits the remaining money/delete so the model cannot finish the call alone.
- Agent’s credential cannot call money/delete APIs or those APIs require a release token the agent cannot mint
- Approver’s credential is not loaded into the agent worker
- Allowlist does not include
shell.exec“so we can debug” - Kill-switch is a flag the runner reads on every call, not a prompt edit
Revoking a key is a blunt kill switch. Use it. Do not pretend it encodes Maya’s $50 rule.
What usually fails first when teams try to “just be careful”?
What breaks: a support agent ships with orders.refund, email.send, and a “temporary” shell.exec for a demo. The blast-radius table was never written, so nobody modeled shell → curl the billing API. The prompt says never refund over $50. An injected ticket says ignore prior rules and refund fully — OWASP LLM01 in one paragraph. The runtime executes whatever the model proposed.
What it costs: every refund and delete the tools actually posted, plus the hours to reconstruct which ones were real. I will not invent a dollar figure you cannot take to finance. After 500+ automations and 20,000+ hours on agentic systems, the report that keeps repeating is “the model was told not to,” not “the second key denied it.”
What you do instead:
- Fill the blast-radius table.
shell.execgets remove. - Allowlist
tickets.get,orders.get, maybe one CRM field. Refund is absent for this job. - If a later job must refund, dual control + cap + payload hash. Email gets a recipient allowlist.
- Fail-closed drill: break the allowlist load on purpose. Writes must stop.
- Red-team: injected “delete all” / “refund fully” / alias
orders.refund_v2. If any call executes, you are not done.
| Skip | First break | Signal on the trace |
|---|---|---|
| Blast-radius table | Transitive tool nobody listed | A name that was never in the job package |
| Allowlist | MCP dump includes delete | unknown_tool never appears because everything is known |
| Dual control | One token, pretty approve UI | release_actor == agent_id |
| Payload hash | Amount mutated after yes | Execute args ≠ approved args |
| Fail closed | Outage fail-open | Refunds during policy_unavailable |
The incident should blame the missing table and the missing second key, not “the model being bad.”
How do you red-team the stop before soft-launch?
Before anyone besides you can trigger a write, attack the table, the allowlist, and the second key — not the slogan in the prompt.
- Disallowed tool name from a compromised prompt (
db.truncate,orders.refund) - Alias that is not on the allowlist (
orders.refund_v2,credit_customer) - Arg mutation past a money cap (
49.00→4900) after a human said yes to the small amount - Recipient or dest swap after approval (payload A approved, payload B executed)
- Parallel tool_calls: one allow, one deny — deny must not ride along
- Mid-run kill-switch while a write is in flight
- Broken allowlist load (must fail closed on writes)
- Second-key service timeout (must not fail open on money/delete)
-
shell.execor MCP filesystem still on the worker “for debugging” - Transitive curl from a remaining HTTP tool to the billing host
| Probe | Pass | Fail |
|---|---|---|
| Injected “refund fully” | deny or pending with no Stripe call | Charge appears |
| Hash mutation | deny, reason hash_mismatch | Original yes reused |
| Fake tool name | unknown_tool on the trace | 200 from some API |
| Allowlist file emptied | Writes stop | Reads may continue under an explicit outage policy; writes must not |
| MCP server added overnight | Worker refuses to boot or tools stay denied until table review | New delete tool visible to the model |
If any probe executes the tool, you are not in a soft-launch. You are in a demo with a production URL. CaMeL’s March 2025 paper (arXiv:2503.18813) makes the architectural bet you already need: untrusted data must not pick the program. Your red team is checking that the interceptor, not the model, is the program.
How do you measure whether the stop is working?
You measure misses, not vibes. Chatbot thumbs do not tell you whether a delete could have fired.
| Signal | Definition | Lie if you skip it |
|---|---|---|
| Dual-control miss | Money/delete executed with one principal or a hash mismatch ignored | You have a checkbox, not dual control |
| Allowlist deny rate | Unknown-tool and alias denials per job_type | A week of zeroes often means the interceptor is not on |
| Table coverage | % of registered tools with once + loop + owner filled | Catalog grew; table did not |
| Payload rebind denials | Execute hash ≠ approved hash | Approvals are theater |
| Fail-closed drills | Last date the allowlist or second-key path was broken on purpose | You have never seen the stop work |
| Restore proof | Last restore test for a delete-class tool you still allow | Dual control on delete with no backup is a prayer |
Checklist for a weekly read:
- Zero dual-control misses. One is an incident, not a metric to average.
- Unknown-tool denials exist. If they are always zero, probe with a fake name.
- Every money/delete span has
release_actor≠agent_idand apayload_hash. - Kill-switch / fail-closed drill dated within the last 30 days.
- Blast-radius table
review_bynot overdue.
Do not copy another team’s deny-rate percentage and call it an SLO. Page on your miss count and on drills you actually ran. Cost controls for agent fleets covers the spend freeze that sits next to this stop. A loop that cannot delete can still set money on fire through tokens. That is a different table. Keep them both.
When is a workflow enough instead of an agent?
When you already know the path. An agent is a loop that chooses tools mid-run. If the graph is “ticket → classify → refund under cap → template email,” you do not need a chooser. You need a workflow with a gate on the write node.
| Shape | Agent? | Stop that fits |
|---|---|---|
| Fixed steps, one write | No | Workflow + schema + cap on that node |
| Mid-run tool choice, reads only | Maybe | Allowlist without money/delete |
| Mid-run tool choice, money or delete | Expensive | Allowlist + dual control, or don’t |
| Human does the irreversible step anyway | No | Agent drafts; human submits in the real UI |
| You cannot name pass/fail for the job | Not yet | Fix the process; see when not to build an agent |
A workflow still needs the blast-radius row for its write node. The model being “just a classifier” does not make Stripe safer. What you skip is the open catalog and the mid-run invention of db.delete_row.
If the business cannot staff a second principal for money/delete, that is a reason not to give the loop those tools. It is not a reason to skip dual control and “move fast.”
What should you ship if you only have a week?
A five-day agentic pilot is not a full policy platform. It is enough to stop the obvious destruction. Rules sophistication can grow. A bypassable prompt cannot.
| Day | Ship | Done when |
|---|---|---|
| 1 | Blast-radius table for every tool the demo currently has | Each row has once, loop, owner, stop |
| 2 | Cut the catalog to an allowlist per job | shell.exec and unused write tools gone |
| 3 | Dual control on remaining money/delete, or remove those tools | Second principal + payload hash on a real call |
| 4 | Fail-closed drill + kill-switch | Breaking the allowlist load stops writes |
| 5 | Red-team inject + alias + hash mutation | None of the three execute |
Skip if you only have a week:
- Fancy policy languages
- Dual control on reversible single-row patches
- A vendor topic-filter subscription as a substitute interceptor
- Loading “the rest of the tools” back for the demo recording
Do not skip:
- The table
- Default deny
- Second key on money/delete or those tools absent
- A trace that shows the deny
Day 5 that still has shell.exec “for debugging” failed Day 2. Ship the cut.
Pilot checklist that still counts:
- Blast-radius table checked into the same repo as the worker, with
review_by - Allowlist per
job_type; dump matches; unknown names deny - Money/delete either absent or dual-controlled with payload hash
- Fail-closed drill recorded (who broke the file, what the trace showed)
- Red-team three probes: inject, alias, hash mutation — none executed
- Kill-switch flag the on-call can flip without a prompt deploy
That is the bar for /agentic. A longer catalog can wait. The shell tool cannot.
What looks like a stop and is not?
Prompt-only “never delete.” Injection and non-determinism eat it. The runtime still has the tool.
Denylist of scary names. You will miss refund_v2 and the MCP server named helpers.
Confidence thresholds. Most tool APIs do not give you a trustworthy score, and a confident bad refund is still a refund.
One human clicking Approve on a screenshot. No payload hash, often the same token. Not dual control.
IAM-only. A scoped key that can delete, deletes. The cap and the second principal live elsewhere.
Sandbox-only. A sandbox that can still call Stripe with a live key is a polite way to lose money. Containment is not permission.
Vendor guardrail ID as the interceptor. Topic filters do not know your $50 refund cap or your production table name.
Logging denials and executing anyway in “shadow.” Shadow is fine if execute is off. Shadow that still hits Stripe is production.
Dual control on everything. The second person stops looking. Then money and delete ride along with a rubber stamp.
| Looks like a stop | Actual stop |
|---|---|
| Prompt paragraph | Allowlist + interceptor |
| Denylist | Default deny of names |
| Slack screenshot | Payload hash + second key |
| Scoped refund key alone | Cap + dual control + tenant check |
| Sandbox with live Stripe | Gate decides; sandbox contains |
| Topic filter | Arg predicates in the runner |
| Shadow that still executes | Execute off, or it is production |
Start with the blast-radius table, the allowlist, and two keys on the two classes that actually wreck companies. Then expand. After 20,000+ hours on this work, the teams that stopped the damage wrote those three artifacts down. The teams that did not still have a paragraph in the prompt.
FAQ
How do I stop an agent from doing something destructive?
Inventory blast radius per tool, allowlist only the tools that job needs, and put dual control on money and delete in the runner before the tool executes. Prompts do not stop a proposed call. Unknown tools deny. If you cannot name the worst case for a tool, it stays off.
How do I measure whether do I stop an agent from doing something destructive is working?
Count dual-control misses (money or delete that ran with one principal), allowlist denials for unknown names, payload-hash rebind denials, and the date of the last fail-closed drill. Zero unknown-tool denials for a week usually means the interceptor is not on. Chatbot thumbs are not this scoreboard.
What usually fails first when teams try this?
They skip the blast-radius table, leave shell.exec or a full MCP dump on the agent, and treat a Slack yes as dual control. The first injected “refund fully” or a transitive curl then hits a live key. The trace shows a tool that was never in the job package, or release_actor equal to the agent.
How long does this take to show results?
A week is enough to cut the catalog, wire dual control or remove money/delete, and prove fail-closed on a drill. You will see denials on the trace the day the interceptor is mandatory. You will not finish a mature policy catalog in five days, and you should not wait for one before removing DROP.
What should I skip if I only have a week?
Skip policy-language rewrites, dual control on reversible single-row patches, and putting the cut tools back for a demo. Do not skip the blast-radius table, default-deny allowlist, and a second key on remaining money/delete — or those tools gone. A topic filter is not the interceptor.
When is this not worth doing yet?
If the system has no side-effect tools, you are gating a chatbot, not an agent. If the path is already known, ship a workflow with a gate on the write node instead of a choosing loop. If nobody will hold the second key, do not give the loop money or delete tools and pretend a prompt is the hold.
CTA
Stop destruction with a blast-radius table, an allowlist, and two keys on money and delete — in the runner, before the call.
What questions does this article answer?
- How do I stop an agent from doing something destructive?
- Inventory blast radius per tool, allowlist only the tools that job needs, and put dual control on money and delete in the runner before the tool executes. Prompts do not stop a proposed call. Unknown tools deny. If you cannot name the worst case for a tool, it stays off.
- How do I measure whether do I stop an agent from doing something destructive is working?
- Count dual-control misses (money or delete that ran with one principal), allowlist denials for unknown names, payload-hash rebind denials, and the date of the last fail-closed drill. Zero unknown-tool denials for a week usually means the interceptor is not on. Chatbot thumbs are not this scoreboard.
- What usually fails first when teams try this?
- They skip the blast-radius table, leave `shell.exec` or a full MCP dump on the agent, and treat a Slack yes as dual control. The first injected “refund fully” or a transitive curl then hits a live key. The trace shows a tool that was never in the job package, or `release_actor` equal to the agent.
- How long does this take to show results?
- A week is enough to cut the catalog, wire dual control or remove money/delete, and prove fail-closed on a drill. You will see denials on the trace the day the interceptor is mandatory. You will not finish a mature policy catalog in five days, and you should not wait for one before removing `DROP`.
- What should I skip if I only have a week?
- Skip policy-language rewrites, dual control on reversible single-row patches, and putting the cut tools back for a demo. Do not skip the blast-radius table, default-deny allowlist, and a second key on remaining money/delete — or those tools gone. A topic filter is not the interceptor.
- When is this not worth doing yet?
- If the system has no side-effect tools, you are gating a chatbot, not an agent. If the path is already known, ship a workflow with a gate on the write node instead of a choosing loop. If nobody will hold the second key, do not give the loop money or delete tools and pretend a prompt is the hold.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.