Spurlock Studios
Contact
Share LinkedIn X
A violet ring. Thesis: STOP AGENT DOING SOMETHING DESTRUCTIVE.

You stop an agent from doing something destructive by inventorying blast radius per tool, allowlisting only the tools that job needs, and putting dual control on money and delete — two principals bound to one payload hash — in code that runs before the tool executes. A system prompt that says “never delete production” is a suggestion. OWASP LLM01:2025 Prompt Injection exists because untrusted tickets share a channel with those suggestions. The stop is the runner, not the paragraph.

This spoke sits under the Agentic Systems Operating Manual. If the job should not be an agent at all, stop at when not to build an agent. Spend spirals that look like damage live in cost controls for agent fleets. The on-ramp for a five-day control-plane pilot is /agentic.

The short answer

  • Build a blast-radius table before credentials: tool, class, worst case once, worst case in a loop, stop to apply.
  • Allowlist tools per principal and job. Default deny. Unknown names and aliases deny.
  • Dual control on money and delete: the identity that proposes cannot be the identity that releases.
  • Bind the second yes to a payload hash. A changed amount or id invalidates the prior approval.
  • Fail closed if the table, the allowlist, or the second key cannot be evaluated.

What counts as destructive — money, delete, and the quiet cousins?

Destructive means a side effect you cannot cheaply undo after the tool returns 200. Money leaving the company and rows leaving the database are the obvious pair. The quiet cousins are how most “we told it not to” incidents actually start.

ClassExamplesUndo pathTreat as
Money outrefund, payout, credit memo, void invoice, cancel-and-refundChargeback / reverse if the processor still allows itDual control
Delete / destroyDROP, truncate, bulk delete, destroy customer, wipe bucket, force-push mainRestore from backup — if you have one and the RPO holdsDual control
Privilegeattach IAM policy, create admin, mint a long-lived keyRevoke, if you noticeDual control or deny
Exfilemail to a new domain, webhook to an unknown host, dump query to chatYou already leakedAllowlist dest; pending on novel
Mass write“update all,” schema-wide patch, fan-out sendPartial restore, messyCap N; pending above N
Reversible writeone CRM field on one id, draft in a sinkEdit the rowAllow with schema
Readfetch allowlisted URL, get order by idN/AAllow if dest is named

OWASP LLM06:2025 Excessive Agency names three roots: too much functionality, too many permissions, too much autonomy. A support agent that can read a ticket and also refund and delete the user has all three until you split them.

Destructive is not “the model said something rude.” Destructive is a tool the runtime still executed.

  • Money tools tagged irreversible at registration
  • Delete tools tagged irreversible — not inferred from the name cleanup
  • Privilege and exfil tagged, even when finance does not own them
  • “Update all” treated as mass write, not as a friendly batch

If you cannot put a tool in one of those rows, it does not ship.

How do you build a blast-radius table before you ship tools?

You write the worst case before the agent holds a key. The table is the artifact. A slide that says “we sandbox tools” is not.

  1. List every tool the runtime can reach, including transitive paths (shell.exec → curl → Stripe).
  2. For each tool, fill once and loop. A single refund is one order. A loop is the till.
  3. Name the stop: remove, allowlist-only, cap, dual control, or deny.
  4. Name the owner of that row. A team channel is not an owner.
  5. If the worst case is “not sure,” the tool stays off.
ToolClassWorst case onceWorst case if it loopsDual controlDefault stop
orders.refundmoneyOne order emptiedProcessor drained until the cap or the freezeYesSecond key + amount cap
billing.payoutmoneyOne payee paidUnauthorized disbursement runYesDeny unless named payee list
billing.void_invoicemoneyOne invoice goneRevenue hole for the periodYesDual control
db.delete_rowdeleteOne customer goneTable goneYes if prodDeny, or dual + backup check
db.truncate / DROPdeleteObject goneSchema gonen/aDeny — never on the agent
files.deletedeleteOne object goneBucket wipeYes if prodDeny by default
git.push (force)delete-adjacentHistory rewritemain wreckedYesDeny for agents
email.sendexfilOne message outInbox forwarded to an attackerPending on novel destRecipient allowlist
http.fetchread / exfilOne scrapeData pulled to a bad hostNoURL allowlist; POST is not fetch
http.request (write)unboundedOne mutating callWhatever the host will doIf money/deleteAllowlist host + method
crm.bulk_updatemass writeOne field, many rowsAll accounts patchedPending if n > capCap N
crm.updatereversible writeOne bad patchStill one row if id requiredNoField allowlist; no “all”
iam.attach_policyprivilegeNew adminFleet ownedYesDeny
secrets.createprivilegeNew credentialLasting bypassYesDeny
shell.execunboundedAnything the box can doSame, fastern/aRemove from the catalog
calendar.cancelmessyOne meeting droppedTour / on-call wreckedNoRequire event id; cap N

Transitive tools belong on the same table. If shell.exec exists, you do not have an allowlist. You have a prompt hoping the model will not notice bash.

Question the row must answerFail if blank
What is the unit of damage?“It depends”
What happens if the agent retries?You only modeled the happy path
Can we restore in the RPO window?Delete shipped with no backup story
Who holds the second key?Dual control is a slogan
What deny reason lands on the trace?On-call will argue in Slack

Bravery is not a restore strategy. Fill the table or cut the tool.

Hunt tools the demo hid. The catalog in the README is rarely the catalog in the worker.

Place to lookWhat you usually findRow it belongs in
MCP servers loaded at bootA filesystem, shell, or browser tool next to searchUnbounded — cut
“Debug” wrapper with a second HTTP clientSame Stripe call, new nameMoney — dual or deny
Workflow node after the modelA write that never hits the agent interceptorSame class as the node
Hosted custom-tool handlerAuto-execute on whatever the vendor returnedMoney/delete if the tool is
Shared worker identityResearch agent inherits support refundsSplit principals
  1. Dump the tool registry the runner actually loads. Diff it against the blast-radius table.
  2. Grep the worker for SDK clients (stripe, boto, octokit) that are not behind tools.invoke.
  3. Any client outside invoke is a bypass. Treat it like a missing row.
  4. Re-run the dump in staging after every MCP or plugin add. Adding a server is a table change.

If a tool cannot be found in the dump, it cannot be on the agent. If it is in the dump and not on the table, it is already a production incident waiting on the first injected ticket.

Why is an allowlist the only tool list that stops damage?

An allowlist is a closed set of names the principal may propose. Everything else is deny — including last Friday’s alias, the MCP server you “just connected,” and the debug wrapper that talks to the same API under a new string.

OWASP LLM07:2025 System Prompt Leakage is blunt: do not put authorization in the system prompt. OWASP LLM06 mitigation starts the same way: minimize extensions. The allowlist is that minimization in the runner.

List typeWhat happens on a new toolStops destruction?
Allowlist (default deny)Unknown name → denyYes, if the interceptor is mandatory
DenylistUnknown name → allowNo. You will miss the next alias
“All MCP tools loaded”Model sees delete next to searchNo. Blast radius inherited by accident
Prompt: “only use these tools”Model may still emit another nameNo. The runtime still executes it

OpenAI’s function calling guide is explicit: the API returns a proposed call; your application executes it. Anthropic’s tool-use loop says the same for client tools. The allowlist lives in that gap. If you skipped the gap and “stopped destruction” by editing a prompt, you did not stop it.

Procedure for the allowlist:

  1. Start from the job package, not from the vendor’s full tool catalog.
  2. Register each kept tool with name, schema (additionalProperties: false), and side-effect class.
  3. Bind the list to principal + job_type. A research agent does not inherit orders.refund.
  4. Deny unknown names and unknown aliases (orders.refund_v2).
  5. Re-review the list when someone adds an MCP server. Adding a server is a change to blast radius, not a convenience flag.
Keep for a support-triage jobCut for that same job
tickets.getorders.refund
orders.getdb.delete_row
crm.update on status, ownercrm.bulk_update
email.send to @support templatesemail.send free-text to any domain
kb.searchshell.exec
—iam.*, secrets.*, git.push

A chat agent that can see a delete tool will eventually try it. Injection does not need to be clever if the catalog already contains the weapon.

MCP and “load everything” are how allowlists die without a commit that says they died.

MoveWhat the model seesStop
Connect a new MCP server for one research taskEvery tool on that server, including write/delete if the server ships themPer-job allowlist; do not inherit the server wholesale
Copy the demo worker’s env into prodDemo shell + prod keysSeparate catalogs; demo tools never meet live keys
Alias a denied tool (refund → credit_customer)A name not on yesterday’s denylistAllowlist of names, not a denylist of vibes
Let the model pass a tool name as a string into http.requestA confused deputyhttp.request is itself a tool that needs host + method allowlist
  • Registry dump matches the allowlist for this job_type
  • New MCP server requires a table review before the worker boots with it
  • No second HTTP client in the process
  • Aliases listed or denied — never “close enough”

The allowlist is a closed set. A server dump is an open set with extra steps.

When does money or delete need dual control?

Whenever a single successful call can move cash or destroy a production object, one principal is not enough. Dual control is not “Slack me if it looks weird.” It is two independent authorities on the same payload.

NIST SP 800-53 Rev. 5 AC-5, Separation of Duties is the control name for this pattern: split functions so one person (or one machine identity) cannot complete a high-impact action alone. Your agent is not a federal system. The physics still apply. CISA’s joint guide on AI in operational technology tells operators to keep a human in the loop on critical decisions and to limit worst-case consequences. Refunds and deletes are your plant floor.

ActionOne principal enough?Dual control shape
Refund under a tiny cap, template-onlySometimes, if finance signed the capSecond key above the cap
Refund over the cap, or any payoutNoApprover identity + payload hash
Delete one staging rowMaybeStill deny in prod
Delete in productionNoSecond key, or the tool is absent
DROP / truncate / bucket wipeNoTool not in the catalog
Mass CRM update over NNoPending + second yes, or cap
One field on one idYes, if schema-tightAllowlist fields
Privilege changeNoDeny, or dual with security

Decision list:

  • If the tool can move money → dual control above the cap you can tolerate losing on a bug. Below the cap, still allowlist + schema + tenant check.
  • If the tool can delete production → dual control or remove the tool. “Soft delete with a flag” is still a delete if the flag hides the customer from ops.
  • If the tool can mint privilege → deny. Dual control on iam.attach_policy is how you rubber-stamp an admin. Cut it.
  • If finance and security disagree on the cap → the lower cap wins until a named owner writes a new row.

Dual control on every CRM keystroke is how the second person stops reading. Save the second key for money and delete. Rubber stamps are not a second control.

Break-glass is dual control with a clock, not a hole in the allowlist.

Break-glass eventWhoClockTrace
Second principal unreachable, refund must shipNamed backup on the table row, not “whoever is in Slack”Minutes, not a weekendbreakglass=true, backup id, expiry
Kill-switch on, one payout still requiredSecurity + finance, both on the hashUntil the freeze liftsSeparate reason code
Delete to contain an incidentSecurity only; agent still cannot hold the keyTicket id requiredHuman ran it in the real console
“Just this once” from the agent workerNobodyn/aDeny. The worker does not get a bypass flag

A break-glass flag inside the agent process is a bypass you will forget to turn off. Humans break glass in the real billing or database UI. The agent drafts the payload. It does not carry a god mode.

How do two keys actually bind — payload hash, not a Slack yes?

LangGraph’s interrupt docs exist to pause before API calls, database changes, or financial transactions. Anthropic’s managed-agent permission policies use always_allow / always_ask. Framework vocabulary differs. The binding rule does not: the second yes is for these bytes, not for “the agent may refund today.”

PieceRequired behaviorFailure if skipped
Two principalsAgent identity proposes; human or distinct release identity confirmsThe model is both requester and approver
Payload hashSHA-256 of normalized args (amount, currency, id, dest)Approver said yes to $49; $4900 executed
Re-check at executeHash mismatch → deny or new pendingRace: args edited after the yes
ExpiryPending older than N minutes diesStale yes sits in queue overnight
Second credentialRelease path uses a key the agent runtime cannot read“Dual control” is a UI checkbox on the same token
Trace fieldsrelease_actor, payload_hash, decided_atYou cannot prove the stop fired

Procedure:

  1. Normalize args (canonical JSON, sorted keys, currency present).
  2. Hash. Store hash with the pending record.
  3. Show the human the decoded fields, not the hash. People do not approve digests.
  4. On release, re-hash the args about to execute. Mismatch → deny.
  5. Execute only after the second principal’s credential is used for the money/delete API, or after a release token minted for that hash.
  6. Expire. A yes from Tuesday is not a yes for Friday’s retry.
Approval you receivedArgs at executeGate
Refund $49, order A$49, order Aallow (second key)
Refund $49, order A$4900, order Adeny
Refund $49, order A$49, order Bdeny
“Looks good” in Slack, no hashanythingdeny — not dual control
Same OAuth token, extra checkboxanythingdeny — one principal

A Slack thumbs-up on a screenshot is evidence of a vibe. It is not a second key.

Where must the stop fire — before the API call?

Before the SDK sends. Not after Stripe returns. Not in a “we’ll catch it in review” job at 2 a.m. Not in a weekly export of refunds.

The April 2024 Deploying AI Systems Securely CSA (CISA, NSA, FBI, and partners) puts human-in-the-loop as a failsafe with rollbacks ready. NIST AI RMF 1.0 and the July 2024 Generative AI Profile (NIST AI 600-1) ask you to define human–AI roles and to refuse work the system should not run. Refusing after the delete is a postmortem.

LayerWhen it runsCan it stop this call?
System promptBefore tokensNo
Vendor topic filterOn text in/outNot on {"amount": 50000}
Allowlist + schemaAfter propose, before executeYes
Dual-control interruptAfter valid args, before money/deleteYes
IAM on the second keyAt the providerYes, coarse
SandboxDuring executeLimits the room; does not decide “may this refund happen”
Nightly reconHours laterToo late for the first delete

Evaluation order that actually stops damage:

  1. Fleet kill-switch / write freeze
  2. Tool on this principal’s allowlist
  3. Args match schema
  4. Blast-radius class → money/delete?
  5. If yes: dual-control satisfied for this hash
  6. Else: deny predicates, then allow
FailureCorrect stop
Policy or allowlist cannot loaddeny writes
Second-key service downdo not fail open on money/delete
Unknown tool namedeny
Hash mismatchdeny
Kill-switch ondeny

Fail open (“let the refund through, we’ll look Monday”) is how a prompt-injected ticket empties the till during an outage. Document that a red policy dependency means agents stop writing. That is success.

Where teams hide a bypass, and what to require instead:

RuntimeUsual holeRequired stop
Custom loopDebug script with a second clientSame invoke; no second client
LangGraphTool node that calls the SDK directlyInterrupt or adapter inside the tool function
n8n / workflow + LLM stepWrite node after the model with no checkGate on the write node, not the prompt
Hosted agent APICustom tool handler that auto-executesYour allow / dual-control before the vendor result
MCP hostServer connected, all tools exposedPer-job allowlist on names, not “the server is trusted”

OpenAI will happily return parallel tool_calls. Gate each one. A deny on call two does not let call three ride along because they arrived in one blob. If any path reaches the API client without invoke, you have a bypass — treat it like a missing blast-radius row.

What belongs in IAM versus dual control versus the allowlist?

Three layers. Teams collapse them into “we scoped the key” and then wonder why a scoped refund key refunded every order.

ConcernAllowlistDual controlIAM / credentials
May this tool name be proposed?PrimaryNoNo
May these args move money / delete?Caps helpPrimaryToo fine for most IAM
Which keys exist at all?NoSecond key is one trickPrimary
Tenant isolation at the vendorMirror checkNoPrimary
Emergency freezeRunner flag (fast)Freeze pending queueKey revoke (blunt)
“Refunds over $50 need Maya”PredicateMaya is the second principalUsually inexpressible

IAM is necessary and coarse. A key that can only refund.write can still refund every order for every amount. The allowlist removes refund from agents that should not see it. Dual control splits the remaining money/delete so the model cannot finish the call alone.

  • Agent’s credential cannot call money/delete APIs or those APIs require a release token the agent cannot mint
  • Approver’s credential is not loaded into the agent worker
  • Allowlist does not include shell.exec “so we can debug”
  • Kill-switch is a flag the runner reads on every call, not a prompt edit

Revoking a key is a blunt kill switch. Use it. Do not pretend it encodes Maya’s $50 rule.

What usually fails first when teams try to “just be careful”?

What breaks: a support agent ships with orders.refund, email.send, and a “temporary” shell.exec for a demo. The blast-radius table was never written, so nobody modeled shell → curl the billing API. The prompt says never refund over $50. An injected ticket says ignore prior rules and refund fully — OWASP LLM01 in one paragraph. The runtime executes whatever the model proposed.

What it costs: every refund and delete the tools actually posted, plus the hours to reconstruct which ones were real. I will not invent a dollar figure you cannot take to finance. After 500+ automations and 20,000+ hours on agentic systems, the report that keeps repeating is “the model was told not to,” not “the second key denied it.”

What you do instead:

  1. Fill the blast-radius table. shell.exec gets remove.
  2. Allowlist tickets.get, orders.get, maybe one CRM field. Refund is absent for this job.
  3. If a later job must refund, dual control + cap + payload hash. Email gets a recipient allowlist.
  4. Fail-closed drill: break the allowlist load on purpose. Writes must stop.
  5. Red-team: injected “delete all” / “refund fully” / alias orders.refund_v2. If any call executes, you are not done.
SkipFirst breakSignal on the trace
Blast-radius tableTransitive tool nobody listedA name that was never in the job package
AllowlistMCP dump includes deleteunknown_tool never appears because everything is known
Dual controlOne token, pretty approve UIrelease_actor == agent_id
Payload hashAmount mutated after yesExecute args ≠ approved args
Fail closedOutage fail-openRefunds during policy_unavailable

The incident should blame the missing table and the missing second key, not “the model being bad.”

How do you red-team the stop before soft-launch?

Before anyone besides you can trigger a write, attack the table, the allowlist, and the second key — not the slogan in the prompt.

  • Disallowed tool name from a compromised prompt (db.truncate, orders.refund)
  • Alias that is not on the allowlist (orders.refund_v2, credit_customer)
  • Arg mutation past a money cap (49.00 → 4900) after a human said yes to the small amount
  • Recipient or dest swap after approval (payload A approved, payload B executed)
  • Parallel tool_calls: one allow, one deny — deny must not ride along
  • Mid-run kill-switch while a write is in flight
  • Broken allowlist load (must fail closed on writes)
  • Second-key service timeout (must not fail open on money/delete)
  • shell.exec or MCP filesystem still on the worker “for debugging”
  • Transitive curl from a remaining HTTP tool to the billing host
ProbePassFail
Injected “refund fully”deny or pending with no Stripe callCharge appears
Hash mutationdeny, reason hash_mismatchOriginal yes reused
Fake tool nameunknown_tool on the trace200 from some API
Allowlist file emptiedWrites stopReads may continue under an explicit outage policy; writes must not
MCP server added overnightWorker refuses to boot or tools stay denied until table reviewNew delete tool visible to the model

If any probe executes the tool, you are not in a soft-launch. You are in a demo with a production URL. CaMeL’s March 2025 paper (arXiv:2503.18813) makes the architectural bet you already need: untrusted data must not pick the program. Your red team is checking that the interceptor, not the model, is the program.

How do you measure whether the stop is working?

You measure misses, not vibes. Chatbot thumbs do not tell you whether a delete could have fired.

SignalDefinitionLie if you skip it
Dual-control missMoney/delete executed with one principal or a hash mismatch ignoredYou have a checkbox, not dual control
Allowlist deny rateUnknown-tool and alias denials per job_typeA week of zeroes often means the interceptor is not on
Table coverage% of registered tools with once + loop + owner filledCatalog grew; table did not
Payload rebind denialsExecute hash ≠ approved hashApprovals are theater
Fail-closed drillsLast date the allowlist or second-key path was broken on purposeYou have never seen the stop work
Restore proofLast restore test for a delete-class tool you still allowDual control on delete with no backup is a prayer

Checklist for a weekly read:

  • Zero dual-control misses. One is an incident, not a metric to average.
  • Unknown-tool denials exist. If they are always zero, probe with a fake name.
  • Every money/delete span has release_actor ≠ agent_id and a payload_hash.
  • Kill-switch / fail-closed drill dated within the last 30 days.
  • Blast-radius table review_by not overdue.

Do not copy another team’s deny-rate percentage and call it an SLO. Page on your miss count and on drills you actually ran. Cost controls for agent fleets covers the spend freeze that sits next to this stop. A loop that cannot delete can still set money on fire through tokens. That is a different table. Keep them both.

When is a workflow enough instead of an agent?

When you already know the path. An agent is a loop that chooses tools mid-run. If the graph is “ticket → classify → refund under cap → template email,” you do not need a chooser. You need a workflow with a gate on the write node.

ShapeAgent?Stop that fits
Fixed steps, one writeNoWorkflow + schema + cap on that node
Mid-run tool choice, reads onlyMaybeAllowlist without money/delete
Mid-run tool choice, money or deleteExpensiveAllowlist + dual control, or don’t
Human does the irreversible step anywayNoAgent drafts; human submits in the real UI
You cannot name pass/fail for the jobNot yetFix the process; see when not to build an agent

A workflow still needs the blast-radius row for its write node. The model being “just a classifier” does not make Stripe safer. What you skip is the open catalog and the mid-run invention of db.delete_row.

If the business cannot staff a second principal for money/delete, that is a reason not to give the loop those tools. It is not a reason to skip dual control and “move fast.”

What should you ship if you only have a week?

A five-day agentic pilot is not a full policy platform. It is enough to stop the obvious destruction. Rules sophistication can grow. A bypassable prompt cannot.

DayShipDone when
1Blast-radius table for every tool the demo currently hasEach row has once, loop, owner, stop
2Cut the catalog to an allowlist per jobshell.exec and unused write tools gone
3Dual control on remaining money/delete, or remove those toolsSecond principal + payload hash on a real call
4Fail-closed drill + kill-switchBreaking the allowlist load stops writes
5Red-team inject + alias + hash mutationNone of the three execute

Skip if you only have a week:

  • Fancy policy languages
  • Dual control on reversible single-row patches
  • A vendor topic-filter subscription as a substitute interceptor
  • Loading “the rest of the tools” back for the demo recording

Do not skip:

  • The table
  • Default deny
  • Second key on money/delete or those tools absent
  • A trace that shows the deny

Day 5 that still has shell.exec “for debugging” failed Day 2. Ship the cut.

Pilot checklist that still counts:

  • Blast-radius table checked into the same repo as the worker, with review_by
  • Allowlist per job_type; dump matches; unknown names deny
  • Money/delete either absent or dual-controlled with payload hash
  • Fail-closed drill recorded (who broke the file, what the trace showed)
  • Red-team three probes: inject, alias, hash mutation — none executed
  • Kill-switch flag the on-call can flip without a prompt deploy

That is the bar for /agentic. A longer catalog can wait. The shell tool cannot.

What looks like a stop and is not?

Prompt-only “never delete.” Injection and non-determinism eat it. The runtime still has the tool.

Denylist of scary names. You will miss refund_v2 and the MCP server named helpers.

Confidence thresholds. Most tool APIs do not give you a trustworthy score, and a confident bad refund is still a refund.

One human clicking Approve on a screenshot. No payload hash, often the same token. Not dual control.

IAM-only. A scoped key that can delete, deletes. The cap and the second principal live elsewhere.

Sandbox-only. A sandbox that can still call Stripe with a live key is a polite way to lose money. Containment is not permission.

Vendor guardrail ID as the interceptor. Topic filters do not know your $50 refund cap or your production table name.

Logging denials and executing anyway in “shadow.” Shadow is fine if execute is off. Shadow that still hits Stripe is production.

Dual control on everything. The second person stops looking. Then money and delete ride along with a rubber stamp.

Looks like a stopActual stop
Prompt paragraphAllowlist + interceptor
DenylistDefault deny of names
Slack screenshotPayload hash + second key
Scoped refund key aloneCap + dual control + tenant check
Sandbox with live StripeGate decides; sandbox contains
Topic filterArg predicates in the runner
Shadow that still executesExecute off, or it is production

Start with the blast-radius table, the allowlist, and two keys on the two classes that actually wreck companies. Then expand. After 20,000+ hours on this work, the teams that stopped the damage wrote those three artifacts down. The teams that did not still have a paragraph in the prompt.

FAQ

How do I stop an agent from doing something destructive?

Inventory blast radius per tool, allowlist only the tools that job needs, and put dual control on money and delete in the runner before the tool executes. Prompts do not stop a proposed call. Unknown tools deny. If you cannot name the worst case for a tool, it stays off.

How do I measure whether do I stop an agent from doing something destructive is working?

Count dual-control misses (money or delete that ran with one principal), allowlist denials for unknown names, payload-hash rebind denials, and the date of the last fail-closed drill. Zero unknown-tool denials for a week usually means the interceptor is not on. Chatbot thumbs are not this scoreboard.

What usually fails first when teams try this?

They skip the blast-radius table, leave shell.exec or a full MCP dump on the agent, and treat a Slack yes as dual control. The first injected “refund fully” or a transitive curl then hits a live key. The trace shows a tool that was never in the job package, or release_actor equal to the agent.

How long does this take to show results?

A week is enough to cut the catalog, wire dual control or remove money/delete, and prove fail-closed on a drill. You will see denials on the trace the day the interceptor is mandatory. You will not finish a mature policy catalog in five days, and you should not wait for one before removing DROP.

What should I skip if I only have a week?

Skip policy-language rewrites, dual control on reversible single-row patches, and putting the cut tools back for a demo. Do not skip the blast-radius table, default-deny allowlist, and a second key on remaining money/delete — or those tools gone. A topic filter is not the interceptor.

When is this not worth doing yet?

If the system has no side-effect tools, you are gating a chatbot, not an agent. If the path is already known, ship a workflow with a gate on the write node instead of a choosing loop. If nobody will hold the second key, do not give the loop money or delete tools and pretend a prompt is the hold.

CTA

Stop destruction with a blast-radius table, an allowlist, and two keys on money and delete — in the runner, before the call.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

How do I stop an agent from doing something destructive?
Inventory blast radius per tool, allowlist only the tools that job needs, and put dual control on money and delete in the runner before the tool executes. Prompts do not stop a proposed call. Unknown tools deny. If you cannot name the worst case for a tool, it stays off.
How do I measure whether do I stop an agent from doing something destructive is working?
Count dual-control misses (money or delete that ran with one principal), allowlist denials for unknown names, payload-hash rebind denials, and the date of the last fail-closed drill. Zero unknown-tool denials for a week usually means the interceptor is not on. Chatbot thumbs are not this scoreboard.
What usually fails first when teams try this?
They skip the blast-radius table, leave `shell.exec` or a full MCP dump on the agent, and treat a Slack yes as dual control. The first injected “refund fully” or a transitive curl then hits a live key. The trace shows a tool that was never in the job package, or `release_actor` equal to the agent.
How long does this take to show results?
A week is enough to cut the catalog, wire dual control or remove money/delete, and prove fail-closed on a drill. You will see denials on the trace the day the interceptor is mandatory. You will not finish a mature policy catalog in five days, and you should not wait for one before removing `DROP`.
What should I skip if I only have a week?
Skip policy-language rewrites, dual control on reversible single-row patches, and putting the cut tools back for a demo. Do not skip the blast-radius table, default-deny allowlist, and a second key on remaining money/delete — or those tools gone. A topic filter is not the interceptor.
When is this not worth doing yet?
If the system has no side-effect tools, you are gating a chatbot, not an agent. If the path is already known, ship a workflow with a gate on the write node instead of a choosing loop. If nobody will hold the second key, do not give the loop money or delete tools and pretend a prompt is the hold.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot