Spurlock Studios
Contact
Share LinkedIn X
A violet ring. Thesis: WHO ACCOUNTABLE AGENT ACTS REFUNDS.

A named human is accountable when an agent refunds, emails, or writes — not the model, not “the agent,” and not the intern who clicked approve without opening the payload. The agent is a tool with a blast radius: money that left, a message that hit an inbox, a record other systems will treat as true. Put a policy gate in code before those tools fire, bind the act to an owner id, and keep a reverse path. This is operating practice for production agents, not legal advice. Counsel owns liability questions. You own the runtime.

The map lives in the Agentic Systems Operating Manual. This spoke owns the accountability layer: who is on the hook, how wide the blast can go, and which gates must fire before an irreversible tool runs.

The short answer

  • Name one human owner per tool class before the first production credential is minted. Job title plus a paging channel. “The AI team” is not a name.
  • Treat refund, external email, and production write as three different blast radii. Caps, gates, and reverse paths differ. Do not copy one policy onto all three.
  • A pre-execution policy gate sees the concrete tool and hashed arguments and returns allow, deny, or pending-approval before the tool runs. If policy is down, fail closed.
  • An approval is bound to one argument snapshot. A changed payload is a new act. Rubber-stamp yes is how you get a named owner and still no accountability.
  • Measure whether this is working with unowned-run count, pending age, rubber-stamp rate, post-act rewrite rate, and time-to-reverse. A green evaluator on the draft is not ownership.

Who is accountable when an agent acts?

The deployer who put the tool in the agent’s hands, plus the named owner of that tool class, plus whoever approved a pending-approval payload if one was required. The model is not a legal person. It cannot take a page, reverse a Stripe capture, or sit in the customer call.

NIST’s AI Risk Management Framework 1.0 puts it in the trustworthy-AI list: accountable and transparent. Accountability, in that document, presupposes transparency — you can see who did what. The AI RMF Core makes the operating translation: Govern 2 wants roles, responsibilities, and lines of communication documented and clear; Govern 3.2 wants policies that differentiate human-AI configurations; Map 3.5 wants human-oversight processes defined, assessed, and documented. None of that names your billing lead. You still have to write the row.

This post will not tell you who a court would hold liable in your jurisdiction. That is counsel’s job. The operating answer you can ship this week is: every side-effecting tool has a human owner, a gate, a cap, and a reverse path, and the trace can prove all four.

PartyWhat they ownWhat they do not own
Named tool-class ownerPolicy for that class, paging, reverse SOPModel weights, vendor uptime
Approver (when pending)The specific hashed payload they signedLater args the agent swapped in
Runtime / platform ownerGate code, fail-closed, credentials, tracesThe business rule for “should we refund”
Model vendorThe API contract they publishedYour tool allowlist and your customer
The agent / workerNothing. It is software.Anything you wish you could blame it for

If you cannot fill the first two rows with real names, do not mint a write credential.

Who is the named human owner?

One person, reachable, who can stop the class and reverse an act. Not a Slack channel with twelve lurkers. Not “shared between CX and finance.” Shared ownership is how Friday night becomes a finger-pointing thread while the refund is already gone.

Write the owner map before the job contract. The owner does not have to execute every approval. They own the policy, the paging roster, and the reverse SOP. If they go on leave, a named deputy is in the same row — not “whoever is on Slack.”

  1. List every side-effecting tool the agent may call. Reads do not get an owner row unless a read can dump PII to a third party.
  2. Group tools into classes: refund / billing, external email, production write, internal draft, shell. Keep the list short.
  3. Put one human and one deputy on each class. Record the paging channel they actually answer.
  4. Bind owner_id onto every run at intake. No owner, no act state.
  5. Tell the owner, in a meeting with a calendar invite, that they will be paged. A wiki row they never saw is theater.
Tool classNamed ownerDeputyPagingThey can halt
Refund / billing changeBilling or finance leadController or founderPhone + on-callRefund tool + payout tool
External email / SMSCX or comms leadFounderPhone + on-callSend tools, recipient allowlist
Production CRM / ticket writeCRM or ops ownerEngineering leadOn-callWrite tools, field allowlist
Internal draft onlySame as the production ownerSameSlack is enoughPromotion from draft → live
Shell / code execEngineering leadFounderPhoneRunner, network egress

If two classes share a credential, they share a blast radius. Split the credential or admit they are one class.

  • Every write tool maps to exactly one owner_id
  • Deputy is named, not “the team”
  • Owner has been told they will be paged, and has reversed a staging act once
  • Runs without owner_id cannot enter act
  • Offboarding the owner rotates the row the same day

RACI for the act — not for the slide deck:

RoleRefundExternal emailProduction write
Responsible (approves when pending)Billing roster on-callCX roster on-callCRM / ops roster
Accountable (one neck)Billing or finance leadCX or comms leadCRM or ops owner
ConsultedFounder; counsel if the class is newFounder; counsel on promisesEngineering (schema, tenants)
InformedCX when money is customer-facingBilling if the body mentions moneyOwners of systems that read the record

Responsible can rotate. Accountable cannot. If the accountable person cannot name the last three production acts in their class without asking Slack, the map is theater.

A title in a deck is not an owner. A person who has reversed a fake refund in staging is.

What is the blast radius for refunds, emails, and writes?

Blast radius is what is already true in the world after the tool returns ok. Not what the prompt hoped would happen. Not what you can “usually” undo.

Refunds move money. Emails hit inboxes you do not control. Writes become facts that billing, shipping, and the next agent will read as ground truth. Those three fail differently, reverse differently, and deserve different default gates. Copy-pasting “human in the loop” onto all three is how you pending-approve an internal draft for six minutes and auto-send a customer email because the queue felt slow.

OWASP’s GenAI LLM Top 10 2026 (published 4 August 2026) treats Excessive Agency as a first-class risk: damaging actions from unexpected, ambiguous, or manipulated model output, via too much functionality, too much permission, or too much autonomy. Your blast-radius table is how you refuse that excess on purpose.

ActionWhat is already true after okWho feels itDefault gate (pilot)Cap to bindReverse path
Refund / capture / payoutMoney moved or a dispute clock startedCustomer, finance, card networkpending-approval at $0.01Per-refund and per-day USDCredit memo + payment-log row; do not “refund the refund” from the agent
External email / SMSMessage is in someone else’s inboxCustomer, brand, sometimes legalpending-approval for any external recipientRecipients per run, domains, no BCC surprisesRecall if the vendor has it; otherwise a human-sent correction SOP
Production write (CRM, ticket, inventory, file)Other systems will read this as trueOps, the next automation, the next agentallow only to a draft field; pending-approval to liveRows, fields, tenantsCompensating write + audit row; never silent overwrite
Internal draftStaff might see itInternal onlyallow with tenant checkField allowlistDelete or edit; still log
Shell / arbitrary codeWhatever the command didInfra, data, secretsdeny unless the job contract names itNetwork egress, cwd, timeoutRestore from backup; this is why it stays denied

Three questions decide the row:

  1. Can a stranger feel this in under a minute?
  2. Can you reverse it from your systems alone?
  3. Will another machine treat the result as fact?

If (1) is yes or (2) is no, the default is pending-approval. If (3) is yes, the live write is a second act, not a side effect of the draft.

Classifying a new tool before it ships:

  1. Write the one-line blast: “After ok, ____ is already true.”
  2. Put it in the table above. If it does not fit, it is a new class — new owner row, new credential.
  3. Name the reverse SOP and run it once in staging. If you cannot, the default stays deny or pending with no autonomy clock.
  4. Split the credential if this tool can also do a hotter class (a “notes” integration that can send is an email tool).
  5. Version the row as policy_version. Quiet staging is not a reason to widen the class.

Caps belong next to the owner, not in the prompt. Dollar, recipient, and row caps are the same family of control as the spend caps in cost controls for agent fleets. A refund agent with no per-day USD cap is a blank check with extra tokens.

Why do policy gates sit before irreversible tools?

Because a system prompt that says “do not refund” is a suggestion, and the tool does not read suggestions. The gate is a function in your runtime that sees the tool name and the arguments and returns a decision before the sandbox executes. If that function times out, the answer is deny, not retry-as-allow.

Anthropic’s computer-use guidance is blunt for a vendor doc: isolate the environment and put a human confirm on consequential actions. Their building effective agents note is the same posture one layer up — start simple, add loops only when a cheaper pattern fails. You do not need their product to take the point. Irreversible work gets a gate outside the model.

The state machine for the loop is the cage: side effects only in act, and only after the gate. plan, evaluate, and revise stay read-only. A worker that “helpfully” writes during evaluate is not being agentic. It is skipping the owner.

DecisionWhenWhat must be in the traceLegal next state
allowPayload matches policy, under cap, tenant matchesTool, args hash, rule id, owner_idExecute once
denyUnknown tool, over cap, wrong tenant, banned recipientReason code, no side effectescalate or abort
pending-approvalIrreversible class, novel shape, over a thresholdQueue id, proposed payload, owner_idWait; no execute
Fail closedPolicy timeout, bad payload, unknown toolAbort reason, no retry-as-allowabort

Procedure for every irreversible tool:

  1. Model proposes tool + args.
  2. Runtime canonicalizes and hashes args.
  3. Policy returns allow / deny / pending-approval.
  4. On allow, execute once. Same (tool, args_hash) cannot fire twice in the run.
  5. On pending-approval, park the hashed payload. Human sees the payload, not a summary the worker wrote.
  6. Approval signs that hash. If args change, the prior yes is void.
  7. On policy outage, abort. The worker does not get a bypass flag.
  • Unknown tools cannot execute
  • Write tools require a tenant id that matches the run
  • Dollar / recipient / row caps live in code
  • Policy outage cannot be skipped by the worker
  • Deny and pending leave a reason code an operator can filter

If the trace cannot prove the gate fired, you do not have a gate. You have a story about a gate.

How should refunds be gated?

Default: the refund tool is pending-approval at a cent. Autonomy is earned, not assumed, and week one of a pilot does not earn it. Finance already knows this. Engineering forgets it because the sandbox refund “worked.”

A refund is not a chat reply with a dollar sign. It is a money movement plus a customer belief plus, often, a card-network clock. The named owner is billing or finance, not the person who prompted the agent. The agent may propose a refund. It does not issue one until the owner’s policy says the hash may fire.

Refund shapePilot defaultCapOwner action
Full refund, original tender, under a tiny test amount in stagingallow in staging onlyStaging merchant accountConfirm the test capture reversed
Partial refund on a live chargepending-approvalPer-refund and per-day USDApprove the hash; check order id
Refund to a different destinationdenyn/aHuman process outside the agent
Duplicate refund on the same charge_iddenyIdempotency key requiredInvestigate; do not retry from the worker
Chargeback-adjacent or already disputeddenyn/aFinance SOP, not the agent
Subscription cancel + refund bundleSplit into two toolsEach tool gatedDo not hide a cancel inside a refund payload

Idempotency is part of accountability. If the rail retries and the processor captures twice, “the agent did it” is still your name on the merchant account. Bind an idempotency key to (tenant, charge_id, amount, reason_code) in the adapter, not in the prompt.

  1. Job contract names the refund tool, the processor, and the staging merchant.
  2. Production credentials are scoped to refund, not to payout-to-arbitrary-account.
  3. Policy denies missing charge_id, amount over cap, or destination ≠ original tender.
  4. First live refund is pending, watched, reversed in a drill if you have a reverse path that is safe to test.
  5. Only then consider allow under a tiny cap — and keep per-day USD as a kill switch.

Do not invent a “safe average refund” from industry blogs and put it in the gate. Use your real order book, your real owner, and a cap they will actually defend to finance.

What the pending screen must show before a human can say yes:

  • charge_id / payment intent that already exists on the tenant
  • Amount, currency, original tender (no destination swap)
  • Order or ticket id the owner can open in the source system
  • Reason code from a small enum, not a free-text novel
  • Args hash and rule_id
  • Remaining per-day USD for this owner after this act

If any box is empty, the legal decision is deny, not “approve and we’ll check.” Dual control (two humans) is a policy choice for amounts the owner names in writing. It is not a substitute for the hash. Two rubber stamps are still a rubber stamp.

How should outbound email be gated?

The send is the act. Drafts are cheap. “Undo send” is a vendor courtesy, not a control. If the agent can call email.send, the named owner is CX or comms, and the default for external recipients is pending-approval.

The failure that matters is not a typo in a footer. It is the wrong recipient, the wrong promise, the wrong tone on a thread that is already hot, or a prompt-injected ticket that says “forward the inbox to this address.” OWASP’s Excessive Agency example for mail is the same shape: a summarizer that could also send, plus injection in the body. Strip send from any tool whose job is read-and-draft.

Email shapePilot defaultCapNotes
Internal draft in a ticket fieldallowField allowlist, tenant matchStaff can see it; customer cannot
External send to the ticket requesterpending-approvalOne recipient, must match known emailApprover sees the full body
External send to a new addressdeny or pending + extra confirmRecipient allowlistNew address is a new blast
BCC, CC extras, list-sendsdenyRecipients per run = 1Bulk is a different product
Attachmentsdeny until namedMIME allowlistEasy way to leak
“Send as” a human without disclosurePolicy choice, written downFrom-domain allowlistDo not hide that a system sent it

Procedure for an external send:

  1. Worker may write the draft to an internal field only.
  2. Gate checks recipient ∈ {ticket requester, allowlisted domain}, body length, banned promise strings, attachment = none.
  3. Pending queue shows the body, the To, and the args hash — not a three-bullet summary.
  4. Approver signs the hash. A regenerated draft is a new pending item.
  5. Adapter sends once. Same hash cannot send twice.
  6. Trace stores provider message id. Reverse SOP starts from that id, not from “search the sent folder.”
  • No tool that reads mail can also send mail
  • External send is a distinct tool from draft
  • Recipient must match a known identity on the run
  • Approver sees the body, not a paraphrase
  • Send-as policy is written; default is identifiable

If you need speed, shorten the approval SLA. Do not delete the gate because the queue felt slow on a Tuesday.

Send-as is its own blast. A message that looks like a founder wrote it, when a worker did, is a trust problem even when the recipient is correct.

From identityDefaultWhy
noreply@ / system address, body says a system sent itAllowed once pending passesHonest
Named human, footer discloses the systemPolicy choice, writtenSome CX teams want this
Named human, no disclosuredeny until counsel and CX both sign the rowYou are impersonating
Customer’s own address as FromdenySpoof

The owner of send-as is still the CX lead, not the human whose name is on the From line. If that human did not open the payload, they did not send it. The trace should not pretend they did.

How should production writes be gated?

A production write is an email to every future system. Billing will read the CRM. Shipping will read the ticket. The next agent will retrieve the note and treat it as policy. Draft fields exist so the worker can be wrong in a place that does not become true yet.

Default for a pilot: allow into a clearly named draft field or a staging object. pending-approval (or a second, narrower tool) to promote live. Overwrite of money, status, or identity fields stays pending longer than a comment.

Write shapePilot defaultCapReverse
Comment / note on the same recordallow after tenant checkRecord must already be in scopeEdit or tombstone the note
Draft field the customer never seesallowField allowlistClear the field
Status change (open → closed, paid, shipped)pending-approvalEnum allowlistCompensating status + audit
Identity / destination / amount fieldspending-approvalPer-field allowlistCompensating write; page owner
Create a new recordpending-approvalOne create per runArchive; do not delete if other systems copied it
Bulk / fan-out writesdenyRows per run = 1You will not reverse fifty quietly

Idempotency again. A retried act that creates a second ticket is an unowned duplicate, not a retry. The state machine should refuse a second write on the same (tool, args_hash) for the run.

  1. Scope the record at intake. The agent cannot search-then-write an arbitrary id.
  2. Split tools: record.draft_write vs record.promote_live.
  3. Promote copies hashed draft → live fields. If the draft changed after approval, void.
  4. Live writes emit an audit row: who, hash, rule id, before/after on the fields that moved.
  5. Compensating writes are a human or a dedicated reverse tool with its own gate — not the worker “fixing” in the same loop.
  • Customer-visible fields are not on the draft tool
  • Tenant id on the write matches the run
  • Row cap is 1 until the owner raises it in writing
  • Before/after is stored for identity and money fields
  • Duplicate creates are denied by idempotency key

A CRM that “usually” lets you undo is not a reverse path. A reverse path is a SOP you have run in staging with the named owner on the call.

How do you measure whether accountability is working?

You measure whether a human can still be found after an act, whether approvals are real, and whether you can reverse. Pass rate on the draft will stay green while ownership rots.

OpenAI’s agent evals guidance is useful here even if you never buy their dashboard: grade the trace — tool calls, guardrails, handoffs — not only the final blob. Accountability is a trace property. If owner_id, gate decision, and args hash are missing, the run failed the accountability eval regardless of how pretty the email was.

MetricWhat it catchesVeto / page if
Unowned-run countIntake without owner_id, tools that bypass the mapAny production run with a side effect and no owner
Gate-miss countTool ok with no allow/pending recordAny
Rubber-stamp rateApprove without viewing the hashed payload (no open event, sub-second yes)Rate leaves the band you set with the owner
Pending ageQueue nobody is staffingSLA the owner agreed to, usually hours not weeks
Post-act rewrite / apology rateThe act was wrong in the worldClimbing while draft pass rate holds
Time-to-reverseReverse SOP is fictionYou cannot reverse a staging drill
Deny-without-reasonOperators cannot filter or learnAny
Duplicate side effectsMissing idempotencyAny second ok on the same hash

Procedure for a weekly ownership review (thirty minutes, named owner in the room):

  1. Sample ten production side effects. For each, open the trace. Name the owner, the gate decision, the hash, the approver if any.
  2. Count unowned, gate-miss, and duplicate ok.
  3. Open five approvals. If two were sub-second with no payload view, the gate is a decoration.
  4. Pick one reverse drill for the class that moved the most money or mail that week. Run it in staging. Time it.
  5. Write one policy change or one cap change. Version the rule id. Do not “tweak the prompt instead.”
  • Traces have owner_id, rule_id, args_hash, decision
  • Approvals have approver_id and payload_opened_at
  • Weekly sample is on the calendar, not “when we have time”
  • Staging reverse drill has a recorded duration
  • Prompt edits are not accepted as a substitute for a rule-id change

If you only chart draft pass rate, you will ship a polite agent that nobody owns.

Failure mode: the agent did it

The failure is not a wrong token. The failure is a side effect in the world and nobody who can reverse it is holding the trace.

Picture the shape — not a named client, not a fabricated loss figure: a refund tool with a production key, a prompt that says “be careful,” no owner_id on the run, no gate, a Slack thread that starts with “the agent did it,” and finance asking which charge_id moved. The model is not going to join that thread. The intern who approved a pending item in a sub-second click without opening the body is not going to have a better answer than you.

What breaks:

  • Money: duplicate refunds, refunds to the wrong tender, refunds on already-disputed charges.
  • Inbox: the wrong customer, a promise you cannot keep, a leak via a tool that could send as well as read.
  • Records: a status flip that shipping trusted, a note the next agent retrieved as policy, a second ticket from a retry.

What it costs: the reverse work, the customer conversation, the freeze on the whole agent program, and the discovery that your “human in the loop” never saw the payload. I will not put a dollar average on that. Your finance lead already has the number they care about — the cap they will sign.

What you do instead:

  1. Kill the tool credential. Not the Slack bot. The credential.
  2. Freeze act for that class. The kill switch and the policy deny can be the same button.
  3. Export every ok for that tool since last known-good. You need charge_ids, message ids, record ids — not chat transcripts.
  4. Named owner runs reverse SOP on each, or marks unreversible and starts the human process.
  5. Do not turn the tool back on until the owner map, gate, caps, idempotency, and a staging reverse drill exist.
SymptomLikely holeFirst move
“The agent refunded twice”No idempotency, retry in actFreeze refund tool; reconcile charge_ids
“It emailed the wrong person”Send tool + searchable arbitrary ToFreeze send; recipient allowlist
“CRM is wrong and shipping went”Live write without promote stepFreeze live write; compensating status
“Someone approved it”Rubber-stamp, hash not boundVoid approvals that lack payload_opened_at
“We thought it was staging”Shared credential, no tenant checkRotate keys; split env

Bravery is not a reverse strategy. A named owner with a drill is.

What should you ship in one week?

You do not ship a fleet. You ship an owner map, three gated tool classes, and a staging drill that has actually reversed one fake act. Multi-agent handoffs, long-term memory, and a pretty ops wall can wait.

NIST Map 3.5 is the week-one spirit: define, assess, and document human oversight. You can do that for one job without a governance committee. You cannot do it for twelve tools and no names.

  1. Pick one sentence-sized job that might refund, email, or write — not all three if you can avoid it.
  2. Fill the owner table for the classes that job needs. Get the humans in a room. Record the deputy.
  3. Split tools: draft vs send, draft vs live write, propose-refund vs execute-refund.
  4. Implement the gate: allow / deny / pending-approval, fail closed, args hash.
  5. Cap at something finance and CX will say out loud. $0.01 pending for live refunds is a valid week-one cap.
  6. Bind owner_id at intake. No owner, no act.
  7. Run two staging drills: one deny (unknown tool or over-cap) and one reverse after a pending approve.
  8. Put the weekly ten-trace sample on a calendar.

Skip on purpose:

Skip this weekWhyDo not skip
Second agent, “reviewer” agentA second worker is not an ownerNamed human + deputy
Memory storesContaminated context is a later fireTrace fields on this run
Fancy dashboardA CSV of ten traces is enoughThe ten-trace review existing
Raising caps because staging was quietQuiet is not earned autonomyPending on refund and external send
Sharing prod credentials with stagingThat is how env mix-ups happenSplit keys, tenant checks

Wire the pending queue on the rail you already trust. In n8n that is a wait / approval node on a durable execution, not a Slack emoji the worker interprets as yes. The worker proposes. The rail holds the hashed payload. The named roster clicks in a UI that shows the payload. The adapter fires once after the hash matches.

Rail pieceJob this weekFailure if skipped
Durable executionPending survives deploysApprovals vanish; people bypass
Approval wait bound to args_hashYes is for this payloadRegenerated args ride an old yes
Error path on timeoutExpire pending → abortSilent execute after the owner went home
Staging merchant / inboxReverse drill is realYou only tested JSON

If the week ends and a production send or refund can still fire without a gate record, you did not ship accountability. You shipped a demo with a wiki.

When is a workflow enough instead of an agent?

When the path is known. Accountability gets cheaper when there is no tool-choice step. A refund that is always “if status=X and reason=Y and amount≤Z, then this Stripe call” is an automation with an approval node. Calling it an agent does not add judgement. It adds a model that can pick the wrong charge_id.

The manual’s split still holds. Default to automation when the path is known. Build an agent only when the path varies and you can still write pass/fail criteria. Accountability does not disappear in the automation — the owner and the cap still exist — but you stop pretending a worker is making a decision you already encoded.

You have…Call it…Accountability shape
Fixed path, rare judgement, known charge_idAutomationApproval node + idempotency + owner on the workflow
Varying path, crisp criteria (which ticket, which template)Agentic systemGate + owner map + evaluator + state machine
Varying path, mushy taste (“make the customer happy”)Do not auto-sendHuman writes; tools stay draft-only
High stakes, weak reverseHuman + checklistDo not mint the credential

Decision list:

  1. Can you write the next tool name without a model? → Workflow.
  2. Can you name pass/fail for the artifact? If no → not an agent yet. Criteria first.
  3. Does the step need to choose among records? → Agent maybe, but scope the record set at intake.
  4. Is the act irreversible and reverse-weak? → Human, or pending with a drill, or do not build it.

n8n is a natural rail for both shapes: durable execution, an approval wait, retries that you actually bound. The agent lives only in the steps where choice is required. The rail owns delivery, idempotency, and the pending queue. If your “agent” is a single prompt that always calls the same three APIs and always returns green, call it an automation and put the owner on the workflow. Language matters because risk reviews change when you admit you shipped non-deterministic software.

What must the trace prove after a side effect?

After ok, an operator who was not in the room must be able to answer: who owned this, what fired, who said yes, what payload, and how we reverse. If any of those are missing, the run is an unowned act even if the customer outcome happened to be fine.

OpenTelemetry-style traces are enough. A chat log is not. The evaluator may score the draft; it does not get to mark done on a write it never saw a gate for.

FieldWhy it existsFail the run if missing after a write
run_id / tenant_idScopeYes
owner_id / deputy_idAccountabilityYes
stateWhether writes were legalYes if state ≠ act
tool / args_hashWhat was attemptedYes
rule_id / decisionGate actually firedYes
approver_id / payload_opened_atPending was realYes when decision was pending
idempotency_keyRetries will not double-fireYes on money and send
provider_object_idReverse starts hereYes
evaluator_verdictDraft quality, not ownershipNo — but done requires it

Checklist before you call a class “in production”:

  • An operator can pull one refund, one send, and one live write from last week and fill every required field above without asking Slack
  • A changed payload after approval cannot execute on the old yes
  • Policy deny in staging is visible as deny in the trace, not as a model apology
  • Reverse SOP names the provider_object_id field it needs
  • The weekly sample uses this schema, not screenshots of the chat UI

If you cannot produce that row, you are not ready to argue with a customer, a processor, or your own finance lead. You are ready to keep the tool on deny.

Operator lookup when Slack is already on fire:

  1. Filter traces by tool and time window. Do not grep chat logs first.
  2. For each ok, copy provider_object_id, owner_id, args_hash, decision.
  3. If decision is missing, treat the act as unowned. Freeze the class.
  4. Hand the list to the named owner. They run reverse SOP. Engineering does not freelance refunds in the processor dashboard “to be helpful.”
  5. Write the incident note against run_id, not against a model vendor’s name.

How do you approve without rubber-stamping?

A pending queue that nobody reads is worse than deny. It prints a name on an act the human never saw. Accountability requires that the yes is about this payload.

Bind every approval to four facts: the args_hash, the rule_id, the approver_id, and a payload_opened_at that is not the same millisecond as approved_at. If the UI can approve from a list of subject lines, you built a rubber-stamp machine.

Approval designWhat the human seesWhat they signUse when
Hash-bound payloadFull args, not a summaryargs_hashDefault for refund, send, live write
Dual controlSame payload, two humansTwo hashes, two idsOwner-named USD or recipient thresholds
Expire unused pendingQueue item + TTLNothing; run abortOwner went home
Replay old yesHiddenA stale hashNever

Procedure:

  1. Queue stores canonical args + hash. The worker cannot edit the parked payload in place.
  2. UI renders the payload. Opening it writes payload_opened_at.
  3. Approve is disabled until opened. Sub-second approve is a bug, not a speed win.
  4. Approve writes approver_id and re-hashes. Mismatch → void.
  5. Adapter executes only if live args hash equals approved hash.
  6. Timeout → abort, page the owner, do not execute.
  • Approve disabled until payload open
  • Changed draft creates a new queue item
  • Approver cannot be the same process that proposed (the worker never signs)
  • Dual control, if you use it, is two humans — not a human and a second model
  • Expired pending cannot be “nudged through” by engineering

A second model that says “looks fine” is not dual control. It is a second worker. The named human still has not seen the charge_id.

FAQ

Who is accountable when an agent acts (refunds, emails, writes)?

A named human who owns that tool class, plus whoever approved a pending payload if the gate required it. The model is not a person and cannot take the page or reverse the charge. The runtime owner is accountable for the gate, the caps, and the trace. This is how you operate the system. It is not a legal opinion.

How do I measure whether accountability for agent actions is working?

Count unowned runs, gate-misses, rubber-stamp approvals, post-act rewrites, and time-to-reverse on a staging drill. Sample ten live side effects a week and require owner_id, rule_id, args_hash, and a real payload-open on pending items. Draft pass rate will not tell you this.

What usually fails first when teams try this?

They skip the owner map and put “human in the loop” on a queue nobody staffs, or they let one credential send, refund, and write. Second: approvals that sign a summary instead of a hashed payload. Third: production keys in the same sandbox that was supposed to be staging.

How long does this take to show results?

The owner map and pending-on-refund-and-send can exist in days, not a quarter. Proof is a staging deny plus a reverse drill with the named owner on the call. Autonomy on live refunds should take longer on purpose. If “results” means fewer Slack incidents, you will see that when gate-miss count hits zero, not when the demo email looks nice.

What should I skip if I only have a week?

Skip extra agents, memory stores, and a polished dashboard. Do not skip: named owner plus deputy, split draft/send and draft/live tools, fail-closed policy, args-hash approvals, caps, and one reverse drill. A week that only writes principles is a week you already had.

When is this not worth doing yet?

When the system cannot refund, email, or write — keep it as a draft generator and do not mint those credentials. It is also not worth a fleet if nobody will take the page or you have no reverse path. The moment a production key can move money or hit an inbox, the owner map and the gate are already due.

CTA

Name the owner. Gate the tool. Then scale the job.

Read the agentic systems manual, then use agentic or book an agentic pilot.

FAQ

What questions does this article answer?

Who is accountable when an agent acts (refunds, emails, writes)?
A named human who owns that tool class, plus whoever approved a pending payload if the gate required it. The model is not a person and cannot take the page or reverse the charge. The runtime owner is accountable for the gate, the caps, and the trace. This is how you operate the system. It is not a legal opinion.
How do I measure whether accountability for agent actions is working?
Count unowned runs, gate-misses, rubber-stamp approvals, post-act rewrites, and time-to-reverse on a staging drill. Sample ten live side effects a week and require `owner_id`, `rule_id`, `args_hash`, and a real payload-open on pending items. Draft pass rate will not tell you this.
What usually fails first when teams try this?
They skip the owner map and put “human in the loop” on a queue nobody staffs, or they let one credential send, refund, and write. Second: approvals that sign a summary instead of a hashed payload. Third: production keys in the same sandbox that was supposed to be staging.
How long does this take to show results?
The owner map and pending-on-refund-and-send can exist in days, not a quarter. Proof is a staging deny plus a reverse drill with the named owner on the call. Autonomy on live refunds should take longer on purpose. If “results” means fewer Slack incidents, you will see that when gate-miss count hits zero, not when the demo email looks nice.
What should I skip if I only have a week?
Skip extra agents, memory stores, and a polished dashboard. Do not skip: named owner plus deputy, split draft/send and draft/live tools, fail-closed policy, args-hash approvals, caps, and one reverse drill. A week that only writes principles is a week you already had.
When is this not worth doing yet?
When the system cannot refund, email, or write — keep it as a draft generator and do not mint those credentials. It is also not worth a fleet if nobody will take the page or you have no reverse path. The moment a production key can move money or hit an inbox, the owner map and the gate are already due.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot