Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
William Spurlock Founder — Spurlock Studios 31 MIN
A named human is accountable when an agent refunds, emails, or writes — not the model, not “the agent,” and not the intern who clicked approve without opening the payload. The agent is a tool with a blast radius: money that left, a message that hit an inbox, a record other systems will treat as true. Put a policy gate in code before those tools fire, bind the act to an owner id, and keep a reverse path. This is operating practice for production agents, not legal advice. Counsel owns liability questions. You own the runtime.
The map lives in the Agentic Systems Operating Manual. This spoke owns the accountability layer: who is on the hook, how wide the blast can go, and which gates must fire before an irreversible tool runs.
The short answer
- Name one human owner per tool class before the first production credential is minted. Job title plus a paging channel. “The AI team” is not a name.
- Treat refund, external email, and production write as three different blast radii. Caps, gates, and reverse paths differ. Do not copy one policy onto all three.
- A pre-execution policy gate sees the concrete tool and hashed arguments and returns
allow,deny, orpending-approvalbefore the tool runs. If policy is down, fail closed. - An approval is bound to one argument snapshot. A changed payload is a new act. Rubber-stamp yes is how you get a named owner and still no accountability.
- Measure whether this is working with unowned-run count, pending age, rubber-stamp rate, post-act rewrite rate, and time-to-reverse. A green evaluator on the draft is not ownership.
Who is accountable when an agent acts?
The deployer who put the tool in the agent’s hands, plus the named owner of that tool class, plus whoever approved a pending-approval payload if one was required. The model is not a legal person. It cannot take a page, reverse a Stripe capture, or sit in the customer call.
NIST’s AI Risk Management Framework 1.0 puts it in the trustworthy-AI list: accountable and transparent. Accountability, in that document, presupposes transparency — you can see who did what. The AI RMF Core makes the operating translation: Govern 2 wants roles, responsibilities, and lines of communication documented and clear; Govern 3.2 wants policies that differentiate human-AI configurations; Map 3.5 wants human-oversight processes defined, assessed, and documented. None of that names your billing lead. You still have to write the row.
This post will not tell you who a court would hold liable in your jurisdiction. That is counsel’s job. The operating answer you can ship this week is: every side-effecting tool has a human owner, a gate, a cap, and a reverse path, and the trace can prove all four.
| Party | What they own | What they do not own |
|---|---|---|
| Named tool-class owner | Policy for that class, paging, reverse SOP | Model weights, vendor uptime |
| Approver (when pending) | The specific hashed payload they signed | Later args the agent swapped in |
| Runtime / platform owner | Gate code, fail-closed, credentials, traces | The business rule for “should we refund” |
| Model vendor | The API contract they published | Your tool allowlist and your customer |
| The agent / worker | Nothing. It is software. | Anything you wish you could blame it for |
If you cannot fill the first two rows with real names, do not mint a write credential.
Who is the named human owner?
One person, reachable, who can stop the class and reverse an act. Not a Slack channel with twelve lurkers. Not “shared between CX and finance.” Shared ownership is how Friday night becomes a finger-pointing thread while the refund is already gone.
Write the owner map before the job contract. The owner does not have to execute every approval. They own the policy, the paging roster, and the reverse SOP. If they go on leave, a named deputy is in the same row — not “whoever is on Slack.”
- List every side-effecting tool the agent may call. Reads do not get an owner row unless a read can dump PII to a third party.
- Group tools into classes: refund / billing, external email, production write, internal draft, shell. Keep the list short.
- Put one human and one deputy on each class. Record the paging channel they actually answer.
- Bind
owner_idonto every run at intake. No owner, noactstate. - Tell the owner, in a meeting with a calendar invite, that they will be paged. A wiki row they never saw is theater.
| Tool class | Named owner | Deputy | Paging | They can halt |
|---|---|---|---|---|
| Refund / billing change | Billing or finance lead | Controller or founder | Phone + on-call | Refund tool + payout tool |
| External email / SMS | CX or comms lead | Founder | Phone + on-call | Send tools, recipient allowlist |
| Production CRM / ticket write | CRM or ops owner | Engineering lead | On-call | Write tools, field allowlist |
| Internal draft only | Same as the production owner | Same | Slack is enough | Promotion from draft → live |
| Shell / code exec | Engineering lead | Founder | Phone | Runner, network egress |
If two classes share a credential, they share a blast radius. Split the credential or admit they are one class.
- Every write tool maps to exactly one
owner_id - Deputy is named, not “the team”
- Owner has been told they will be paged, and has reversed a staging act once
- Runs without
owner_idcannot enteract - Offboarding the owner rotates the row the same day
RACI for the act — not for the slide deck:
| Role | Refund | External email | Production write |
|---|---|---|---|
| Responsible (approves when pending) | Billing roster on-call | CX roster on-call | CRM / ops roster |
| Accountable (one neck) | Billing or finance lead | CX or comms lead | CRM or ops owner |
| Consulted | Founder; counsel if the class is new | Founder; counsel on promises | Engineering (schema, tenants) |
| Informed | CX when money is customer-facing | Billing if the body mentions money | Owners of systems that read the record |
Responsible can rotate. Accountable cannot. If the accountable person cannot name the last three production acts in their class without asking Slack, the map is theater.
A title in a deck is not an owner. A person who has reversed a fake refund in staging is.
What is the blast radius for refunds, emails, and writes?
Blast radius is what is already true in the world after the tool returns ok. Not what the prompt hoped would happen. Not what you can “usually” undo.
Refunds move money. Emails hit inboxes you do not control. Writes become facts that billing, shipping, and the next agent will read as ground truth. Those three fail differently, reverse differently, and deserve different default gates. Copy-pasting “human in the loop” onto all three is how you pending-approve an internal draft for six minutes and auto-send a customer email because the queue felt slow.
OWASP’s GenAI LLM Top 10 2026 (published 4 August 2026) treats Excessive Agency as a first-class risk: damaging actions from unexpected, ambiguous, or manipulated model output, via too much functionality, too much permission, or too much autonomy. Your blast-radius table is how you refuse that excess on purpose.
| Action | What is already true after ok | Who feels it | Default gate (pilot) | Cap to bind | Reverse path |
|---|---|---|---|---|---|
| Refund / capture / payout | Money moved or a dispute clock started | Customer, finance, card network | pending-approval at $0.01 | Per-refund and per-day USD | Credit memo + payment-log row; do not “refund the refund” from the agent |
| External email / SMS | Message is in someone else’s inbox | Customer, brand, sometimes legal | pending-approval for any external recipient | Recipients per run, domains, no BCC surprises | Recall if the vendor has it; otherwise a human-sent correction SOP |
| Production write (CRM, ticket, inventory, file) | Other systems will read this as true | Ops, the next automation, the next agent | allow only to a draft field; pending-approval to live | Rows, fields, tenants | Compensating write + audit row; never silent overwrite |
| Internal draft | Staff might see it | Internal only | allow with tenant check | Field allowlist | Delete or edit; still log |
| Shell / arbitrary code | Whatever the command did | Infra, data, secrets | deny unless the job contract names it | Network egress, cwd, timeout | Restore from backup; this is why it stays denied |
Three questions decide the row:
- Can a stranger feel this in under a minute?
- Can you reverse it from your systems alone?
- Will another machine treat the result as fact?
If (1) is yes or (2) is no, the default is pending-approval. If (3) is yes, the live write is a second act, not a side effect of the draft.
Classifying a new tool before it ships:
- Write the one-line blast: “After
ok, ____ is already true.” - Put it in the table above. If it does not fit, it is a new class — new owner row, new credential.
- Name the reverse SOP and run it once in staging. If you cannot, the default stays
denyor pending with no autonomy clock. - Split the credential if this tool can also do a hotter class (a “notes” integration that can send is an email tool).
- Version the row as
policy_version. Quiet staging is not a reason to widen the class.
Caps belong next to the owner, not in the prompt. Dollar, recipient, and row caps are the same family of control as the spend caps in cost controls for agent fleets. A refund agent with no per-day USD cap is a blank check with extra tokens.
Why do policy gates sit before irreversible tools?
Because a system prompt that says “do not refund” is a suggestion, and the tool does not read suggestions. The gate is a function in your runtime that sees the tool name and the arguments and returns a decision before the sandbox executes. If that function times out, the answer is deny, not retry-as-allow.
Anthropic’s computer-use guidance is blunt for a vendor doc: isolate the environment and put a human confirm on consequential actions. Their building effective agents note is the same posture one layer up — start simple, add loops only when a cheaper pattern fails. You do not need their product to take the point. Irreversible work gets a gate outside the model.
The state machine for the loop is the cage: side effects only in act, and only after the gate. plan, evaluate, and revise stay read-only. A worker that “helpfully” writes during evaluate is not being agentic. It is skipping the owner.
| Decision | When | What must be in the trace | Legal next state |
|---|---|---|---|
allow | Payload matches policy, under cap, tenant matches | Tool, args hash, rule id, owner_id | Execute once |
deny | Unknown tool, over cap, wrong tenant, banned recipient | Reason code, no side effect | escalate or abort |
pending-approval | Irreversible class, novel shape, over a threshold | Queue id, proposed payload, owner_id | Wait; no execute |
| Fail closed | Policy timeout, bad payload, unknown tool | Abort reason, no retry-as-allow | abort |
Procedure for every irreversible tool:
- Model proposes
tool+args. - Runtime canonicalizes and hashes
args. - Policy returns
allow/deny/pending-approval. - On
allow, execute once. Same(tool, args_hash)cannot fire twice in the run. - On
pending-approval, park the hashed payload. Human sees the payload, not a summary the worker wrote. - Approval signs that hash. If args change, the prior yes is void.
- On policy outage,
abort. The worker does not get a bypass flag.
- Unknown tools cannot execute
- Write tools require a tenant id that matches the run
- Dollar / recipient / row caps live in code
- Policy outage cannot be skipped by the worker
- Deny and pending leave a reason code an operator can filter
If the trace cannot prove the gate fired, you do not have a gate. You have a story about a gate.
How should refunds be gated?
Default: the refund tool is pending-approval at a cent. Autonomy is earned, not assumed, and week one of a pilot does not earn it. Finance already knows this. Engineering forgets it because the sandbox refund “worked.”
A refund is not a chat reply with a dollar sign. It is a money movement plus a customer belief plus, often, a card-network clock. The named owner is billing or finance, not the person who prompted the agent. The agent may propose a refund. It does not issue one until the owner’s policy says the hash may fire.
| Refund shape | Pilot default | Cap | Owner action |
|---|---|---|---|
| Full refund, original tender, under a tiny test amount in staging | allow in staging only | Staging merchant account | Confirm the test capture reversed |
| Partial refund on a live charge | pending-approval | Per-refund and per-day USD | Approve the hash; check order id |
| Refund to a different destination | deny | n/a | Human process outside the agent |
Duplicate refund on the same charge_id | deny | Idempotency key required | Investigate; do not retry from the worker |
| Chargeback-adjacent or already disputed | deny | n/a | Finance SOP, not the agent |
| Subscription cancel + refund bundle | Split into two tools | Each tool gated | Do not hide a cancel inside a refund payload |
Idempotency is part of accountability. If the rail retries and the processor captures twice, “the agent did it” is still your name on the merchant account. Bind an idempotency key to (tenant, charge_id, amount, reason_code) in the adapter, not in the prompt.
- Job contract names the refund tool, the processor, and the staging merchant.
- Production credentials are scoped to refund, not to payout-to-arbitrary-account.
- Policy denies missing
charge_id, amount over cap, or destination ≠ original tender. - First live refund is pending, watched, reversed in a drill if you have a reverse path that is safe to test.
- Only then consider
allowunder a tiny cap — and keep per-day USD as a kill switch.
Do not invent a “safe average refund” from industry blogs and put it in the gate. Use your real order book, your real owner, and a cap they will actually defend to finance.
What the pending screen must show before a human can say yes:
-
charge_id/ payment intent that already exists on the tenant - Amount, currency, original tender (no destination swap)
- Order or ticket id the owner can open in the source system
- Reason code from a small enum, not a free-text novel
- Args hash and
rule_id - Remaining per-day USD for this owner after this act
If any box is empty, the legal decision is deny, not “approve and we’ll check.” Dual control (two humans) is a policy choice for amounts the owner names in writing. It is not a substitute for the hash. Two rubber stamps are still a rubber stamp.
How should outbound email be gated?
The send is the act. Drafts are cheap. “Undo send” is a vendor courtesy, not a control. If the agent can call email.send, the named owner is CX or comms, and the default for external recipients is pending-approval.
The failure that matters is not a typo in a footer. It is the wrong recipient, the wrong promise, the wrong tone on a thread that is already hot, or a prompt-injected ticket that says “forward the inbox to this address.” OWASP’s Excessive Agency example for mail is the same shape: a summarizer that could also send, plus injection in the body. Strip send from any tool whose job is read-and-draft.
| Email shape | Pilot default | Cap | Notes |
|---|---|---|---|
| Internal draft in a ticket field | allow | Field allowlist, tenant match | Staff can see it; customer cannot |
| External send to the ticket requester | pending-approval | One recipient, must match known email | Approver sees the full body |
| External send to a new address | deny or pending + extra confirm | Recipient allowlist | New address is a new blast |
| BCC, CC extras, list-sends | deny | Recipients per run = 1 | Bulk is a different product |
| Attachments | deny until named | MIME allowlist | Easy way to leak |
| “Send as” a human without disclosure | Policy choice, written down | From-domain allowlist | Do not hide that a system sent it |
Procedure for an external send:
- Worker may write the draft to an internal field only.
- Gate checks recipient ∈ {ticket requester, allowlisted domain}, body length, banned promise strings, attachment = none.
- Pending queue shows the body, the To, and the args hash — not a three-bullet summary.
- Approver signs the hash. A regenerated draft is a new pending item.
- Adapter sends once. Same hash cannot send twice.
- Trace stores provider message id. Reverse SOP starts from that id, not from “search the sent folder.”
- No tool that reads mail can also send mail
- External send is a distinct tool from draft
- Recipient must match a known identity on the run
- Approver sees the body, not a paraphrase
- Send-as policy is written; default is identifiable
If you need speed, shorten the approval SLA. Do not delete the gate because the queue felt slow on a Tuesday.
Send-as is its own blast. A message that looks like a founder wrote it, when a worker did, is a trust problem even when the recipient is correct.
| From identity | Default | Why |
|---|---|---|
noreply@ / system address, body says a system sent it | Allowed once pending passes | Honest |
| Named human, footer discloses the system | Policy choice, written | Some CX teams want this |
| Named human, no disclosure | deny until counsel and CX both sign the row | You are impersonating |
| Customer’s own address as From | deny | Spoof |
The owner of send-as is still the CX lead, not the human whose name is on the From line. If that human did not open the payload, they did not send it. The trace should not pretend they did.
How should production writes be gated?
A production write is an email to every future system. Billing will read the CRM. Shipping will read the ticket. The next agent will retrieve the note and treat it as policy. Draft fields exist so the worker can be wrong in a place that does not become true yet.
Default for a pilot: allow into a clearly named draft field or a staging object. pending-approval (or a second, narrower tool) to promote live. Overwrite of money, status, or identity fields stays pending longer than a comment.
| Write shape | Pilot default | Cap | Reverse |
|---|---|---|---|
| Comment / note on the same record | allow after tenant check | Record must already be in scope | Edit or tombstone the note |
| Draft field the customer never sees | allow | Field allowlist | Clear the field |
| Status change (open → closed, paid, shipped) | pending-approval | Enum allowlist | Compensating status + audit |
| Identity / destination / amount fields | pending-approval | Per-field allowlist | Compensating write; page owner |
| Create a new record | pending-approval | One create per run | Archive; do not delete if other systems copied it |
| Bulk / fan-out writes | deny | Rows per run = 1 | You will not reverse fifty quietly |
Idempotency again. A retried act that creates a second ticket is an unowned duplicate, not a retry. The state machine should refuse a second write on the same (tool, args_hash) for the run.
- Scope the record at intake. The agent cannot search-then-write an arbitrary id.
- Split tools:
record.draft_writevsrecord.promote_live. - Promote copies hashed draft → live fields. If the draft changed after approval, void.
- Live writes emit an audit row: who, hash, rule id, before/after on the fields that moved.
- Compensating writes are a human or a dedicated reverse tool with its own gate — not the worker “fixing” in the same loop.
- Customer-visible fields are not on the draft tool
- Tenant id on the write matches the run
- Row cap is 1 until the owner raises it in writing
- Before/after is stored for identity and money fields
- Duplicate creates are denied by idempotency key
A CRM that “usually” lets you undo is not a reverse path. A reverse path is a SOP you have run in staging with the named owner on the call.
How do you measure whether accountability is working?
You measure whether a human can still be found after an act, whether approvals are real, and whether you can reverse. Pass rate on the draft will stay green while ownership rots.
OpenAI’s agent evals guidance is useful here even if you never buy their dashboard: grade the trace — tool calls, guardrails, handoffs — not only the final blob. Accountability is a trace property. If owner_id, gate decision, and args hash are missing, the run failed the accountability eval regardless of how pretty the email was.
| Metric | What it catches | Veto / page if |
|---|---|---|
| Unowned-run count | Intake without owner_id, tools that bypass the map | Any production run with a side effect and no owner |
| Gate-miss count | Tool ok with no allow/pending record | Any |
| Rubber-stamp rate | Approve without viewing the hashed payload (no open event, sub-second yes) | Rate leaves the band you set with the owner |
| Pending age | Queue nobody is staffing | SLA the owner agreed to, usually hours not weeks |
| Post-act rewrite / apology rate | The act was wrong in the world | Climbing while draft pass rate holds |
| Time-to-reverse | Reverse SOP is fiction | You cannot reverse a staging drill |
| Deny-without-reason | Operators cannot filter or learn | Any |
| Duplicate side effects | Missing idempotency | Any second ok on the same hash |
Procedure for a weekly ownership review (thirty minutes, named owner in the room):
- Sample ten production side effects. For each, open the trace. Name the owner, the gate decision, the hash, the approver if any.
- Count unowned, gate-miss, and duplicate
ok. - Open five approvals. If two were sub-second with no payload view, the gate is a decoration.
- Pick one reverse drill for the class that moved the most money or mail that week. Run it in staging. Time it.
- Write one policy change or one cap change. Version the rule id. Do not “tweak the prompt instead.”
- Traces have
owner_id,rule_id,args_hash,decision - Approvals have
approver_idandpayload_opened_at - Weekly sample is on the calendar, not “when we have time”
- Staging reverse drill has a recorded duration
- Prompt edits are not accepted as a substitute for a rule-id change
If you only chart draft pass rate, you will ship a polite agent that nobody owns.
Failure mode: the agent did it
The failure is not a wrong token. The failure is a side effect in the world and nobody who can reverse it is holding the trace.
Picture the shape — not a named client, not a fabricated loss figure: a refund tool with a production key, a prompt that says “be careful,” no owner_id on the run, no gate, a Slack thread that starts with “the agent did it,” and finance asking which charge_id moved. The model is not going to join that thread. The intern who approved a pending item in a sub-second click without opening the body is not going to have a better answer than you.
What breaks:
- Money: duplicate refunds, refunds to the wrong tender, refunds on already-disputed charges.
- Inbox: the wrong customer, a promise you cannot keep, a leak via a tool that could send as well as read.
- Records: a status flip that shipping trusted, a note the next agent retrieved as policy, a second ticket from a retry.
What it costs: the reverse work, the customer conversation, the freeze on the whole agent program, and the discovery that your “human in the loop” never saw the payload. I will not put a dollar average on that. Your finance lead already has the number they care about — the cap they will sign.
What you do instead:
- Kill the tool credential. Not the Slack bot. The credential.
- Freeze
actfor that class. The kill switch and the policy deny can be the same button. - Export every
okfor that tool since last known-good. You need charge_ids, message ids, record ids — not chat transcripts. - Named owner runs reverse SOP on each, or marks unreversible and starts the human process.
- Do not turn the tool back on until the owner map, gate, caps, idempotency, and a staging reverse drill exist.
| Symptom | Likely hole | First move |
|---|---|---|
| “The agent refunded twice” | No idempotency, retry in act | Freeze refund tool; reconcile charge_ids |
| “It emailed the wrong person” | Send tool + searchable arbitrary To | Freeze send; recipient allowlist |
| “CRM is wrong and shipping went” | Live write without promote step | Freeze live write; compensating status |
| “Someone approved it” | Rubber-stamp, hash not bound | Void approvals that lack payload_opened_at |
| “We thought it was staging” | Shared credential, no tenant check | Rotate keys; split env |
Bravery is not a reverse strategy. A named owner with a drill is.
What should you ship in one week?
You do not ship a fleet. You ship an owner map, three gated tool classes, and a staging drill that has actually reversed one fake act. Multi-agent handoffs, long-term memory, and a pretty ops wall can wait.
NIST Map 3.5 is the week-one spirit: define, assess, and document human oversight. You can do that for one job without a governance committee. You cannot do it for twelve tools and no names.
- Pick one sentence-sized job that might refund, email, or write — not all three if you can avoid it.
- Fill the owner table for the classes that job needs. Get the humans in a room. Record the deputy.
- Split tools: draft vs send, draft vs live write, propose-refund vs execute-refund.
- Implement the gate:
allow/deny/pending-approval, fail closed, args hash. - Cap at something finance and CX will say out loud. $0.01 pending for live refunds is a valid week-one cap.
- Bind
owner_idat intake. No owner, noact. - Run two staging drills: one deny (unknown tool or over-cap) and one reverse after a pending approve.
- Put the weekly ten-trace sample on a calendar.
Skip on purpose:
| Skip this week | Why | Do not skip |
|---|---|---|
| Second agent, “reviewer” agent | A second worker is not an owner | Named human + deputy |
| Memory stores | Contaminated context is a later fire | Trace fields on this run |
| Fancy dashboard | A CSV of ten traces is enough | The ten-trace review existing |
| Raising caps because staging was quiet | Quiet is not earned autonomy | Pending on refund and external send |
| Sharing prod credentials with staging | That is how env mix-ups happen | Split keys, tenant checks |
Wire the pending queue on the rail you already trust. In n8n that is a wait / approval node on a durable execution, not a Slack emoji the worker interprets as yes. The worker proposes. The rail holds the hashed payload. The named roster clicks in a UI that shows the payload. The adapter fires once after the hash matches.
| Rail piece | Job this week | Failure if skipped |
|---|---|---|
| Durable execution | Pending survives deploys | Approvals vanish; people bypass |
Approval wait bound to args_hash | Yes is for this payload | Regenerated args ride an old yes |
| Error path on timeout | Expire pending → abort | Silent execute after the owner went home |
| Staging merchant / inbox | Reverse drill is real | You only tested JSON |
If the week ends and a production send or refund can still fire without a gate record, you did not ship accountability. You shipped a demo with a wiki.
When is a workflow enough instead of an agent?
When the path is known. Accountability gets cheaper when there is no tool-choice step. A refund that is always “if status=X and reason=Y and amount≤Z, then this Stripe call” is an automation with an approval node. Calling it an agent does not add judgement. It adds a model that can pick the wrong charge_id.
The manual’s split still holds. Default to automation when the path is known. Build an agent only when the path varies and you can still write pass/fail criteria. Accountability does not disappear in the automation — the owner and the cap still exist — but you stop pretending a worker is making a decision you already encoded.
| You have… | Call it… | Accountability shape |
|---|---|---|
| Fixed path, rare judgement, known charge_id | Automation | Approval node + idempotency + owner on the workflow |
| Varying path, crisp criteria (which ticket, which template) | Agentic system | Gate + owner map + evaluator + state machine |
| Varying path, mushy taste (“make the customer happy”) | Do not auto-send | Human writes; tools stay draft-only |
| High stakes, weak reverse | Human + checklist | Do not mint the credential |
Decision list:
- Can you write the next tool name without a model? → Workflow.
- Can you name pass/fail for the artifact? If no → not an agent yet. Criteria first.
- Does the step need to choose among records? → Agent maybe, but scope the record set at intake.
- Is the act irreversible and reverse-weak? → Human, or pending with a drill, or do not build it.
n8n is a natural rail for both shapes: durable execution, an approval wait, retries that you actually bound. The agent lives only in the steps where choice is required. The rail owns delivery, idempotency, and the pending queue. If your “agent” is a single prompt that always calls the same three APIs and always returns green, call it an automation and put the owner on the workflow. Language matters because risk reviews change when you admit you shipped non-deterministic software.
What must the trace prove after a side effect?
After ok, an operator who was not in the room must be able to answer: who owned this, what fired, who said yes, what payload, and how we reverse. If any of those are missing, the run is an unowned act even if the customer outcome happened to be fine.
OpenTelemetry-style traces are enough. A chat log is not. The evaluator may score the draft; it does not get to mark done on a write it never saw a gate for.
| Field | Why it exists | Fail the run if missing after a write |
|---|---|---|
run_id / tenant_id | Scope | Yes |
owner_id / deputy_id | Accountability | Yes |
state | Whether writes were legal | Yes if state ≠ act |
tool / args_hash | What was attempted | Yes |
rule_id / decision | Gate actually fired | Yes |
approver_id / payload_opened_at | Pending was real | Yes when decision was pending |
idempotency_key | Retries will not double-fire | Yes on money and send |
provider_object_id | Reverse starts here | Yes |
evaluator_verdict | Draft quality, not ownership | No — but done requires it |
Checklist before you call a class “in production”:
- An operator can pull one refund, one send, and one live write from last week and fill every required field above without asking Slack
- A changed payload after approval cannot execute on the old yes
- Policy deny in staging is visible as
denyin the trace, not as a model apology - Reverse SOP names the
provider_object_idfield it needs - The weekly sample uses this schema, not screenshots of the chat UI
If you cannot produce that row, you are not ready to argue with a customer, a processor, or your own finance lead. You are ready to keep the tool on deny.
Operator lookup when Slack is already on fire:
- Filter traces by
tooland time window. Do not grep chat logs first. - For each
ok, copyprovider_object_id,owner_id,args_hash,decision. - If
decisionis missing, treat the act as unowned. Freeze the class. - Hand the list to the named owner. They run reverse SOP. Engineering does not freelance refunds in the processor dashboard “to be helpful.”
- Write the incident note against
run_id, not against a model vendor’s name.
How do you approve without rubber-stamping?
A pending queue that nobody reads is worse than deny. It prints a name on an act the human never saw. Accountability requires that the yes is about this payload.
Bind every approval to four facts: the args_hash, the rule_id, the approver_id, and a payload_opened_at that is not the same millisecond as approved_at. If the UI can approve from a list of subject lines, you built a rubber-stamp machine.
| Approval design | What the human sees | What they sign | Use when |
|---|---|---|---|
| Hash-bound payload | Full args, not a summary | args_hash | Default for refund, send, live write |
| Dual control | Same payload, two humans | Two hashes, two ids | Owner-named USD or recipient thresholds |
| Expire unused pending | Queue item + TTL | Nothing; run abort | Owner went home |
| Replay old yes | Hidden | A stale hash | Never |
Procedure:
- Queue stores canonical args + hash. The worker cannot edit the parked payload in place.
- UI renders the payload. Opening it writes
payload_opened_at. - Approve is disabled until opened. Sub-second approve is a bug, not a speed win.
- Approve writes
approver_idand re-hashes. Mismatch → void. - Adapter executes only if live args hash equals approved hash.
- Timeout →
abort, page the owner, do not execute.
- Approve disabled until payload open
- Changed draft creates a new queue item
- Approver cannot be the same process that proposed (the worker never signs)
- Dual control, if you use it, is two humans — not a human and a second model
- Expired pending cannot be “nudged through” by engineering
A second model that says “looks fine” is not dual control. It is a second worker. The named human still has not seen the charge_id.
FAQ
Who is accountable when an agent acts (refunds, emails, writes)?
A named human who owns that tool class, plus whoever approved a pending payload if the gate required it. The model is not a person and cannot take the page or reverse the charge. The runtime owner is accountable for the gate, the caps, and the trace. This is how you operate the system. It is not a legal opinion.
How do I measure whether accountability for agent actions is working?
Count unowned runs, gate-misses, rubber-stamp approvals, post-act rewrites, and time-to-reverse on a staging drill. Sample ten live side effects a week and require owner_id, rule_id, args_hash, and a real payload-open on pending items. Draft pass rate will not tell you this.
What usually fails first when teams try this?
They skip the owner map and put “human in the loop” on a queue nobody staffs, or they let one credential send, refund, and write. Second: approvals that sign a summary instead of a hashed payload. Third: production keys in the same sandbox that was supposed to be staging.
How long does this take to show results?
The owner map and pending-on-refund-and-send can exist in days, not a quarter. Proof is a staging deny plus a reverse drill with the named owner on the call. Autonomy on live refunds should take longer on purpose. If “results” means fewer Slack incidents, you will see that when gate-miss count hits zero, not when the demo email looks nice.
What should I skip if I only have a week?
Skip extra agents, memory stores, and a polished dashboard. Do not skip: named owner plus deputy, split draft/send and draft/live tools, fail-closed policy, args-hash approvals, caps, and one reverse drill. A week that only writes principles is a week you already had.
When is this not worth doing yet?
When the system cannot refund, email, or write — keep it as a draft generator and do not mint those credentials. It is also not worth a fleet if nobody will take the page or you have no reverse path. The moment a production key can move money or hit an inbox, the owner map and the gate are already due.
CTA
Name the owner. Gate the tool. Then scale the job.
Read the agentic systems manual, then use agentic or book an agentic pilot.
What questions does this article answer?
- Who is accountable when an agent acts (refunds, emails, writes)?
- A named human who owns that tool class, plus whoever approved a pending payload if the gate required it. The model is not a person and cannot take the page or reverse the charge. The runtime owner is accountable for the gate, the caps, and the trace. This is how you operate the system. It is not a legal opinion.
- How do I measure whether accountability for agent actions is working?
- Count unowned runs, gate-misses, rubber-stamp approvals, post-act rewrites, and time-to-reverse on a staging drill. Sample ten live side effects a week and require `owner_id`, `rule_id`, `args_hash`, and a real payload-open on pending items. Draft pass rate will not tell you this.
- What usually fails first when teams try this?
- They skip the owner map and put “human in the loop” on a queue nobody staffs, or they let one credential send, refund, and write. Second: approvals that sign a summary instead of a hashed payload. Third: production keys in the same sandbox that was supposed to be staging.
- How long does this take to show results?
- The owner map and pending-on-refund-and-send can exist in days, not a quarter. Proof is a staging deny plus a reverse drill with the named owner on the call. Autonomy on live refunds should take longer on purpose. If “results” means fewer Slack incidents, you will see that when gate-miss count hits zero, not when the demo email looks nice.
- What should I skip if I only have a week?
- Skip extra agents, memory stores, and a polished dashboard. Do not skip: named owner plus deputy, split draft/send and draft/live tools, fail-closed policy, args-hash approvals, caps, and one reverse drill. A week that only writes principles is a week you already had.
- When is this not worth doing yet?
- When the system cannot refund, email, or write — keep it as a draft generator and do not mint those credentials. It is also not worth a fleet if nobody will take the page or you have no reverse path. The moment a production key can move money or hit an inbox, the owner map and the gate are already due.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
AI Agents Agentic Systems: An Operating Manual for Multi-Agent Work That Ships
An agentic system is evaluators, policy gates, sandboxes, and kill switches — not a chat window — so production tool work survives contact with real data.
AI Agents Why Agent Demos Die in Production: Control Gaps, Not Model IQ
Demo success proves a staged happy path. Production fails when the control loop — schemas, auth, evaluators, and kill switches — was never in the harness.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.