When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
William Spurlock Founder — Spurlock Studios 26 MIN
Retry when the failure is transient and the write is idempotent. Escalate when the failure is auth, a policy deny, ambiguous user intent, the same tool failing with the same args, or anything that moves money. The worker does not get a vote. The harness maps a typed error class onto retry, escalate, or abort before the next tool execute fires.
This spoke sits under the Agentic Systems Operating Manual. The cage that makes the branch legal is state machines for agent loops. The judge that must still see a failed write is build the evaluator before the agent. What follows is the split itself: when another attempt is diligence, and when it is how you double-charge a card.
The short answer
- Retry is for blips:
429,503, timeouts, connect resets — on a read, or on a write that already holds a stable idempotency key. - Escalate is for everything a second identical call cannot fix:
401/403, policydeny, two legal customer ids, the same fingerprint after the bound, refunds and payouts with unknown status. - Abort is not escalate. Abort is poison, out of contract, or a kill switch. Do not page a human to finish work you already decided is illegal.
- The model proposes. The runner refuses. A prompt that says “don’t retry auth” is a hint. A pre-execute gate is a control.
- The escalate package is the product when the loop stops: goal, tools tried, error class, last artifact, budget remaining, and the exact question the human must answer.
When should the agent escalate instead of retrying?
Escalate when another attempt cannot change the outcome without a human, a credential, or a different intent. Retry when the outcome can change because the world blinked — and you can prove the next call is the same logical write.
| Signal | Retry? | Escalate? | Why |
|---|---|---|---|
429 / 503 with Retry-After | Yes, capped | After the cap | Transient; HTTP already gave you a clock |
| Timeout / connect reset on a read | Yes, capped | After the cap | No side effect to duplicate |
| Timeout on a write with a stored idempotency key | Yes: reuse the key, then poll | If status stays unknown past TTL | The risk is a second identity, not waiting |
| Timeout on a write with no key | No | Yes | A second POST can double-apply |
401 / 403 / missing scope | No | Yes | Tokens do not heal themselves inside the loop |
Policy deny or pending-approval | No | Yes, or wait on the approval queue | Retry-as-allow is how gates die |
| Ambiguous user intent (0 or N matches) | One clarify ask, then stop | Yes if still ambiguous | Guessing the account is a wrong write |
Same (tool, args_hash, error_class) after the bound | No | Yes | That is no-progress, not persistence |
Money movement (refund, payout, charge, transfer) | Only receipt poll on the same key | Default yes | Irreversible + unknown is a human job |
| Schema / validation fail after one corrected payload | No | Yes | The second “fix” is usually invention |
| Revision ceiling hit, criteria still fail | No | Yes | Grind-to-green is not a retry policy |
| Out of job contract / kill switch | No | No — abort | Do not escalate illegal work |
If you cannot put a row on this table for a tool, that tool is not in the loop. It is a hope with an HTTP client.
Decision list the runner should apply in order:
- Is the error class
auth,policy,ambiguous_intent, orfatal? →escalateorabort. No backoff. - Is the tool in the money class and status
unknownwithout a key? →escalate. - Is this the same fingerprint at or past the bound? →
escalatewithtool_retry_exhausted. - Is it
retryableand idempotent and budget remains? → backoff, then one moreact. - Else →
escalatewithunclassified_fail_closed.
Unknown is not retryable in production. Loosen a class after you have traces, not after a demo felt “almost.”
What makes a retry legal?
A legal retry changes the world’s chance of success without changing the identity of the write. That is a narrower idea than “the model wants to try again.”
| Must be true | If false |
|---|---|
Error class is retryable (429, 503, timeout, some 5xx) | Escalate or abort |
Side-effect class is read, or a ledger already claimed an idempotency key | Escalate on writes |
Fingerprint is new, or this is attempt k of k_max on the same retryable fingerprint | Escalate on no-progress |
Backoff honors Retry-After when present | Immediate re-call is a hammer |
| Per-run tool budget still has room for this call | Escalate budget_exhausted |
Policy gate would still allow the same payload | You are not retrying a deny |
RFC 9110 defines Retry-After as the wait before a follow-up request. MDN’s Retry-After notes the usual pairing: 503 (temporary unavailability) and 429 (rate limit). Those two are retries with a clock. 401, 403, 404, and 422 are not a clock. They are a stop.
Temporal’s own design note is the same split in workflow language: mark a failure non-retryable at the throw site — invalid input, missing records, authorization — because repeating the call will never succeed (non-retryable errors). Your tool adapter needs that flag. An error string the model can reread is not a policy.
Worked 429 on a read tool:
- Adapter returns
rate_limitedwithretry_after_sfrom the header, or a jittered default if the header is missing. - Runner stays in
actbut does not execute until the clock fires. n8n Wait is the rail version of this. - Same fingerprint is allowed only because the class is retryable and attempt
<bound. - Still
429at the bound →escalatetool_retry_exhausted. Do not pivot to a sibling write tool as “another way.” - Next result is
401→ new class. Escalate auth. The 429 budget does not transfer.
Worked 429 on a write tool is the same clock plus the ledger key already claimed. If you cannot show the key id on the span, treat the write as unknown and escalate. A clock without an identity is how a rate limit becomes a duplicate row.
- Tool adapter returns
ok/retryable/fatal/auth/not_found/rate_limited/policy - Writes register an idempotency key before the HTTP client
-
Retry-Afteris copied onto the span, not discarded - Unknown error class defaults to
retryable: false - n8n Retry On Fail (if you use it on the rail) has Max Tries and then Stop — it is not a second unbounded agent loop
n8n’s node Settings give you Retry On Fail for workflow nodes. That is the rail. The agent loop still needs its own bound. Two retry owners without a shared key is how a 503 becomes two CRM rows.
Why does auth escalate on the first failure?
Auth failures do not become valid credentials because the model asked again. A 401, a 403, an expired refresh, or a missing scope is an operator problem. Retrying it is how you lock an account and how you leak a “try a different token” plan into the transcript.
| Auth signal | Legal next state | Illegal next state |
|---|---|---|
401 unauthorized | escalate tool_auth_error | Mint a token from the prompt |
403 forbidden / missing scope | escalate tool_forbidden | Call a sibling write tool “just to see” |
| Refresh token rejected | escalate tool_auth_expired | Loop the refresh grant |
| Tenant id mismatch on the gate | escalate tenant_mismatch | Swap tenant from ticket text |
| Secret missing at startup | abort | Retry the tool hoping env “settles” |
Procedure the adapter should run on any auth-class result:
- Do not execute the same tool again in this run.
- Redact tokens from the observation the model sees. Leave the error class.
- Freeze further write tools for this run.
- Open
escalatewith the tool name, status, scope that failed, and the tenant id the gate saw. - Page the owner of the credential, not the owner of the prompt.
The model will offer to “re-authenticate.” That sentence is not a control. If a human must rotate a secret, the run is already past self-help.
A retry that “works” after auth is usually a race: the operator fixed the secret while the loop was still burning tokens. Your dashboard will call that success. It was an escalate you failed to issue.
Why does a policy deny never retry as allow?
A policy gate already answered. deny means the payload is illegal. pending-approval means a human has to say yes. Re-proposing the same tool with synonyms, shuffled JSON keys, or a “pretty please” system-prompt addendum is an attack on your own gate.
Anthropic’s note on building effective agents is the cheap version of this: start simple, add loops only when a cheaper pattern fails. A deny is not a cheaper pattern failing. It is the pattern working.
| Gate result | Runner action | Model visibility |
|---|---|---|
allow | Execute in the sandbox | Observation from the tool |
deny | No execute; escalate or abort by reason | Reason code, not a puzzle to solve |
pending-approval | Park on the approval queue | “Waiting on operator”; no second propose |
| Policy timeout / unknown tool | Fail closed; abort | No retry-as-allow |
- Unknown tools cannot execute
- Deny leaves a reason code an operator can filter
- The worker cannot transition
deny→acton the same payload -
pending-approvalis not a sleep-and-retry of the tool - Policy outage cannot be bypassed by “the model is sure”
If your traces show deny followed by a successful execute of the same tool in the same run, you do not have a gate. You have a log line.
The only legal “retry” after policy is a new payload that is a different job: a smaller row cap, a draft instead of a send, a read instead of a write. That is a new fingerprint. It still has to pass the gate. It is not a second try at the forbidden call.
Why is ambiguous user intent an escalate?
If the agent cannot name a single in-scope target, it must not pick a favorite. Multi-match CRM accounts, two email addresses in the request, a refund with no invoice id, “the Acme one” when you have Acme Inc and Acme LLC — those are not retrieval problems. They are intent problems. Another search_accounts with the same string is how you write the wrong tenant.
| Intake shape | One retry allowed? | Then |
|---|---|---|
| Exactly one candidate after normalize | No retry needed | act on that id |
| Zero candidates | One revised query or one clarify question | escalate no_match |
| Two or more candidates | One disambiguation ask | escalate ambiguous_intent if still N |
| Required slot missing (amount, id, date) | One ask | escalate missing_slot |
| User said two actions that contradict policy | No | escalate intent_conflict |
Procedure:
- Normalize names, emails, and external ids before search.
- If
candidates.length != 1, do not call a write tool. - Ask once, with the candidate list, if the channel can take a question.
- If the channel is async and nobody answers inside the job’s SLA,
escalatewith the list attached. - Never let the worker “choose the closest embedding.” Closest is not unique.
This is the failure that looks like intelligence. The draft is fluent. The wrong crm_account_id is already in the payload. The evaluator has to reject unsupported identity, not grade the prose. Criteria that say “account id came from a unique match or the run escalated” belong in the golden set — see evaluators before agents.
Guessing under ambiguity is not a retry. It is a write you will spend the afternoon reversing.
When does repeated same-tool fail become escalate?
The first retryable failure may be a blip. The second identical call is a loop. Fingerprint tool + normalized args + error class before execute. If that fingerprint is already at the bound, refuse the call and escalate. Do not wait for the token budget to notice.
| Attempt | Same fingerprint, retryable: true | Same fingerprint, retryable: false |
|---|---|---|
| 1 | Execute (first try) | Do not retry; escalate now |
| 2 | Execute after backoff | Already illegal |
| 3 (bound) | escalate tool_retry_exhausted | — |
| Distinct fingerprint, same tool | Allowed until the per-tool cap | Still no if class is auth/policy |
A legitimate retry changes something material: args, a page cursor, a backoff clock. Equivalent JSON with keys reshuffled is not material. Strip request_id and injected timestamps before you hash, or every retry looks unique and the bound never fires.
- Fingerprint computed before the HTTP client
- Normalize has a unit test per write tool
- Bound is a number in runner config, not a vibe in the prompt
- Empty result
[]on the same query counts as a fingerprint - Failure ledger lives outside the chat so summarization cannot resurrect the call
The state machine spoke is the cage: retryable may return to act once, then escalate. The model does not get skip_eval as a consolation prize.
If you cannot draw “same fingerprint → refuse,” you will call this diligence in the standup and a stampede in the bill.
How should money movement retry?
Money tools default to escalate. The exception is narrow: you already minted a stable idempotency key, the error is transient or indeterminate, and you are retrying the receipt — the same key, or a retrieve-by-metadata — not a second PaymentIntent.
Stripe’s low-level error handling is the cleanest vendor split I will cite here because the headers are explicit:
| Stripe / HTTP signal | Same key? | Agent action |
|---|---|---|
Stripe-Should-Retry: true | Yes | Backoff; retry the same identity |
Stripe-Should-Retry: false | — | Stop; escalate |
| Timeout / no body | Yes | Treat as unknown; retry same key or poll |
500 / 502 / 503 / 504 | Yes | Indeterminate; never a new key |
Content 400 you will fix | New key only after the body changes | One corrected payload, then escalate if it still fails |
401 | After credentials exist | That is auth. Escalate. Do not “retry the charge.” |
429 | Wait; same key if the limiter ran after idempotency | Clock, then bound, then escalate |
Procedure for any refund, payout, charge, or transfer tool:
- Claim a ledger row with a unique idempotency key before the provider call.
- Put your local id in provider metadata so a webhook can finish a timed-out HTTP.
- On timeout: status
unknown. Poll. Do not mint a second key. - On
Should-Retry: falseor a non-retryable 4xx that is not a fixable schema miss:escalatemoney_unknownormoney_rejected. - On success: store the receipt; the evaluator must see it. A worker story about the refund is not a receipt.
If the upstream API has no idempotency header and no client-supplied external id, money movement is escalate-only until you build a local outbox. “The model will be careful” is not an outbox.
Poll loop for unknown money (this is not a second POST):
| Step | Action | Stop if |
|---|---|---|
| Retrieve by key or metadata | Found receipt | Continue to evaluate with the receipt |
| Webhook arrives with your local id | Reconcile the ledger | done path, or escalate on mismatch |
TTL exceeded, still unknown | escalate money_unknown | Human checks the provider dashboard |
| Retrieve says not found and TTL is still short | Stay unknown; do not POST again | Clock not done |
| Operator marks duplicate in the ledger | abort or escalate money_rejected | Never “retry to make it right” |
Do not ask the model whether the refund worked. The model will narrate. The ledger and the provider are the only two oracles.
A retry that opens a new write identity on a charge is not a retry. It is a second payment.
Who decides retry versus escalate?
The runner. Not the worker, not the evaluator’s prose, not a Slack emoji on the trace.
| Decision | Owner | What the model may do |
|---|---|---|
| Propose a tool and args | Worker, inside act | Suggest allowlisted calls |
| Classify the tool result | Tool adapter | Nothing — it receives a typed observation |
| Map class → next state | State machine / harness | Nothing |
| Allow the next execute | Policy gate + fingerprint bound | Propose only |
Declare done | Evaluator + machine | Cannot self-pass a failed write |
Open escalate | Machine, on listed events | Cannot veto |
LangGraph’s Graph API says this with different nouns: nodes do work, edges decide what is next. Your edge table has to include error class, remaining retries, and money/auth/policy flags. If the only edge is “model said continue,” you built a chatbot with side effects.
n8n is a natural rail for the branch you already trust: backoff, durable execution, a dead path that pages. Put the agent in the step that must choose a tool. Put retry-vs-escalate in the graph around it. I have collaborated with the n8n team; the canvas is not the control plane. Max Tries on a node plus an escalate sub-workflow is.
-
(state, event, guards) → next_stateis data, not a paragraph -
escalateandabortare named terminals - The worker cannot increment its own retry counter downward
- Evaluator
failwith ceiling remaining goes torevise, not a raw tool hammer - Evaluator
failat ceiling goes toescalate, notact
Freedom stays inside the state. The split lives on the edge. That is the whole point of the operating manual stack: policy, sandbox, machine, then the model.
What must the escalate package contain?
Escalate is a designed success of a kind: the system stopped while a human could still finish. If the package is “something broke, good luck,” you trained operators to ignore the queue.
| Field | Required | Why |
|---|---|---|
job_contract + hard nos | Yes | The human should not rediscover scope |
reason_code from a catalog | Yes | Filter, alert, golden-set harvest |
| Tools tried: name, fingerprint, error class | Yes | No-progress evidence |
| Last artifact (redacted) | Yes | Resume without re-doing reads |
| Evaluator failures + evidence | Yes | What “done” still lacks |
| Budget remaining / spent | Yes | Whether finishing is even cheap |
| The question the human must answer | Yes | One decision, not a novel |
Candidate list (if ambiguous_intent) | When relevant | Do not hide the N-match |
| Receipt / ledger id (if money) | When relevant | Poll, do not re-charge |
| Trace id | Yes | Ops, not archaeology |
Catalog starters you can copy:
tool_auth_errortool_forbiddenpolicy_denyambiguous_intentmissing_slottool_retry_exhaustedmoney_unknownmoney_rejectedrevision_ceilingbudget_exhaustedunclassified_fail_closed
Minimum package shape (illustrative — your schema can be stricter):
terminal: escalate
reason_code: ambiguous_intent
job: "Update billing email for the account named in the ticket"
hard_nos: ["no refund", "no send to customer"]
question_for_human: "Which crm_account_id is in scope — 441 or 882?"
candidates:
- { id: 441, name: "Acme Inc" }
- { id: 882, name: "Acme LLC" }
tools_tried:
- { tool: search_accounts, fingerprint: "…", error_class: ok, note: "N=2" }
last_artifact: internal_draft.md
evaluator_unmet: ["account id from unique match"]
budget: { spent_usd: null, calls_left: 4 }
trace_id: "run_…"
ledger_key: null
Leave spent_usd null if you do not meter dollars yet. Do not invent a cost. Calls-left and reason code are enough to start.
Procedure when the machine opens escalate:
- Stop all write tools for the run.
- Persist the package before you notify.
- Notify the queue you actually staff — not a channel nobody owns.
- Leave the run resumable from
actonly after a human supplies the missing slot or approval. Do not auto-resume on a timer. - Harvest the case into the golden set the same week, with expected terminal
escalate.
If operators keep finishing jobs from the raw transcript instead of the package, the package is wrong. Fix the schema. Do not add another summary model on top.
What fails first when teams skip this split?
The first failure is almost never “the model is dumb.” It is a retry storm that looks like uptime: green executions, climbing token spend, duplicate CRM writes, and a refund that landed twice because timeout was treated like “try again with a fresh id.”
Concrete shape, no invented client numbers:
- Billing tool times out. Adapter returns
"failed"with no class. - Worker retries with a new client request id.
- Provider had already captured the first POST.
- Evaluator never sees a receipt; worker marks the run done from the second
200. - Finance finds the duplicate on the card. Ops finds a “successful” agent run.
Second shape: ambiguous intent dressed as search.
- Ticket says “update Acme.” Search returns Acme Inc and Acme LLC.
- Worker calls search again with the same string, then picks the first row.
- Write lands on the wrong tenant. Evaluator grades the email tone. Tone passed.
- You learn from the angry reply, not from the trace.
Fix: candidates.length != 1 is illegal for act. The golden case expects escalate ambiguous_intent with both ids in the package. Closest embedding is not unique.
| What broke | What it costs | What you do instead |
|---|---|---|
Untyped "failed" | Retry of auth, 404, and money alike | Typed classes; unknown → fail closed |
| New idempotency key on timeout | Duplicate charge / duplicate ticket | Ledger first; reuse key; poll |
| Prompt-only “don’t loop” | Returns after the next helpfulness edit | Pre-execute fingerprint bound |
| Continue-on-error on the rail | Green run, Error Workflow never fires | Stop or branch; do not mute writes |
| Self-pass after a failed write | Silent wrong done | Evaluator sees tool outcome |
| Escalate as a Slack dump | Humans ignore the queue | Catalog reason + one question |
| Punishing escalate rate in reviews | Hidden retries to look autonomous | Treat escalate as a designed terminal |
The culture failure is real even when the code is not. If leadership rewards “the bot handled it” and punishes pages, the harness will grow illegal retries. Healthy ops treat escalate like a caught exception: cheaper than the duplicate write.
Do not invent a “retry success rate” to paper over this. Count duplicates, unknown-money rows, and same-fingerprint refusals. Those are receipts. A percentage without a denominator is a slide.
How do I evaluate this in production?
You evaluate the branch, not the vibes. A high pass rate with hidden retries is grind. A high escalate rate can be the system working — if the reasons are ambiguous_intent and policy_deny, not unclassified_fail_closed on every 503 you forgot to mark retryable.
| Signal | What it tells you | Veto / inspect if… |
|---|---|---|
Escalate rate by reason_code | Mix of designed stops vs leaks | unclassified_fail_closed dominates |
| Mean retry depth on retryable classes | Whether backoff is doing work | Depth sits at the cap on every run |
| Same-fingerprint refuse count | No-progress detector is alive | Count is zero while 5xx are not |
Money unknown rows open past TTL | Ledger / poll is broken | Any row older than your poll window |
| Duplicate provider ids per local key | Second write identity leaked | Count is not zero |
| Online vs golden escalate mix | Production drift | New codes with no fixture |
| Human finish time from package | Package quality | Operators still open raw traces first |
| Pass rate beside the above | Quality, not a cover story | Pass holds while duplicates rise |
Golden-set the split. For each fixture, expected_terminal is done, escalate, or abort, and the reason code is part of the assertion. Auth fixtures must not done. Multi-match fixtures must not pick. Timeout-on-charge fixtures must not mint a second key.
| Fixture | Input trap | Expected terminal | Fail if |
|---|---|---|---|
auth-401-write | Write tool returns 401 | escalate tool_auth_error | Any second execute of that tool |
policy-deny-refund | Gate denies issue_refund | escalate policy_deny or abort | Refund tool HTTP happens |
ambiguous-acme | Two CRM ids for one name | escalate ambiguous_intent | Write on either id |
retry-429-then-ok | First read 429, second 200 | done if criteria pass | Immediate re-call with no wait |
timeout-charge-same-key | Charge HTTP times out | Poll / retry same key, or escalate money_unknown | New key appears in the ledger |
same-fingerprint-3x | Identical args, retryable 5xx | escalate tool_retry_exhausted at bound | Attempt 4 of the same hash |
schema-fail-twice | 422, model “fixes,” 422 again | escalate | Third invented payload |
OpenAI’s agent evals treat tool choice and trajectory as first-class grades. Use that idea even if you never touch their dashboard: a run that “passes” after three identical refund calls is a trajectory fail.
Cadence that does not need a fabricated SLA:
- Unit: adapter classification + fingerprint normalize, every commit.
- Task: golden set including escalate fixtures, every behavior change.
- Online: sample production terminals daily or weekly; alert on new reason codes.
- Silent-fail: once a week, sample “done” runs that finance or support later rewrote.
If you cannot name expected_terminal for a case, you cannot tell retry from escalate. Write the fixture before you widen the tool.
What guardrails do I need?
You need the same stack the operating manual already named — but aimed at this branch. Without them, “retry vs escalate” is a slide in a design doc.
| Guardrail | Job on this split |
|---|---|
| Typed tool errors | Give the machine a class, not a paragraph |
| Pre-execute policy gate | deny cannot become allow by persistence |
| Fingerprint + bound | Same call dies on purpose |
| Idempotency ledger | Writes can wait without cloning |
| Money tool allowlist | Default pending or deny in week one |
| Revision ceiling | Eval fail does not become infinite act |
| Per-run budget | 429 storms still halt |
| Escalate package schema | Humans finish; they do not excavate |
| Kill switch | Abort illegal work; do not escalate it |
| Trace fields | reason_code, retry depth, key id, gate decision |
Week-one checklist — skip the fleet, keep the brake:
- Every write tool has an error class enum the runner understands
- Auth and policy classes have zero retry budget
- Money tools are
pending-approvalor escalate-on-unknown - Multi-match intake escalates; it does not “pick closest”
- Same-fingerprint bound is 2 for retryable, 0 for the rest
- n8n Retry On Fail (rail) is bounded and does not Continue the leftovers into a charge
- Golden set has at least one fixture per escalate reason you claim to support
- You have tripped
escalatein staging and opened the package as an operator
If a box is empty, do not add another model. Fill the box.
When is a workflow enough instead of an agent?
When the retry/escalate split is the whole job. Known path, structured input, rare judgement: n8n (or any workflow engine) already has retries, Wait, and a dead-letter. Buying an agent so a model can “decide” to retry a 429 is how you pay token tax for a backoff you already owned.
| Signal | Build this | Retry/escalate lives where |
|---|---|---|
| Fixed steps, typed I/O, rare exceptions | Workflow | Node Retry On Fail + Error Workflow + DLQ |
| Path varies, criteria exist, tools are many | Agentic loop in a machine | Adapter classes + edge table |
| Path varies, criteria are mush | Workshop, not a build | Nowhere — do not automate the argument |
| One irreversible write, no key | Human + checklist | Escalate-only; no agent “careful POST” |
| Volume is tiny | Human + template | Page, don’t loop |
| You only wanted a chat UI | Wrong lane | — |
Default to the workflow when you can write the branch in a table without a model. Default to the agent when tool choice is the uncertainty and the table still holds for errors. The brake pedal in the parent manual is the same: no evaluator, no agent.
A hybrid that ships: workflow owns delivery, idempotency, and the escalate queue; a bounded agent owns the step where the tool is not known in advance. The agent returns done / escalate / abort. The workflow never lets the model own HTTP retry.
- I can draw the retry/escalate table without naming a model
- 429/503 already have a Wait or Retry On Fail on the rail
- 401 already pages, it does not loop
- Money writes already have a key or they are human-only
- The remaining uncertainty is which allowlisted tool, not whether to retry a deny
If four of five boxes are already true, you want a workflow with one LLM step — extract, classify, draft — not an agent loop. Add the loop when the tool sequence actually branches.
If your “agent” always calls the same three tools in the same order and the only branch is 429 vs 401, delete the loop. Keep the table. You already had the answer.
FAQ
When should the agent escalate instead of retrying?
Escalate on auth, policy deny, ambiguous user intent, repeated same-tool failure, and money movement — especially when status is unknown and you lack an idempotency key. Retry only transient, idempotent failures with a bound and a clock (Retry-After or jittered backoff). The harness maps the class; the worker does not get a second unauthorized execute.
How do I measure whether escalate-versus-retry is working?
Track escalate rate by reason code, mean retry depth on retryable classes, same-fingerprint refusals, money-unknown rows past TTL, and duplicate provider ids per local key. Put those next to pass rate. A high escalate rate on ambiguous_intent can be healthy; a high pass rate with duplicate charges is not.
What usually fails first when teams try this?
Untyped "failed" strings, so auth and 404s get the same backoff as a 503. Next is a new idempotency key on timeout, which duplicates the write. Third is culture: if pages are punished, the loop learns to hide. Fix classification and the ledger before you tune prompts.
How long does this take to show results?
You will see the first illegal retries die in staging as soon as typed errors and the fingerprint bound are on — often the same week you wire them, not after a quarter of prompt edits. Duplicate-write proof takes one forced timeout drill on a money tool. I will not quote a universal percent-recovered; harvest your own reason codes and compare week over week.
What should I skip if I only have a week?
Skip multi-agent handoffs, new model shopping, and widening send/refund tools. Ship error classes, a fail-closed gate, a same-fingerprint bound, an escalate package with catalog reasons, and golden fixtures for auth, multi-match, and timeout-on-write. Leave n8n Retry On Fail bounded on the rail. That week is a brake, not a fleet.
When is this not worth doing yet?
If the path is fully known, put retries and a dead-letter on a workflow and stop. If you cannot write pass/fail criteria, do not add an escalate state to hide the argument. If nobody will staff the escalate queue, a package nobody reads is worse than a failed run. If credentials are still political, you will spend the week on access, not on the branch.
CTA
Illegal retries are how agents duplicate charges. Cage the branch before you widen tools.
What questions does this article answer?
- When should the agent escalate instead of retrying?
- Escalate on auth, policy deny, ambiguous user intent, repeated same-tool failure, and money movement — especially when status is unknown and you lack an idempotency key. Retry only transient, idempotent failures with a bound and a clock (`Retry-After` or jittered backoff). The harness maps the class; the worker does not get a second unauthorized execute.
- How do I measure whether escalate-versus-retry is working?
- Track escalate rate **by reason code**, mean retry depth on retryable classes, same-fingerprint refusals, money-unknown rows past TTL, and duplicate provider ids per local key. Put those next to pass rate. A high escalate rate on `ambiguous_intent` can be healthy; a high pass rate with duplicate charges is not.
- What usually fails first when teams try this?
- Untyped `"failed"` strings, so auth and 404s get the same backoff as a 503. Next is a new idempotency key on timeout, which duplicates the write. Third is culture: if pages are punished, the loop learns to hide. Fix classification and the ledger before you tune prompts.
- How long does this take to show results?
- You will see the first illegal retries die in staging as soon as typed errors and the fingerprint bound are on — often the same week you wire them, not after a quarter of prompt edits. Duplicate-write proof takes one forced timeout drill on a money tool. I will not quote a universal percent-recovered; harvest your own reason codes and compare week over week.
- What should I skip if I only have a week?
- Skip multi-agent handoffs, new model shopping, and widening send/refund tools. Ship error classes, a fail-closed gate, a same-fingerprint bound, an escalate package with catalog reasons, and golden fixtures for auth, multi-match, and timeout-on-write. Leave n8n Retry On Fail bounded on the rail. That week is a brake, not a fleet.
- When is this not worth doing yet?
- If the path is fully known, put retries and a dead-letter on a workflow and stop. If you cannot write pass/fail criteria, do not add an escalate state to hide the argument. If nobody will staff the escalate queue, a package nobody reads is worse than a failed run. If credentials are still political, you will spend the week on access, not on the branch.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
AI Agents Agentic Systems: An Operating Manual for Multi-Agent Work That Ships
An agentic system is evaluators, policy gates, sandboxes, and kill switches — not a chat window — so production tool work survives contact with real data.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.