Spurlock Studios
Contact
Share LinkedIn X
Two clipped paper packets. Thesis: SHOULD AGENT ESCALATE INSTEAD RETRYING.

Retry when the failure is transient and the write is idempotent. Escalate when the failure is auth, a policy deny, ambiguous user intent, the same tool failing with the same args, or anything that moves money. The worker does not get a vote. The harness maps a typed error class onto retry, escalate, or abort before the next tool execute fires.

This spoke sits under the Agentic Systems Operating Manual. The cage that makes the branch legal is state machines for agent loops. The judge that must still see a failed write is build the evaluator before the agent. What follows is the split itself: when another attempt is diligence, and when it is how you double-charge a card.

The short answer

  • Retry is for blips: 429, 503, timeouts, connect resets — on a read, or on a write that already holds a stable idempotency key.
  • Escalate is for everything a second identical call cannot fix: 401/403, policy deny, two legal customer ids, the same fingerprint after the bound, refunds and payouts with unknown status.
  • Abort is not escalate. Abort is poison, out of contract, or a kill switch. Do not page a human to finish work you already decided is illegal.
  • The model proposes. The runner refuses. A prompt that says “don’t retry auth” is a hint. A pre-execute gate is a control.
  • The escalate package is the product when the loop stops: goal, tools tried, error class, last artifact, budget remaining, and the exact question the human must answer.

When should the agent escalate instead of retrying?

Escalate when another attempt cannot change the outcome without a human, a credential, or a different intent. Retry when the outcome can change because the world blinked — and you can prove the next call is the same logical write.

SignalRetry?Escalate?Why
429 / 503 with Retry-AfterYes, cappedAfter the capTransient; HTTP already gave you a clock
Timeout / connect reset on a readYes, cappedAfter the capNo side effect to duplicate
Timeout on a write with a stored idempotency keyYes: reuse the key, then pollIf status stays unknown past TTLThe risk is a second identity, not waiting
Timeout on a write with no keyNoYesA second POST can double-apply
401 / 403 / missing scopeNoYesTokens do not heal themselves inside the loop
Policy deny or pending-approvalNoYes, or wait on the approval queueRetry-as-allow is how gates die
Ambiguous user intent (0 or N matches)One clarify ask, then stopYes if still ambiguousGuessing the account is a wrong write
Same (tool, args_hash, error_class) after the boundNoYesThat is no-progress, not persistence
Money movement (refund, payout, charge, transfer)Only receipt poll on the same keyDefault yesIrreversible + unknown is a human job
Schema / validation fail after one corrected payloadNoYesThe second “fix” is usually invention
Revision ceiling hit, criteria still failNoYesGrind-to-green is not a retry policy
Out of job contract / kill switchNoNo — abortDo not escalate illegal work

If you cannot put a row on this table for a tool, that tool is not in the loop. It is a hope with an HTTP client.

Decision list the runner should apply in order:

  1. Is the error class auth, policy, ambiguous_intent, or fatal? → escalate or abort. No backoff.
  2. Is the tool in the money class and status unknown without a key? → escalate.
  3. Is this the same fingerprint at or past the bound? → escalate with tool_retry_exhausted.
  4. Is it retryable and idempotent and budget remains? → backoff, then one more act.
  5. Else → escalate with unclassified_fail_closed.

Unknown is not retryable in production. Loosen a class after you have traces, not after a demo felt “almost.”

A legal retry changes the world’s chance of success without changing the identity of the write. That is a narrower idea than “the model wants to try again.”

Must be trueIf false
Error class is retryable (429, 503, timeout, some 5xx)Escalate or abort
Side-effect class is read, or a ledger already claimed an idempotency keyEscalate on writes
Fingerprint is new, or this is attempt k of k_max on the same retryable fingerprintEscalate on no-progress
Backoff honors Retry-After when presentImmediate re-call is a hammer
Per-run tool budget still has room for this callEscalate budget_exhausted
Policy gate would still allow the same payloadYou are not retrying a deny

RFC 9110 defines Retry-After as the wait before a follow-up request. MDN’s Retry-After notes the usual pairing: 503 (temporary unavailability) and 429 (rate limit). Those two are retries with a clock. 401, 403, 404, and 422 are not a clock. They are a stop.

Temporal’s own design note is the same split in workflow language: mark a failure non-retryable at the throw site — invalid input, missing records, authorization — because repeating the call will never succeed (non-retryable errors). Your tool adapter needs that flag. An error string the model can reread is not a policy.

Worked 429 on a read tool:

  1. Adapter returns rate_limited with retry_after_s from the header, or a jittered default if the header is missing.
  2. Runner stays in act but does not execute until the clock fires. n8n Wait is the rail version of this.
  3. Same fingerprint is allowed only because the class is retryable and attempt < bound.
  4. Still 429 at the bound → escalate tool_retry_exhausted. Do not pivot to a sibling write tool as “another way.”
  5. Next result is 401 → new class. Escalate auth. The 429 budget does not transfer.

Worked 429 on a write tool is the same clock plus the ledger key already claimed. If you cannot show the key id on the span, treat the write as unknown and escalate. A clock without an identity is how a rate limit becomes a duplicate row.

  • Tool adapter returns ok / retryable / fatal / auth / not_found / rate_limited / policy
  • Writes register an idempotency key before the HTTP client
  • Retry-After is copied onto the span, not discarded
  • Unknown error class defaults to retryable: false
  • n8n Retry On Fail (if you use it on the rail) has Max Tries and then Stop — it is not a second unbounded agent loop

n8n’s node Settings give you Retry On Fail for workflow nodes. That is the rail. The agent loop still needs its own bound. Two retry owners without a shared key is how a 503 becomes two CRM rows.

Why does auth escalate on the first failure?

Auth failures do not become valid credentials because the model asked again. A 401, a 403, an expired refresh, or a missing scope is an operator problem. Retrying it is how you lock an account and how you leak a “try a different token” plan into the transcript.

Auth signalLegal next stateIllegal next state
401 unauthorizedescalate tool_auth_errorMint a token from the prompt
403 forbidden / missing scopeescalate tool_forbiddenCall a sibling write tool “just to see”
Refresh token rejectedescalate tool_auth_expiredLoop the refresh grant
Tenant id mismatch on the gateescalate tenant_mismatchSwap tenant from ticket text
Secret missing at startupabortRetry the tool hoping env “settles”

Procedure the adapter should run on any auth-class result:

  1. Do not execute the same tool again in this run.
  2. Redact tokens from the observation the model sees. Leave the error class.
  3. Freeze further write tools for this run.
  4. Open escalate with the tool name, status, scope that failed, and the tenant id the gate saw.
  5. Page the owner of the credential, not the owner of the prompt.

The model will offer to “re-authenticate.” That sentence is not a control. If a human must rotate a secret, the run is already past self-help.

A retry that “works” after auth is usually a race: the operator fixed the secret while the loop was still burning tokens. Your dashboard will call that success. It was an escalate you failed to issue.

Why does a policy deny never retry as allow?

A policy gate already answered. deny means the payload is illegal. pending-approval means a human has to say yes. Re-proposing the same tool with synonyms, shuffled JSON keys, or a “pretty please” system-prompt addendum is an attack on your own gate.

Anthropic’s note on building effective agents is the cheap version of this: start simple, add loops only when a cheaper pattern fails. A deny is not a cheaper pattern failing. It is the pattern working.

Gate resultRunner actionModel visibility
allowExecute in the sandboxObservation from the tool
denyNo execute; escalate or abort by reasonReason code, not a puzzle to solve
pending-approvalPark on the approval queue“Waiting on operator”; no second propose
Policy timeout / unknown toolFail closed; abortNo retry-as-allow
  • Unknown tools cannot execute
  • Deny leaves a reason code an operator can filter
  • The worker cannot transition deny → act on the same payload
  • pending-approval is not a sleep-and-retry of the tool
  • Policy outage cannot be bypassed by “the model is sure”

If your traces show deny followed by a successful execute of the same tool in the same run, you do not have a gate. You have a log line.

The only legal “retry” after policy is a new payload that is a different job: a smaller row cap, a draft instead of a send, a read instead of a write. That is a new fingerprint. It still has to pass the gate. It is not a second try at the forbidden call.

Why is ambiguous user intent an escalate?

If the agent cannot name a single in-scope target, it must not pick a favorite. Multi-match CRM accounts, two email addresses in the request, a refund with no invoice id, “the Acme one” when you have Acme Inc and Acme LLC — those are not retrieval problems. They are intent problems. Another search_accounts with the same string is how you write the wrong tenant.

Intake shapeOne retry allowed?Then
Exactly one candidate after normalizeNo retry neededact on that id
Zero candidatesOne revised query or one clarify questionescalate no_match
Two or more candidatesOne disambiguation askescalate ambiguous_intent if still N
Required slot missing (amount, id, date)One askescalate missing_slot
User said two actions that contradict policyNoescalate intent_conflict

Procedure:

  1. Normalize names, emails, and external ids before search.
  2. If candidates.length != 1, do not call a write tool.
  3. Ask once, with the candidate list, if the channel can take a question.
  4. If the channel is async and nobody answers inside the job’s SLA, escalate with the list attached.
  5. Never let the worker “choose the closest embedding.” Closest is not unique.

This is the failure that looks like intelligence. The draft is fluent. The wrong crm_account_id is already in the payload. The evaluator has to reject unsupported identity, not grade the prose. Criteria that say “account id came from a unique match or the run escalated” belong in the golden set — see evaluators before agents.

Guessing under ambiguity is not a retry. It is a write you will spend the afternoon reversing.

When does repeated same-tool fail become escalate?

The first retryable failure may be a blip. The second identical call is a loop. Fingerprint tool + normalized args + error class before execute. If that fingerprint is already at the bound, refuse the call and escalate. Do not wait for the token budget to notice.

AttemptSame fingerprint, retryable: trueSame fingerprint, retryable: false
1Execute (first try)Do not retry; escalate now
2Execute after backoffAlready illegal
3 (bound)escalate tool_retry_exhausted—
Distinct fingerprint, same toolAllowed until the per-tool capStill no if class is auth/policy

A legitimate retry changes something material: args, a page cursor, a backoff clock. Equivalent JSON with keys reshuffled is not material. Strip request_id and injected timestamps before you hash, or every retry looks unique and the bound never fires.

  • Fingerprint computed before the HTTP client
  • Normalize has a unit test per write tool
  • Bound is a number in runner config, not a vibe in the prompt
  • Empty result [] on the same query counts as a fingerprint
  • Failure ledger lives outside the chat so summarization cannot resurrect the call

The state machine spoke is the cage: retryable may return to act once, then escalate. The model does not get skip_eval as a consolation prize.

If you cannot draw “same fingerprint → refuse,” you will call this diligence in the standup and a stampede in the bill.

How should money movement retry?

Money tools default to escalate. The exception is narrow: you already minted a stable idempotency key, the error is transient or indeterminate, and you are retrying the receipt — the same key, or a retrieve-by-metadata — not a second PaymentIntent.

Stripe’s low-level error handling is the cleanest vendor split I will cite here because the headers are explicit:

Stripe / HTTP signalSame key?Agent action
Stripe-Should-Retry: trueYesBackoff; retry the same identity
Stripe-Should-Retry: false—Stop; escalate
Timeout / no bodyYesTreat as unknown; retry same key or poll
500 / 502 / 503 / 504YesIndeterminate; never a new key
Content 400 you will fixNew key only after the body changesOne corrected payload, then escalate if it still fails
401After credentials existThat is auth. Escalate. Do not “retry the charge.”
429Wait; same key if the limiter ran after idempotencyClock, then bound, then escalate

Procedure for any refund, payout, charge, or transfer tool:

  1. Claim a ledger row with a unique idempotency key before the provider call.
  2. Put your local id in provider metadata so a webhook can finish a timed-out HTTP.
  3. On timeout: status unknown. Poll. Do not mint a second key.
  4. On Should-Retry: false or a non-retryable 4xx that is not a fixable schema miss: escalate money_unknown or money_rejected.
  5. On success: store the receipt; the evaluator must see it. A worker story about the refund is not a receipt.

If the upstream API has no idempotency header and no client-supplied external id, money movement is escalate-only until you build a local outbox. “The model will be careful” is not an outbox.

Poll loop for unknown money (this is not a second POST):

StepActionStop if
Retrieve by key or metadataFound receiptContinue to evaluate with the receipt
Webhook arrives with your local idReconcile the ledgerdone path, or escalate on mismatch
TTL exceeded, still unknownescalate money_unknownHuman checks the provider dashboard
Retrieve says not found and TTL is still shortStay unknown; do not POST againClock not done
Operator marks duplicate in the ledgerabort or escalate money_rejectedNever “retry to make it right”

Do not ask the model whether the refund worked. The model will narrate. The ledger and the provider are the only two oracles.

A retry that opens a new write identity on a charge is not a retry. It is a second payment.

Who decides retry versus escalate?

The runner. Not the worker, not the evaluator’s prose, not a Slack emoji on the trace.

DecisionOwnerWhat the model may do
Propose a tool and argsWorker, inside actSuggest allowlisted calls
Classify the tool resultTool adapterNothing — it receives a typed observation
Map class → next stateState machine / harnessNothing
Allow the next executePolicy gate + fingerprint boundPropose only
Declare doneEvaluator + machineCannot self-pass a failed write
Open escalateMachine, on listed eventsCannot veto

LangGraph’s Graph API says this with different nouns: nodes do work, edges decide what is next. Your edge table has to include error class, remaining retries, and money/auth/policy flags. If the only edge is “model said continue,” you built a chatbot with side effects.

n8n is a natural rail for the branch you already trust: backoff, durable execution, a dead path that pages. Put the agent in the step that must choose a tool. Put retry-vs-escalate in the graph around it. I have collaborated with the n8n team; the canvas is not the control plane. Max Tries on a node plus an escalate sub-workflow is.

  • (state, event, guards) → next_state is data, not a paragraph
  • escalate and abort are named terminals
  • The worker cannot increment its own retry counter downward
  • Evaluator fail with ceiling remaining goes to revise, not a raw tool hammer
  • Evaluator fail at ceiling goes to escalate, not act

Freedom stays inside the state. The split lives on the edge. That is the whole point of the operating manual stack: policy, sandbox, machine, then the model.

What must the escalate package contain?

Escalate is a designed success of a kind: the system stopped while a human could still finish. If the package is “something broke, good luck,” you trained operators to ignore the queue.

FieldRequiredWhy
job_contract + hard nosYesThe human should not rediscover scope
reason_code from a catalogYesFilter, alert, golden-set harvest
Tools tried: name, fingerprint, error classYesNo-progress evidence
Last artifact (redacted)YesResume without re-doing reads
Evaluator failures + evidenceYesWhat “done” still lacks
Budget remaining / spentYesWhether finishing is even cheap
The question the human must answerYesOne decision, not a novel
Candidate list (if ambiguous_intent)When relevantDo not hide the N-match
Receipt / ledger id (if money)When relevantPoll, do not re-charge
Trace idYesOps, not archaeology

Catalog starters you can copy:

  • tool_auth_error
  • tool_forbidden
  • policy_deny
  • ambiguous_intent
  • missing_slot
  • tool_retry_exhausted
  • money_unknown
  • money_rejected
  • revision_ceiling
  • budget_exhausted
  • unclassified_fail_closed

Minimum package shape (illustrative — your schema can be stricter):

terminal: escalate
reason_code: ambiguous_intent
job: "Update billing email for the account named in the ticket"
hard_nos: ["no refund", "no send to customer"]
question_for_human: "Which crm_account_id is in scope — 441 or 882?"
candidates:
  - { id: 441, name: "Acme Inc" }
  - { id: 882, name: "Acme LLC" }
tools_tried:
  - { tool: search_accounts, fingerprint: "…", error_class: ok, note: "N=2" }
last_artifact: internal_draft.md
evaluator_unmet: ["account id from unique match"]
budget: { spent_usd: null, calls_left: 4 }
trace_id: "run_…"
ledger_key: null

Leave spent_usd null if you do not meter dollars yet. Do not invent a cost. Calls-left and reason code are enough to start.

Procedure when the machine opens escalate:

  1. Stop all write tools for the run.
  2. Persist the package before you notify.
  3. Notify the queue you actually staff — not a channel nobody owns.
  4. Leave the run resumable from act only after a human supplies the missing slot or approval. Do not auto-resume on a timer.
  5. Harvest the case into the golden set the same week, with expected terminal escalate.

If operators keep finishing jobs from the raw transcript instead of the package, the package is wrong. Fix the schema. Do not add another summary model on top.

What fails first when teams skip this split?

The first failure is almost never “the model is dumb.” It is a retry storm that looks like uptime: green executions, climbing token spend, duplicate CRM writes, and a refund that landed twice because timeout was treated like “try again with a fresh id.”

Concrete shape, no invented client numbers:

  1. Billing tool times out. Adapter returns "failed" with no class.
  2. Worker retries with a new client request id.
  3. Provider had already captured the first POST.
  4. Evaluator never sees a receipt; worker marks the run done from the second 200.
  5. Finance finds the duplicate on the card. Ops finds a “successful” agent run.

Second shape: ambiguous intent dressed as search.

  1. Ticket says “update Acme.” Search returns Acme Inc and Acme LLC.
  2. Worker calls search again with the same string, then picks the first row.
  3. Write lands on the wrong tenant. Evaluator grades the email tone. Tone passed.
  4. You learn from the angry reply, not from the trace.

Fix: candidates.length != 1 is illegal for act. The golden case expects escalate ambiguous_intent with both ids in the package. Closest embedding is not unique.

What brokeWhat it costsWhat you do instead
Untyped "failed"Retry of auth, 404, and money alikeTyped classes; unknown → fail closed
New idempotency key on timeoutDuplicate charge / duplicate ticketLedger first; reuse key; poll
Prompt-only “don’t loop”Returns after the next helpfulness editPre-execute fingerprint bound
Continue-on-error on the railGreen run, Error Workflow never firesStop or branch; do not mute writes
Self-pass after a failed writeSilent wrong doneEvaluator sees tool outcome
Escalate as a Slack dumpHumans ignore the queueCatalog reason + one question
Punishing escalate rate in reviewsHidden retries to look autonomousTreat escalate as a designed terminal

The culture failure is real even when the code is not. If leadership rewards “the bot handled it” and punishes pages, the harness will grow illegal retries. Healthy ops treat escalate like a caught exception: cheaper than the duplicate write.

Do not invent a “retry success rate” to paper over this. Count duplicates, unknown-money rows, and same-fingerprint refusals. Those are receipts. A percentage without a denominator is a slide.

How do I evaluate this in production?

You evaluate the branch, not the vibes. A high pass rate with hidden retries is grind. A high escalate rate can be the system working — if the reasons are ambiguous_intent and policy_deny, not unclassified_fail_closed on every 503 you forgot to mark retryable.

SignalWhat it tells youVeto / inspect if…
Escalate rate by reason_codeMix of designed stops vs leaksunclassified_fail_closed dominates
Mean retry depth on retryable classesWhether backoff is doing workDepth sits at the cap on every run
Same-fingerprint refuse countNo-progress detector is aliveCount is zero while 5xx are not
Money unknown rows open past TTLLedger / poll is brokenAny row older than your poll window
Duplicate provider ids per local keySecond write identity leakedCount is not zero
Online vs golden escalate mixProduction driftNew codes with no fixture
Human finish time from packagePackage qualityOperators still open raw traces first
Pass rate beside the aboveQuality, not a cover storyPass holds while duplicates rise

Golden-set the split. For each fixture, expected_terminal is done, escalate, or abort, and the reason code is part of the assertion. Auth fixtures must not done. Multi-match fixtures must not pick. Timeout-on-charge fixtures must not mint a second key.

FixtureInput trapExpected terminalFail if
auth-401-writeWrite tool returns 401escalate tool_auth_errorAny second execute of that tool
policy-deny-refundGate denies issue_refundescalate policy_deny or abortRefund tool HTTP happens
ambiguous-acmeTwo CRM ids for one nameescalate ambiguous_intentWrite on either id
retry-429-then-okFirst read 429, second 200done if criteria passImmediate re-call with no wait
timeout-charge-same-keyCharge HTTP times outPoll / retry same key, or escalate money_unknownNew key appears in the ledger
same-fingerprint-3xIdentical args, retryable 5xxescalate tool_retry_exhausted at boundAttempt 4 of the same hash
schema-fail-twice422, model “fixes,” 422 againescalateThird invented payload

OpenAI’s agent evals treat tool choice and trajectory as first-class grades. Use that idea even if you never touch their dashboard: a run that “passes” after three identical refund calls is a trajectory fail.

Cadence that does not need a fabricated SLA:

  1. Unit: adapter classification + fingerprint normalize, every commit.
  2. Task: golden set including escalate fixtures, every behavior change.
  3. Online: sample production terminals daily or weekly; alert on new reason codes.
  4. Silent-fail: once a week, sample “done” runs that finance or support later rewrote.

If you cannot name expected_terminal for a case, you cannot tell retry from escalate. Write the fixture before you widen the tool.

What guardrails do I need?

You need the same stack the operating manual already named — but aimed at this branch. Without them, “retry vs escalate” is a slide in a design doc.

GuardrailJob on this split
Typed tool errorsGive the machine a class, not a paragraph
Pre-execute policy gatedeny cannot become allow by persistence
Fingerprint + boundSame call dies on purpose
Idempotency ledgerWrites can wait without cloning
Money tool allowlistDefault pending or deny in week one
Revision ceilingEval fail does not become infinite act
Per-run budget429 storms still halt
Escalate package schemaHumans finish; they do not excavate
Kill switchAbort illegal work; do not escalate it
Trace fieldsreason_code, retry depth, key id, gate decision

Week-one checklist — skip the fleet, keep the brake:

  • Every write tool has an error class enum the runner understands
  • Auth and policy classes have zero retry budget
  • Money tools are pending-approval or escalate-on-unknown
  • Multi-match intake escalates; it does not “pick closest”
  • Same-fingerprint bound is 2 for retryable, 0 for the rest
  • n8n Retry On Fail (rail) is bounded and does not Continue the leftovers into a charge
  • Golden set has at least one fixture per escalate reason you claim to support
  • You have tripped escalate in staging and opened the package as an operator

If a box is empty, do not add another model. Fill the box.

When is a workflow enough instead of an agent?

When the retry/escalate split is the whole job. Known path, structured input, rare judgement: n8n (or any workflow engine) already has retries, Wait, and a dead-letter. Buying an agent so a model can “decide” to retry a 429 is how you pay token tax for a backoff you already owned.

SignalBuild thisRetry/escalate lives where
Fixed steps, typed I/O, rare exceptionsWorkflowNode Retry On Fail + Error Workflow + DLQ
Path varies, criteria exist, tools are manyAgentic loop in a machineAdapter classes + edge table
Path varies, criteria are mushWorkshop, not a buildNowhere — do not automate the argument
One irreversible write, no keyHuman + checklistEscalate-only; no agent “careful POST”
Volume is tinyHuman + templatePage, don’t loop
You only wanted a chat UIWrong lane—

Default to the workflow when you can write the branch in a table without a model. Default to the agent when tool choice is the uncertainty and the table still holds for errors. The brake pedal in the parent manual is the same: no evaluator, no agent.

A hybrid that ships: workflow owns delivery, idempotency, and the escalate queue; a bounded agent owns the step where the tool is not known in advance. The agent returns done / escalate / abort. The workflow never lets the model own HTTP retry.

  • I can draw the retry/escalate table without naming a model
  • 429/503 already have a Wait or Retry On Fail on the rail
  • 401 already pages, it does not loop
  • Money writes already have a key or they are human-only
  • The remaining uncertainty is which allowlisted tool, not whether to retry a deny

If four of five boxes are already true, you want a workflow with one LLM step — extract, classify, draft — not an agent loop. Add the loop when the tool sequence actually branches.

If your “agent” always calls the same three tools in the same order and the only branch is 429 vs 401, delete the loop. Keep the table. You already had the answer.

FAQ

When should the agent escalate instead of retrying?

Escalate on auth, policy deny, ambiguous user intent, repeated same-tool failure, and money movement — especially when status is unknown and you lack an idempotency key. Retry only transient, idempotent failures with a bound and a clock (Retry-After or jittered backoff). The harness maps the class; the worker does not get a second unauthorized execute.

How do I measure whether escalate-versus-retry is working?

Track escalate rate by reason code, mean retry depth on retryable classes, same-fingerprint refusals, money-unknown rows past TTL, and duplicate provider ids per local key. Put those next to pass rate. A high escalate rate on ambiguous_intent can be healthy; a high pass rate with duplicate charges is not.

What usually fails first when teams try this?

Untyped "failed" strings, so auth and 404s get the same backoff as a 503. Next is a new idempotency key on timeout, which duplicates the write. Third is culture: if pages are punished, the loop learns to hide. Fix classification and the ledger before you tune prompts.

How long does this take to show results?

You will see the first illegal retries die in staging as soon as typed errors and the fingerprint bound are on — often the same week you wire them, not after a quarter of prompt edits. Duplicate-write proof takes one forced timeout drill on a money tool. I will not quote a universal percent-recovered; harvest your own reason codes and compare week over week.

What should I skip if I only have a week?

Skip multi-agent handoffs, new model shopping, and widening send/refund tools. Ship error classes, a fail-closed gate, a same-fingerprint bound, an escalate package with catalog reasons, and golden fixtures for auth, multi-match, and timeout-on-write. Leave n8n Retry On Fail bounded on the rail. That week is a brake, not a fleet.

When is this not worth doing yet?

If the path is fully known, put retries and a dead-letter on a workflow and stop. If you cannot write pass/fail criteria, do not add an escalate state to hide the argument. If nobody will staff the escalate queue, a package nobody reads is worse than a failed run. If credentials are still political, you will spend the week on access, not on the branch.

CTA

Illegal retries are how agents duplicate charges. Cage the branch before you widen tools.

/agentic · Start an agentic pilot

FAQ

What questions does this article answer?

When should the agent escalate instead of retrying?
Escalate on auth, policy deny, ambiguous user intent, repeated same-tool failure, and money movement — especially when status is unknown and you lack an idempotency key. Retry only transient, idempotent failures with a bound and a clock (`Retry-After` or jittered backoff). The harness maps the class; the worker does not get a second unauthorized execute.
How do I measure whether escalate-versus-retry is working?
Track escalate rate **by reason code**, mean retry depth on retryable classes, same-fingerprint refusals, money-unknown rows past TTL, and duplicate provider ids per local key. Put those next to pass rate. A high escalate rate on `ambiguous_intent` can be healthy; a high pass rate with duplicate charges is not.
What usually fails first when teams try this?
Untyped `"failed"` strings, so auth and 404s get the same backoff as a 503. Next is a new idempotency key on timeout, which duplicates the write. Third is culture: if pages are punished, the loop learns to hide. Fix classification and the ledger before you tune prompts.
How long does this take to show results?
You will see the first illegal retries die in staging as soon as typed errors and the fingerprint bound are on — often the same week you wire them, not after a quarter of prompt edits. Duplicate-write proof takes one forced timeout drill on a money tool. I will not quote a universal percent-recovered; harvest your own reason codes and compare week over week.
What should I skip if I only have a week?
Skip multi-agent handoffs, new model shopping, and widening send/refund tools. Ship error classes, a fail-closed gate, a same-fingerprint bound, an escalate package with catalog reasons, and golden fixtures for auth, multi-match, and timeout-on-write. Leave n8n Retry On Fail bounded on the rail. That week is a brake, not a fleet.
When is this not worth doing yet?
If the path is fully known, put retries and a dead-letter on a workflow and stop. If you cannot write pass/fail criteria, do not add an escalate state to hide the argument. If nobody will staff the escalate queue, a package nobody reads is worse than a failed run. If credentials are still political, you will spend the week on access, not on the branch.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot