Spurlock Studios
Contact
Share LinkedIn X
A cracked amber fuse. Thesis: IDEMPOTENT AGENT TOOL WRITES RETRIES.

When an agent tool write times out, the dangerous question is not “did the model fail?” — it is “did the side effect already land?” Make writes safe by minting a stable idempotency key in the runtime before any retry layer can fire, then reusing that same key for model retries, harness retries, and HTTP client retries.

This spoke sits inside the Agentic Systems Operating Manual. It owns keys so retries do not double-write. Argument shape lives in tool schemas agents follow. Packaging lives on /agentic.

The short answer

  • Timeouts are ambiguous: the upstream may have committed while your agent saw a network error. Stripe’s own essay names the three failure cuts — connect never happens, work dies mid-flight, work finishes and the response never returns.
  • Agents stack retries (model loop + harness + HTTP). One timed-out write can become two charges or two emails if each layer invents a new identity.
  • Birth the key in the runtime from (run_id, tool_name, intent_fingerprint), not in the prompt and not in the HTTP client.
  • Prefer a native Idempotency-Key header when the API supports it. Stripe documents the header, a 255-character cap, and a 24-hour v1 replay window.
  • Without a header, claim the key in a local ledger (or a client-supplied external_id) before the write. No claim, no HTTP.
  • Test duplicate delivery in staging with forced timeouts before you grant write autonomy.

What is the agent-specific idempotency failure mode?

Workflow automation usually has one retry owner (the workflow engine). Agents have three:

LayerWhat retriesTypical triggerHow it double-writes
Model loop“Tool failed, try again”Tool error text in the next turnNew tool call, often a new UUID in args
Harness / runnerRe-invoke the act stepTimeout, crash, checkpoint resumeSame intent, new HTTP attempt, new key if the runtime forgot
HTTP clientSame request again408 / 429 / 5xx, connection resetSafe only if the header/key is stable

If each layer invents its own “try again” without a shared key, a single ambiguous timeout becomes stacked side effects. That is the agent-specific failure — not “webhooks can fire twice,” but “three systems each think they are being helpful.”

HTTP semantics already treat GET, PUT, and DELETE as idempotent methods. Agent tools are usually POST-shaped: create a charge, send a message, open a ticket. Those are the verbs that need a key.

Why “only call once” in the prompt fails

Prompts do not control TCP. A model that obediently “calls send_email once” still loses when:

  1. The HTTP call hangs past the client timeout after the provider accepted the message.
  2. The harness resumes the run after a deploy and re-enters state:act.
  3. The model sees a generic timeout string and emits a second tool call with slightly different arguments — and a new message_id it invented.

Instructional discipline is not a transport guarantee. Treat “call once” as documentation for humans, not as a safety control.

If the tool schema exposes idempotency_key as a model-writable string, you made the failure easier. The model will “helpfully” rotate the field after a timeout. Strip or overwrite that property at the adapter. The schema can document that a key exists; it must not let the model own it.

Checklist for the prompt-and-schema layer:

  • System prompt does not ask the model to generate UUIDs for writes
  • Tool schema does not list idempotency_key as a required model argument
  • Adapter injects the runtime key after validation, before HTTP
  • Tool error text distinguishes timeout / unknown from validation_failed

How do I make agent tool writes safe when the call times out?

Procedure that holds up in production:

  1. Classify the tool as read, write_idempotent, or write_irreversible before registration.
  2. Mint a key in the runtime when the act step decides to call a write tool — before the HTTP request starts.
  3. Persist key → status (pending | succeeded | failed_poison) in a ledger keyed by tenant.
  4. Pass the same key into every retry of that logical write: harness replay, HTTP retry, and any model re-emit for the same intent.
  5. On timeout: leave status pending (or unknown), do not mint a new key, and either poll for receipt or escalate — never “just send again” with a fresh identity.
  6. On success: store the upstream receipt id beside the key; mark succeeded.
  7. On definitive failure (4xx that will not succeed on retry): mark failed_poison so retries stop.

Timeout means unknown. Unknown means reuse the key or escalate — never invent a second write identity.

Observed outcomeStatus to storeNext action
HTTP 2xx + receipt idsucceededReturn cached receipt on later claims
Client timeout / connection resetpending / unknownReuse key; poll or escalate
Hard 4xx (bad args, declined card you will not resubmit)failed_poisonStop; new intent needs a new key
409 / “already exists” with a receiptsucceededTreat as the first write landing
500 with no receiptunknownReuse key; do not rotate

How do I generate stable keys across retry layers?

Key material should be stable for the business intent, not for the HTTP attempt:

key = hash(tenant_id + run_id + tool_name + intent_fingerprint)

intent_fingerprint is a canonical hash of the fields that define the side effect (to, template_id, invoice_id, amount_cents) — not of ephemeral fields like requested_at or random UUIDs the model invents.

Stripe’s API reference suggests V4 UUIDs or another high-entropy string, caps keys at 255 characters, and tells you not to put email addresses or other personal identifiers in the key. A SHA-256 hex of the tuple above stays under that cap and stays free of PII. If you prefer a UUID, mint it once in the runtime and persist it next to the intent — do not mint a UUID per HTTP attempt.

Source of keySafe?Why
Model-generated UUID in tool argsNoNew UUID on every re-emit
HTTP attempt id / x-request-id per tryNoNew per transport retry
Runtime: run_id + tool + intent hashYesSurvives all three layers
Upstream event id (when writing because of an event)YesAligns with business identity
Shopping-cart / invoice id (Stripe’s second strategy)YesStripe names this as a way to block double submit

The IETF Idempotency-Key draft (revision 07, 15 October 2025; the listed expiry was 18 April 2026, and as of August 2026 it is still a draft, not an RFC) says the same thing in standards language: the client generates a unique key, must not reuse it with a different payload, and the server may fingerprint the body. Stripe is listed in that draft’s implementation status. Until it is an RFC, treat vendor docs as the contract you actually get.

Where the key is born — model vs runtime

BirthplaceOutcome
Model fills idempotency_keyModel invents a new key after timeout; duplicates ship
HTTP library mints a UUID on each request()Harness resume looks like a new charge
Runtime injects key into the tool-call envelopeRetries reuse identity even if the model rephrases args
Runtime + schema forbids model overrideStrongest: model cannot “helpfully” rotate the key

Default: the harness owns the field. If the tool schema exposes idempotency_key, strip or overwrite model-supplied values before dispatch.

This is the same reason tool schemas should keep write-path enums tight. A model that can invent message_id or charge_ref has a second way to break identity even when your header is correct.

Procedure at the adapter:

  1. Validate args against the schema (types, enums, required).
  2. Compute intent_fingerprint from the canonical write fields.
  3. Load or insert idempotency_key from the ledger for (run_id, tool_name, intent_fingerprint).
  4. Overwrite any model-supplied key.
  5. Dispatch HTTP with that key in the header or external_id field.

What does Stripe’s Idempotency-Key actually guarantee?

Copy Stripe’s contract into the adapter. Do not invent a friendlier one.

Idempotent requests (API v1):

  • Send Idempotency-Key on POST. GET and DELETE ignore it; those verbs are already idempotent.
  • Stripe saves the status code and body of the first request for that key — including 500 errors — and replays that result.
  • Keys are removed after they are at least 24 hours old. Reusing a pruned key starts a new request.
  • Incoming parameters are compared to the original. A mismatch errors so you cannot accidentally reuse a key for a different charge.
  • Results are saved only after endpoint execution begins. Validation failures and some concurrent conflicts are not stored, and you may retry those.

API v2 changes the replay rules:

RuleAPI v1API v2
Verbs that accept keysPOST onlyPOST and DELETE
Replay window≥ 24 hours30 days, same account or sandbox, same API
Failed-request replayReturns the cached first result, including 5xxRe-executes the failed work without extra side effects, or explains why it cannot
If you omit a keyYou own the riskStripe mints a UUID for you — useless for agent retries unless you persist it

If the SDK mints a key you never stored, the next harness resume cannot replay. Always set the key yourself from the ledger.

Low-level error handling adds two headers worth wiring into the tool span:

HeaderMeaning
Idempotent-Replayed: trueThis response is the stored first result, not a second charge
Stripe-Should-Retry: trueRetry this request (still back off)
Stripe-Should-Retry: falseDo not retry; another attempt will not help
Header absentFall back to status-code rules

A replayed 200 with Idempotent-Replayed: true is success. Log it as a retry hit, not as a second write.

How should I retry Stripe 4xx, 5xx, and timeouts?

Do not use one retry policy for every status. Stripe splits errors into content, network, and server — and the idempotency layer does not treat them the same.

ClassExamplesSame key?New key?Agent action
Network / timeoutSocket reset, read timeout, no bodyYesNeverRetry until you have a Stripe result
Content 4xx you will fixMissing param, bad enumOnly if you did not change the bodyYes, after you change paramsStripe caches a 400 once the method started; a fixed body needs a new key
429 rate limitToo many requestsMaybePrefer wait, same keyStripe notes the limiter can run before the idempotency layer, so the same key may not replay
401 missing keyAuth header omittedDo not trust cacheAfter you attach credentialsSame caveat as 429
409 conflictConcurrent request with the same keyYesNoWait; this is the in-flight case the IETF draft also maps to 409
500 / 502 / 503 / 504Stripe-side failureYesNoTreat as indeterminate. A new key can double-charge if the first attempt already hit the card network

On 500, Stripe says they will try to reconcile — roll forward if the payment network already saw the charge, roll back if not — and may fire a webhook later for an object you never saw on the API response. That is why you send a local identifier in metadata on create: the webhook can carry it even when the HTTP client timed out.

Agent mapping:

  1. Timeout or network error → reuse key, exponential backoff with jitter (Stripe’s 2017 write-up is still the clean description of backoff + jitter vs a thundering herd).
  2. Stripe-Should-Retry: true → same key, wait.
  3. Stripe-Should-Retry: false → stop; escalate.
  4. 500 → same key, mark unknown, watch webhooks / retrieve by metadata. Do not open a second PaymentIntent identity.
  5. Validation 400 you will correct → new key only after the body changes, and only if you are sure the first attempt never executed.

The HTTP method table in RFC 9110 will not save you here. A Stripe POST /v1/payment_intents is not a PUT. The header is the whole safety story.

What if the upstream API has no Idempotency-Key header?

Many CRM, email, and internal APIs do not speak Stripe-style idempotency. Some vendors use a different name for the same idea. The IETF draft’s implementation list is a useful map: Adyen uses Idempotency-Key; PayPal uses PayPal-Request-Id; Twilio webhooks use I-Twilio-Idempotency-Token; Square puts idempotency_key in the JSON body. Read the vendor page. Do not send a header they ignore and call the job done.

Options, in order of preference:

  1. Native unique constraint — if the API accepts a client-supplied external id (external_id, reference, client_ref, PayPal-Request-Id), use your runtime key there.
  2. Pre-write ledger gate — before calling the API, claim the key in your DB with a unique index. If claim fails because status is succeeded, return the stored receipt and skip the call. If pending and younger than TTL, wait or poll. If older than TTL, escalate.
  3. Read-before-write with a stable lookup — only when the domain has a natural unique query (invoice already paid, ticket already has comment hash X). Fragile; document the race.
  4. Outbox + single worker — enqueue the write once; a single consumer performs the HTTP call. Agent retries enqueue the same outbox id.
API behaviorWhat you do
Speaks Idempotency-Key (Stripe, Adyen, others)Send the runtime key; still keep a local ledger for timeouts
Speaks a renamed header / body fieldMap the same key into that field
Offers external_id uniquenessPut the key there; handle the duplicate error as success-with-receipt
Offers nothingLedger claim is the control. No claim, no HTTP

Do not pretend a header exists. Build the ledger. Workflow graphs have the same storage problem; agents need the idea on the tool boundary, not on the webhook trigger.

Read tools vs write tools — retry rules

Tool classRetry on timeout?Key required?
Read / searchYes, usually safeOptional (cache key helps)
Write with server idempotencyYes, same keyRequired
Write without server idempotencyOnly after ledger claim or escalateRequired locally
Irreversible external (wire, legal notice)Human or outbox onlyRequired + approval

RFC 9110 is why reads feel easy: a repeated GET is supposed to leave server state alone. Your “search_contacts” tool is a GET in spirit even if the wrapper is POST. Your “send_invoice_reminder” tool is not.

Checklist before marking a tool write_idempotent in the registry:

  • Side-effect class documented (email, charge, crm_write, ticket)
  • Key birthplace = runtime
  • Ledger or native unique field wired
  • Timeout path leaves status unknown / pending, not failed
  • Stripe (or vendor) error classes mapped — especially 500 and 429
  • Forced duplicate-delivery test exists
  • Metadata / correlation id set so a late webhook can join the run

If any box is open, the tool is write_irreversible until you close it.

Compensating actions that stay idempotent

Compensations (void charge, send apology, delete draft) are also writes. They need their own keys, derived from the original:

compensate_key = hash(original_key + ":compensate:" + action)

Rules:

  1. Never compensate twice for the same original key.
  2. Never compensate if the original write status is still pending — resolve unknown first. A void fired against a charge that never landed is how you get a later “success” webhook and a confused ledger.
  3. Log compensation under the same run_id with a distinct tool span.
  4. If the original was a Stripe 500, wait for the webhook or a retrieve-by-metadata before you refund. Refunding a charge that Stripe rolled back is a second mess.
Original statusCompensate?
succeeded + receiptYes, with compensate_key
pending / unknownNo — resolve first
failed_poisonNo — nothing to undo
succeeded and already compensatedReturn cached compensation receipt

Blind “undo” loops are how you get a charge, a void, and a second charge.

Failure example: double invoice email

Job: Agent drafts and sends invoice reminder for inv_8841.

What happened:

  1. Runtime minted no key; model called email.send.
  2. Provider accepted the message; client timed out at 30s.
  3. Harness retried state:act. Model called email.send again with a new message_id it invented.
  4. Customer received two reminders; support spent a day on “your system is broken.”

Cost: trust, not just SMTP fees.

Fix:

  • Runtime key: hash(tenant + run + email.send + inv_8841 + template_reminder_v2)
  • Ledger claim before SMTP
  • On timeout, poll provider by key/metadata or escalate — do not re-emit with a new message id

The charge version of the same story is worse. Swap email.send for stripe.payment_intents.create, skip the header, and a 30-second timeout becomes two PaymentIntents. Stripe’s docs exist because that exact network cut is normal, not exotic.

I have watched this class of bug in production agent loops across a book of 500+ automations and 20,000+ hours on agentic systems. The model is rarely the interesting part. The missing key is.

How do I test duplicate delivery before production?

Staging drills that catch the stacked-retry bug:

  1. Inject latency past the HTTP timeout after the mock server records the write.
  2. Confirm harness retry reuses the same key and the mock sees one logical commit.
  3. Force model re-emit by returning a fake timeout string once; assert the second tool call carries the injected key (or is blocked).
  4. Crash mid-pending and resume from checkpoint; assert no second charge.
  5. Poison 409 / duplicate from upstream; assert the agent treats it as success-with-receipt, not endless retry.
  6. Stripe-shaped 500 from the mock, then a webhook with your metadata; assert the ledger flips to succeeded without a second POST.
DrillPass criterion
Slow success + client timeoutExactly one side effect
Double harness resumeLedger blocks second HTTP
Model invents new args, same intentSame key; one effect
Upstream duplicate errorMaps to succeeded
Mock 500 + later webhookOne object; ledger joins on metadata
Same key, different amountAdapter refuses; Stripe would error the mismatch

If you have not run the timeout drill, you have not tested agent writes.

Ledger fields that belong on the trace

Put these on the tool span and in the ledger row:

FieldPurpose
idempotency_keyShared identity across retries
intent_fingerprintProve which args defined the write
statuspending / succeeded / failed_poison / unknown
attemptTransport attempt count (not a new key)
upstream_receipt_idCorrelate to CRM / email / PaymentIntent
idempotent_replayedStripe (or vendor) said this body was a replay
stripe_should_retryVendor retry hint, when present
first_seen_at / succeeded_atDispute timeline
run_id / tool_call_idJoin to the agent trace
metadata_local_idThe id you sent so a late webhook can find you

Without receipt correlation, ops cannot answer “which run sent the second email?”

Anti-patterns

UUID in the prompt template. Guarantees uniqueness per emit — the opposite of idempotency.

Retrying irreversible tools on any error string. Distinguish timeout / unknown from validation_failed.

Per-layer keys. Model key ≠ harness key ≠ HTTP key means three charges.

Deleting ledger rows on failure. If the write may have landed, keep the key until you know.

Treating HTTP 200 as the only success. Some APIs return errors after committing; prefer receipt ids. Stripe’s Idempotent-Replayed: true on a cached body is success too.

Minting a new Stripe key after a 500. Stripe advises against it because the original key may already have produced side effects.

Letting the official SDK “just retry.” Libraries can attach keys and back off — Stripe says you must configure that. If the SDK mints a key per process start, your harness resume is a new customer.

Hashing the entire request body as the key. Two legitimate sends with the same body (two invoices, same amount, different customers) collapse. Fingerprint the intent fields, then store a unique key for that intent.

Decision list: ship write autonomy?

Ship autonomous writes only when all are true:

  1. Tool is classified and keyed in the runtime.
  2. Upstream supports idempotency or a local ledger gate is live.
  3. Timeout drill passed in staging.
  4. 500 / webhook path is defined for money-moving tools.
  5. Evaluator or policy gate can block high-risk tools (operating manual).
  6. Kill switch can freeze the write tool class without redeploying prompts.

If any box is open, keep the tool behind human approval or an outbox.

Worked ledger claim (pseudo)

claim(key):
  insert ledger(key, status=pending) on conflict do nothing
  if conflict and status=succeeded: return cached_receipt
  if conflict and status=pending and age < TTL: wait or escalate
  if conflict and status=pending and age >= TTL: escalate unknown
  if inserted: call upstream with Idempotency-Key=key (or external_id=key)
  on success: status=succeeded, store receipt
  on timeout: leave pending, schedule resolve job
  on hard 4xx: status=failed_poison
  on 500: leave unknown, attach metadata_local_id, watch webhook

Agents call claim through the tool adapter — never raw HTTP from the model. Hash intent fields (to, template_id, invoice_id); do not put raw customer bodies into the key string. Stripe’s 255-character limit and “no PII in the key” rule are the right constraints even when the vendor is not Stripe.

TTL should outlive your retry window and sit inside the vendor’s replay window. For Stripe v1 that window is 24 hours; for v2 it is 30 days. A ledger TTL of a few hours is enough for agent retries and still forces a human on a key that stayed pending overnight.

The resolve job is not optional for money tools. After a timeout it should:

  1. Re-GET or retrieve-by-metadata using the local id you stored.
  2. If a receipt exists, mark succeeded and return it — do not POST again.
  3. If nothing exists and the vendor key is still inside the replay window, POST once more with the same Idempotency-Key.
  4. If the window is gone and status is still unknown, page a human. A pruned Stripe v1 key is a new request if you reuse it.

That job is how you turn “the client hung up” into a closed ledger row instead of a second charge.

Pilot minimum

A Spurlock Studios $1,500 · 5-day agentic pilot that includes write tools ships: runtime key injection, a thin ledger, timeout classification, Stripe-header (or vendor-field) mapping, and one forced-duplicate drill in staging — not a promise that “the model will be careful.”

/agentic · /contact?intent=agentic-pilot

FAQ

How is this different from n8n idempotency keys?

n8n idempotency keys dedupe workflow executions and webhook redeliveries inside an automation graph. Agent idempotency keys dedupe tool writes across model loops, harness resumes, and HTTP clients. Same idea, different boundary: the tool adapter, not the workflow trigger. If both run in one system, they still need distinct keys — a workflow retry and a tool retry are not the same intent.

Read tools vs write tools — retry rules?

Reads can usually retry freely; writes need a stable key and a ledger or native idempotency before any retry. Irreversible writes should escalate or use a single-consumer outbox when status is unknown after timeout. Classify the tool at registration time so the harness cannot “just retry” a wire transfer because the model said the tool failed.

Where should the key be born — model or runtime?

Runtime. Keys born in the model get rotated on every re-emit after a timeout, which causes the duplicates you are trying to prevent. Inject and overwrite at the harness boundary. If the schema lets the model pass idempotency_key, treat that field as hostile input.

What if the upstream API has no Idempotency-Key header?

Use a client-supplied unique field if the API has one, or claim the key in your own ledger before the call and skip or replay from stored receipts. Some vendors rename the header (PayPal-Request-Id, Square’s body idempotency_key). Map your runtime key into whatever they actually honor. Do not invent a header the vendor ignores.

How do compensating actions stay idempotent?

Derive a compensation key from the original key plus action name, refuse to compensate while the original is still pending, and record compensation on the same run trace so you never void twice. If the original was a Stripe 500, wait for retrieve-or-webhook before you refund.

What ledger fields belong on the trace?

At minimum: idempotency_key, intent_fingerprint, status, attempt, upstream_receipt_id, timestamps, and run_id. Add idempotent_replayed and metadata_local_id when the vendor is Stripe. Those fields let ops prove one logical write across stacked retries.

CTA

Timeouts without keys are how agents earn a reputation for double-billing. Wire runtime idempotency before you widen write autonomy — /agentic · /contact?intent=agentic-pilot.

FAQ

What questions does this article answer?

How is this different from n8n idempotency keys?
n8n idempotency keys dedupe workflow executions and webhook redeliveries inside an automation graph. Agent idempotency keys dedupe *tool writes* across model loops, harness resumes, and HTTP clients. Same idea, different boundary: the tool adapter, not the workflow trigger. If both run in one system, they still need distinct keys — a workflow retry and a tool retry are not the same intent.
Read tools vs write tools — retry rules?
Reads can usually retry freely; writes need a stable key and a ledger or native idempotency before any retry. Irreversible writes should escalate or use a single-consumer outbox when status is unknown after timeout. Classify the tool at registration time so the harness cannot “just retry” a wire transfer because the model said the tool failed.
Where should the key be born — model or runtime?
Runtime. Keys born in the model get rotated on every re-emit after a timeout, which causes the duplicates you are trying to prevent. Inject and overwrite at the harness boundary. If the schema lets the model pass `idempotency_key`, treat that field as hostile input.
What if the upstream API has no Idempotency-Key header?
Use a client-supplied unique field if the API has one, or claim the key in your own ledger before the call and skip or replay from stored receipts. Some vendors rename the header (`PayPal-Request-Id`, Square’s body `idempotency_key`). Map your runtime key into whatever they actually honor. Do not invent a header the vendor ignores.
How do compensating actions stay idempotent?
Derive a compensation key from the original key plus action name, refuse to compensate while the original is still pending, and record compensation on the same run trace so you never void twice. If the original was a Stripe `500`, wait for retrieve-or-webhook before you refund.
What ledger fields belong on the trace?
At minimum: `idempotency_key`, `intent_fingerprint`, `status`, `attempt`, `upstream_receipt_id`, timestamps, and `run_id`. Add `idempotent_replayed` and `metadata_local_id` when the vendor is Stripe. Those fields let ops prove one logical write across stacked retries.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot