Idempotent Agent Tool Writes: Retries Without Double Emails or Double Charges
Timeouts are unknown, not failed. Mint one runtime idempotency key per write intent and reuse it across retries so you cannot double-email or double-charge.
William Spurlock Founder — Spurlock Studios Updated 16 MIN
When an agent tool write times out, the dangerous question is not “did the model fail?” — it is “did the side effect already land?” Make writes safe by minting a stable idempotency key in the runtime before any retry layer can fire, then reusing that same key for model retries, harness retries, and HTTP client retries.
This spoke sits inside the Agentic Systems Operating Manual. It owns keys so retries do not double-write. Argument shape lives in tool schemas agents follow. Packaging lives on /agentic.
The short answer
- Timeouts are ambiguous: the upstream may have committed while your agent saw a network error. Stripe’s own essay names the three failure cuts — connect never happens, work dies mid-flight, work finishes and the response never returns.
- Agents stack retries (model loop + harness + HTTP). One timed-out write can become two charges or two emails if each layer invents a new identity.
- Birth the key in the runtime from
(run_id, tool_name, intent_fingerprint), not in the prompt and not in the HTTP client. - Prefer a native
Idempotency-Keyheader when the API supports it. Stripe documents the header, a 255-character cap, and a 24-hour v1 replay window. - Without a header, claim the key in a local ledger (or a client-supplied
external_id) before the write. No claim, no HTTP. - Test duplicate delivery in staging with forced timeouts before you grant write autonomy.
What is the agent-specific idempotency failure mode?
Workflow automation usually has one retry owner (the workflow engine). Agents have three:
| Layer | What retries | Typical trigger | How it double-writes |
|---|---|---|---|
| Model loop | “Tool failed, try again” | Tool error text in the next turn | New tool call, often a new UUID in args |
| Harness / runner | Re-invoke the act step | Timeout, crash, checkpoint resume | Same intent, new HTTP attempt, new key if the runtime forgot |
| HTTP client | Same request again | 408 / 429 / 5xx, connection reset | Safe only if the header/key is stable |
If each layer invents its own “try again” without a shared key, a single ambiguous timeout becomes stacked side effects. That is the agent-specific failure — not “webhooks can fire twice,” but “three systems each think they are being helpful.”
HTTP semantics already treat GET, PUT, and DELETE as idempotent methods. Agent tools are usually POST-shaped: create a charge, send a message, open a ticket. Those are the verbs that need a key.
Why “only call once” in the prompt fails
Prompts do not control TCP. A model that obediently “calls send_email once” still loses when:
- The HTTP call hangs past the client timeout after the provider accepted the message.
- The harness resumes the run after a deploy and re-enters
state:act. - The model sees a generic
timeoutstring and emits a second tool call with slightly different arguments — and a newmessage_idit invented.
Instructional discipline is not a transport guarantee. Treat “call once” as documentation for humans, not as a safety control.
If the tool schema exposes idempotency_key as a model-writable string, you made the failure easier. The model will “helpfully” rotate the field after a timeout. Strip or overwrite that property at the adapter. The schema can document that a key exists; it must not let the model own it.
Checklist for the prompt-and-schema layer:
- System prompt does not ask the model to generate UUIDs for writes
- Tool schema does not list
idempotency_keyas a required model argument - Adapter injects the runtime key after validation, before HTTP
- Tool error text distinguishes
timeout/unknownfromvalidation_failed
How do I make agent tool writes safe when the call times out?
Procedure that holds up in production:
- Classify the tool as
read,write_idempotent, orwrite_irreversiblebefore registration. - Mint a key in the runtime when the act step decides to call a write tool — before the HTTP request starts.
- Persist
key → status(pending|succeeded|failed_poison) in a ledger keyed by tenant. - Pass the same key into every retry of that logical write: harness replay, HTTP retry, and any model re-emit for the same intent.
- On timeout: leave status
pending(orunknown), do not mint a new key, and either poll for receipt or escalate — never “just send again” with a fresh identity. - On success: store the upstream receipt id beside the key; mark
succeeded. - On definitive failure (4xx that will not succeed on retry): mark
failed_poisonso retries stop.
Timeout means unknown. Unknown means reuse the key or escalate — never invent a second write identity.
| Observed outcome | Status to store | Next action |
|---|---|---|
| HTTP 2xx + receipt id | succeeded | Return cached receipt on later claims |
| Client timeout / connection reset | pending / unknown | Reuse key; poll or escalate |
| Hard 4xx (bad args, declined card you will not resubmit) | failed_poison | Stop; new intent needs a new key |
| 409 / “already exists” with a receipt | succeeded | Treat as the first write landing |
| 500 with no receipt | unknown | Reuse key; do not rotate |
How do I generate stable keys across retry layers?
Key material should be stable for the business intent, not for the HTTP attempt:
key = hash(tenant_id + run_id + tool_name + intent_fingerprint)
intent_fingerprint is a canonical hash of the fields that define the side effect (to, template_id, invoice_id, amount_cents) — not of ephemeral fields like requested_at or random UUIDs the model invents.
Stripe’s API reference suggests V4 UUIDs or another high-entropy string, caps keys at 255 characters, and tells you not to put email addresses or other personal identifiers in the key. A SHA-256 hex of the tuple above stays under that cap and stays free of PII. If you prefer a UUID, mint it once in the runtime and persist it next to the intent — do not mint a UUID per HTTP attempt.
| Source of key | Safe? | Why |
|---|---|---|
| Model-generated UUID in tool args | No | New UUID on every re-emit |
HTTP attempt id / x-request-id per try | No | New per transport retry |
Runtime: run_id + tool + intent hash | Yes | Survives all three layers |
| Upstream event id (when writing because of an event) | Yes | Aligns with business identity |
| Shopping-cart / invoice id (Stripe’s second strategy) | Yes | Stripe names this as a way to block double submit |
The IETF Idempotency-Key draft (revision 07, 15 October 2025; the listed expiry was 18 April 2026, and as of August 2026 it is still a draft, not an RFC) says the same thing in standards language: the client generates a unique key, must not reuse it with a different payload, and the server may fingerprint the body. Stripe is listed in that draft’s implementation status. Until it is an RFC, treat vendor docs as the contract you actually get.
Where the key is born — model vs runtime
| Birthplace | Outcome |
|---|---|
Model fills idempotency_key | Model invents a new key after timeout; duplicates ship |
HTTP library mints a UUID on each request() | Harness resume looks like a new charge |
| Runtime injects key into the tool-call envelope | Retries reuse identity even if the model rephrases args |
| Runtime + schema forbids model override | Strongest: model cannot “helpfully” rotate the key |
Default: the harness owns the field. If the tool schema exposes idempotency_key, strip or overwrite model-supplied values before dispatch.
This is the same reason tool schemas should keep write-path enums tight. A model that can invent message_id or charge_ref has a second way to break identity even when your header is correct.
Procedure at the adapter:
- Validate args against the schema (types, enums, required).
- Compute
intent_fingerprintfrom the canonical write fields. - Load or insert
idempotency_keyfrom the ledger for(run_id, tool_name, intent_fingerprint). - Overwrite any model-supplied key.
- Dispatch HTTP with that key in the header or
external_idfield.
What does Stripe’s Idempotency-Key actually guarantee?
Copy Stripe’s contract into the adapter. Do not invent a friendlier one.
Idempotent requests (API v1):
- Send
Idempotency-KeyonPOST.GETandDELETEignore it; those verbs are already idempotent. - Stripe saves the status code and body of the first request for that key — including
500errors — and replays that result. - Keys are removed after they are at least 24 hours old. Reusing a pruned key starts a new request.
- Incoming parameters are compared to the original. A mismatch errors so you cannot accidentally reuse a key for a different charge.
- Results are saved only after endpoint execution begins. Validation failures and some concurrent conflicts are not stored, and you may retry those.
API v2 changes the replay rules:
| Rule | API v1 | API v2 |
|---|---|---|
| Verbs that accept keys | POST only | POST and DELETE |
| Replay window | ≥ 24 hours | 30 days, same account or sandbox, same API |
| Failed-request replay | Returns the cached first result, including 5xx | Re-executes the failed work without extra side effects, or explains why it cannot |
| If you omit a key | You own the risk | Stripe mints a UUID for you — useless for agent retries unless you persist it |
If the SDK mints a key you never stored, the next harness resume cannot replay. Always set the key yourself from the ledger.
Low-level error handling adds two headers worth wiring into the tool span:
| Header | Meaning |
|---|---|
Idempotent-Replayed: true | This response is the stored first result, not a second charge |
Stripe-Should-Retry: true | Retry this request (still back off) |
Stripe-Should-Retry: false | Do not retry; another attempt will not help |
| Header absent | Fall back to status-code rules |
A replayed 200 with Idempotent-Replayed: true is success. Log it as a retry hit, not as a second write.
How should I retry Stripe 4xx, 5xx, and timeouts?
Do not use one retry policy for every status. Stripe splits errors into content, network, and server — and the idempotency layer does not treat them the same.
| Class | Examples | Same key? | New key? | Agent action |
|---|---|---|---|---|
| Network / timeout | Socket reset, read timeout, no body | Yes | Never | Retry until you have a Stripe result |
| Content 4xx you will fix | Missing param, bad enum | Only if you did not change the body | Yes, after you change params | Stripe caches a 400 once the method started; a fixed body needs a new key |
429 rate limit | Too many requests | Maybe | Prefer wait, same key | Stripe notes the limiter can run before the idempotency layer, so the same key may not replay |
401 missing key | Auth header omitted | Do not trust cache | After you attach credentials | Same caveat as 429 |
409 conflict | Concurrent request with the same key | Yes | No | Wait; this is the in-flight case the IETF draft also maps to 409 |
500 / 502 / 503 / 504 | Stripe-side failure | Yes | No | Treat as indeterminate. A new key can double-charge if the first attempt already hit the card network |
On 500, Stripe says they will try to reconcile — roll forward if the payment network already saw the charge, roll back if not — and may fire a webhook later for an object you never saw on the API response. That is why you send a local identifier in metadata on create: the webhook can carry it even when the HTTP client timed out.
Agent mapping:
- Timeout or network error → reuse key, exponential backoff with jitter (Stripe’s 2017 write-up is still the clean description of backoff + jitter vs a thundering herd).
Stripe-Should-Retry: true→ same key, wait.Stripe-Should-Retry: false→ stop; escalate.500→ same key, markunknown, watch webhooks / retrieve by metadata. Do not open a second PaymentIntent identity.- Validation
400you will correct → new key only after the body changes, and only if you are sure the first attempt never executed.
The HTTP method table in RFC 9110 will not save you here. A Stripe POST /v1/payment_intents is not a PUT. The header is the whole safety story.
What if the upstream API has no Idempotency-Key header?
Many CRM, email, and internal APIs do not speak Stripe-style idempotency. Some vendors use a different name for the same idea. The IETF draft’s implementation list is a useful map: Adyen uses Idempotency-Key; PayPal uses PayPal-Request-Id; Twilio webhooks use I-Twilio-Idempotency-Token; Square puts idempotency_key in the JSON body. Read the vendor page. Do not send a header they ignore and call the job done.
Options, in order of preference:
- Native unique constraint — if the API accepts a client-supplied external id (
external_id,reference,client_ref,PayPal-Request-Id), use your runtime key there. - Pre-write ledger gate — before calling the API, claim the key in your DB with a unique index. If claim fails because status is
succeeded, return the stored receipt and skip the call. Ifpendingand younger than TTL, wait or poll. If older than TTL, escalate. - Read-before-write with a stable lookup — only when the domain has a natural unique query (invoice already paid, ticket already has comment hash X). Fragile; document the race.
- Outbox + single worker — enqueue the write once; a single consumer performs the HTTP call. Agent retries enqueue the same outbox id.
| API behavior | What you do |
|---|---|
Speaks Idempotency-Key (Stripe, Adyen, others) | Send the runtime key; still keep a local ledger for timeouts |
| Speaks a renamed header / body field | Map the same key into that field |
Offers external_id uniqueness | Put the key there; handle the duplicate error as success-with-receipt |
| Offers nothing | Ledger claim is the control. No claim, no HTTP |
Do not pretend a header exists. Build the ledger. Workflow graphs have the same storage problem; agents need the idea on the tool boundary, not on the webhook trigger.
Read tools vs write tools — retry rules
| Tool class | Retry on timeout? | Key required? |
|---|---|---|
| Read / search | Yes, usually safe | Optional (cache key helps) |
| Write with server idempotency | Yes, same key | Required |
| Write without server idempotency | Only after ledger claim or escalate | Required locally |
| Irreversible external (wire, legal notice) | Human or outbox only | Required + approval |
RFC 9110 is why reads feel easy: a repeated GET is supposed to leave server state alone. Your “search_contacts” tool is a GET in spirit even if the wrapper is POST. Your “send_invoice_reminder” tool is not.
Checklist before marking a tool write_idempotent in the registry:
- Side-effect class documented (
email,charge,crm_write,ticket) - Key birthplace = runtime
- Ledger or native unique field wired
- Timeout path leaves status
unknown/pending, notfailed - Stripe (or vendor) error classes mapped — especially
500and429 - Forced duplicate-delivery test exists
- Metadata / correlation id set so a late webhook can join the run
If any box is open, the tool is write_irreversible until you close it.
Compensating actions that stay idempotent
Compensations (void charge, send apology, delete draft) are also writes. They need their own keys, derived from the original:
compensate_key = hash(original_key + ":compensate:" + action)
Rules:
- Never compensate twice for the same original key.
- Never compensate if the original write status is still
pending— resolve unknown first. A void fired against a charge that never landed is how you get a later “success” webhook and a confused ledger. - Log compensation under the same
run_idwith a distinct tool span. - If the original was a Stripe
500, wait for the webhook or a retrieve-by-metadata before you refund. Refunding a charge that Stripe rolled back is a second mess.
| Original status | Compensate? |
|---|---|
succeeded + receipt | Yes, with compensate_key |
pending / unknown | No — resolve first |
failed_poison | No — nothing to undo |
succeeded and already compensated | Return cached compensation receipt |
Blind “undo” loops are how you get a charge, a void, and a second charge.
Failure example: double invoice email
Job: Agent drafts and sends invoice reminder for inv_8841.
What happened:
- Runtime minted no key; model called
email.send. - Provider accepted the message; client timed out at 30s.
- Harness retried
state:act. Model calledemail.sendagain with a newmessage_idit invented. - Customer received two reminders; support spent a day on “your system is broken.”
Cost: trust, not just SMTP fees.
Fix:
- Runtime key:
hash(tenant + run + email.send + inv_8841 + template_reminder_v2) - Ledger claim before SMTP
- On timeout, poll provider by key/metadata or escalate — do not re-emit with a new message id
The charge version of the same story is worse. Swap email.send for stripe.payment_intents.create, skip the header, and a 30-second timeout becomes two PaymentIntents. Stripe’s docs exist because that exact network cut is normal, not exotic.
I have watched this class of bug in production agent loops across a book of 500+ automations and 20,000+ hours on agentic systems. The model is rarely the interesting part. The missing key is.
How do I test duplicate delivery before production?
Staging drills that catch the stacked-retry bug:
- Inject latency past the HTTP timeout after the mock server records the write.
- Confirm harness retry reuses the same key and the mock sees one logical commit.
- Force model re-emit by returning a fake timeout string once; assert the second tool call carries the injected key (or is blocked).
- Crash mid-pending and resume from checkpoint; assert no second charge.
- Poison 409 / duplicate from upstream; assert the agent treats it as success-with-receipt, not endless retry.
- Stripe-shaped 500 from the mock, then a webhook with your metadata; assert the ledger flips to
succeededwithout a second POST.
| Drill | Pass criterion |
|---|---|
| Slow success + client timeout | Exactly one side effect |
| Double harness resume | Ledger blocks second HTTP |
| Model invents new args, same intent | Same key; one effect |
| Upstream duplicate error | Maps to succeeded |
Mock 500 + later webhook | One object; ledger joins on metadata |
| Same key, different amount | Adapter refuses; Stripe would error the mismatch |
If you have not run the timeout drill, you have not tested agent writes.
Ledger fields that belong on the trace
Put these on the tool span and in the ledger row:
| Field | Purpose |
|---|---|
idempotency_key | Shared identity across retries |
intent_fingerprint | Prove which args defined the write |
status | pending / succeeded / failed_poison / unknown |
attempt | Transport attempt count (not a new key) |
upstream_receipt_id | Correlate to CRM / email / PaymentIntent |
idempotent_replayed | Stripe (or vendor) said this body was a replay |
stripe_should_retry | Vendor retry hint, when present |
first_seen_at / succeeded_at | Dispute timeline |
run_id / tool_call_id | Join to the agent trace |
metadata_local_id | The id you sent so a late webhook can find you |
Without receipt correlation, ops cannot answer “which run sent the second email?”
Anti-patterns
UUID in the prompt template. Guarantees uniqueness per emit — the opposite of idempotency.
Retrying irreversible tools on any error string. Distinguish timeout / unknown from validation_failed.
Per-layer keys. Model key ≠ harness key ≠ HTTP key means three charges.
Deleting ledger rows on failure. If the write may have landed, keep the key until you know.
Treating HTTP 200 as the only success. Some APIs return errors after committing; prefer receipt ids. Stripe’s Idempotent-Replayed: true on a cached body is success too.
Minting a new Stripe key after a 500. Stripe advises against it because the original key may already have produced side effects.
Letting the official SDK “just retry.” Libraries can attach keys and back off — Stripe says you must configure that. If the SDK mints a key per process start, your harness resume is a new customer.
Hashing the entire request body as the key. Two legitimate sends with the same body (two invoices, same amount, different customers) collapse. Fingerprint the intent fields, then store a unique key for that intent.
Decision list: ship write autonomy?
Ship autonomous writes only when all are true:
- Tool is classified and keyed in the runtime.
- Upstream supports idempotency or a local ledger gate is live.
- Timeout drill passed in staging.
500/ webhook path is defined for money-moving tools.- Evaluator or policy gate can block high-risk tools (operating manual).
- Kill switch can freeze the write tool class without redeploying prompts.
If any box is open, keep the tool behind human approval or an outbox.
Worked ledger claim (pseudo)
claim(key):
insert ledger(key, status=pending) on conflict do nothing
if conflict and status=succeeded: return cached_receipt
if conflict and status=pending and age < TTL: wait or escalate
if conflict and status=pending and age >= TTL: escalate unknown
if inserted: call upstream with Idempotency-Key=key (or external_id=key)
on success: status=succeeded, store receipt
on timeout: leave pending, schedule resolve job
on hard 4xx: status=failed_poison
on 500: leave unknown, attach metadata_local_id, watch webhook
Agents call claim through the tool adapter — never raw HTTP from the model. Hash intent fields (to, template_id, invoice_id); do not put raw customer bodies into the key string. Stripe’s 255-character limit and “no PII in the key” rule are the right constraints even when the vendor is not Stripe.
TTL should outlive your retry window and sit inside the vendor’s replay window. For Stripe v1 that window is 24 hours; for v2 it is 30 days. A ledger TTL of a few hours is enough for agent retries and still forces a human on a key that stayed pending overnight.
The resolve job is not optional for money tools. After a timeout it should:
- Re-GET or retrieve-by-metadata using the local id you stored.
- If a receipt exists, mark
succeededand return it — do not POST again. - If nothing exists and the vendor key is still inside the replay window, POST once more with the same
Idempotency-Key. - If the window is gone and status is still unknown, page a human. A pruned Stripe v1 key is a new request if you reuse it.
That job is how you turn “the client hung up” into a closed ledger row instead of a second charge.
Pilot minimum
A Spurlock Studios $1,500 · 5-day agentic pilot that includes write tools ships: runtime key injection, a thin ledger, timeout classification, Stripe-header (or vendor-field) mapping, and one forced-duplicate drill in staging — not a promise that “the model will be careful.”
/agentic · /contact?intent=agentic-pilot
FAQ
How is this different from n8n idempotency keys?
n8n idempotency keys dedupe workflow executions and webhook redeliveries inside an automation graph. Agent idempotency keys dedupe tool writes across model loops, harness resumes, and HTTP clients. Same idea, different boundary: the tool adapter, not the workflow trigger. If both run in one system, they still need distinct keys — a workflow retry and a tool retry are not the same intent.
Read tools vs write tools — retry rules?
Reads can usually retry freely; writes need a stable key and a ledger or native idempotency before any retry. Irreversible writes should escalate or use a single-consumer outbox when status is unknown after timeout. Classify the tool at registration time so the harness cannot “just retry” a wire transfer because the model said the tool failed.
Where should the key be born — model or runtime?
Runtime. Keys born in the model get rotated on every re-emit after a timeout, which causes the duplicates you are trying to prevent. Inject and overwrite at the harness boundary. If the schema lets the model pass idempotency_key, treat that field as hostile input.
What if the upstream API has no Idempotency-Key header?
Use a client-supplied unique field if the API has one, or claim the key in your own ledger before the call and skip or replay from stored receipts. Some vendors rename the header (PayPal-Request-Id, Square’s body idempotency_key). Map your runtime key into whatever they actually honor. Do not invent a header the vendor ignores.
How do compensating actions stay idempotent?
Derive a compensation key from the original key plus action name, refuse to compensate while the original is still pending, and record compensation on the same run trace so you never void twice. If the original was a Stripe 500, wait for retrieve-or-webhook before you refund.
What ledger fields belong on the trace?
At minimum: idempotency_key, intent_fingerprint, status, attempt, upstream_receipt_id, timestamps, and run_id. Add idempotent_replayed and metadata_local_id when the vendor is Stripe. Those fields let ops prove one logical write across stacked retries.
CTA
Timeouts without keys are how agents earn a reputation for double-billing. Wire runtime idempotency before you widen write autonomy — /agentic · /contact?intent=agentic-pilot.
What questions does this article answer?
- How is this different from n8n idempotency keys?
- n8n idempotency keys dedupe workflow executions and webhook redeliveries inside an automation graph. Agent idempotency keys dedupe *tool writes* across model loops, harness resumes, and HTTP clients. Same idea, different boundary: the tool adapter, not the workflow trigger. If both run in one system, they still need distinct keys — a workflow retry and a tool retry are not the same intent.
- Read tools vs write tools — retry rules?
- Reads can usually retry freely; writes need a stable key and a ledger or native idempotency before any retry. Irreversible writes should escalate or use a single-consumer outbox when status is unknown after timeout. Classify the tool at registration time so the harness cannot “just retry” a wire transfer because the model said the tool failed.
- Where should the key be born — model or runtime?
- Runtime. Keys born in the model get rotated on every re-emit after a timeout, which causes the duplicates you are trying to prevent. Inject and overwrite at the harness boundary. If the schema lets the model pass `idempotency_key`, treat that field as hostile input.
- What if the upstream API has no Idempotency-Key header?
- Use a client-supplied unique field if the API has one, or claim the key in your own ledger before the call and skip or replay from stored receipts. Some vendors rename the header (`PayPal-Request-Id`, Square’s body `idempotency_key`). Map your runtime key into whatever they actually honor. Do not invent a header the vendor ignores.
- How do compensating actions stay idempotent?
- Derive a compensation key from the original key plus action name, refuse to compensate while the original is still pending, and record compensation on the same run trace so you never void twice. If the original was a Stripe `500`, wait for retrieve-or-webhook before you refund.
- What ledger fields belong on the trace?
- At minimum: `idempotency_key`, `intent_fingerprint`, `status`, `attempt`, `upstream_receipt_id`, timestamps, and `run_id`. Add `idempotent_replayed` and `metadata_local_id` when the vendor is Stripe. Those fields let ops prove one logical write across stacked retries.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
AI Agents Why Agents Loop on Failed Tools: No-Progress Detection Beats Longer Prompts
Agents loop on failed tools because the harness never detects no-progress. Fingerprint calls, honor retryable:false, cap turns, terminate with a reason code.
AI Agents Agentic Systems: An Operating Manual for Multi-Agent Work That Ships
An agentic system is evaluators, policy gates, sandboxes, and kill switches — not a chat window — so production tool work survives contact with real data.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.