Spurlock Studios
Contact
Share LinkedIn X
A scuffed work smartphone with a blank glowing circular button. Thesis: DURABLE AGENT RUNTIMES SURVIVE RESTARTS.

A long-running agent that survives crashes and waits for humans needs a durable runtime: serializable control state outside the process, a resume path that continues from the last safe boundary, and tool writes that stay safe when that boundary re-runs. That is not “memory.” Memory is what the agent recalls. Durability is whether the run still exists after the host dies.

This spoke sits under the Agentic Systems Operating Manual. Pair it with state machines for agent loops for the legal states the loop may occupy. Durability is the substrate that keeps those states alive across restarts. The agentic lane is where we prove that substrate on one real job.

The short answer

  • Durability ≠ memory. Memory stores facts and conversation. Durability stores execution progress so a new process can resume the same job.
  • In-process loops die under real ops. Deploys, OOM kills, Durable Object eviction, and overnight approvals all outlive a single Node or Python process.
  • Resume re-enters a boundary. Snapshot engines, event-history engines, and fiber/stash engines all re-run the last unfinished step. Design for that, or you double-send.
  • HITL pauses must be durable. A pause that lives only in RAM is a memory leak with a polite name.
  • Measure resume correctness, not “it remembered.” After a kill, does the same run id continue without a second email or refund?

What is a durable agent runtime (and what is not)?

A durable agent runtime is the store plus the resume protocol that lets a new process continue a named run. If you cannot kill the worker mid-approval and still resume the same job tomorrow without inventing a new run id, you do not have durability. You have a longer prompt.

LangGraph’s own persistence docs split this cleanly: a checkpointer tracks thread-scoped graph progress (fault tolerance, interrupts, time travel); a store tracks cross-thread facts (preferences, knowledge). The first is a runtime concern. The second is memory. Teams collapse them and then wonder why a vector index did not survive a deploy.

ConcernDurable runtimeChat / agent “memory”
Question it answersWhere was this run when the process died?What facts should the model see next turn?
Typical storeCheckpoints, event history, DO SQLite, workflow journalVector store, profile rows, transcript slices
Success metricResume correctness, no duplicate side effectsAnswer quality, fewer re-asks
Survives a 3-day human wait?RequiredOptional
Survives process death mid-tool?RequiredIrrelevant

If you only keep a longer prompt history, you have memory. If you can evict the worker and still resume the same handle, you have durability.

Why do in-process agent loops die under real ops?

A demo agent is a while loop in one process. Production invents ways to murder that process, and none of them wait for the model to finish a thought.

  1. Rolling deploys mid-tool-call
  2. Platform eviction (Cloudflare Durable Objects typically idle-evict after ~70–140 seconds without keep-alive, alarms, or inbound events — see the Durable Object lifecycle)
  3. Spot instances and scale-to-zero on serverless
  4. Operator restart after a bad release
  5. Human approval that arrives after the original HTTP request is gone
  6. Uncaught exception that tears down in-memory state while the disk or SQLite row is still honest

The failure we see constantly after 500+ automations and 20,000+ hours on agentic systems: the agent emailed the customer, the process died before writing “done,” and a retry emailed again. The model did nothing wrong. The runtime treated “in memory” as “committed.”

KillWhat diesWhat must survive
Deploy / pod restartProcess, RAM, open socketsRun handle + last committed boundary
DO hibernate / evictIn-memory vars, local timers, open fetchessetState / SQL / fiber stash / alarm
Overnight HITLOriginal request threadPause record + approval binding
OOM / crashEverything in the heapWhatever you flushed before the kill

Cloudflare’s rules of Durable Objects say the quiet part: in-memory state is not preserved on eviction or an uncaught exception. Persist first. Then do the thing that might fail.

What actually gets persisted on a crash-resume path?

Three records have to exist before you call the run durable. Miss one and you will reconstruct a story from logs instead of resuming a job.

RecordJobIf missing
Run handleStable id (thread_id, fiber id, workflow id)Resume invents a sibling run
Boundary snapshot or journalLast safe step + inputs/resultsResume rewinds or skips
Side-effect ledgerIdempotency keys + receipts for writesCorrect resume becomes a duplicate charge

Checklist for the ledger you actually ship:

  • One durable handle per job, minted before the first tool write
  • Boundary writes are committed before the next side effect starts
  • Tool receipts are keyed by that handle, not by “whatever the model said last”
  • Secrets are pointers, not blobs in the checkpoint
  • The UI loads the handle, not “the latest conversation”

Webhook retries are a crash with better manners. The provider will POST again. If accept is not durable, you start a second run that looks like recovery. Cloudflare startFiber() exists for that boundary: persist acceptance and an idempotency key before work begins, then return status on the retry. Same rule on any host — mint the handle on the first byte you acknowledge.

Procedure for minting the handle:

  1. Receive the job (HTTP, queue, email, cron).
  2. Derive or generate run_id from an idempotency key the caller already sent, or from a key you persist before ack.
  3. Write the handle row in the same transaction as “accepted.”
  4. Only then call a model or a write tool.
  5. On a duplicate delivery, return the existing handle. Do not create a sibling.

Restate’s key concepts name the same contract in journal language: record the step and its result; on crash, replay the journal and skip completed work. Temporal’s event history is the same idea with a different product name. The product is optional. The three records are not.

Snapshot checkpoints vs event-history replay — what is the difference?

Crash-resume engines do not all restore the same way. Confusing the mechanism is how you write a “resume” that is actually a replay of the entire graph with half the side effects live.

Snapshot / checkpoint. Persist graph or entity state at a named boundary. A new process loads the blob and continues from that super-step. LangGraph checkpointers do this at super-step boundaries. Cloudflare fibers do a close cousin: stash() a snapshot, recover via onFiberRecovered.

Event-history / journal replay. Persist an append-only log of decisions and activity results. A new worker re-runs the control code from the start, substituting recorded results for work that already happened. Temporal replays from event history rather than restoring a heap snapshot. Azure Durable Functions use the same event-sourcing replay. Restate’s request lifecycle restarts the handler and skips journaled steps.

MechanismRestore methodWhat re-runsTypical failure if you ignore it
Snapshot checkpointLoad last state blobThe interrupted node / fiber from its startSide effects before the interrupt fire twice
Event-history replayRe-execute control code; reuse recorded resultsControl flow (must stay deterministic)Non-deterministic clocks/IDs fork the history
Fiber stashLoad last stash(), then your recover hookWhatever you did not stashRecover starts the job from zero

This is not a vendor bake-off. It is the difference between “load the save file” and “re-read the campaign log.” Both can be correct. Both will punish you if you treat RAM as the source of truth.

LangGraph durability modes make the snapshot tradeoff explicit. From the current checkpointer docs, least to most durable:

ModeWhen it writesCrash mid-graph
"exit"Only when the graph exits (success, error, or interrupt)Intermediate steps are gone
"async"In the background while the next step runsSmall window where the checkpoint never lands
"sync"Before the next step startsHighest durability, more write cost

"exit" is faster for long graphs and does not protect mid-execution crashes the way "sync" does. If you say “we checkpoint,” name the mode.

Event-history engines add a second tax: control code must be deterministic, because it will run again. Azure’s orchestrator constraints and Temporal’s replay rules are the same list in different SDKs.

In control codeSafeUnsafe
TimeWorkflow/orchestrator clock APIsDate.now(), wall-clock sleeps
IDsJournaled / history-recorded IDsFresh UUID on every replay
I/OActivities / ctx.run / tool adaptersDirect HTTP from the replayed function
BranchingBased on recorded resultsBased on “what the model might say this time” inside control

If you need a non-deterministic call — LLM, HTTP, Math.random — put it behind a recorded boundary. Replay then returns the first result. Doing it inline is how a recovered run takes a different fork and the history engine throws a non-determinism error, or worse, silently diverges on a snapshot host.

Why does resume re-enter the last step?

Because almost every durable engine treats a boundary as a restartable unit, not a line number in your source file.

LangGraph interrupts are the clearest public contract: on Command(resume=...), the node restarts from the beginning. Code before interrupt() runs again. The interrupt call then returns the human’s value instead of pausing. That is intentional, not a bug.

Event-history engines do the same at a different grain. Temporal and Azure Durable Functions re-execute the workflow or orchestrator from the top and fill awaits from history. Activities that already completed return the recorded result. Activities that did not complete may be scheduled again. Cloudflare fibers call your recover hook with the last stash; they do not rewind your JavaScript to the exact line.

Engine familyResume grainWhat you must make safe
Graph + interruptWhole nodeAnything before interrupt()
Event historyWhole workflow functionNon-deterministic control code; unrecorded side effects
Fiber / stashFrom last stashWork after the last stash()

Procedure that keeps resume from becoming a second send:

  1. Split “prepare” and “write” into separate boundaries.
  2. Mint the idempotency key in the prepare boundary and persist it.
  3. On the write boundary, send that same key. Never generate a new UUID on resume.
  4. Treat logs before an interrupt as expected duplicates unless you gate them on a committed flag.

If you need the legal states around that split, that is the state machine spoke. This post only owns the fact that resume will walk back into the last room.

How do human-in-the-loop pauses stay durable?

A durable HITL pause is three parts. Skip any one and the human is approving a ghost.

  1. Persist the run at a named boundary (checkpoint, fiber stash, workflow wait, or alarm).
  2. Return a handle the UI or ops channel can load (thread_id, fiber id, workflow id).
  3. Resume by feeding the decision into that handle — not by starting a new chat.

Cloudflare’s long-running agents docs are blunt: agents are durable identities, not always-on processes. State, SQL, schedules, and fiber checkpoints survive hibernation. In-memory variables, timers, open fetches, and local closures do not. A “wait for the manager” that is a setTimeout in RAM is already dead.

Durable Object alarms are the wake primitive when the human may take hours: at-least-once alarm() delivery, exponential backoff on throw, and a hard cap of six automatic retries unless you reschedule. Design the escalate path as if the human never answers.

Checklist:

  • Pause state is in durable storage, not the request thread
  • Approval payload is bound to that run id plus a payload hash (no “approve whatever is latest”)
  • Resume path is exercised in staging with a process kill mid-wait
  • Tool nodes after resume are idempotent
  • Timeout / escalate path exists if the human never answers
  • The original HTTP connection is treated as already gone

If the human waits three days, the original request is archaeology. Only the durable handle matters.

ClockWhat must be true
T0 — pauseBoundary flushed; handle returned; approval UI bound to run_id + payload hash
T1 — host diesIn-memory queue is gone; SQLite / checkpoint / journal / alarm still hold the wait
T2 — human actsResume loads T0 payload, not “latest ticket”; decision written on the same handle
T3 — writeRefund/email tool uses the key minted at T0, not a key minted at T2
T4 — silenceEscalate alarm fires; the run moves to a named escalate state, not a new chat

A Slack message with a raw “approve” button and no handle is T0 skipped. Everything after it is folklore.

What happens when the host hibernates or evicts the agent?

Hibernation is a first-class kill, not an edge case for people who picked Cloudflare. Any scale-to-zero host will do a version of this. Cloudflare just publishes the clocks.

From the Durable Object lifecycle:

StateWhat is true
Active, in-memoryHandling a request or event
Idle, hibernateableAfter ~10 seconds of inactivity, the runtime may hibernate
HibernatedRemoved from memory; hibernated WebSockets can stay connected
Evicted / inactiveAfter ~70–140 seconds idle without keep-alive conditions; next wake is a cold start

Cloudflare Agents durable execution draws the line teams miss: keepAlive() reduces the chance of eviction. runFiber() makes eviction survivable by registering the work in SQLite, letting you stash(), and calling onFiberRecovered on the next activation. startFiber() adds durable acceptance, an idempotency key, and retained status for callers that retry (webhooks).

PrimitiveWhat it buysWhat it does not buy
keepAlive() / keepAliveWhile()Fewer evictions during active LLM/tool workRecovery if eviction still happens
runFiber() / stash()Crash-recoverable work + recover hookAutomatic replay of your JS from line 47
startFiber()Durable accept + inspect + cancel + dedupeA free pass on tool idempotency
schedule() / alarmsWake later for HITL and pollsIn-memory closures from the previous activation

keepAlive without a fiber is how “it worked in staging” becomes “lost the job overnight.” Fibers without idempotent writes are how a recovered job double-books.

A recover hook is only as good as the last stash. Write a small, typed snapshot — not “whatever was in RAM.”

Field in stash()Why
run_idSo recover cannot attach to a sibling job
stepNamed boundary (drafted, awaiting_approval, refund_sent)
idempotency_keyReused on the next write; never regenerated
payload_hashBinds the human decision to the draft they saw
attemptCaps recover loops the same way a state machine caps revisions

If onFiberRecovered cannot find those fields, fail the fiber and page a human. Guessing the step from the transcript is memory cosplay.

What state should never live only in the prompt?

The prompt may summarize control state. The prompt must not be the only copy. Summarization is how agents re-call a tool that already succeeded, then tell you they “remembered” the ticket.

StateWhere it belongsWhy the prompt is a bad home
Run / thread / fiber idsRuntime ledgerSummarizers drop or invent ids
Tool ledger (what ran, keys, result hashes)Harness DBModel will retry a “failed” success
Auth tokens and IAM scopeSecret store + policy gateCheckpoints leak; prompts get logged
Approval decisionsDurable HITL recordA new chat cannot see a signed decision
Budget / kill-switch countersControl planeSoft numbers in context do not stop a loop
Full raw tool dumpsArtifact store with redactionBlow the window; hide the receipt

Keep these out of the model context as source of truth:

  • Run handle
  • Idempotency keys
  • Approval payload hash
  • Token material
  • Cost and turn counters that trip the kill switch

The model can see a redacted summary. The resume path must not ask the model where the run left off.

How do you measure resume correctness?

“The model recalled the ticket number” is a memory win. It is not a durability win. Run this drill before you argue about model quality, and again after every runtime change.

  1. Start a real job that reaches a HITL pause or a slow tool write
  2. Kill the worker / evict the Durable Object / restart the pod
  3. Resume from the stored handle only — no new chat
  4. Score the outcome against the table below
MetricPassFail
Same run id continuesYesNew run invented
Side effects onceOne email / one chargeDuplicate
State-machine positionSame state as before the killRewound or skipped
Human sees prior contextApproval UI shows the pending payloadEmpty or wrong job
Ledger matches realityReceipt hash equals the upstream object“Done” in RAM, nothing upstream

Monthly is enough for a stable harness. After a durability-mode change, a fiber migration, or a new write tool, run it before the next deploy. Record the drill in the handoff. If you cannot produce the recording, you do not have a durable runtime. You have a story.

Staging drill, minimum:

  • Job reaches a real write or a real HITL pause (not a mocked sleep)
  • Kill is the kill you will actually see (deploy, wrangler restart, pod delete, DO idle-evict)
  • Resume uses only the stored handle — no “start over” button
  • Upstream is checked for a second object, not just your own “done” flag
  • Trace shows the same run_id before and after the kill
  • Recording (screen or log bundle) is attached to the handoff

Skip the drill and you will learn the same facts from a customer. That is a more expensive lab.

How does durability interact with idempotent tool writes?

Durability increases how often a step re-executes after a crash. That makes idempotency mandatory, not optional. A correct resume that re-enters a refund node without a stable key is a second refund with better logging.

Stripe’s idempotent requests are the public contract most payment tools already speak: send an Idempotency-Key, and retries return the original result instead of creating a second object. Keys are kept at least 24 hours. Generate the key before the first attempt and persist it on the run handle. Do not mint a fresh UUID inside the resumed node.

On resume:

  1. Re-enter the node / fiber / activity
  2. Tool layer sees the same idempotency key from the ledger
  3. Upstream returns the original receipt — no second charge
LayerOwnsMust not own
RuntimeRun handle, boundary, key mint + persistThe English of the email
Tool adapterSending the key, mapping receiptsInventing a new key on retry
ModelDrafts and proposed args“I think we already refunded” as control

Treat durability and idempotency as one control loop. The operating manual frames the rest of that stack. This post only owns the resume substrate.

Worked failure: the overnight approval that double-booked

What broke: A support agent drafted a refund and paused for manager approval in an in-memory queue. A deploy restarted the API. The manager approved what looked like a stuck ticket. A new agent run also resumed from a stale Redis key that was never bound to a single run_id. Two refunds left.

Cost: Finance cleanup, a customer-trust hit, and two days of “why agents suck” in Slack. No model regression. A missing handle.

Instead:

  1. Persist the pause under a single durable run_id before the draft is shown
  2. Bind the approval button to that id plus a payload hash
  3. Issue the refund tool with an idempotency key owned by the runtime
  4. Kill-test the pause path before soft-launch — deploy mid-wait, then approve once
ControlBefore (failed)After (holds)
Pause storeIn-memory queueCheckpoint / fiber / workflow wait
Approval target“Latest stuck ticket”run_id + payload hash
Refund keyNonePersisted idempotency key
Kill testNoneStaging drill recorded

Bravery is not a restore strategy.

What is the pilot minimum for crash-resume?

A $1,500 · 5-day agentic pilot does not require Temporal, Restate, or a fiber rewrite on day one. It does require a resume path you can kill.

  1. One durable handle per job (thread_id or equivalent), minted before the first write
  2. One kill-and-resume test recorded in the handoff
  3. Idempotent write tools on the critical path
  4. HITL pause that survives process death, or an explicit “no human wait in v1” written down
  5. A named boundary flush before the side effect, not after
DayCrash-resume proof
1Name the handle, the store, and the one write that must not double
2–3Wire persist → kill → resume on a staging job
4Bind HITL or document why the job has no pause
5Record the drill; list the two kills you still do not survive

Framework fashion is optional. Resume correctness is not. If the drill fails, you do not have a durable agent. You have a demo that remembers things until the host blinks.

Which durability anti-patterns fail first?

These fail in the first real week. They fail in the same order almost every time.

Calling a vector store “our durable agent.” That is memory. It will not resume a run.

In-memory checkpointer in production. Fine for unit tests. Worthless for crashes. LangGraph’s InMemorySaver is the example everyone ships by accident.

HITL as “email the ops channel and hope.” No run handle, no resume, no payload hash. The human approves a vibe.

Checkpointing the entire blob of secrets into Postgres. Redact. Store pointers. Treat the checkpoint as a document an intern can dump.

Assuming the agent is an always-on process. Cloudflare Agents hibernate. Kubernetes pods die. Serverless scales to zero. Design for wake/sleep.

Minting a new idempotency key on resume. The ledger must reuse the key from the prepare boundary. A new UUID is a new charge with a clean conscience.

Scoring “it remembered.” Recalling a ticket number after restart is memory. Same run id, one refund, held state — that is durability.

Anti-patternFirst symptomFix
Vector store as runtimeNew run id after deployAdd a handle + boundary store
In-memory saverEmpty thread after restartDurable checkpointer / SQL / journal
Slack-only HITLDuplicate work after approvalBind UI to run_id
Secrets in checkpointsToken in a support dumpPointers + redaction
New key on resumeDouble charge with matching logsPersist the key at prepare

Durability is not the diagnosis when the model picks the wrong tool, the evaluator is mush, or the state machine has no escalate. Those are other spokes. This one only answers: after the host dies, does the same run continue without a second write?

SymptomFirst place to look
New run id after deployHandle + store
Same run, second refundIdempotency ledger
Same run, wrong next stateState machine, not the runtime
Recalled the ticket, lost the pauseYou built memory and called it durable

FAQ

What is a durable agent runtime versus memory?

A durable runtime stores execution progress so a new process can resume the same named run after a crash, deploy, eviction, or human wait. Memory stores facts and conversation the model may see next turn. You need both. A longer prompt history is not a resume path, and a vector index will not keep a refund from firing twice.

What happens when the process dies mid-tool-call?

Whatever you had not committed is gone. A durable engine reloads the last boundary or replays the journal and re-enters the unfinished step. If that step already sent an email or a charge, only an idempotency key on the ledger prevents a duplicate. If you had no handle, the next wake starts a sibling run and you debug from logs.

How do human-in-the-loop pauses stay durable?

Persist the run at a boundary, return a stable handle, and resume that same handle with the human’s decision. The original HTTP request will be gone. Bind approvals to run id plus payload hash, schedule a wake if the host hibernates, and kill-test the pause path in staging before anyone important waits on it.

Why does resume re-run the last step?

Durable engines treat a boundary as a restartable unit. Graph interrupts restart the node. Event-history engines re-execute control code and fill completed work from the log. Fibers call your recover hook from the last stash. Put writes behind a persisted key, or split prepare and write into two boundaries, so the re-entry cannot create a second side effect.

What state should never be stuffed into the prompt?

Run ids, tool ledgers, auth material, approval records, budgets, and raw secret-bearing tool dumps. Summaries may appear in context; the durable store remains authoritative. Prompt-only “state” evaporates on summarization and restart, which is how a recovered agent re-calls a tool that already succeeded.

How does durability interact with idempotent tool writes?

Resume re-executes boundaries. Without idempotency keys outside the model, a correct resume becomes a duplicate side effect. Persist the key on the run handle before the first attempt, reuse it on every retry, and measure “side effects once” in the kill-and-resume drill. Durability without that ledger just makes the second charge more reliable.

CTA

Need a durable agent that survives the first real deploy — not just the demo loop? Start on /agentic or book the pilot at /contact?intent=agentic-pilot.

FAQ

What questions does this article answer?

What is a durable agent runtime versus memory?
A durable runtime stores execution progress so a new process can resume the same named run after a crash, deploy, eviction, or human wait. Memory stores facts and conversation the model may see next turn. You need both. A longer prompt history is not a resume path, and a vector index will not keep a refund from firing twice.
What happens when the process dies mid-tool-call?
Whatever you had not committed is gone. A durable engine reloads the last boundary or replays the journal and re-enters the unfinished step. If that step already sent an email or a charge, only an idempotency key on the ledger prevents a duplicate. If you had no handle, the next wake starts a sibling run and you debug from logs.
How do human-in-the-loop pauses stay durable?
Persist the run at a boundary, return a stable handle, and resume that same handle with the human’s decision. The original HTTP request will be gone. Bind approvals to run id plus payload hash, schedule a wake if the host hibernates, and kill-test the pause path in staging before anyone important waits on it.
Why does resume re-run the last step?
Durable engines treat a boundary as a restartable unit. Graph interrupts restart the node. Event-history engines re-execute control code and fill completed work from the log. Fibers call your recover hook from the last stash. Put writes behind a persisted key, or split prepare and write into two boundaries, so the re-entry cannot create a second side effect.
What state should never be stuffed into the prompt?
Run ids, tool ledgers, auth material, approval records, budgets, and raw secret-bearing tool dumps. Summaries may appear in context; the durable store remains authoritative. Prompt-only “state” evaporates on summarization and restart, which is how a recovered agent re-calls a tool that already succeeded.
How does durability interact with idempotent tool writes?
Resume re-executes boundaries. Without idempotency keys outside the model, a correct resume becomes a duplicate side effect. Persist the key on the run handle before the first attempt, reuse it on every retry, and measure “side effects once” in the kill-and-resume drill. Durability without that ledger just makes the second charge more reliable.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot