Durable Agent Runtimes: Survive Restarts Without Calling It "Memory"
A durable agent runtime stores resume-correct execution progress outside the process so a crash or human wait continues the same named run—not a new one.
William Spurlock Founder — Spurlock Studios Updated 20 MIN
A long-running agent that survives crashes and waits for humans needs a durable runtime: serializable control state outside the process, a resume path that continues from the last safe boundary, and tool writes that stay safe when that boundary re-runs. That is not “memory.” Memory is what the agent recalls. Durability is whether the run still exists after the host dies.
This spoke sits under the Agentic Systems Operating Manual. Pair it with state machines for agent loops for the legal states the loop may occupy. Durability is the substrate that keeps those states alive across restarts. The agentic lane is where we prove that substrate on one real job.
The short answer
- Durability ≠ memory. Memory stores facts and conversation. Durability stores execution progress so a new process can resume the same job.
- In-process loops die under real ops. Deploys, OOM kills, Durable Object eviction, and overnight approvals all outlive a single Node or Python process.
- Resume re-enters a boundary. Snapshot engines, event-history engines, and fiber/stash engines all re-run the last unfinished step. Design for that, or you double-send.
- HITL pauses must be durable. A pause that lives only in RAM is a memory leak with a polite name.
- Measure resume correctness, not “it remembered.” After a kill, does the same run id continue without a second email or refund?
What is a durable agent runtime (and what is not)?
A durable agent runtime is the store plus the resume protocol that lets a new process continue a named run. If you cannot kill the worker mid-approval and still resume the same job tomorrow without inventing a new run id, you do not have durability. You have a longer prompt.
LangGraph’s own persistence docs split this cleanly: a checkpointer tracks thread-scoped graph progress (fault tolerance, interrupts, time travel); a store tracks cross-thread facts (preferences, knowledge). The first is a runtime concern. The second is memory. Teams collapse them and then wonder why a vector index did not survive a deploy.
| Concern | Durable runtime | Chat / agent “memory” |
|---|---|---|
| Question it answers | Where was this run when the process died? | What facts should the model see next turn? |
| Typical store | Checkpoints, event history, DO SQLite, workflow journal | Vector store, profile rows, transcript slices |
| Success metric | Resume correctness, no duplicate side effects | Answer quality, fewer re-asks |
| Survives a 3-day human wait? | Required | Optional |
| Survives process death mid-tool? | Required | Irrelevant |
If you only keep a longer prompt history, you have memory. If you can evict the worker and still resume the same handle, you have durability.
Why do in-process agent loops die under real ops?
A demo agent is a while loop in one process. Production invents ways to murder that process, and none of them wait for the model to finish a thought.
- Rolling deploys mid-tool-call
- Platform eviction (Cloudflare Durable Objects typically idle-evict after ~70–140 seconds without keep-alive, alarms, or inbound events — see the Durable Object lifecycle)
- Spot instances and scale-to-zero on serverless
- Operator restart after a bad release
- Human approval that arrives after the original HTTP request is gone
- Uncaught exception that tears down in-memory state while the disk or SQLite row is still honest
The failure we see constantly after 500+ automations and 20,000+ hours on agentic systems: the agent emailed the customer, the process died before writing “done,” and a retry emailed again. The model did nothing wrong. The runtime treated “in memory” as “committed.”
| Kill | What dies | What must survive |
|---|---|---|
| Deploy / pod restart | Process, RAM, open sockets | Run handle + last committed boundary |
| DO hibernate / evict | In-memory vars, local timers, open fetches | setState / SQL / fiber stash / alarm |
| Overnight HITL | Original request thread | Pause record + approval binding |
| OOM / crash | Everything in the heap | Whatever you flushed before the kill |
Cloudflare’s rules of Durable Objects say the quiet part: in-memory state is not preserved on eviction or an uncaught exception. Persist first. Then do the thing that might fail.
What actually gets persisted on a crash-resume path?
Three records have to exist before you call the run durable. Miss one and you will reconstruct a story from logs instead of resuming a job.
| Record | Job | If missing |
|---|---|---|
| Run handle | Stable id (thread_id, fiber id, workflow id) | Resume invents a sibling run |
| Boundary snapshot or journal | Last safe step + inputs/results | Resume rewinds or skips |
| Side-effect ledger | Idempotency keys + receipts for writes | Correct resume becomes a duplicate charge |
Checklist for the ledger you actually ship:
- One durable handle per job, minted before the first tool write
- Boundary writes are committed before the next side effect starts
- Tool receipts are keyed by that handle, not by “whatever the model said last”
- Secrets are pointers, not blobs in the checkpoint
- The UI loads the handle, not “the latest conversation”
Webhook retries are a crash with better manners. The provider will POST again. If accept is not durable, you start a second run that looks like recovery. Cloudflare startFiber() exists for that boundary: persist acceptance and an idempotency key before work begins, then return status on the retry. Same rule on any host — mint the handle on the first byte you acknowledge.
Procedure for minting the handle:
- Receive the job (HTTP, queue, email, cron).
- Derive or generate
run_idfrom an idempotency key the caller already sent, or from a key you persist before ack. - Write the handle row in the same transaction as “accepted.”
- Only then call a model or a write tool.
- On a duplicate delivery, return the existing handle. Do not create a sibling.
Restate’s key concepts name the same contract in journal language: record the step and its result; on crash, replay the journal and skip completed work. Temporal’s event history is the same idea with a different product name. The product is optional. The three records are not.
Snapshot checkpoints vs event-history replay — what is the difference?
Crash-resume engines do not all restore the same way. Confusing the mechanism is how you write a “resume” that is actually a replay of the entire graph with half the side effects live.
Snapshot / checkpoint. Persist graph or entity state at a named boundary. A new process loads the blob and continues from that super-step. LangGraph checkpointers do this at super-step boundaries. Cloudflare fibers do a close cousin: stash() a snapshot, recover via onFiberRecovered.
Event-history / journal replay. Persist an append-only log of decisions and activity results. A new worker re-runs the control code from the start, substituting recorded results for work that already happened. Temporal replays from event history rather than restoring a heap snapshot. Azure Durable Functions use the same event-sourcing replay. Restate’s request lifecycle restarts the handler and skips journaled steps.
| Mechanism | Restore method | What re-runs | Typical failure if you ignore it |
|---|---|---|---|
| Snapshot checkpoint | Load last state blob | The interrupted node / fiber from its start | Side effects before the interrupt fire twice |
| Event-history replay | Re-execute control code; reuse recorded results | Control flow (must stay deterministic) | Non-deterministic clocks/IDs fork the history |
| Fiber stash | Load last stash(), then your recover hook | Whatever you did not stash | Recover starts the job from zero |
This is not a vendor bake-off. It is the difference between “load the save file” and “re-read the campaign log.” Both can be correct. Both will punish you if you treat RAM as the source of truth.
LangGraph durability modes make the snapshot tradeoff explicit. From the current checkpointer docs, least to most durable:
| Mode | When it writes | Crash mid-graph |
|---|---|---|
"exit" | Only when the graph exits (success, error, or interrupt) | Intermediate steps are gone |
"async" | In the background while the next step runs | Small window where the checkpoint never lands |
"sync" | Before the next step starts | Highest durability, more write cost |
"exit" is faster for long graphs and does not protect mid-execution crashes the way "sync" does. If you say “we checkpoint,” name the mode.
Event-history engines add a second tax: control code must be deterministic, because it will run again. Azure’s orchestrator constraints and Temporal’s replay rules are the same list in different SDKs.
| In control code | Safe | Unsafe |
|---|---|---|
| Time | Workflow/orchestrator clock APIs | Date.now(), wall-clock sleeps |
| IDs | Journaled / history-recorded IDs | Fresh UUID on every replay |
| I/O | Activities / ctx.run / tool adapters | Direct HTTP from the replayed function |
| Branching | Based on recorded results | Based on “what the model might say this time” inside control |
If you need a non-deterministic call — LLM, HTTP, Math.random — put it behind a recorded boundary. Replay then returns the first result. Doing it inline is how a recovered run takes a different fork and the history engine throws a non-determinism error, or worse, silently diverges on a snapshot host.
Why does resume re-enter the last step?
Because almost every durable engine treats a boundary as a restartable unit, not a line number in your source file.
LangGraph interrupts are the clearest public contract: on Command(resume=...), the node restarts from the beginning. Code before interrupt() runs again. The interrupt call then returns the human’s value instead of pausing. That is intentional, not a bug.
Event-history engines do the same at a different grain. Temporal and Azure Durable Functions re-execute the workflow or orchestrator from the top and fill awaits from history. Activities that already completed return the recorded result. Activities that did not complete may be scheduled again. Cloudflare fibers call your recover hook with the last stash; they do not rewind your JavaScript to the exact line.
| Engine family | Resume grain | What you must make safe |
|---|---|---|
| Graph + interrupt | Whole node | Anything before interrupt() |
| Event history | Whole workflow function | Non-deterministic control code; unrecorded side effects |
| Fiber / stash | From last stash | Work after the last stash() |
Procedure that keeps resume from becoming a second send:
- Split “prepare” and “write” into separate boundaries.
- Mint the idempotency key in the prepare boundary and persist it.
- On the write boundary, send that same key. Never generate a new UUID on resume.
- Treat logs before an interrupt as expected duplicates unless you gate them on a committed flag.
If you need the legal states around that split, that is the state machine spoke. This post only owns the fact that resume will walk back into the last room.
How do human-in-the-loop pauses stay durable?
A durable HITL pause is three parts. Skip any one and the human is approving a ghost.
- Persist the run at a named boundary (checkpoint, fiber stash, workflow wait, or alarm).
- Return a handle the UI or ops channel can load (
thread_id, fiber id, workflow id). - Resume by feeding the decision into that handle — not by starting a new chat.
Cloudflare’s long-running agents docs are blunt: agents are durable identities, not always-on processes. State, SQL, schedules, and fiber checkpoints survive hibernation. In-memory variables, timers, open fetches, and local closures do not. A “wait for the manager” that is a setTimeout in RAM is already dead.
Durable Object alarms are the wake primitive when the human may take hours: at-least-once alarm() delivery, exponential backoff on throw, and a hard cap of six automatic retries unless you reschedule. Design the escalate path as if the human never answers.
Checklist:
- Pause state is in durable storage, not the request thread
- Approval payload is bound to that run id plus a payload hash (no “approve whatever is latest”)
- Resume path is exercised in staging with a process kill mid-wait
- Tool nodes after resume are idempotent
- Timeout / escalate path exists if the human never answers
- The original HTTP connection is treated as already gone
If the human waits three days, the original request is archaeology. Only the durable handle matters.
| Clock | What must be true |
|---|---|
| T0 — pause | Boundary flushed; handle returned; approval UI bound to run_id + payload hash |
| T1 — host dies | In-memory queue is gone; SQLite / checkpoint / journal / alarm still hold the wait |
| T2 — human acts | Resume loads T0 payload, not “latest ticket”; decision written on the same handle |
| T3 — write | Refund/email tool uses the key minted at T0, not a key minted at T2 |
| T4 — silence | Escalate alarm fires; the run moves to a named escalate state, not a new chat |
A Slack message with a raw “approve” button and no handle is T0 skipped. Everything after it is folklore.
What happens when the host hibernates or evicts the agent?
Hibernation is a first-class kill, not an edge case for people who picked Cloudflare. Any scale-to-zero host will do a version of this. Cloudflare just publishes the clocks.
From the Durable Object lifecycle:
| State | What is true |
|---|---|
| Active, in-memory | Handling a request or event |
| Idle, hibernateable | After ~10 seconds of inactivity, the runtime may hibernate |
| Hibernated | Removed from memory; hibernated WebSockets can stay connected |
| Evicted / inactive | After ~70–140 seconds idle without keep-alive conditions; next wake is a cold start |
Cloudflare Agents durable execution draws the line teams miss: keepAlive() reduces the chance of eviction. runFiber() makes eviction survivable by registering the work in SQLite, letting you stash(), and calling onFiberRecovered on the next activation. startFiber() adds durable acceptance, an idempotency key, and retained status for callers that retry (webhooks).
| Primitive | What it buys | What it does not buy |
|---|---|---|
keepAlive() / keepAliveWhile() | Fewer evictions during active LLM/tool work | Recovery if eviction still happens |
runFiber() / stash() | Crash-recoverable work + recover hook | Automatic replay of your JS from line 47 |
startFiber() | Durable accept + inspect + cancel + dedupe | A free pass on tool idempotency |
schedule() / alarms | Wake later for HITL and polls | In-memory closures from the previous activation |
keepAlive without a fiber is how “it worked in staging” becomes “lost the job overnight.” Fibers without idempotent writes are how a recovered job double-books.
A recover hook is only as good as the last stash. Write a small, typed snapshot — not “whatever was in RAM.”
Field in stash() | Why |
|---|---|
run_id | So recover cannot attach to a sibling job |
step | Named boundary (drafted, awaiting_approval, refund_sent) |
idempotency_key | Reused on the next write; never regenerated |
payload_hash | Binds the human decision to the draft they saw |
attempt | Caps recover loops the same way a state machine caps revisions |
If onFiberRecovered cannot find those fields, fail the fiber and page a human. Guessing the step from the transcript is memory cosplay.
What state should never live only in the prompt?
The prompt may summarize control state. The prompt must not be the only copy. Summarization is how agents re-call a tool that already succeeded, then tell you they “remembered” the ticket.
| State | Where it belongs | Why the prompt is a bad home |
|---|---|---|
| Run / thread / fiber ids | Runtime ledger | Summarizers drop or invent ids |
| Tool ledger (what ran, keys, result hashes) | Harness DB | Model will retry a “failed” success |
| Auth tokens and IAM scope | Secret store + policy gate | Checkpoints leak; prompts get logged |
| Approval decisions | Durable HITL record | A new chat cannot see a signed decision |
| Budget / kill-switch counters | Control plane | Soft numbers in context do not stop a loop |
| Full raw tool dumps | Artifact store with redaction | Blow the window; hide the receipt |
Keep these out of the model context as source of truth:
- Run handle
- Idempotency keys
- Approval payload hash
- Token material
- Cost and turn counters that trip the kill switch
The model can see a redacted summary. The resume path must not ask the model where the run left off.
How do you measure resume correctness?
“The model recalled the ticket number” is a memory win. It is not a durability win. Run this drill before you argue about model quality, and again after every runtime change.
- Start a real job that reaches a HITL pause or a slow tool write
- Kill the worker / evict the Durable Object / restart the pod
- Resume from the stored handle only — no new chat
- Score the outcome against the table below
| Metric | Pass | Fail |
|---|---|---|
| Same run id continues | Yes | New run invented |
| Side effects once | One email / one charge | Duplicate |
| State-machine position | Same state as before the kill | Rewound or skipped |
| Human sees prior context | Approval UI shows the pending payload | Empty or wrong job |
| Ledger matches reality | Receipt hash equals the upstream object | “Done” in RAM, nothing upstream |
Monthly is enough for a stable harness. After a durability-mode change, a fiber migration, or a new write tool, run it before the next deploy. Record the drill in the handoff. If you cannot produce the recording, you do not have a durable runtime. You have a story.
Staging drill, minimum:
- Job reaches a real write or a real HITL pause (not a mocked
sleep) - Kill is the kill you will actually see (deploy,
wranglerrestart, pod delete, DO idle-evict) - Resume uses only the stored handle — no “start over” button
- Upstream is checked for a second object, not just your own “done” flag
- Trace shows the same
run_idbefore and after the kill - Recording (screen or log bundle) is attached to the handoff
Skip the drill and you will learn the same facts from a customer. That is a more expensive lab.
How does durability interact with idempotent tool writes?
Durability increases how often a step re-executes after a crash. That makes idempotency mandatory, not optional. A correct resume that re-enters a refund node without a stable key is a second refund with better logging.
Stripe’s idempotent requests are the public contract most payment tools already speak: send an Idempotency-Key, and retries return the original result instead of creating a second object. Keys are kept at least 24 hours. Generate the key before the first attempt and persist it on the run handle. Do not mint a fresh UUID inside the resumed node.
On resume:
- Re-enter the node / fiber / activity
- Tool layer sees the same idempotency key from the ledger
- Upstream returns the original receipt — no second charge
| Layer | Owns | Must not own |
|---|---|---|
| Runtime | Run handle, boundary, key mint + persist | The English of the email |
| Tool adapter | Sending the key, mapping receipts | Inventing a new key on retry |
| Model | Drafts and proposed args | “I think we already refunded” as control |
Treat durability and idempotency as one control loop. The operating manual frames the rest of that stack. This post only owns the resume substrate.
Worked failure: the overnight approval that double-booked
What broke: A support agent drafted a refund and paused for manager approval in an in-memory queue. A deploy restarted the API. The manager approved what looked like a stuck ticket. A new agent run also resumed from a stale Redis key that was never bound to a single run_id. Two refunds left.
Cost: Finance cleanup, a customer-trust hit, and two days of “why agents suck” in Slack. No model regression. A missing handle.
Instead:
- Persist the pause under a single durable
run_idbefore the draft is shown - Bind the approval button to that id plus a payload hash
- Issue the refund tool with an idempotency key owned by the runtime
- Kill-test the pause path before soft-launch — deploy mid-wait, then approve once
| Control | Before (failed) | After (holds) |
|---|---|---|
| Pause store | In-memory queue | Checkpoint / fiber / workflow wait |
| Approval target | “Latest stuck ticket” | run_id + payload hash |
| Refund key | None | Persisted idempotency key |
| Kill test | None | Staging drill recorded |
Bravery is not a restore strategy.
What is the pilot minimum for crash-resume?
A $1,500 · 5-day agentic pilot does not require Temporal, Restate, or a fiber rewrite on day one. It does require a resume path you can kill.
- One durable handle per job (
thread_idor equivalent), minted before the first write - One kill-and-resume test recorded in the handoff
- Idempotent write tools on the critical path
- HITL pause that survives process death, or an explicit “no human wait in v1” written down
- A named boundary flush before the side effect, not after
| Day | Crash-resume proof |
|---|---|
| 1 | Name the handle, the store, and the one write that must not double |
| 2–3 | Wire persist → kill → resume on a staging job |
| 4 | Bind HITL or document why the job has no pause |
| 5 | Record the drill; list the two kills you still do not survive |
Framework fashion is optional. Resume correctness is not. If the drill fails, you do not have a durable agent. You have a demo that remembers things until the host blinks.
Which durability anti-patterns fail first?
These fail in the first real week. They fail in the same order almost every time.
Calling a vector store “our durable agent.” That is memory. It will not resume a run.
In-memory checkpointer in production. Fine for unit tests. Worthless for crashes. LangGraph’s InMemorySaver is the example everyone ships by accident.
HITL as “email the ops channel and hope.” No run handle, no resume, no payload hash. The human approves a vibe.
Checkpointing the entire blob of secrets into Postgres. Redact. Store pointers. Treat the checkpoint as a document an intern can dump.
Assuming the agent is an always-on process. Cloudflare Agents hibernate. Kubernetes pods die. Serverless scales to zero. Design for wake/sleep.
Minting a new idempotency key on resume. The ledger must reuse the key from the prepare boundary. A new UUID is a new charge with a clean conscience.
Scoring “it remembered.” Recalling a ticket number after restart is memory. Same run id, one refund, held state — that is durability.
| Anti-pattern | First symptom | Fix |
|---|---|---|
| Vector store as runtime | New run id after deploy | Add a handle + boundary store |
| In-memory saver | Empty thread after restart | Durable checkpointer / SQL / journal |
| Slack-only HITL | Duplicate work after approval | Bind UI to run_id |
| Secrets in checkpoints | Token in a support dump | Pointers + redaction |
| New key on resume | Double charge with matching logs | Persist the key at prepare |
Durability is not the diagnosis when the model picks the wrong tool, the evaluator is mush, or the state machine has no escalate. Those are other spokes. This one only answers: after the host dies, does the same run continue without a second write?
| Symptom | First place to look |
|---|---|
| New run id after deploy | Handle + store |
| Same run, second refund | Idempotency ledger |
| Same run, wrong next state | State machine, not the runtime |
| Recalled the ticket, lost the pause | You built memory and called it durable |
FAQ
What is a durable agent runtime versus memory?
A durable runtime stores execution progress so a new process can resume the same named run after a crash, deploy, eviction, or human wait. Memory stores facts and conversation the model may see next turn. You need both. A longer prompt history is not a resume path, and a vector index will not keep a refund from firing twice.
What happens when the process dies mid-tool-call?
Whatever you had not committed is gone. A durable engine reloads the last boundary or replays the journal and re-enters the unfinished step. If that step already sent an email or a charge, only an idempotency key on the ledger prevents a duplicate. If you had no handle, the next wake starts a sibling run and you debug from logs.
How do human-in-the-loop pauses stay durable?
Persist the run at a boundary, return a stable handle, and resume that same handle with the human’s decision. The original HTTP request will be gone. Bind approvals to run id plus payload hash, schedule a wake if the host hibernates, and kill-test the pause path in staging before anyone important waits on it.
Why does resume re-run the last step?
Durable engines treat a boundary as a restartable unit. Graph interrupts restart the node. Event-history engines re-execute control code and fill completed work from the log. Fibers call your recover hook from the last stash. Put writes behind a persisted key, or split prepare and write into two boundaries, so the re-entry cannot create a second side effect.
What state should never be stuffed into the prompt?
Run ids, tool ledgers, auth material, approval records, budgets, and raw secret-bearing tool dumps. Summaries may appear in context; the durable store remains authoritative. Prompt-only “state” evaporates on summarization and restart, which is how a recovered agent re-calls a tool that already succeeded.
How does durability interact with idempotent tool writes?
Resume re-executes boundaries. Without idempotency keys outside the model, a correct resume becomes a duplicate side effect. Persist the key on the run handle before the first attempt, reuse it on every retry, and measure “side effects once” in the kill-and-resume drill. Durability without that ledger just makes the second charge more reliable.
CTA
Need a durable agent that survives the first real deploy — not just the demo loop? Start on /agentic or book the pilot at /contact?intent=agentic-pilot.
What questions does this article answer?
- What is a durable agent runtime versus memory?
- A durable runtime stores execution progress so a new process can resume the same named run after a crash, deploy, eviction, or human wait. Memory stores facts and conversation the model may see next turn. You need both. A longer prompt history is not a resume path, and a vector index will not keep a refund from firing twice.
- What happens when the process dies mid-tool-call?
- Whatever you had not committed is gone. A durable engine reloads the last boundary or replays the journal and re-enters the unfinished step. If that step already sent an email or a charge, only an idempotency key on the ledger prevents a duplicate. If you had no handle, the next wake starts a sibling run and you debug from logs.
- How do human-in-the-loop pauses stay durable?
- Persist the run at a boundary, return a stable handle, and resume that same handle with the human’s decision. The original HTTP request will be gone. Bind approvals to run id plus payload hash, schedule a wake if the host hibernates, and kill-test the pause path in staging before anyone important waits on it.
- Why does resume re-run the last step?
- Durable engines treat a boundary as a restartable unit. Graph interrupts restart the node. Event-history engines re-execute control code and fill completed work from the log. Fibers call your recover hook from the last stash. Put writes behind a persisted key, or split prepare and write into two boundaries, so the re-entry cannot create a second side effect.
- What state should never be stuffed into the prompt?
- Run ids, tool ledgers, auth material, approval records, budgets, and raw secret-bearing tool dumps. Summaries may appear in context; the durable store remains authoritative. Prompt-only “state” evaporates on summarization and restart, which is how a recovered agent re-calls a tool that already succeeded.
- How does durability interact with idempotent tool writes?
- Resume re-executes boundaries. Without idempotency keys outside the model, a correct resume becomes a duplicate side effect. Persist the key on the run handle before the first attempt, reuse it on every retry, and measure “side effects once” in the kill-and-resume drill. Durability without that ledger just makes the second charge more reliable.
- docs.langchain.com
- developers.cloudflare.com
- developers.cloudflare.com
- docs.restate.dev
- docs.temporal.io
- docs.langchain.com
- developers.cloudflare.com
- docs.temporal.io
- learn.microsoft.com
- docs.restate.dev
- learn.microsoft.com
- docs.langchain.com
- developers.cloudflare.com
- developers.cloudflare.com
- docs.stripe.com
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.