State Machines for Agent Loops: Determinism Where It Matters
Give the model freedom inside a named state. Legal transitions, revision ceilings, and escalate paths keep operations deterministic when tokens are not.
William Spurlock Founder — Spurlock Studios Updated 26 MIN
Models are non-deterministic. Operations cannot be. The way you reconcile those facts is a state machine: the model gets freedom to plan, draft, and pick tools inside a named state; the runner decides which transitions are legal, when a write may happen, how many revisions you buy, and when the run must escalate or abort.
This spoke sits under the Agentic Systems Operating Manual. If you are still shopping for whether you need an agent at all, read When Not to Build an Agent first.
The short answer
- Freedom is intra-state. Inside
plan,act, orrevise, the model may propose content and tool calls. It does not get to invent a new state or skip evaluation. - Transitions are extra-state and legal or they do not happen. The runner maps
(state, event, guards)to a next state. Illegal wishes are ignored. - Revision has a ceiling. Three is a good pilot default. Hit it and you escalate with evidence, not another hidden loop.
- Escalate is a terminal success of a kind. The system knew it was out of its depth while a human could still help.
- Identical tokens are not the bar. Deterministic multi-agent means control flow, policy, accounting, and terminals. Chasing bit-identical model output wastes budget.
What is an agent loop state machine?
An agent loop state machine is an explicit set of named states and allowed transitions that wrap model calls and tool use. The model proposes inside a state. The machine decides whether the next state is legal given the evaluator verdict, the remaining budget, the revision counter, and the error class.
Without that wrapper you get unbounded retries, duplicate side effects, “done” declared by the worker, and no clean place for a human to intervene. I have spent 20,000+ hours on agentic systems and built 500+ automations. The loops that survived contact with production all had a graph an engineer could read without speaking prompt.
| Object | Who owns it | What “freedom” means |
|---|---|---|
| Prompt / draft / tool args | Model, inside the current state | Propose content and allowlisted calls |
| Next state | Runner / graph | Only listed edges fire |
| Side-effect permission | State definition | Writes only in act (or a named write child) |
| Stop | Counters + terminals | done, escalate, or abort — never “keep going” |
| Evidence | Trace store | Every transition leaves a reason code |
LangGraph’s own Graph API says the same thing in vendor vocabulary: you model the workflow as a graph of state, nodes, and edges — “nodes do the work, edges tell what to do next.” See Graph API overview. XState’s primer is blunter about the control plane: transitions are deterministic; each combination of state and event always points to the same next state. That line lives on What are state machines and statecharts?.
- Named states written down (not implied by a mega prompt)
- Allowed edges written down (
from,event,to,guards) - Side-effect class attached to each state
- At least one terminal that is not
done - A counter the model cannot reset
If any box is empty, you have a directed suggestion. Not a machine.
Why do models get freedom inside states but not between them?
Because the useful non-determinism is local. You want the model to try a different retrieval query, a tighter draft, or a different allowlisted tool. You do not want it to invent skip_eval, email a customer from plan, or declare victory because the last sentence sounded confident.
XState calls the between-states rule by its real name. A transition is a change from one finite state to another, triggered by an event. Enabled transitions are selected from the current state; if none are enabled, the state does not change. That is the cage. The model emits artifacts. An adapter turns outcomes into events the machine already understands.
| Layer | Non-deterministic is fine | Must be deterministic |
|---|---|---|
| Tokens | Wording, tool-arg phrasing, plan order | Schema of the artifact you accept |
| Tool choice | Among the allowlist for this state | Whether the state may call tools at all |
| Timing | Model latency inside a timeout | Timeout → named event (timeout) |
| Quality | First-pass draft | Evaluator verdict → eval_pass / eval_fail |
| Stop | Never | Ceiling, budget, policy, human |
LangGraph splits the same idea into workflow vs agent. Workflows and agents says workflows have predetermined code paths and “operate in a certain order”; agents are dynamic and “define their own processes and tool usage.” Production agent loops are a mix: predetermined edges, dynamic work inside nodes. If the whole job is a predetermined path, you do not need this spoke. You need a workflow. That no lives in When Not to Build an Agent.
- Name the state the model is in.
- Give it a contract: input schema, output schema, tool allowlist, timeout.
- Translate its output into a small event (
planned,tool_ok,tool_err,eval_fail). - Let guards pick the next state. Do not let the model name the next state as a string you then
eval.
The model is a guest in the room. The floor plan is yours.
What are the legal transitions on a thin machine?
Start with eight states. Fancy parallel regions wait until this graph clears a golden set.
| State | Allowed work | Forbidden work |
|---|---|---|
intake | Validate job contract, attach budget and tool allowlist | Model calls, production writes |
plan | Model proposes steps | Production writes |
act | Allowlisted tools; side effects only here | Declaring done |
evaluate | Independent judge; no writes | Patching CRM, sending mail |
revise | Worker edits with failure evidence; increment counter | Resetting the counter |
done | Terminal success; store receipts | Further tools |
escalate | Terminal handoff with a package | Quiet retry |
abort | Terminal stop on budget, policy, or unrecoverable error | “Best effort” write on the way out |
Legal transitions (simplified):
| From | Event | To | Guard |
|---|---|---|---|
intake | start | plan | schema valid, budget > 0 |
intake | reject | abort | out of scope or auth fail |
plan | planned | act | steps in allowlist |
plan | timeout / policy | escalate or abort | — |
act | tool_ok | evaluate | — |
act | tool_err | escalate or abort | by error class |
evaluate | eval_pass | done | — |
evaluate | eval_fail | revise | revisions < ceiling |
evaluate | eval_fail | escalate | revisions >= ceiling |
revise | revised | act | ceiling not hit |
| anything serious | budget_low / policy | abort | — |
The exact graph can vary. The requirement is that it is written down and enforced in code or in the workflow rail. AWS Step Functions makes the same demand in Amazon States Language: execution starts at StartAt, each state names Next or ends, and the run stops at Succeed, Fail, or "End": true. See Transitions in state machines. If your “graph” is a while-loop in a system prompt, you do not have Next. You have hope.
W3C already standardized the older, vendor-neutral version of this idea. SCXML (Recommendation, 1 September 2015) is a general-purpose event-based state machine language. You do not need to author SCXML. You do need the same objects: states, events, transitions, and a legal configuration.
What does “deterministic multi-agent” actually mean?
People say “deterministic multi-agent” when they want predictability. You will not get bit-identical model outputs. You can get four things that matter to ops.
| Kind of determinism | What it looks like | What it is not |
|---|---|---|
| Control flow | Same (state, event, guards) → same next state | Same tokens |
| Policy | Allowlists, caps, ceilings, deny classes | “The model usually behaves” |
| Accounting | Cost, latency, tools, and reason codes always recorded | A chat log you grep later |
| Terminality | Every run ends in done, escalate, or abort | A hanging Wait with no limit |
That is the bar. Temporal’s workflow docs make the split operational: workflow definitions must be deterministic so replay works; LLM calls, API calls, and other external work go in Activities, outside the replay path. Put the non-deterministic guest in a room with a door. Do not let it rebuild the building.
When a second role joins later (librarian, writer, critic), it enters a known state with a typed package. It does not join a free-form shared chat and “figure it out.” Coordination is an edge, not a vibe.
- I can unit-test
(state, event) → nextwithout calling a model - I can name the three terminals
- I can name the budget owner (one, not three retry layers)
- A new agent role would enter through
intakeor a child machine, not through a Slack thread
If you cannot tick those, you are still in demo land. The parent map for evaluators, sandboxes, and kill switches is the operating manual.
Where do side effects live?
Only in act — or in a named write child of act. This rule prevents half the horror stories. Planning tokens do not email customers. Evaluation does not patch CRM. Revision drafts land in scratch until act applies them under sandbox rules.
XState splits effects the same way the rest of serious FSM work does. Guards are pure, synchronous predicates that return true or false; they decide whether a transition is enabled. Actions are fire-and-forget side effects that run because a transition was taken. If your architecture lets any state call any tool, you do not have a state machine. You have a directed suggestion.
| State | May call tools? | May write customer-visible systems? | Notes |
|---|---|---|---|
intake | No | No | Schema and auth only |
plan | Read-only retrieval if you must | No | Prefer retrieval as a child with no write tools |
act | Yes, allowlisted | Yes, under sandbox + idempotency | The only write room |
evaluate | Read-only checks | No | Independent of the worker |
revise | No writes | No | Edits the artifact, increments the counter |
done / escalate / abort | Receipts / package only | No new business writes | Terminal |
If a tool result must change the next state, the tool returns to act, act emits tool_ok or tool_err, and the edge decides. The tool does not jump the graph.
- Classify every tool: read, write, irreversible.
- Bind write and irreversible tools only to
act. - Put irreversible classes behind a human gate or a deny-by-default policy.
- Refuse to compile the graph if a write tool is reachable from
planorevaluate.
LangGraph will let a node perform a side-effect — the Graph API says nodes “perform some computation or side-effect.” That is permission to be careful, not permission to be sloppy. You still decide which nodes are allowed to write.
How do revision ceilings and escalate packages work?
Unbounded revise loops are how you discover a task is impossible after the bill already moved. Cap revisions. Three is a good pilot default. On ceiling, stop buying another turn.
| Counter | Pilot default | On trip |
|---|---|---|
| Revisions | 3 | evaluate → escalate (not another revise) |
Tool calls in act | Per-job allowlist + max | tool_err / budget_low → escalate or abort |
| Wall-clock per state | Explicit timeout | timeout event |
| Spend | Hard cap on the job | abort or escalate with cost in the package |
| Human wait | Limit Wait Time / interrupt SLA | timeout → escalate, never hang forever |
XState guards are the right shape for the ceiling: revisions < 3 is a pure predicate. First matching guarded transition wins; put the specific guard before the default. Do not hide the counter in the prompt (“try a few times”). The model will try until the card declines.
Escalate package — ship all of this, every time:
- Job contract (what “done” meant)
- Artifacts so far
- Evaluator failures with evidence
- Tools called and outcomes
- Cost and latency
- Reason code (
ceiling,policy,timeout,auth,schema) - Recommended human action, if you have one
Escalation is success of a kind. The system knew it was out of its depth while someone could still help. A silent retry after the ceiling is a lie.
LangGraph’s interrupt is the vendor form of “pause and wait for a human.” interrupt() saves graph state through the persistence layer and waits until you resume with Command. That is an escalate gate when you still might return to act. A true escalate terminal is different: the machine is done; a human owns the next business action. Do not confuse a pause with a handoff. Name both.
- Ceiling is an integer in state, not a vibe in the prompt
-
eval_fail+ ceiling →escalate, tested - Package schema is versioned
- On-call can open a run and see why it stopped in under a minute
How do LangGraph, XState, and Step Functions encode the same cage?
You do not owe any one framework. You owe the objects. The vendors already documented them.
| Object you need | LangGraph | XState | Step Functions |
|---|---|---|---|
| Shared snapshot | State schema + reducers | Context + finite state | Execution input / variables |
| Work | Nodes | Actions / invoked actors | Task |
| Next step | Edges, including conditional | Transitions on events | Next, Choice |
| Predicate | Your function on a conditional edge | Guards | Choice rules |
| Side effect | Node you marked as write-capable | Actions on the taken transition | Task that calls the write |
| Persist / resume | Checkpointers + thread id | Persisted actor snapshot (your store) | Service-managed execution history |
| Human pause | Interrupts | Event from the human, or a wait state | Wait-for-callback / .waitForTaskToken patterns |
| Terminal | END / your done node | Final state | Succeed, Fail, "End": true |
LangGraph persistence is explicit: checkpointers save a thread’s graph state as checkpoints (short-term, for HITL, time travel, fault tolerance); stores hold long-term data across threads. Compile without a checkpointer and you cannot pause or survive a crash. That is not an implementation detail. That is the difference between a notebook demo and a run you can page.
XState’s machines doc lists the same kit: actions, actors, guards, delayed transitions. Use it as a checklist even if you never import the library. If you cannot point to each object in your runner, you are missing a piece.
Step Functions bills and traces by transition. Error handling treats retries as transitions; Catch names the next state when a Task fails. That is the grown-up version of “the model will retry.” Name the catcher. Name Next.
Pick the rail that matches where the customer already lives. The cage does not change.
Where does n8n fit as the cage?
n8n is a strong cage for this pattern when the customer already lives there or when webhook and ops glue dominate. The state machine concept still applies if you use Workers, Step Functions, Temporal, or a custom runner. The canvas is not the control plane. The control plane is named states, guards, a ceiling, and an escalate path you actually page.
| Thin-machine state | n8n shape | Watch-out |
|---|---|---|
intake | Webhook or queue + validation | Lease the job_id before you fan out |
plan / act model calls | Bounded AI nodes with timeouts | Do not let one Agent node own the whole lifecycle |
| Tools | HTTP Request through your runner | No ad-hoc credentials in the prompt |
| Guards | IF / Switch on verdict, counters, error class | Keep the top-level readable |
escalate / abort | Error workflow + Stop and Error | Error Trigger is a separate workflow; set it in Workflow Settings |
| Human gate | Wait on webhook / form, or Send-and-Wait | Turn on Limit Wait Time; hanging Wait is a silent escalate you forgot |
| Idempotency / leases | Static data, Redis, or your store | Duplicate delivery must no-op |
n8n’s Wait node pauses and offloads execution data to the database, then reloads when the resume condition is met — time, webhook, or form. That is a real pause, not a sleep in a prompt. Short waits historically stayed in memory (the docs still describe a sub-65-second in-process path on some versions); do not bet a human approval on an in-memory timer. Set a limit. On expiry, emit timeout and escalate.
Keep model calls inside bounded nodes. Do not let a single “AI node” own intake through done with hidden retries. The rail should be readable by an engineer who does not speak prompt. I have collaborated with the n8n team. Readable beats clever. If only one contractor understands the workflow, you do not have an operable machine.
- Each state is a node group or sub-workflow with a name a stranger can say out loud
- Error workflow attached and tested with Stop and Error
- Wait / approval has Limit Wait Time
- Agent node (if you use one) has max iterations and sits inside
act, not around the whole job
Spurlock Studios uses n8n heavily as that rail. The offer page is /agentic when the job is a loop; the cheaper shape is still a workflow.
How do you persist, pause, and resume without losing the run?
A machine that forgets its state on restart is a script. Production loops pause for humans, die with the process, and get replayed by the queue. Persist at state boundaries.
| Need | LangGraph | n8n | Temporal / Step Functions |
|---|---|---|---|
| Snapshot after each step | Checkpointer + thread_id | Execution data; Wait offloads to DB | Platform history |
| Crash resume | Replay from last checkpoint | Waiting executions reload | Replay / redrive |
| Human pause | interrupt + Command(resume=…) | Wait + $execution.resumeUrl | Callback token / signal |
| Time travel / debug | Checkpoint history | Execution log | Event history |
| Cross-run memory | Store (not the checkpointer) | Your DB, not the canvas | Your DB |
LangGraph is explicit: a checkpoint is a snapshot at each super-step, organized into threads. Compile with a checkpointer or you cannot do HITL, time travel, or fault-tolerant execution. In-memory savers are for tests. Production wants Postgres (or the Agent Server that hides the same job).
Resume rules that keep ops deterministic:
- Resume enters a named event (
human_approve,human_reject,timeout), not “continue the vibe.” human_approvereturns toactor toevaluateon the final artifact. Approvals that skip evaluation recreate self-grading with a human rubber stamp.human_rejectgoes torevise(if ceiling remains) orescalate.- Duplicate resume delivery is idempotent. Two clicks on Approve do not double-write.
If you cannot draw those four arrows, do not ship the Wait node yet.
How do you keep writes idempotent at state boundaries?
Webhooks and retries will re-enter states. Two webhooks can start two runs for one logical job. Determinism at the control plane dies the moment both act sequences look “correct” in isolation and wrong together.
| Risk | What you do at intake | What you do at act |
|---|---|---|
| Double start | Lease on job_id; second arrival joins or no-ops | Never start a second write sequence |
| Double delivery of the same write | — | Idempotency key = job + state + intent |
| Retry after timeout | Same lease, same thread / execution id | Check store; success-without-redo |
| Child machine re-entry | Inherit parent job_id + child suffix | Child writes only if parent sandbox allows |
Compute the key from job identity + state + intent. Before a hard write in act, check the store. Duplicate delivery returns success-without-redo, not a second charge or a second email.
Temporal’s saga pattern is the same discipline with compensation attached: register the undo before the forward activity, run compensations in reverse, and make every compensation idempotent — including the case where the forward activity never completed. Agents do not remove distributed-systems thinking. They add a non-deterministic planner on top of it.
-
job_idlease atintake - Idempotency key on every hard write
- Compensation named for every irreversible-adjacent write, or the write is behind a human
- Duplicate webhook test in CI
Pair leases with keys. That pairing is what people mean when they ask for deterministic multi-agent systems.
What happens when act must touch two systems?
Use a saga mindset. Write A with an idempotency key. Write B. If B fails, compensate A if you can. If you cannot compensate (email already sent, public post already live), escalate with a repair checklist. Do not let the model “try the other order” as a surprise. Order is an edge.
| Pattern | When it is legal | Failure move |
|---|---|---|
Single write in act | One system, one intent | Retry with same key, or escalate |
| Saga (A then B) | B depends on A; A is compensatable | Compensate A; if compensation fails, escalate |
| Parallel writes | Independent, both idempotent | Join; any fail → compensate the successes you can |
| Irreversible then anything | Almost never in a pilot | Human gate before the irreversible write |
Temporal’s saga guide is the citation I want on the wall: compensations registered before execution, reverse order, idempotent, able to no-op if the forward step never happened. See Saga pattern. If your agent runner cannot say those sentences, do not give it two write tools in one act.
Define whether revise may re-enter plan or only act. Re-planning is powerful and expensive. Acting on the same plan with failure evidence is usually enough for a pilot. Document the choice in the graph id, not in Slack.
- List the writes in order.
- Mark each compensatable or irreversible.
- Irreversible writes get a human state or they do not ship in the pilot.
- Encode the saga as edges, not as a paragraph in the system prompt.
How do you test the machine, not the model?
Unit-test transitions. Fault-inject tool errors. Load-test that budgets trip. These tests catch regressions prompts will never show.
XState’s primer lists “simple to test” as a benefit because the machine is deterministic: you can test all possible states and transitions. That is the point of the cage. See state machines and statecharts. LangGraph’s compile step even checks structure (no orphaned nodes) before you run — Graph API. Steal both habits.
| Test | Input | Expected |
|---|---|---|
| Happy path | valid intake, eval_pass first time | intake → plan → act → evaluate → done |
| One revise | eval_fail then eval_pass | counter = 1; ends done |
| Ceiling | eval_fail × (ceiling + 1) | escalate, never a fourth revise |
| Illegal wish | model emits skip_eval | ignored; stay or legal path only |
| Tool auth fail | tool_err class auth | abort or escalate, not done |
| Timeout | no model return | timeout → escalate / abort |
| Double webhook | two start for one job_id | one act sequence |
| Duplicate write | retry same key | success-without-redo |
| Human reject | human_reject | revise or escalate, not done |
| Budget trip | spend at cap mid-act | abort or escalate with cost |
Do not substitute a golden-set pass rate for this table. A model can look good on samples while the graph silently gained a new edge. Version the graph id in traces so you can compare scores across versions.
- Transition table is executable (JSON/YAML → tests)
- Fault injection for
tool_errandtimeout - Ceiling test fails the build if someone “helpfully” raises it in a prompt
- Diagram generated from the same file the runner loads
When someone asks “can evaluate call email.send?” the answer is in the table, not in folklore.
What anti-patterns break the cage?
These are the ones I keep seeing after 20,000+ hours in this work. Each one turns a machine back into a chat log.
| Anti-pattern | What it looks like | What you do instead |
|---|---|---|
| Hidden loop in the prompt | “Keep going until perfect” | Counter on the machine; ceiling → escalate |
| Worker-set terminal | Worker prints DONE | Only evaluator or human marks done |
| God state | One mega-state that plans, acts, and judges | Split; writes only in act |
| Retry storm | Provider retries + your retries + model retries | One budget owner |
| Illegal edge as a string | Model returns next_state: "done" and you honor it | Adapter → events; graph picks to |
| Guard with side effects | “Guard” that sends Slack | Guards stay pure; actions fire after the transition |
| Pause without a limit | Wait / interrupt forever | Limit Wait Time; timeout is an event |
| Nesting without budget inheritance | Child librarian with its own unlimited loop | Child returns a package; parent owns spend |
| Hot-edit of production edges | One bad run → live rewire | Golden case + versioned graph id |
| Self-grade | Same model marks eval_pass | Independent judge; no writes in evaluate |
“While loop in the prompt” fails audits. Auditors and careful buyers ask where the stop condition lives. “The model decides” is not an answer. A state machine with counters is. That distinction is what “deterministic multi-agent systems” has to mean if you want to operate them.
Teams start with one mega AI node. Refactor into named states when they cannot answer where side effects occur. Do the refactor before volume rises.
How do you evolve the graph without hot-editing production?
Additive changes are safer than rewiring terminals. A new optional human_approve state is an add. Changing evaluate → done so the worker can skip the judge is a regression wearing a ticket.
| Change | Risk | Ship path |
|---|---|---|
| New optional state on an existing edge | Low | Version graph id; add golden case |
| New event in the vocabulary | Low if unused edges stay closed | Document; test unknown events no-op |
New write tool on act | Medium | Allowlist + idempotency + sandbox test |
| Rewire a terminal | High | Treat like a prompt change: golden set, then promote |
| Raise the revision ceiling | High | Needs a failure review, not a Friday tweak |
Let evaluate write | Forbidden | Do not |
Version the graph id in every trace. Never hot-edit production transitions based on one bad run. Add a golden case and ship through the same path you use for prompt changes. Temporal’s workflow docs say the quiet part: once executions depend on a definition, code changes can cause non-deterministic replay; you need a versioning story. Agent graphs are the same, even if your rail is n8n.
Encode the machine so humans can read it. Write a small JSON/YAML table of from, event, to, guards. Generate the diagram from that file. Guards failing should produce explicit reason codes on the way to escalate or abort.
Event vocabulary worth keeping small: start, planned, tool_ok, tool_err, eval_pass, eval_fail, budget_low, timeout, human_approve, human_reject. Models emit artifacts. Adapters translate. A small vocabulary is what makes the graph testable.
A parent job may spawn a child machine (retrieval-only librarian, for example). Children return a package. They do not write customer-visible systems unless the parent’s sandbox allows it. Nesting without budget inheritance is a cost bug.
What does a five-day pilot actually ship?
In a $1,500 · 5-day Spurlock Studios pilot we implement a thin machine for one job: intake, act, evaluate, revise×N, done/escalate. Fancy parallel states wait until the thin machine clears the golden set. You should feel the cage before we elaborate it.
| Day | What exists at end of day | What does not |
|---|---|---|
| 1 | Job contract, states, edges, allowlist, ceiling | Fleet, multi-agent chat |
| 2 | Runner or n8n rail; act sandboxed; idempotency keys | Production writes on real customers |
| 3 | Evaluator hooked; traces; reason codes | Self-grading |
| 4 | Golden set + transition tests + one fault-injection | Raised ceilings “just in case” |
| 5 | Escalate package + kill switch + you can operate it | A rewrite of a flowchart that already fit on one slide |
Walkthrough that means the machine is doing its job — not a chat log:
A ticket arrives at intake; schema validates; a budget is attached. plan proposes two tool calls. act runs them under sandbox. evaluate fails a citation rule. revise 1 redrafts. evaluate passes. done writes an internal note idempotently. The trace shows each state. Finance sees the spend. If evaluate had failed through the ceiling, escalate would package evidence for a human. Deterministic multi-agent systems use the same terminals when more roles join later.
- One job, real samples
- Thin graph enforced in the rail
- Ceiling and escalate package
- You can answer “what state is this run in?” without opening a prompt file
If the path is already a flowchart, do not buy this week. Read When Not to Build an Agent and automate. If the path branches and you can still write pass/fail, the cage is how you ship. Parent map: operating manual. Offer: /agentic.
FAQ
What is an agent loop state machine?
It is an explicit graph of states and transitions that wraps model and tool calls so side effects, retries, evaluation, and termination follow rules you enforce — not vibes from the model. The model gets freedom inside a state. The runner owns the edges. If you cannot name the next state from (state, event, guards), you do not have a machine yet.
How do you get deterministic multi-agent systems?
Make control flow, policy, accounting, and terminality deterministic. Accept that tokens vary. Coordinate agents with typed handoffs that enter known states, not free-form shared chats. LangGraph, XState, Step Functions, and Temporal all document that split; none of them promise identical tokens.
Why use n8n for agent state machines?
n8n gives durable execution, webhooks, error workflows, and Wait / approval nodes — the boring reliability layer around non-deterministic model steps. It is one good rail, not the only one. The Agent node must sit inside act with a max-iteration cap; it must not own the whole lifecycle.
How many states do we need?
Enough to separate intake, planning, acting, evaluating, revising, and terminal outcomes. Start with the thin eight. Split a state when a node becomes untestable or when side-effect classes collide. Add parallel regions only after the thin graph clears a golden set.
What happens when the model wants an illegal transition?
The runner ignores the wish and either stays, forces a legal path, asks for a replan inside the current state, or escalates. The model does not get to rewrite the graph at runtime. Adapters emit events; guards pick to. A string like next_state: "done" is evidence for the trace, not a command.
How does Spurlock Studios implement this on a pilot?
We ship a minimal enforced graph with budgets, a revision ceiling, and escalate in five days for one job, then expand. The week is $1,500. See /agentic and the operating manual. If the path is already known, we will tell you to automate instead.
CTA
If the loop is real, buy the cage — named states, legal transitions, a ceiling, and an escalate path — not another prompt that says “keep going.”
What questions does this article answer?
- What is an agent loop state machine?
- It is an explicit graph of states and transitions that wraps model and tool calls so side effects, retries, evaluation, and termination follow rules you enforce — not vibes from the model. The model gets freedom inside a state. The runner owns the edges. If you cannot name the next state from `(state, event, guards)`, you do not have a machine yet.
- How do you get deterministic multi-agent systems?
- Make control flow, policy, accounting, and terminality deterministic. Accept that tokens vary. Coordinate agents with typed handoffs that enter known states, not free-form shared chats. LangGraph, XState, Step Functions, and Temporal all document that split; none of them promise identical tokens.
- Why use n8n for agent state machines?
- n8n gives durable execution, webhooks, error workflows, and Wait / approval nodes — the boring reliability layer around non-deterministic model steps. It is one good rail, not the only one. The Agent node must sit inside `act` with a max-iteration cap; it must not own the whole lifecycle.
- How many states do we need?
- Enough to separate intake, planning, acting, evaluating, revising, and terminal outcomes. Start with the thin eight. Split a state when a node becomes untestable or when side-effect classes collide. Add parallel regions only after the thin graph clears a golden set.
- What happens when the model wants an illegal transition?
- The runner ignores the wish and either stays, forces a legal path, asks for a replan inside the current state, or escalates. The model does not get to rewrite the graph at runtime. Adapters emit events; guards pick `to`. A string like `next_state: "done"` is evidence for the trace, not a command.
- How does Spurlock Studios implement this on a pilot?
- We ship a minimal enforced graph with budgets, a revision ceiling, and escalate in five days for one job, then expand. The week is $1,500. See [/agentic](/agentic) and the [operating manual](/blog/agentic-systems-operating-manual). If the path is already known, we will tell you to automate instead.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.