Spurlock Studios
Contact
Share LinkedIn X
Two clipped paper packets. Thesis: STATE MACHINES AGENT LOOPS DETERMINISM.

Models are non-deterministic. Operations cannot be. The way you reconcile those facts is a state machine: the model gets freedom to plan, draft, and pick tools inside a named state; the runner decides which transitions are legal, when a write may happen, how many revisions you buy, and when the run must escalate or abort.

This spoke sits under the Agentic Systems Operating Manual. If you are still shopping for whether you need an agent at all, read When Not to Build an Agent first.

The short answer

  • Freedom is intra-state. Inside plan, act, or revise, the model may propose content and tool calls. It does not get to invent a new state or skip evaluation.
  • Transitions are extra-state and legal or they do not happen. The runner maps (state, event, guards) to a next state. Illegal wishes are ignored.
  • Revision has a ceiling. Three is a good pilot default. Hit it and you escalate with evidence, not another hidden loop.
  • Escalate is a terminal success of a kind. The system knew it was out of its depth while a human could still help.
  • Identical tokens are not the bar. Deterministic multi-agent means control flow, policy, accounting, and terminals. Chasing bit-identical model output wastes budget.

What is an agent loop state machine?

An agent loop state machine is an explicit set of named states and allowed transitions that wrap model calls and tool use. The model proposes inside a state. The machine decides whether the next state is legal given the evaluator verdict, the remaining budget, the revision counter, and the error class.

Without that wrapper you get unbounded retries, duplicate side effects, “done” declared by the worker, and no clean place for a human to intervene. I have spent 20,000+ hours on agentic systems and built 500+ automations. The loops that survived contact with production all had a graph an engineer could read without speaking prompt.

ObjectWho owns itWhat “freedom” means
Prompt / draft / tool argsModel, inside the current statePropose content and allowlisted calls
Next stateRunner / graphOnly listed edges fire
Side-effect permissionState definitionWrites only in act (or a named write child)
StopCounters + terminalsdone, escalate, or abort — never “keep going”
EvidenceTrace storeEvery transition leaves a reason code

LangGraph’s own Graph API says the same thing in vendor vocabulary: you model the workflow as a graph of state, nodes, and edges — “nodes do the work, edges tell what to do next.” See Graph API overview. XState’s primer is blunter about the control plane: transitions are deterministic; each combination of state and event always points to the same next state. That line lives on What are state machines and statecharts?.

  • Named states written down (not implied by a mega prompt)
  • Allowed edges written down (from, event, to, guards)
  • Side-effect class attached to each state
  • At least one terminal that is not done
  • A counter the model cannot reset

If any box is empty, you have a directed suggestion. Not a machine.

Why do models get freedom inside states but not between them?

Because the useful non-determinism is local. You want the model to try a different retrieval query, a tighter draft, or a different allowlisted tool. You do not want it to invent skip_eval, email a customer from plan, or declare victory because the last sentence sounded confident.

XState calls the between-states rule by its real name. A transition is a change from one finite state to another, triggered by an event. Enabled transitions are selected from the current state; if none are enabled, the state does not change. That is the cage. The model emits artifacts. An adapter turns outcomes into events the machine already understands.

LayerNon-deterministic is fineMust be deterministic
TokensWording, tool-arg phrasing, plan orderSchema of the artifact you accept
Tool choiceAmong the allowlist for this stateWhether the state may call tools at all
TimingModel latency inside a timeoutTimeout → named event (timeout)
QualityFirst-pass draftEvaluator verdict → eval_pass / eval_fail
StopNeverCeiling, budget, policy, human

LangGraph splits the same idea into workflow vs agent. Workflows and agents says workflows have predetermined code paths and “operate in a certain order”; agents are dynamic and “define their own processes and tool usage.” Production agent loops are a mix: predetermined edges, dynamic work inside nodes. If the whole job is a predetermined path, you do not need this spoke. You need a workflow. That no lives in When Not to Build an Agent.

  1. Name the state the model is in.
  2. Give it a contract: input schema, output schema, tool allowlist, timeout.
  3. Translate its output into a small event (planned, tool_ok, tool_err, eval_fail).
  4. Let guards pick the next state. Do not let the model name the next state as a string you then eval.

The model is a guest in the room. The floor plan is yours.

Start with eight states. Fancy parallel regions wait until this graph clears a golden set.

StateAllowed workForbidden work
intakeValidate job contract, attach budget and tool allowlistModel calls, production writes
planModel proposes stepsProduction writes
actAllowlisted tools; side effects only hereDeclaring done
evaluateIndependent judge; no writesPatching CRM, sending mail
reviseWorker edits with failure evidence; increment counterResetting the counter
doneTerminal success; store receiptsFurther tools
escalateTerminal handoff with a packageQuiet retry
abortTerminal stop on budget, policy, or unrecoverable error“Best effort” write on the way out

Legal transitions (simplified):

FromEventToGuard
intakestartplanschema valid, budget > 0
intakerejectabortout of scope or auth fail
planplannedactsteps in allowlist
plantimeout / policyescalate or abort—
acttool_okevaluate—
acttool_errescalate or abortby error class
evaluateeval_passdone—
evaluateeval_failreviserevisions < ceiling
evaluateeval_failescalaterevisions >= ceiling
reviserevisedactceiling not hit
anything seriousbudget_low / policyabort—

The exact graph can vary. The requirement is that it is written down and enforced in code or in the workflow rail. AWS Step Functions makes the same demand in Amazon States Language: execution starts at StartAt, each state names Next or ends, and the run stops at Succeed, Fail, or "End": true. See Transitions in state machines. If your “graph” is a while-loop in a system prompt, you do not have Next. You have hope.

W3C already standardized the older, vendor-neutral version of this idea. SCXML (Recommendation, 1 September 2015) is a general-purpose event-based state machine language. You do not need to author SCXML. You do need the same objects: states, events, transitions, and a legal configuration.

What does “deterministic multi-agent” actually mean?

People say “deterministic multi-agent” when they want predictability. You will not get bit-identical model outputs. You can get four things that matter to ops.

Kind of determinismWhat it looks likeWhat it is not
Control flowSame (state, event, guards) → same next stateSame tokens
PolicyAllowlists, caps, ceilings, deny classes“The model usually behaves”
AccountingCost, latency, tools, and reason codes always recordedA chat log you grep later
TerminalityEvery run ends in done, escalate, or abortA hanging Wait with no limit

That is the bar. Temporal’s workflow docs make the split operational: workflow definitions must be deterministic so replay works; LLM calls, API calls, and other external work go in Activities, outside the replay path. Put the non-deterministic guest in a room with a door. Do not let it rebuild the building.

When a second role joins later (librarian, writer, critic), it enters a known state with a typed package. It does not join a free-form shared chat and “figure it out.” Coordination is an edge, not a vibe.

  • I can unit-test (state, event) → next without calling a model
  • I can name the three terminals
  • I can name the budget owner (one, not three retry layers)
  • A new agent role would enter through intake or a child machine, not through a Slack thread

If you cannot tick those, you are still in demo land. The parent map for evaluators, sandboxes, and kill switches is the operating manual.

Where do side effects live?

Only in act — or in a named write child of act. This rule prevents half the horror stories. Planning tokens do not email customers. Evaluation does not patch CRM. Revision drafts land in scratch until act applies them under sandbox rules.

XState splits effects the same way the rest of serious FSM work does. Guards are pure, synchronous predicates that return true or false; they decide whether a transition is enabled. Actions are fire-and-forget side effects that run because a transition was taken. If your architecture lets any state call any tool, you do not have a state machine. You have a directed suggestion.

StateMay call tools?May write customer-visible systems?Notes
intakeNoNoSchema and auth only
planRead-only retrieval if you mustNoPrefer retrieval as a child with no write tools
actYes, allowlistedYes, under sandbox + idempotencyThe only write room
evaluateRead-only checksNoIndependent of the worker
reviseNo writesNoEdits the artifact, increments the counter
done / escalate / abortReceipts / package onlyNo new business writesTerminal

If a tool result must change the next state, the tool returns to act, act emits tool_ok or tool_err, and the edge decides. The tool does not jump the graph.

  1. Classify every tool: read, write, irreversible.
  2. Bind write and irreversible tools only to act.
  3. Put irreversible classes behind a human gate or a deny-by-default policy.
  4. Refuse to compile the graph if a write tool is reachable from plan or evaluate.

LangGraph will let a node perform a side-effect — the Graph API says nodes “perform some computation or side-effect.” That is permission to be careful, not permission to be sloppy. You still decide which nodes are allowed to write.

How do revision ceilings and escalate packages work?

Unbounded revise loops are how you discover a task is impossible after the bill already moved. Cap revisions. Three is a good pilot default. On ceiling, stop buying another turn.

CounterPilot defaultOn trip
Revisions3evaluate → escalate (not another revise)
Tool calls in actPer-job allowlist + maxtool_err / budget_low → escalate or abort
Wall-clock per stateExplicit timeouttimeout event
SpendHard cap on the jobabort or escalate with cost in the package
Human waitLimit Wait Time / interrupt SLAtimeout → escalate, never hang forever

XState guards are the right shape for the ceiling: revisions < 3 is a pure predicate. First matching guarded transition wins; put the specific guard before the default. Do not hide the counter in the prompt (“try a few times”). The model will try until the card declines.

Escalate package — ship all of this, every time:

  • Job contract (what “done” meant)
  • Artifacts so far
  • Evaluator failures with evidence
  • Tools called and outcomes
  • Cost and latency
  • Reason code (ceiling, policy, timeout, auth, schema)
  • Recommended human action, if you have one

Escalation is success of a kind. The system knew it was out of its depth while someone could still help. A silent retry after the ceiling is a lie.

LangGraph’s interrupt is the vendor form of “pause and wait for a human.” interrupt() saves graph state through the persistence layer and waits until you resume with Command. That is an escalate gate when you still might return to act. A true escalate terminal is different: the machine is done; a human owns the next business action. Do not confuse a pause with a handoff. Name both.

  • Ceiling is an integer in state, not a vibe in the prompt
  • eval_fail + ceiling → escalate, tested
  • Package schema is versioned
  • On-call can open a run and see why it stopped in under a minute

How do LangGraph, XState, and Step Functions encode the same cage?

You do not owe any one framework. You owe the objects. The vendors already documented them.

Object you needLangGraphXStateStep Functions
Shared snapshotState schema + reducersContext + finite stateExecution input / variables
WorkNodesActions / invoked actorsTask
Next stepEdges, including conditionalTransitions on eventsNext, Choice
PredicateYour function on a conditional edgeGuardsChoice rules
Side effectNode you marked as write-capableActions on the taken transitionTask that calls the write
Persist / resumeCheckpointers + thread idPersisted actor snapshot (your store)Service-managed execution history
Human pauseInterruptsEvent from the human, or a wait stateWait-for-callback / .waitForTaskToken patterns
TerminalEND / your done nodeFinal stateSucceed, Fail, "End": true

LangGraph persistence is explicit: checkpointers save a thread’s graph state as checkpoints (short-term, for HITL, time travel, fault tolerance); stores hold long-term data across threads. Compile without a checkpointer and you cannot pause or survive a crash. That is not an implementation detail. That is the difference between a notebook demo and a run you can page.

XState’s machines doc lists the same kit: actions, actors, guards, delayed transitions. Use it as a checklist even if you never import the library. If you cannot point to each object in your runner, you are missing a piece.

Step Functions bills and traces by transition. Error handling treats retries as transitions; Catch names the next state when a Task fails. That is the grown-up version of “the model will retry.” Name the catcher. Name Next.

Pick the rail that matches where the customer already lives. The cage does not change.

Where does n8n fit as the cage?

n8n is a strong cage for this pattern when the customer already lives there or when webhook and ops glue dominate. The state machine concept still applies if you use Workers, Step Functions, Temporal, or a custom runner. The canvas is not the control plane. The control plane is named states, guards, a ceiling, and an escalate path you actually page.

Thin-machine staten8n shapeWatch-out
intakeWebhook or queue + validationLease the job_id before you fan out
plan / act model callsBounded AI nodes with timeoutsDo not let one Agent node own the whole lifecycle
ToolsHTTP Request through your runnerNo ad-hoc credentials in the prompt
GuardsIF / Switch on verdict, counters, error classKeep the top-level readable
escalate / abortError workflow + Stop and ErrorError Trigger is a separate workflow; set it in Workflow Settings
Human gateWait on webhook / form, or Send-and-WaitTurn on Limit Wait Time; hanging Wait is a silent escalate you forgot
Idempotency / leasesStatic data, Redis, or your storeDuplicate delivery must no-op

n8n’s Wait node pauses and offloads execution data to the database, then reloads when the resume condition is met — time, webhook, or form. That is a real pause, not a sleep in a prompt. Short waits historically stayed in memory (the docs still describe a sub-65-second in-process path on some versions); do not bet a human approval on an in-memory timer. Set a limit. On expiry, emit timeout and escalate.

Keep model calls inside bounded nodes. Do not let a single “AI node” own intake through done with hidden retries. The rail should be readable by an engineer who does not speak prompt. I have collaborated with the n8n team. Readable beats clever. If only one contractor understands the workflow, you do not have an operable machine.

  • Each state is a node group or sub-workflow with a name a stranger can say out loud
  • Error workflow attached and tested with Stop and Error
  • Wait / approval has Limit Wait Time
  • Agent node (if you use one) has max iterations and sits inside act, not around the whole job

Spurlock Studios uses n8n heavily as that rail. The offer page is /agentic when the job is a loop; the cheaper shape is still a workflow.

How do you persist, pause, and resume without losing the run?

A machine that forgets its state on restart is a script. Production loops pause for humans, die with the process, and get replayed by the queue. Persist at state boundaries.

NeedLangGraphn8nTemporal / Step Functions
Snapshot after each stepCheckpointer + thread_idExecution data; Wait offloads to DBPlatform history
Crash resumeReplay from last checkpointWaiting executions reloadReplay / redrive
Human pauseinterrupt + Command(resume=…)Wait + $execution.resumeUrlCallback token / signal
Time travel / debugCheckpoint historyExecution logEvent history
Cross-run memoryStore (not the checkpointer)Your DB, not the canvasYour DB

LangGraph is explicit: a checkpoint is a snapshot at each super-step, organized into threads. Compile with a checkpointer or you cannot do HITL, time travel, or fault-tolerant execution. In-memory savers are for tests. Production wants Postgres (or the Agent Server that hides the same job).

Resume rules that keep ops deterministic:

  1. Resume enters a named event (human_approve, human_reject, timeout), not “continue the vibe.”
  2. human_approve returns to act or to evaluate on the final artifact. Approvals that skip evaluation recreate self-grading with a human rubber stamp.
  3. human_reject goes to revise (if ceiling remains) or escalate.
  4. Duplicate resume delivery is idempotent. Two clicks on Approve do not double-write.

If you cannot draw those four arrows, do not ship the Wait node yet.

How do you keep writes idempotent at state boundaries?

Webhooks and retries will re-enter states. Two webhooks can start two runs for one logical job. Determinism at the control plane dies the moment both act sequences look “correct” in isolation and wrong together.

RiskWhat you do at intakeWhat you do at act
Double startLease on job_id; second arrival joins or no-opsNever start a second write sequence
Double delivery of the same write—Idempotency key = job + state + intent
Retry after timeoutSame lease, same thread / execution idCheck store; success-without-redo
Child machine re-entryInherit parent job_id + child suffixChild writes only if parent sandbox allows

Compute the key from job identity + state + intent. Before a hard write in act, check the store. Duplicate delivery returns success-without-redo, not a second charge or a second email.

Temporal’s saga pattern is the same discipline with compensation attached: register the undo before the forward activity, run compensations in reverse, and make every compensation idempotent — including the case where the forward activity never completed. Agents do not remove distributed-systems thinking. They add a non-deterministic planner on top of it.

  • job_id lease at intake
  • Idempotency key on every hard write
  • Compensation named for every irreversible-adjacent write, or the write is behind a human
  • Duplicate webhook test in CI

Pair leases with keys. That pairing is what people mean when they ask for deterministic multi-agent systems.

What happens when act must touch two systems?

Use a saga mindset. Write A with an idempotency key. Write B. If B fails, compensate A if you can. If you cannot compensate (email already sent, public post already live), escalate with a repair checklist. Do not let the model “try the other order” as a surprise. Order is an edge.

PatternWhen it is legalFailure move
Single write in actOne system, one intentRetry with same key, or escalate
Saga (A then B)B depends on A; A is compensatableCompensate A; if compensation fails, escalate
Parallel writesIndependent, both idempotentJoin; any fail → compensate the successes you can
Irreversible then anythingAlmost never in a pilotHuman gate before the irreversible write

Temporal’s saga guide is the citation I want on the wall: compensations registered before execution, reverse order, idempotent, able to no-op if the forward step never happened. See Saga pattern. If your agent runner cannot say those sentences, do not give it two write tools in one act.

Define whether revise may re-enter plan or only act. Re-planning is powerful and expensive. Acting on the same plan with failure evidence is usually enough for a pilot. Document the choice in the graph id, not in Slack.

  1. List the writes in order.
  2. Mark each compensatable or irreversible.
  3. Irreversible writes get a human state or they do not ship in the pilot.
  4. Encode the saga as edges, not as a paragraph in the system prompt.

How do you test the machine, not the model?

Unit-test transitions. Fault-inject tool errors. Load-test that budgets trip. These tests catch regressions prompts will never show.

XState’s primer lists “simple to test” as a benefit because the machine is deterministic: you can test all possible states and transitions. That is the point of the cage. See state machines and statecharts. LangGraph’s compile step even checks structure (no orphaned nodes) before you run — Graph API. Steal both habits.

TestInputExpected
Happy pathvalid intake, eval_pass first timeintake → plan → act → evaluate → done
One reviseeval_fail then eval_passcounter = 1; ends done
Ceilingeval_fail × (ceiling + 1)escalate, never a fourth revise
Illegal wishmodel emits skip_evalignored; stay or legal path only
Tool auth failtool_err class authabort or escalate, not done
Timeoutno model returntimeout → escalate / abort
Double webhooktwo start for one job_idone act sequence
Duplicate writeretry same keysuccess-without-redo
Human rejecthuman_rejectrevise or escalate, not done
Budget tripspend at cap mid-actabort or escalate with cost

Do not substitute a golden-set pass rate for this table. A model can look good on samples while the graph silently gained a new edge. Version the graph id in traces so you can compare scores across versions.

  • Transition table is executable (JSON/YAML → tests)
  • Fault injection for tool_err and timeout
  • Ceiling test fails the build if someone “helpfully” raises it in a prompt
  • Diagram generated from the same file the runner loads

When someone asks “can evaluate call email.send?” the answer is in the table, not in folklore.

What anti-patterns break the cage?

These are the ones I keep seeing after 20,000+ hours in this work. Each one turns a machine back into a chat log.

Anti-patternWhat it looks likeWhat you do instead
Hidden loop in the prompt“Keep going until perfect”Counter on the machine; ceiling → escalate
Worker-set terminalWorker prints DONEOnly evaluator or human marks done
God stateOne mega-state that plans, acts, and judgesSplit; writes only in act
Retry stormProvider retries + your retries + model retriesOne budget owner
Illegal edge as a stringModel returns next_state: "done" and you honor itAdapter → events; graph picks to
Guard with side effects“Guard” that sends SlackGuards stay pure; actions fire after the transition
Pause without a limitWait / interrupt foreverLimit Wait Time; timeout is an event
Nesting without budget inheritanceChild librarian with its own unlimited loopChild returns a package; parent owns spend
Hot-edit of production edgesOne bad run → live rewireGolden case + versioned graph id
Self-gradeSame model marks eval_passIndependent judge; no writes in evaluate

“While loop in the prompt” fails audits. Auditors and careful buyers ask where the stop condition lives. “The model decides” is not an answer. A state machine with counters is. That distinction is what “deterministic multi-agent systems” has to mean if you want to operate them.

Teams start with one mega AI node. Refactor into named states when they cannot answer where side effects occur. Do the refactor before volume rises.

How do you evolve the graph without hot-editing production?

Additive changes are safer than rewiring terminals. A new optional human_approve state is an add. Changing evaluate → done so the worker can skip the judge is a regression wearing a ticket.

ChangeRiskShip path
New optional state on an existing edgeLowVersion graph id; add golden case
New event in the vocabularyLow if unused edges stay closedDocument; test unknown events no-op
New write tool on actMediumAllowlist + idempotency + sandbox test
Rewire a terminalHighTreat like a prompt change: golden set, then promote
Raise the revision ceilingHighNeeds a failure review, not a Friday tweak
Let evaluate writeForbiddenDo not

Version the graph id in every trace. Never hot-edit production transitions based on one bad run. Add a golden case and ship through the same path you use for prompt changes. Temporal’s workflow docs say the quiet part: once executions depend on a definition, code changes can cause non-deterministic replay; you need a versioning story. Agent graphs are the same, even if your rail is n8n.

Encode the machine so humans can read it. Write a small JSON/YAML table of from, event, to, guards. Generate the diagram from that file. Guards failing should produce explicit reason codes on the way to escalate or abort.

Event vocabulary worth keeping small: start, planned, tool_ok, tool_err, eval_pass, eval_fail, budget_low, timeout, human_approve, human_reject. Models emit artifacts. Adapters translate. A small vocabulary is what makes the graph testable.

A parent job may spawn a child machine (retrieval-only librarian, for example). Children return a package. They do not write customer-visible systems unless the parent’s sandbox allows it. Nesting without budget inheritance is a cost bug.

What does a five-day pilot actually ship?

In a $1,500 · 5-day Spurlock Studios pilot we implement a thin machine for one job: intake, act, evaluate, revise×N, done/escalate. Fancy parallel states wait until the thin machine clears the golden set. You should feel the cage before we elaborate it.

DayWhat exists at end of dayWhat does not
1Job contract, states, edges, allowlist, ceilingFleet, multi-agent chat
2Runner or n8n rail; act sandboxed; idempotency keysProduction writes on real customers
3Evaluator hooked; traces; reason codesSelf-grading
4Golden set + transition tests + one fault-injectionRaised ceilings “just in case”
5Escalate package + kill switch + you can operate itA rewrite of a flowchart that already fit on one slide

Walkthrough that means the machine is doing its job — not a chat log:

A ticket arrives at intake; schema validates; a budget is attached. plan proposes two tool calls. act runs them under sandbox. evaluate fails a citation rule. revise 1 redrafts. evaluate passes. done writes an internal note idempotently. The trace shows each state. Finance sees the spend. If evaluate had failed through the ceiling, escalate would package evidence for a human. Deterministic multi-agent systems use the same terminals when more roles join later.

  • One job, real samples
  • Thin graph enforced in the rail
  • Ceiling and escalate package
  • You can answer “what state is this run in?” without opening a prompt file

If the path is already a flowchart, do not buy this week. Read When Not to Build an Agent and automate. If the path branches and you can still write pass/fail, the cage is how you ship. Parent map: operating manual. Offer: /agentic.

FAQ

What is an agent loop state machine?

It is an explicit graph of states and transitions that wraps model and tool calls so side effects, retries, evaluation, and termination follow rules you enforce — not vibes from the model. The model gets freedom inside a state. The runner owns the edges. If you cannot name the next state from (state, event, guards), you do not have a machine yet.

How do you get deterministic multi-agent systems?

Make control flow, policy, accounting, and terminality deterministic. Accept that tokens vary. Coordinate agents with typed handoffs that enter known states, not free-form shared chats. LangGraph, XState, Step Functions, and Temporal all document that split; none of them promise identical tokens.

Why use n8n for agent state machines?

n8n gives durable execution, webhooks, error workflows, and Wait / approval nodes — the boring reliability layer around non-deterministic model steps. It is one good rail, not the only one. The Agent node must sit inside act with a max-iteration cap; it must not own the whole lifecycle.

How many states do we need?

Enough to separate intake, planning, acting, evaluating, revising, and terminal outcomes. Start with the thin eight. Split a state when a node becomes untestable or when side-effect classes collide. Add parallel regions only after the thin graph clears a golden set.

What happens when the model wants an illegal transition?

The runner ignores the wish and either stays, forces a legal path, asks for a replan inside the current state, or escalates. The model does not get to rewrite the graph at runtime. Adapters emit events; guards pick to. A string like next_state: "done" is evidence for the trace, not a command.

How does Spurlock Studios implement this on a pilot?

We ship a minimal enforced graph with budgets, a revision ceiling, and escalate in five days for one job, then expand. The week is $1,500. See /agentic and the operating manual. If the path is already known, we will tell you to automate instead.

CTA

If the loop is real, buy the cage — named states, legal transitions, a ceiling, and an escalate path — not another prompt that says “keep going.”

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

What is an agent loop state machine?
It is an explicit graph of states and transitions that wraps model and tool calls so side effects, retries, evaluation, and termination follow rules you enforce — not vibes from the model. The model gets freedom inside a state. The runner owns the edges. If you cannot name the next state from `(state, event, guards)`, you do not have a machine yet.
How do you get deterministic multi-agent systems?
Make control flow, policy, accounting, and terminality deterministic. Accept that tokens vary. Coordinate agents with typed handoffs that enter known states, not free-form shared chats. LangGraph, XState, Step Functions, and Temporal all document that split; none of them promise identical tokens.
Why use n8n for agent state machines?
n8n gives durable execution, webhooks, error workflows, and Wait / approval nodes — the boring reliability layer around non-deterministic model steps. It is one good rail, not the only one. The Agent node must sit inside `act` with a max-iteration cap; it must not own the whole lifecycle.
How many states do we need?
Enough to separate intake, planning, acting, evaluating, revising, and terminal outcomes. Start with the thin eight. Split a state when a node becomes untestable or when side-effect classes collide. Add parallel regions only after the thin graph clears a golden set.
What happens when the model wants an illegal transition?
The runner ignores the wish and either stays, forces a legal path, asks for a replan inside the current state, or escalates. The model does not get to rewrite the graph at runtime. Adapters emit events; guards pick `to`. A string like `next_state: "done"` is evidence for the trace, not a command.
How does Spurlock Studios implement this on a pilot?
We ship a minimal enforced graph with budgets, a revision ceiling, and escalate in five days for one job, then expand. The week is $1,500. See [/agentic](/agentic) and the [operating manual](/blog/agentic-systems-operating-manual). If the path is already known, we will tell you to automate instead.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot