Agent Loop vs LLM-in-Workflow: Pick the Shape That Matches the Uncertainty
Most jobs need a workflow with one LLM step, not an agent loop. Use a three-tier test — certainty, branching, blast radius — then prove hybrid is enough.
William Spurlock Founder — Spurlock Studios Updated 18 MIN
You usually need a workflow with one LLM step — not a full agent loop. An agent earns its keep only when the next action cannot be named until mid-run, the long-tail branches explode a scripted graph, and you can afford the cost and failure modes that come with open-ended tool use. Pick the shape that matches the uncertainty, not the buzzword on the slide.
This spoke sits under the Agentic Systems Operating Manual. It owns the hybrid middle: when a scripted graph plus one bounded model call is enough, and what evidence graduates you. For the hard “no agent” cases, see when not to build an agent.
The short answer
- Tier-1: deterministic workflow — no model, or a model used only offline to design the flow.
- Tier-2 (default): workflow + one bounded LLM step (classify, extract, draft) with schema-checked I/O.
- Tier-3: agent loop — plan → act → observe → decide, with tools, budgets, and evaluators.
- Wrong shape taxes you twice: Tier-3 burns tokens on solvable graphs; Tier-1/2 silently fails when the long tail needs mid-run decisions.
- Prove Tier-2 first. Graduate to Tier-3 only when you have failure traces that a scripted graph cannot absorb without becoming a second product.
What is an agent loop vs a scripted workflow?
| Shape | Who picks the next step | Tool calls | Typical failure |
|---|---|---|---|
| Scripted workflow | You, at design time | Fixed edges in n8n / code | Missing branch, silent skip |
| Workflow + LLM step | You for routing; model for one transformation | Fixed; model never chooses tools | Schema miss, bad extract |
| Agent loop | Model + harness at runtime | Dynamic, policy-gated | Loops, wrong tool, cost blowups |
An LLM step is a node: input in, structured output out, next edge known. An agent loop is a control plane: the model may propose tools, the harness may allow or deny them, and termination is earned (done, escalate, abort) — not a final webhook hop.
If you can draw every edge on a whiteboard before the first production ticket arrives, you are still in workflow land.
That is not a Spurlock-only distinction. Anthropic’s Building effective agents essay (Dec 2024, still the citation most teams use) draws the same line: workflows are “systems where LLMs and tools are orchestrated through predefined code paths”; agents are systems where the model “dynamically direct[s] [its] own processes and tool usage.” LangGraph’s docs repeat it almost verbatim — workflows have predetermined code paths; agents “define their own processes and tool usage.”
The product name on the node does not change the physics. A Basic LLM Chain that secretly retries five times with different tools is a loop. An “agent” that always calls the same three functions in the same order is a workflow wearing a costume.
What the vendors actually say
Read the primary pages before you inherit a framework’s default node.
| Source | Workflow / LLM step | Agent loop | The line they draw |
|---|---|---|---|
| Anthropic, Building effective agents | Predefined code paths; start with one augmented LLM call | Model directs process and tools | “Find the simplest solution possible… This might mean not building agentic systems at all.” |
| OpenAI, A practical guide to building agents | Deterministic / rule-based automation | Complex judgment, brittle rulesets, heavy unstructured data | “Otherwise, a deterministic solution may suffice.” |
| LangGraph, Workflows and agents | Predetermined order | Continuous feedback loops, unpredictable problems | Mix both in one graph; do not default the whole job to a loop |
| n8n, What agents do | Chain: predetermined sequence of calls | “A chain that knows how to make decisions” | Agent node runs multiple times per execution |
Anthropic is explicit about the tax: agentic systems “often trade latency and cost for better task performance,” and autonomy means “higher costs, and the potential for compounding errors.” They tell you to optimize a single LLM call with retrieval and examples first; add multi-step agentic systems only when simpler solutions fall short.
OpenAI’s guide is the same climb from the other side. Agents belong where traditional automation already failed — refund judgment, vendor reviews that outgrew a rules file, intake that arrives as a pile of PDFs. If you cannot point at that friction, stay deterministic.
I have built 500+ automations and spent 20,000+ hours on agentic systems. The jobs that paid for a loop were the ones where a human already changed course mid-task. The jobs that did not were the ones a classifier plus three IF nodes finished on Tuesday.
Who owns the next step?
This is the only question that matters. Everything else is implementation.
| Question | Tier-2 answer | Tier-3 answer |
|---|---|---|
| What runs after this model call? | A named edge you wrote | Whatever the model proposes and the harness allows |
| How many model calls per successful job? | One (or a fixed chain you counted) | Unknown until the run ends |
| Who may call a write tool? | The workflow, after schema checks | The loop, after policy + evaluator |
| What does “done” mean? | Last node succeeded | Status in a result package: done / escalate / abort |
| Where do credentials live? | Workflow nodes with scoped secrets | Tool runner with an allowlist — never “the model has the CRM key” |
Write the answers down before you pick a node. If column two is honest, you do not have an agent problem. You have a schema problem, a prompt problem, or a missing IF node.
- I can name the next system call without looking at the payload → stay Tier-1/2
- The next call depends on free-form content I have not seen → Tier-3 candidate
- I already know the write I want; I just need fields extracted → Tier-2
- A human would open a second tab, then a third, then change the plan → Tier-3 candidate
Candidate is not permission. Candidate means “instrument Tier-2 and see if the long tail stays ugly.”
Why wrong shape blows cost and reliability
Agent loops pay a planning tax on every turn: context, tool schemas, retries, and evaluator rounds. That is fine when the alternative is a human. It is wasteful when a classifier plus three IF nodes would finish the job.
Anthropic’s warning is operational, not aesthetic. Each extra autonomous turn adds latency, tokens, and a chance that an early mistake becomes the next prompt. A workflow bounds that: you decide how many model calls a run is allowed to make. A loop does not, unless you install a budget.
Workflows fail the other way. You encode the happy path, miss the ugly path, and ship a “successful” run that wrote the wrong CRM field because no model was allowed to notice the exception.
| Wrong choice | What you feel in week two |
|---|---|
| Agent for a form extract | Token bill vs a single structured-output call; flaky tool retries |
| Workflow for messy exceptions | Escalation pile grows; humans rewrite “automation” output |
| Hybrid without schema checks | LLM step drifts; downstream nodes trust garbage |
| Agent with irreversible tools and no deny gate | Duplicate charges, public posts, deleted rows |
Cost is not only tokens. Wrong-shape agents also burn eng time debugging loops that a state machine never should have entered. I have watched teams spend a sprint chasing “the agent got stuck” when the graph had six known branches the whole time.
The three-tier test
Run every candidate job through these questions in order. Stop at the first tier that fits.
1. Certainty of the next action
- Can you name the next system call before looking at the payload? → Tier-1 or Tier-2.
- Does the next call depend on free-form content you have not seen yet? → Candidate Tier-3.
2. Branch count and long-tail rate
- Under ~10 stable branches, update the graph. Prefer Tier-2.
- Long-tail exceptions that keep inventing new branches after every release → Tier-3 may earn itself.
3. Blast radius if the model is wrong
| Side-effect class | Prefer |
|---|---|
| Read-only / draft-only | Tier-2 LLM step is fine |
| Reversible write (draft email, note) | Tier-2 with human review, or Tier-3 with tight policy |
| Irreversible (charge, delete, public post) | Default Tier-2 + human gate; Tier-3 only with pre-execution deny |
If blast radius is high and certainty is low, you still might not want an agent — you might want a human. Agents are not a courage substitute.
Print this and score the job in a meeting, not in a Slack thread after someone already bought an “agent” template.
| Axis | Score 1 (workflow) | Score 3 (hybrid) | Score 5 (loop) |
|---|---|---|---|
| Next-action certainty | Named before payload | Named after one extract | Unknown until mid-run |
| Branch / long-tail | ≤10, stable | Growing slowly | New branch every release |
| Blast radius | Read / draft | Reversible write | Irreversible |
| Eval readiness | Schema + idempotency | Sampled human audit | Golden-set trajectories |
| Ops readiness | n8n alerts | Escalate package | Budgets, kill switch, policy |
Sum of 5–9: stay Tier-1/2. Sum of 10–15: ship hybrid, measure. Sum of 16–25: you may have earned a bounded loop — still prove it on fixtures first.
Tier-2: workflow + one LLM step (the default)
Pattern that ships:
- Trigger (webhook, form, inbox).
- Normalize + validate input mechanically.
- One model call with a strict schema (JSON Schema / Zod / structured output).
- Mechanical checks on the schema (required fields, enums, ranges).
- Deterministic routing and writes in n8n or code.
- Escalate path when checks fail — no silent “best effort” write.
The model owns a transformation. The graph owns control flow. That is the whole design.
Example jobs that stay Tier-2 for a long time:
- Intent classify → route ticket
- Extract fields from an invoice PDF → Airtable row
- Draft a reply → human send
- Summarize a Zoom transcript → Notion page with fixed template
OpenAI’s Structured Outputs exist for this shape: the model must adhere to your JSON Schema, not merely emit something that looks like JSON. JSON mode guarantees valid JSON. Structured Outputs guarantees the keys, types, and enums you declared. Downstream n8n nodes should still re-check required fields — vendor guarantees are not your only gate.
n8n is a natural host: the graph owns control flow; the model owns one transformation. The Basic LLM Chain node is the Tier-2 primitive — prompt in, optional output parser, no tool picker. Do not let that step call tools “just in case.” That silently becomes Tier-3 without the harness.
| Tier-2 contract | Pass | Fail |
|---|---|---|
| Model call count | 1 per job (or a counted chain) | Hidden retries that call tools |
| Output | Schema-valid object | Prose the next node “figures out” |
| Writes | Workflow node after checks | Model holds the write credential |
| Failure | Escalate with the raw payload | Best-effort CRM patch |
If you cannot tick Pass on all four, you do not have a default. You have a demo.
How n8n draws the same line
n8n’s own docs are unusually honest. A chain “set[s] up a sequence of calls.” An agent is “a chain that knows how to make decisions.” When you execute a workflow with an AI Agent node, “the agent runs multiple times” — setup, tool call, evaluate, respond. That is a loop inside one canvas node.
n8n also documents a hard product difference: chain nodes cannot use memory. If you need a continuing conversation, they tell you to use an agent. That is a product fact, not a recommendation to promote every extract job.
| n8n node | What it is | Use it when |
|---|---|---|
| Basic LLM Chain | One prompt, optional parser, no memory | Classify / extract / draft |
| Q&A or Summarization Chain | Fixed sequence + retriever or summary | RAG question or template summary |
| AI Agent | Tools Agent loop; model picks tools | You already failed the three-tier test |
| HTTP / Code / IF | Deterministic edges | Everything after the model returns |
n8n’s comparison page with LangChain is also clear: n8n wraps LangChain agent abstractions as visual nodes; LangChain in code is what you reach for when the visual loop hides a failure you need to inspect. I have collaborated with the n8n team. The canvas is not the control plane. The control plane is max iterations, tool allowlist, structured output, and an escalate path you actually page.
Their agents vs chains example workflow even routes on the words “agent” and “chain” so you can feel the difference. Use that as a teaching canvas. Do not use it as a production architecture.
- Happy path is a Chain (or HTTP + structured output), not an Agent
- Agent node, if present, sits only on the exception lane
- Max iterations is set and alerted
- Tools on the Agent are read-first; writes stay on workflow nodes
- Failed tool results return structured errors, not empty success
If those boxes are empty, you shipped a loop and called it a workflow.
When you’ve earned a real agent loop
Signals from production, not from a demo:
- Same exception class keeps adding branches to the workflow after three releases
- Humans already do multi-step research across tools with mid-course corrections
- You can stub tools and score trajectories on a golden set
- You have budgets, kill switches, and a policy gate before side effects
- Cost of a failed autonomous run is bounded and recoverable
OpenAI’s guide names the same cluster: nuanced judgment, rulesets too tangled to maintain, heavy unstructured data. Anthropic names the open-ended case: you cannot predict the number of steps, and you cannot hardcode a fixed path. Coding agents are their canonical example — which files to touch is unknown until the model is partway in.
If those boxes stay unchecked, keep Tier-2. Fashion is not an acceptance criterion.
A useful filter: would a competent operator write a 12-box flowchart this week and still be right next quarter? If yes, update the flowchart. If they would throw the flowchart away after three tickets, you are looking at a loop.
What ReAct is — and what it is not
The loop most frameworks wrap is ReAct (Yao et al., 2022): interleaved reasoning traces and actions. Thought, act, observation, repeat. The paper’s point was grounding — actions pull evidence from an environment so the next thought is not pure chain-of-thought drift.
That is a pattern, not a product requirement. You do not “need ReAct” the way you need idempotency keys. You need a loop only when the next action depends on the last observation and you cannot write the sequence in advance.
| Pattern | Control flow | When it is enough |
|---|---|---|
| Single structured call | You | Extract, classify, draft |
| Prompt chain (Anthropic) | You; fixed sequence | Outline → check → write |
| Routing | You; classifier picks a branch | Ticket types, model tiers |
| Evaluator–optimizer | You; fixed critique loop | Drafts with clear criteria |
| ReAct / Tools Agent | Model; observation-driven | Path unknown until mid-run |
Anthropic’s named workflows — prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer — are still workflows. The model fills nodes. You wrote the edges. Do not promote a routing classifier to “our agent” in the board deck. You will inherit agent ops (budgets, traces, policy) for a job that never needed them.
Once you do need a loop, do not leave it as free-form ReAct forever. Cage it: intake → plan → act → evaluate → revise | done | escalate. The cage comes after you prove you need autonomy — it is not a reason to skip the three-tier test.
Hybrid: n8n owns the spine, loop owns the long tail
A clean hybrid:
| Layer | Owns |
|---|---|
| n8n / workflow | Triggers, auth, deterministic writes, SLAs, retries with idempotency |
| Bounded agent | Only the exception lane: “research + propose” or “triage + draft” |
| Policy + evaluator | Allow / deny / escalate before irreversible tools |
Contract between layers:
- Workflow calls the agent with a job package (goal, allowed tools, budget, deadline).
- Agent returns a result package (status, artifacts, reason codes) — never raw chat.
- Workflow decides the write. The agent does not hold production credentials for blast-radius tools unless the pilot explicitly scopes them.
This is how you keep ops familiar (n8n runs, alerts, retries) while still using a loop where uncertainty lives. LangGraph’s pitch is the same idea in code: mix deterministic steps with LLM-driven steps in one graph. n8n can host the deterministic shell and call a small loop over HTTP when the exception lane fires.
| Field | Job package (in) | Result package (out) |
|---|---|---|
| Identity | job_id, trigger source | Same job_id |
| Goal | One sentence + done definition | status: done / escalate / abort |
| Tools | Allowlist + read/write/irreversible tags | tools_used[] with counts |
| Budget | Max turns, max tokens, deadline | turns, tokens, elapsed_ms |
| Output | Expected schema name | Artifacts matching that schema |
| Failure | — | reason_codes[], raw evidence links |
If the loop cannot fill that package, it is not ready to sit behind a production webhook.
Measuring whether hybrid is enough
Do not argue architecture. Instrument a two-week trial of Tier-2 and score:
| Metric | Tier-2 is enough if… |
|---|---|
| Human rewrite rate | Under your job’s tolerance (often <15% for drafts) |
| Silent wrong writes | Near zero on sampled audits |
| New branch requests | Not growing week over week |
| Cost per successful job | Inside the band finance already approved |
| Time-to-escalate | Humans get a package faster than doing the job cold |
If rewrite rate stays high and the failures are “needed another tool / another look,” you have evidence for Tier-3. If failures are schema and template issues, fix Tier-2 — do not promote the model to CEO.
Run the trial like an experiment, not a vibe check.
- Freeze the job definition and the done criteria.
- Ship Tier-2 with schema checks and an escalate path.
- Sample every Nth run plus every escalate (do not sample only the pretty ones).
- Tag each failure:
schema,template,missing_branch,needed_another_look,needed_another_tool. - After two weeks, count tags. Promote only if the last two dominate and the three-tier score still points at a loop.
| Failure tag | What you do |
|---|---|
schema | Tighten JSON Schema / parser; do not add tools |
template | Fix the prompt and the Notion/CRM mapping |
missing_branch | Add an IF node; stay Tier-2 |
needed_another_look | Hybrid candidate — bounded research loop |
needed_another_tool | Hybrid candidate — allowlist that one tool on the exception lane |
I will not invent a “typical” rewrite rate for your shop. Set the tolerance from the human baseline: how often does the current operator redo their own first draft? If you do not know that number, you are not ready to compare architectures.
Failure mode: the faux agent
What breaks: an “agent” that is really while true: call model; call every tool with no state machine, no evaluator, and no policy gate.
What it costs: duplicate emails, duplicate CRM notes, token bills that make the chatbot demo look cheap, and a team that stops trusting automation.
What you do instead:
- Collapse to Tier-2 for the happy path.
- Put the long tail behind an escalate package.
- Only then stand up a bounded loop with max turns, tool allowlist, and offline golden-set gate.
Bravery is not a restore strategy.
The cousin failure is the secret loop: one n8n node that looks like Tier-2, then chains five model calls and two HTTP tools inside a Code node. Count turns. If the model chose the second call, you are in Tier-3 without the harness — no budget, no deny, no result package. Ops will treat it like a workflow until the bill arrives.
| Smell | What it usually is | Fix |
|---|---|---|
| Agent node on the happy path | Faux agent | Replace with Chain + IF |
| “We’ll add tools later” | Threat-model debt | Design the tier with the tools you will enable |
| Max iterations left at default | Unbounded loop | Set it; alert at 70% |
| Model has the Stripe key | Missing policy gate | Workflow holds the write |
| Success metric is “demo worked” | Theater | Rewrite rate + blast radius |
Decision checklist (print this)
- I can state the job in one sentence with a done definition
- I tried Tier-2 with schema-checked I/O for two weeks of real traffic (or a dense fixture pack)
- I know which tools are read vs write vs irreversible
- I have an escalate path that humans will actually use
- If Tier-3: budgets, traces, evaluator, and pre-execution policy exist before soft-launch
- I am not choosing Tier-3 because a competitor’s landing page used the word “agent”
- I can point at Anthropic’s or OpenAI’s criterion this job actually meets
- n8n (or code) owns writes; the loop returns a package
If the last two are unchecked, you are choosing a brand, not a shape.
Acceptance criteria when there is no agent
Tier-2 still needs a definition of done:
- Schema validation pass rate on the LLM step
- Downstream write success with idempotency keys
- Sampled human audit score (or mechanical checks where possible)
- Explicit escalate rate — not “errors hidden in Slack”
No agent does not mean no eval. It means the eval is cheaper and mostly mechanical.
| Check | How you measure | Red line |
|---|---|---|
| Schema pass | Validator on every LLM output | Silent coerce-and-continue |
| Write success | Idempotency key + downstream ACK | Retry storms, duplicate rows |
| Audit | Weekly sample with a rubric | “Looks fine” with no rubric |
| Escalate | Count + time-to-human | Escalates that nobody opens |
A workflow that hides misses in a Slack channel is not simpler than an agent. It is an agent’s failure mode with worse traces.
Mapping common jobs to tiers
| Job | Starting tier | Graduate when |
|---|---|---|
| Lead enrich + CRM field fill | Tier-2 | Enrichment vendors disagree and need multi-hop research |
| Support macro reply | Tier-2 | Refunds / account changes need tool sequencing under policy |
| Ops research brief | Tier-3 candidate | Humans already juggle 4+ sources per brief |
| Invoice → bill pay | Tier-2 + human approve | Never fully autonomous without dual control |
| Content repurpose pipeline | Tier-2 | Brand-risk drafts need iterative critique loops |
| Inbox → calendar hold | Tier-2 | Ambiguous threads need a research pass before booking |
| Vendor security questionnaire | Tier-3 candidate | OpenAI’s “ruleset too tangled” case — if you can still score answers |
| Nightly report from one SQL view | Tier-1 | You do not need a model |
Start left. Move right only with traces that justify it.
Notice the pattern: money movement and public posting stay gated even when the research lane is a loop. Shape and blast radius are independent axes. A Tier-3 research brief that returns markdown is a different animal from a Tier-3 that can refund a card.
Cost sketch without fake precision
You do not need a vendor’s $/task fantasy. Compare architectures on the same job:
- Tokens + tool fees for 100 real cases under Tier-2
- Same 100 under a prototype loop (even if stubbed tools)
- Human minutes saved vs human minutes spent reviewing
If Tier-3 does not beat Tier-2 on successful outcomes per dollar after review cost, the loop is a science project. Pin models and keep the comparison honest when you re-run — floating aliases contaminate the experiment.
Anthropic’s published rule is qualitative: expect higher cost and latency, and only pay it when task performance improves. I will not invent a multiplier for your workload. Measure the 100-case pack. If someone quotes “agents are 10×” without showing the pack, they are selling a slide.
| Line | Tier-2 | Prototype loop |
|---|---|---|
| Model calls / 100 jobs | Fixed (you counted) | Observed (you log) |
| Tool fees | Known HTTP | Known + surprise retries |
| Review minutes | Escalate-only | Every irreversible act |
| Silent wrong writes | Audit sample | Audit sample |
| Outcomes / dollar | Compute last | Compute last |
The last row is the only row that decides. Everything above is input.
How a five-day pilot settles the shape
A Spurlock Studios $1,500 · 5-day agentic pilot is often a shape decision with receipts, not a forced Tier-3 build:
| Day | Output |
|---|---|
| 1 | Job map + three-tier score |
| 2 | Tier-2 spike in n8n (or existing stack) |
| 3 | Failure harvest from fixtures / shadows |
| 4 | Go / no-go for bounded loop; if go, thin harness |
| 5 | Metrics panel + recommendation writeup |
You leave knowing whether to keep shipping hybrid or to fund a real agent build. Scope detail lives in agent pilot scope. The $1,500 credits toward a later build; you keep the spike either way.
That week is not “build us an agent.” It is “prove which column of the comparison table this job lives in.” Plenty of pilots end with a sharper workflow and a documented no. That is a successful week.
Anti-patterns for this decision
“We’ll add tools later.” Tools change the threat model. Design the tier with the tools you will actually enable.
One LLM step that secretly chains five model calls. That is a loop without a harness. Count turns.
Replacing a working workflow because the board wants “AI agents.” Keep the workflow; put agents on the exception lane if anywhere.
Measuring only demo success. Demos are Tier-3 theater. Production is rewrite rate and blast radius.
Calling routing an agent. A classifier that picks queue A or queue B is Anthropic’s routing workflow. It is a good workflow. It is not autonomy.
Giving the loop production write credentials “to move faster.” Speed here is the time until the first irreversible mistake. The workflow holds the key; the loop proposes.
Skipping when not to build an agent. This spoke assumes the job survived that filter. If the path is fully known, or you cannot write pass/fail criteria, stop. Do not three-tier a job that should stay a checklist.
FAQ
When is Tier-2 (workflow + one LLM) the right default?
Whenever the graph of next actions is mostly known and the model’s job is transform, classify, or draft inside a schema. That covers a large share of SMB automation: tickets, extracts, summaries, and draft replies. Anthropic and OpenAI both tell you to start there — simplest solution, deterministic when it suffices. Escalate the exceptions; do not promote every exception into an open tool loop on day one.
What signals mean you’ve earned a real agent loop?
Repeated long-tail branches that make the workflow unmaintainable, multi-step tool work humans already do with mid-run decisions, and the control plane pieces (eval, budget, policy) ready before autonomy. OpenAI’s version is complex judgment, tangled rules, or unstructured data that already beat a rules engine. Demo applause is not a signal. Failure traces are.
How does this differ from “when not to build an agent”?
That spoke owns refusal — jobs that should stay human or stay deterministic. This spoke owns the middle: when hybrid is enough, and how to graduate. Read both. Many teams need the refusal post first, then this decision tree for the remainder.
Can n8n host the workflow while a bounded loop handles the long tail?
Yes — and that is often the production shape. n8n owns triggers, credentials for deterministic writes, and SLAs; the loop returns a result package for the exception lane. n8n’s own docs treat the Agent node as a multi-run decision loop and the Chain nodes as a predetermined sequence. Keep irreversible tools behind policy gates either way.
What acceptance criteria still apply if there’s no agent?
Schema pass rate, write success with idempotency, sampled audit quality, and a visible escalate rate. “No agent” is not “no measurement.” It is a cheaper measurement surface, and it is the surface you need before you have any business funding a loop.
What does a Spurlock pilot prove in five days on this decision?
Which tier fits the job, with a Tier-2 spike, failure harvest, and a go / no-go for a bounded loop — plus the minimum metrics so the recommendation is not a vibe. Details: /agentic and pilot scope.
CTA
Pick the shape before you pick the framework.
What questions does this article answer?
- When is Tier-2 (workflow + one LLM) the right default?
- Whenever the graph of next actions is mostly known and the model’s job is transform, classify, or draft inside a schema. That covers a large share of SMB automation: tickets, extracts, summaries, and draft replies. Anthropic and OpenAI both tell you to start there — simplest solution, deterministic when it suffices. Escalate the exceptions; do not promote every exception into an open tool loop on day one.
- What signals mean you’ve earned a real agent loop?
- Repeated long-tail branches that make the workflow unmaintainable, multi-step tool work humans already do with mid-run decisions, and the control plane pieces (eval, budget, policy) ready before autonomy. OpenAI’s version is complex judgment, tangled rules, or unstructured data that already beat a rules engine. Demo applause is not a signal. Failure traces are.
- How does this differ from “when not to build an agent”?
- That spoke owns refusal — jobs that should stay human or stay deterministic. This spoke owns the middle: when hybrid is enough, and how to graduate. Read both. Many teams need the refusal post first, then this decision tree for the remainder.
- Can n8n host the workflow while a bounded loop handles the long tail?
- Yes — and that is often the production shape. n8n owns triggers, credentials for deterministic writes, and SLAs; the loop returns a result package for the exception lane. n8n’s own docs treat the Agent node as a multi-run decision loop and the Chain nodes as a predetermined sequence. Keep irreversible tools behind policy gates either way.
- What acceptance criteria still apply if there’s no agent?
- Schema pass rate, write success with idempotency, sampled audit quality, and a visible escalate rate. “No agent” is not “no measurement.” It is a cheaper measurement surface, and it is the surface you need before you have any business funding a loop.
- What does a Spurlock pilot prove in five days on this decision?
- Which tier fits the job, with a Tier-2 spike, failure harvest, and a go / no-go for a bounded loop — plus the minimum metrics so the recommendation is not a vibe. Details: [/agentic](/agentic) and [pilot scope](/blog/agent-pilot-scope).
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.