When Not to Build an Agent (And What to Build Instead)
Skip the agent when the path is known or criteria are mush. A workflow plus one schema-checked LLM step is the default. Build the loop only after that fails.
William Spurlock Founder — Spurlock Studios Updated 28 MIN
Do not build an agent when a workflow — or a workflow with one LLM step — will finish the job. Agents are for uncertain paths with criteria you can write down. If the path is known, ship automation. If “good” is taste with no test, fix the process. Non-determinism is a cost. Pay it only when it buys a branch you cannot draw in advance.
This spoke is the brake pedal for the Agentic Systems Operating Manual. The cousin that owns the architecture test — certainty, branching, blast radius — is agent loop vs LLM workflow step. Here I own the no: when the cheaper shape wins, what to build instead, and how Spurlock Studios talks a team out of a loop during a paid pilot conversation.
The short answer
- Known path. Same steps, same systems, rare exceptions. That is a workflow. No agent.
- One messy field. Invoice PDF, email body, ticket text. One schema-checked LLM step, then deterministic routing. Still no agent.
- Mushy criteria. You cannot write pass/fail. Do not launder disagreement through a planner. Workshop the process.
- High blast radius, weak recovery. Irreversible legal, financial, or public actions without a deny gate and a human. Stay on a workflow with an approval node.
- Prove the cheap shape first. Graduate to a loop only when traces show the graph exploding and you already have evaluators, a sandbox, a stop condition, and a kill switch.
When does a workflow or LLM step beat an agent?
Most “we need an agent” tickets are a classifier, an extractor, or a draft node wearing a costume. The model should transform one payload. The graph should own what happens next.
| Job you were sold | Shape that actually ships | Why the loop loses |
|---|---|---|
| “Agentize invoice intake” | PDF → structured extract → schema check → QuickBooks / Airtable | Path is known; exceptions are validation failures |
| “Agent for weekly report” | Scheduled SQL → template → optional narrative paragraph after numbers exist | Figures must come from code, not a planner |
| “Agent to classify tickets” | Text classifier → Switch → queue | Next queue is a named edge |
| “Agent to draft replies” | Retrieve macros → one draft call → human send | Write is irreversible; human is the gate |
| “Agent to sync CRM fields” | Webhook → map → idempotent write → DLQ | Choice is a lookup table, not a tool picker |
A workflow is a graph you drew before the first production ticket. An LLM step is a node on that graph: input in, schema out, next edge known. An agent is a control plane: the model may propose tools, the harness may allow or deny them, and “done” is a status you evaluate — not the last webhook hop.
If you can name the next system call without looking at the payload, you are still in workflow land. I have built 500+ automations and spent 20,000+ hours on agentic systems. The jobs that paid for a loop were the ones where a human already changed course mid-task. The jobs that did not were a classifier plus three IF nodes, finished on Tuesday.
- I can whiteboard every edge without saying “it depends” more than twice → workflow
- I need the model only to turn messy text into a schema or a draft → workflow + one LLM step
- The next tool is unknown until mid-run and I can score the result → agent candidate
- I want the model to “just handle it” → stop. That is theater.
Candidate is not permission. Candidate means instrument the cheap shape and see whether the long tail stays ugly.
Install the LLM step like a contract, not like a brainstorm:
- Trigger (webhook, form, inbox, schedule).
- Normalize and validate input in code — required fields, types, size caps.
- One model call with a JSON Schema / structured-output parser.
- Mechanical re-check: required keys, enums, ranges. Fail closed.
- Deterministic Switch / IF routing and writes in n8n or code.
- Escalate path that carries the raw payload. No silent “best effort” write.
If step 3 starts calling tools, you left this shape. Put the tools back on the graph or admit you are now in loop land and go read the operating manual.
When should you not use AI agents?
Skip the loop — for now, and often for good — when any of these are true.
| Hard no | What you feel | What to build |
|---|---|---|
| The path is fully known | Same five steps every time | Deterministic automation |
| You cannot write pass/fail | Reviewers argue about “good” | Process workshop, then a rubric |
| Stakes are high and recovery is hard | One bad write is a lawsuit or a chargeback | Workflow + human approval; no write tools on a model |
| Data and tool access are political | Month one is credentials, not learning | Unblock access; do not start a pilot |
| Volume is tiny | Ten items a month | Human + template |
| Nobody owns the SOP | The bot will become the scapegoat | Named owner first |
| You want magic, not measurement | Deck says “autonomous” | Decline. Theater budgets are not a production plan |
In those cases, build something else. The rest of this post is the menu.
Google Cloud’s architecture guide (last reviewed May 2026) is blunt about the cheap cases: if the workload is “predictable or highly structured, or if it can be executed with a single call to an AI model,” stay non-agentic — summarizing a document, translating text, classifying feedback. See Choose a design pattern for your agentic AI system. Microsoft’s decision framework opens the same way: if the work is “deterministic, repeatable, or can be expressed as a clear function/workflow, don’t use an agent.” That line lives on their feature comparison page, not in a blog aside.
What do the vendors actually say?
You do not need my opinion to decline a loop. The people selling the models already wrote the no.
| Source | When they say stay simple | The line |
|---|---|---|
| Anthropic, Building effective agents | Start with one augmented LLM call; add multi-step systems only when simpler solutions fall short | “This might mean not building agentic systems at all.” Agentic systems “trade latency and cost for better task performance” and risk “compounding errors.” |
| OpenAI, A practical guide to building agents | Validate the use case first | “Otherwise, a deterministic solution may suffice.” Agents belong where rule-based automation already failed. |
| LangGraph, Workflows and agents | Predetermined code paths | Workflows “operate in a certain order.” Agents “define their own processes and tool usage.” Mix both; do not default the whole job to a loop. |
| n8n, What agents do | Chains follow a predetermined sequence | An agent is “a chain that knows how to make decisions.” The Agent node “runs multiple times” per execution. |
| AWS, Orchestration models | Deterministic business process | Step Functions for controlled processes. Bedrock Agents when the path is dynamic or conversational. Deterministic workflow: agents “not needed.” |
| Google Cloud, agent design patterns | Predictable, sequential, single-call jobs | Non-agentic first. Sequential “workflow agents” run predefined logic without asking a model which step is next. |
Anthropic is the citation I hand executives. They sell the models and still tell you the simplest solution might be no agent. If your vendor deck disagrees with that paragraph, keep your vocabulary. The product name on the node does not change the physics.
Print the table. In the meeting, ask which row the job actually matches — not which logo is on the slide. If three vendors say “stay deterministic” and one SE says “agent,” believe the docs. Google even names a sequential workflow agent that “operates on predefined logic without having to consult an AI model for the orchestration of its subagents.” That is a pipeline. Call it a pipeline.
| Meeting move | Do this | Do not do this |
|---|---|---|
| Open | Read Anthropic’s “might mean not building agentic systems at all” out loud | Start from the vendor demo |
| Classify | Put the job on one row of the table above | Let “AI” mean “agent” by default |
| Price | Ask how many model calls a successful run is allowed | Accept “it depends, the model decides” |
| Close | Assign a cheap-shape owner and a date | Schedule a six-week agent POC to avoid the no |
How do you choose agent vs automation workflow?
Draw the signal table in the intake call. Do not draw it after someone already bought seats.
| Signal | Prefer automation / LLM step | Prefer agent |
|---|---|---|
| Path | Fixed flowchart | Varies with input in ways you cannot economically encode |
| Judgement | Rare, or one bounded transform | Frequent, and you can still write criteria |
| Tools | Few, predictable, called by the graph | Many, chosen mid-run |
| Failure mode | Retry + dead-letter queue | Evaluate + revise + escalate |
| Ops skill | Workflow engineering | Evaluation, sandboxing, budgets, kill switch |
| Cost shape | Known nodes per run | Unknown turns until a stop condition fires |
| Example | Invoice PDF → line items → ledger with schema checks | Research brief from messy sources with citation rules |
Automation — often on n8n — wins when determinism is available. A single LLM step wins when one field is messy and everything after it is a Switch. Agents win when choice is required and you can still evaluate outcomes. Hybrid wins most weeks: automation rail, model only in the judgement step.
n8n’s own docs make the product difference boring and useful. A Basic LLM Chain is prompt in, optional output parser, no tool picker. An Agent node is a loop inside one canvas box. If you attach tools “just in case,” you silently bought the expensive shape without the harness. I have collaborated with the n8n team. The canvas is not the control plane. The control plane is max iterations, an allowlist, structured output, and an escalate path you actually page.
AWS says the same thing in cloud vocabulary. Step Functions when you need a full state trace and a machine-readable path. An agent when the user arrives in natural language and the next action cannot be named yet. Do not replace a deterministic process with a planner because the slide said “AI-native.”
What should you build instead?
Pick the cheapest object that clears the job. In order:
1. A written SOP plus a human checklist
If the work is rare or politically sensitive, clarity beats silicon. Agents amplify existing process. They do not invent accountability.
- Job fits on one line
- Steps a new hire could follow
- Named reviewer
- Definition of done a second person can score
2. Deterministic automation
Webhooks, queues, schema validation, idempotency keys, dead-letter queues, human approval nodes. Boring, shippable, auditable. Spurlock Studios’ automation lane exists for this reason. Point at /automation when the tree says so.
3. Search, a dashboard, or a better form
Sometimes “we need an agent” means “we cannot find the number” or “intake is a paragraph in Slack.” Fix retrieval UX, reporting, or the form before you add a planner. Structured intake deletes more agent tickets than any framework.
| Intake smell | What people ask for | What actually fixes it |
|---|---|---|
| “The agent will figure out the request” | Planner + ten tools | Required fields, enums, file upload |
| “We cannot find last month’s number” | Research agent | A dashboard with a saved query |
| “Customers write novels in the form” | Classifier agent | Split the form; make the novel optional |
| “Nobody knows which queue” | Router agent | One dropdown the submitter already knows |
4. Rules engine plus one LLM step
Use code for decisions that are actually rules. Use a model only to turn messy text into a schema, then continue deterministically. OpenAI’s Structured Outputs exist for this shape: the model must match your JSON Schema, not emit something that looks like JSON. Downstream n8n nodes still re-check required fields. Vendor guarantees are not your only gate.
5. Draft, then a human
The model writes. A person sends, posts, refunds, or files. That is not cowardice. That is a write-tool you refused to hand a loop.
6. A narrow pilot — later
If you are close but not ready, schedule the readiness work (criteria, data access, sandbox) and then book the $1,500 · 5-day pilot. Do not force the week early to satisfy a board slide.
| Instead-of | Ships when | Still not an agent |
|---|---|---|
| SOP + checklist | Volume is tiny or political | Correct |
| n8n / code workflow | Path is drawable | Correct |
| Workflow + one LLM step | One messy transform | Correct |
| Draft + human | Write is irreversible | Correct |
| Process workshop | Criteria are mush | Correct |
| Agent pilot | Path branches, criteria exist, sandbox is real | Only then |
How do you run the decision tree in an intake call?
Keep this tree honest even when the vendor demo was pretty.
Can you draw the full flowchart without hand-waving?
yes → automation, or automation + one LLM step
no → Can you write pass/fail criteria for outputs?
no → process workshop / product definition
yes → Are tools and data accessible in a sandbox this month?
no → unblock access first
yes → Is volume and value worth the non-determinism tax?
no → human + templates
yes → agent candidate → read the operating manual
Facilitator asks, in this order:
- Can anyone whiteboard the steps without “it depends” more than twice?
- What would a unit test assert on the output?
- Which actions are irreversible?
- How many items per week?
- Who owns criteria after the build?
Scoring: mostly clear steps + rare judgement → automation. One messy field + named next edge → LLM step. Frequent “it depends” + crisp tests → agent candidate. Frequent “it depends” + no tests → process work. Do not skip step five. An orphan agent is a scapegoat with an API key.
What are false reasons to build an agent?
These show up in real intake calls. None of them are a use case.
| False reason | What it actually is | Honest move |
|---|---|---|
| “Competitors say they have agents.” | Marketing | Ask what job, what evaluator, what kill switch |
| “Our CEO tried ChatGPT and liked it.” | A chat, not a control plane | Offer a workflow + one draft step |
| “The automation tool has an AI node.” | A product SKU | Use the Basic LLM Chain, not the Agent node, until the tree says otherwise |
| “We already bought seats.” | Sunk cost | Seats do not create criteria |
| “It will replace the team by Q3.” | A headcount fantasy | It will not if criteria and data are weak |
| “The vendor’s deck says agent.” | Agent-washing | If the demo is a Zap with a model in the middle, keep calling it a workflow |
True reasons, and there are not many:
- Paths branch in ways flowcharting cannot economically capture.
- Tool choice depends on content you have not seen yet.
- Evaluation can be made crisp.
- Unit economics beat human handling at expected volume.
- You can name the sandbox, the stop condition, and the person who pulls the kill switch.
If you only have the left column, you do not have an agent problem. You have a procurement problem.
What does the non-determinism tax actually buy?
You pay three bills a workflow does not: extra model turns, extra latency, and a chance that an early mistake becomes the next prompt. Anthropic names that trade in the open — latency and cost for task performance, plus compounding errors. Buy it when a human already changes course mid-task and you can still score the landing. Do not buy it to classify a ticket.
| You pay | Workflow / LLM step | Agent loop |
|---|---|---|
| Model calls per success | One, or a counted chain | Unknown until stop |
| Latency | Bounded by named nodes | Bounded only if you install a budget |
| Failure | Schema miss, retry, DLQ | Wrong tool, loop, fluent wrong write |
| Debug | Open the node, read the payload | Reconstruct a trajectory |
| Audit | The graph is the audit | You must log every proposed tool |
Volume is a veto, not a vibe. Ten items a month wants a human and a template. A hundred items with a drawable path wants n8n. A thousand items with a drawable path still wants n8n — plus capacity, not a planner. Agents enter when volume is high and the path will not sit still and a wrong write is cheaper than a human on every item.
- I can price a successful run before it starts → stay on the graph
- I cannot name the call count and I have no max-iterations cap → do not ship
- The tax is less than a human handle at expected volume → candidate, still prove it
- The tax is a slide metric with no volume → kill the ticket
What breaks when you ignore the no?
Wrong shape taxes you twice. A loop on a solvable graph burns tokens and eng time. A workflow that pretends the long tail does not exist silently writes the wrong field.
Failure mode I keep seeing: an “agent” wrapped around a fixed five-step Zap. Invoice lands, model “plans,” calls the same three tools in the same order, writes the row. You paid a planning tax for a flowchart. When the model invents a fourth tool call, you get a duplicate charge or a skipped validation. The honest version is a chain plus schema checks.
The inverse lie is worse: automation with hidden model calls and no evaluator. You built an agent and lied on the diagram. Ops pages a “workflow” that looped eight times and posted to a customer.
| Wrong choice | Week-two symptom | Repair |
|---|---|---|
| Agent for a form extract | Token bill vs one structured-output call; flaky tool retries | Delete the Agent node; keep one extract |
| Workflow for a messy exception pile | Humans rewrite “automation” output every morning | Bound a loop on the exception lane only — or keep a human |
| Agent with irreversible tools and no deny gate | Duplicate refunds, public posts, deleted rows | Strip write tools; approval node; read OWASP LLM06 |
| Endless POCs to avoid deciding | Six demos, zero owner | Decide: automate, agentize, or change the process |
OWASP’s 2025 Top 10 still lists LLM06 Excessive Agency as its own vulnerability: too much functionality, too many permissions, too much autonomy. The example is an assistant that needed to read mail and was handed a plugin that could send. That is not an edge case. That is the default if you attach every tool “for later.” Least privilege is the substitute for bravery.
NIST’s AI 600-1 Generative AI Profile (July 2024) names the human failure that follows: automation bias — excessive deference to a confident system — plus confabulation risk when people believe fluent false content. An agent that sounds sure is more dangerous than a workflow that fails a schema check in the open. If you cannot deactivate the system, you already failed GOVERN 1.7 in that profile.
Cost is not only tokens. Wrong-shape agents also burn a sprint chasing “the agent got stuck” when the graph had six known branches the whole time.
Which case patterns should stay off the agent backlog?
These four keep arriving. Treat them as a filter, not a vibe.
“Agentize the weekly report.” The report is fixed SQL plus a template. Build a scheduled automation. Add a model only to draft a narrative paragraph after numbers are computed in code, with a check that every figure appears in the numeric source. If a number is missing, fail the run. Do not let a planner invent a metric.
“Agent to replace Tier-1 support.” Usually means missing macros, a thin help center, and no SLA. Fix knowledge and macros. Consider a draft-reply step with citation rules later — still gated. A FAQ bot with retrieval is RAG plus rules, not a full loop. Google’s guide puts classify-and-summarize in the non-agentic bucket for a reason.
“Agent to manage the calendar.” Permissions and irreversible invites make this a sandbox nightmare. Start with draft suggestions a human confirms. If the model can send the invite, you have already granted excessive autonomy.
“We automated 80% already; the last 20% is messy.” That last 20% is often the correct agent boundary — if criteria exist. Do not rewrite the working 80% into a nondeterministic loop. Keep the rail. Bound the exception lane. Return a result package (done / escalate / abort) to the workflow. The architecture for that hybrid lives in agent loop vs LLM workflow step.
| Pattern | Default | Promote only if |
|---|---|---|
| Weekly report | Schedule + template + optional narrative | Narrative must cite live numbers and still fails checks |
| Tier-1 support | Macros + help center + draft | Citation rules exist and a human still sends |
| Calendar | Draft holds | You have a deny list and a confirm step that cannot be skipped |
| Last 20% messy | Keep the 80% rail | Criteria, sandbox, and a stop condition exist |
Which hybrid designs are usually enough?
These three deliver value without pretending the whole flowchart is nondeterministic.
| Hybrid | Model owns | Graph owns | Use when |
|---|---|---|---|
| Extract then automate | Messy email / PDF → schema | Routing, writes, retries, DLQ | Intake is unstructured; the rest is rules |
| Draft then human | Reply / post / summary | Send, publish, refund | Write is irreversible |
| Retrieve then template | Nothing, or a one-line rewrite | Librarian + template engine | You needed a clause, not a planner |
Use a full agent loop only for the segment that still branches after structure exists. LangGraph’s docs allow mixing a predetermined graph with an agentic node. That is a feature. It is not a reason to promote the whole job.
AWS’s prescriptive guidance draws the same split: Step Functions when you need an audit trail; an agent when the user goal is inferred. If you need both, the durable outer loop stays deterministic and the model sits inside one state. Do not invert that. An LLM inferring “I think we finished step 3” is not a state machine.
- Trigger, normalize, and validate are mechanical
- At most one model call on the happy path, or a counted chain you can price
- Writes happen in workflow nodes after checks
- Escalate carries the raw payload, not a “best effort” patch
- The model does not hold the CRM key
If you cannot tick those, you do not have a hybrid. You have a demo.
How do high stakes and culture veto an agent?
Technology is the easy half. Culture is the veto.
If the org punishes escalation and rewards “the bot handled it,” agents will hide failure. Healthy cultures treat escalate as a designed outcome. If leadership forbids that, do not build an agent until they agree a failed criterion is a success when the human is paged.
NIST’s Human-AI Configuration risk is the academic name for this: over-reliance, anthropomorphism, automation bias. People trust fluent systems. Your job is to make the cheap, visible failure (schema miss, approval hold) more attractive than the expensive, hidden one (wrong refund, public post).
OWASP’s mitigation list is the engineering name: minimize tools, minimize what each tool can do, require a human on high-impact actions. If your “agent” can delete, charge, or publish without a confirm, you are not early. You are exposed.
| Stake | Default | Agent only if |
|---|---|---|
| Read / draft | LLM step is fine | You somehow need tool choice to draft — rare |
| Reversible write (internal note) | Workflow + optional model | Criteria and an audit log exist |
| Irreversible (charge, delete, public) | Workflow + human gate | Pre-execution deny, sandbox, named owner, kill switch |
| Legal / medical / identity | Human. Full stop. | You have a control plane the operating manual would accept |
Translate the tree into money for executives: cost of a wrong irreversible action × frequency, versus cost of human handling × frequency, versus automation engineering cost. Agents enter when branching complexity makes encoding the graph more expensive than an agent-with-controls and criteria are available. If the executive wants speed above all, still put hard nos in writing. Speed without hard nos is how brands get surprised in public.
When not to use AI agents is often a culture answer wearing a technology costume. Spurlock Studios will say so.
How should you plot the portfolio so you stop funding the wrong POCs?
Plot each idea on two axes: path uncertainty and criteria clarity.
| Quadrant | Uncertainty | Criteria | Fund this |
|---|---|---|---|
| Automate | Low | High or low | Workflow. Optional LLM step. |
| Workshop | High | Low | Process work. No silicon yet. |
| Agent candidate | High | High | Pilot — after sandbox and owner |
| Vanity | Low | Low | Neither. Kill the ticket. |
This view stops the organization from funding a dozen agent POCs that belong in different quadrants. Vendors will redefine automation as agents for marketing. Keep your words. When not to use AI agents includes “when the vendor’s deck says agent but the demo is a Zap.”
If you cannot name the evaluator, sandbox, stop condition, and kill switch, you are not building an agent yet — regardless of model brand. Read the operating manual. If the filter passes, the measured next step is a pilot, not a rewrite of the working rail.
How does Spurlock Studios handle “maybe agent”?
On /agentic we still start many relationships with a pilot — but only after the job sentence and criteria clear the bar. If automation is the fit, we say so and point at that lane. Integrity is part of the product. We would rather route you to a workflow or a process workshop than sell an agent week into a bad fit.
I will talk you out of an agent when that is the honest move, including during a paid conversation. That is not a humility bit. It is how you avoid spending 20,000+ hours of agentic work on a job a webhook would have finished.
When the tree says agent, the $1,500 · 5-day pilot is the measured next step: one job, one evaluator, one sandbox, a kill switch. Doctrine stays in the operating manual. Shape details stay in agent loop vs LLM workflow step.
Contact when the tree says go: /contact?intent=agentic-pilot.
What must be true before you open the pilot form?
- Job sentence fits one line
- Pass/fail criteria drafted (a second person can score a sample)
- Hard nos listed (what the system must never do)
- Real sample data available this month
- Automation or one LLM step is clearly insufficient — you tried, or you can show why the graph explodes
- Sponsor and domain reviewer named
- Write tools are off the model until a deny gate exists
- You can name the evaluator, sandbox, stop condition, and kill switch
If three boxes are empty, do the readiness work first. The operating manual and /agentic will still be there. The courageous product decision is often declining the loop. When not to use AI agents includes most weeks when criteria are mush and paths are fixed. Agent vs automation workflow clarity saves quarters. If your tree still points to an agent after a hard look, we will meet you on /agentic and prove one job — or tell you to automate instead.
Readiness work is not a stall. It is the week you write criteria, get the credentials, and collect twenty real samples. That week is cheaper than a pilot you will have to redo. If access is still political after that week, you do not have an agent project. You have an IT ticket.
Bring the twenty samples to the first call. If you cannot, we will spend the week on intake, not on a loop. That is still a good week. It is not an agent week.
If your flowchart already fits on one slide without hand-waving, celebrate and automate. Agents are for the messy remainder after structure exists — not a costume for the easy eighty percent.
FAQ
When should you not use AI agents?
When the path is known, criteria are mush, stakes and recovery are worse than the upside, access is blocked, volume is tiny, or nobody owns the process. Build an SOP, a workflow, or a single schema-checked LLM step instead. Anthropic, OpenAI, Microsoft, AWS, and Google all publish some version of that no. If you cannot write pass/fail, do not launder the argument through a planner.
How do you choose agent vs automation workflow?
Prefer automation for fixed paths with schema checks and dead-letter queues. Prefer a workflow plus one LLM step when only one field is messy and the next edge is named. Prefer an agent when the path branches, tool choice depends on content you have not seen, and you can still write evaluator criteria. Hybrid designs — model extracts or drafts, graph writes — are common and healthy.
Is a chatbot an agent?
Not by itself. A chatbot becomes agentic when it chooses tools under constraints with evaluation and terminal states. A FAQ bot with retrieval may only need RAG plus rules — not a full agent loop. If every turn is “retrieve, then template,” you built search, not an agent. Keep the word for systems that pick the next action at runtime.
Can we start with automation and add an agent later?
Yes — and often should. Put deterministic glue in place first. Insert an evaluated agent step only where judgement concentrates. That sequencing reduces blast radius and keeps the working eighty percent out of a nondeterministic loop. The graduation test lives in agent loop vs LLM workflow step.
Will Spurlock Studios decline an agent pilot?
Yes, if scope is theater or automation is clearly enough. We would rather keep trust than sell a five-day week into a bad fit. When it is a fit, the pilot is $1,500 for five days on /agentic: one job, criteria, sandbox, kill switch.
Where should I read next if we are building an agent?
Start with the Agentic Systems Operating Manual for evaluators, sandboxes, state, cost caps, and observability. Then read agent loop vs LLM workflow step for the three-tier shape test. Choosing “agent” commits you to that operating manual — not to a single prompt file.
CTA
If the tree still says agent after a hard look, book the measured week — not a rewrite of a flowchart that already fits on one slide.
What questions does this article answer?
- When should you not use AI agents?
- When the path is known, criteria are mush, stakes and recovery are worse than the upside, access is blocked, volume is tiny, or nobody owns the process. Build an SOP, a workflow, or a single schema-checked LLM step instead. Anthropic, OpenAI, Microsoft, AWS, and Google all publish some version of that no. If you cannot write pass/fail, do not launder the argument through a planner.
- How do you choose agent vs automation workflow?
- Prefer automation for fixed paths with schema checks and dead-letter queues. Prefer a workflow plus one LLM step when only one field is messy and the next edge is named. Prefer an agent when the path branches, tool choice depends on content you have not seen, and you can still write evaluator criteria. Hybrid designs — model extracts or drafts, graph writes — are common and healthy.
- Is a chatbot an agent?
- Not by itself. A chatbot becomes agentic when it chooses tools under constraints with evaluation and terminal states. A FAQ bot with retrieval may only need RAG plus rules — not a full agent loop. If every turn is “retrieve, then template,” you built search, not an agent. Keep the word for systems that pick the next action at runtime.
- Can we start with automation and add an agent later?
- Yes — and often should. Put deterministic glue in place first. Insert an evaluated agent step only where judgement concentrates. That sequencing reduces blast radius and keeps the working eighty percent out of a nondeterministic loop. The graduation test lives in [agent loop vs LLM workflow step](/blog/agent-loop-vs-llm-workflow-step).
- Will Spurlock Studios decline an agent pilot?
- Yes, if scope is theater or automation is clearly enough. We would rather keep trust than sell a five-day week into a bad fit. When it *is* a fit, the pilot is $1,500 for five days on [/agentic](/agentic): one job, criteria, sandbox, kill switch.
- Where should I read next if we are building an agent?
- Start with the [Agentic Systems Operating Manual](/blog/agentic-systems-operating-manual) for evaluators, sandboxes, state, cost caps, and observability. Then read [agent loop vs LLM workflow step](/blog/agent-loop-vs-llm-workflow-step) for the three-tier shape test. Choosing “agent” commits you to that operating manual — not to a single prompt file.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.