Single Agent First: Split Only When Trust, Audience, or Timing Conflicts
Start with one agent and many tools. Split only when trust, audience, or timing conflict—and prove that split with pass rate, cost, and escalate rate.
William Spurlock Founder — Spurlock Studios Updated 26 MIN
Should you start with a single agent or go multi-agent? Start single. One agent with a clear job, a bounded tool set, an evaluator, and a kill switch beats a committee of prompts that hand work to each other for theater. Split only when trust, audience, or timing actually conflict — then prove the split helped with pass rate, cost per pass, and escalate rate on the same golden set.
This spoke belongs to the Agentic Systems Operating Manual. Once you do split, package the boundary with multi-agent handoffs. This post owns the decision to split at all.
That is not a taste take. Anthropic tells teams to find the simplest solution and add complexity only when it demonstrably improves outcomes. OpenAI says define the smallest agent that can own a clear task, then add agents only for separate ownership, tools, or approval policies. Microsoft calls a single agent with tools the usual enterprise default. LangChain says a single agent with the right tools and prompt often matches a crew. I have spent 20,000+ hours on agentic systems and shipped 500+ automations. The expensive failures were almost always a second write path I did not need.
The short answer
- Default topology: one agent, many tools, one evaluator, one write authority.
- Still single-agent: planner / tool-user / reviser states inside one loop — not three products.
- Split when tools need different trust levels, outputs serve conflicting audiences, or timing and durability requirements diverge.
- Premature multi-agent creates coordination bugs: lost context, ABAB loops, shared-data races, and debug across process walls.
- Proof of a good split: better pass rate or lower escalate rate or lower cost per pass on the same golden set — not a prettier diagram.
What still counts as a single agent with many tools?
A single agent is one control loop with one terminal authority. Tools are capabilities. States are phases. A human pause is a gate. None of those are teammates.
| Pattern | Why it is still one agent |
|---|---|
| Many tools (CRM, search, calendar) | Tools are capabilities, not teammates |
| Plan → act → evaluate → revise states | States are phases of one job |
| Router that picks a tool subset | Filtering tools ≠ spawning agents |
| Human approval pause mid-run | HITL is a gate, not a second agent |
Specialist prompts swapped by job_type | Config, not a multi-agent system |
| Skills loaded on demand | LangChain keeps one agent in control while it loads context |
You have multi-agent when another autonomous loop can take actions with its own tool rights, memory, or stop conditions — especially across process or queue boundaries.
If your “researcher agent” cannot terminate the job and cannot write, it may be a function with a costume. Costumes are fine; do not bill them as architecture.
- One
run_id/trace_idowns the job from trigger to terminal write - One allowlist decides which tools may fire this turn
- One evaluator scores the composed result, not each costume
- One kill switch stops every side effect
Microsoft’s complexity ladder starts at a direct model call, then a single agent with tools, and only then multi-agent. Skip a rung only when a measured conflict forces it.
Why does fashion push multi-agent?
Demos love casts of characters. Role names make slides readable. Frameworks make crews easy to spin up. None of that proves you needed more than one write path.
Multi-agent hype usually optimizes for:
- Narrative clarity in a demo
- Parallelism you have not measured
- Mimicking an org chart
Production optimizes for:
- Correct side effects
- Debuggable failures
- Cost per passing job
When those conflict, ship the boring single loop.
Cognition put the failure in two sentences: share context, and remember that actions carry implicit decisions. Two writers who never saw each other’s choices produce a merge conflict with extra tokens. Their later note is narrower, not a reversal: extra agents can add intelligence, but writes stay single-threaded. That is the same rule I use on client work.
| Signal you are buying fashion | What to do instead |
|---|---|
| The deck has a “researcher / writer / critic” triangle | One loop with an evaluator and a deny gate |
| The framework scaffolded three agents in ten minutes | Keep the scaffold; delete two writers |
| Someone said “we need parallelism” with no p95 | Measure the single-agent bottleneck first |
| Role names match the org chart | Map trust tiers, not job titles |
Fashion is cheap in a slide. Coordination is expensive in a queue.
Why does premature multi-agent create coordination bugs?
These are system bugs mislabeled as “the model is dumb.” Fix topology before you buy a bigger model.
| Bug | How it shows up | What actually broke |
|---|---|---|
| Lost intent | Agent B never sees the constraint Agent A “agreed” in prose | Context was not packaged |
| Dual write authority | Two agents update the same CRM field with different drafts | Two wallets, one row |
| ABAB oscillation | A hands to B; B rejects; A “fixes”; infinite courtesy | No terminal owner |
| Shared-data races | Both read stale state; both write; last write wins silently | No compare-and-set |
| Trace fracture | New run_id per agent; nobody can reconstruct the story | No propagated traceparent |
| Eval gaps | Each agent “looks fine”; the composed job fails | Evaluator scored costumes, not the job |
Cognition’s Flappy Bird example is the same bug in costume: one subagent builds Mario pipes, the other builds a bird that is not a game asset, and the merger inherits two implicit decisions that never met. Copying the original task into both prompts does not fix it. The conversation already made decisions the subagents never saw.
Anthropic’s research writeup is the other side of the same coin. Their lead-plus-subagent setup beat a single-agent baseline by 90.2% on an internal research eval — breadth-first, parallel search, high value. They also said agents use about 4× the tokens of chat, multi-agent about 15×, and that coding work with shared context and tight dependencies is a poor fit. Token spend explained 80% of BrowseComp variance. That is a research receipt, not a license to spawn a refund committee.
Microsoft’s Cloud Adoption Framework is blunt about the tax: every extra agent adds protocol design, error handling, state sync, prompt work, monitoring, credentials, and handoff latency. Pay that tax for a conflict. Do not pay it for a slide.
When do conflicting audiences force a split?
Split when one loop cannot honestly serve two masters.
| Conflict | Example | Split shape |
|---|---|---|
| Audience | Internal ops notes vs customer-facing email | Drafter (internal tools) → Sender (email-only tools) |
| Trust | Read-only research vs irreversible refunds | Researcher (no wallet) → Actor (refund tool + HITL) |
| Timing | Fast FAQ answers vs overnight batch enrichment | Online agent vs batch worker with different SLOs |
| Compliance | PII-heavy retrieval vs public content generation | Librarian in a restricted VPC → Writer with redacted packs |
If you can solve the conflict with tool allowlists and policy gates inside one agent, prefer that. A split is for when allowlists still leave a trust or SLO collision.
OpenAI’s split list matches this, not a casting call: different tool or MCP surface, different approval policy or guardrail, different output style, or explicit routing in traces. “It felt cleaner as three agents” is not on that list.
Checklist before you split on audience:
- The two outputs have contradictory pass criteria (ops honesty vs customer tone)
- One template or style guide cannot hold both without lying
- The sender’s tool allowlist can be narrower than the drafter’s
- One writer still owns the irreversible send
Two audiences with one success metric is a prompt problem. Two audiences with two success metrics is a topology problem.
When does a trust boundary beat roleplay?
Trust is the reason. Roleplay is the costume.
Decision procedure:
- List every side-effecting tool.
- Tag each:
read,draft,write_reversible,write_irreversible. - Ask: should one persona ever hold
write_irreversibleand broadreadover sensitive stores in the same turn without a gate? - If no, either add a pre-execution policy gate or split the actor.
| Keep single | Split |
|---|---|
| Same trust tier; gate irreversible calls | Irreversible tools must never see raw untrusted retrieval in-prompt |
| One audience; tone handled by templates | Two audiences with contradictory success criteria |
| One SLO | Interactive vs batch cannot share budgets |
Skills / prompt swap by job_type | Separate tenancy, VPC, or credential set |
Microsoft’s architecture guide justifies multi-agent when a single agent cannot hold the work because of prompt complexity, tool overload, or security requirements. Security is the trust row. Tool overload is often an allowlist problem you have not tried yet.
Cognition’s 2026 follow-up is the production version of the same rule: extra agents may review, search, or advise; they do not get a second write. A clean-context reviewer is intelligence. A second refund tool is a second wallet.
- Irreversible tools sit behind a policy gate or a second, narrower agent
- The researcher cannot call the wallet, the refund API, or the send tool
- The actor cannot see raw forbidden documents — only a redacted pack
- HITL sits on the irreversible call, not on every search
If the only reason to split is “the critic should be a different vibe,” keep one loop and write a better evaluator.
When do timing and durability force a split?
Sometimes the conflict is clocks, not vibes.
- Agent A must answer in 8 seconds with retrieval only.
- Agent B must wait 6 hours for a human approval, then write.
Forcing both into one in-process loop creates either timeouts or heroic thread parking. Here a split — or a durable runtime with a clear handoff — is justified. See durability needs in the operating manual, and package the boundary like a handoff.
| Clock conflict | Stay single if | Split if |
|---|---|---|
| Interactive FAQ vs overnight enrich | One job, one SLO, pause is rare | Two SLOs, two budgets, two failure pages |
| Human approval mid-run | Durable runtime can park one loop | Approval lives in another process with its own tools |
| Burst parallelism | Parallel tool calls inside one agent | Parallel writers on the same entity |
Anthropic’s research system is the measured version of parallelism: independent search directions, separate context windows, a lead that synthesizes, and a citation pass that does not write to your CRM. They also said most coding tasks have fewer truly parallelizable pieces and that agents are still weak at real-time coordination. Breadth-first research is a timing-and-context conflict. A three-tool CRM note is not.
Checklist before splitting on timing:
- Same
job_id/trace_idacross the pause - Explicit handoff package (inputs, constraints, artifacts)
- One owner of the final write
- Idempotency keys on both sides
- A deadline and a kill switch on the slow side
If you can park one loop and resume it, you do not yet have two agents. You have one agent with a nap.
How do you prove a split helped?
Freeze the golden set. Run A/B. Do not argue from the diagram.
| Metric | Single baseline | Multi after split | Win condition |
|---|---|---|---|
| Pass rate | — | — | ≥ baseline |
| Cost per pass | — | — | ≤ baseline + agreed band |
| Escalate rate | — | — | ≤ baseline (or justified by safety) |
| p95 latency | — | — | Meets SLO |
| Debug minutes / incident | — | — | Down |
| Handoff reject rate | — | — | Not a new junk drawer |
| Revision depth | — | — | Down or justified |
If multi-agent raises cost and escalate rate while pass rate is flat, you bought coordination debt. Roll back.
Anthropic’s own economics are the warning label: 4× tokens for an agent versus chat, 15× for multi-agent, and a 90% research win that only pays when the task value covers the spend. Token usage explained 80% of their BrowseComp variance. A split that merely spends more tokens to look busy is not a win. A split that raises pass rate or cuts escalate rate or cuts cost per pass on the same cases is a win.
LangChain’s pattern table is the other receipt. On a one-shot “buy coffee” job, a subagent topology costs 4 model calls where skills, handoffs, or a router cost 3. On a repeat request, stateful skills or handoffs drop to 2 calls; stateless subagents stay at 4. On a multi-domain compare, isolated subagents can beat a skills dump on tokens — ~9K versus ~15K in their worked example — because each worker sees only its pack. Use that table the way it was written: pick the cheaper pattern for the job, then measure your golden set.
Procedure:
- Freeze N golden cases and the cost band on the single agent.
- Record pass rate, cost per pass, escalate rate, p95, debug minutes.
- Ship the split behind a flag. One new responsibility. One writer.
- Re-run the same N cases. No cherry-picks.
- Keep the split only if at least one primary metric wins and none of the three primaries blow the band.
Also track revision depth and handoff reject rate. A split that merely moves failures into handoff NACK spam is not a win.
Are supervisor patterns automatically better?
No. A supervisor looks like leadership. It often adds extra LLM calls for routing, another place prompts can drift, and a new loop that can disagree with the evaluator.
| Use a supervisor when | Skip it when |
|---|---|
| Routing is complex and changes often | A static job_type → tool allowlist works |
| Workers are truly autonomous services | Workers are functions you could call directly |
| You measured routing accuracy | You want org-chart cosplay |
| Workers need separate release trains | One prompt and one allowlist still pass |
LangChain’s own comparison is the telephone problem: a naive supervisor can lose quality because the worker cannot talk to the user, so the supervisor translates and burns more tokens. Their later docs still recommend a single agent with skills for simple, focused tasks. Many “supervisor multi-agent” systems are a switch statement with token overhead. Prefer the switch until the switch hurts.
Microsoft’s Cloud Adoption Framework says the same thing in enterprise language: distinct roles (planner, reviewer, executor) do not automatically justify multiple agents. Prototype persona switching, tool permissioning, and context gating first. Move to multi-agent only when that prototype fails in a way you cannot fix with prompts, retrieval, or policy.
- I can name the routing error rate on a labeled set
- A static map from
job_typeto allowlist is worse on that set - The supervisor cannot write; workers that write have one owner per entity
- Killing the supervisor does not leave two writers racing
If the supervisor is the only thing that can say no, you built a second evaluator and hid it. Put the evaluator on the job, not on the org chart.
When is a librarian or retriever agent worth it?
A separate librarian (retrieval-only agent) is worth it when:
- Retrieval needs a different model, index set, or tenancy boundary
- You must prove the writer never received raw forbidden documents
- Retrieval quality has its own evaluator and release train
It is not worth it when the “librarian” is one search tool call wrapped in a persona. That is a tool. Call the tool.
| Signal | Action |
|---|---|
| Same index, same rights, same latency budget | Keep retrieval as tools on the single agent |
| Cross-trust retrieval → generation | Librarian emits a redacted evidence pack; writer consumes only the pack |
| Retrieval has its own golden set and drift | Separate release train; still no write tools on the librarian |
| Writer must never see raw PII or secrets | Pack is the contract; raw docs stay in the restricted store |
Cognition’s later note is useful here: most working “multi-agent” setups in the wild are read-only subagents — web search, code search — that look like tool calls with extra context isolation. Anthropic’s citation pass is the same idea: a second loop that attributes claims, not a second loop that refunds a card.
If you cannot describe the evidence pack schema on one slide, you are not ready to split the librarian. You are ready to write a better search tool description.
How do shared-data races show up?
Classic race:
- Agent A reads ticket status
open - Agent B reads ticket status
open - A writes comment + status
pending - B writes comment + status
open(stale plan) - Customer sees contradictory updates
That is the lost-update problem. HTTP already has the primitive: send the version you observed. RFC 9110 If-Match / ETag is compare-and-set for the web. If the tag does not match, the write fails with 412 and the agent re-reads. Two agents that ignore the tag are not collaborating. They are overwriting.
Retries without identity create a second race: the network dies after the write, the other agent retries, and you get two refunds. Stripe stores the first result for an Idempotency-Key so a retry is a replay, not a second charge. Your tool writes need the same key. A new key per hop is how you double-spend.
| Mitigation | What it stops | What it does not stop |
|---|---|---|
| One writer agent per entity type | Dual wallets on the same row | A single writer with a bad plan |
If-Match / etag / observed_version | Silent last-write-wins | A writer that never sends the version |
| Idempotency keys on tool writes | Duplicate side effects on retry | Two different keys for one intent |
Handoff package includes observed_version | Stale plans crossing a queue | A pack that omits the version |
If two agents can write the same row, you do not have collaboration — you have a distributed systems homework assignment. Assign it on purpose or don’t.
Cognition’s principle 2 is this race in prose: actions carry implicit decisions. Two writers who never saw each other’s choices will disagree in the database, not in the chat.
What is the debug cost of crossing process boundaries?
Every process boundary multiplies work you already under-budgeted.
| Cost | Symptom | Fix before you split |
|---|---|---|
| Observability | Missing trace_id propagation | Inject W3C traceparent on every hop |
| Repro | Can’t replay without both queues warm | One replay harness, both sides |
| Ownership | “Their agent failed” pages in Slack | One on-call for the job, not the costume |
| Latency | Serialization + queue wait | Measure p95 of the hop itself |
| Security | Broader network attack surface | Narrow credentials per hop |
OpenTelemetry defaults to those W3C headers so a backend can stitch one trace from two processes. If your second agent mints a new run_id and drops traceparent, you did not gain modularity. You gained a murder mystery.
Anthropic said the quiet part in the research post: without production tracing they could not tell whether “not finding obvious information” was a bad query, a bad source, or a tool failure. Multi-agent makes that worse because the failure lives in the interaction, not in one prompt.
Budget an extra day of harness work per boundary. If the pilot is five days, that is a real fraction of the calendar — which is why Spurlock defaults to single-agent in pilot scope.
- Same
trace_idfrom trigger to terminal write - Replay works with one command
- One owner pages, even if two processes ran
- Credentials on hop B cannot do hop A’s reads
If you cannot afford the harness, you cannot afford the split.
Failure mode: five agents, one missing criterion
What breaks: Research, draft, critique, SEO, and send agents form a pipeline. The critique agent praises tone. Nobody checks “correct refund amount.” The send agent has email rights. A wrong refund notice ships.
What it costs: Customer trust, finance cleanup, and a week of blame aimed at “hallucination.”
What you do instead: One agent with an evaluator that includes the amount check; email tool behind HITL until the golden set is green. Add agents only if a trust split requires it — and keep a single write authority.
| Costume | What it actually checked | What it missed |
|---|---|---|
| Research | “Sources exist” | Amount in the ledger |
| Draft | “Reads like us” | Amount in the ledger |
| Critique | “Tone is fine” | Amount in the ledger |
| SEO | “Title is punchy” | Amount in the ledger |
| Send | “SMTP accepted” | Amount in the ledger |
Five pass rates of 100% and a composed fail of 100%. That is an eval gap, not a model gap. Cognition’s later review-loop note is the exception that proves the rule: a clean-context reviewer can catch bugs the writer is blind to, if writes stay single-threaded and the communication bridge filters out-of-scope nits. A fifth writer with SMTP is not a reviewer.
I have watched this exact shape on production automations: the pipeline looks busy, the customer-facing send is the only tool that matters, and the missing criterion was never in anyone’s rubric. Fix the rubric. Then decide if you still need a second loop.
What is the Spurlock default topology?
| Stage | Topology |
|---|---|
| Pilot (5-day) | Single agent, tool allowlist, evaluator, sandbox, kill switch |
| First production job | Still single unless a trust / audience / timing conflict is documented |
| Scale | Split along those conflicts; handoff packages; shared trace ids |
| Never default | Supervisor cosplay for a three-tool CRM note |
Default: expand tools and tighten gates before inventing colleagues.
That matches the vendor ladder. Anthropic: simplest solution, add complexity when it improves outcomes. OpenAI: smallest agent that can own the task. Microsoft: single agent with tools is the enterprise default. LangChain: try one agent with tools and skills first. Cognition: single-threaded writes, even after you add intelligence around the writer.
On a Spurlock agentic pilot I will not staff a crew because the deck has three stick figures. I will staff one loop, one golden set, and three numbers: pass rate, cost per pass, escalate rate. If those numbers later demand a librarian or a sender, we split on purpose.
| Question | If yes → |
|---|---|
| Can one allowlist + policy gate express the trust model? | Stay single |
| Do two audiences need contradictory “good” outputs? | Split by audience |
| Must irreversible tools be isolated from raw retrieval? | Split librarian / actor or harden gates until equivalent |
| Is parallelism measured and bottlenecked? | Consider parallel workers with one merger + one writer |
| Is the only reason an org-chart slide? | Stay single |
How do you migrate from single to multi without regret?
Do not rewrite the system. Move one responsibility. Keep the writer singular. Keep the golden set frozen.
- Freeze golden cases and the cost band on the single agent.
- Document the conflict (trust / audience / timing) in one paragraph.
- Define the handoff schema before writing the second agent. See multi-agent handoffs.
- Move one responsibility; keep write authority singular.
- Re-run the golden set; compare pass rate, cost per pass, and escalate rate.
- Only then add a third agent.
Rollback plan: feature-flag the second agent and route back to the single loop in one config change.
| Step | Done when |
|---|---|
| Baseline | Pass / cost / escalate recorded on N cases |
| Conflict note | One paragraph a skeptic can disagree with |
| Pack schema | Inputs, constraints, artifacts, observed_version, trace_id |
| First split | One new loop, zero new writers |
| Proof | At least one primary metric wins; none blow the band |
| Flag | Off switch returns traffic to the single loop |
Anthropic’s early research agents spawned 50 subagents for simple queries and duplicated searches. Their fix was not “more agents.” It was effort scaling, task boundaries, and evals on about 20 real queries before anyone built a hundred-case harness. Start that small. If the split does not move the three primaries, you do not have a migration. You have a branch to delete.
What are the anti-patterns?
Agent per function. FormatDateAgent is a function. Call format_date.
Critique without authority is expensive commentary. If the critic cannot stop the write, it is a log line with an invoice.
New run ids per hop make production undebuggable. Propagate traceparent or do not split.
Multi-agent to fix a missing evaluator adds speakers, not truth. Write the amount check. Then decide if you still need a second loop.
Parallel writers on one entity is a race with a product name. One writer. Read-only helpers if you must.
Supervisor as the only deny path hides the evaluator inside a router. Put deny on the tool.
Research-eval cargo cult copies Anthropic’s 90.2% headline onto a refund workflow. Their own post said shared-context, high-dependency jobs — coding is the example — are a poor fit, and that multi-agent is a token-spending strategy for breadth-first work whose value covers 15× chat.
| Anti-pattern | Tell | Replacement |
|---|---|---|
| Agent per function | Noun ends in Agent, body is ten lines | Function + schema |
| Critique without teeth | Critic cannot NACK the write | Evaluator on the job |
Fresh run_id per hop | Two traces for one customer | W3C traceparent |
| Crew as missing rubric | Five 100% scores, one bad send | One evaluator, one writer |
| Parallel wallets | Two tools can PATCH the same row | One writer + If-Match |
Bravery is not a topology.
FAQ
Is a supervisor pattern always better?
No. Supervisors add routing calls and another failure point. LangChain’s supervisor benchmarks show quality can drop because the supervisor translates and the worker cannot talk to the user. Prefer a static job_type router or tool allowlist until measured complexity forces a learned supervisor — and keep a single write authority either way.
How does this relate to multi-agent handoffs?
This post decides whether to split. Multi-agent handoffs defines the package, ownership, and trace propagation once you split. Do not implement handoff theater for a single loop. If you cannot name the conflict in one paragraph, you are not ready for a pack schema.
When is a librarian/retriever agent worth it?
When retrieval needs a different trust boundary, index set, or release train than the writer — and the librarian emits a constrained evidence pack. If retrieval is one tool with the same rights, keep it on the single agent. Cognition’s working setups are mostly read-only search helpers, not second writers.
How do shared-data races show up?
Two agents read the same entity, both plan, both write, and the last write wins. Customers see contradictory updates. Fix with one writer per entity class, compare-and-set tools (If-Match / etag), idempotency keys, and versioned handoff packages. If two loops can PATCH the same row, you assigned distributed-systems homework.
What’s the debug cost of crossing process boundaries?
You pay in tracing, reproduction, ownership, latency, and security surface. W3C Trace Context and OpenTelemetry exist so two processes can still be one story. Budget real engineering time per boundary; on a short pilot, that cost alone argues for staying single until a conflict is proven.
What’s the Spurlock default topology?
Single agent with many tools, evaluator-in-the-loop, sandboxed writes, and a kill switch. Split only on documented trust, audience, or timing conflicts — then prove the split on the same golden set with pass rate, cost per pass, and escalate rate. Start that path on /agentic.
CTA
One loop until a real conflict shows up. Then hand off on purpose.
What questions does this article answer?
- Is a supervisor pattern always better?
- No. Supervisors add routing calls and another failure point. LangChain’s supervisor benchmarks show quality can drop because the supervisor translates and the worker cannot talk to the user. Prefer a static `job_type` router or tool allowlist until measured complexity forces a learned supervisor — and keep a single write authority either way.
- How does this relate to multi-agent handoffs?
- This post decides whether to split. [Multi-agent handoffs](/blog/multi-agent-handoffs) defines the package, ownership, and trace propagation once you split. Do not implement handoff theater for a single loop. If you cannot name the conflict in one paragraph, you are not ready for a pack schema.
- When is a librarian/retriever agent worth it?
- When retrieval needs a different trust boundary, index set, or release train than the writer — and the librarian emits a constrained evidence pack. If retrieval is one tool with the same rights, keep it on the single agent. Cognition’s working setups are mostly read-only search helpers, not second writers.
- How do shared-data races show up?
- Two agents read the same entity, both plan, both write, and the last write wins. Customers see contradictory updates. Fix with one writer per entity class, compare-and-set tools (`If-Match` / etag), idempotency keys, and versioned handoff packages. If two loops can PATCH the same row, you assigned distributed-systems homework.
- What’s the debug cost of crossing process boundaries?
- You pay in tracing, reproduction, ownership, latency, and security surface. W3C Trace Context and OpenTelemetry exist so two processes can still be one story. Budget real engineering time per boundary; on a short pilot, that cost alone argues for staying single until a conflict is proven.
- What’s the Spurlock default topology?
- Single agent with many tools, evaluator-in-the-loop, sandboxed writes, and a kill switch. Split only on documented trust, audience, or timing conflicts — then prove the split on the same golden set with pass rate, cost per pass, and escalate rate. Start that path on [/agentic](/agentic).
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.