Spurlock Studios
Contact
Share LinkedIn X
A small stack of coins. Thesis: SINGLE AGENT FIRST SPLIT ONLY.

Should you start with a single agent or go multi-agent? Start single. One agent with a clear job, a bounded tool set, an evaluator, and a kill switch beats a committee of prompts that hand work to each other for theater. Split only when trust, audience, or timing actually conflict — then prove the split helped with pass rate, cost per pass, and escalate rate on the same golden set.

This spoke belongs to the Agentic Systems Operating Manual. Once you do split, package the boundary with multi-agent handoffs. This post owns the decision to split at all.

That is not a taste take. Anthropic tells teams to find the simplest solution and add complexity only when it demonstrably improves outcomes. OpenAI says define the smallest agent that can own a clear task, then add agents only for separate ownership, tools, or approval policies. Microsoft calls a single agent with tools the usual enterprise default. LangChain says a single agent with the right tools and prompt often matches a crew. I have spent 20,000+ hours on agentic systems and shipped 500+ automations. The expensive failures were almost always a second write path I did not need.

The short answer

  • Default topology: one agent, many tools, one evaluator, one write authority.
  • Still single-agent: planner / tool-user / reviser states inside one loop — not three products.
  • Split when tools need different trust levels, outputs serve conflicting audiences, or timing and durability requirements diverge.
  • Premature multi-agent creates coordination bugs: lost context, ABAB loops, shared-data races, and debug across process walls.
  • Proof of a good split: better pass rate or lower escalate rate or lower cost per pass on the same golden set — not a prettier diagram.

What still counts as a single agent with many tools?

A single agent is one control loop with one terminal authority. Tools are capabilities. States are phases. A human pause is a gate. None of those are teammates.

PatternWhy it is still one agent
Many tools (CRM, search, calendar)Tools are capabilities, not teammates
Plan → act → evaluate → revise statesStates are phases of one job
Router that picks a tool subsetFiltering tools ≠ spawning agents
Human approval pause mid-runHITL is a gate, not a second agent
Specialist prompts swapped by job_typeConfig, not a multi-agent system
Skills loaded on demandLangChain keeps one agent in control while it loads context

You have multi-agent when another autonomous loop can take actions with its own tool rights, memory, or stop conditions — especially across process or queue boundaries.

If your “researcher agent” cannot terminate the job and cannot write, it may be a function with a costume. Costumes are fine; do not bill them as architecture.

  • One run_id / trace_id owns the job from trigger to terminal write
  • One allowlist decides which tools may fire this turn
  • One evaluator scores the composed result, not each costume
  • One kill switch stops every side effect

Microsoft’s complexity ladder starts at a direct model call, then a single agent with tools, and only then multi-agent. Skip a rung only when a measured conflict forces it.

Why does fashion push multi-agent?

Demos love casts of characters. Role names make slides readable. Frameworks make crews easy to spin up. None of that proves you needed more than one write path.

Multi-agent hype usually optimizes for:

  1. Narrative clarity in a demo
  2. Parallelism you have not measured
  3. Mimicking an org chart

Production optimizes for:

  1. Correct side effects
  2. Debuggable failures
  3. Cost per passing job

When those conflict, ship the boring single loop.

Cognition put the failure in two sentences: share context, and remember that actions carry implicit decisions. Two writers who never saw each other’s choices produce a merge conflict with extra tokens. Their later note is narrower, not a reversal: extra agents can add intelligence, but writes stay single-threaded. That is the same rule I use on client work.

Signal you are buying fashionWhat to do instead
The deck has a “researcher / writer / critic” triangleOne loop with an evaluator and a deny gate
The framework scaffolded three agents in ten minutesKeep the scaffold; delete two writers
Someone said “we need parallelism” with no p95Measure the single-agent bottleneck first
Role names match the org chartMap trust tiers, not job titles

Fashion is cheap in a slide. Coordination is expensive in a queue.

Why does premature multi-agent create coordination bugs?

These are system bugs mislabeled as “the model is dumb.” Fix topology before you buy a bigger model.

BugHow it shows upWhat actually broke
Lost intentAgent B never sees the constraint Agent A “agreed” in proseContext was not packaged
Dual write authorityTwo agents update the same CRM field with different draftsTwo wallets, one row
ABAB oscillationA hands to B; B rejects; A “fixes”; infinite courtesyNo terminal owner
Shared-data racesBoth read stale state; both write; last write wins silentlyNo compare-and-set
Trace fractureNew run_id per agent; nobody can reconstruct the storyNo propagated traceparent
Eval gapsEach agent “looks fine”; the composed job failsEvaluator scored costumes, not the job

Cognition’s Flappy Bird example is the same bug in costume: one subagent builds Mario pipes, the other builds a bird that is not a game asset, and the merger inherits two implicit decisions that never met. Copying the original task into both prompts does not fix it. The conversation already made decisions the subagents never saw.

Anthropic’s research writeup is the other side of the same coin. Their lead-plus-subagent setup beat a single-agent baseline by 90.2% on an internal research eval — breadth-first, parallel search, high value. They also said agents use about 4× the tokens of chat, multi-agent about 15×, and that coding work with shared context and tight dependencies is a poor fit. Token spend explained 80% of BrowseComp variance. That is a research receipt, not a license to spawn a refund committee.

Microsoft’s Cloud Adoption Framework is blunt about the tax: every extra agent adds protocol design, error handling, state sync, prompt work, monitoring, credentials, and handoff latency. Pay that tax for a conflict. Do not pay it for a slide.

When do conflicting audiences force a split?

Split when one loop cannot honestly serve two masters.

ConflictExampleSplit shape
AudienceInternal ops notes vs customer-facing emailDrafter (internal tools) → Sender (email-only tools)
TrustRead-only research vs irreversible refundsResearcher (no wallet) → Actor (refund tool + HITL)
TimingFast FAQ answers vs overnight batch enrichmentOnline agent vs batch worker with different SLOs
CompliancePII-heavy retrieval vs public content generationLibrarian in a restricted VPC → Writer with redacted packs

If you can solve the conflict with tool allowlists and policy gates inside one agent, prefer that. A split is for when allowlists still leave a trust or SLO collision.

OpenAI’s split list matches this, not a casting call: different tool or MCP surface, different approval policy or guardrail, different output style, or explicit routing in traces. “It felt cleaner as three agents” is not on that list.

Checklist before you split on audience:

  • The two outputs have contradictory pass criteria (ops honesty vs customer tone)
  • One template or style guide cannot hold both without lying
  • The sender’s tool allowlist can be narrower than the drafter’s
  • One writer still owns the irreversible send

Two audiences with one success metric is a prompt problem. Two audiences with two success metrics is a topology problem.

When does a trust boundary beat roleplay?

Trust is the reason. Roleplay is the costume.

Decision procedure:

  1. List every side-effecting tool.
  2. Tag each: read, draft, write_reversible, write_irreversible.
  3. Ask: should one persona ever hold write_irreversible and broad read over sensitive stores in the same turn without a gate?
  4. If no, either add a pre-execution policy gate or split the actor.
Keep singleSplit
Same trust tier; gate irreversible callsIrreversible tools must never see raw untrusted retrieval in-prompt
One audience; tone handled by templatesTwo audiences with contradictory success criteria
One SLOInteractive vs batch cannot share budgets
Skills / prompt swap by job_typeSeparate tenancy, VPC, or credential set

Microsoft’s architecture guide justifies multi-agent when a single agent cannot hold the work because of prompt complexity, tool overload, or security requirements. Security is the trust row. Tool overload is often an allowlist problem you have not tried yet.

Cognition’s 2026 follow-up is the production version of the same rule: extra agents may review, search, or advise; they do not get a second write. A clean-context reviewer is intelligence. A second refund tool is a second wallet.

  • Irreversible tools sit behind a policy gate or a second, narrower agent
  • The researcher cannot call the wallet, the refund API, or the send tool
  • The actor cannot see raw forbidden documents — only a redacted pack
  • HITL sits on the irreversible call, not on every search

If the only reason to split is “the critic should be a different vibe,” keep one loop and write a better evaluator.

When do timing and durability force a split?

Sometimes the conflict is clocks, not vibes.

  • Agent A must answer in 8 seconds with retrieval only.
  • Agent B must wait 6 hours for a human approval, then write.

Forcing both into one in-process loop creates either timeouts or heroic thread parking. Here a split — or a durable runtime with a clear handoff — is justified. See durability needs in the operating manual, and package the boundary like a handoff.

Clock conflictStay single ifSplit if
Interactive FAQ vs overnight enrichOne job, one SLO, pause is rareTwo SLOs, two budgets, two failure pages
Human approval mid-runDurable runtime can park one loopApproval lives in another process with its own tools
Burst parallelismParallel tool calls inside one agentParallel writers on the same entity

Anthropic’s research system is the measured version of parallelism: independent search directions, separate context windows, a lead that synthesizes, and a citation pass that does not write to your CRM. They also said most coding tasks have fewer truly parallelizable pieces and that agents are still weak at real-time coordination. Breadth-first research is a timing-and-context conflict. A three-tool CRM note is not.

Checklist before splitting on timing:

  • Same job_id / trace_id across the pause
  • Explicit handoff package (inputs, constraints, artifacts)
  • One owner of the final write
  • Idempotency keys on both sides
  • A deadline and a kill switch on the slow side

If you can park one loop and resume it, you do not yet have two agents. You have one agent with a nap.

How do you prove a split helped?

Freeze the golden set. Run A/B. Do not argue from the diagram.

MetricSingle baselineMulti after splitWin condition
Pass rate——≥ baseline
Cost per pass——≤ baseline + agreed band
Escalate rate——≤ baseline (or justified by safety)
p95 latency——Meets SLO
Debug minutes / incident——Down
Handoff reject rate——Not a new junk drawer
Revision depth——Down or justified

If multi-agent raises cost and escalate rate while pass rate is flat, you bought coordination debt. Roll back.

Anthropic’s own economics are the warning label: 4× tokens for an agent versus chat, 15× for multi-agent, and a 90% research win that only pays when the task value covers the spend. Token usage explained 80% of their BrowseComp variance. A split that merely spends more tokens to look busy is not a win. A split that raises pass rate or cuts escalate rate or cuts cost per pass on the same cases is a win.

LangChain’s pattern table is the other receipt. On a one-shot “buy coffee” job, a subagent topology costs 4 model calls where skills, handoffs, or a router cost 3. On a repeat request, stateful skills or handoffs drop to 2 calls; stateless subagents stay at 4. On a multi-domain compare, isolated subagents can beat a skills dump on tokens — ~9K versus ~15K in their worked example — because each worker sees only its pack. Use that table the way it was written: pick the cheaper pattern for the job, then measure your golden set.

Procedure:

  1. Freeze N golden cases and the cost band on the single agent.
  2. Record pass rate, cost per pass, escalate rate, p95, debug minutes.
  3. Ship the split behind a flag. One new responsibility. One writer.
  4. Re-run the same N cases. No cherry-picks.
  5. Keep the split only if at least one primary metric wins and none of the three primaries blow the band.

Also track revision depth and handoff reject rate. A split that merely moves failures into handoff NACK spam is not a win.

Are supervisor patterns automatically better?

No. A supervisor looks like leadership. It often adds extra LLM calls for routing, another place prompts can drift, and a new loop that can disagree with the evaluator.

Use a supervisor whenSkip it when
Routing is complex and changes oftenA static job_type → tool allowlist works
Workers are truly autonomous servicesWorkers are functions you could call directly
You measured routing accuracyYou want org-chart cosplay
Workers need separate release trainsOne prompt and one allowlist still pass

LangChain’s own comparison is the telephone problem: a naive supervisor can lose quality because the worker cannot talk to the user, so the supervisor translates and burns more tokens. Their later docs still recommend a single agent with skills for simple, focused tasks. Many “supervisor multi-agent” systems are a switch statement with token overhead. Prefer the switch until the switch hurts.

Microsoft’s Cloud Adoption Framework says the same thing in enterprise language: distinct roles (planner, reviewer, executor) do not automatically justify multiple agents. Prototype persona switching, tool permissioning, and context gating first. Move to multi-agent only when that prototype fails in a way you cannot fix with prompts, retrieval, or policy.

  • I can name the routing error rate on a labeled set
  • A static map from job_type to allowlist is worse on that set
  • The supervisor cannot write; workers that write have one owner per entity
  • Killing the supervisor does not leave two writers racing

If the supervisor is the only thing that can say no, you built a second evaluator and hid it. Put the evaluator on the job, not on the org chart.

When is a librarian or retriever agent worth it?

A separate librarian (retrieval-only agent) is worth it when:

  1. Retrieval needs a different model, index set, or tenancy boundary
  2. You must prove the writer never received raw forbidden documents
  3. Retrieval quality has its own evaluator and release train

It is not worth it when the “librarian” is one search tool call wrapped in a persona. That is a tool. Call the tool.

SignalAction
Same index, same rights, same latency budgetKeep retrieval as tools on the single agent
Cross-trust retrieval → generationLibrarian emits a redacted evidence pack; writer consumes only the pack
Retrieval has its own golden set and driftSeparate release train; still no write tools on the librarian
Writer must never see raw PII or secretsPack is the contract; raw docs stay in the restricted store

Cognition’s later note is useful here: most working “multi-agent” setups in the wild are read-only subagents — web search, code search — that look like tool calls with extra context isolation. Anthropic’s citation pass is the same idea: a second loop that attributes claims, not a second loop that refunds a card.

If you cannot describe the evidence pack schema on one slide, you are not ready to split the librarian. You are ready to write a better search tool description.

How do shared-data races show up?

Classic race:

  1. Agent A reads ticket status open
  2. Agent B reads ticket status open
  3. A writes comment + status pending
  4. B writes comment + status open (stale plan)
  5. Customer sees contradictory updates

That is the lost-update problem. HTTP already has the primitive: send the version you observed. RFC 9110 If-Match / ETag is compare-and-set for the web. If the tag does not match, the write fails with 412 and the agent re-reads. Two agents that ignore the tag are not collaborating. They are overwriting.

Retries without identity create a second race: the network dies after the write, the other agent retries, and you get two refunds. Stripe stores the first result for an Idempotency-Key so a retry is a replay, not a second charge. Your tool writes need the same key. A new key per hop is how you double-spend.

MitigationWhat it stopsWhat it does not stop
One writer agent per entity typeDual wallets on the same rowA single writer with a bad plan
If-Match / etag / observed_versionSilent last-write-winsA writer that never sends the version
Idempotency keys on tool writesDuplicate side effects on retryTwo different keys for one intent
Handoff package includes observed_versionStale plans crossing a queueA pack that omits the version

If two agents can write the same row, you do not have collaboration — you have a distributed systems homework assignment. Assign it on purpose or don’t.

Cognition’s principle 2 is this race in prose: actions carry implicit decisions. Two writers who never saw each other’s choices will disagree in the database, not in the chat.

What is the debug cost of crossing process boundaries?

Every process boundary multiplies work you already under-budgeted.

CostSymptomFix before you split
ObservabilityMissing trace_id propagationInject W3C traceparent on every hop
ReproCan’t replay without both queues warmOne replay harness, both sides
Ownership“Their agent failed” pages in SlackOne on-call for the job, not the costume
LatencySerialization + queue waitMeasure p95 of the hop itself
SecurityBroader network attack surfaceNarrow credentials per hop

OpenTelemetry defaults to those W3C headers so a backend can stitch one trace from two processes. If your second agent mints a new run_id and drops traceparent, you did not gain modularity. You gained a murder mystery.

Anthropic said the quiet part in the research post: without production tracing they could not tell whether “not finding obvious information” was a bad query, a bad source, or a tool failure. Multi-agent makes that worse because the failure lives in the interaction, not in one prompt.

Budget an extra day of harness work per boundary. If the pilot is five days, that is a real fraction of the calendar — which is why Spurlock defaults to single-agent in pilot scope.

  • Same trace_id from trigger to terminal write
  • Replay works with one command
  • One owner pages, even if two processes ran
  • Credentials on hop B cannot do hop A’s reads

If you cannot afford the harness, you cannot afford the split.

Failure mode: five agents, one missing criterion

What breaks: Research, draft, critique, SEO, and send agents form a pipeline. The critique agent praises tone. Nobody checks “correct refund amount.” The send agent has email rights. A wrong refund notice ships.

What it costs: Customer trust, finance cleanup, and a week of blame aimed at “hallucination.”

What you do instead: One agent with an evaluator that includes the amount check; email tool behind HITL until the golden set is green. Add agents only if a trust split requires it — and keep a single write authority.

CostumeWhat it actually checkedWhat it missed
Research“Sources exist”Amount in the ledger
Draft“Reads like us”Amount in the ledger
Critique“Tone is fine”Amount in the ledger
SEO“Title is punchy”Amount in the ledger
Send“SMTP accepted”Amount in the ledger

Five pass rates of 100% and a composed fail of 100%. That is an eval gap, not a model gap. Cognition’s later review-loop note is the exception that proves the rule: a clean-context reviewer can catch bugs the writer is blind to, if writes stay single-threaded and the communication bridge filters out-of-scope nits. A fifth writer with SMTP is not a reviewer.

I have watched this exact shape on production automations: the pipeline looks busy, the customer-facing send is the only tool that matters, and the missing criterion was never in anyone’s rubric. Fix the rubric. Then decide if you still need a second loop.

What is the Spurlock default topology?

StageTopology
Pilot (5-day)Single agent, tool allowlist, evaluator, sandbox, kill switch
First production jobStill single unless a trust / audience / timing conflict is documented
ScaleSplit along those conflicts; handoff packages; shared trace ids
Never defaultSupervisor cosplay for a three-tool CRM note

Default: expand tools and tighten gates before inventing colleagues.

That matches the vendor ladder. Anthropic: simplest solution, add complexity when it improves outcomes. OpenAI: smallest agent that can own the task. Microsoft: single agent with tools is the enterprise default. LangChain: try one agent with tools and skills first. Cognition: single-threaded writes, even after you add intelligence around the writer.

On a Spurlock agentic pilot I will not staff a crew because the deck has three stick figures. I will staff one loop, one golden set, and three numbers: pass rate, cost per pass, escalate rate. If those numbers later demand a librarian or a sender, we split on purpose.

QuestionIf yes →
Can one allowlist + policy gate express the trust model?Stay single
Do two audiences need contradictory “good” outputs?Split by audience
Must irreversible tools be isolated from raw retrieval?Split librarian / actor or harden gates until equivalent
Is parallelism measured and bottlenecked?Consider parallel workers with one merger + one writer
Is the only reason an org-chart slide?Stay single

How do you migrate from single to multi without regret?

Do not rewrite the system. Move one responsibility. Keep the writer singular. Keep the golden set frozen.

  1. Freeze golden cases and the cost band on the single agent.
  2. Document the conflict (trust / audience / timing) in one paragraph.
  3. Define the handoff schema before writing the second agent. See multi-agent handoffs.
  4. Move one responsibility; keep write authority singular.
  5. Re-run the golden set; compare pass rate, cost per pass, and escalate rate.
  6. Only then add a third agent.

Rollback plan: feature-flag the second agent and route back to the single loop in one config change.

StepDone when
BaselinePass / cost / escalate recorded on N cases
Conflict noteOne paragraph a skeptic can disagree with
Pack schemaInputs, constraints, artifacts, observed_version, trace_id
First splitOne new loop, zero new writers
ProofAt least one primary metric wins; none blow the band
FlagOff switch returns traffic to the single loop

Anthropic’s early research agents spawned 50 subagents for simple queries and duplicated searches. Their fix was not “more agents.” It was effort scaling, task boundaries, and evals on about 20 real queries before anyone built a hundred-case harness. Start that small. If the split does not move the three primaries, you do not have a migration. You have a branch to delete.

What are the anti-patterns?

Agent per function. FormatDateAgent is a function. Call format_date.

Critique without authority is expensive commentary. If the critic cannot stop the write, it is a log line with an invoice.

New run ids per hop make production undebuggable. Propagate traceparent or do not split.

Multi-agent to fix a missing evaluator adds speakers, not truth. Write the amount check. Then decide if you still need a second loop.

Parallel writers on one entity is a race with a product name. One writer. Read-only helpers if you must.

Supervisor as the only deny path hides the evaluator inside a router. Put deny on the tool.

Research-eval cargo cult copies Anthropic’s 90.2% headline onto a refund workflow. Their own post said shared-context, high-dependency jobs — coding is the example — are a poor fit, and that multi-agent is a token-spending strategy for breadth-first work whose value covers 15× chat.

Anti-patternTellReplacement
Agent per functionNoun ends in Agent, body is ten linesFunction + schema
Critique without teethCritic cannot NACK the writeEvaluator on the job
Fresh run_id per hopTwo traces for one customerW3C traceparent
Crew as missing rubricFive 100% scores, one bad sendOne evaluator, one writer
Parallel walletsTwo tools can PATCH the same rowOne writer + If-Match

Bravery is not a topology.

FAQ

Is a supervisor pattern always better?

No. Supervisors add routing calls and another failure point. LangChain’s supervisor benchmarks show quality can drop because the supervisor translates and the worker cannot talk to the user. Prefer a static job_type router or tool allowlist until measured complexity forces a learned supervisor — and keep a single write authority either way.

How does this relate to multi-agent handoffs?

This post decides whether to split. Multi-agent handoffs defines the package, ownership, and trace propagation once you split. Do not implement handoff theater for a single loop. If you cannot name the conflict in one paragraph, you are not ready for a pack schema.

When is a librarian/retriever agent worth it?

When retrieval needs a different trust boundary, index set, or release train than the writer — and the librarian emits a constrained evidence pack. If retrieval is one tool with the same rights, keep it on the single agent. Cognition’s working setups are mostly read-only search helpers, not second writers.

How do shared-data races show up?

Two agents read the same entity, both plan, both write, and the last write wins. Customers see contradictory updates. Fix with one writer per entity class, compare-and-set tools (If-Match / etag), idempotency keys, and versioned handoff packages. If two loops can PATCH the same row, you assigned distributed-systems homework.

What’s the debug cost of crossing process boundaries?

You pay in tracing, reproduction, ownership, latency, and security surface. W3C Trace Context and OpenTelemetry exist so two processes can still be one story. Budget real engineering time per boundary; on a short pilot, that cost alone argues for staying single until a conflict is proven.

What’s the Spurlock default topology?

Single agent with many tools, evaluator-in-the-loop, sandboxed writes, and a kill switch. Split only on documented trust, audience, or timing conflicts — then prove the split on the same golden set with pass rate, cost per pass, and escalate rate. Start that path on /agentic.

CTA

One loop until a real conflict shows up. Then hand off on purpose.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

Is a supervisor pattern always better?
No. Supervisors add routing calls and another failure point. LangChain’s supervisor benchmarks show quality can drop because the supervisor translates and the worker cannot talk to the user. Prefer a static `job_type` router or tool allowlist until measured complexity forces a learned supervisor — and keep a single write authority either way.
How does this relate to multi-agent handoffs?
This post decides whether to split. [Multi-agent handoffs](/blog/multi-agent-handoffs) defines the package, ownership, and trace propagation once you split. Do not implement handoff theater for a single loop. If you cannot name the conflict in one paragraph, you are not ready for a pack schema.
When is a librarian/retriever agent worth it?
When retrieval needs a different trust boundary, index set, or release train than the writer — and the librarian emits a constrained evidence pack. If retrieval is one tool with the same rights, keep it on the single agent. Cognition’s working setups are mostly read-only search helpers, not second writers.
How do shared-data races show up?
Two agents read the same entity, both plan, both write, and the last write wins. Customers see contradictory updates. Fix with one writer per entity class, compare-and-set tools (`If-Match` / etag), idempotency keys, and versioned handoff packages. If two loops can PATCH the same row, you assigned distributed-systems homework.
What’s the debug cost of crossing process boundaries?
You pay in tracing, reproduction, ownership, latency, and security surface. W3C Trace Context and OpenTelemetry exist so two processes can still be one story. Budget real engineering time per boundary; on a short pilot, that cost alone argues for staying single until a conflict is proven.
What’s the Spurlock default topology?
Single agent with many tools, evaluator-in-the-loop, sandboxed writes, and a kill switch. Split only on documented trust, audience, or timing conflicts — then prove the split on the same golden set with pass rate, cost per pass, and escalate rate. Start that path on [/agentic](/agentic).
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot