Multi-Agent Handoffs Without Lost Context
More agents are not maturity. Typed handoff packages — goal, artifacts, failures, remaining budget — keep context without a shared contaminated transcript.
William Spurlock Founder — Spurlock Studios Updated 24 MIN
The fastest way to lose trust in a multi-agent demo is a hop that drops the ticket ID, double-sends the email, or “remembers” a plan the evaluator already rejected. More agents are not a maturity model. A typed handoff contract is.
This spoke sits under the Agentic Systems Operating Manual. The decision to split at all lives in single vs multi-agent. This page owns the package that crosses the split.
The short answer
- Default hop: Agent A finishes a stage, emits a schema-valid package, Agent B starts clean. No shared scratch pad.
- Minimum fields:
job_id, goal, constraints, typed artifacts, open questions, tools tried, evaluator verdict, remaining budget, memory refs, next hint. - Transport ≠ contract. Framework
transfer_to_*tools,input_filters, and graphCommands move control. They do not decide what “done” means. - Fail closed. Invalid package →
escalate. Do not let B invent missing IDs. - One writer per external resource per run. Handoffs retry; CRM patches and customer sends must not.
When is a handoff contract worth more than another agent?
Split when work naturally specializes and you can name the interface. Do not split because a slide said “multi-agent.” One worker plus one evaluator beats five chatty peers that share a muddy transcript.
| You have a real split when… | You still have one job when… |
|---|---|
| Research vs draft vs compliance check | The same person would do all three in one sitting |
| Intake classification vs enrichment vs write-back | Write-back is just the last function in one loop |
| Librarian (retrieval) vs worker (prose) vs evaluator (judge) | The “librarian” cannot terminate and cannot write |
| Tools need different trust levels | Every hop holds the same tools |
| Outputs serve conflicting audiences | Everyone is writing the same internal note |
Microsoft’s architecture guide is blunt about the cheaper path: if the right agent or sequence is identifiable from the initial input, use deterministic routing. Do not invent a chairman agent to pick a path you already know.
Anthropic’s research system is the other pole — breadth-first work that cannot be hardcoded. Their lead agent writes a plan to memory, then spawns subagents with a self-contained task: objective, output format, tools, and a stop condition. That is a contract. It is not a group chat.
- You can name the output artifact type in one noun
- You can name the tools this hop may hold
- You can name the evaluator criteria before anyone writes
- Two stages that share tools and criteria collapse into one stage
- The two-agent path already meets pass-rate and cost targets
If those boxes stay empty, you need a better single loop, not a third persona.
What belongs in a typed handoff package?
A handoff package is a form. The next agent reads the form. It does not inherit a novel.
{
"schema": "handoff.v1",
"job_id": "tkt_18422",
"goal": "Draft internal triage summary; do not send to customer",
"constraints": ["no refund promises", "cite help center or no_match"],
"artifacts": [{ "type": "summary_md", "uri": "s3://ops/tkt_18422/summary.md" }],
"open_questions": ["customer timezone unknown"],
"tools_tried": [{ "name": "tickets.get", "ok": true }],
"evaluator": { "last_verdict": "fail", "failures": ["citation_missing"] },
"budget": { "usd_remaining": 1.2, "revisions_remaining": 2 },
"memory_refs": ["customer_id:cus_9"],
"trace": { "traceparent": "00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01" },
"next_hint": "re-retrieve refund policy; rewrite summary"
}
| Field | Why it exists | Failure if missing |
|---|---|---|
schema | Version the contract | Silent field reuse |
job_id | One story across hops | B invents or asks again |
goal | Bound the job | B renegotiates the ask |
constraints | Policy that survives the hop | Refund promises sneak back in |
artifacts[] | Typed outputs with URIs | “See chat above” |
open_questions | Known unknowns | B hallucinates a timezone |
tools_tried | Burned calls travel | B retries a dead endpoint |
evaluator | Failures travel | B repeats a rejected plan |
budget | One wallet for the run | Each hop “stayed under cap” |
memory_refs | Pointers, not payloads | PII in every hop |
trace | Same timeline | Three traces, no blame |
next_hint | Optional, never authoritative | Hint treated as a new goal |
Rules that do not move:
- Typed artifacts, not “see chat above.”
- Budget remaining travels with the work. A fresh wallet per agent is how fleets overspend.
- Evaluator state travels so the next hop does not blindly retry a failed approach — or so a reviser does retry with evidence.
- No private chain-of-thought required. If B needs reasoning, regenerate it from artifacts and failures.
- Pass refs, not dumps. If B needs raw tool JSON, B re-calls a read tool under its own sandbox.
Anthropic’s long-running harness makes the same bet in another shape: structured artifacts between sessions, plus a sprint contract the generator and evaluator agree on before anyone writes code. The package is the product. The persona name is not.
Which handoff patterns survive a Tuesday outage?
Four patterns. Start with the first. Earn the rest.
| Pattern | What moves | Who advances state | Use when |
|---|---|---|---|
| Relay | A typed package | The workflow rail | Default for business ops |
| Hub | Packages in and out of a thin coordinator | Deterministic code | Fan-in / fan-out with a merge |
| Critique loop | Artifact + failure list | Worker revises; critic never writes | Quality gates |
| Parallel specialists | Disjoint packages + idempotent merge | Rail after both return | Hard partitions only |
Relay. Agent A finishes a stage, emits a package, Agent B starts clean. No shared scratch. This is the default.
Hub. A thin coordinator assigns sub-jobs, collects packages, and decides next states. The hub should be workflow logic, not a free-form gossip model. Anthropic’s research lead is this shape for research: it delegates, waits, synthesizes. Their own write-up notes the cost — multi-agent systems used about 15× more tokens than chats, and some domains (most coding) are a poor fit because the work is not actually parallel.
Critique loop. Worker produces; critic returns failures; worker revises. The critic must not hold write tools. This is still “multi-agent” even when people call the critic a module.
Parallel specialists. Two specialists work disjoint subproblems, then a merge step reconciles. Requires hard partitions and an idempotent merge. Easy to get wrong. Great when it fits.
Microsoft’s handoff pattern is a fifth cousin: one active agent at a time, full control transfers, no parallel work. They tell you to avoid it when routing is rule-based, when a bad hop ruins the customer, or when you cannot stop an infinite bounce. That last one is the ABAB courtesy loop from the single-vs-multi-agent post, now with extra invoices.
Decision list:
- Can you name the artifact type? If no, you do not have a hop.
- Does the next hop need write tools? If yes, it is a writer, not a critic.
- Is the partition hard (no shared fields)? If no, do not run parallel.
- Is the sequence known at intake? If yes, use a rail, not a chairman.
How do vendor handoffs differ from a business contract?
Vendors sell transfer. You still have to own meaning.
OpenAI’s Agents SDK treats a handoff as a tool. A hop to a refund specialist shows up as transfer_to_refund_agent. You can attach input_type so the model must emit a small JSON payload (reason, priority) that the SDK validates locally before on_handoff runs. That is useful metadata. It is not the next agent’s job contract. Their docs say so: input_type does not replace the receiving agent’s main input, and by default the next agent sees the entire conversation history unless you set an input_filter or the opt-in nested-history beta.
That default is the opposite of a business hop. A refund agent that inherits the triage novel will re-argue the goal, re-try burned tools, and spend the rest of the budget restating what A already knew.
LangChain / LangGraph are more honest about the blast radius. Their handoff docs say the term itself was coined by OpenAI for tool-call transfers. When you hop across subgraphs you must pass a valid pair: the AIMessage that called the transfer tool and a ToolMessage that acknowledges it. Skip the pair and the next model sees a broken transcript. Pass the entire subagent conversation and you get “context bloat” — their words — plus a specialist confused by someone else’s scratch.
Google’s ADK does the same transfer as a function call: transfer_to_agent(agent_name=...). The framework swaps InvocationContext. That is control flow. It is not a schema.
AutoGen’s older core guide starts in the right place: define the message protocol first (UserTask, AgentResponse, topic types), then write agents that publish those types. The protocol is the contract. The personas are costumes.
| Vendor primitive | What it actually does | What you still owe |
|---|---|---|
OpenAI handoff() + input_type | Validates hop metadata; transfers control | Package schema, budget, evaluator state |
OpenAI input_filter | Trims history the next agent sees | A real artifact URI, not a shorter novel |
LangGraph Command + ToolMessage | Keeps the transcript well-formed | What content crosses the hop |
ADK transfer_to_agent | Switches the active agent | Role cards and write authority |
| AutoGen topic messages | Pub/sub typed events | Versioned fields and ack |
Treat every SDK hop as a carrier. Put your handoff.v1 JSON on the carrier. If the SDK will not take a payload, write the package to object storage and pass the URI.
- Hop metadata (
reason) is not a substitute forjob_id+ goal - History filters are on; full-transcript default is off
- Guardrails: OpenAI applies input guardrails only to the first agent and output guardrails only to the last — so policy in the middle is your rail
- Side effects are not hidden inside
transfer_to_sales(); ticket creation is a separate, idempotent write
Why shared transcripts lose the job
Lost context is usually a schema problem, not a model problem.
| Failure | Symptom | Fix |
|---|---|---|
| Dropped IDs | B invents or asks again | Required fields; reject the package |
| Stale plan | B follows A’s rejected plan | Evaluator failures travel; ban plan reuse without re-validate |
| Double write | Two agents patch the same CRM field | Single-writer rule; idempotency keys |
| Overshare | PII in every hop | Redact; pass refs, not payloads |
| Undocumented tool use | B retries burned tools | tools_tried with error codes |
| Trace fracture | New run_id per agent | Propagate W3C traceparent |
| Silent reset | Evaluator state wiped at the hop | Invalid if last_verdict is missing after a fail |
| Goal drift | B “improves” the ask | Goal is intake-only; B cannot rewrite it |
LangGraph’s warning is the same physics: do not pass all subagent messages. Summarize the work in the acknowledgement, or point at an artifact. Anthropic persists the plan outside the window because a 200,000-token context still truncates. If the lead’s plan lives only in the transcript, the next wave starts lost.
OpenTelemetry already has the join key. A trace is the path of a request; a span is one unit of work. Across queues you propagate W3C Trace Context (traceparent, tracestate) on the package. The GenAI conventions for invoke_agent spans are still Development-stability as of August 2026 — own an internal span schema and map outward. Do not wait for the SIG to invent job_id for you.
Procedure when context “disappears”:
- Open the hop span. Is
job_idthe same? - Validate the package against
handoff.v1. Which required field is empty? - Check
evaluator.last_verdict. Did B start as if A had passed? - Check
tools_tried. Did B call a tool A already burned? - Check artifact URIs. Did B read the file, or the chat?
If you cannot answer those five without a raw dump, the hop is not observable. Fix the package before you add another agent.
How should the next agent acknowledge the hop?
B emits accepted or rejected_schema before heavy work. Fire-and-forget relays hide poison packages until the money is gone.
| Ack | Meaning | Next action |
|---|---|---|
accepted | Schema valid; role matches; budget > 0 | Start the stage |
rejected_schema | Required field missing or wrong version | Escalate; do not guess |
rejected_role | This agent does not hold the named tools | Hub reroutes or escalates |
rejected_budget | Wallet is empty | Escalate; do not “just finish” |
duplicate | Same job_id + stage already completed | Return the prior artifact; no new write |
Acknowledgements also stitch timelines. The ack span is a child of the hop span. Ops can see “package left A at 14:02, B rejected schema at 14:02” without reading prompts.
LangGraph’s mandatory ToolMessage is the framework-shaped version of this ack. OpenAI’s on_handoff callback is a place to persist the package and emit the ack before the specialist spends tokens. Use them. Do not confuse “the SDK transferred control” with “B agreed the form is valid.”
Checklist:
- Ack is a first-class package field, not a log line
- Reject paths are cheaper than a full specialist run
- Duplicates return the prior artifact
- Humans see the same ack states in the escalate UI
How do you version and validate the contract?
Handoff schemas get versions (handoff.v1, handoff.v2). Consumers declare accepted versions. Breaking changes need a migration. Silent field reuse — “notes used to be free text, now it is JSON” — is how production bleeds.
Validate with JSON Schema (Draft 2020-12) at every hop. Invalid package → escalate, not “improvise.” If a hop cannot be validated in CI, it is not ready for production volume.
| Change | Compatible? | What you do |
|---|---|---|
| Add optional field | Yes | Ship; old consumers ignore it |
| Add required field | No | handoff.v2; dual-write until consumers catch up |
| Rename a field | No | New version; never reuse the old name |
| Change a type | No | New version |
| Tighten an enum | Maybe | Only if production data already complies |
Reuse notes for a new meaning | Never | That is how you get two truths |
Contract tests in CI:
- Fixture packages (pass, fail, duplicate, empty budget) validate or reject as specified.
- Sample worker outputs validate before merge.
- A consumer that declares
handoff.v1still accepts today’s fixtures. - Chaos: drop each optional field; the system fails closed on required ones.
Anthropic’s early research agents spawned 50 subagents for simple queries and duplicated searches because the lead’s instructions were a sentence, not a contract. They fixed it by requiring objective, output format, tools, and boundaries on every spawn. That is versioned interface design with a prompt-shaped skin. Put the same fields in JSON so CI can fail the build.
How do you keep writes from firing twice?
Handoffs retry. Side effects must not.
Temporal’s rule is the one to steal even if you never run Temporal: Activities may execute more than once. A worker can finish the write, crash before reporting, and the platform will retry. Idempotence is “same result whether you run it once or five times.” Their recommended key is Workflow Run ID + Activity ID — stable across retries, unique across runs. The downstream system enforces the key, not the activity.
Map that onto a hop:
| Hop concept | Temporal analog | Rule |
|---|---|---|
job_id | Workflow Id | Stable for the business event |
stage + attempt | Activity Id / attempt | Attempt is not in the external key |
idempotency_key | Run ID + Activity ID | job_id:stage:artifact_type |
| Write tool | Activity | At-least-once; must upsert |
Ack duplicate | Deduped complete | Return the first receipt |
- Single writer for each external resource per run when you can.
- Idempotency key on every customer-visible write: email, CRM patch, refund, ticket comment.
- Do not put
attemptin the external key. Retries must hit the same slot. - Writeback never runs on evaluator fail. The rail owns that gate, not the writer’s mood.
If you cannot draw which agent may write which system, stop adding agents.
How much should agents talk to each other?
Less than vendors imply. Prefer package relay through the workflow rail. Free-form agent-to-agent chat is hard to audit, hard to budget, and easy to poison.
Anthropic is explicit that domains where every agent must share the same context, or where agents must coordinate in real time, are a poor fit today. Their lead waits for a wave of subagents; it does not let them gossip mid-search. That is a feature. Gossip is how two specialists “agree” a refund the policy forbids.
When you need negotiation, constrain it:
| Allowed talk | Forbidden talk |
|---|---|
| Structured proposal JSON, fixed rounds | Open chat until someone “feels done” |
| Evaluator on the merge | Peers voting without a judge |
| Hub collects packages | Specialists writing to each other’s memory |
| One escalate path | Three agents emailing the same manager |
Azure’s cost note is the same math in enterprise clothing: handoff patterns invoke agents one at a time, so spend accumulates across the chain. Concurrent patterns spike. Magentic-style managers iterate until a plan appears — the least predictable bill. If you cannot bound rounds, you do not have a hop. You have a salon.
Business users already understand tickets moving across statuses. Mirror that. Status = state. Assignee = agent role. Attachment = artifact URI. That is agent coordination for business without the swarm vocabulary.
How do you test a hop without a live customer?
Demos lie. Contract tests do not.
| Test | Input | Pass condition |
|---|---|---|
| Schema accept | Valid handoff.v1 | Ack accepted |
| Schema reject | Missing job_id | Ack rejected_schema; no model call |
| Replay | Frozen package P | B’s artifact meets criteria |
| Stale plan | Package with last_verdict: fail | B does not reuse A’s rejected draft |
| Duplicate delivery | Same hop twice | One write; second ack is duplicate |
| Chaos optional | Drop each optional field | Still accepts or fails closed as specified |
| Empty wallet | usd_remaining: 0 | rejected_budget; no write |
| Guardrail gap | Policy only on first agent | Middle hop still cannot send email |
Load shape worth running before a customer sees it:
- Replay 1,000 packages through B, including 10% invalid.
- Duplicate 5% of deliveries; confirm idempotent writes.
- Confirm reject paths spend near-zero tokens.
- Confirm
traceparentis identical on A’s emit and B’s ack.
Anthropic’s eval lesson still applies: start with ~20 real queries before you wait for a 400-case suite. For hops, those twenty should be packages, not prompts. Grade the artifact against criteria. Do not grade whether B “sounded like a specialist.”
Spurlock Studios pilots ($1,500 · 5 days) default to one worker + one evaluator. A librarian is the first extra agent when retrieval is in scope. Fancy mesh topologies wait until the thin path clears the golden set. That is 500+ automations and 20,000+ hours talking: the hop you can replay is the hop you can sell.
What does a human escalation look like as a handoff?
Escalation to a human is a hop. Use the same package schema plus a UI that shows failures and proposed next actions. Do not dump a raw transcript and call it collaboration.
| Human UI field | Source | Why |
|---|---|---|
| Job and customer refs | job_id, memory_refs | They know which ticket |
| What A tried | tools_tried | They do not repeat a burned call |
| Why it failed | evaluator.failures | They see criteria, not vibes |
| What to do next | next_hint + constraints | Bounded action |
| Budget spent | budget | They know if this is a money fire |
| Proposed artifact | Artifact URI | They edit a draft, not a chat |
Azure treats humans as escalation targets in handoff patterns and tells you to persist state at those checkpoints so you do not replay prior work. Same rule: the package is the checkpoint.
One escalation owner. Not three agents emailing the same manager. If escalate queues are deep, intake sheds load or degrades to human-only. Spawning more agents into a full queue is how you buy a bigger outage.
- Same schema as agent hops
- Human can reject schema too (bad package, not bad person)
- Human write uses the same idempotency key family
- Returning from human to agent is another typed hop, not “paste the Slack thread”
Latency, fan-out, and one shared budget
Each extra agent adds model latency and hop overhead. Parallel specialists only pay off when wall-clock matters and merge is cheap. For most internal ops jobs, a serial relay is faster to operate even if slightly slower to run, because traces are linear and blame is obvious.
Anthropic’s BrowseComp analysis is the receipt: token usage alone explained 80% of performance variance; tokens, tool calls, and model choice together explained 95%. Multi-agent architectures scale tokens. They also scale the bill. Their agents used about 4× chat tokens; multi-agent used about 15×. If the job is not valuable enough to pay that, you do not have a multi-agent problem. You have a pricing problem.
| Topology | Latency shape | Cost shape | Operate on a Tuesday? |
|---|---|---|---|
| Single worker + evaluator | One loop | Predictable | Yes |
| Serial relay (2–3 hops) | Sum of hops | Sum of hops | Yes, if packages are small |
| Hub + parallel specialists | Max(specialists) + merge | Sum of specialists | Only with a hard partition |
| Free-form mesh | Unbounded | Unbounded | No |
Budget remaining must decrement across the whole relay. A fresh budget per agent is how fleets overspend while each hop claims compliance. Pass usd_remaining and revisions_remaining on every hop. No exceptions.
Keep the package small. Point to artifact URIs instead of inlining megabytes of tool dumps. Large packages tempt the next model to ignore the middle. If B needs raw JSON, B re-reads under its own sandbox.
Anti-patterns that look like coordination
Shared infinite transcript as bus. Contamination as a service. OpenAI’s default — next agent sees the whole history — is this unless you filter.
Agents that renegotiate the job goal. Goal changes are human or intake-only.
Peer agents with identical tools. Confusion and double writes. If they hold the same tools and the same criteria, they are one stage.
Handoff via screenshots or Slack vibes. Not a system.
Coordinator that is “just an LLM.” Prefer code or workflow logic for routing and budget. Use a model only for soft classification inside intake, with hard policy checks after.
Handoff tool that also writes. transfer_to_sales() must not open a Salesforce opportunity. LangChain’s own guidance: clarify whether the hop only updates routing state or also performs side effects. Split those.
Fresh wallet per hop. See above. This is how you blow the month while every span looks green.
Naming agents after people. retrieve, draft, judge, writeback beat alice, bob, and genius. Verb names clarify tools and keep vanity headcount down.
Infinite bounce. Azure lists avoiding an infinite handoff loop as a reason to skip the pattern. Cap hops. After N transfers, escalate.
If a section of this list would still make sense in a post about image compression, it would not belong here. These failures are hop-shaped. That is the point.
Interface-first design workshop (90 minutes)
Before naming agents, fill a table. This workshop kills vanity topologies faster than any model comparison.
| Stage | Input package | Output artifact | Tools allowed | Writer? | Evaluator criteria |
|---|---|---|---|---|---|
| intake | raw ticket | intake.v1 | read-only ticketing | no | required fields present |
| retrieve | intake.v1 | chunks.v1 or no_hit | search, KB | no | citation keys or no_hit |
| draft | chunks.v1 | summary_md | none | no | cites keys; no refund promises |
| judge | summary_md | verdict + failures | none | no | rubric scores |
| writeback | passing verdict | CRM receipt | CRM patch | yes | idempotent upsert |
Rules for the room:
- If two stages share tools and criteria, they are one stage.
- If a stage cannot name its output artifact type, it is not ready to be an agent.
- Exactly one row may say
Writer? = yesfor a given external system. - The judge row holds no write tools.
- Every output artifact has a schema version before anyone opens a framework quickstart.
Example relay once the table is honest:
- Librarian returns chunks or a
no_hitpackage. - Drafter writes summary JSON with citation keys.
- Evaluator judges criteria; on fail, drafter revises with evidence; on pass, writeback patches an internal field only.
Each hop validates schema. Writeback never runs on fail. Budget decrements along the path. Trace ids stay constant. That is multi-agent handoff patterns without a chat room of agents arguing.
FAQ
What are multi-agent handoff patterns that work?
Relay with typed packages, hub-and-spoke with a deterministic coordinator, worker–critic loops, and carefully partitioned parallel specialists. Start with relay plus critic. Framework transfer_to_* tools are transport for those patterns, not a substitute for a schema. If you cannot name the artifact that crosses the hop, you do not have a pattern yet — you have a demo.
What does agent coordination for business require?
Clear role cards, a rail that owns state transitions, idempotent writes, budget that decrements across hops, and one escalation path. Coordination is systems engineering. Tickets, statuses, and attachments are the vocabulary operators already trust. A chairman model that renegotiates the goal is not coordination; it is another source of drift.
How do you avoid lost context between agents?
Schema-validate a handoff package that carries job IDs, artifact URIs, evaluator failures, tools tried, remaining budget, and a W3C traceparent. Pass references to durable facts instead of dumping the transcript. Anthropic persists plans outside the window for the same reason: the next hop cannot depend on a context that might have been truncated.
When should we add a third agent?
When a second specialty has a crisp interface and the two-agent path already meets pass-rate and cost targets on the same golden set. Specialty without an interface is headcount for models. Read single vs multi-agent before you split; come back here to package the boundary.
Can Spurlock Studios build multi-agent systems?
Yes. Tier 2+ builds split workflows across agents with evaluation harnesses. Pilots stay thin on purpose: worker plus evaluator, librarian only when retrieval is in scope. See /agentic and the operating manual. The hop contract is part of the deliverable, not a later cleanup ticket.
Should the coordinator be an LLM?
Prefer code or workflow logic for routing, schema validation, budget enforcement, and write gates. Use a model coordinator only for soft classification inside intake, with hard policy checks after. Microsoft’s own guidance: if the sequence is knowable from the first input, use a dispatcher, not a handoff swarm.
CTA
Prove one hop pair after a thin path — not a mesh.
What questions does this article answer?
- What are multi-agent handoff patterns that work?
- Relay with typed packages, hub-and-spoke with a deterministic coordinator, worker–critic loops, and carefully partitioned parallel specialists. Start with relay plus critic. Framework `transfer_to_*` tools are transport for those patterns, not a substitute for a schema. If you cannot name the artifact that crosses the hop, you do not have a pattern yet — you have a demo.
- What does agent coordination for business require?
- Clear role cards, a rail that owns state transitions, idempotent writes, budget that decrements across hops, and one escalation path. Coordination is systems engineering. Tickets, statuses, and attachments are the vocabulary operators already trust. A chairman model that renegotiates the goal is not coordination; it is another source of drift.
- How do you avoid lost context between agents?
- Schema-validate a handoff package that carries job IDs, artifact URIs, evaluator failures, tools tried, remaining budget, and a W3C `traceparent`. Pass references to durable facts instead of dumping the transcript. Anthropic persists plans outside the window for the same reason: the next hop cannot depend on a context that might have been truncated.
- When should we add a third agent?
- When a second specialty has a crisp interface and the two-agent path already meets pass-rate and cost targets on the same golden set. Specialty without an interface is headcount for models. Read [single vs multi-agent](/blog/single-vs-multi-agent) before you split; come back here to package the boundary.
- Can Spurlock Studios build multi-agent systems?
- Yes. Tier 2+ builds split workflows across agents with evaluation harnesses. Pilots stay thin on purpose: worker plus evaluator, librarian only when retrieval is in scope. See [/agentic](/agentic) and the [operating manual](/blog/agentic-systems-operating-manual). The hop contract is part of the deliverable, not a later cleanup ticket.
- Should the coordinator be an LLM?
- Prefer code or workflow logic for routing, schema validation, budget enforcement, and write gates. Use a model coordinator only for soft classification inside `intake`, with hard policy checks after. Microsoft’s own guidance: if the sequence is knowable from the first input, use a dispatcher, not a handoff swarm.
Last reviewed
AI Agents
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.