Spurlock Studios
Contact
Share LinkedIn X
An expired brass key. Thesis: MANAGE MULTIPLE AI AGENTS PRODUCTION.

How do you manage multiple AI agents in production? Isolate them. Each agent gets its own credentials, its own write surface, a named human owner, and evals that score that hop and the composed job. Handoffs move a typed contract, not a shared brain. Do not share one god-agent that can read anything and write anything because the deck called it a crew.

This spoke sits under the Agentic Systems Operating Manual. If you have not decided you need more than one loop, stop — when not to build an agent is the cheaper question. If a demo already “has a crew” and production is on fire, the gap is usually the same one why agent demos fail in production names: missing isolation, not missing IQ.

I have spent 20,000+ hours architecting agentic systems and shipped 500+ automations. The expensive multi-agent weeks were almost never “we needed a third persona.” They were a researcher holding a refund key, a critic that could PATCH the CRM, or one eval suite scoring costumes while the customer-facing send was the only tool that mattered.

The short answer

  • Isolate four things per agent: credentials, write surface, owner, evals. If any of those is shared “for convenience,” you still have one agent with extra latency.
  • Handoff with a contract. Typed package in, clean start, ACK or NACK. No shared scratch pad, no inherited chain-of-thought.
  • One writer per entity class. Extra agents may search, draft, or review. They do not get a second wallet, a second SMTP, or a second CRM PATCH.
  • Kill one without killing the fleet. Per-agent flags, budgets, and kill switches. A shared .env is how a researcher outage takes down refunds.
  • Register the fleet. If you cannot list agent_id, tools, owner, eval set, and kill path on one page, you do not have a fleet. You have a pile.
Shared thingWhat you actually bought
One admin API keyOne leak, every write
One write allowlistDual wallets on the same row
One Slack channel “owner”Nobody pages, everyone comments
One “looks good” evalFive 100% hops, one bad send
One transcript for every hopA god-agent with extra hops

A costume is not an operating unit. Isolation is.

What does managing a fleet actually mean?

Managing multiple agents is not adding personas. It is running a small fleet of isolated loops that can fail independently and still produce one customer-correct result.

OpenAI tells teams to define the smallest agent that can own a clear task, then add agents only for separate ownership, tools, or approval policies. Microsoft is blunt about the tax: every extra agent adds protocol design, error handling, state sync, prompt work, monitoring, credentials, and handoff latency. You pay that tax in operations, not in the slide.

Operating objectExists whenFake version
AgentOwn loop, own tools, own stop conditionsA prompt alias in a shared runtime
Credential setSecret store entry this hop alone can readOne .env copied into every container
Write surfaceAllowlist of side-effecting tools“The crew can use Stripe if needed”
OwnerNamed human who can kill it and explain the evalA Slack channel
Eval setFrozen cases for this hop plus a composed job setOne vibe check at the end
Kill pathFlag or secret revocation that stops this hopRestart the whole stack

If you cannot fill that table for two agents, you are not ready to run two. Stay on one loop until you can.

  • Every agent has an agent_id that shows up in traces
  • Every agent has a tool allowlist checked in, not described in Slack
  • Every agent has a human owner with a backup
  • Every agent has a kill switch that does not take down siblings
  • Every customer-facing job has one composed evaluator

A fleet you cannot inventory is not a fleet. It is an incident with a product name.

Why does a shared god-agent fail first?

A god-agent is one identity with every tool, every secret, and every eval. Teams keep it because it is fast to demo. Production hates it because one hop’s mistake is every hop’s blast radius.

Anthropic tells teams to find the simplest solution and add complexity only when it demonstrably improves outcomes. A god-agent is complexity with none of the isolation. Cognition is the write rule: extra agents may add intelligence; they do not get a second write. A god-agent is the opposite — one identity with every write.

God-agent shortcutTuesday failure
One Stripe key on every hopResearcher can refund; intern prompt injection spends money
One CRM token with PATCHLibrarian “fixes” a field it was only supposed to read
One shared vector storeTenant A’s retrieval lands in Tenant B’s draft
One eval named “quality”Tone passes; amount is wrong; send still fires
One memory dumpConstraints from hop A vanish; hop B renegotiates the goal

What it costs: a leaked researcher identity that can also send, charge, and overwrite. You spend the week proving which costume did it, because every costume had the same key.

What you do instead: least privilege per hop, one writer, a contract between hops, and evals that cannot be satisfied by looking busy.

Root on the researcher is not a shortcut. It is the incident.

How do you isolate credentials per agent?

Credentials are the isolation you can prove. Prompt text is not.

Give each agent an identity in the secret store. That identity may hold only the keys for the tools on its allowlist. Google Cloud IAM calls this least privilege: grant only the permissions required to perform the intended function. Stripe restricted keys exist so a process can charge without also being able to change the account. Your researcher should not hold the charge key. Your actor should not hold the warehouse admin key.

AgentMay holdMust never hold
Librarian / retrieverRead credentials for the allowed indexesWallet, SMTP, CRM write, admin IAM
DrafterRead tools + draft storeSend, refund, production PATCH
Actor / writerOne write class (CRM or email or refunds)Broad read over raw PII plus a second write class
EvaluatorScore artifacts; maybe read-only toolsAny side-effecting tool
Supervisor / routerRoute metadata onlyAny write, any secret the workers hold

Procedure:

  1. List every secret the fleet currently shares. Treat a copied .env as one identity, even if the filenames differ.
  2. Map each secret to the tools that need it. If two hops need the same write secret, you still have one writer — collapse them.
  3. Mint per-agent identities. Rotate the shared key. Do not “add” isolation on top of a live admin key.
  4. Prove the negative: run the researcher against the refund tool and expect deny, not a model apology.
  5. Put rotation and revocation on the owner’s checklist. An identity nobody can rotate is not isolated. It is abandoned.
  • No production secret is in git, a shared Notion page, or a demo laptop
  • Revoking the librarian identity cannot stop the actor mid-send
  • Staging secrets cannot mint production writes
  • Tool allowlists and IAM allowlists match. A prompt “do not refund” is not a credential boundary

Prove isolation with deny tests, not with a policy paragraph in the prompt:

TestExpectedFail tell
Librarian calls refund toolPolicy deny before the model “decides”Model says “I shouldn’t” after the HTTP 200
Drafter loads actor identitySecret-store missSame STRIPE_KEY in both containers
Actor uses staging key in prodHard failCharge lands in the live account
Revoke librarian, run actorActor still writesShared key; both hops die

A shared key is a shared blast radius.

How do you isolate write surfaces?

A write surface is every tool that changes the world: CRM PATCH, Stripe charge, SMTP, ticket status, inventory decrement, calendar insert. Isolation means one hop is allowed to call that class for a given entity.

Cognition’s production note is the same sentence I use on pilots: extra agents can review, search, or advise; writes stay single-threaded. RFC 9110 If-Match / ETag is compare-and-set for HTTP. Stripe stores the first result for an Idempotency-Key so a retry is a replay, not a second charge. Those primitives only help if one identity is the writer.

SurfaceOne writer meansShared-surface tell
CRM account rowOne agent_id may PATCH account:*Librarian and actor both PATCH “notes”
Refund / chargeActor only, HITL until the golden set is greenResearcher “confirms” by calling Stripe
Customer emailSender only; drafter emits an artifact URICritic has SMTP “to save a hop”
Ticket statusActor with observed_versionTwo hops comment and both set status
CalendarScheduler agent or a deterministic workflowEvery persona can insert events

Decision list — keep the write on one hop if:

  1. Two hops would PATCH the same entity class.
  2. The second hop only exists to “be careful.” Put care in the evaluator and the policy gate.
  3. You cannot name a compare-and-set field or an idempotency key for the tool.
  4. You cannot NACK a write without paging a human who is not in the loop.
  • Tool registry lists read / draft / write_reversible / write_irreversible per tool
  • Irreversible tools sit on one agent or behind HITL
  • Idempotency keys are stable for the job, not minted per hop
  • Handoff packages carry observed_version for any row the next hop might write

If you already have two writers, drain before you invent a lock:

  1. Freeze new jobs that hit the contested surface.
  2. Pick one agent_id as the writer. Strip the tool from the other hop in config, not in prose.
  3. Replay in-flight jobs with the same idempotency key and If-Match on the version they observed.
  4. Re-run composed cases that include duplicate retries and stale versions.
  5. Only then turn the second hop back on as read or draft.

Two writers on one row is not collaboration. It is a race with a product name.

Who owns each agent on a Tuesday?

An owner is the human who can kill the agent, rotate its keys, explain its eval, and take the page at 2 a.m. A Slack channel is a comment thread. It is not an owner.

Microsoft’s tax list includes monitoring and credentials for a reason: extra agents multiply on-call. If you cannot name one person per agent_id, the fleet will page “platform” while four costumes blame each other.

RoleOwnsDoes not own
Agent ownerKill switch, keys, eval set, prompt/allowlist releaseThe customer relationship
Job ownerComposed result, customer-facing SLO, escalate pathEvery hop’s prompt wording
Security / IAMIdentity minting, rotation policy, deny testsPrompt copy
On-call (job)The run that failed, even if two processes ran“Their agent,” as a way to bounce the page

Procedure for naming owners:

  1. Write agent_id → owner → backup on one page. No “shared” cells.
  2. Tie the pager to the job, not the costume. Hop B failing still pages the job owner.
  3. Require the owner to run the deny test after every allowlist change.
  4. Retire an agent by revoking its identity and deleting its write tools. Leaving a zombie prompt with a live key is how god-agents return.
QuestionIf you cannot answer
Who can rotate this key tonight?Do not ship the hop
Who can flip the kill flag without a deploy?Do not ship the hop
Who changes the eval when the rubric was wrong?You have a costume, not a product
Who says the composed job is done?You will ship on vibe

Page routing that actually isolates blame:

PageGoes toMust not bounce to
Composed job failedJob owner“Whoever wrote the last prompt”
Deny test red after allowlist changeAgent ownerSecurity as a comment thread
Key leak / unexpected Stripe callAgent owner + IAMThe model vendor
Handoff NACK stormJob owner, with from_agent in the payloadA new supervisor agent

A costume is not an owner.

How do you eval each agent without a shared scorecard lie?

Each agent gets evals for its job. The fleet gets a composed eval for the customer result. Scoring only the costumes is how five hops pass and the customer still gets the wrong refund.

Anthropic did not declare multi-agent a win because the subagents “looked busy.” They measured a research eval, then published the cost: agents use about 4× the tokens of chat, multi-agent about 15×, and token spend explained 80% of BrowseComp variance. That is a research receipt, not a license to skip composed scoring on a refund workflow.

HopIts eval asksMust not count as a pass
LibrarianPack complete, redacted, citations resolvable“Sources exist” with raw PII in the pack
DrafterConstraints honored, artifact schema validTone only
ActorCorrect entity, correct amount, idempotent writeSMTP 250 or HTTP 200
EvaluatorRubric hits on the jobAgreeing with the drafter’s vibe
Composed jobCustomer-correct end stateAny hop-level 100%

Procedure:

  1. Freeze N golden cases for the job (the customer-correct end state).
  2. Freeze a smaller set per hop: librarian pack quality, drafter constraint violations, actor deny tests.
  3. A hop may ship only if its set is green and it cannot tank the composed set.
  4. Track hop pass rate, composed pass rate, cost per passing job, escalate rate, and handoff NACK rate.
  5. If hop pass rates rise while composed pass rate is flat, you are scoring theater. Fix the composed rubric before you add a hop.
  • Deny tests exist: researcher vs refund tool, evaluator vs SMTP
  • Composed cases include the ugly ones: stale version, duplicate retry, missing constraint
  • Cost is per passing job, not per hop that “stayed under cap”
  • Handoff NACK spam is a failing metric, not a sign the critic is thorough

Do not share one golden set across hops “to save labeling.” A librarian case is an evidence pack. An actor case is a write. Mixing them is how you ship a beautiful pack and a wrong charge.

Shared scorecardWhat it hides
One “quality” rubric for every hopAmount errors scored as tone
Same N cases, no hop-specific deniesResearcher never sees a refund tool in eval
Eval owned by the framework vendorNobody updates the rubric when policy changes
Only happy-path ticketsDuplicate retries and stale versions never fail CI

Costume pass rates do not count.

How do handoffs with contracts keep isolation honest?

Isolation dies the moment hop B inherits hop A’s tools, secrets, or novel. The contract is how you keep the cut.

A handoff is: A finishes a stage, emits a schema-valid package, B starts clean. Fail closed on an invalid package — escalate, do not let B invent missing IDs. Anthropic’s research lead writes a plan, then spawns subagents with a self-contained task: objective, output format, tools, and a stop condition. That is a contract. It is not a group chat.

Minimum fields the operating page should name even if the schema lives elsewhere:

FieldIsolation job
job_id + traceparentOne story across hops
from_agent / to_agentExplicit identities, not “next”
goal + constraintsB cannot renegotiate policy
artifacts[] with URIsNo “see chat above”
tools_allowed for BB does not inherit A’s allowlist
budget_remainingOne wallet for the run
evaluator failuresB does not retry a rejected plan blindly
observed_versionStale writes die with 412, not last-write-wins

Rules that do not move:

  1. B’s runtime loads B’s credentials. The package never contains secrets.
  2. B’s tool allowlist is B’s. A “next_hint” is not a new grant.
  3. Invalid package → escalate. Missing job_id is not a creative writing prompt.
  4. The writer of the irreversible tool is named in the contract. If two hops can write, the contract is a lie.
  5. Humans escalate as a handoff too: same fields, plus why the machine stopped.

Microsoft’s architecture guide prefers deterministic routing when the next hop is knowable from the input. A chairman agent that “decides who talks” while holding every key is a god-agent with a gavel.

The contract is the hop. The transcript is gossip.

How do you register the fleet so you can count it?

If the agents only exist in a framework graph, you will lose one. Production needs a registry a skeptic can read without opening the repo.

ColumnWhy it exists
agent_idStable name in traces, logs, and pages
Job in one sentencePrevents persona drift
Tool allowlist (hash or version)Diffable; “usually CRM” is not a list
Credential identitySecret-store path, not a password
Write surfacesEmpty for read-only hops
Owner + backupTuesday coverage
Eval suite idWhich golden set gates a release
Kill flag / revocation pathIndependent stop
SLO classInteractive vs batch; different budgets
Statuspilot / prod / retired

Procedure to stand the registry up in a week:

  1. Inventory every loop that can call a tool. If it can call a tool, it is an agent or it is a function. Functions do not get rows.
  2. Fill the table. Empty write-surface cells are a feature for librarians and evaluators.
  3. Delete rows that are prompt aliases. FormatDateAgent is a function. Call format_date.
  4. Put the table in the same release train as the allowlists. A wiki that lags prod is how god-agents hide.
  5. Review retired rows monthly. Live keys on retired is the classic leak.
  • Count of registry rows equals count of identities in the secret store
  • A new write tool requires a registry diff, not a prompt edit
  • Demos cannot mint a row without an owner
  • The composed job names which agent_ids it may call

Example rows for a three-hop support job (roles, not a client):

agent_idWritesOwner
support-librariannoneLibrarian owner + backup
support-drafterdrafts:* onlySame owner as librarian, no write key
support-actortickets.patch + SMTPJob owner + backup
support-evalnoneJob owner — rubric, not tools

If you cannot list the agents, you do not have multi-agent. You have a pile.

How do you budget tokens, latency, and blast radius?

Each hop will spend “a reasonable amount.” Summed, that is how fleets blow the wallet and the SLO. The run carries one budget. Isolation does not mean a fresh wallet per costume.

Anthropic’s economics are the warning label: 4× tokens for an agent versus chat, 15× for multi-agent, and a research win that only pays when the task value covers the spend. Token usage explained 80% of their BrowseComp variance. A hop that looks cheap while the job is expensive is not a well-managed hop.

BudgetWho holds itIsolation rule
USD / tokens for the jobTravels in the handoffNo hop mints a new cap
Revisions remainingJob ownerCritic loops cannot infinite-courtesy
p95 latencySLO class on the registryBatch hops cannot steal the interactive budget
Blast radiusWrite surfaces + credential scopeLibrarian timeout must not unlock actor writes
Fan-outHard cap on sibling workersResearch-style fan-out is opt-in, not default

Decision list:

  1. Interactive jobs get a hard deadline. Crossing it escalates; it does not spawn three more workers.
  2. Batch jobs get a wall clock and a kill. “Still thinking” is not a status.
  3. Fan-out is for breadth-first work you measured. A three-tool CRM note is not BrowseComp.
  4. Cost per passing job is the number you review, not cost per hop.
  5. If adding a hop raises cost and escalate rate while composed pass rate is flat, roll the hop back. You bought coordination debt.
  • Handoff includes usd_remaining and revisions_remaining
  • Dashboards can slice spend by agent_id and by job_id
  • A hop that hits $0 remaining cannot call paid tools
  • Staging has a cheaper model and a blocked write path — cheaper is not “prod writes with a discount model”

A fresh wallet per agent is how fleets overspend while every hop “stayed under cap.”

Failure mode: one god-agent, three jobs, one leak

What breaks: Support-triage, refund-draft, and internal-research “agents” are three prompts on one runtime. They share a Stripe key, a CRM PAT, and a help-center index. A prompt-injected ticket asks the researcher to “verify the refund by processing it.” The researcher has the key. The charge fires. Traces show agent=support-bot because that is the only identity.

What it costs: Money out the door, a CRM row that no longer matches finance, and a week of blame aimed at “the model.” The model did what a god-agent is built to do: it used the tools it was given.

What you do instead: Split identities before you split prompts. Librarian reads. Drafter emits an artifact. Actor holds Stripe behind HITL until deny tests and the composed golden set are green. One owner per identity. One eval per hop plus composed.

CostumeTools it actually hadWhat isolation required
ResearcherStripe + CRM PATCH + searchSearch / index read only
DrafterSame as researcherArtifact write only
“Refund agent”Same key, different promptStripe restricted key, HITL, HITL owner
EvaluatorNone — it was a comment in SlackRubric on amount, entity, idempotency

I have watched this shape on production automations: the prompts were different, the key was not, and finance found it before the eval did. Isolation is the product. The persona names were decoration.

Do not invent a rate for how often this happens. There is no honest studio-wide “X% of multi-agent leaks” number I will put on this page. The failure is mechanical: if the researcher identity can charge, it will eventually charge.

Three prompts, one key, one blast radius. That is not a fleet.

How do you kill, pause, and rollback one agent without taking the fleet down?

A fleet you can only stop by killing the cluster is not isolated. You need a per-agent stop that leaves siblings running.

ControlStopsMust not stop
Kill flag on agent_idNew turns for that hop; in-flight tool calls after the gateOther agents’ in-flight writes
Credential revocationFuture API calls for that identityIdentities you did not rotate
Traffic split / flagRoutes the job back to a single loop or a workflowLeaves two writers racing during the flip
Version pinRolls the prompt/allowlist to last greenShared config that also pins the actor
HITL force-onIrreversible tools require a humanRead tools the librarian still needs

Procedure:

  1. Ship every agent behind a flag keyed by agent_id. Default off in a new environment.
  2. On incident: flip the flag, then revoke if you suspect key theft. Flag-only is for bad prompts. Revocation is for bad identities.
  3. Drain: let in-flight jobs finish or NACK with escalate. Do not dual-write during drain.
  4. Rollback the allowlist/prompt version. Re-run deny tests and the composed golden set before the flag goes on again.
  5. Record the composed pass rate, cost per pass, and escalate rate on the same N cases. If the hop was the problem, the single-loop path should recover the job.
  • You can name the flag in one sentence: “turn off refund-actor”
  • Revoking the librarian cannot cancel a sender’s in-flight SMTP after the policy gate passed — or you accept that and document it
  • Rollback is a config change, not a three-hour rebuild
  • Post-incident, the registry status moves to paused with an owner note

If you cannot turn off one agent, you never isolated it.

How do you debug a hop without a murder mystery?

Process boundaries multiply work you already under-budgeted. Isolation without a shared story is how you get three traces and no blame.

OpenTelemetry defaults to W3C traceparent so a backend can stitch one trace from two processes. If hop B mints a new run_id and drops traceparent, you did not gain modularity. You gained a murder mystery. Anthropic said the quiet part in the research post: without production tracing they could not tell whether “not finding obvious information” was a bad query, a bad source, or a tool failure. Multi-agent makes that worse because the failure lives in the interaction.

SignalIsolated-and-debuggableMurder mystery
TraceOne trace_id trigger → terminal writeNew run_id per hop
Identityagent_id on every span“support-bot” for every costume
Tool callsAllowlist + deny reason in the spanModel apology, no policy log
HandoffPackage id + schema version on the span“see chat above” in a blob
ReplayOne command, both hops, same job_idNeeds both queues warm and a prayer

Checklist before you call the fleet production:

  • Same trace_id from trigger to terminal write
  • Spans name agent_id, tool, and policy decision (allow / deny / escalate)
  • Replay works with one command on a frozen fixture
  • One job owner pages, even if two processes ran
  • Credential on hop B cannot do hop A’s reads — and that deny is visible in the trace

Budget real engineering time per boundary. If the agentic pilot is five days, that cost is a real fraction of the calendar — which is why isolation work comes before the third persona.

If you cannot replay the hop, you cannot manage the hop.

What operating checklist ships before a second writer?

Do not wait for a pretty registry. Ship the isolation that makes a second writer survivable. If this list is empty, you are still on a god-agent, even if the slide has two stick figures.

  1. Freeze the composed golden set and the cost band on the current path (single loop or current hop).
  2. Name the second writer’s only write surface. If you cannot name one noun, you do not get a second writer.
  3. Mint a new identity. Deny-test the old identity against that surface and the new identity against everything else.
  4. Write the handoff contract fields before the prompt. Invalid package → escalate.
  5. Put traceparent, job_id, and agent_id on every span.
  6. Name owner + backup. Wire the job pager.
  7. Feature-flag the hop. Re-run composed cases. Keep the hop only if composed pass rate holds or rises and cost per pass stays in band.
  8. Write the rollback sentence: “flip agent_id off; traffic returns to ____.”
DayDone when
1Registry rows + identity map; shared keys listed
2Writer allowlist is one surface; deny tests red on the wrong hop
3Handoff schema validating in staging; secrets not in the package
4Traces stitch; replay command exists
5Composed golden set compared; flag off-switch practiced

Skip this week and you will “manage” the fleet by reading Slack. That is not management. That is narration.

If a workflow would still do the job, do not spend the week on a second writer. Read when not to build an agent and keep the loop small.

FAQ

How do I manage multiple AI agents in production?

Isolate each agent: credentials, write surface, named owner, and evals. Handoff with a typed contract so hop B does not inherit hop A’s tools, secrets, or transcript. Keep one writer per entity class. If those four isolations are shared, you still have one god-agent, no matter how many prompt files you added.

How do I measure whether managing multiple AI agents in production is working?

Composed pass rate on a frozen golden set, cost per passing job, escalate rate, handoff NACK rate, and whether you can kill one agent_id without taking siblings down. Hop-level “looks good” scores are not a fleet metric. If hop pass rates rise while composed pass rate is flat, you are scoring costumes.

What usually fails first when teams try this?

Shared credentials and dual write surfaces. The researcher can refund, the critic can PATCH, and traces cannot tell the costumes apart because they share one identity. Prompt boundaries fail after that. Fix identities and writers before you hire another persona.

How long does this take to show results?

You should see operational control as soon as one hop can be killed, denied, and scored without taking the fleet down — often inside a focused week if the inventory is honest. Do not wait on a revenue number to prove isolation. I will not invent a studio-wide days-to-ROI figure; the early result is a deny test that holds and a composed set that still passes after the split.

What should I skip if I only have a week?

Skip the third agent, the LLM supervisor, shared memory, and fan-out. Do the registry, split the write key off the researcher, name owners, wire traceparent, and put a kill flag on each agent_id. A week of isolation beats a week of new personas.

When is this not worth doing yet?

When you still have one honest loop, or when a workflow or a single LLM step would do. Isolation tax is real — credentials, evals, tracing, on-call. Pay it after a documented split, not to make a demo look like a department. If the job should not be an agent, do not manage a fleet of them.

CTA

Isolate the fleet. Then add a writer only when the contract and the owner exist.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

How do I manage multiple AI agents in production?
Isolate each agent: credentials, write surface, named owner, and evals. Handoff with a typed contract so hop B does not inherit hop A’s tools, secrets, or transcript. Keep one writer per entity class. If those four isolations are shared, you still have one god-agent, no matter how many prompt files you added.
How do I measure whether managing multiple AI agents in production is working?
Composed pass rate on a frozen golden set, cost per passing *job*, escalate rate, handoff NACK rate, and whether you can kill one `agent_id` without taking siblings down. Hop-level “looks good” scores are not a fleet metric. If hop pass rates rise while composed pass rate is flat, you are scoring costumes.
What usually fails first when teams try this?
Shared credentials and dual write surfaces. The researcher can refund, the critic can PATCH, and traces cannot tell the costumes apart because they share one identity. Prompt boundaries fail after that. Fix identities and writers before you hire another persona.
How long does this take to show results?
You should see operational control as soon as one hop can be killed, denied, and scored without taking the fleet down — often inside a focused week if the inventory is honest. Do not wait on a revenue number to prove isolation. I will not invent a studio-wide days-to-ROI figure; the early result is a deny test that holds and a composed set that still passes after the split.
What should I skip if I only have a week?
Skip the third agent, the LLM supervisor, shared memory, and fan-out. Do the registry, split the write key off the researcher, name owners, wire `traceparent`, and put a kill flag on each `agent_id`. A week of isolation beats a week of new personas.
When is this not worth doing yet?
When you still have one honest loop, or when a workflow or a single LLM step would do. Isolation tax is real — credentials, evals, tracing, on-call. Pay it after a documented split, not to make a demo look like a department. If the job should not be an agent, do not manage a fleet of them.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot