Spurlock Studios
Contact
Share LinkedIn X
Nested brass frames. Thesis: AGENT DEMOS DIE PRODUCTION CONTROL.

Your AI agent works in the demo but fails in production because the demo optimized for a staged happy path, not a control loop. The model did not get dumber overnight. Production introduced real schemas, real tenants, real tool failures, and real side effects — and your harness never owned those constraints.

This spoke sits under the Agentic Systems Operating Manual. If the job is still a known path, stop here and read when not to build an agent. The rest of this page assumes you already earned the loop and now need it to survive contact with live tools.

The short answer

  • Demos prove “can it look smart once?” Production asks “can it stay correct under drift, load, and bad tools?”
  • Most “hallucinations” in launch week are system bugs: schema drift, stale tool results, authorization bleed, cascade after one bad tool call.
  • Readiness is a checklist on the harness — evaluators, sandboxes, reason codes, kill switches — not a higher model tier.
  • Soft-launch without a reliability audit is how you buy expensive chaos.
  • Fix the control loop first; then argue about prompts.

What does a demo prove that production does not?

A demo is a theater set. Inputs are curated. Tools return clean JSON. Credentials are a single sandbox tenant. Latency is low. Nobody else is writing to the CRM while the agent runs. The audience watches one path succeed.

Production is adversarial by accident. Vendor fields rename. Tokens expire mid-run. Two companies share a name. A write tool returns 200 with an empty body. Overnight, there is no human in the room to say “that note went on the wrong account.”

Demo assumptionProduction reality
Fixed tool schemasVendors ship breaking field renames
One tenant, one roleMulti-tenant auth with bleed risk
Tools always succeedPartial failures, timeouts, empty results
Single operator watchingOvernight runs, no human in the room
“Looks right” is enoughEvaluator criteria or a customer complaint
Shared sandbox keyPer-tenant tokens, rotation, expiry

Fiddler’s production write-up makes the same split without dressing it up: demo and test environments use curated data and controlled conditions; production adds rate limits, authentication expiry, concurrent sessions, and edge cases the demo never saw. See Why do 70-95% of AI agent projects fail in production?. Treat the headline range as a vendor estimate as of 2026, not a law. The useful claim is the gap, not the percent.

If your success metric was applause, you measured the wrong thing.

Why is the control loop the bottleneck, not model IQ?

When the demo dies, the instinct is to swap models or rewrite the system prompt. That treats intelligence as the bottleneck. For business agents, the bottleneck is almost always the control loop: plan → act → evaluate → revise → terminate, with deterministic guards around non-deterministic steps.

I have spent 20,000+ hours architecting agentic systems and built 500+ automations. The runs that survived launch week all had the same boring furniture: typed tool contracts, tenant-bound credentials, an evaluator that can refuse “done,” and a write freeze that does not wait on a prompt deploy.

Name the gap before you touch the model:

  1. No evaluator — terminal success is “HTTP 200” or “agent said done.”
  2. No side-effect classes — write tools run with the same trust as read tools.
  3. No schema contracts — tool args and results are free-form blobs.
  4. No tenant binding — run context does not pin authorization.
  5. No cascade brake — one bad tool result feeds the next plan as truth.
  6. No kill switch — stopping writes means shipping a new prompt.

AWS’s own agent-building series is blunt about the first of those: poorly defined tool schemas are a leading source of agent failures, because the schema is the contract between reasoning and action. That line lives on AI agent frameworks and building blocks. Upgrade the harness. Then, if quality is still soft, change the model. Blame order matters.

Which four controls close the demo-to-prod gap?

Four objects. If any one is missing, you are still demoing.

ControlWhat it ownsDemo smellProduction bar
SchemasTool args and results“The model usually fills this in”Versioned JSON Schema, fail closed
AuthWho the run may act asOne long-lived API keyTenant + role bound per call
EvalsWhether “done” is legalA human nodded in the roomCriteria the runner can enforce
Kill switchStopping writes nowRestart the process / redeployFreeze writes without a prompt change

NIST’s AI Risk Management Framework is the governance vocabulary for this, not a product pitch. The Core is GOVERN, MAP, MEASURE, MANAGE. A demo that never MEASURES tool-shape errors or MANAGES a freeze path has not left GOVERN-as-slide-deck.

  • Every tool has a pinned input schema and a pinned output schema
  • Every tool call carries tenant_id and role from the runner, not from the model
  • “Done” is an evaluator verdict, not a sentence the worker typed
  • Writes can be frozen without touching prompts

Four boxes. Empty any one and the rest become theater.

How do unversioned tool schemas invent fields?

The tool once returned customer.email. Three months later the API returns contact.primaryEmail. The agent invents an email field that “should” exist. That is not a creative model. That is an unversioned contract.

OpenAI’s function-calling guide tells you to define tools with JSON Schema and to turn on strict so argument objects adhere to the schema instead of “best effort.” See Function calling. Anthropic’s tool-use docs make the same demand on the other side of the fence: every custom tool needs an input_schema object. See Implement tool use. Neither vendor will version your CRM adapter for you.

JSON Schema itself is the vocabulary. The current dialect is 2020-12; declare it. The spec index is JSON Schema — Specification.

LayerPin thisFail closed when
Model → tool argsVendor strict / input_schemaExtra keys, missing required, wrong types
Tool → adapterYour versioned output schemaUnknown fields, null where required
Adapter → runnerReason code + spanParse error rate spikes
Runner → writeIdempotency key + tenantSchema miss on a write tool

What breaks: CRM writes with null emails, silent skips, or fabricated values.

What you do: Pin schemas, version adapters with dates, refuse unknown shapes, alert on adapter-error rate — not on “the model sounded unsure.”

  1. Snapshot today’s live payload for each tool. That is the contract, not the vendor’s marketing docs.
  2. Wrap it in a schema. Mark required fields. Ban extras.
  3. Store schema_version on every span.
  4. On mismatch: no write, schema_mismatch reason code, page the adapter owner.

A prompt that says “be careful with fields” is not a contract.

How does authorization bleed across tenants?

Demo used one API key. Production shares a worker pool. A run for Tenant A accidentally carries Tenant B’s token, or a tool accepts an id without checking ownership. The model did not “decide” to leak. The harness never enforced tenancy.

This is an incident, not a quality ticket.

OWASP’s Top 10 for LLM Applications names the shape: LLM06 Excessive Agency is damaging action from unexpected, ambiguous, or manipulated model output, usually because of excessive functionality, excessive permissions, or excessive autonomy. A shared god-key is all three.

Microsoft’s Zero Trust note on agents is the same rule in identity language: define identity, scope, tool access, and auditability before you widen autonomy, and test revocation — disable the agent, rotate credentials, invalidate tokens. See Least privilege for AI agents. Foundry’s agent-identity write-up adds the implementation: provision an agent identity, assign RBAC to that identity, and exchange a token at tool time instead of stuffing secrets into prompts. See Agent identity. You do not need Azure to steal the pattern.

BindingWho sets itWho must not set it
tenant_idRunner, from the job ticketModel, tool args, retrieved text
Role / scopePolicy table for the job type“The agent asked for admin”
CredentialSecret store / token exchangePrompt, memory, tool description
Resource idVerified against tenant ownershipTrust the model’s account_id

What breaks: Cross-customer reads or writes.

What you do:

  • Bind tenant_id at the harness for every tool call
  • Refuse tools that ignore the binding
  • Execute writes in the caller’s authorization scope, not a shared god-key
  • Test bleed on purpose in staging (see drills below)

Authorization is not a system-prompt paragraph. It is a check the model cannot skip.

What happens when one bad tool result cascades?

First tool returns an empty list or a wrong match. The agent treats that as ground truth, invents a narrative, and writes it downstream. Later steps look like hallucination. The root cause was trusting a bad observation without a verify step.

StepObservationIf you trust itIf you brake
1Enrichment returns []Agent “recalls” a companyEscalate: empty_enrichment
2Two CRM ids share a nameWrites the first hitGate: candidates.length == 1
3Write tool 500s once“Done” on the retry storyBounded retry or escalate
4Read-after-write missesInvents a confirmationVerify tool or fail

What breaks: Plausible wrong CRM notes, tickets, or emails.

What you do: Require verification tools for high-stakes entities. Evaluator criteria must reject “asserted without evidence.” Escalate when criteria fail — not when the model feels unsure.

Genuine model error exists. Lead with harness bugs first. You will be right more often.

Why do evaluators have to own “done”?

Because the worker is incentivized to finish. Left alone, it will narrate success after a partial write, a skipped verify, or a tool that returned 200 with garbage.

LangSmith’s agent-eval docs split the job into three scores you can actually implement: final response, single step, and trajectory — whether the path of tool calls was the path you meant. Start at Evaluate a complex agent and trajectory evaluations. You do not have to buy LangSmith. You do have to score more than the last sentence.

Eval typeQuestion it answersDemo substitute
FinalDid the terminal artifact meet criteria?“Looks good in the chat”
StepWas this tool the right one, with legal args?None
TrajectoryDid the path stay inside the allowlist?Happy-path recording
PolicyDid any call violate tenant, budget, or write class?Hope

Rules that belong in code, not in a rubric paragraph:

  1. The worker cannot mark done. Only the evaluator can.
  2. eval_pass requires evidence pointers (ids, etags, screenshots of the write), not vibes.
  3. eval_fail with revisions remaining goes to revise. Ceiling hit goes to escalate.
  4. Policy fail goes to abort or freeze-writes. It does not get another try.

If “done” is a string the model is allowed to emit, you do not have an evaluator. You have a narrator.

What does a kill switch actually freeze?

A kill switch is a runtime flag that stops side effects without a deploy. It is not “we will revert the prompt.” It is not “restart the worker.” It is a gate in front of every write tool that a human can flip in one place.

Amazon’s AgentOps write-up on Bedrock AgentCore treats evaluation, observability, and governance as the production pillars — not a prettier planner. See AgentOps: operationalize agentic AI at scale. A freeze flag is the smallest governance object that still matters at 2 a.m.

SwitchScopeWhen you flip it
freeze_writesAll write tools, all tenantsUnknown blast radius
freeze_writes:job_typeOne jobOne flow is poisoning CRM
freeze_writes:tenantOne customerSuspected bleed or bad data
freeze_modelNew model callsCost spike or provider outage
shadow_onlyPropose, do not executeYou are still scoring

Requirements:

  • Flip does not require a prompt change or a container rebuild
  • In-flight runs see the flag before the next write, not after
  • Reads may continue so you can diagnose
  • Every blocked call gets killed_by_switch plus the switch name
  • On-call knows who is allowed to flip it

A kill switch you cannot find in the runbook is a slide.

How do launch-week “hallucinations” map to harness bugs?

Treat these as system bugs until proven otherwise.

Schema drift

Unversioned adapters plus invented fields. Covered above. Reason code: schema_mismatch.

Stale tool results

A cache or previous span result is reused after the underlying record changed. The agent plans from yesterday’s pipeline stage and books the wrong follow-up.

What breaks: Wrong next actions that look internally consistent.

What you do: TTL on tool results, etags or updated_at checks before writes, span metadata that marks result_stale=true.

Authorization bleed

Shared keys, missing tenant bind, resource ids the model supplied. Reason code: tool_auth_error. Incident process, not a prompt tweak.

Cascade after one bad observation

Empty or wrong tool result treated as ground truth. Reason code: unverified_assertion or empty_enrichment.

Intake garbage treated as narrative

Empty strings, HTML in “plain text,” dual-language names, archived ids that still look valid. Reason code: intake_schema_fail.

Label you will hearCheck this firstOnly then
“It hallucinated an email”Output schema + adapter versionPrompt / model
“It wrote the wrong account”Disambiguation gate + tenant bindPrompt / model
“It ignored the policy”Whether policy lives outside the modelPrompt / model
“It looped all night”Step ceiling + kill switchPrompt / model
“It leaked another customer”Credential binding — incidentDo not “prompt harder”

If your postmortem starts at the system prompt, you skipped the table.

Illustrative walkthrough: demo green, prod red

Illustrative — not a client result. Staging demo: agent looks up a lead, enriches firmographics, writes a CRM note. Tools are stubbed to always return the same Acme Corp payload. Soft-launch: real CRM has duplicate company names; enrichment returns two candidates; the agent picks the wrong one and writes a note on the wrong account.

Misdiagnosis: “the model hallucinated the company.”

Actual failure mode: no disambiguation gate, no evaluator check that crm_account_id matched the enrichment candidate ids, no escalate path for multi-match.

The fix is a control: if candidates.length != 1 → escalate, plus a golden case for duplicate company names. The model upgrade is optional.

Walk the same incident through the four controls:

ControlWhat was missingWhat you add
SchemaEnrichment result was an untyped listcandidates: array, minItems / maxItems rules
AuthWrite used a shared CRM keyTenant-scoped token; id must belong to tenant
Eval“Wrote a note” counted as donecrm_account_id ∈ candidate_ids
Kill switchWrong-account writes kept flowingfreeze_writes:crm_note until the gate ships

That is a reliability audit in one incident. Do it in staging on purpose, not in week one of soft-launch by accident.

How do you run a reliability audit before soft-launch?

Run this before the agent can write in production. Timebox it. Do not wait for a perfect platform.

  • Inventory every tool: read / write / irreversible; name the side-effect class
  • Pin and version each tool schema; record adapters with dates
  • Bind tenant + role to every tool invocation in the harness
  • Require evaluator criteria for the job type; worker cannot self-certify
  • Define terminal reason codes (eval_pass, tool_auth_error, policy_violation, killed_by_switch)
  • Kill switch that freezes writes without redeploying prompts
  • Golden set with at least one case per known failure mode above
  • Staging soak with real schemas (not stubs) for 48 hours
  • On-call owner for freeze-writes decisions
  • Forced drills: schema rename, empty enrichment, wrong tenant token, mid-run 500

If any checkbox is empty, you are still in demo mode with a production URL.

Score the audit the way you will score the run:

GatePassFail
SchemaLive payloads validate; unknown fields refuseStubs, or extras silently dropped
AuthBleed drill hard-refusesShared key “for now”
EvalGolden set has fail cases, not only winsOnly happy-path recordings
KillFlip blocks the next write in stagingRestart-the-box is the plan

Pass all four or do not soft-launch writes.

What readiness signals beat a green demo?

Do not use “demo succeeded” as a gate. Use a thin readiness scorecard.

SignalDemo-only smellReady signal
Tool fidelityStubs / fixturesLive schemas + recorded adapters
AuthSingle sandbox keyTenant-bound tokens under test
EvalHuman nodsAutomated criteria + escalate
Failure drillsNoneForced bad tool + auth bleed tests
ObservabilityChat consoleRun ids, tool spans, reason codes
Cost / budgetUnlimitedCap + kill switch wired
IntakeClean stringsTyped schema, dirty-field golden cases
On-call“We’ll watch Slack”Named freeze owner

Ship when the right column is true for the job types you are soft-launching — not when the slide deck looks clean.

OpenAI’s structured-outputs guide is the vendor version of the fidelity row: JSON mode gives you valid JSON; schema-constrained outputs give you the keys you asked for. See Structured model outputs. Valid JSON that invented primaryEmail is still a failed contract.

How do edge-case inputs collapse an agent?

Demos avoid messy inputs. Production receives empty strings, HTML in “plain text” fields, dual-language names, and ids that look valid but point at archived records. Agents without input validation treat garbage as narrative fuel.

Procedure for hardening intake:

  1. Define a typed intake schema per job type.
  2. Reject or escalate on schema fail before any model call.
  3. Normalize known dirty fields (strip HTML, trim, canonicalize phones).
  4. Add golden cases from the last ten production complaints.
  5. Never let the model “fix” an id that failed ownership check.
Dirty inputDemo fateProduction fate without a gateWith a gate
"" emailNever shownAgent invents oneintake_schema_fail
HTML in notesNever shownTags become “facts”Strip or reject
Archived CRM idHiddenWrite lands on a ghostOwnership + status check
Two legal namesOne fixtureWrong-account writeMulti-match escalate

Bravery at intake is not a product strategy.

What sequence closes the demo-to-prod gap?

Skipping steps compresses the gap into a single outage. Hold each stage until you can answer: what failed, which reason code, who owns the fix.

  1. Shadow mode — agent proposes; humans execute. Score proposals offline.
  2. Write with human approve — irreversible tools behind a button.
  3. Narrow autonomy — one job type, one tenant cohort, hard budget.
  4. Widen only after online scores hold for a defined window.
StageWritesWhat you must be able to show
ShadowNoneProposal vs human action, scored
ApproveHuman-gatedTime-to-approve, reject reasons
NarrowOne job, one cohortReason-code histogram, freeze drill
WidenMore jobs or tenantsScores held; no open bleed bugs

If shadow mode only produces vibes, you are still demoing. The parent operating picture — evaluators, sandboxes, stop conditions — lives in the operating manual. If you cannot name those objects yet, you may not want an agent at all; that decision is when not to build an agent.

What belongs in the five-day pilot?

A Spurlock Studios $1,500 · 5-day agentic pilot is not a longer demo. Its job is to close the control gap on one real job: evaluator criteria, tool contracts, tenant binding, traces with reason codes, and a kill switch. You leave with a reliability audit trail, not applause.

DayObject you should be able to point at
1Job sentence, tool inventory, side-effect classes
2Pinned schemas + tenant bind on the runner
3Evaluator criteria + golden set (including fail cases)
4Forced drills + kill switch flip in staging
5Soak on live schemas; reason-code histogram; go / no-go

Widen autonomy only after those pieces exist. Context lives on /agentic and in the operating manual.

Which anti-patterns keep the gap open?

“We’ll add evals after launch.” Then launch is the eval — paid for by customers.

Stubbing tools forever. Schema drift never appears until it hurts.

Treating every wrong write as a prompt bug. You will rewrite prompts while the auth bug remains.

Measuring only success demos. Sample failures. Force them in staging.

Confusing model upgrade with harness upgrade. Different levers, different costs.

Sharing one long-lived API key across tenants “until SSO is ready.” Authorization bleed is not a backlog item once writes are live. It is an incident waiting for a run id.

Kill switch equals Slack message. If the only freeze is “tell the intern to stop the box,” you will eat the next cascade.

Anti-patternWhat it costsReplacement
Evals laterCustomers become the golden setCriteria before writes
Eternal stubsFirst live schema is an outageSoak on real payloads
Prompt-first postmortemsAuth bugs surviveBlame table, harness first
Shared god-keyCross-tenant incidentPer-tenant bind
No freezeYou debug while it keeps writingRuntime write flag

Model or harness first?

Ask in order. Do not skip to the fun lever.

  1. Did a tool return an unexpected shape or error? → harness / adapter
  2. Did the run cross tenant boundaries? → harness / auth (incident)
  3. Did a bad observation cascade into a write? → verify gates + evaluator
  4. Did criteria fail but the agent still terminated “success”? → evaluator authority
  5. Only then: did the model choose a wrong plan under correct observations? → prompt / model
QuestionIf yesIf you start at the model
Unexpected shape?Adapter + fail closedYou will “fix” a field the API renamed
Cross-tenant?Incident + bindYou will prompt “don’t leak”
Cascade?Verify gateYou will add “double-check” to the prompt
False done?Evaluator owns terminalYou will ask the worker to be honest
Plan wrong on good data?Now you may touch the modelThis is the only row that belongs here

If you start at step 5 every time, you will never close the demo→prod gap.

Forced failure drills (staging only)

Before soft-launch, break the agent on purpose.

DrillInjectExpect
Schema renameAdapter returns new field namesFail closed + alert, no invented fields
Empty enrichmentTool returns []Escalate or verify — no CRM write from fiction
Wrong tenant tokenHarness omits or binds bad tenant_idHard refuse + tool_auth_error
Mid-run 500Write tool errors onceBounded retry or escalate; no silent success
Stale recordupdated_at older than cacheresult_stale, no write
Kill switchFlip freeze_writes mid-runNext write blocked, killed_by_switch
Dirty intakeHTML + empty emailReject before first model call

If a drill does not produce the expected reason code, you found a control gap while the blast radius is still staging.

Procedure:

  1. Pick one job type. Do not drill a fleet.
  2. Run the table above against live schemas, not stubs.
  3. File a ticket per missed reason code. Do not “note it.”
  4. Re-run until every row matches.
  5. Keep the traces. They are the start of the golden set.

A drill you never re-run after a harness change is a memory, not a control.

FAQ

Why do edge-case inputs collapse agents?

Because demos never train the harness on dirty intake. Empty fields, HTML junk, and ambiguous ids become model narrative instead of schema rejects. Validate and escalate before the first tool call, then add those cases to the golden set so the next dirty payload fails the same way.

How does schema drift break tools months later?

APIs rename fields and change nullability without your prompt noticing. The agent fills gaps with invented structure that “should” exist. Version adapters, fail closed on unknown shapes, and alert when adapter errors spike — that is a contract failure, not a sudden drop in model IQ.

What’s cascade failure after a bad tool result?

One wrong or empty tool observation becomes “truth” for later planning, so downstream writes look like hallucination. Require verification for high-stakes entities and evaluator criteria that reject unsupported assertions. If candidates.length is not one, escalate; do not pick a favorite.

How do I test authorization bleed across tenants?

In staging, run two tenants with distinct data and deliberately swap or omit tenant_id on tool calls. The harness must refuse. Add automated cases that attempt cross-tenant reads and writes and expect hard failure plus a tool_auth_error reason code.

Should I blame the model or the harness first?

Harness first: schemas, auth binding, evaluators, cascade brakes, kill switches. Genuine model error is real, but launch-week failures are usually control gaps mislabeled as intelligence failures. Touch the prompt only after observations, tenancy, and terminals check out.

What’s the five-day pilot’s job in closing this gap?

Install the minimum control loop on one job — criteria, contracts, tenant binding, traces, kill switch — against live schemas. The pilot proves production-readiness machinery, not a prettier demo path. You should leave able to freeze writes and name last night’s reason codes.

CTA

Close the control loop before you scale the demo.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

Why do edge-case inputs collapse agents?
Because demos never train the harness on dirty intake. Empty fields, HTML junk, and ambiguous ids become model narrative instead of schema rejects. Validate and escalate before the first tool call, then add those cases to the golden set so the next dirty payload fails the same way.
How does schema drift break tools months later?
APIs rename fields and change nullability without your prompt noticing. The agent fills gaps with invented structure that “should” exist. Version adapters, fail closed on unknown shapes, and alert when adapter errors spike — that is a contract failure, not a sudden drop in model IQ.
What's cascade failure after a bad tool result?
One wrong or empty tool observation becomes “truth” for later planning, so downstream writes look like hallucination. Require verification for high-stakes entities and evaluator criteria that reject unsupported assertions. If `candidates.length` is not one, escalate; do not pick a favorite.
How do I test authorization bleed across tenants?
In staging, run two tenants with distinct data and deliberately swap or omit `tenant_id` on tool calls. The harness must refuse. Add automated cases that attempt cross-tenant reads and writes and expect hard failure plus a `tool_auth_error` reason code.
Should I blame the model or the harness first?
Harness first: schemas, auth binding, evaluators, cascade brakes, kill switches. Genuine model error is real, but launch-week failures are usually control gaps mislabeled as intelligence failures. Touch the prompt only after observations, tenancy, and terminals check out.
What's the five-day pilot’s job in closing this gap?
Install the minimum control loop on one job — criteria, contracts, tenant binding, traces, kill switch — against live schemas. The pilot proves production-readiness machinery, not a prettier demo path. You should leave able to freeze writes and name last night’s reason codes.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot