Why Agent Demos Die in Production: Control Gaps, Not Model IQ
Demo success proves a staged happy path. Production fails when the control loop — schemas, auth, evaluators, and kill switches — was never in the harness.
William Spurlock Founder — Spurlock Studios Updated 24 MIN
Your AI agent works in the demo but fails in production because the demo optimized for a staged happy path, not a control loop. The model did not get dumber overnight. Production introduced real schemas, real tenants, real tool failures, and real side effects — and your harness never owned those constraints.
This spoke sits under the Agentic Systems Operating Manual. If the job is still a known path, stop here and read when not to build an agent. The rest of this page assumes you already earned the loop and now need it to survive contact with live tools.
The short answer
- Demos prove “can it look smart once?” Production asks “can it stay correct under drift, load, and bad tools?”
- Most “hallucinations” in launch week are system bugs: schema drift, stale tool results, authorization bleed, cascade after one bad tool call.
- Readiness is a checklist on the harness — evaluators, sandboxes, reason codes, kill switches — not a higher model tier.
- Soft-launch without a reliability audit is how you buy expensive chaos.
- Fix the control loop first; then argue about prompts.
What does a demo prove that production does not?
A demo is a theater set. Inputs are curated. Tools return clean JSON. Credentials are a single sandbox tenant. Latency is low. Nobody else is writing to the CRM while the agent runs. The audience watches one path succeed.
Production is adversarial by accident. Vendor fields rename. Tokens expire mid-run. Two companies share a name. A write tool returns 200 with an empty body. Overnight, there is no human in the room to say “that note went on the wrong account.”
| Demo assumption | Production reality |
|---|---|
| Fixed tool schemas | Vendors ship breaking field renames |
| One tenant, one role | Multi-tenant auth with bleed risk |
| Tools always succeed | Partial failures, timeouts, empty results |
| Single operator watching | Overnight runs, no human in the room |
| “Looks right” is enough | Evaluator criteria or a customer complaint |
| Shared sandbox key | Per-tenant tokens, rotation, expiry |
Fiddler’s production write-up makes the same split without dressing it up: demo and test environments use curated data and controlled conditions; production adds rate limits, authentication expiry, concurrent sessions, and edge cases the demo never saw. See Why do 70-95% of AI agent projects fail in production?. Treat the headline range as a vendor estimate as of 2026, not a law. The useful claim is the gap, not the percent.
If your success metric was applause, you measured the wrong thing.
Why is the control loop the bottleneck, not model IQ?
When the demo dies, the instinct is to swap models or rewrite the system prompt. That treats intelligence as the bottleneck. For business agents, the bottleneck is almost always the control loop: plan → act → evaluate → revise → terminate, with deterministic guards around non-deterministic steps.
I have spent 20,000+ hours architecting agentic systems and built 500+ automations. The runs that survived launch week all had the same boring furniture: typed tool contracts, tenant-bound credentials, an evaluator that can refuse “done,” and a write freeze that does not wait on a prompt deploy.
Name the gap before you touch the model:
- No evaluator — terminal success is “HTTP 200” or “agent said done.”
- No side-effect classes — write tools run with the same trust as read tools.
- No schema contracts — tool args and results are free-form blobs.
- No tenant binding — run context does not pin authorization.
- No cascade brake — one bad tool result feeds the next plan as truth.
- No kill switch — stopping writes means shipping a new prompt.
AWS’s own agent-building series is blunt about the first of those: poorly defined tool schemas are a leading source of agent failures, because the schema is the contract between reasoning and action. That line lives on AI agent frameworks and building blocks. Upgrade the harness. Then, if quality is still soft, change the model. Blame order matters.
Which four controls close the demo-to-prod gap?
Four objects. If any one is missing, you are still demoing.
| Control | What it owns | Demo smell | Production bar |
|---|---|---|---|
| Schemas | Tool args and results | “The model usually fills this in” | Versioned JSON Schema, fail closed |
| Auth | Who the run may act as | One long-lived API key | Tenant + role bound per call |
| Evals | Whether “done” is legal | A human nodded in the room | Criteria the runner can enforce |
| Kill switch | Stopping writes now | Restart the process / redeploy | Freeze writes without a prompt change |
NIST’s AI Risk Management Framework is the governance vocabulary for this, not a product pitch. The Core is GOVERN, MAP, MEASURE, MANAGE. A demo that never MEASURES tool-shape errors or MANAGES a freeze path has not left GOVERN-as-slide-deck.
- Every tool has a pinned input schema and a pinned output schema
- Every tool call carries
tenant_idand role from the runner, not from the model - “Done” is an evaluator verdict, not a sentence the worker typed
- Writes can be frozen without touching prompts
Four boxes. Empty any one and the rest become theater.
How do unversioned tool schemas invent fields?
The tool once returned customer.email. Three months later the API returns contact.primaryEmail. The agent invents an email field that “should” exist. That is not a creative model. That is an unversioned contract.
OpenAI’s function-calling guide tells you to define tools with JSON Schema and to turn on strict so argument objects adhere to the schema instead of “best effort.” See Function calling. Anthropic’s tool-use docs make the same demand on the other side of the fence: every custom tool needs an input_schema object. See Implement tool use. Neither vendor will version your CRM adapter for you.
JSON Schema itself is the vocabulary. The current dialect is 2020-12; declare it. The spec index is JSON Schema — Specification.
| Layer | Pin this | Fail closed when |
|---|---|---|
| Model → tool args | Vendor strict / input_schema | Extra keys, missing required, wrong types |
| Tool → adapter | Your versioned output schema | Unknown fields, null where required |
| Adapter → runner | Reason code + span | Parse error rate spikes |
| Runner → write | Idempotency key + tenant | Schema miss on a write tool |
What breaks: CRM writes with null emails, silent skips, or fabricated values.
What you do: Pin schemas, version adapters with dates, refuse unknown shapes, alert on adapter-error rate — not on “the model sounded unsure.”
- Snapshot today’s live payload for each tool. That is the contract, not the vendor’s marketing docs.
- Wrap it in a schema. Mark required fields. Ban extras.
- Store
schema_versionon every span. - On mismatch: no write,
schema_mismatchreason code, page the adapter owner.
A prompt that says “be careful with fields” is not a contract.
How does authorization bleed across tenants?
Demo used one API key. Production shares a worker pool. A run for Tenant A accidentally carries Tenant B’s token, or a tool accepts an id without checking ownership. The model did not “decide” to leak. The harness never enforced tenancy.
This is an incident, not a quality ticket.
OWASP’s Top 10 for LLM Applications names the shape: LLM06 Excessive Agency is damaging action from unexpected, ambiguous, or manipulated model output, usually because of excessive functionality, excessive permissions, or excessive autonomy. A shared god-key is all three.
Microsoft’s Zero Trust note on agents is the same rule in identity language: define identity, scope, tool access, and auditability before you widen autonomy, and test revocation — disable the agent, rotate credentials, invalidate tokens. See Least privilege for AI agents. Foundry’s agent-identity write-up adds the implementation: provision an agent identity, assign RBAC to that identity, and exchange a token at tool time instead of stuffing secrets into prompts. See Agent identity. You do not need Azure to steal the pattern.
| Binding | Who sets it | Who must not set it |
|---|---|---|
tenant_id | Runner, from the job ticket | Model, tool args, retrieved text |
| Role / scope | Policy table for the job type | “The agent asked for admin” |
| Credential | Secret store / token exchange | Prompt, memory, tool description |
| Resource id | Verified against tenant ownership | Trust the model’s account_id |
What breaks: Cross-customer reads or writes.
What you do:
- Bind
tenant_idat the harness for every tool call - Refuse tools that ignore the binding
- Execute writes in the caller’s authorization scope, not a shared god-key
- Test bleed on purpose in staging (see drills below)
Authorization is not a system-prompt paragraph. It is a check the model cannot skip.
What happens when one bad tool result cascades?
First tool returns an empty list or a wrong match. The agent treats that as ground truth, invents a narrative, and writes it downstream. Later steps look like hallucination. The root cause was trusting a bad observation without a verify step.
| Step | Observation | If you trust it | If you brake |
|---|---|---|---|
| 1 | Enrichment returns [] | Agent “recalls” a company | Escalate: empty_enrichment |
| 2 | Two CRM ids share a name | Writes the first hit | Gate: candidates.length == 1 |
| 3 | Write tool 500s once | “Done” on the retry story | Bounded retry or escalate |
| 4 | Read-after-write misses | Invents a confirmation | Verify tool or fail |
What breaks: Plausible wrong CRM notes, tickets, or emails.
What you do: Require verification tools for high-stakes entities. Evaluator criteria must reject “asserted without evidence.” Escalate when criteria fail — not when the model feels unsure.
Genuine model error exists. Lead with harness bugs first. You will be right more often.
Why do evaluators have to own “done”?
Because the worker is incentivized to finish. Left alone, it will narrate success after a partial write, a skipped verify, or a tool that returned 200 with garbage.
LangSmith’s agent-eval docs split the job into three scores you can actually implement: final response, single step, and trajectory — whether the path of tool calls was the path you meant. Start at Evaluate a complex agent and trajectory evaluations. You do not have to buy LangSmith. You do have to score more than the last sentence.
| Eval type | Question it answers | Demo substitute |
|---|---|---|
| Final | Did the terminal artifact meet criteria? | “Looks good in the chat” |
| Step | Was this tool the right one, with legal args? | None |
| Trajectory | Did the path stay inside the allowlist? | Happy-path recording |
| Policy | Did any call violate tenant, budget, or write class? | Hope |
Rules that belong in code, not in a rubric paragraph:
- The worker cannot mark
done. Only the evaluator can. eval_passrequires evidence pointers (ids, etags, screenshots of the write), not vibes.eval_failwith revisions remaining goes torevise. Ceiling hit goes toescalate.- Policy fail goes to
abortor freeze-writes. It does not get another try.
If “done” is a string the model is allowed to emit, you do not have an evaluator. You have a narrator.
What does a kill switch actually freeze?
A kill switch is a runtime flag that stops side effects without a deploy. It is not “we will revert the prompt.” It is not “restart the worker.” It is a gate in front of every write tool that a human can flip in one place.
Amazon’s AgentOps write-up on Bedrock AgentCore treats evaluation, observability, and governance as the production pillars — not a prettier planner. See AgentOps: operationalize agentic AI at scale. A freeze flag is the smallest governance object that still matters at 2 a.m.
| Switch | Scope | When you flip it |
|---|---|---|
freeze_writes | All write tools, all tenants | Unknown blast radius |
freeze_writes:job_type | One job | One flow is poisoning CRM |
freeze_writes:tenant | One customer | Suspected bleed or bad data |
freeze_model | New model calls | Cost spike or provider outage |
shadow_only | Propose, do not execute | You are still scoring |
Requirements:
- Flip does not require a prompt change or a container rebuild
- In-flight runs see the flag before the next write, not after
- Reads may continue so you can diagnose
- Every blocked call gets
killed_by_switchplus the switch name - On-call knows who is allowed to flip it
A kill switch you cannot find in the runbook is a slide.
How do launch-week “hallucinations” map to harness bugs?
Treat these as system bugs until proven otherwise.
Schema drift
Unversioned adapters plus invented fields. Covered above. Reason code: schema_mismatch.
Stale tool results
A cache or previous span result is reused after the underlying record changed. The agent plans from yesterday’s pipeline stage and books the wrong follow-up.
What breaks: Wrong next actions that look internally consistent.
What you do: TTL on tool results, etags or updated_at checks before writes, span metadata that marks result_stale=true.
Authorization bleed
Shared keys, missing tenant bind, resource ids the model supplied. Reason code: tool_auth_error. Incident process, not a prompt tweak.
Cascade after one bad observation
Empty or wrong tool result treated as ground truth. Reason code: unverified_assertion or empty_enrichment.
Intake garbage treated as narrative
Empty strings, HTML in “plain text,” dual-language names, archived ids that still look valid. Reason code: intake_schema_fail.
| Label you will hear | Check this first | Only then |
|---|---|---|
| “It hallucinated an email” | Output schema + adapter version | Prompt / model |
| “It wrote the wrong account” | Disambiguation gate + tenant bind | Prompt / model |
| “It ignored the policy” | Whether policy lives outside the model | Prompt / model |
| “It looped all night” | Step ceiling + kill switch | Prompt / model |
| “It leaked another customer” | Credential binding — incident | Do not “prompt harder” |
If your postmortem starts at the system prompt, you skipped the table.
Illustrative walkthrough: demo green, prod red
Illustrative — not a client result. Staging demo: agent looks up a lead, enriches firmographics, writes a CRM note. Tools are stubbed to always return the same Acme Corp payload. Soft-launch: real CRM has duplicate company names; enrichment returns two candidates; the agent picks the wrong one and writes a note on the wrong account.
Misdiagnosis: “the model hallucinated the company.”
Actual failure mode: no disambiguation gate, no evaluator check that crm_account_id matched the enrichment candidate ids, no escalate path for multi-match.
The fix is a control: if candidates.length != 1 → escalate, plus a golden case for duplicate company names. The model upgrade is optional.
Walk the same incident through the four controls:
| Control | What was missing | What you add |
|---|---|---|
| Schema | Enrichment result was an untyped list | candidates: array, minItems / maxItems rules |
| Auth | Write used a shared CRM key | Tenant-scoped token; id must belong to tenant |
| Eval | “Wrote a note” counted as done | crm_account_id ∈ candidate_ids |
| Kill switch | Wrong-account writes kept flowing | freeze_writes:crm_note until the gate ships |
That is a reliability audit in one incident. Do it in staging on purpose, not in week one of soft-launch by accident.
How do you run a reliability audit before soft-launch?
Run this before the agent can write in production. Timebox it. Do not wait for a perfect platform.
- Inventory every tool: read / write / irreversible; name the side-effect class
- Pin and version each tool schema; record adapters with dates
- Bind tenant + role to every tool invocation in the harness
- Require evaluator criteria for the job type; worker cannot self-certify
- Define terminal reason codes (
eval_pass,tool_auth_error,policy_violation,killed_by_switch) - Kill switch that freezes writes without redeploying prompts
- Golden set with at least one case per known failure mode above
- Staging soak with real schemas (not stubs) for 48 hours
- On-call owner for freeze-writes decisions
- Forced drills: schema rename, empty enrichment, wrong tenant token, mid-run 500
If any checkbox is empty, you are still in demo mode with a production URL.
Score the audit the way you will score the run:
| Gate | Pass | Fail |
|---|---|---|
| Schema | Live payloads validate; unknown fields refuse | Stubs, or extras silently dropped |
| Auth | Bleed drill hard-refuses | Shared key “for now” |
| Eval | Golden set has fail cases, not only wins | Only happy-path recordings |
| Kill | Flip blocks the next write in staging | Restart-the-box is the plan |
Pass all four or do not soft-launch writes.
What readiness signals beat a green demo?
Do not use “demo succeeded” as a gate. Use a thin readiness scorecard.
| Signal | Demo-only smell | Ready signal |
|---|---|---|
| Tool fidelity | Stubs / fixtures | Live schemas + recorded adapters |
| Auth | Single sandbox key | Tenant-bound tokens under test |
| Eval | Human nods | Automated criteria + escalate |
| Failure drills | None | Forced bad tool + auth bleed tests |
| Observability | Chat console | Run ids, tool spans, reason codes |
| Cost / budget | Unlimited | Cap + kill switch wired |
| Intake | Clean strings | Typed schema, dirty-field golden cases |
| On-call | “We’ll watch Slack” | Named freeze owner |
Ship when the right column is true for the job types you are soft-launching — not when the slide deck looks clean.
OpenAI’s structured-outputs guide is the vendor version of the fidelity row: JSON mode gives you valid JSON; schema-constrained outputs give you the keys you asked for. See Structured model outputs. Valid JSON that invented primaryEmail is still a failed contract.
How do edge-case inputs collapse an agent?
Demos avoid messy inputs. Production receives empty strings, HTML in “plain text” fields, dual-language names, and ids that look valid but point at archived records. Agents without input validation treat garbage as narrative fuel.
Procedure for hardening intake:
- Define a typed intake schema per job type.
- Reject or escalate on schema fail before any model call.
- Normalize known dirty fields (strip HTML, trim, canonicalize phones).
- Add golden cases from the last ten production complaints.
- Never let the model “fix” an id that failed ownership check.
| Dirty input | Demo fate | Production fate without a gate | With a gate |
|---|---|---|---|
"" email | Never shown | Agent invents one | intake_schema_fail |
| HTML in notes | Never shown | Tags become “facts” | Strip or reject |
| Archived CRM id | Hidden | Write lands on a ghost | Ownership + status check |
| Two legal names | One fixture | Wrong-account write | Multi-match escalate |
Bravery at intake is not a product strategy.
What sequence closes the demo-to-prod gap?
Skipping steps compresses the gap into a single outage. Hold each stage until you can answer: what failed, which reason code, who owns the fix.
- Shadow mode — agent proposes; humans execute. Score proposals offline.
- Write with human approve — irreversible tools behind a button.
- Narrow autonomy — one job type, one tenant cohort, hard budget.
- Widen only after online scores hold for a defined window.
| Stage | Writes | What you must be able to show |
|---|---|---|
| Shadow | None | Proposal vs human action, scored |
| Approve | Human-gated | Time-to-approve, reject reasons |
| Narrow | One job, one cohort | Reason-code histogram, freeze drill |
| Widen | More jobs or tenants | Scores held; no open bleed bugs |
If shadow mode only produces vibes, you are still demoing. The parent operating picture — evaluators, sandboxes, stop conditions — lives in the operating manual. If you cannot name those objects yet, you may not want an agent at all; that decision is when not to build an agent.
What belongs in the five-day pilot?
A Spurlock Studios $1,500 · 5-day agentic pilot is not a longer demo. Its job is to close the control gap on one real job: evaluator criteria, tool contracts, tenant binding, traces with reason codes, and a kill switch. You leave with a reliability audit trail, not applause.
| Day | Object you should be able to point at |
|---|---|
| 1 | Job sentence, tool inventory, side-effect classes |
| 2 | Pinned schemas + tenant bind on the runner |
| 3 | Evaluator criteria + golden set (including fail cases) |
| 4 | Forced drills + kill switch flip in staging |
| 5 | Soak on live schemas; reason-code histogram; go / no-go |
Widen autonomy only after those pieces exist. Context lives on /agentic and in the operating manual.
Which anti-patterns keep the gap open?
“We’ll add evals after launch.” Then launch is the eval — paid for by customers.
Stubbing tools forever. Schema drift never appears until it hurts.
Treating every wrong write as a prompt bug. You will rewrite prompts while the auth bug remains.
Measuring only success demos. Sample failures. Force them in staging.
Confusing model upgrade with harness upgrade. Different levers, different costs.
Sharing one long-lived API key across tenants “until SSO is ready.” Authorization bleed is not a backlog item once writes are live. It is an incident waiting for a run id.
Kill switch equals Slack message. If the only freeze is “tell the intern to stop the box,” you will eat the next cascade.
| Anti-pattern | What it costs | Replacement |
|---|---|---|
| Evals later | Customers become the golden set | Criteria before writes |
| Eternal stubs | First live schema is an outage | Soak on real payloads |
| Prompt-first postmortems | Auth bugs survive | Blame table, harness first |
| Shared god-key | Cross-tenant incident | Per-tenant bind |
| No freeze | You debug while it keeps writing | Runtime write flag |
Model or harness first?
Ask in order. Do not skip to the fun lever.
- Did a tool return an unexpected shape or error? → harness / adapter
- Did the run cross tenant boundaries? → harness / auth (incident)
- Did a bad observation cascade into a write? → verify gates + evaluator
- Did criteria fail but the agent still terminated “success”? → evaluator authority
- Only then: did the model choose a wrong plan under correct observations? → prompt / model
| Question | If yes | If you start at the model |
|---|---|---|
| Unexpected shape? | Adapter + fail closed | You will “fix” a field the API renamed |
| Cross-tenant? | Incident + bind | You will prompt “don’t leak” |
| Cascade? | Verify gate | You will add “double-check” to the prompt |
False done? | Evaluator owns terminal | You will ask the worker to be honest |
| Plan wrong on good data? | Now you may touch the model | This is the only row that belongs here |
If you start at step 5 every time, you will never close the demo→prod gap.
Forced failure drills (staging only)
Before soft-launch, break the agent on purpose.
| Drill | Inject | Expect |
|---|---|---|
| Schema rename | Adapter returns new field names | Fail closed + alert, no invented fields |
| Empty enrichment | Tool returns [] | Escalate or verify — no CRM write from fiction |
| Wrong tenant token | Harness omits or binds bad tenant_id | Hard refuse + tool_auth_error |
| Mid-run 500 | Write tool errors once | Bounded retry or escalate; no silent success |
| Stale record | updated_at older than cache | result_stale, no write |
| Kill switch | Flip freeze_writes mid-run | Next write blocked, killed_by_switch |
| Dirty intake | HTML + empty email | Reject before first model call |
If a drill does not produce the expected reason code, you found a control gap while the blast radius is still staging.
Procedure:
- Pick one job type. Do not drill a fleet.
- Run the table above against live schemas, not stubs.
- File a ticket per missed reason code. Do not “note it.”
- Re-run until every row matches.
- Keep the traces. They are the start of the golden set.
A drill you never re-run after a harness change is a memory, not a control.
FAQ
Why do edge-case inputs collapse agents?
Because demos never train the harness on dirty intake. Empty fields, HTML junk, and ambiguous ids become model narrative instead of schema rejects. Validate and escalate before the first tool call, then add those cases to the golden set so the next dirty payload fails the same way.
How does schema drift break tools months later?
APIs rename fields and change nullability without your prompt noticing. The agent fills gaps with invented structure that “should” exist. Version adapters, fail closed on unknown shapes, and alert when adapter errors spike — that is a contract failure, not a sudden drop in model IQ.
What’s cascade failure after a bad tool result?
One wrong or empty tool observation becomes “truth” for later planning, so downstream writes look like hallucination. Require verification for high-stakes entities and evaluator criteria that reject unsupported assertions. If candidates.length is not one, escalate; do not pick a favorite.
How do I test authorization bleed across tenants?
In staging, run two tenants with distinct data and deliberately swap or omit tenant_id on tool calls. The harness must refuse. Add automated cases that attempt cross-tenant reads and writes and expect hard failure plus a tool_auth_error reason code.
Should I blame the model or the harness first?
Harness first: schemas, auth binding, evaluators, cascade brakes, kill switches. Genuine model error is real, but launch-week failures are usually control gaps mislabeled as intelligence failures. Touch the prompt only after observations, tenancy, and terminals check out.
What’s the five-day pilot’s job in closing this gap?
Install the minimum control loop on one job — criteria, contracts, tenant binding, traces, kill switch — against live schemas. The pilot proves production-readiness machinery, not a prettier demo path. You should leave able to freeze writes and name last night’s reason codes.
CTA
Close the control loop before you scale the demo.
What questions does this article answer?
- Why do edge-case inputs collapse agents?
- Because demos never train the harness on dirty intake. Empty fields, HTML junk, and ambiguous ids become model narrative instead of schema rejects. Validate and escalate before the first tool call, then add those cases to the golden set so the next dirty payload fails the same way.
- How does schema drift break tools months later?
- APIs rename fields and change nullability without your prompt noticing. The agent fills gaps with invented structure that “should” exist. Version adapters, fail closed on unknown shapes, and alert when adapter errors spike — that is a contract failure, not a sudden drop in model IQ.
- What's cascade failure after a bad tool result?
- One wrong or empty tool observation becomes “truth” for later planning, so downstream writes look like hallucination. Require verification for high-stakes entities and evaluator criteria that reject unsupported assertions. If `candidates.length` is not one, escalate; do not pick a favorite.
- How do I test authorization bleed across tenants?
- In staging, run two tenants with distinct data and deliberately swap or omit `tenant_id` on tool calls. The harness must refuse. Add automated cases that attempt cross-tenant reads and writes and expect hard failure plus a `tool_auth_error` reason code.
- Should I blame the model or the harness first?
- Harness first: schemas, auth binding, evaluators, cascade brakes, kill switches. Genuine model error is real, but launch-week failures are usually control gaps mislabeled as intelligence failures. Touch the prompt only after observations, tenancy, and terminals check out.
- What's the five-day pilot’s job in closing this gap?
- Install the minimum control loop on one job — criteria, contracts, tenant binding, traces, kill switch — against live schemas. The pilot proves production-readiness machinery, not a prettier demo path. You should leave able to freeze writes and name last night’s reason codes.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.