Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
William Spurlock Founder — Spurlock Studios 34 MIN
Your agent looks 10× more expensive than the chatbot demo because it is a different product with a different bill. The demo is one model completion: a system prompt, a user message, an answer. The agent is a loop. Each turn re-sends the system prompt, the tool schemas, the growing transcript, and the last tool dump so the model can decide what to do next. Operators feel that as “10×.” That ratio is a complaint to diagnose on your traces. It is not a measured industry tax, and this post will not invent one.
This spoke sits under the Agentic Systems Operating Manual. Shape choice — workflow versus loop — is owned by agent loop vs LLM workflow step. If you still need to prove one job on real data before you argue about invoices, scope it with agent pilot scope.
The short answer
- A chatbot demo bills one turn, zero tools, a short prefix. An agent bills N turns × (static prefix + growing context), plus tool APIs, plus any evaluator rounds.
- Anthropic’s Building effective agents (Dec 19, 2024) is explicit: agentic systems “often trade latency and cost for better task performance,” and autonomy means “higher costs, and the potential for compounding errors.” They do not publish a universal 10×.
- Input token-exposures are not “tokens × turns.” Each turn re-reads what every prior turn left behind. That triangular term is the planning/tool-loop tax.
- Prompt caching discounts a stable prefix. It does not flatten fat tool results you append every step.
- Cut the tax with turn caps, truncated tool digests, a smaller allowlist, and demoting known paths to a workflow. Swap models last.
Why is the chatbot demo a different bill?
The demo you showed finance is almost always a single-turn chat completion. One request. One response. No tools, or a fake “search” that never ran. The system prompt is a paragraph. The user message is one question. That bill is honest for that product. It is not a forecast for an agent that can call CRM, search, and a code runner until it feels done.
n8n’s AI Agent node is the other product: connect a chat model and tools, and the node “decides which tools to call to complete a task.” One workflow execution is not one model call. LangGraph draws the same line in workflows vs agents: predetermined code paths versus a model that directs process and tool usage.
| What finance saw | What production ran | Why the bills diverge |
|---|---|---|
| Playground chat, one question | Job with tools and retries | Demo never paid for a loop |
| System prompt ~a paragraph | System + full tool JSON schemas | Prefix is 3–10× before the first tool |
| No tools, or mocked search | Live CRM / search / sandbox | Tool dumps become next-turn context |
| You clicked Stop | No turn cap, no kill switch | N is unbounded |
| “It answered” | Evaluator + revisions | Extra model calls per job |
| Cost per message | Cost per job | Denominator changed |
If you cannot point to the demo’s trace and show N = 1, tools = 0, stop comparing it to the agent. You are comparing a completion to a control plane.
Checklist before you quote a multiplier in a meeting:
- Demo trace exported: one request, one response, token usage object
- Agent trace exported for the same job sentence
- Turn count, tool-call count, and input tokens per turn on both
- Same model pin, or you admit the bake-off is contaminated
- You are comparing cost per job, not cost per API request
The demo is a poster. The agent is the machine the poster was selling.
What is the planning tax?
Planning tax is every token you spend so the model can choose the next action instead of you encoding that choice in a graph. It is not “the model thinking too much” as a vibe. It is billed context: the job contract, the tool catalog, the scratch of prior plans, and the “what should I do now?” completion on every turn.
Anthropic’s essay treats agents as LLMs using tools from environmental feedback in a loop. That loop is the tax. A workflow with one structured-output call pays planning once. An agent pays it on turn 1, turn 2, and turn 12, and turn 12 still has to carry the catalog that made turn 1 possible.
| Planning ingredient | Where it lives | Paid again every turn? |
|---|---|---|
| Job / policy text | System prompt | Yes, unless cached |
| Tool names, descriptions, JSON schemas | Tools array / MCP catalog | Yes, unless cached |
| “Think about the next step” completion | Assistant output | Yes — that output becomes later input |
| Prior plan fragments | Transcript | Yes, until you drop them |
| Evaluator rubric (if in-band) | Extra messages | If you stuffed it into the worker |
MCP makes this worse when you are sloppy. A host that dumps a giant tool catalog into every request is paying a planning tax for tools the job will never call. Allowlist the job. Do not “connect the registry and see.”
Procedure to separate planning tax from work tax on one trace:
- Sum input tokens that are system prompt + tool definitions. Call that S (static prefix).
- Sum assistant tokens that are plans, not tool arguments or the final artifact. Call that P.
- Sum tool-result tokens. Call that R.
- If S + P dominates R on a job that only needed three known HTTP calls, you bought an agent for a workflow.
- If R dominates and still grows each turn, you have a tool-loop tax, next section.
Planning is the fee for uncertainty. If the next action is nameable before the payload arrives, stop paying it.
What is the tool-loop tax?
Tool-loop tax is the re-billing of tool results (and the assistant turns that produced them) on every subsequent model call. The first search result was useful at step 2. At step 9 you are still paying to re-read the raw JSON unless you truncated it to a digest.
This is ordinary API physics, not a secret vendor trick. Chat Completions / Messages-style APIs send the full conversation you provide. If your harness appends tool messages verbatim, turn k includes turns 1…k−1. Independent tools called in sequence become a stack of blobs. Parallel tool calls in one turn are cheaper than the same calls as four separate turns, because you pay the prefix once for that round.
| Loop behavior | What gets appended | Tax shape |
|---|---|---|
| One tool, digest of 80 tokens | Small R | Manageable |
| One tool, raw 8,000-token API body | Fat R | Quadratic in later turns |
| Four sequential dependent tools | Four round trips | Four prefix payments |
| Four independent tools, parallel | One round trip | One prefix payment |
| Failed tool, full error page | Poison + size | Retry multiplies both |
| MCP resource dumped as context | Files you did not ask to keep | Prefix + R both jump |
n8n is honest about the shape: the agent uses tools to retrieve and act, and it “determine[s] which tool to use depending on the task.” That determination is a model call. The observation is another message. Repeat until done — or until you forgot to define done.
Decision list:
- Does the next tool depend on the last result? Sequential is fine; still digest the last result.
- Can two tools run with the same inputs? Parallel that turn.
- Is the raw body needed after the next decision? If no, store it outside the prompt and keep a 40–80 token digest in the transcript.
- Did the tool fail? Put a short error class in context, not the vendor’s HTML 500 page.
The loop is how agents earn their keep on messy jobs. Unbounded raw tool I/O is how they earn their reputation with finance.
Cost model: turns × tools × context
Do not multiply “tokens per demo message × expected tools.” That is the naive linear model, and it is wrong the moment the transcript grows. A diagnostic worksheet:
Input token-exposures (uncached, illustration — not a bill):
E = N × S + g × N × (N − 1) / 2
- N = model turns in the run (each time you call the model, including “I’ll call a tool” turns)
- S = static prefix tokens: system + tool schemas (+ frozen policy)
- g = average new tokens you append per turn after the first (tool digest or dump + assistant)
Output tokens are roughly linear in N. They matter, but for tool-heavy agents the compounding line is usually input re-reads. This table uses labeled toy numbers so you can see how an operator complaint like “10×” can appear without anyone measuring a universal tax. Replace S, g, and N with your p50 traces. Do not quote these rows as Spurlock Studios invoices or as industry benchmarks.
| Shape | N (turns) | Tools called | S (prefix) | g (added / turn) | E (input token-exposures) | E vs 1,200-token chatbot |
|---|---|---|---|---|---|---|
| Chatbot demo | 1 | 0 | 1,200 | 0 | 1,200 | 1× (baseline) |
| Agent, first look, no tool | 1 | 0 | 4,000 | 0 | 4,000 | 3.3× |
| Agent, short loop | 3 | 3 | 4,000 | 600 | 13,800 | 11.5× |
| Agent, typical write job | 5 | 6 | 4,000 | 600 | 26,000 | 21.7× |
| Agent, wander | 8 | 10 | 4,000 | 800 | 54,400 | 45.3× |
| Same wander, fat dumps | 8 | 10 | 8,000 | 2,000 | 120,000 | 100× |
Read the 3-turn row slowly. On these assumptions, input-token exposure already crosses the neighborhood of 10× versus a 1,200-token chatbot prefix. Change S or g and the crossing move. That is why “my agent is 10× the demo” is a useful ticket and a terrible statistic.
Second axis: tools. Tools change g (result size) and often N (extra turns to call them). A catalog of 40 MCP tools inflates S before anyone calls one.
| Lever | Hits | Direction |
|---|---|---|
| Turn cap | N | Linear on S, quadratic on g |
| Truncate tool results | g | Quadratic savings on later turns |
| Shrink tool catalog | S | Linear on every turn |
| Parallel independent tools | N | Fewer prefix payments |
| Prompt cache on S | Effective S | Discount, not a flatten of g |
| Evaluator as separate call | Extra N | Adds a line; do not hide it in the worker |
Procedure to fill the worksheet from last week’s production:
- Pick one
job_type. Ignore mixed dashboards. - For 30 runs, record N, tool-call count, input tokens per turn, output tokens, cache read/write tokens if the provider sends them.
- Estimate S from turn-1 input minus the user payload.
- Estimate g as median (input tokens at turn k − input tokens at turn k−1), after subtracting cache-only artifacts if you must.
- Compute E from the formula and from summed provider
input_tokens. If they disagree a lot, your harness is dropping or injecting context you did not model. - Divide total model spend for those runs by successes, not by runs. Cheap failures flatter cost per run.
Worked loop on the same labeled assumptions (S = 4,000, g = 600, N = 5). Turn k input ≈ S + g×(k − 1). These are token counts, not dollars.
| Turn | Tools this turn | Context the model re-reads | Input tokens this call | Cumulative E |
|---|---|---|---|---|
| 1 | 0–1 | Prefix + user | 4,000 | 4,000 |
| 2 | 1 | Prefix + prior assistant + tool digest | 4,600 | 8,600 |
| 3 | 1 | Everything above + last digest | 5,200 | 13,800 |
| 4 | 1 | Growing tail | 5,800 | 19,600 |
| 5 | 1 | Growing tail | 6,400 | 26,000 |
Five requests did not cost 5 × the first request. Request 5 is already 1.6× request 1, and the sum is 6.5× request 1. Compare that sum to a 1,200-token chatbot completion and you are in the “why is this ~20×” conversation — still a worksheet, still not a published benchmark.
If you cannot fill N, S, and g, you do not have a cost model. You have a credit-card alert.
Why does cost grow faster than turn count?
Because turn k re-reads turns 1…k−1. Sum of 1 through N is N(N+1)/2. The g term in the worksheet is that triangle. A 16-turn run is not “twice an 8-turn run.” If g is stable, the triangular piece is about four times larger, and you still pay N×S.
OpenAI’s tokenizer note is a sizing heuristic, not a bill: about four characters per token for common English, roughly 100 tokens ≈ 75 words. Use it to sanity-check a 20,000-character CRM payload you were about to paste into the transcript. Do not use it to forecast invoices.
| If you only watch… | You will think… | What you missed |
|---|---|---|
| Requests / minute | Traffic is flat | N per job climbed |
| Tokens / request | Each call is “about the same” | Later calls are fatter |
| Output tokens | “We asked it to be brief” | Input re-reads dominate |
| Demo vs agent request count | “It’s only 5× the calls” | Each call is not the demo’s size |
| Monthly $ | “Models got cheaper” | N and g grew into the discount |
Plot prefill / input tokens against step number for a handful of long runs. If that line slopes up, you are paying the square. If it is roughly flat, your harness is already digesting or resetting context. Flat per-step input is the difference between a job you can run and a job you can only demo.
- Per-turn input tokens logged, not only totals
- Step index on the span
- Tool result bytes before and after truncation
- A chart someone in finance can see without opening JSON
Output tokens do not save you here. Five turns at 400 output tokens is 2,000 generated tokens against 26,000 input-exposures on the worksheet above. Output is often priced higher per token; it is still usually the smaller pile when tools dump JSON. “Answer briefly” is a latency tweak, not a cost strategy, until g is small.
| N | N×S (S=4,000) | Triangle g×N×(N−1)/2 (g=600) | E | E / E(N=4) |
|---|---|---|---|---|
| 4 | 16,000 | 3,600 | 19,600 | 1.0× |
| 8 | 32,000 | 16,800 | 48,800 | 2.5× |
| 16 | 64,000 | 72,000 | 136,000 | 6.9× |
Doubling turns from 8 to 16 is not 2× tokens. The triangle takes over. That is why a “quick extra retry” policy is a finance event.
Slope up is the tax. Argue with the slope, not with the adjective “expensive.”
What else gets billed besides the worker loop?
The worker’s N×S+g is the line operators notice first. Production jobs grow other meters that the chatbot demo never had.
| Meter | Who bills it | Demo had it? | Hide risk |
|---|---|---|---|
| Worker model calls | LLM provider | Yes, once | You compare this to the demo and stop |
| Evaluator model calls | LLM provider | No | “Quality” looks free |
| Embeddings / retrieval | LLM or vector host | Maybe | RAG on every turn |
| Paid search / enrich APIs | Tool vendor | No | Per-call, not tokens |
| Code-exec / browser sandbox | Tool vendor | No | Minutes × machine |
| Prompt-cache writes | LLM provider | Rarely | First-turn premium |
| Human rewrite time | Payroll | You were the demo | Cost per clean ship |
| Retries / 429 storms | All of the above | No | Multiplies N |
Braintrust’s public arithmetic is the one to steal for the denominator: cost per success is cost per task divided by success rate, and cost per resolved request only counts work that clears quality gates. A config that succeeds one-in-six pays for roughly six attempts per win. I will not paste their study’s dollar figures as if they were yours.
Procedure for a single job type:
- Sum worker model cost.
- Sum evaluator model cost (separate pin, separate spans).
- Sum metered tools.
- Add human minutes × loaded rate for rewrites on the same jobs.
- Divide by successes, then by clean ships.
- If evaluator cost is a rounding error, either your evaluator is mechanical (good) or you are not actually evaluating (bad).
When “10× vs the demo” is the wrong diagnosis, the model loop is fine and a different meter moved. Check these before you rewrite prompts.
| Symptom | Loop looks | Actual driver | First move |
|---|---|---|---|
| Token $ flat, vendor $ up | N=3, small g | Paid search / enrich per call | Cap tool calls; cache results |
| Token $ up, N=1 | One huge completion | You stuffed the CRM dump into a “chatbot” | That is not an agent tax; that is a prompt dump |
| Model $ fine, payroll up | Pass rate high | Humans rewrite every artifact | Cost per clean ship |
| Embeddings line item new | N unchanged | Retrieval on every turn | Retrieve once; pass a digest |
| Weekend spike | Few jobs, huge N | One runaway without abort | Kill switch, then postmortem |
The chatbot demo had meter 1, once. Your agent has a stack. Compare stacks.
Does prompt caching erase the tax?
No. Caching discounts a stable prefix. The growing tail is still full price unless you keep it small.
Anthropic’s prompt caching docs (verify live; this class of number moves) bill cache reads at 0.1× base input, 5-minute cache writes at 1.25×, and 1-hour writes at 2×. OpenAI’s prompt caching guide, as of this writing, uses the same 1.25× write / 0.1× read shape for GPT-5.6 and later when you place breakpoints; older caching modes on that page are model-dependent. I am not pasting per-model dollar tables here. They rot. Read the live page before you build a finance model on them.
| What you cache | Helps | Does not help |
|---|---|---|
| System prompt + tool schemas (S) | Every later turn’s prefix | Fat tool dumps after the breakpoint |
| Frozen few-shot policy | Stable jobs | Per-ticket CRM JSON |
| Nothing, hoping the vendor “just caches” | OpenAI automatic modes, if you qualify | Anthropic-style explicit breakpoints you never set |
| The entire growing transcript | Only the unchanged leading tokens | Anything you prepend that busts the prefix |
Caching is a discount on S. Truncation is a cut to g. Turn caps cut N. Teams that “turn on caching” and still paste 8k-token search hits are applying a 90% discount to the wrong pile.
Checklist:
- Tool schemas sit in a cached prefix, not below a timestamp that changes every request
- Usage shows cache-read tokens on turn 2+, not only on marketing slides
- User payload and tool results sit after the breakpoint
- TTL matches your cadence (slow human-in-the-loop vs tight loop)
- You still have a turn cap, because a cached runaway is still a runaway
Same N=5 worksheet with S cached after turn 1, using the 1.25× write / 0.1× read multipliers from the vendor pages above. Tail tokens stay full price. Input-equivalent (not a $ invoice):
| Piece | Tokens | Multiplier | Input-equivalent |
|---|---|---|---|
| S cache write (turn 1) | 4,000 | 1.25× | 5,000 |
| S cache reads (turns 2–5) | 4 × 4,000 | 0.1× | 1,600 |
| Growing tail (triangle) | 6,000 | 1.0× | 6,000 |
| Cached total | — | — | 12,600 |
| Uncached E from earlier | — | — | 26,000 |
Caching roughly halved the prefix tax on this toy run. It did nothing to the 6,000 tail tokens. If g is 2,000 because you paste search hits, the tail dominates and the coupon looks disappointing. Verify with cache_read / cache_write on the usage object. If those fields are zero, you are paying list price for S every turn.
Cache is a coupon. It is not a kill switch.
What usually blows the bill first?
The failure mode is not “the model got pricier.” It is unbounded N plus untruncated R, usually after a demo that never hit either.
I have spent 20,000+ hours on agentic systems and shipped 500+ automations into production. The invoice surprise is almost always the same plot: playground chat looks cheap, someone wires ten tools, a search tool returns a novel, retries sit on “until success,” and there is no abort on token budget. Nobody fabricated a 10× study. The card just moved.
| Failure | What the trace shows | What you do instead |
|---|---|---|
| No turn cap | N = 20, 40, 80 on one ticket | Hard max turns → abort with reason |
| Raw tool dump | Single R of thousands of tokens reused every step | Digest; store raw outside the prompt |
| Giant MCP catalog | S already huge at turn 1 | Allowlist 2–5 tools for the job |
| Retry without budget | Same tool, same args, three fat errors | Error class + cap; then escalate |
| Evaluator inside the worker | Rubric tokens on every plan turn | Separate evaluator call, mechanical checks first |
| Demo model pin ≠ prod pin | Flagship on the loop, small model in the screenshot | Same pin or admit the comparison is theater |
| Cost per request dashboard | Looks “fine” while jobs get longer | Cost per success by job_type |
Worked pattern (labeled, not a client bill): a support-looking agent calls search, gets a 6,000-token hit, then calls CRM, then “thinks,” then searches again. N=8. g stays fat because nobody truncated. The chatbot demo for “what is our return policy?” was N=1, S≈1,200, tools=0. Finance asks why prod is “about 10×.” On a worksheet like the one above, you may already be past that at N=3. At N=8 you are arguing about a different sport.
Do not fix this with a hotter model. Hotter models that take more steps can lose even if the per-token rate is lower. Steps enter the triangle. Price enters linearly.
n8n-specific version of the same blow-up: a Chat node or a single Basic LLM Chain in a playground workflow is N=1. An AI Agent with five tool sub-nodes can sit in one execution and still be N=8 model calls. The execution list looks like “1 run.” The provider bill looks like a conversation. If you estimate cost from n8n execution count, you will always lose the argument with finance.
| n8n shape | What “1 execution” contains | Cost model |
|---|---|---|
| Chat / LLM node, no tools | One completion | Chatbot-demo math |
| HTTP + IF graph, one LLM extract | One bounded call | Workflow + LLM step |
| AI Agent, 2 allowlisted tools, max iterations set | Bounded loop | Fill N, S, g |
| AI Agent, MCP dump, default iterations | Unbounded loop | The surprise |
If you cannot trip the kill switch in staging with a fake infinite loop, you do not have a kill switch. You have a comment in the prompt.
How do I measure this on real traces?
Measure per job, with N, tools, tokens, cache, and terminal reason on one timeline. Provider token charts are traffic. Microsoft Foundry’s agent-tracing note is blunt about why chat logs fail here: many steps, changing order, long payloads, nesting — capture inputs, outputs, tool usage, token consumption, duration. Langfuse’s rule is the same spirit: prefer ingested provider usage over homemade tokenizer math.
| Field | Why | Fake if missing |
|---|---|---|
job_type | Mix hides the expensive shape | Blended $ / run |
turn_index | Slope of input tokens | Totals only |
tool_name + result bytes | g diagnosis | “Tools were used” |
input_tokens / output_tokens | Provider usage | Your guess |
cache_read / cache_write | Whether S is actually discounted | Caching folklore |
terminal (done / escalate / abort) | Denominator for success | Endless “running” |
revision count | Grind-to-green | Pass rate theater |
eval_cost | Hidden second model | Worker-only bill |
Procedure (one afternoon on a live job):
- Sample 30 production traces of one sentence-sized job.
- Chart input tokens vs
turn_index. Note p50 N and p95 N. - List tool result byte sizes. Mark anything > ~1–2k tokens that was re-sent.
- Compute cost per run, cost per pass, cost per success.
- Diff the playground demo trace for the same question, if it still exists.
- Write the gap as “N, S, g” — not as a mood.
- Same evaluator online and offline if you already have one
- Abort runs counted, not dropped from the average
- Tool vendor invoices mapped to
job_type, not a junk drawer
Decision list when the pack is in front of you:
- p50 N ≤ 3 and g is a digest → tax is probably S (catalog). Shrink tools before you panic about loops.
- Slope of input tokens vs turn is steep → tax is g. Truncate before you cap.
- p95 N is 4× p50 N → the average is lying. Budget the tail or the tail is the product.
- Cache-read share ≈ 0 on turn 2+ → you are rebuying S. Fix the prefix before you buy a smaller model.
- Cost per success >> cost per run → failures or rewrites are the business. Do not celebrate cheap crashes.
If p95 N is 4× p50 N, your average is lying. Budget for the tail or the tail is the product.
How do I cut the tax without killing the job?
Cut g first (quadratic), then S (every turn), then N (caps and better tool choice), then model pin. That order is the opposite of what slide decks recommend.
| Move | Cuts | Risk if you overdo it |
|---|---|---|
| Truncate tool results to a digest | g | Lost fields the next decision needed |
| Scratchpad / state object outside the transcript | g | Stale state if you forget to update it |
| Allowlist tools per job | S and bad tool choice | Missing a tool you actually need |
| Parallel independent calls | N | Coupled calls you forced in parallel |
| Turn + token budget | N | Early abort on hard tickets — that can be correct |
| Cache S | Effective prefix $ | Broken prefix hashing, 0% hits |
| Demote to workflow + one LLM step | N and S | Long-tail exceptions you refused to see |
| Smaller / cheaper pin | Rate | More steps, worse tool choice — measure |
Numbered cutover that does not require a platform rewrite:
- Log per-turn tokens for one job (yesterday).
- Cap N at p95 + a small buffer. Send overflow to
escalate, not to “one more try.” - Wrap every tool with a max-bytes digest. Keep raw in blob storage keyed on
run_id. - Drop unused tools from the catalog. If MCP, do not load the whole registry.
- Put cache breakpoints on system + tools. Confirm cache-read tokens on turn 2.
- Re-run the 30-trace pack. Compare E, cost per success, pass rate, escalate rate.
- Only then A/B a cheaper pin on the same pack. If N jumps, you did not save money.
If pass rate holds only because you loosened the evaluator, you did not cut the tax. You hid it.
When is a workflow cheaper than an agent?
When you can name the next call before the payload arrives. Then a graph plus one schema-checked LLM step beats a loop that re-plans the same three HTTP calls. That is the whole point of agent loop vs LLM workflow step, and Anthropic’s “simplest solution” line in Building effective agents. LangGraph will let you mix both in one graph. Defaulting the entire job to a loop is how chatbot-demo math dies.
| Signal | Prefer | Cost reason |
|---|---|---|
| < ~10 stable branches | Workflow + LLM step | N is 1–2 by design |
| Next tool known from a field | n8n IF + HTTP | Zero planning tax |
| Long-tail exceptions, crisp pass/fail | Bounded agent on the exception lane | Pay the loop only when the graph misses |
| Irreversible writes, mushy taste | Human + checklist | Tokens are not the expensive part |
| You cannot write criteria | Do not build | You will pay for wander and still argue about quality |
Hybrid that usually wins: n8n (or equivalent) owns trigger, auth, retries, writes, and SLA. A bounded loop owns only the slice where the path varies. The loop returns a result package. The workflow writes.
- Every edge you can draw is drawn
- The loop has max turns and a digest policy
- Writes happen in the workflow, not in a free tool
- You compared cost per success on the same 30 cases, both shapes
Bake-off that ends the argument:
- Freeze 30 real cases.
- Run A: workflow + one structured LLM step (classifier / extract / draft).
- Run B: current agent loop, same pin, same tools allowed.
- Record pass, escalate, N, E, cost per success.
- Graduate to B only if A fails the long tail and B’s cost per success is a number you would defend to finance.
If the agent’s p50 N is 2 and both calls are “extract then done,” you bought a control plane for a node. Put the node back.
What spend guardrails belong in the harness?
A spend guardrail is code, not a polite system-prompt sentence. The operating manual’s stack already names cost + kill switches as a layer. This spoke is the numbers those switches should see.
| Guardrail | Fires when | Terminal |
|---|---|---|
| Max model turns | turn_index ≥ cap | abort |
| Max input tokens per turn | Prefill exceeds band (runaway dump) | abort or force digest |
| Max spend per run | Provider usage × your rate card | abort |
Max spend per job_type / hour | Burst | Shed load, then page |
| Tool result byte cap | Wrapper truncates; over-cap is an error class | Continue with digest or escalate |
| Duplicate tool+args | Same call twice | Deny + log; do not rebill the novel |
| Evaluator fail budget | Too many revisions | escalate |
| Policy deny | Illegal tool | deny — not a retry |
OpenAI’s agent-eval material treats the trace (model calls, tool calls, guardrails) as the unit you grade, not a single completion. Grade spend on that same object. If the policy service is down, fail closed. A loop that cannot read its budget will not honor a prompt that says “be economical.”
Checklist you should be able to demo in staging:
- Forced infinite-tool loop hits max turns
- Forced 50k-character tool result hits byte cap
- Dashboard shows abort reason
budgetvspolicyvseval - Rate card is dated; cache tokens billed at cache rates, not at full input
- Human approval path does not reset N to zero without logging the spend already burned
If the only budget is “we’ll watch the invoice,” you will watch it after the weekend job.
What should a week of diagnosis look like?
A week is enough to name the tax on one job. It is not enough to rebuild the platform. Pair it with the five-day proof spike in agent pilot scope: one sentence, real data, a cage, a score. Cost instrumentation is part of the cage, not a phase-two luxury.
| Day | Output | Done when |
|---|---|---|
| 1 | Job sentence + demo trace vs prod trace | You can say N, tools, S for both |
| 2 | Per-turn token chart on 30 runs | Slope is visible |
| 3 | Digest wrappers + turn cap in staging | Kill switch trips on purpose |
| 4 | Cache prefix verified (usage fields) | Turn 2+ shows cache reads |
| 5 | Same 30 cases re-run | Cost per success and pass rate side by side |
| 6–7 | Shape decision | Loop stays, hybrid, or workflow — written |
Skip if you only have a week: model bake-offs, multi-agent diagrams, “connect all of MCP,” custom cost dashboards with seven colors. Do not skip: N cap, digest, allowlist, abort reason, cost per success on one job_type.
When it is not worth doing yet: there is no production job, only a playground. Instrument when the loop can call a tool you care about. Until then the chatbot demo is the product, and it should stay billed like one.
| End of week | Ship | Rescope | Park |
|---|---|---|---|
| Slope down, cost per success in band, pass holds | Keep the loop, freeze the cap | — | — |
| Slope down only after you demoted to workflow | — | Workflow + LLM step is the product | — |
| Cannot get traces, no kill switch, tools still unbounded | — | — | Do not scale volume |
| Pass up, cost per success up more | — | Tighten g / N before celebrating | — |
Spurlock Studios will not quote a fake “average agent is 10× chat.” Fill the worksheet. If someone sells you 10× without N, S, and g, they are selling a slide.
FAQ
Why is my agent 10× more expensive than the chatbot demo?
Because the demo is one completion and the agent is a loop that re-bills a larger prefix plus growing tool results on every turn. “10×” is the complaint that shows up when N, S, and g are unbounded compared to that demo — not a universal measured tax. Pull both traces and fill the turns × tools × context worksheet before you quote a multiplier.
How do I measure whether the planning tax is actually shrinking?
Chart input tokens against turn index and watch p50 N, S, g, cache-read share, and cost per success on a frozen 30-case pack after each change. If N and g fall and pass rate plus escalate rate hold, the tax shrank. If dollars fall only because you sampled easier jobs or loosened the evaluator, it did not.
What usually fails first when teams try this?
Untruncated tool dumps and no turn cap, often on a catalog that loaded every MCP tool “for later.” The second miss is comparing cost per request to a playground chat that never called tools. The third is swapping to a cheaper pin that takes more steps and loses on the triangle.
How long does this take to show results?
A slope chart and a turn cap are usually days, not a quarter, if traces already exist. Digest wrappers show up on the next run. Cache hits are visible as soon as usage fields show cache-read tokens. Cost-per-success movement needs a frozen case pack; do not wait for month-end to see whether the tax moved.
What should I skip if I only have a week?
Skip multi-agent rewrites and vendor bake-offs. Do not skip: export demo vs prod traces, cap N, truncate tool results, shrink the allowlist, and compute cost per success on one job. If you cannot trip abort in staging, spend the week on the kill switch, not on a new model card.
When is this not worth doing yet?
When there is no loop in production — only a chatbot — there is no planning tax to hunt. The moment the agent can call tools on real tickets, the demo comparison is already a lie and the worksheet is worth filling. If you cannot write pass/fail criteria, pause the agent spend and write criteria; wander is the most expensive way to discover you had no job.
CTA
The demo was one completion. Price the loop.
Read the Agentic Systems Operating Manual, then use agentic or book an agentic pilot.
What questions does this article answer?
- Why is my agent 10× more expensive than the chatbot demo?
- Because the demo is one completion and the agent is a loop that re-bills a larger prefix plus growing tool results on every turn. “10×” is the complaint that shows up when N, S, and g are unbounded compared to that demo — not a universal measured tax. Pull both traces and fill the turns × tools × context worksheet before you quote a multiplier.
- How do I measure whether the planning tax is actually shrinking?
- Chart input tokens against turn index and watch p50 N, S, g, cache-read share, and cost per success on a frozen 30-case pack after each change. If N and g fall and pass rate plus escalate rate hold, the tax shrank. If dollars fall only because you sampled easier jobs or loosened the evaluator, it did not.
- What usually fails first when teams try this?
- Untruncated tool dumps and no turn cap, often on a catalog that loaded every MCP tool “for later.” The second miss is comparing cost per request to a playground chat that never called tools. The third is swapping to a cheaper pin that takes more steps and loses on the triangle.
- How long does this take to show results?
- A slope chart and a turn cap are usually days, not a quarter, if traces already exist. Digest wrappers show up on the next run. Cache hits are visible as soon as usage fields show cache-read tokens. Cost-per-success movement needs a frozen case pack; do not wait for month-end to see whether the tax moved.
- What should I skip if I only have a week?
- Skip multi-agent rewrites and vendor bake-offs. Do not skip: export demo vs prod traces, cap N, truncate tool results, shrink the allowlist, and compute cost per success on one job. If you cannot trip abort in staging, spend the week on the kill switch, not on a new model card.
- When is this not worth doing yet?
- When there is no loop in production — only a chatbot — there is no planning tax to hunt. The moment the agent can call tools on real tickets, the demo comparison is already a lie and the worksheet is worth filling. If you cannot write pass/fail criteria, pause the agent spend and write criteria; wander is the most expensive way to discover you had no job.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
AI Agents Agentic Systems: An Operating Manual for Multi-Agent Work That Ships
An agentic system is evaluators, policy gates, sandboxes, and kill switches — not a chat window — so production tool work survives contact with real data.
AI Agents Why Agent Demos Die in Production: Control Gaps, Not Model IQ
Demo success proves a staged happy path. Production fails when the control loop — schemas, auth, evaluators, and kill switches — was never in the harness.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.