Spurlock Studios
Contact
Share LinkedIn X
A small text-file card with no glyphs. Thesis: AGENT 10 MORE EXPENSIVE THAN.

Your agent looks 10× more expensive than the chatbot demo because it is a different product with a different bill. The demo is one model completion: a system prompt, a user message, an answer. The agent is a loop. Each turn re-sends the system prompt, the tool schemas, the growing transcript, and the last tool dump so the model can decide what to do next. Operators feel that as “10×.” That ratio is a complaint to diagnose on your traces. It is not a measured industry tax, and this post will not invent one.

This spoke sits under the Agentic Systems Operating Manual. Shape choice — workflow versus loop — is owned by agent loop vs LLM workflow step. If you still need to prove one job on real data before you argue about invoices, scope it with agent pilot scope.

The short answer

  • A chatbot demo bills one turn, zero tools, a short prefix. An agent bills N turns × (static prefix + growing context), plus tool APIs, plus any evaluator rounds.
  • Anthropic’s Building effective agents (Dec 19, 2024) is explicit: agentic systems “often trade latency and cost for better task performance,” and autonomy means “higher costs, and the potential for compounding errors.” They do not publish a universal 10×.
  • Input token-exposures are not “tokens × turns.” Each turn re-reads what every prior turn left behind. That triangular term is the planning/tool-loop tax.
  • Prompt caching discounts a stable prefix. It does not flatten fat tool results you append every step.
  • Cut the tax with turn caps, truncated tool digests, a smaller allowlist, and demoting known paths to a workflow. Swap models last.

Why is the chatbot demo a different bill?

The demo you showed finance is almost always a single-turn chat completion. One request. One response. No tools, or a fake “search” that never ran. The system prompt is a paragraph. The user message is one question. That bill is honest for that product. It is not a forecast for an agent that can call CRM, search, and a code runner until it feels done.

n8n’s AI Agent node is the other product: connect a chat model and tools, and the node “decides which tools to call to complete a task.” One workflow execution is not one model call. LangGraph draws the same line in workflows vs agents: predetermined code paths versus a model that directs process and tool usage.

What finance sawWhat production ranWhy the bills diverge
Playground chat, one questionJob with tools and retriesDemo never paid for a loop
System prompt ~a paragraphSystem + full tool JSON schemasPrefix is 3–10× before the first tool
No tools, or mocked searchLive CRM / search / sandboxTool dumps become next-turn context
You clicked StopNo turn cap, no kill switchN is unbounded
“It answered”Evaluator + revisionsExtra model calls per job
Cost per messageCost per jobDenominator changed

If you cannot point to the demo’s trace and show N = 1, tools = 0, stop comparing it to the agent. You are comparing a completion to a control plane.

Checklist before you quote a multiplier in a meeting:

  • Demo trace exported: one request, one response, token usage object
  • Agent trace exported for the same job sentence
  • Turn count, tool-call count, and input tokens per turn on both
  • Same model pin, or you admit the bake-off is contaminated
  • You are comparing cost per job, not cost per API request

The demo is a poster. The agent is the machine the poster was selling.

What is the planning tax?

Planning tax is every token you spend so the model can choose the next action instead of you encoding that choice in a graph. It is not “the model thinking too much” as a vibe. It is billed context: the job contract, the tool catalog, the scratch of prior plans, and the “what should I do now?” completion on every turn.

Anthropic’s essay treats agents as LLMs using tools from environmental feedback in a loop. That loop is the tax. A workflow with one structured-output call pays planning once. An agent pays it on turn 1, turn 2, and turn 12, and turn 12 still has to carry the catalog that made turn 1 possible.

Planning ingredientWhere it livesPaid again every turn?
Job / policy textSystem promptYes, unless cached
Tool names, descriptions, JSON schemasTools array / MCP catalogYes, unless cached
“Think about the next step” completionAssistant outputYes — that output becomes later input
Prior plan fragmentsTranscriptYes, until you drop them
Evaluator rubric (if in-band)Extra messagesIf you stuffed it into the worker

MCP makes this worse when you are sloppy. A host that dumps a giant tool catalog into every request is paying a planning tax for tools the job will never call. Allowlist the job. Do not “connect the registry and see.”

Procedure to separate planning tax from work tax on one trace:

  1. Sum input tokens that are system prompt + tool definitions. Call that S (static prefix).
  2. Sum assistant tokens that are plans, not tool arguments or the final artifact. Call that P.
  3. Sum tool-result tokens. Call that R.
  4. If S + P dominates R on a job that only needed three known HTTP calls, you bought an agent for a workflow.
  5. If R dominates and still grows each turn, you have a tool-loop tax, next section.

Planning is the fee for uncertainty. If the next action is nameable before the payload arrives, stop paying it.

What is the tool-loop tax?

Tool-loop tax is the re-billing of tool results (and the assistant turns that produced them) on every subsequent model call. The first search result was useful at step 2. At step 9 you are still paying to re-read the raw JSON unless you truncated it to a digest.

This is ordinary API physics, not a secret vendor trick. Chat Completions / Messages-style APIs send the full conversation you provide. If your harness appends tool messages verbatim, turn k includes turns 1…k−1. Independent tools called in sequence become a stack of blobs. Parallel tool calls in one turn are cheaper than the same calls as four separate turns, because you pay the prefix once for that round.

Loop behaviorWhat gets appendedTax shape
One tool, digest of 80 tokensSmall RManageable
One tool, raw 8,000-token API bodyFat RQuadratic in later turns
Four sequential dependent toolsFour round tripsFour prefix payments
Four independent tools, parallelOne round tripOne prefix payment
Failed tool, full error pagePoison + sizeRetry multiplies both
MCP resource dumped as contextFiles you did not ask to keepPrefix + R both jump

n8n is honest about the shape: the agent uses tools to retrieve and act, and it “determine[s] which tool to use depending on the task.” That determination is a model call. The observation is another message. Repeat until done — or until you forgot to define done.

Decision list:

  1. Does the next tool depend on the last result? Sequential is fine; still digest the last result.
  2. Can two tools run with the same inputs? Parallel that turn.
  3. Is the raw body needed after the next decision? If no, store it outside the prompt and keep a 40–80 token digest in the transcript.
  4. Did the tool fail? Put a short error class in context, not the vendor’s HTML 500 page.

The loop is how agents earn their keep on messy jobs. Unbounded raw tool I/O is how they earn their reputation with finance.

Cost model: turns × tools × context

Do not multiply “tokens per demo message × expected tools.” That is the naive linear model, and it is wrong the moment the transcript grows. A diagnostic worksheet:

Input token-exposures (uncached, illustration — not a bill):

E = N × S + g × N × (N − 1) / 2

  • N = model turns in the run (each time you call the model, including “I’ll call a tool” turns)
  • S = static prefix tokens: system + tool schemas (+ frozen policy)
  • g = average new tokens you append per turn after the first (tool digest or dump + assistant)

Output tokens are roughly linear in N. They matter, but for tool-heavy agents the compounding line is usually input re-reads. This table uses labeled toy numbers so you can see how an operator complaint like “10×” can appear without anyone measuring a universal tax. Replace S, g, and N with your p50 traces. Do not quote these rows as Spurlock Studios invoices or as industry benchmarks.

ShapeN (turns)Tools calledS (prefix)g (added / turn)E (input token-exposures)E vs 1,200-token chatbot
Chatbot demo101,20001,2001× (baseline)
Agent, first look, no tool104,00004,0003.3×
Agent, short loop334,00060013,80011.5×
Agent, typical write job564,00060026,00021.7×
Agent, wander8104,00080054,40045.3×
Same wander, fat dumps8108,0002,000120,000100×

Read the 3-turn row slowly. On these assumptions, input-token exposure already crosses the neighborhood of 10× versus a 1,200-token chatbot prefix. Change S or g and the crossing move. That is why “my agent is 10× the demo” is a useful ticket and a terrible statistic.

Second axis: tools. Tools change g (result size) and often N (extra turns to call them). A catalog of 40 MCP tools inflates S before anyone calls one.

LeverHitsDirection
Turn capNLinear on S, quadratic on g
Truncate tool resultsgQuadratic savings on later turns
Shrink tool catalogSLinear on every turn
Parallel independent toolsNFewer prefix payments
Prompt cache on SEffective SDiscount, not a flatten of g
Evaluator as separate callExtra NAdds a line; do not hide it in the worker

Procedure to fill the worksheet from last week’s production:

  1. Pick one job_type. Ignore mixed dashboards.
  2. For 30 runs, record N, tool-call count, input tokens per turn, output tokens, cache read/write tokens if the provider sends them.
  3. Estimate S from turn-1 input minus the user payload.
  4. Estimate g as median (input tokens at turn k − input tokens at turn k−1), after subtracting cache-only artifacts if you must.
  5. Compute E from the formula and from summed provider input_tokens. If they disagree a lot, your harness is dropping or injecting context you did not model.
  6. Divide total model spend for those runs by successes, not by runs. Cheap failures flatter cost per run.

Worked loop on the same labeled assumptions (S = 4,000, g = 600, N = 5). Turn k input ≈ S + g×(k − 1). These are token counts, not dollars.

TurnTools this turnContext the model re-readsInput tokens this callCumulative E
10–1Prefix + user4,0004,000
21Prefix + prior assistant + tool digest4,6008,600
31Everything above + last digest5,20013,800
41Growing tail5,80019,600
51Growing tail6,40026,000

Five requests did not cost 5 × the first request. Request 5 is already 1.6× request 1, and the sum is 6.5× request 1. Compare that sum to a 1,200-token chatbot completion and you are in the “why is this ~20×” conversation — still a worksheet, still not a published benchmark.

If you cannot fill N, S, and g, you do not have a cost model. You have a credit-card alert.

Why does cost grow faster than turn count?

Because turn k re-reads turns 1…k−1. Sum of 1 through N is N(N+1)/2. The g term in the worksheet is that triangle. A 16-turn run is not “twice an 8-turn run.” If g is stable, the triangular piece is about four times larger, and you still pay N×S.

OpenAI’s tokenizer note is a sizing heuristic, not a bill: about four characters per token for common English, roughly 100 tokens ≈ 75 words. Use it to sanity-check a 20,000-character CRM payload you were about to paste into the transcript. Do not use it to forecast invoices.

If you only watch…You will think…What you missed
Requests / minuteTraffic is flatN per job climbed
Tokens / requestEach call is “about the same”Later calls are fatter
Output tokens“We asked it to be brief”Input re-reads dominate
Demo vs agent request count“It’s only 5× the calls”Each call is not the demo’s size
Monthly $“Models got cheaper”N and g grew into the discount

Plot prefill / input tokens against step number for a handful of long runs. If that line slopes up, you are paying the square. If it is roughly flat, your harness is already digesting or resetting context. Flat per-step input is the difference between a job you can run and a job you can only demo.

  • Per-turn input tokens logged, not only totals
  • Step index on the span
  • Tool result bytes before and after truncation
  • A chart someone in finance can see without opening JSON

Output tokens do not save you here. Five turns at 400 output tokens is 2,000 generated tokens against 26,000 input-exposures on the worksheet above. Output is often priced higher per token; it is still usually the smaller pile when tools dump JSON. “Answer briefly” is a latency tweak, not a cost strategy, until g is small.

NN×S (S=4,000)Triangle g×N×(N−1)/2 (g=600)EE / E(N=4)
416,0003,60019,6001.0×
832,00016,80048,8002.5×
1664,00072,000136,0006.9×

Doubling turns from 8 to 16 is not 2× tokens. The triangle takes over. That is why a “quick extra retry” policy is a finance event.

Slope up is the tax. Argue with the slope, not with the adjective “expensive.”

What else gets billed besides the worker loop?

The worker’s N×S+g is the line operators notice first. Production jobs grow other meters that the chatbot demo never had.

MeterWho bills itDemo had it?Hide risk
Worker model callsLLM providerYes, onceYou compare this to the demo and stop
Evaluator model callsLLM providerNo“Quality” looks free
Embeddings / retrievalLLM or vector hostMaybeRAG on every turn
Paid search / enrich APIsTool vendorNoPer-call, not tokens
Code-exec / browser sandboxTool vendorNoMinutes × machine
Prompt-cache writesLLM providerRarelyFirst-turn premium
Human rewrite timePayrollYou were the demoCost per clean ship
Retries / 429 stormsAll of the aboveNoMultiplies N

Braintrust’s public arithmetic is the one to steal for the denominator: cost per success is cost per task divided by success rate, and cost per resolved request only counts work that clears quality gates. A config that succeeds one-in-six pays for roughly six attempts per win. I will not paste their study’s dollar figures as if they were yours.

Procedure for a single job type:

  1. Sum worker model cost.
  2. Sum evaluator model cost (separate pin, separate spans).
  3. Sum metered tools.
  4. Add human minutes × loaded rate for rewrites on the same jobs.
  5. Divide by successes, then by clean ships.
  6. If evaluator cost is a rounding error, either your evaluator is mechanical (good) or you are not actually evaluating (bad).

When “10× vs the demo” is the wrong diagnosis, the model loop is fine and a different meter moved. Check these before you rewrite prompts.

SymptomLoop looksActual driverFirst move
Token $ flat, vendor $ upN=3, small gPaid search / enrich per callCap tool calls; cache results
Token $ up, N=1One huge completionYou stuffed the CRM dump into a “chatbot”That is not an agent tax; that is a prompt dump
Model $ fine, payroll upPass rate highHumans rewrite every artifactCost per clean ship
Embeddings line item newN unchangedRetrieval on every turnRetrieve once; pass a digest
Weekend spikeFew jobs, huge NOne runaway without abortKill switch, then postmortem

The chatbot demo had meter 1, once. Your agent has a stack. Compare stacks.

Does prompt caching erase the tax?

No. Caching discounts a stable prefix. The growing tail is still full price unless you keep it small.

Anthropic’s prompt caching docs (verify live; this class of number moves) bill cache reads at 0.1× base input, 5-minute cache writes at 1.25×, and 1-hour writes at 2×. OpenAI’s prompt caching guide, as of this writing, uses the same 1.25× write / 0.1× read shape for GPT-5.6 and later when you place breakpoints; older caching modes on that page are model-dependent. I am not pasting per-model dollar tables here. They rot. Read the live page before you build a finance model on them.

What you cacheHelpsDoes not help
System prompt + tool schemas (S)Every later turn’s prefixFat tool dumps after the breakpoint
Frozen few-shot policyStable jobsPer-ticket CRM JSON
Nothing, hoping the vendor “just caches”OpenAI automatic modes, if you qualifyAnthropic-style explicit breakpoints you never set
The entire growing transcriptOnly the unchanged leading tokensAnything you prepend that busts the prefix

Caching is a discount on S. Truncation is a cut to g. Turn caps cut N. Teams that “turn on caching” and still paste 8k-token search hits are applying a 90% discount to the wrong pile.

Checklist:

  • Tool schemas sit in a cached prefix, not below a timestamp that changes every request
  • Usage shows cache-read tokens on turn 2+, not only on marketing slides
  • User payload and tool results sit after the breakpoint
  • TTL matches your cadence (slow human-in-the-loop vs tight loop)
  • You still have a turn cap, because a cached runaway is still a runaway

Same N=5 worksheet with S cached after turn 1, using the 1.25× write / 0.1× read multipliers from the vendor pages above. Tail tokens stay full price. Input-equivalent (not a $ invoice):

PieceTokensMultiplierInput-equivalent
S cache write (turn 1)4,0001.25×5,000
S cache reads (turns 2–5)4 × 4,0000.1×1,600
Growing tail (triangle)6,0001.0×6,000
Cached total——12,600
Uncached E from earlier——26,000

Caching roughly halved the prefix tax on this toy run. It did nothing to the 6,000 tail tokens. If g is 2,000 because you paste search hits, the tail dominates and the coupon looks disappointing. Verify with cache_read / cache_write on the usage object. If those fields are zero, you are paying list price for S every turn.

Cache is a coupon. It is not a kill switch.

What usually blows the bill first?

The failure mode is not “the model got pricier.” It is unbounded N plus untruncated R, usually after a demo that never hit either.

I have spent 20,000+ hours on agentic systems and shipped 500+ automations into production. The invoice surprise is almost always the same plot: playground chat looks cheap, someone wires ten tools, a search tool returns a novel, retries sit on “until success,” and there is no abort on token budget. Nobody fabricated a 10× study. The card just moved.

FailureWhat the trace showsWhat you do instead
No turn capN = 20, 40, 80 on one ticketHard max turns → abort with reason
Raw tool dumpSingle R of thousands of tokens reused every stepDigest; store raw outside the prompt
Giant MCP catalogS already huge at turn 1Allowlist 2–5 tools for the job
Retry without budgetSame tool, same args, three fat errorsError class + cap; then escalate
Evaluator inside the workerRubric tokens on every plan turnSeparate evaluator call, mechanical checks first
Demo model pin ≠ prod pinFlagship on the loop, small model in the screenshotSame pin or admit the comparison is theater
Cost per request dashboardLooks “fine” while jobs get longerCost per success by job_type

Worked pattern (labeled, not a client bill): a support-looking agent calls search, gets a 6,000-token hit, then calls CRM, then “thinks,” then searches again. N=8. g stays fat because nobody truncated. The chatbot demo for “what is our return policy?” was N=1, S≈1,200, tools=0. Finance asks why prod is “about 10×.” On a worksheet like the one above, you may already be past that at N=3. At N=8 you are arguing about a different sport.

Do not fix this with a hotter model. Hotter models that take more steps can lose even if the per-token rate is lower. Steps enter the triangle. Price enters linearly.

n8n-specific version of the same blow-up: a Chat node or a single Basic LLM Chain in a playground workflow is N=1. An AI Agent with five tool sub-nodes can sit in one execution and still be N=8 model calls. The execution list looks like “1 run.” The provider bill looks like a conversation. If you estimate cost from n8n execution count, you will always lose the argument with finance.

n8n shapeWhat “1 execution” containsCost model
Chat / LLM node, no toolsOne completionChatbot-demo math
HTTP + IF graph, one LLM extractOne bounded callWorkflow + LLM step
AI Agent, 2 allowlisted tools, max iterations setBounded loopFill N, S, g
AI Agent, MCP dump, default iterationsUnbounded loopThe surprise

If you cannot trip the kill switch in staging with a fake infinite loop, you do not have a kill switch. You have a comment in the prompt.

How do I measure this on real traces?

Measure per job, with N, tools, tokens, cache, and terminal reason on one timeline. Provider token charts are traffic. Microsoft Foundry’s agent-tracing note is blunt about why chat logs fail here: many steps, changing order, long payloads, nesting — capture inputs, outputs, tool usage, token consumption, duration. Langfuse’s rule is the same spirit: prefer ingested provider usage over homemade tokenizer math.

FieldWhyFake if missing
job_typeMix hides the expensive shapeBlended $ / run
turn_indexSlope of input tokensTotals only
tool_name + result bytesg diagnosis“Tools were used”
input_tokens / output_tokensProvider usageYour guess
cache_read / cache_writeWhether S is actually discountedCaching folklore
terminal (done / escalate / abort)Denominator for successEndless “running”
revision countGrind-to-greenPass rate theater
eval_costHidden second modelWorker-only bill

Procedure (one afternoon on a live job):

  1. Sample 30 production traces of one sentence-sized job.
  2. Chart input tokens vs turn_index. Note p50 N and p95 N.
  3. List tool result byte sizes. Mark anything > ~1–2k tokens that was re-sent.
  4. Compute cost per run, cost per pass, cost per success.
  5. Diff the playground demo trace for the same question, if it still exists.
  6. Write the gap as “N, S, g” — not as a mood.
  • Same evaluator online and offline if you already have one
  • Abort runs counted, not dropped from the average
  • Tool vendor invoices mapped to job_type, not a junk drawer

Decision list when the pack is in front of you:

  1. p50 N ≤ 3 and g is a digest → tax is probably S (catalog). Shrink tools before you panic about loops.
  2. Slope of input tokens vs turn is steep → tax is g. Truncate before you cap.
  3. p95 N is 4× p50 N → the average is lying. Budget the tail or the tail is the product.
  4. Cache-read share ≈ 0 on turn 2+ → you are rebuying S. Fix the prefix before you buy a smaller model.
  5. Cost per success >> cost per run → failures or rewrites are the business. Do not celebrate cheap crashes.

If p95 N is 4× p50 N, your average is lying. Budget for the tail or the tail is the product.

How do I cut the tax without killing the job?

Cut g first (quadratic), then S (every turn), then N (caps and better tool choice), then model pin. That order is the opposite of what slide decks recommend.

MoveCutsRisk if you overdo it
Truncate tool results to a digestgLost fields the next decision needed
Scratchpad / state object outside the transcriptgStale state if you forget to update it
Allowlist tools per jobS and bad tool choiceMissing a tool you actually need
Parallel independent callsNCoupled calls you forced in parallel
Turn + token budgetNEarly abort on hard tickets — that can be correct
Cache SEffective prefix $Broken prefix hashing, 0% hits
Demote to workflow + one LLM stepN and SLong-tail exceptions you refused to see
Smaller / cheaper pinRateMore steps, worse tool choice — measure

Numbered cutover that does not require a platform rewrite:

  1. Log per-turn tokens for one job (yesterday).
  2. Cap N at p95 + a small buffer. Send overflow to escalate, not to “one more try.”
  3. Wrap every tool with a max-bytes digest. Keep raw in blob storage keyed on run_id.
  4. Drop unused tools from the catalog. If MCP, do not load the whole registry.
  5. Put cache breakpoints on system + tools. Confirm cache-read tokens on turn 2.
  6. Re-run the 30-trace pack. Compare E, cost per success, pass rate, escalate rate.
  7. Only then A/B a cheaper pin on the same pack. If N jumps, you did not save money.

If pass rate holds only because you loosened the evaluator, you did not cut the tax. You hid it.

When is a workflow cheaper than an agent?

When you can name the next call before the payload arrives. Then a graph plus one schema-checked LLM step beats a loop that re-plans the same three HTTP calls. That is the whole point of agent loop vs LLM workflow step, and Anthropic’s “simplest solution” line in Building effective agents. LangGraph will let you mix both in one graph. Defaulting the entire job to a loop is how chatbot-demo math dies.

SignalPreferCost reason
< ~10 stable branchesWorkflow + LLM stepN is 1–2 by design
Next tool known from a fieldn8n IF + HTTPZero planning tax
Long-tail exceptions, crisp pass/failBounded agent on the exception lanePay the loop only when the graph misses
Irreversible writes, mushy tasteHuman + checklistTokens are not the expensive part
You cannot write criteriaDo not buildYou will pay for wander and still argue about quality

Hybrid that usually wins: n8n (or equivalent) owns trigger, auth, retries, writes, and SLA. A bounded loop owns only the slice where the path varies. The loop returns a result package. The workflow writes.

  • Every edge you can draw is drawn
  • The loop has max turns and a digest policy
  • Writes happen in the workflow, not in a free tool
  • You compared cost per success on the same 30 cases, both shapes

Bake-off that ends the argument:

  1. Freeze 30 real cases.
  2. Run A: workflow + one structured LLM step (classifier / extract / draft).
  3. Run B: current agent loop, same pin, same tools allowed.
  4. Record pass, escalate, N, E, cost per success.
  5. Graduate to B only if A fails the long tail and B’s cost per success is a number you would defend to finance.

If the agent’s p50 N is 2 and both calls are “extract then done,” you bought a control plane for a node. Put the node back.

What spend guardrails belong in the harness?

A spend guardrail is code, not a polite system-prompt sentence. The operating manual’s stack already names cost + kill switches as a layer. This spoke is the numbers those switches should see.

GuardrailFires whenTerminal
Max model turnsturn_index ≥ capabort
Max input tokens per turnPrefill exceeds band (runaway dump)abort or force digest
Max spend per runProvider usage × your rate cardabort
Max spend per job_type / hourBurstShed load, then page
Tool result byte capWrapper truncates; over-cap is an error classContinue with digest or escalate
Duplicate tool+argsSame call twiceDeny + log; do not rebill the novel
Evaluator fail budgetToo many revisionsescalate
Policy denyIllegal tooldeny — not a retry

OpenAI’s agent-eval material treats the trace (model calls, tool calls, guardrails) as the unit you grade, not a single completion. Grade spend on that same object. If the policy service is down, fail closed. A loop that cannot read its budget will not honor a prompt that says “be economical.”

Checklist you should be able to demo in staging:

  • Forced infinite-tool loop hits max turns
  • Forced 50k-character tool result hits byte cap
  • Dashboard shows abort reason budget vs policy vs eval
  • Rate card is dated; cache tokens billed at cache rates, not at full input
  • Human approval path does not reset N to zero without logging the spend already burned

If the only budget is “we’ll watch the invoice,” you will watch it after the weekend job.

What should a week of diagnosis look like?

A week is enough to name the tax on one job. It is not enough to rebuild the platform. Pair it with the five-day proof spike in agent pilot scope: one sentence, real data, a cage, a score. Cost instrumentation is part of the cage, not a phase-two luxury.

DayOutputDone when
1Job sentence + demo trace vs prod traceYou can say N, tools, S for both
2Per-turn token chart on 30 runsSlope is visible
3Digest wrappers + turn cap in stagingKill switch trips on purpose
4Cache prefix verified (usage fields)Turn 2+ shows cache reads
5Same 30 cases re-runCost per success and pass rate side by side
6–7Shape decisionLoop stays, hybrid, or workflow — written

Skip if you only have a week: model bake-offs, multi-agent diagrams, “connect all of MCP,” custom cost dashboards with seven colors. Do not skip: N cap, digest, allowlist, abort reason, cost per success on one job_type.

When it is not worth doing yet: there is no production job, only a playground. Instrument when the loop can call a tool you care about. Until then the chatbot demo is the product, and it should stay billed like one.

End of weekShipRescopePark
Slope down, cost per success in band, pass holdsKeep the loop, freeze the cap——
Slope down only after you demoted to workflow—Workflow + LLM step is the product—
Cannot get traces, no kill switch, tools still unbounded——Do not scale volume
Pass up, cost per success up more—Tighten g / N before celebrating—

Spurlock Studios will not quote a fake “average agent is 10× chat.” Fill the worksheet. If someone sells you 10× without N, S, and g, they are selling a slide.

FAQ

Why is my agent 10× more expensive than the chatbot demo?

Because the demo is one completion and the agent is a loop that re-bills a larger prefix plus growing tool results on every turn. “10×” is the complaint that shows up when N, S, and g are unbounded compared to that demo — not a universal measured tax. Pull both traces and fill the turns × tools × context worksheet before you quote a multiplier.

How do I measure whether the planning tax is actually shrinking?

Chart input tokens against turn index and watch p50 N, S, g, cache-read share, and cost per success on a frozen 30-case pack after each change. If N and g fall and pass rate plus escalate rate hold, the tax shrank. If dollars fall only because you sampled easier jobs or loosened the evaluator, it did not.

What usually fails first when teams try this?

Untruncated tool dumps and no turn cap, often on a catalog that loaded every MCP tool “for later.” The second miss is comparing cost per request to a playground chat that never called tools. The third is swapping to a cheaper pin that takes more steps and loses on the triangle.

How long does this take to show results?

A slope chart and a turn cap are usually days, not a quarter, if traces already exist. Digest wrappers show up on the next run. Cache hits are visible as soon as usage fields show cache-read tokens. Cost-per-success movement needs a frozen case pack; do not wait for month-end to see whether the tax moved.

What should I skip if I only have a week?

Skip multi-agent rewrites and vendor bake-offs. Do not skip: export demo vs prod traces, cap N, truncate tool results, shrink the allowlist, and compute cost per success on one job. If you cannot trip abort in staging, spend the week on the kill switch, not on a new model card.

When is this not worth doing yet?

When there is no loop in production — only a chatbot — there is no planning tax to hunt. The moment the agent can call tools on real tickets, the demo comparison is already a lie and the worksheet is worth filling. If you cannot write pass/fail criteria, pause the agent spend and write criteria; wander is the most expensive way to discover you had no job.

CTA

The demo was one completion. Price the loop.

Read the Agentic Systems Operating Manual, then use agentic or book an agentic pilot.

FAQ

What questions does this article answer?

Why is my agent 10× more expensive than the chatbot demo?
Because the demo is one completion and the agent is a loop that re-bills a larger prefix plus growing tool results on every turn. “10×” is the complaint that shows up when N, S, and g are unbounded compared to that demo — not a universal measured tax. Pull both traces and fill the turns × tools × context worksheet before you quote a multiplier.
How do I measure whether the planning tax is actually shrinking?
Chart input tokens against turn index and watch p50 N, S, g, cache-read share, and cost per success on a frozen 30-case pack after each change. If N and g fall and pass rate plus escalate rate hold, the tax shrank. If dollars fall only because you sampled easier jobs or loosened the evaluator, it did not.
What usually fails first when teams try this?
Untruncated tool dumps and no turn cap, often on a catalog that loaded every MCP tool “for later.” The second miss is comparing cost per request to a playground chat that never called tools. The third is swapping to a cheaper pin that takes more steps and loses on the triangle.
How long does this take to show results?
A slope chart and a turn cap are usually days, not a quarter, if traces already exist. Digest wrappers show up on the next run. Cache hits are visible as soon as usage fields show cache-read tokens. Cost-per-success movement needs a frozen case pack; do not wait for month-end to see whether the tax moved.
What should I skip if I only have a week?
Skip multi-agent rewrites and vendor bake-offs. Do not skip: export demo vs prod traces, cap N, truncate tool results, shrink the allowlist, and compute cost per success on one job. If you cannot trip abort in staging, spend the week on the kill switch, not on a new model card.
When is this not worth doing yet?
When there is no loop in production — only a chatbot — there is no planning tax to hunt. The moment the agent can call tools on real tickets, the demo comparison is already a lie and the worksheet is worth filling. If you cannot write pass/fail criteria, pause the agent spend and write criteria; wander is the most expensive way to discover you had no job.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot