Cost Controls for Agent Fleets: Budgets, Caps, and Kill Switches
Control agent-fleet cost with runner-enforced caps, policy-versioned prompt caches, and state-based model routing — not a prompt that says to be frugal.
William Spurlock Founder — Spurlock Studios Updated 25 MIN
An agent without a budget is a blank check. A fleet without caps, cache keys, and a freeze button is how finance learns about the program from the card statement. Cost control is part of the runtime: per-run counters, revision and tool ceilings, model routing by state, and kill switches that abort with a reason code.
This spoke sits under the Agentic Systems Operating Manual. It owns the money path. Observability for agents owns the screen that shows cost next to the pass. Doctrine lives in the manual. Caps, caches, and routing live here.
The short answer
- Attach
budget_usdandbudget_tokensat intake. Decrement on every billable model and tool call. Refuse the next step when remaining budget cannot cover it. - Cap revisions, tool calls, retrievals, parallel children, and context assembly. Spirals look like “the agent is trying.” Ops sees a melting wallet.
- Route by state: cheap for narrow classify, mid for tool choice, flagship only where quality is the product. Score cost per pass, not cost per token.
- Cache the stable prefix. Key it with
doc_versionand evaluator criterion versions. Short TTL on knowledge-grounded answers. - Provider dashboards and monthly spend limits are backstops. The kill switch that freezes a
job_typelives in your runner.
What does cost control mean for an agent fleet?
Observation is a chart. Control is a refused transition. After 500+ automations and 20,000+ hours on agentic systems, the invoice surprises I see are almost never “the flagship model is expensive.” They are uncapped revisions, unbounded fan-out, and a shared key nobody owned.
You need four layers. Miss one and the others become theater.
| Layer | What it is | Failure if missing |
|---|---|---|
| Unit economics | Expected cost per successful job, agreed with the buyer | You optimize tokens and still lose money |
| Runtime budgets | Hard caps per run, per day, per tenant | A single loop writes the month |
| Shape controls | Max tool calls, revisions, retrievals, tokens in/out | “Trying” becomes a spiral |
| Kill switches | Automatic abort + page when a trip fires | You find out in the billing email |
If you only have dashboards, you have observation. If the runner can refuse act and freeze a job_type, you have control.
- Buyer signed a cost-per-success band, not a vibe
- Every run carries remaining budget from intake
- Caps are integers in config, not adjectives in a prompt
- Someone can freeze one job type without taking the fleet down
Anthropic’s own note on building effective agents starts in the same place: add multi-step loops only when a cheaper pattern fails. That is the cheapest cost control you will ever ship — not building a fleet for a known path.
How do you attach a token budget the runner can enforce?
A token budget is useless as a sticky note. Implement it as counters the runner checks before every model call and every paid tool call.
Practical scheme:
- Attach
budget_usdandbudget_tokensatintake. - Decrement after each billable call. Prefer provider usage fields. Estimate with a documented margin when the provider is silent.
- Refuse transition into
actorrevisewhen remaining budget is below the cost of the next step. - Land in
escalateorabortwith a catalog reason:budget_exhausted. - Carry remaining budget in the handoff package so child agents cannot each assume a full wallet.
| Counter | When it moves | Who may override |
|---|---|---|
budget_usd | After each model + paid tool | Budget owner, time-boxed |
budget_tokens | After each completion | Technical owner only |
revisions_used | After each evaluator fail → retry | Nobody mid-run |
tool_calls_used | After each tool span | Nobody mid-run |
children_spawned | After each fan-out | Hub only |
OpenAI’s spend limits are a monthly org or project backstop. When tracked spend hits the hard cap, requests return 429 with organization_spend_limit_exceeded or project_spend_limit_exceeded. Enforcement is not instantaneous — recorded spend can overshoot. That is why the per-run counter lives in your runner. By the time the provider 429s, the loop may have already written.
LiteLLM’s provider, model, and tag budgets are the same idea at the proxy: $N per day on a provider, a model, or a product: tag. Use them as a second fence, not the only fence. A proxy that cannot see your revision ceiling will still let one job grind.
- Usage ingested from the provider response, not guessed after the fact
- Estimates marked
estimated=trueon the span - Child allowances allocated from the parent, never a silent second wallet
- Staging keys cannot draw prod quota
Which caps stop revision and tool spirals?
Token price is the number people argue about. Spirals are the number that shows up on the card. A mid-tier model that revises eight times will beat a flagship that passes on the first evaluator.
| Cap | Typical starting integer | What it stops |
|---|---|---|
| Revision ceiling | 3 | Evaluator ping-pong |
| Max tool calls per run | 20 | Search-and-retry grind |
| Max retrieval calls | 8 | Over-retrieval as “memory” |
| Max parallel children | 1–3 | One job becoming a fleet |
| Context assembly | Hard truncate | Ancient scratch crowding the job contract |
| Fan-out | One hop, named types | Unbounded child graphs |
Encode the integers. Do not trust a prompt that says “be frugal.”
per_run_usd_max: <set with the buyer>
per_day_tenant_usd_max: <set with finance>
max_revisions: 3
max_tool_calls: 20
max_retrieval_calls: 8
on_budget_exhaust: escalate
on_day_cap: freeze_new_runs + alert
Those dollar fields are placeholders you fill from your unit economics. They are not a claimed saving and not a studio default that magically fits every job.
Publish the per-run cap next to the job contract so builders see the number while they prompt. Invisible caps get treated as suggestions. Visible caps shape design.
- Revision ceiling trips to
escalate, not a fourth silent retry - Tool and retrieval counters share the same budget object
- Fan-out requires a named child
job_typewith its own allowance - Context truncate prefers job contract + last failures over old scratch
How should model routing work by state?
Not every state needs the flagship. Routing is a table the runner consults, not a hope that the model will “pick a cheaper path.”
| State | Typical tier | Why |
|---|---|---|
intake classify | Small / cheap | Narrow schema output |
plan | Mid | Judgment, not essays |
act tool choice | Mid | Schema-constrained |
| Draft prose | Mid or high | Quality-sensitive |
evaluate mechanical | Code | Free |
evaluate judgment | Mid/high, short context | Accuracy over creativity |
| Fallback | Explicit mid, own multiplier | Availability, not a silent downgrade |
Measure cost per passing run on a golden set before you call a cheaper tier a win. A cheaper model that needs eight revisions can lose. Anthropic’s cost guide is blunt about this: compare models on cost per completed task, not per token. Their published figures are directional vendor benchmarks, not your invoice. Measure on your cases.
LiteLLM budget fallbacks will silently reroute to the next model in a chain once a per-model cap is exhausted. That is useful for availability. It is dangerous for quality if the golden set does not still pass on the fallback. Wire fallback as an explicit state transition with its own budget multiplier. Blind failover to a cheap model is how silent wrongness spikes while the cost chart looks healthy.
| Routing move | Allowed when | Forbidden when |
|---|---|---|
| State → cheaper tier | Golden-set pass rate holds inside the cost band | You have not scored the cheaper tier |
| Budget fallback | Fallback has its own cap + eval | The proxy swaps models with no receipt |
| Flagship only on hard branch | Advisor consult rate is measured | Every turn “just in case” |
| Batch / delayed tier | The job can wait | A user is staring at a spinner |
- Each state names a model id, not a family nickname
- Fallback is a logged transition, not a proxy surprise
- Cost-per-pass and pass rate ship in the same review
- “Slightly better, 4× cost” is a product decision, not an automatic ship
When does prompt caching pay — and when does it lie?
Agent loops resend the growing prefix every turn: system prompt, tool schemas, and prior messages. Anthropic’s prompt caching docs price that reality: 5-minute cache writes at 1.25× base input, 1-hour writes at 2×, cache reads at 0.1×. Default TTL is five minutes and refreshes on hit. Their cost-and-intelligence guide reports cache cutting their measured agent-loop cost by a factor of 2.5 to 3.7, with 81–90% hit rates on those runs. That is Anthropic’s benchmark, not a Spurlock client result. Treat it as a reason to turn caching on and measure — not as a savings you can put in a deck.
OpenAI’s prompt caching is prefix-exact. Eligible prompts start at 1,024 tokens. Cached input bills at 0.1× the uncached input rate. On the GPT-5.6 family and later, cache writes bill at 1.25× and you can set a 30-minute TTL with prompt_cache_options.ttl. Place static instructions, tools, and schemas first. Put the volatile user turn last. Set a stable prompt_cache_key per tenant or session so repeats land on the same cache shard.
Google’s Gemini API documents implicit and explicit context caching. Implicit is automatic and does not guarantee a discount. Explicit lets you set a TTL (default one hour if unset) and bills storage for as long as the cache lives. Amazon Bedrock’s prompt caching page is the same idea on the hosted side: cache the repeated agent prefix; pay write vs read at the model family’s published rates.
| Cache shape | Pays when | Lies when |
|---|---|---|
| Provider prompt cache | Stable prefix, turns seconds apart | You mutate tools or effort mid-loop and keep missing |
| 1-hour / long TTL | Human waits between turns | You cache policy that changed an hour ago |
| Explicit Gemini cache | Same corpus, many short asks | TTL storage outruns the reuse |
| App-level semantic cache | Identical question, same doc_version | You cache “similar” answers across tenants |
Vendor multipliers, as published in August 2026 — cite the page, do not turn them into a client saving:
| Vendor | Write | Read | Lifetime you actually set |
|---|---|---|---|
| Anthropic Claude API | 1.25× (5m) or 2× (1h) | 0.1× | Default 5m, refreshes on hit |
| OpenAI GPT-5.6+ | 1.25× | 0.1× | prompt_cache_options.ttl, 30m default |
| Gemini explicit cache | Create at input rate + storage | Discounted cached input | TTL you choose; default 1h if unset |
| Gemini implicit cache | None extra | Only if a hit happens | Not guaranteed |
Anthropic’s pricing page states the break-even in multipliers, not dollars: a 5-minute write pays off after one cache read; a 1-hour write pays off after two. If your loop waits on a human, the 5-minute default dies and you re-pay full input plus another write. That is when the 1-hour TTL is the cheaper structure, still not a promised percent.
Three settings break a Claude cache mid-task, per that same cost guide: changing effort between requests, changing a task budget partway through, and context-editing in many small passes. Make those changes at a natural break, then confirm cache reads did not drop.
Caching does not stop the resend. It makes the resend cheaper. A 40-turn task still sends turn one forty times. Anthropic says that cost grows with roughly the square of turn count unless the prefix is a cache read. Caps on turn count still matter.
Do not invent a percent saved for the buyer. Report your hit rate, write tokens, and read tokens from provider usage. OpenAI tells you to watch cached_tokens and, on GPT-5.6+, cache_write_tokens so write cost can be compared with later reads. If writes exceed reads, you are paying a tax for a cache nobody hits.
How do you key caches so they do not serve stale policy?
A hot cache that serves last week’s refund rule is not a saving. It is a cheaper wrong answer.
Cache keys for high-stakes jobs should include:
| Key part | Why it is there | What breaks if omitted |
|---|---|---|
job_type | Different jobs, different prefixes | Cross-job contamination |
prompt_version | Prompt edits must miss | Old instructions, new traffic |
tools_version | Schema edits must miss | Cached tools that no longer exist |
doc_version / policy hash | Knowledge and policy drift | Stale refund, pricing, or legal text |
eval_criteria_version | Evaluator changes change the job | Cache hides a criteria swap |
tenant_id when prompts differ | Isolation | Tenant A’s prefix billed as Tenant B’s hit |
Prefer short TTLs on knowledge-grounded answers. Refresh on every policy deploy. If you cannot bust the cache when doc_version moves, you do not have a cache. You have a rumor.
- Policy deploy includes a cache-bust of that
doc_version - Evaluator criterion edits increment
eval_criteria_version - Tenant-specific system text is not in a shared prefix
- Semantic “near match” caches are off for anything that can write
Observability for agents is where cache hits, writes, and misses belong on the same timeline as the evaluator verdict. A cache-hit rate with no verdict is a vanity panel.
What belongs in a kill switch versus a provider spend limit?
Provider limits stop their API. They do not stop your loop from looking healthy while it retries. OpenAI’s production best practices tell you to isolate staging and production projects and set project spend limits so a test key cannot burn prod quota. Do that. Then put the freeze in the runner anyway.
OpenAI rate limits are org- and project-level request and token ceilings, separate from the spend cap and from the usage-tier monthly quota OpenAI assigns. Hitting a rate limit is a 429. Hitting a hard spend limit is also a 429, with a different error code. Your runner should branch on the code. Treating every 429 as “retry with backoff” is how a spend trip becomes a retry storm.
| Trip | Soft action | Hard action |
|---|---|---|
| Spend > X in 10 minutes for a tenant | Warn + degrade noncritical jobs | Freeze new runs, page owner |
| Single run > Z× p95 cost | Escalate, keep receipt | Abort further model calls |
| Error rate > Y% over N runs | Degrade writes | Freeze that job_type |
| Evaluator fail spike after deploy | Hold the canary | Rollback pin + freeze |
| Daily tenant cap at 70 / 90 / 100% | Warn / degrade / freeze | Hard-stop writes if errors ride along |
Provider *_spend_limit_exceeded | Stop scheduling | Do not retry as a rate limit |
Kill-switch tiers you can actually test:
- Warn — page at 70% of the daily cap.
- Degrade — disable noncritical
job_types at 90%. - Freeze — stop new runs at 100%; finish in-flight only if the next step is read-only.
- Hard stop — abort in-flight writes if an error-rate trip accompanies the spend trip.
Test each tier in staging. Untested switches do not exist. Document who can override, for how long, and which receipt the override writes.
Soft mode: degrade to a human-only queue. Hard mode: freeze tool writes. Both leave a reason code. Silent aborts look like “AI is flaky” in the business’s mouth.
How do you forecast fleet spend without inventing a savings number?
Do not put a made-up “we will save $X” in the launch doc. Put a method, then fill it with your golden-set cost per pass.
expected_daily_jobs
× cost_per_pass_p95
× safety_margin
≈ expected daily model+tool spend
Add escalate cost: human minutes × loaded rate. An agent that “saves” five minutes but escalates 40% at fifteen minutes each is a loss — and that 40% has to come from your measured escalate rate, not a slide.
| Input | Source | Do not |
|---|---|---|
| Daily job volume | Historical ticket / queue counts | Invent a hockey stick |
| Cost per pass p95 | Golden set + a canary week | Use a vendor blog’s percent |
| Safety margin | Your variance, often ~1.3 | Call the margin a saving |
| Escalate rate | Same golden set | Assume escalate is free |
| Tool invoices | Vendor price list × expected calls | Count tokens only |
If the forecast scares finance, you have three honest moves: raise automation share on the known path, narrow autonomy, or raise the business-value threshold for which jobs enter the agent path. Hope is not a forecast.
When you change prompts or models, require the golden set to hold pass rate and the cost band. Ship or don’t. “Slightly better, four times the spend” is a product call.
Re-benchmark mid and small tiers on that set every quarter. Provider price cuts do not matter if your revision rate doubles. Record the decision in the same log the rest of the agentic stack uses.
What do you do when a tool costs more than the model?
Enrichment APIs, scrapers, and search calls often outspend the completion. Budgets that only decrement tokens are lying.
Put estimated USD on each tool definition. Decrement the same budget_usd counter. If the tool has no price on the card, it does not ship.
| Tool class | Bill from | Cap |
|---|---|---|
| Model completion | Provider usage | Per-run token + USD |
| Retrieval / search | Vendor per-call or per-credit | max_retrieval_calls |
| Enrichment / scrape | Vendor invoice | Per-run USD + allowlist |
| Write to CRM / inbox | Usually free, high blast radius | Policy gate, not a price cap |
| Child agent | Child’s remaining allowance | Fan-out cap |
- Every tool schema has
est_usdorest_tokens - Paid tools cannot run when remaining budget <
est_usd - Retrieval is not a substitute for memory
- Writes are gated by policy even when they are “free”
A cheap model plus an expensive scrape loop is still an expensive run. Route the model and cap the tool.
Abort or escalate when the budget dies?
Abort when continuing cannot help. Escalate when a human might finish cheaper than another model turn.
| Remaining situation | Terminal | Reason code |
|---|---|---|
| Auth broken, provider 401/403 | abort | upstream_auth |
| Daily tenant cap hit | abort new runs | day_cap |
| Per-run budget < next step, human can finish | escalate | budget_exhausted |
| Revision ceiling hit, criteria still fail | escalate | revision_ceiling |
| Kill switch hard stop | abort | kill_switch |
| Fallback golden set fail | escalate | fallback_quality |
Do not abort quietly. Write an ops event. Carry the receipt on the run: spend, revisions, last state, terminal reason. That is what observability for agents is for — cost on the same timeline as the verdict, not a finance export two weeks later.
OpenTelemetry’s GenAI conventions moved to the semantic-conventions-genai repo. Record gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, and the cache-create / cache-read attributes when the provider sends them. Convert to dollars with a versioned price table you own. There is no standard cost_usd attribute that stays true across vendors. If you emit dollars, stamp price_table_version.
How do you charge back spend without a shared-key black hole?
If product teams do not see spend, they will externalize it onto a shared key. Per-job_type and per-tenant tags make chargeback possible. Incentives should reward cost per successful outcome, not raw call-count cuts that tank quality.
| Practice | Why | Failure mode |
|---|---|---|
| Separate keys/projects per environment | Staging cannot burn prod | Shared key, shared surprise |
Label every run job_type + tenant | You can freeze one slice | You can only freeze the fleet |
| Weekly cost + quality, same meeting | Stops expensive mediocrity | Finance vs eng talking past each other |
| Budget owner + technical owner | Dollars and caps have names | Orphan budget, everyone’s problem |
| Show $/successful job to non-engineers | Right debate | Raw token charts invite the wrong one |
Every job_type has a budget owner (usually product or ops) and a technical owner (engineering). The budget owner sets the dollars. The technical owner implements caps and kill switches. When spend spikes, both are in the thread.
Document who can freeze a job_type. Practice the freeze in staging. Control AI agent costs is an ops skill you rehearse, like restores.
What fails when you only watch cache hit rate?
This is the failure mode that looks like competence.
A team turns on provider caching, hit rate climbs, the token chart droops, and they declare victory. Meanwhile revision count is up, a cheap fallback is writing wrong CRM notes, and a semantic cache is serving a policy that legal retired on Tuesday.
| What you watched | What was actually on fire | What to do instead |
|---|---|---|
| Cache hit rate | Revision spiral | Cap revisions; score cost per pass |
| Token chart | Tool invoices | Decrement tools on the same budget |
| Provider spend limit | Mid-run 429 on a write | Runner freeze before the provider cap |
| Cheaper model share | Pass rate down, retries up | Golden-set gate on the routed tier |
| Shared-key “savings” | One tenant ate the month | Per-tenant caps + chargeback tags |
Optimizing only cache hit rate is useful and secondary. Revision spirals and over-retrieval write the invoice. Paying for giant contexts as a memory strategy is the same class of mistake: you are renting a longer prompt instead of capping what the runner is allowed to remember.
Run three scenarios when you revisit caps: volume 2×, model price 0.5×, pass rate −10%. Update the integers. Fleets that only plan for the happy cost curve get surprised by success (more volume) as often as by failure (more revisions).
Who owns the budget when spend spikes?
Orphan budgets become everyone’s problem and nobody’s priority. Name the pair before the first production week.
| Role | Sets | Does not set |
|---|---|---|
| Budget owner | Dollars, daily tenant cap, which jobs may run | Model ids, cache keys |
| Technical owner | Caps, routing table, kill-switch wiring | The business-value threshold |
| On-call | Freeze / degrade / override with a receipt | Quiet “just this once” without a log |
| Finance | Chargeback tags, monthly review | Runtime integers |
Demo keys need caps too. Unlimited “agent days” for internal demos are how staging teaches the fleet a bad habit. Hide cost from builders and they will not optimize it. Show $ / successful job and weekly spend next to jobs completed. Invite the right debate: is this job still worth calling an agent?
The $1,500 · 5-day agentic pilot includes wiring budgets and a revision ceiling for one job so you see real unit cost on your data before a larger build. Surprises belong in week one, not month three. No invented percent. The number you leave with is the measured cost per pass on your golden set.
How do parent and child budgets work on a multi-agent relay?
Allocate one parent wallet at intake. Children receive allowances. They do not mint money.
If three specialists each assume the full per-run cap, you have tripled the budget the buyer signed. The graph will hide the burn in specialist hops and the parent will look “in band” until the card does not.
| Rule | Parent | Child |
|---|---|---|
| Wallet | Created at intake | Allowance sliced from remaining |
| Request more | Hub decides | Ask the hub; never spend silently |
| Exhausted | Escalate or abort the graph | Return budget_exhausted to hub |
| Caps | Fan-out + daily tenant | Own revision + tool ceilings |
| Receipt | Sum of children + parent calls | Own spend, own reason code |
Procedure the hub actually runs:
- Intake attaches the parent
budget_usd/budget_tokens. - Hub names each child
job_typeand writes an allowance ≤ remaining. - Child decrements only its allowance. It cannot see the parent remainder as spendable.
- Child returns unused allowance on
done,escalate, orabort. - Hub refuses a fourth child when
children_spawnedhits the fan-out cap.
- Parent remaining is visible on the hub span
- Child spans show
allowance_usdandspent_usd - A child that needs more must request; a silent overspend is a bug
- Unused allowance returns; it does not leak to the next job
Control AI agent costs across handoffs or the graph will hide the burn. The freeze still lives on the parent job_type. Freezing a child type without freezing the hub just reroutes the spend.
What must a week-one cost receipt show?
If a human cannot answer these from one screen, you are not ready to widen volume. This is the same weekly screen observability for agents argues for — cost is a field on that screen, not a sidecar spreadsheet.
| Field | Why it is on the receipt | Red flag |
|---|---|---|
job_type + tenant | Chargeback and freeze target | “misc” or missing tenant |
| Terminal reason | done / escalate / abort + code | Free text, or silent abort |
| Cost per pass (p50 / p95) | Unit economics | Only tokens, no dollars |
| Revision count | Spiral detector | Hidden in traces |
| Cache write / read tokens | Whether caching is paying | Hit rate with no write/read split |
| Tool USD | Non-token burn | Tools omitted |
| Model id actually called | Routing truth | Family nickname, or fallback unmarked |
| Price table version | Dollars stay auditable | Hard-coded rates from memory |
Week-one install, in order:
- Separate staging and prod keys. Put a hard project spend limit on staging.
- Attach per-run counters and a revision ceiling of 3 for the one job in the pilot.
- Turn on provider prompt caching with versioned keys. Do not turn on semantic cache for writes.
- Route
intakeand mechanical eval off the flagship. Leave draft/judge where the golden set says they belong. - Wire warn / degrade / freeze. Trip them on purpose in staging.
- Review cost per pass next to pass rate in the same meeting. Once. Then weekly.
You will not leave week one with a percent saved. You will leave with a measured cost-per-pass band and a freeze that works. That is the whole point of the thin pilot.
FAQ
How do you control AI agent costs in production?
Enforce per-run and per-tenant budgets in the runner, cap revisions and tool calls, route models by state, cache the stable prefix with versioned keys, and install kill switches that abort and alert. Review cost beside quality weekly. Provider spend limits are a backstop, not the control plane.
What is a practical token budget for agents?
Start from the business-acceptable cost per successful job, convert to tokens and USD with a documented margin, attach both counters at intake, decrement on each billable call, and escalate when the next step cannot fit. Carry remaining budget across handoffs so children cannot each assume a full wallet.
Will cheaper models always save money?
No. If pass rate drops and revisions spike, total cost rises. Always measure cost per passing run on a golden set before switching tiers or enabling a budget fallback. A silent proxy downgrade that the evaluator never scored is how cheap tokens become expensive wrong writes.
How should prompt caching be keyed on an agent fleet?
Include job_type, prompt and tool versions, policy doc_version, evaluator criterion version, and tenant when prefixes differ. Bust the cache on policy deploy. Short TTL on knowledge-grounded answers. A hit on last week’s rule is not a saving.
How fast should a kill switch react?
Minutes, not months. Spend anomalies and error spikes should warn, degrade, or freeze automatically, then page a human. Batch monthly reviews are too slow for a runaway loop. Test the switch in staging; an untested freeze is a dashboard.
Does Spurlock Studios include cost controls in the agentic pilot?
Yes. Thin budgets, a revision ceiling, and a freeze path ship in the $1,500 · 5-day pilot so unit cost is visible on your data before a larger build. See /agentic and the operating manual.
CTA
If you cannot freeze a job_type in one action, you do not control agent costs — you observe them. Put the freeze next to the dashboard ops already opens.
What questions does this article answer?
- How do you control AI agent costs in production?
- Enforce per-run and per-tenant budgets in the runner, cap revisions and tool calls, route models by state, cache the stable prefix with versioned keys, and install kill switches that abort and alert. Review cost beside quality weekly. Provider spend limits are a backstop, not the control plane.
- What is a practical token budget for agents?
- Start from the business-acceptable cost per successful job, convert to tokens and USD with a documented margin, attach both counters at intake, decrement on each billable call, and escalate when the next step cannot fit. Carry remaining budget across handoffs so children cannot each assume a full wallet.
- Will cheaper models always save money?
- No. If pass rate drops and revisions spike, total cost rises. Always measure cost per passing run on a golden set before switching tiers or enabling a budget fallback. A silent proxy downgrade that the evaluator never scored is how cheap tokens become expensive wrong writes.
- How should prompt caching be keyed on an agent fleet?
- Include `job_type`, prompt and tool versions, policy `doc_version`, evaluator criterion version, and tenant when prefixes differ. Bust the cache on policy deploy. Short TTL on knowledge-grounded answers. A hit on last week’s rule is not a saving.
- How fast should a kill switch react?
- Minutes, not months. Spend anomalies and error spikes should warn, degrade, or freeze automatically, then page a human. Batch monthly reviews are too slow for a runaway loop. Test the switch in staging; an untested freeze is a dashboard.
- Does Spurlock Studios include cost controls in the agentic pilot?
- Yes. Thin budgets, a revision ceiling, and a freeze path ship in the **$1,500 · 5-day** pilot so unit cost is visible on your data before a larger build. See [/agentic](/agentic) and the [operating manual](/blog/agentic-systems-operating-manual).
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
AI Agents Why Agents Loop on Failed Tools: No-Progress Detection Beats Longer Prompts
Agents loop on failed tools because the harness never detects no-progress. Fingerprint calls, honor retryable:false, cap turns, terminate with a reason code.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.