Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
William Spurlock Founder — Spurlock Studios Updated 25 MIN
Pass rate looks fine while the agent is failing in ways that matter because “pass” is usually a thin binary on the final artifact. It ignores how many revisions it took, whether the tools were right, how much of the job space you never tested, and what each successful task actually costs. You need a metric panel that can veto a deploy — not a vanity percentage.
This spoke sits under the Agentic Systems Operating Manual. Criteria and golden-set design live in Build the Evaluator Before the Agent. This post owns which numbers actually gate a ship.
The short answer
- Task success for agents means: correct outcome, acceptable trajectory, bounded cost, and a human rewrite rate you can live with.
- High pass with high revision rate means the agent is grinding to green — customers feel the latency and you feel the spend.
- Score trajectories: tool choice, arguments, step count — not only the final blob.
- Eval coverage asks what fraction of real job shapes your golden set touches.
- Gate deploys on a small panel: pass, revision, trajectory, coverage, cost per success, online/offline gap.
What does task success actually mean for agents?
A CRM note can “pass” an evaluator and still be the wrong account, written after twelve tool calls, rewritten by sales, and three times the cost of a human doing it cold. Binary pass hides that story.
The research community already treats “pass” as a family of numbers, not one. Chen et al. defined pass@k on HumanEval as the chance that at least one of k samples is functionally correct — and showed Codex jumping from 28.8% at a single sample to 70.2% with 100 samples. That is a useful research signal. It is a terrible production gate. SWE-bench reports % Resolved — a patch either passes the designated tests or it does not. Sierra’s τ-bench went the other direction and asked whether an agent succeeds on all of k independent trials. Same word, four different questions.
Define success as a bundle before you chart anything:
| Dimension | Question | Pass rate sees it? |
|---|---|---|
| Outcome | Did criteria pass on the artifact? | Yes — this is usually all it sees |
| Trajectory | Were tools and args appropriate? | No |
| Efficiency | Steps and tokens within band? | No |
| Human load | Did a human rewrite or reject? | No |
| Economics | Cost per successful task in band? | No |
| Coverage | Did you even test this job shape? | No — it is the missing denominator |
If you only chart the first row, your agent is optimized for looking done.
- Outcome criteria written down and independent of the worker
- Trajectory rules named (must-use tools, forbidden tools, order)
- Revision ceiling and rewrite flag on every run
- Cost attached to the run, not only the model invoice
- Job-shape fingerprint so coverage has a unit
A green percentage with empty boxes is a press release.
Why can 90% pass still mean heavy human rewrites?
Illustrative — not a measured fleet statistic. Imagine an offline set where 90 of 100 cases meet criteria on the final artifact. On 40 of those passes, a human still edits tone, adds a missing field, or fixes a wrong link before the note goes out. Your pass rate says “ship.” Your revision and rewrite rates say “copilot with expensive thrash.”
A 2026 coding-agent study asked the same question in public: Does Pass Rate Tell the Whole Story? Agents cleared a large share of benchmark tests while design-satisfaction stayed in the 30–50% band. Tests went green. The patch was still not the work a maintainer would merge. That gap is the rewrite queue wearing a lab coat.
Sources of flattering pass:
| Source | What it hides | First counter-metric |
|---|---|---|
| Evaluator too soft | Fields humans actually edit | Human rewrite / reject rate |
| Pass after N revisions | Grind, latency, token spend | Revision depth p50 / p95 |
| Golden set too friendly | Weird tenants, messy CRMs | Coverage vs last-30-day fingerprints |
| Humans silently fix | Online truth never reaches the chart | Rewrite flag on the shipped artifact |
| Pass@k reported as “works” | Lottery tickets as reliability | Pass@1 with a revision ceiling |
| One blended pass number | Easy jobs hiding a broken expensive one | Pass and cost sliced by job type |
Track human rewrite rate and agent revision depth beside pass. When rewrite rate stays high while pass climbs, you improved the judge or the grind — not the product.
I have spent 20,000+ hours on agentic systems and shipped 500+ automations. The fleets that survived contact with a sales team all had a rewrite flag the dashboard could not ignore. The ones that died had a pretty percentage and a Slack channel full of “can you just fix this one.”
What is revision rate, and why does pass hide it?
Revision rate (or revision depth) asks: how many evaluate→revise cycles ran before terminal?
Pass rate treats a clean hit and a twelve-cycle grind as the same one. They are not. The grind burns tokens, adds latency, and usually means the first draft was wrong in a way the evaluator only half-caught.
| Pattern | Pass rate | Revision depth | Read |
|---|---|---|---|
| Clean hit | High | Low | Healthy |
| Grind to green | High | High | Latent failure |
| Early escalate | Lower | Low | Honest control |
| Flail then fail | Low | High | Broken loop |
Compute it from traces, not from a vibe:
- Count
evaluate → reviseedges per run (the model cannot reset this counter). - Store depth on the run record:
revision_depth, plusterminal(done,escalate,abort). - Chart p50 and p95 by job type, not a fleet average.
- Pair with human rewrite rate: share of
doneruns a human still edited before the artifact shipped. - Gate: a deploy may keep pass flat but must not raise p50/p95 revision depth beyond an agreed band.
| Band (pilot default) | p50 depth | p95 depth | Action |
|---|---|---|---|
| Healthy | 0–1 | ≤2 | Hold |
| Watch | 2 | 3 | Inspect top trajectory fail codes |
| Block | ≥3 | ≥4 | Do not widen autonomy |
Three is a good pilot ceiling because it matches how operators already think about retries. Crossing it is not “the model tried hard.” It is a quality bug with a cost costume.
Gate idea: revision depth is a veto, not a footnote. Grind that eventually passes is still a fail for a ship decision.
How do you score trajectories?
Trajectory scoring grades the path, not only the destination. LangSmith’s own docs put this in vendor vocabulary: a trajectory evaluator compares the steps the agent took against the steps you expected, and their trajectory-evals package splits the job into hard trajectory match versus an LLM judge. You do not need their product. You need the idea: the sequence is a first-class artifact.
Minimum dimensions:
- Tool choice — required tools used; forbidden tools never called
- Arguments — ids and filters match the job; no invented keys
- Step count — within band for the job type
- Order constraints — read-before-write, verify-before-irreversible
- No-progress events — fingerprint blocks should be zero on happy paths
Exact-match against one golden path is brittle. Several correct orders exist. Grade properties of the path, not a single movie script.
| Mode | What it checks | When it works | When it lies |
|---|---|---|---|
| Checklist pass/fail | Must-use / must-not-use tools | Pilot, clear policy | Multiple valid orders |
| Subsequence match | Expected steps appear in order, extras allowed | Known required backbone | Over-specified expert traces |
| Weighted deductions | Known anti-patterns cost points | Mature job types | Weights that nobody owns |
| Compare to expert trace | Distance from a recorded human path | Tiny golden set | Treats a better path as a fail |
| LLM judge on the trace | “Was this a reasonable path?” | Soft jobs, no reference | Judge drift, extra spend |
Simple scoring modes that work in practice start with a checklist you can explain to finance:
- Required read happened before the first write
- Verify / dry-run tool fired before any irreversible call
- No tool outside the allowlist
- No repeated
(tool, args)pair that already returned the same payload - Step count inside the job-type band
- Terminal status matches the outcome check (
donecannot hide a failed criterion)
You do not need a research benchmark. You need “called enrichment before CRM write” as a first-class fail even when the note text looks fine.
A trajectory fail with a passing artifact is the most under-reported result in agent eval. It is also the one that turns into a silent data incident.
What is eval coverage — the denominator pass rate skips?
Pass rate is passes / evaluated. Coverage asks evaluated shapes / shapes that appear in production.
Without coverage, you can have 95% pass on a toy set and collapse on the first weird tenant. OpenAI’s own eval cookbook says the same thing in operator language: start with a gold seed of 10–50 cases, then expand from production failures, and keep a holdout so you can see when you are training for the test. Their evaluation best practices call the last step continuous evaluation — grow the set, or the number rots.
SWE-bench Verified is the public version of this lesson. OpenAI and the benchmark authors released a 500-task human-validated subset because the original set included underspecified issues and badly scoped tests — a pass/fail number on a dirty denominator. Later they stopped reporting Verified for frontier launches, citing contamination and design issues that made the resolve percentage stop measuring the thing people thought it measured. If a famous industry pass number can stop answering the question, yours can too.
Build a coverage map:
- Job types in production vs job types in golden set
- Tenant size bands (solo, mid, messy CRM)
- Known failure modes (duplicates, empty enrichments, auth errors)
- Languages / locales you actually serve
- Write vs read-only paths
- Escalate / refuse paths — jobs the agent should not complete
- Holdout slice that prompt changes are not allowed to touch
Report coverage % as “share of last 30 days’ production job fingerprints that match at least one golden case family.” Exact formulas vary; the point is to stop celebrating pass on an unrepresentative slice.
| Coverage slice | How to fingerprint | Add a case when… |
|---|---|---|
| Job type | job_type on the trigger | A new type exceeds N runs / week |
| Tenant mess | record count, null rate, duplicate rate | A tenant lands outside existing bands |
| Failure mode | error class + tool name | The same class repeats in online samples |
| Locale | language + region | You take a market you never tested |
| Authority | read-only vs write vs irreversible | You widen the write allowlist |
| Refusal | policy-block vs “cannot do this” | You add a policy the set never exercises |
An eval set made only of solvable happy paths never exercises the give-up path. That is where agents are least tested and most damaging.
Cost per successful task vs cost per run — which number is honest?
Cost per run flatters agents that fail cheap and pass expensive — or the reverse. Finance cares about cost per successful task (and per task that ships without human rewrite, if that is your bar).
Braintrust published the arithmetic in the open. On a large agent-trace study they showed cost per task and cost per success rank configs differently: cost per success = cost per task ÷ success rate. A config that succeeds one-in-six pays for roughly six attempts per win. Their follow-up on cost per resolved request is stricter still: a request only counts as resolved if it clears quality gates — correct tools, no unsafe actions, complete response. That is the production cost that matters.
| Metric | Flatters | Use for | Gate? |
|---|---|---|---|
| Cost / run | Cheap failures | Capacity planning | No |
| Cost / pass | Grind that eventually passes | Unit economics of “green” | Watch |
| Cost / success | Nothing — this is the honest attempt cost | Bake-offs, model swaps | Yes |
| Cost / shipped without rewrite | Nothing — this includes human time | Go / no-go on autonomy | Yes |
Illustrative arithmetic: if average cost/run is low but only one in three runs ships without rewrite, your true cost is closer to cost_per_run / ship_rate plus human time. Chart the honest number next to pass rate or you will “save money” into a support queue.
Procedure you can run on last week’s traces:
- Sum model + tool + evaluator spend for the job type (
cost_all). - Count runs that met outcome and trajectory and stayed inside the revision band (
successes). - Count runs a human shipped with no edit (
shipped_clean). cost_per_success = cost_all / successes.cost_per_clean_ship = cost_all / shipped_cleanplushuman_minutes * loaded_rate.- Slice by job type. A cheap FAQ job must not hide an expensive write job.
If you cannot produce those two ratios from the same traces you use for pass rate, you do not have unit economics. You have a cloud bill.
How do online and offline scores diverge?
Offline golden sets are stubs, frozen tools, and known answers. Online is live schemas, live latency, and distribution shift.
Braintrust’s online scoring model is the vendor version of a rule you should already have: score a sample of production traces asynchronously, then promote the misses into the offline set. Offline is how you iterate fast. It is also how you overfit.
Typical divergence patterns:
| Pattern | Offline | Online | Likely cause | First move |
|---|---|---|---|---|
| Offline high, online low | Strong | Weak | Drift, stubs too clean | Diff schemas and tool errors |
| Both high, rewrite high | Strong | Strong | Soft criteria | Tighten the judge; watch rewrite |
| Offline low, online “fine” | Weak | Strong | Prod sampling biased to easy jobs | Re-sample by job type, not volume |
| Spike after deploy | Drop | Drop | Real regression | Roll back; do not “tune the judge” |
| Offline climbs, holdout flat | Strong | Mixed | You trained for the test | Freeze the holdout; add prod cases |
Rule: never promote on offline alone. Sample online through the same evaluator you used in CI. Alert when the online/offline gap widens past your tolerance — that gap is often the first smoke of schema drift or retrieval rot.
- Same evaluator binary / rubric online and offline
- Sample rate written down (1–10% on high volume; higher on write paths)
- Online misses get a fingerprint and a ticket, not a shrug
- Holdout set that prompt PRs cannot add cases to
- Gap alert:
|online_pass - offline_pass|beyond band
A widening gap is a release smell, not noise.
Pass@1 vs Pass@k — which number ships?
Pass@k (success if any of k samples works) is a research comfort metric. The HumanEval harness reports pass@1, pass@10, and pass@100 because the Codex paper was measuring whether some sample in a budget was correct. Production agents usually get one billed trajectory per job unless you explicitly budget parallel attempts.
| Metric | Meaning | Ship decision |
|---|---|---|
| Pass@1 | First trajectory meets criteria | Default gate for autonomy |
| Pass@k | Best of k meets criteria | Research / model compare only |
| Pass@1 with revisions | Success inside a revision ceiling | Allowed if depth stays in band |
| Pass^k | All of k independent trials succeed | Reliability / consistency check |
If you report Pass@k to executives as “the agent works,” you are selling lottery tickets as reliability. Use Pass@k for model bake-offs; ship on Pass@1 (with bounded revisions) and rewrite rate.
A useful bake-off protocol:
- Freeze the golden set and the evaluator.
- Report Pass@1, revision p95, trajectory fail rate, and cost per success — four numbers, same rows.
- You may also report Pass@8 or Pass@16 in the appendix so research-minded stakeholders can compare to papers.
- The ship column is Pass@1 + revision band + trajectory checklist. Not the appendix.
Pass@k answers “can this system ever get it right.” Shipping answers “does it get it right on the attempt we will bill.”
What is pass^k, and why should ops care?
τ-bench introduced pass^k (pass-hat-k) because customer-service agents have to be consistent, not occasionally brilliant. Pass@k asks whether at least one of k trials works. Pass^k asks whether all k trials work. The original paper reported then-frontier function-calling agents succeeding on under 50% of tasks, with pass^8 under 25% in retail. That is the reliability shape operators actually feel: “it worked yesterday” is not a control.
| If Pass@1 is | Then Pass^4 is roughly | Read |
|---|---|---|
| 90% | ~66% | Still a lot of “sometimes” |
| 80% | ~41% | Coin-flip consistency at four tries |
| 70% | ~24% | Do not widen write access |
| 50% | ~6% | Demo, not a teammate |
The table is independent Bernoulli arithmetic, not a fleet measurement. Real runs are correlated — same tenant, same stale schema — so production pass^k is often worse than the coin-flip picture.
When to pull this lever:
- The job is a write, a refund, a booking, or anything you cannot quietly undo
- Two operators already disagree about whether “it usually works”
- You are comparing two prompts that look tied on Pass@1
- A vendor slide shows Pass@k and calls it reliability
Pass^k will make your agent look worse. That is the point. A metric that cannot veto a deploy is decoration.
Which metrics gate a deploy?
Minimum veto panel before widening autonomy or merging prompt/tool changes:
- Offline pass@1 — no drop beyond agreed delta on golden set
- Revision depth — p50/p95 inside band
- Trajectory checklist — no new systematic tool-choice fails
- Eval coverage — not reduced; new failure modes get cases
- Cost per success — inside band
- Online sample — after canary, online pass and rewrite rate hold
Any single green light is insufficient. Pass rate alone is never a gate.
| Signal | Block the merge when… | Who owns the band |
|---|---|---|
| Offline pass@1 | Drops more than the agreed delta | Eval owner |
| Revision p95 | Crosses the job-type band | Loop owner |
| Trajectory fails | A new systematic code appears | Tool / policy owner |
| Coverage | Share of last-30-day fingerprints falls | Eval owner |
| Cost / success | Spikes after the change | whoever pays the bill |
| Online rewrite | Climbs while pass stays flat | Operator closest to the humans |
| Online/offline gap | Widens past tolerance | Release owner |
Write the bands down before the first canary. Negotiating them after a pretty CI run is how vanity pass wins.
A release checklist that references the panel:
- Golden set hash recorded on the PR
- Four gate numbers pasted in the PR (pass@1, revision p95, trajectory fail rate, cost / success)
- Coverage note: fingerprints added or explicitly deferred
- Canary plan: sample rate, duration, rewrite watch
- Rollback owner named
If the PR only has a pass percentage, bounce it.
What weekly panel will ops trust?
Keep this separate from the full trace wall. One screen, business-readable:
| Panel | Owner acts when… |
|---|---|
| Pass@1 (online sample) | Drops vs trailing baseline |
| Human rewrite / reject rate | Climbs while pass flat |
| p95 revision depth | Crosses band |
| Trajectory fail codes (top 3) | Same code repeats |
| Cost / success by job type | Spikes after deploy |
| Coverage gaps (new prod shapes) | Untested families appear |
| Online/offline gap | Widens past band |
Deep traces live one click away. If the weekly meeting needs a data scientist to interpret the primary screen, the panel failed.
Ritual that fits a 20-minute standup:
- Read the seven cells. No slides.
- Open traces for the top trajectory fail code only.
- Pick one fingerprint to add to the golden set, or explicitly defer it.
- If rewrite climbed and pass did not, the judge or the grind moved — treat that as an incident, not a content tweak.
The panel is a decision surface. The traces are evidence. Do not invert them.
What does green CI plus angry sales look like?
Illustrative scenario. CI shows 92% offline pass after a prompt change. Trajectory scoring was not wired. Production: the agent stops calling the verify tool, still produces notes that meet soft criteria, sales rewrites account links daily. Pass rate did not lie — it answered a weaker question than the business asked.
| What CI saw | What production felt | Missing gate |
|---|---|---|
| 92% offline pass | Account links wrong | Trajectory: verify_before_crm_write |
| Latency “fine” in stubs | Notes land late after silent revisions | Revision p95 |
| Cost / run flat | Humans fix 40% of “passes” | Rewrite rate + cost / clean ship |
| Golden set still green | New tenant shape never tested | Coverage fingerprint |
Fix, in order:
- Add the trajectory rule
verify_before_crm_writeas a hard fail. - Measure rewrite rate on the shipped note, not on the evaluator blob.
- Block deploy on trajectory checklist regressions even when pass is flat.
- Add the tenant shape that broke to the golden set before the next prompt PR.
This is the failure mode pass rate is built to hide: the artifact looks done, the path was illegal, and the humans quietly became the runtime.
How do you instrument the panel without a second dashboard?
Observability stores runs, tool spans, scores, and cost. This spoke decides thresholds and veto rules. Practical split:
- Telemetry: emit revision count, trajectory sub-scores, rewrite flags, cost fields on each run
- Release: compare those fields to bands; block or canary
- Weekly ritual: read the panel; open traces for top trajectory fail codes
Do not build two dashboards with the same six charts. Build one telemetry path and a release checklist that references it.
Minimum fields on every run record:
| Field | Type | Feeds |
|---|---|---|
pass_outcome | bool | Pass@1 |
revision_depth | int | Revision band |
trajectory_codes | string[] | Checklist + weekly top-3 |
cost_usd | number | Cost / success |
job_fingerprint | string | Coverage |
rewrite_flag | bool / enum | Human load |
terminal | enum | Honest stop vs hidden retry |
eval_set_id | string | Offline vs online join |
If a field is missing, the corresponding gate is theater. You cannot veto what you did not emit.
What is the pilot minimum?
A Spurlock Studios $1,500 · 5-day agentic pilot should leave you with more than a pass percentage: evaluator criteria, a small golden set with at least one trajectory checklist, revision depth on traces, and a written deploy gate. Fancy coverage math can grow later; “pass rate alone” should already be dead as a ship criterion.
Day-by-day, the metric work is small and non-negotiable:
| Day | Metric outcome |
|---|---|
| 1 | Success bundle written: outcome, trajectory, rewrite, cost |
| 2 | Golden seed (10–50) with fingerprints; holdout split named |
| 3 | Revision counter and trajectory checklist on every trace |
| 4 | Offline pass@1 + revision p95 + cost / success on the seed |
| 5 | Written veto panel and canary rule; one online sample plan |
/agentic · full stack in the operating manual. Criteria first: evaluators before agents.
If day five still reports a single percentage, the pilot failed the only test that matters.
How do you decide if the agent is actually working?
Ask weekly:
- Is online pass holding without a climb in rewrite rate?
- Is revision depth stable or creeping?
- Are trajectory fails concentrated in one tool or job type?
- Did coverage grow with new production shapes?
- Is cost per success inside the band you would defend to finance?
- Is the online/offline gap inside tolerance?
- Would you widen write access on this panel — yes or no, no speech?
If you cannot answer from one panel, you do not know — you are hoping.
| Answer pattern | Verdict |
|---|---|
| Pass up, rewrite up | Judge got softer or grind got longer — do not ship |
| Pass flat, revision p95 up | Latent failure — cap autonomy |
| Pass down, trajectory new code | Real regression — roll back |
| Pass up, coverage down | You deleted hard cases — bounce the PR |
| Pass up, cost / success up | You bought the percentage — finance veto |
| All six stable | You may widen one notch, then re-measure |
Bravery is not a metric strategy.
Anti-patterns
Optimizing only the judge until pass hits a target. You invented grade inflation.
Averaging all job types into one pass number. A tiny easy job hides a broken expensive one.
No ownership on rewrite rate. If sales suffers in silence, metrics stay pretty.
Shipping on Pass@k screenshots. Not how production runs.
Treating % Resolved on a public coding board as your internal gate. Even SWE-bench Verified had to be human-filtered, then later set aside for frontier reporting. Your CRM note is not a GitHub issue with a unit test.
Celebrating cost / run after a week of cheap failures. Failures are not a discount.
Duplicating a trace wall into the weekly panel. Traces are necessary; gates are the deploy policy this spoke owns.
FAQ
Pass@1 vs Pass@k — which ships?
Ship on Pass@1 with a bounded revision ceiling. Use Pass@k for model comparisons and research-style evals. Production jobs rarely get k free attempts, so Pass@k overstates reliability for autonomy decisions. If you need a consistency number, look at pass^k — the chance every trial works — not the chance that one of them did.
What is eval coverage?
It is how much of real production job diversity your golden set and evaluators actually touch. High pass on a narrow set is not readiness — track uncovered job shapes and add cases as they appear online. A practical report is the share of last-30-day production fingerprints that match at least one golden case family.
Cost per successful task vs cost per run?
Cost per run includes cheap failures and hides grind. Cost per successful task (ideally per task shipped without human rewrite) is the unit economic signal that should sit beside pass rate on the gate panel. A common form is cost_all / successes; the stricter form divides by clean ships and adds human time.
How do online and offline scores diverge?
Offline sets use stubs and frozen distributions; production drifts. Promote only when online samples through the same evaluator stay within tolerance of offline — a widening gap is a release smell, not noise. Feed online misses back into the golden set or you will keep passing a test the world already left.
Which metrics gate a deploy?
At minimum: offline pass@1 delta, revision depth band, trajectory checklist, coverage not reduced, cost per success, and a post-canary online hold on pass and rewrite rate. Pass rate alone never ships. Write the bands before the canary, not after a pretty CI run.
How does this connect to the observability dashboard without duplicating it?
Observability emits the fields (scores, revisions, trajectory fails, cost). This spoke sets the veto thresholds and weekly panel. One telemetry path, one release checklist — not two competing dashboards. If a gate field is not on the run record, the gate is theater.
CTA
Stop shipping on a flattering percentage.
What questions does this article answer?
- Pass@1 vs Pass@k — which ships?
- Ship on Pass@1 with a bounded revision ceiling. Use Pass@k for model comparisons and research-style evals. Production jobs rarely get *k* free attempts, so Pass@k overstates reliability for autonomy decisions. If you need a consistency number, look at pass^k — the chance *every* trial works — not the chance that one of them did.
- What is eval coverage?
- It is how much of real production job diversity your golden set and evaluators actually touch. High pass on a narrow set is not readiness — track uncovered job shapes and add cases as they appear online. A practical report is the share of last-30-day production fingerprints that match at least one golden case family.
- Cost per successful task vs cost per run?
- Cost per run includes cheap failures and hides grind. Cost per successful task (ideally per task shipped without human rewrite) is the unit economic signal that should sit beside pass rate on the gate panel. A common form is `cost_all / successes`; the stricter form divides by clean ships and adds human time.
- How do online and offline scores diverge?
- Offline sets use stubs and frozen distributions; production drifts. Promote only when online samples through the same evaluator stay within tolerance of offline — a widening gap is a release smell, not noise. Feed online misses back into the golden set or you will keep passing a test the world already left.
- Which metrics gate a deploy?
- At minimum: offline pass@1 delta, revision depth band, trajectory checklist, coverage not reduced, cost per success, and a post-canary online hold on pass and rewrite rate. Pass rate alone never ships. Write the bands before the canary, not after a pretty CI run.
- How does this connect to the observability dashboard without duplicating it?
- Observability emits the fields (scores, revisions, trajectory fails, cost). This spoke sets the veto thresholds and weekly panel. One telemetry path, one release checklist — not two competing dashboards. If a gate field is not on the run record, the gate is theater.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.