Spurlock Studios
Contact
Share LinkedIn X
A small stack of coins. Thesis: PASS RATE LIES REVISION RATE.

Pass rate looks fine while the agent is failing in ways that matter because “pass” is usually a thin binary on the final artifact. It ignores how many revisions it took, whether the tools were right, how much of the job space you never tested, and what each successful task actually costs. You need a metric panel that can veto a deploy — not a vanity percentage.

This spoke sits under the Agentic Systems Operating Manual. Criteria and golden-set design live in Build the Evaluator Before the Agent. This post owns which numbers actually gate a ship.

The short answer

  • Task success for agents means: correct outcome, acceptable trajectory, bounded cost, and a human rewrite rate you can live with.
  • High pass with high revision rate means the agent is grinding to green — customers feel the latency and you feel the spend.
  • Score trajectories: tool choice, arguments, step count — not only the final blob.
  • Eval coverage asks what fraction of real job shapes your golden set touches.
  • Gate deploys on a small panel: pass, revision, trajectory, coverage, cost per success, online/offline gap.

What does task success actually mean for agents?

A CRM note can “pass” an evaluator and still be the wrong account, written after twelve tool calls, rewritten by sales, and three times the cost of a human doing it cold. Binary pass hides that story.

The research community already treats “pass” as a family of numbers, not one. Chen et al. defined pass@k on HumanEval as the chance that at least one of k samples is functionally correct — and showed Codex jumping from 28.8% at a single sample to 70.2% with 100 samples. That is a useful research signal. It is a terrible production gate. SWE-bench reports % Resolved — a patch either passes the designated tests or it does not. Sierra’s τ-bench went the other direction and asked whether an agent succeeds on all of k independent trials. Same word, four different questions.

Define success as a bundle before you chart anything:

DimensionQuestionPass rate sees it?
OutcomeDid criteria pass on the artifact?Yes — this is usually all it sees
TrajectoryWere tools and args appropriate?No
EfficiencySteps and tokens within band?No
Human loadDid a human rewrite or reject?No
EconomicsCost per successful task in band?No
CoverageDid you even test this job shape?No — it is the missing denominator

If you only chart the first row, your agent is optimized for looking done.

  • Outcome criteria written down and independent of the worker
  • Trajectory rules named (must-use tools, forbidden tools, order)
  • Revision ceiling and rewrite flag on every run
  • Cost attached to the run, not only the model invoice
  • Job-shape fingerprint so coverage has a unit

A green percentage with empty boxes is a press release.

Why can 90% pass still mean heavy human rewrites?

Illustrative — not a measured fleet statistic. Imagine an offline set where 90 of 100 cases meet criteria on the final artifact. On 40 of those passes, a human still edits tone, adds a missing field, or fixes a wrong link before the note goes out. Your pass rate says “ship.” Your revision and rewrite rates say “copilot with expensive thrash.”

A 2026 coding-agent study asked the same question in public: Does Pass Rate Tell the Whole Story? Agents cleared a large share of benchmark tests while design-satisfaction stayed in the 30–50% band. Tests went green. The patch was still not the work a maintainer would merge. That gap is the rewrite queue wearing a lab coat.

Sources of flattering pass:

SourceWhat it hidesFirst counter-metric
Evaluator too softFields humans actually editHuman rewrite / reject rate
Pass after N revisionsGrind, latency, token spendRevision depth p50 / p95
Golden set too friendlyWeird tenants, messy CRMsCoverage vs last-30-day fingerprints
Humans silently fixOnline truth never reaches the chartRewrite flag on the shipped artifact
Pass@k reported as “works”Lottery tickets as reliabilityPass@1 with a revision ceiling
One blended pass numberEasy jobs hiding a broken expensive onePass and cost sliced by job type

Track human rewrite rate and agent revision depth beside pass. When rewrite rate stays high while pass climbs, you improved the judge or the grind — not the product.

I have spent 20,000+ hours on agentic systems and shipped 500+ automations. The fleets that survived contact with a sales team all had a rewrite flag the dashboard could not ignore. The ones that died had a pretty percentage and a Slack channel full of “can you just fix this one.”

What is revision rate, and why does pass hide it?

Revision rate (or revision depth) asks: how many evaluate→revise cycles ran before terminal?

Pass rate treats a clean hit and a twelve-cycle grind as the same one. They are not. The grind burns tokens, adds latency, and usually means the first draft was wrong in a way the evaluator only half-caught.

PatternPass rateRevision depthRead
Clean hitHighLowHealthy
Grind to greenHighHighLatent failure
Early escalateLowerLowHonest control
Flail then failLowHighBroken loop

Compute it from traces, not from a vibe:

  1. Count evaluate → revise edges per run (the model cannot reset this counter).
  2. Store depth on the run record: revision_depth, plus terminal (done, escalate, abort).
  3. Chart p50 and p95 by job type, not a fleet average.
  4. Pair with human rewrite rate: share of done runs a human still edited before the artifact shipped.
  5. Gate: a deploy may keep pass flat but must not raise p50/p95 revision depth beyond an agreed band.
Band (pilot default)p50 depthp95 depthAction
Healthy0–1≤2Hold
Watch23Inspect top trajectory fail codes
Block≥3≥4Do not widen autonomy

Three is a good pilot ceiling because it matches how operators already think about retries. Crossing it is not “the model tried hard.” It is a quality bug with a cost costume.

Gate idea: revision depth is a veto, not a footnote. Grind that eventually passes is still a fail for a ship decision.

How do you score trajectories?

Trajectory scoring grades the path, not only the destination. LangSmith’s own docs put this in vendor vocabulary: a trajectory evaluator compares the steps the agent took against the steps you expected, and their trajectory-evals package splits the job into hard trajectory match versus an LLM judge. You do not need their product. You need the idea: the sequence is a first-class artifact.

Minimum dimensions:

  1. Tool choice — required tools used; forbidden tools never called
  2. Arguments — ids and filters match the job; no invented keys
  3. Step count — within band for the job type
  4. Order constraints — read-before-write, verify-before-irreversible
  5. No-progress events — fingerprint blocks should be zero on happy paths

Exact-match against one golden path is brittle. Several correct orders exist. Grade properties of the path, not a single movie script.

ModeWhat it checksWhen it worksWhen it lies
Checklist pass/failMust-use / must-not-use toolsPilot, clear policyMultiple valid orders
Subsequence matchExpected steps appear in order, extras allowedKnown required backboneOver-specified expert traces
Weighted deductionsKnown anti-patterns cost pointsMature job typesWeights that nobody owns
Compare to expert traceDistance from a recorded human pathTiny golden setTreats a better path as a fail
LLM judge on the trace“Was this a reasonable path?”Soft jobs, no referenceJudge drift, extra spend

Simple scoring modes that work in practice start with a checklist you can explain to finance:

  • Required read happened before the first write
  • Verify / dry-run tool fired before any irreversible call
  • No tool outside the allowlist
  • No repeated (tool, args) pair that already returned the same payload
  • Step count inside the job-type band
  • Terminal status matches the outcome check (done cannot hide a failed criterion)

You do not need a research benchmark. You need “called enrichment before CRM write” as a first-class fail even when the note text looks fine.

A trajectory fail with a passing artifact is the most under-reported result in agent eval. It is also the one that turns into a silent data incident.

What is eval coverage — the denominator pass rate skips?

Pass rate is passes / evaluated. Coverage asks evaluated shapes / shapes that appear in production.

Without coverage, you can have 95% pass on a toy set and collapse on the first weird tenant. OpenAI’s own eval cookbook says the same thing in operator language: start with a gold seed of 10–50 cases, then expand from production failures, and keep a holdout so you can see when you are training for the test. Their evaluation best practices call the last step continuous evaluation — grow the set, or the number rots.

SWE-bench Verified is the public version of this lesson. OpenAI and the benchmark authors released a 500-task human-validated subset because the original set included underspecified issues and badly scoped tests — a pass/fail number on a dirty denominator. Later they stopped reporting Verified for frontier launches, citing contamination and design issues that made the resolve percentage stop measuring the thing people thought it measured. If a famous industry pass number can stop answering the question, yours can too.

Build a coverage map:

  • Job types in production vs job types in golden set
  • Tenant size bands (solo, mid, messy CRM)
  • Known failure modes (duplicates, empty enrichments, auth errors)
  • Languages / locales you actually serve
  • Write vs read-only paths
  • Escalate / refuse paths — jobs the agent should not complete
  • Holdout slice that prompt changes are not allowed to touch

Report coverage % as “share of last 30 days’ production job fingerprints that match at least one golden case family.” Exact formulas vary; the point is to stop celebrating pass on an unrepresentative slice.

Coverage sliceHow to fingerprintAdd a case when…
Job typejob_type on the triggerA new type exceeds N runs / week
Tenant messrecord count, null rate, duplicate rateA tenant lands outside existing bands
Failure modeerror class + tool nameThe same class repeats in online samples
Localelanguage + regionYou take a market you never tested
Authorityread-only vs write vs irreversibleYou widen the write allowlist
Refusalpolicy-block vs “cannot do this”You add a policy the set never exercises

An eval set made only of solvable happy paths never exercises the give-up path. That is where agents are least tested and most damaging.

Cost per successful task vs cost per run — which number is honest?

Cost per run flatters agents that fail cheap and pass expensive — or the reverse. Finance cares about cost per successful task (and per task that ships without human rewrite, if that is your bar).

Braintrust published the arithmetic in the open. On a large agent-trace study they showed cost per task and cost per success rank configs differently: cost per success = cost per task ÷ success rate. A config that succeeds one-in-six pays for roughly six attempts per win. Their follow-up on cost per resolved request is stricter still: a request only counts as resolved if it clears quality gates — correct tools, no unsafe actions, complete response. That is the production cost that matters.

MetricFlattersUse forGate?
Cost / runCheap failuresCapacity planningNo
Cost / passGrind that eventually passesUnit economics of “green”Watch
Cost / successNothing — this is the honest attempt costBake-offs, model swapsYes
Cost / shipped without rewriteNothing — this includes human timeGo / no-go on autonomyYes

Illustrative arithmetic: if average cost/run is low but only one in three runs ships without rewrite, your true cost is closer to cost_per_run / ship_rate plus human time. Chart the honest number next to pass rate or you will “save money” into a support queue.

Procedure you can run on last week’s traces:

  1. Sum model + tool + evaluator spend for the job type (cost_all).
  2. Count runs that met outcome and trajectory and stayed inside the revision band (successes).
  3. Count runs a human shipped with no edit (shipped_clean).
  4. cost_per_success = cost_all / successes.
  5. cost_per_clean_ship = cost_all / shipped_clean plus human_minutes * loaded_rate.
  6. Slice by job type. A cheap FAQ job must not hide an expensive write job.

If you cannot produce those two ratios from the same traces you use for pass rate, you do not have unit economics. You have a cloud bill.

How do online and offline scores diverge?

Offline golden sets are stubs, frozen tools, and known answers. Online is live schemas, live latency, and distribution shift.

Braintrust’s online scoring model is the vendor version of a rule you should already have: score a sample of production traces asynchronously, then promote the misses into the offline set. Offline is how you iterate fast. It is also how you overfit.

Typical divergence patterns:

PatternOfflineOnlineLikely causeFirst move
Offline high, online lowStrongWeakDrift, stubs too cleanDiff schemas and tool errors
Both high, rewrite highStrongStrongSoft criteriaTighten the judge; watch rewrite
Offline low, online “fine”WeakStrongProd sampling biased to easy jobsRe-sample by job type, not volume
Spike after deployDropDropReal regressionRoll back; do not “tune the judge”
Offline climbs, holdout flatStrongMixedYou trained for the testFreeze the holdout; add prod cases

Rule: never promote on offline alone. Sample online through the same evaluator you used in CI. Alert when the online/offline gap widens past your tolerance — that gap is often the first smoke of schema drift or retrieval rot.

  • Same evaluator binary / rubric online and offline
  • Sample rate written down (1–10% on high volume; higher on write paths)
  • Online misses get a fingerprint and a ticket, not a shrug
  • Holdout set that prompt PRs cannot add cases to
  • Gap alert: |online_pass - offline_pass| beyond band

A widening gap is a release smell, not noise.

Pass@1 vs Pass@k — which number ships?

Pass@k (success if any of k samples works) is a research comfort metric. The HumanEval harness reports pass@1, pass@10, and pass@100 because the Codex paper was measuring whether some sample in a budget was correct. Production agents usually get one billed trajectory per job unless you explicitly budget parallel attempts.

MetricMeaningShip decision
Pass@1First trajectory meets criteriaDefault gate for autonomy
Pass@kBest of k meets criteriaResearch / model compare only
Pass@1 with revisionsSuccess inside a revision ceilingAllowed if depth stays in band
Pass^kAll of k independent trials succeedReliability / consistency check

If you report Pass@k to executives as “the agent works,” you are selling lottery tickets as reliability. Use Pass@k for model bake-offs; ship on Pass@1 (with bounded revisions) and rewrite rate.

A useful bake-off protocol:

  1. Freeze the golden set and the evaluator.
  2. Report Pass@1, revision p95, trajectory fail rate, and cost per success — four numbers, same rows.
  3. You may also report Pass@8 or Pass@16 in the appendix so research-minded stakeholders can compare to papers.
  4. The ship column is Pass@1 + revision band + trajectory checklist. Not the appendix.

Pass@k answers “can this system ever get it right.” Shipping answers “does it get it right on the attempt we will bill.”

What is pass^k, and why should ops care?

τ-bench introduced pass^k (pass-hat-k) because customer-service agents have to be consistent, not occasionally brilliant. Pass@k asks whether at least one of k trials works. Pass^k asks whether all k trials work. The original paper reported then-frontier function-calling agents succeeding on under 50% of tasks, with pass^8 under 25% in retail. That is the reliability shape operators actually feel: “it worked yesterday” is not a control.

If Pass@1 isThen Pass^4 is roughlyRead
90%~66%Still a lot of “sometimes”
80%~41%Coin-flip consistency at four tries
70%~24%Do not widen write access
50%~6%Demo, not a teammate

The table is independent Bernoulli arithmetic, not a fleet measurement. Real runs are correlated — same tenant, same stale schema — so production pass^k is often worse than the coin-flip picture.

When to pull this lever:

  • The job is a write, a refund, a booking, or anything you cannot quietly undo
  • Two operators already disagree about whether “it usually works”
  • You are comparing two prompts that look tied on Pass@1
  • A vendor slide shows Pass@k and calls it reliability

Pass^k will make your agent look worse. That is the point. A metric that cannot veto a deploy is decoration.

Which metrics gate a deploy?

Minimum veto panel before widening autonomy or merging prompt/tool changes:

  1. Offline pass@1 — no drop beyond agreed delta on golden set
  2. Revision depth — p50/p95 inside band
  3. Trajectory checklist — no new systematic tool-choice fails
  4. Eval coverage — not reduced; new failure modes get cases
  5. Cost per success — inside band
  6. Online sample — after canary, online pass and rewrite rate hold

Any single green light is insufficient. Pass rate alone is never a gate.

SignalBlock the merge when…Who owns the band
Offline pass@1Drops more than the agreed deltaEval owner
Revision p95Crosses the job-type bandLoop owner
Trajectory failsA new systematic code appearsTool / policy owner
CoverageShare of last-30-day fingerprints fallsEval owner
Cost / successSpikes after the changewhoever pays the bill
Online rewriteClimbs while pass stays flatOperator closest to the humans
Online/offline gapWidens past toleranceRelease owner

Write the bands down before the first canary. Negotiating them after a pretty CI run is how vanity pass wins.

A release checklist that references the panel:

  • Golden set hash recorded on the PR
  • Four gate numbers pasted in the PR (pass@1, revision p95, trajectory fail rate, cost / success)
  • Coverage note: fingerprints added or explicitly deferred
  • Canary plan: sample rate, duration, rewrite watch
  • Rollback owner named

If the PR only has a pass percentage, bounce it.

What weekly panel will ops trust?

Keep this separate from the full trace wall. One screen, business-readable:

PanelOwner acts when…
Pass@1 (online sample)Drops vs trailing baseline
Human rewrite / reject rateClimbs while pass flat
p95 revision depthCrosses band
Trajectory fail codes (top 3)Same code repeats
Cost / success by job typeSpikes after deploy
Coverage gaps (new prod shapes)Untested families appear
Online/offline gapWidens past band

Deep traces live one click away. If the weekly meeting needs a data scientist to interpret the primary screen, the panel failed.

Ritual that fits a 20-minute standup:

  1. Read the seven cells. No slides.
  2. Open traces for the top trajectory fail code only.
  3. Pick one fingerprint to add to the golden set, or explicitly defer it.
  4. If rewrite climbed and pass did not, the judge or the grind moved — treat that as an incident, not a content tweak.

The panel is a decision surface. The traces are evidence. Do not invert them.

What does green CI plus angry sales look like?

Illustrative scenario. CI shows 92% offline pass after a prompt change. Trajectory scoring was not wired. Production: the agent stops calling the verify tool, still produces notes that meet soft criteria, sales rewrites account links daily. Pass rate did not lie — it answered a weaker question than the business asked.

What CI sawWhat production feltMissing gate
92% offline passAccount links wrongTrajectory: verify_before_crm_write
Latency “fine” in stubsNotes land late after silent revisionsRevision p95
Cost / run flatHumans fix 40% of “passes”Rewrite rate + cost / clean ship
Golden set still greenNew tenant shape never testedCoverage fingerprint

Fix, in order:

  1. Add the trajectory rule verify_before_crm_write as a hard fail.
  2. Measure rewrite rate on the shipped note, not on the evaluator blob.
  3. Block deploy on trajectory checklist regressions even when pass is flat.
  4. Add the tenant shape that broke to the golden set before the next prompt PR.

This is the failure mode pass rate is built to hide: the artifact looks done, the path was illegal, and the humans quietly became the runtime.

How do you instrument the panel without a second dashboard?

Observability stores runs, tool spans, scores, and cost. This spoke decides thresholds and veto rules. Practical split:

  • Telemetry: emit revision count, trajectory sub-scores, rewrite flags, cost fields on each run
  • Release: compare those fields to bands; block or canary
  • Weekly ritual: read the panel; open traces for top trajectory fail codes

Do not build two dashboards with the same six charts. Build one telemetry path and a release checklist that references it.

Minimum fields on every run record:

FieldTypeFeeds
pass_outcomeboolPass@1
revision_depthintRevision band
trajectory_codesstring[]Checklist + weekly top-3
cost_usdnumberCost / success
job_fingerprintstringCoverage
rewrite_flagbool / enumHuman load
terminalenumHonest stop vs hidden retry
eval_set_idstringOffline vs online join

If a field is missing, the corresponding gate is theater. You cannot veto what you did not emit.

What is the pilot minimum?

A Spurlock Studios $1,500 · 5-day agentic pilot should leave you with more than a pass percentage: evaluator criteria, a small golden set with at least one trajectory checklist, revision depth on traces, and a written deploy gate. Fancy coverage math can grow later; “pass rate alone” should already be dead as a ship criterion.

Day-by-day, the metric work is small and non-negotiable:

DayMetric outcome
1Success bundle written: outcome, trajectory, rewrite, cost
2Golden seed (10–50) with fingerprints; holdout split named
3Revision counter and trajectory checklist on every trace
4Offline pass@1 + revision p95 + cost / success on the seed
5Written veto panel and canary rule; one online sample plan

/agentic · full stack in the operating manual. Criteria first: evaluators before agents.

If day five still reports a single percentage, the pilot failed the only test that matters.

How do you decide if the agent is actually working?

Ask weekly:

  1. Is online pass holding without a climb in rewrite rate?
  2. Is revision depth stable or creeping?
  3. Are trajectory fails concentrated in one tool or job type?
  4. Did coverage grow with new production shapes?
  5. Is cost per success inside the band you would defend to finance?
  6. Is the online/offline gap inside tolerance?
  7. Would you widen write access on this panel — yes or no, no speech?

If you cannot answer from one panel, you do not know — you are hoping.

Answer patternVerdict
Pass up, rewrite upJudge got softer or grind got longer — do not ship
Pass flat, revision p95 upLatent failure — cap autonomy
Pass down, trajectory new codeReal regression — roll back
Pass up, coverage downYou deleted hard cases — bounce the PR
Pass up, cost / success upYou bought the percentage — finance veto
All six stableYou may widen one notch, then re-measure

Bravery is not a metric strategy.

Anti-patterns

Optimizing only the judge until pass hits a target. You invented grade inflation.

Averaging all job types into one pass number. A tiny easy job hides a broken expensive one.

No ownership on rewrite rate. If sales suffers in silence, metrics stay pretty.

Shipping on Pass@k screenshots. Not how production runs.

Treating % Resolved on a public coding board as your internal gate. Even SWE-bench Verified had to be human-filtered, then later set aside for frontier reporting. Your CRM note is not a GitHub issue with a unit test.

Celebrating cost / run after a week of cheap failures. Failures are not a discount.

Duplicating a trace wall into the weekly panel. Traces are necessary; gates are the deploy policy this spoke owns.

FAQ

Pass@1 vs Pass@k — which ships?

Ship on Pass@1 with a bounded revision ceiling. Use Pass@k for model comparisons and research-style evals. Production jobs rarely get k free attempts, so Pass@k overstates reliability for autonomy decisions. If you need a consistency number, look at pass^k — the chance every trial works — not the chance that one of them did.

What is eval coverage?

It is how much of real production job diversity your golden set and evaluators actually touch. High pass on a narrow set is not readiness — track uncovered job shapes and add cases as they appear online. A practical report is the share of last-30-day production fingerprints that match at least one golden case family.

Cost per successful task vs cost per run?

Cost per run includes cheap failures and hides grind. Cost per successful task (ideally per task shipped without human rewrite) is the unit economic signal that should sit beside pass rate on the gate panel. A common form is cost_all / successes; the stricter form divides by clean ships and adds human time.

How do online and offline scores diverge?

Offline sets use stubs and frozen distributions; production drifts. Promote only when online samples through the same evaluator stay within tolerance of offline — a widening gap is a release smell, not noise. Feed online misses back into the golden set or you will keep passing a test the world already left.

Which metrics gate a deploy?

At minimum: offline pass@1 delta, revision depth band, trajectory checklist, coverage not reduced, cost per success, and a post-canary online hold on pass and rewrite rate. Pass rate alone never ships. Write the bands before the canary, not after a pretty CI run.

How does this connect to the observability dashboard without duplicating it?

Observability emits the fields (scores, revisions, trajectory fails, cost). This spoke sets the veto thresholds and weekly panel. One telemetry path, one release checklist — not two competing dashboards. If a gate field is not on the run record, the gate is theater.

CTA

Stop shipping on a flattering percentage.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

Pass@1 vs Pass@k — which ships?
Ship on Pass@1 with a bounded revision ceiling. Use Pass@k for model comparisons and research-style evals. Production jobs rarely get *k* free attempts, so Pass@k overstates reliability for autonomy decisions. If you need a consistency number, look at pass^k — the chance *every* trial works — not the chance that one of them did.
What is eval coverage?
It is how much of real production job diversity your golden set and evaluators actually touch. High pass on a narrow set is not readiness — track uncovered job shapes and add cases as they appear online. A practical report is the share of last-30-day production fingerprints that match at least one golden case family.
Cost per successful task vs cost per run?
Cost per run includes cheap failures and hides grind. Cost per successful task (ideally per task shipped without human rewrite) is the unit economic signal that should sit beside pass rate on the gate panel. A common form is `cost_all / successes`; the stricter form divides by clean ships and adds human time.
How do online and offline scores diverge?
Offline sets use stubs and frozen distributions; production drifts. Promote only when online samples through the same evaluator stay within tolerance of offline — a widening gap is a release smell, not noise. Feed online misses back into the golden set or you will keep passing a test the world already left.
Which metrics gate a deploy?
At minimum: offline pass@1 delta, revision depth band, trajectory checklist, coverage not reduced, cost per success, and a post-canary online hold on pass and rewrite rate. Pass rate alone never ships. Write the bands before the canary, not after a pretty CI run.
How does this connect to the observability dashboard without duplicating it?
Observability emits the fields (scores, revisions, trajectory fails, cost). This spoke sets the veto thresholds and weekly panel. One telemetry path, one release checklist — not two competing dashboards. If a gate field is not on the run record, the gate is theater.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot