What broke when I tried to evaluate an AI agent in production
Production eval failed on four fronts: no gold labels, biased samples, judge drift, eval tools that wrote. Isolate the harness before you trust the live score.
William Spurlock Founder — Spurlock Studios 29 MIN
What broke when I tried to evaluate an AI agent in production was the measurement, not the model. The score went green while humans still rewrote the work. Four failures stacked: I had no golden labels, I sampled the wrong traffic, the judge drifted, and the eval harness itself called write tools. A live dashboard without those four locked down is a second agent you forgot to sandbox.
This spoke sits under the Agentic Systems Operating Manual. Build order — criteria, independent judge, revision ceiling — lives in evaluators before agents. How to calibrate the scorer lives in LLM-as-judge reliability. Demo-to-prod control gaps live in why agent demos fail production. This page owns the night the eval hit live traffic and lied.
The short answer
- Treat production eval as a second system with its own blast radius. If it can write, it is not an eval.
- Label a stratified gold slice before the online judge ships. A judge with no gold is a vibe with an API bill.
- Sample by job type, tool class, and time window — not by whichever traces finished before the cron.
- Re-score a frozen slice when the worker pin, judge pin, or rubric moves. Pass-rate climbs are not proof.
- Stub tools for replay. Score the stored trace. Do not re-execute refunds, sends, or CRM writes to “check quality.”
What did “evaluate in production” actually mean?
I meant: sample live runs, score them, and use the number to decide whether writes could stay on. I did not mean a CI suite on fixtures. I did not mean a human sitting on every ticket. I meant an online loop.
That loop has four objects. Confusing them is how the number lies.
| Object | Job | Not the job |
|---|---|---|
| Worker run | Do the customer task | Grade itself |
| Stored trace | Evidence of what happened | A license to re-hit live APIs |
| Offline golden | Repeatable gate on known cases | A substitute for live mix |
| Online sample | Watch the mix you did not write fixtures for | Ground truth without labels |
OpenAI’s agent evals and trace grading treat the trace as the unit you grade — model calls, tool calls, guardrails, handoffs — then tell you to promote graded traces into a dataset when you need repeatability. Anthropic’s Demystifying evals for AI agents splits the same object into a task, a grader, a transcript, and an outcome. The outcome is the state of the world, not the last sentence the model uttered.
Microsoft Foundry’s agent evaluators make the vendor cut: system (did the job finish with a usable deliverable) versus process (did the steps stay inside the contract). I needed both. I shipped a chat score and called it production eval.
- I can name the population the sample is supposed to represent
- Scoring reads a stored trace; it does not re-call write tools
- Offline gold and online sample are different objects with different owners
- “Pass” names a criterion, not a thumbs-up
- The eval identity (
judge_pin,rubric_id, sample policy) is logged on every scored run
If the last box is empty, you cannot explain next week’s number. You have a chart.
What broke first: I had no golden labels
I turned on an LLM judge and treated its scores as labels. That is circular. The judge is a second model call. Without human gold, “agreement” is the judge agreeing with itself.
LangChain’s agent-eval note is blunt about the bottleneck: humans review on the order of 50–100 traces per hour, so at production volume you cannot label everything. Judges fill the gap. They also list scoring drift as a failure mode and tell you to start around 20+ labeled examples and grow. OpenAI’s evaluation best practices put the same rule in the anti-pattern list: ignoring human feedback — not calibrating automated metrics against human evals.
I skipped the labels because they felt slow. The dashboard arrived in an afternoon. The lying took a week to notice.
| What I had | What I thought it was | What it actually was |
|---|---|---|
| Judge prompt + rubric paragraph | Ground truth | Another stochastic grader |
| High online pass rate | The agent got good | The judge was lenient, or the sample was easy |
| Human rewrite queue still full | “Support is picky” | False passes the judge never saw |
| No frozen gold slice | “We’ll add labels later” | No way to detect judge drift |
| Worker and judge in the same family | Cost savings | Shared blind spots |
Procedure I now refuse to skip:
- Write pass/fail criteria the buyer will sign. Mechanical checks first.
- Pull 30–50 traces stratified by job type and difficulty — including escalations, not only completions.
- Blind-label with humans on the same package the judge will see. Two raters on the hard slice.
- Run the judge. Store the confusion matrix. The dangerous cell is judge-pass / human-fail.
- Freeze gold-slice ids. Never use them for prompt fiddling.
- Only then let the judge score a live sample.
How to design that calibration is LLM-as-judge reliability. This page’s point is narrower: if step 3 never happened, production eval is theater. You cannot measure whether the eval is working, because you have no independent axis.
A labeled set of fifty hard cases beats an unlabeled firehose of five thousand.
What “no labels” looks like in the first week, so you can recognize it from Slack instead of from a postmortem:
| Slack sentence | Translation |
|---|---|
| “The judge is pretty aligned with what we’d say” | Nobody labeled a stratified set |
| “We’ll use thumbs-down as gold” | Selection bias: silent failures never vote |
| “Pass rate is 90, we’re good to send” | The judge is the only axis |
| “Support is still rewriting, but eval is green” | False-pass cell is the product |
| “We can label after launch” | You launched without a measurement |
Thumbs and CSAT are not golden labels. They are sparse, skewed toward people who bother, and they arrive after the write. OpenAI’s own guidance still says to maintain agreement with humans and to mine logs for cases — not to skip the humans because a model is cheaper. Sparse feedback as a quality estimator is a known selection-bias problem; I do not treat a 5% thumbs-down stream as a calibrated gold set.
Failure mode: sampling bias
The second break was who got into the sample. I scored whatever the nightly job grabbed. Completions. Short traces. Overnight US traffic. The money jobs, the angry jobs, and the traces that hit the escalate path were under-counted.
OpenAI’s anti-pattern is exactly this: biased design — eval datasets that do not faithfully reproduce production traffic patterns. “It seems like it’s working” is on the same list. I hit both in one cron.
| Bias | How it got into my sample | What the score hid |
|---|---|---|
| Completion bias | Dropped traces that escalated or timed out | The jobs that already failed |
| Time-of-day bias | Cron at 03:00 Eastern on last night’s window | Daytime mix, other locales |
| Length bias | Truncated long tool traces to save judge tokens | The runs where agents actually do damage |
| Volume bias | Sampled proportional to count, not blast radius | Refund and PII jobs drowned in FAQs |
| Success bias | Preferred traces with a final artifact | Partial writes with no “done” message |
| New-tool blindness | Strata frozen before the last schema add | Silent false passes on a new write |
What I do instead:
- Name the population: last 7 days, all
job_typevalues, including escalate and budget-stop. - Stratify. Floor per stratum — happy, auth fail, empty tool, policy deny, money, ambiguous — even if that means oversampling a rare refund path.
- Reweight when you report a global mean. Or skip the global mean and report per stratum.
- Keep a mix histogram next to the pass rate. If the sample’s
job_typebar chart moved, the number is not comparable. - Never drop a trace because it is long, ugly, or incomplete. Those are the ones the demo never showed.
- Sample window is at least one full week, not one overnight dump
- Escalated and timed-out runs are in the frame, not filtered as “invalid”
- Each irreversible tool has a minimum count in the sample or a named waiver
- Mix (
job_type, locale, channel) is stored on every scored row - Global pass rate is labeled “reweighted” or is not shown
A 94% on last night’s completed chats is not a production eval. It is a weather report from a street you like.
Coverage I now require before I believe an online number. Count filled cells, not traces.
| Stratum | Completions | Escalations | Empty-tool | Money / PII |
|---|---|---|---|---|
| FAQ / lookup | Must appear | Must appear | Must appear | n/a unless the tool exists |
| Account change | Must appear | Must appear | Must appear | If the job can mutate |
| Refund / payment | Oversample | Oversample | Must appear | Required |
| Ambiguous / missing policy | Must appear | Must appear | Must appear | If a write is possible |
A column of zeros is a named hole. Name it, waive it in writing, or keep that write tool off. Do not average it away.
If two strata moved more than the others week over week, I do not compare pass rates until I reweight. OpenAI’s “real-world distributions” line is the whole job. A sample that cannot name its distribution is a biased design with a prettier query.
Failure mode: judge drift
The third break: worker prompts unchanged, online pass rate climbed, humans still rewrote. The scorer moved.
LangChain says to recalibrate regularly, because judges drift just like the agents they evaluate. Their monitoring note treats online judges as a smoke detector on sampled traffic — useful only if you keep routing disagreements into an annotation queue. I had the detector. I never checked whether the detector was still aimed at the fire.
Triggers that forced a re-score of the frozen slice, once I had one:
- Judge model pin or provider snapshot changed
- Rubric “clarification” shipped (a version bump even when the English looks tiny)
- Worker model upgraded — the trace distribution moved
- New tool or side-effect class added
- Human override rate diverged from judge pass rate for two weeks
- Judge JSON fail rate or empty-evidence rate jumped
| Signal | Read as | Do not read as |
|---|---|---|
| Gold-slice agreement holds, online rewrite rate up | Worker or mix moved; judge may still be fine | “Quality improved” |
| Gold-slice false-pass up | Judge got lenient, or rubric got vague | A model win |
| Gold-slice false-fail up | Judge got harsh; check pin and prompt | The agent collapsed |
| Online pass up, gold slice untouched | Sample mix or judge pin moved | Proof you can widen writes |
| Judge evidence slot empty | Harness bug | A quality signal |
Zheng et al. documented position bias, verbosity bias, and self-enhancement on chat judges (MT-Bench). Agent traces make the costume worse: a three-paragraph “I verified in the CRM” speech with no ledger span. I watched verbosity win. The judge loved essays. The humans wanted the correct tool call.
I will not copy a paper’s headline agreement rate into a runbook. Complementary calibration — human agreement, false-pass bounds, gold-slice ritual — is owned by LLM-as-judge reliability. Here the production receipt is simpler: if you cannot query last week’s judge-pass / human-fail, you will learn about drift from a customer.
Judge spans belong on the same run id as the tool calls. A score without a pin is a rumor.
Failure mode: eval tools with side effects
The break that still makes me angry: the measurement wrote.
I wanted “outcome eval” — did the ticket actually close, did the note land on the right account. So the eval agent got tools. Some of those tools were the same write surface as the worker. Replay meant “run it again.” Running it again meant a second CRM note, a second email, a second draft refund in a staging project that was not as staging as the ticket claimed.
Anthropic’s split is the warning I ignored: the outcome is the state of the world. If your grader changes that state, you are not measuring. You are acting. OpenAI’s graders warning about grader hacking is the cousin: a policy that scores high on the grader and low with experts. Side-effect eval is worse than hacking. It is a second production actor with a research excuse.
| Eval move | Side effect I have seen | Isolation |
|---|---|---|
| Re-run the worker on a live trace | Duplicate send, duplicate CRM write, double charge | Replay against stubs; never re-call writes |
| Judge with a search or ticket-read tool | Rate limits, noisy audit log, accidental write if the client SDK is fat | Read-only credentials; no send/refund/create in the judge role |
| Shadow traffic to the real endpoint | Customers see two agents; idempotency keys collide | Shadow against a sandbox tenant |
| Fixture that “checks the email arrived” by sending one | You sent the email | Assert on the stored outbound payload |
| Eval job using the worker’s API key | Same blast radius as the agent | Separate principal, deny-by-default writes |
| Sampling that retries failed tools “to be fair” | You just became the retry storm | Score the original ledger |
OWASP’s LLM06 Excessive Agency is about workers. It applies to graders that inherited the same tools. Least privilege is not a slogan for the agent only.
Hard rules I now write on the eval service account before anyone adds a “helpful” tool:
- The evaluator does not call write tools. Ever. Same line as evaluators before agents.
- Replay uses recorded tool results. Missing results are a harness fail, not a license to hit production.
- If you need world-state, read a replica or an anonymized snapshot, not the live writable tenant.
- Idempotency keys on the worker do not protect you from a second actor. The eval is a second actor.
- A “dry run” flag in a prompt is not isolation. Isolation is credentials and network policy.
- Eval principal cannot
create,send,refund,delete, orupdateon production - Replay fixtures stub HTTP; CI has no CRM/SMTP/billing credentials
- Shadow jobs have a tenant allowlist that is not a customer
- Judge evidence package is a digest, not a live tool loop
- Side-effect diffs are computed from the stored ledger, not from a second write
If the eval can email, you do not have an eval. You have a quieter intern with a cron.
Decision tree before any “replay” button exists:
- Is the original tool ledger complete? If no, fail the harness. Do not fill gaps live.
- Would this call create, update, send, refund, or delete? If yes, stub or deny.
- Do you need world-state? Read a snapshot keyed by the original run id.
- Still tempted to hit production “just this once”? That is the incident. Stop.
| Temptation | What people say | What to do |
|---|---|---|
| “Confirm the ticket closed” | Outcome eval | Read the stored ticket.status on the snapshot |
| “See if the email looks right” | Qualitative check | Grade the stored outbound MIME / payload |
| “The sandbox is stale” | Need fresh data | Refresh the snapshot on a schedule; do not punch through |
| “Shadow 1% of traffic” | Safe canary | Shadow a non-customer tenant, or do not shadow writes |
| “Judge should browse the order” | Evidence | Put the redacted order JSON in the package |
Replay is a read of history. The moment it becomes a write, you left eval and entered production with worse observability.
Worked example: the support-draft eval that sent twice
Job: draft a support reply, optionally send. Worker had tickets.reply behind an approval flag. I added an online judge “to watch quality after send.” The judge got a tickets.get tool so it could “see the thread.” The SDK’s default client was the same bot user as the worker. One helper on that client was reply. Someone wired “fetch thread” as a generic run_tool(name, args) and the judge, looking for evidence, posted a second reply that quoted its own rubric.
No customer name. No invented ticket volume. The shape is enough: eval principal inherited write methods, and a tool-loop judge treated production as a retrieval index.
| Layer | What I shipped | What I should have shipped |
|---|---|---|
| Credentials | Worker bot token reused | Eval token, tickets.get only, no reply |
| Judge package | “Go look at the ticket” | Redacted thread digest already on the trace |
| Replay | Re-run worker on sampled ids | Score stored artifacts + ledger |
| Labels | None the first week | 40 stratified drafts, two raters on the “send” slice |
| Sample | Last 200 sent tickets | Sent + unsent + escalate, seven days, by queue |
| Dashboard | Pass rate on sent tickets | False-pass vs humans; mix histogram; gold slice |
What broke, mapped to the four failures:
- No gold. The judge learned that long, careful-sounding replies were “good.” Humans still rewrote short correct ones that cited the order id.
- Sampling. Unsent drafts and escalations never entered the sample, so the send-path bugs were the whole chart and still looked fine.
- Drift. A “clarified” rubric added “be thorough.” Verbosity bias got a permission slip. Pass rate climbed.
- Side effects. The second reply was the eval. Customers saw a duplicate. Idempotency on the worker did not apply to a second principal.
Fix that actually held:
- Eval token scoped to read
- Judge package is a digest; zero tools on the judge
- Gold slice includes short-correct and long-wrong exemplars
- Sample includes unsent and escalate
- Duplicate-outbound detector on
thread_id+eval_run_id
The send flag stayed behind approval until those boxes were checked. That is the opposite of “we’ll eval in production and then open the gate.”
Why did the dashboard stay green?
Because every lying number had a friend. Sampling bias fed the judge easy traces. The uncalibrated judge passed essays. Drift made the pass rate climb. Side-effect replay created “successful” world states that were actually duplicates. The chart was consistent. It was consistently wrong.
This is the production costume of the pass-rate lie. The metric panel — revision rate, trajectory, coverage, cost per success — is owned elsewhere. Here I only need the eval-side version:
| Green number | How it was manufactured | What to show instead |
|---|---|---|
| Online pass rate | Easy-stratum sample + lenient judge | Hard-stratum pass + false-pass vs humans |
| “N traces scored” | Volume, not coverage | Filled cells in the mix × tool matrix |
| Judge latency down | Truncated traces | Evidence-quote rate on full ledgers |
| Cost per eval down | Same-family judge, no gold | Cost per labeled disagreement |
| Week-over-week lift | Rubric loosened or mix shifted | Gold-slice delta with pins frozen |
OpenAI’s best-practices line still holds: combine metrics with human judgment so you are answering the right questions. A single online percentage answers a question nobody asked: “did the judge like this week’s easy completions?”
I now refuse to put online pass rate on a ship/no-ship toggle unless the gold-slice false-pass bound is next to it. An advisory score drawn as a gate is how teams get surprised.
Offline vs online: the split I skipped
I tried to make one loop do both jobs. CI wanted freeze. Production wanted mix. The compromise was a dataset that thawed every night and a “golden” that was just yesterday’s sample with the labels still missing.
LangChain’s improvement-loop writeup draws the split I needed: online evaluators watch live behavior and flag traces; offline evaluations are controlled experiments on curated datasets before a change ships. Online tells you what is going wrong. Offline tells you whether the fix actually addressed it.
| Offline golden | Online sample | |
|---|---|---|
| Source | Harvested failures + signed fixtures | Stratified live traces |
| Tools | Stubbed | Already happened; do not re-run writes |
| Labels | Required | Sparse; overlay judge + human queue |
| Cadence | Every behavior change | Rolling window |
| Decision | Block a merge | Alert, harvest, do not auto-widen writes |
| Failure if skipped | You ship regressions you already paid for | You miss mix the fixtures never covered |
Procedure that finally held:
- Offline suite is append-only except for documented deletions. Pass-rate jumps after you deleted hard cases are rot.
- Online sample never mutates the golden set automatically. A human promotes a trace into a fixture.
- A PR can fail on offline gold. A PR cannot “pass” because online pass ticked up.
- Disagreements (judge-pass / human-fail) enter the annotation queue the same day, not the next quarter.
-
dataset_idfor gold is different fromsample_idfor online - Promotion to gold requires a stubbed tool world and an expected terminal
- Online alerts page a human; they do not flip a write allowlist
- The last three harvested misses have
case_ids in git or a signed waiver
Harvest — how a bad run becomes a fixture — is a different spoke. This page only insists the two loops do not share a write key.
How do I evaluate this in production without lying?
Target query, operational version: score a stratified sample of stored traces, against criteria you already froze, with a judge you have calibrated on gold, with a principal that cannot write.
| Step | Owner | Done when |
|---|---|---|
1. Freeze criteria + rubric_id | Job owner + eng | Buyer signed pass/fail lines |
| 2. Mechanical graders on the ledger | Eng | Schema, allowlist, ids, budgets never hit the judge |
| 3. Label a gold slice | Job owner | Confusion matrix exists; false-pass bound written |
| 4. Sample policy | Eng | Strata, window, reweight rule in config, not a comment |
| 5. Score stored traces | Eval service | Judge pin + rubric on every span |
| 6. Annotation queue | Job owner | Disagreements SLA in hours, not “when we can” |
| 7. Harvest | Eng | Fixture stubs, not live replay |
Guardrails that belong on that loop:
- Separate eval principal. Deny writes at the token.
- Cap judge package size. A 50k-token chat dump grades prose. A one-page ledger with ids the judge must quote grades actions.
- Independent scoring for revisions. Position bias is not a chatbot-only problem.
- Evidence quotes tied to tool ids. “Looks careful” is not a criterion.
- Pin worker, judge, tools, and sample policy. Log the identity line on every run.
If step 3 is missing, stop. You can still collect traces. You cannot trust scores.
Evidence package for an online row — minimum, not a chat dump:
| Field | Why |
|---|---|
| Job goal + acceptance lines | The question the run was supposed to answer |
rubric_id + criterion list | Same text the humans used |
| Final artifacts | What would ship |
| Tool ledger digest | Names, redacted args, result ids, timestamps |
| Harness terminal reason | pass / revise / escalate / budget |
| Sample metadata | job_type, locale, channel, window, stratum |
| Identity line | Worker pin, judge pin, tools version, sample policy |
Leave out worker chain-of-thought, secrets, full RAG dumps, and prior judge rationales. Default no on giving the judge live tools. If a field is missing, the row is unscorable — that is a harness incident, not a 0.0 quality score.
What guardrails do I need on the eval itself?
The worker needs a sandbox. The eval needs one too. Why agent demos fail production owns schemas, auth bleed, and kill switches for the agent. Copy the furniture onto the grader.
| Guardrail | Worker | Eval harness |
|---|---|---|
| Tool allowlist | Job-scoped writes | Read-only or none |
| Tenant bind | Customer tenant | Snapshot / sandbox / stubs |
| Kill switch | Freeze writes without a prompt deploy | Disable scoring and replay without a prompt deploy |
| Identity log | model_id, tools_version | Those plus judge_pin, rubric_id, sample policy |
| Budget | Turn cap, dollar cap | Judge-token cap so you stop truncating the hard traces |
| Auth | Per-tenant credentials | Eval credentials that cannot mint worker tokens |
Checklist before the first live sample:
- Eval service account reviewed like a production actor
- Network policy: no SMTP, no billing write, no CRM mutate
- Replay path unit-tested with a fixture that would have written, and did not
- Kill switch is a config flag, not a judge-prompt paragraph
- PII redaction on packages that leave the tenant boundary (judge vendor, logs, gold store)
A prompt that says “you are only evaluating” is not a guardrail. Credentials are.
When is a workflow enough instead of an agent?
When you cannot isolate the eval, you cannot earn the loop.
If the job is a known path with known tools, a workflow already gives you deterministic tests: assert the payload, assert the idempotency key, assert the destination. You do not need a live judge to tell you a mapping job worked. You need a fixture.
| You still want an agent | You want a workflow |
|---|---|
| The next action depends on messy evidence | The path is known before the run starts |
| You can name pass/fail criteria and label gold | You cannot get two humans to agree after seeing the ledger |
| Eval can score stored traces with stubs | The only “eval” you can imagine is re-running writes |
| Blast radius is gated (draft-only, approval on sends) | Irreversible tools with no replica and no stubs |
| Mix is wide enough that fixtures will miss it | Mix is a handful of shapes you can enumerate |
Anthropic’s Building effective agents is blunt: the evaluator-refiner pattern only fits when you have clear evaluation criteria and iterative refinement actually improves the artifact. No criteria, no loop. No isolated eval, same conclusion.
After 20,000+ hours on agentic systems and 500+ automations, the honest fork is: if production eval requires a second copy of the blast radius, you are not ready for an agent. Keep the workflow. Collect traces in draft mode. Build gold. Then argue about autonomy.
A demo that cannot be evaluated without sending a real email was never a demo of an agent. It was a demo of a send button.
One-week honesty test I run with the buyer:
- Can we stub the irreversible tools today? If no, the agent stays draft-only.
- Can two humans label twenty hard traces without splitting on the rubric? If no, workshop criteria; do not hire a judge.
- Can the eval principal be denied writes in IAM, not in a prompt? If no, you do not have a harness.
- Is the job a lookup-plus-map that a workflow already does? If yes, stop the agent ticket.
- Draft-only worker is acceptable for a week of harvest
- “We’ll be careful in production” is not a stub
- A workflow spike is a valid outcome of an agentic pilot, not a failed sale
If those four answers are no, the week should produce a workflow sketch and a criteria workshop, not an online leaderboard.
What I freeze before the next live sample
This is the packing list I use on a Spurlock Studios agentic pilot. Five days will not finish an academic inter-annotator study. Five days will tell you whether the measurement is usable or dangerous.
| Day | Freeze | Do not do |
|---|---|---|
| 1 | Criteria split: mechanical vs judge. Eval principal with writes denied. | Give the judge the worker’s tools “just for outcome checks” |
| 2 | Sample policy: strata, 7-day window, mix histogram. Pull 30–50 traces. | Score last night’s completions and call it a baseline |
| 3 | Human labels on the hard slice. First confusion matrix. | Ship the online judge because the prompt “looks strict” |
| 4 | Stubbed replay on three known-bad traces. Prove no live writes. | Re-run the worker against production to “verify the note landed” |
| 5 | Gold-slice ids frozen. Judge span on run id. False-pass bound written next to blast radius. | Put online pass rate on a ship toggle |
What “done” means on day five:
- Mechanical checks no longer go through the judge
- Eval credentials cannot write
- Gold-slice ids exist and are not used for prompt fiddling
- Sample mix is queryable next to the score
- A named owner for the next re-calibration trigger
- Three harvested misses have a path into fixtures, not into folklore
That is the same discipline as the rest of the control plane. The scorer is a component with a version. The sample is a component with a policy. The eval principal is a component with a deny list. None of those are a paragraph in a system prompt.
Anti-patterns that recreate the original incident:
| Anti-pattern | What it costs |
|---|---|
| Online pass rate as the ship toggle | You widen writes on a lenient week |
| Gold slice = last week’s happy paths | Easy-case accuracy in a new costume |
| Judge tools “for evidence” | Side effects; you are running a second agent |
| Sample = completed chats after 10pm | Mix lie |
| Rubric edits with no version bump | Silent drift |
| Same token for worker and eval | Shared blast radius |
| Deleting hard fixtures to “clean the suite” | Pass-rate rot |
If an anti-pattern is already in the dashboard, turn writes down before you argue about models.
How this spoke sits next to the cousins
Do not collapse four posts into one dashboard.
| Question | This spoke | Somewhere else |
|---|---|---|
| What broke when live eval hit production? | Labels, sampling, drift, eval side effects | — |
| What do I build first? | Assume you already believe this | Evaluators before agents |
| Is the judge calibrated? | I needed it; I did not have it | LLM-as-judge reliability |
| Why did the demo die? | Not this page | Why agent demos fail production |
| What else has to exist? | Eval as a second sandboxed system | Operating manual |
If you skip the first column and only add an online judge, you have a second agent with a clipboard. The clipboard will write if you let it.
The only question here is: when production eval says pass, did the measurement survive contact with live tools — or did you just run the agent twice?
FAQ
What broke when I tried to evaluate an AI agent in production?
The measurement: no golden labels, a biased sample of easy completions, a drifting judge, and an eval harness that called write tools. The judge graded into a void while humans still rewrote. Scoring mutated the world it was supposed to observe. Fix those four before you trust a live score.
How do I measure whether broke when I tried to evaluate an AI agent in production is working?
Look at gold-slice false-pass versus humans, mix histograms next to the online number, and whether replay can mutate production. If judge-pass / human-fail is unqueryable, the eval is not working. A rising online pass rate with a frozen gold slice and a stable mix is the only lift I believe.
What usually fails first when teams try this?
Golden labels fail first: teams ship a judge prompt in an afternoon and skip 30–50 stratified human labels. Sampling bias is second — overnight completions, no escalations. Eval side effects show up the first time someone “replays” a write, and drift shows up later, which is why it gets blamed on the model.
How long does this take to show results?
A five-day pilot can tell you whether the harness is usable or dangerous: criteria split, denied writes, a labeled slice, stubbed replay, frozen gold ids. Trustworthy online numbers take longer — you need a week of stratified mix and an annotation SLA, not a weekend of green charts. I will not invent a universal week-three accuracy number.
What should I skip if I only have a week?
Skip widening writes, a same-family judge as the only gate, and re-running the worker to “check outcomes.” Do not skip labels, sample strata, or eval-credential deny lists. A thin week that freezes those three beats a theatrical week that scores five thousand unlabeled chats.
When is this not worth doing yet?
When two trained humans cannot agree on the ledger, when you have no replica or stubs for irreversible tools, or when the job is a known path a workflow already covers. Online eval of an agent you cannot isolate is how you buy a second copy of the blast radius. Keep writes off, or stay on a workflow, until gold and isolation exist.
CTA
Sandbox the measurement before you trust the live score.
What questions does this article answer?
- What broke when I tried to evaluate an AI agent in production?
- The measurement: no golden labels, a biased sample of easy completions, a drifting judge, and an eval harness that called write tools. The judge graded into a void while humans still rewrote. Scoring mutated the world it was supposed to observe. Fix those four before you trust a live score.
- How do I measure whether broke when I tried to evaluate an AI agent in production is working?
- Look at gold-slice false-pass versus humans, mix histograms next to the online number, and whether replay can mutate production. If judge-pass / human-fail is unqueryable, the eval is not working. A rising online pass rate with a frozen gold slice and a stable mix is the only lift I believe.
- What usually fails first when teams try this?
- Golden labels fail first: teams ship a judge prompt in an afternoon and skip 30–50 stratified human labels. Sampling bias is second — overnight completions, no escalations. Eval side effects show up the first time someone “replays” a write, and drift shows up later, which is why it gets blamed on the model.
- How long does this take to show results?
- A five-day pilot can tell you whether the harness is usable or dangerous: criteria split, denied writes, a labeled slice, stubbed replay, frozen gold ids. Trustworthy online numbers take longer — you need a week of stratified mix and an annotation SLA, not a weekend of green charts. I will not invent a universal week-three accuracy number.
- What should I skip if I only have a week?
- Skip widening writes, a same-family judge as the only gate, and re-running the worker to “check outcomes.” Do not skip labels, sample strata, or eval-credential deny lists. A thin week that freezes those three beats a theatrical week that scores five thousand unlabeled chats.
- When is this not worth doing yet?
- When two trained humans cannot agree on the ledger, when you have no replica or stubs for irreversible tools, or when the job is a known path a workflow already covers. Online eval of an agent you cannot isolate is how you buy a second copy of the blast radius. Keep writes off, or stay on a workflow, until gold and isolation exist.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.