Spurlock Studios
Contact
Share LinkedIn X
A cracked amber fuse. Thesis: BROKE TRIED EVALUATE AI AGENT.

What broke when I tried to evaluate an AI agent in production was the measurement, not the model. The score went green while humans still rewrote the work. Four failures stacked: I had no golden labels, I sampled the wrong traffic, the judge drifted, and the eval harness itself called write tools. A live dashboard without those four locked down is a second agent you forgot to sandbox.

This spoke sits under the Agentic Systems Operating Manual. Build order — criteria, independent judge, revision ceiling — lives in evaluators before agents. How to calibrate the scorer lives in LLM-as-judge reliability. Demo-to-prod control gaps live in why agent demos fail production. This page owns the night the eval hit live traffic and lied.

The short answer

  • Treat production eval as a second system with its own blast radius. If it can write, it is not an eval.
  • Label a stratified gold slice before the online judge ships. A judge with no gold is a vibe with an API bill.
  • Sample by job type, tool class, and time window — not by whichever traces finished before the cron.
  • Re-score a frozen slice when the worker pin, judge pin, or rubric moves. Pass-rate climbs are not proof.
  • Stub tools for replay. Score the stored trace. Do not re-execute refunds, sends, or CRM writes to “check quality.”

What did “evaluate in production” actually mean?

I meant: sample live runs, score them, and use the number to decide whether writes could stay on. I did not mean a CI suite on fixtures. I did not mean a human sitting on every ticket. I meant an online loop.

That loop has four objects. Confusing them is how the number lies.

ObjectJobNot the job
Worker runDo the customer taskGrade itself
Stored traceEvidence of what happenedA license to re-hit live APIs
Offline goldenRepeatable gate on known casesA substitute for live mix
Online sampleWatch the mix you did not write fixtures forGround truth without labels

OpenAI’s agent evals and trace grading treat the trace as the unit you grade — model calls, tool calls, guardrails, handoffs — then tell you to promote graded traces into a dataset when you need repeatability. Anthropic’s Demystifying evals for AI agents splits the same object into a task, a grader, a transcript, and an outcome. The outcome is the state of the world, not the last sentence the model uttered.

Microsoft Foundry’s agent evaluators make the vendor cut: system (did the job finish with a usable deliverable) versus process (did the steps stay inside the contract). I needed both. I shipped a chat score and called it production eval.

  • I can name the population the sample is supposed to represent
  • Scoring reads a stored trace; it does not re-call write tools
  • Offline gold and online sample are different objects with different owners
  • “Pass” names a criterion, not a thumbs-up
  • The eval identity (judge_pin, rubric_id, sample policy) is logged on every scored run

If the last box is empty, you cannot explain next week’s number. You have a chart.

What broke first: I had no golden labels

I turned on an LLM judge and treated its scores as labels. That is circular. The judge is a second model call. Without human gold, “agreement” is the judge agreeing with itself.

LangChain’s agent-eval note is blunt about the bottleneck: humans review on the order of 50–100 traces per hour, so at production volume you cannot label everything. Judges fill the gap. They also list scoring drift as a failure mode and tell you to start around 20+ labeled examples and grow. OpenAI’s evaluation best practices put the same rule in the anti-pattern list: ignoring human feedback — not calibrating automated metrics against human evals.

I skipped the labels because they felt slow. The dashboard arrived in an afternoon. The lying took a week to notice.

What I hadWhat I thought it wasWhat it actually was
Judge prompt + rubric paragraphGround truthAnother stochastic grader
High online pass rateThe agent got goodThe judge was lenient, or the sample was easy
Human rewrite queue still full“Support is picky”False passes the judge never saw
No frozen gold slice“We’ll add labels later”No way to detect judge drift
Worker and judge in the same familyCost savingsShared blind spots

Procedure I now refuse to skip:

  1. Write pass/fail criteria the buyer will sign. Mechanical checks first.
  2. Pull 30–50 traces stratified by job type and difficulty — including escalations, not only completions.
  3. Blind-label with humans on the same package the judge will see. Two raters on the hard slice.
  4. Run the judge. Store the confusion matrix. The dangerous cell is judge-pass / human-fail.
  5. Freeze gold-slice ids. Never use them for prompt fiddling.
  6. Only then let the judge score a live sample.

How to design that calibration is LLM-as-judge reliability. This page’s point is narrower: if step 3 never happened, production eval is theater. You cannot measure whether the eval is working, because you have no independent axis.

A labeled set of fifty hard cases beats an unlabeled firehose of five thousand.

What “no labels” looks like in the first week, so you can recognize it from Slack instead of from a postmortem:

Slack sentenceTranslation
“The judge is pretty aligned with what we’d say”Nobody labeled a stratified set
“We’ll use thumbs-down as gold”Selection bias: silent failures never vote
“Pass rate is 90, we’re good to send”The judge is the only axis
“Support is still rewriting, but eval is green”False-pass cell is the product
“We can label after launch”You launched without a measurement

Thumbs and CSAT are not golden labels. They are sparse, skewed toward people who bother, and they arrive after the write. OpenAI’s own guidance still says to maintain agreement with humans and to mine logs for cases — not to skip the humans because a model is cheaper. Sparse feedback as a quality estimator is a known selection-bias problem; I do not treat a 5% thumbs-down stream as a calibrated gold set.

Failure mode: sampling bias

The second break was who got into the sample. I scored whatever the nightly job grabbed. Completions. Short traces. Overnight US traffic. The money jobs, the angry jobs, and the traces that hit the escalate path were under-counted.

OpenAI’s anti-pattern is exactly this: biased design — eval datasets that do not faithfully reproduce production traffic patterns. “It seems like it’s working” is on the same list. I hit both in one cron.

BiasHow it got into my sampleWhat the score hid
Completion biasDropped traces that escalated or timed outThe jobs that already failed
Time-of-day biasCron at 03:00 Eastern on last night’s windowDaytime mix, other locales
Length biasTruncated long tool traces to save judge tokensThe runs where agents actually do damage
Volume biasSampled proportional to count, not blast radiusRefund and PII jobs drowned in FAQs
Success biasPreferred traces with a final artifactPartial writes with no “done” message
New-tool blindnessStrata frozen before the last schema addSilent false passes on a new write

What I do instead:

  1. Name the population: last 7 days, all job_type values, including escalate and budget-stop.
  2. Stratify. Floor per stratum — happy, auth fail, empty tool, policy deny, money, ambiguous — even if that means oversampling a rare refund path.
  3. Reweight when you report a global mean. Or skip the global mean and report per stratum.
  4. Keep a mix histogram next to the pass rate. If the sample’s job_type bar chart moved, the number is not comparable.
  5. Never drop a trace because it is long, ugly, or incomplete. Those are the ones the demo never showed.
  • Sample window is at least one full week, not one overnight dump
  • Escalated and timed-out runs are in the frame, not filtered as “invalid”
  • Each irreversible tool has a minimum count in the sample or a named waiver
  • Mix (job_type, locale, channel) is stored on every scored row
  • Global pass rate is labeled “reweighted” or is not shown

A 94% on last night’s completed chats is not a production eval. It is a weather report from a street you like.

Coverage I now require before I believe an online number. Count filled cells, not traces.

StratumCompletionsEscalationsEmpty-toolMoney / PII
FAQ / lookupMust appearMust appearMust appearn/a unless the tool exists
Account changeMust appearMust appearMust appearIf the job can mutate
Refund / paymentOversampleOversampleMust appearRequired
Ambiguous / missing policyMust appearMust appearMust appearIf a write is possible

A column of zeros is a named hole. Name it, waive it in writing, or keep that write tool off. Do not average it away.

If two strata moved more than the others week over week, I do not compare pass rates until I reweight. OpenAI’s “real-world distributions” line is the whole job. A sample that cannot name its distribution is a biased design with a prettier query.

Failure mode: judge drift

The third break: worker prompts unchanged, online pass rate climbed, humans still rewrote. The scorer moved.

LangChain says to recalibrate regularly, because judges drift just like the agents they evaluate. Their monitoring note treats online judges as a smoke detector on sampled traffic — useful only if you keep routing disagreements into an annotation queue. I had the detector. I never checked whether the detector was still aimed at the fire.

Triggers that forced a re-score of the frozen slice, once I had one:

  • Judge model pin or provider snapshot changed
  • Rubric “clarification” shipped (a version bump even when the English looks tiny)
  • Worker model upgraded — the trace distribution moved
  • New tool or side-effect class added
  • Human override rate diverged from judge pass rate for two weeks
  • Judge JSON fail rate or empty-evidence rate jumped
SignalRead asDo not read as
Gold-slice agreement holds, online rewrite rate upWorker or mix moved; judge may still be fine“Quality improved”
Gold-slice false-pass upJudge got lenient, or rubric got vagueA model win
Gold-slice false-fail upJudge got harsh; check pin and promptThe agent collapsed
Online pass up, gold slice untouchedSample mix or judge pin movedProof you can widen writes
Judge evidence slot emptyHarness bugA quality signal

Zheng et al. documented position bias, verbosity bias, and self-enhancement on chat judges (MT-Bench). Agent traces make the costume worse: a three-paragraph “I verified in the CRM” speech with no ledger span. I watched verbosity win. The judge loved essays. The humans wanted the correct tool call.

I will not copy a paper’s headline agreement rate into a runbook. Complementary calibration — human agreement, false-pass bounds, gold-slice ritual — is owned by LLM-as-judge reliability. Here the production receipt is simpler: if you cannot query last week’s judge-pass / human-fail, you will learn about drift from a customer.

Judge spans belong on the same run id as the tool calls. A score without a pin is a rumor.

Failure mode: eval tools with side effects

The break that still makes me angry: the measurement wrote.

I wanted “outcome eval” — did the ticket actually close, did the note land on the right account. So the eval agent got tools. Some of those tools were the same write surface as the worker. Replay meant “run it again.” Running it again meant a second CRM note, a second email, a second draft refund in a staging project that was not as staging as the ticket claimed.

Anthropic’s split is the warning I ignored: the outcome is the state of the world. If your grader changes that state, you are not measuring. You are acting. OpenAI’s graders warning about grader hacking is the cousin: a policy that scores high on the grader and low with experts. Side-effect eval is worse than hacking. It is a second production actor with a research excuse.

Eval moveSide effect I have seenIsolation
Re-run the worker on a live traceDuplicate send, duplicate CRM write, double chargeReplay against stubs; never re-call writes
Judge with a search or ticket-read toolRate limits, noisy audit log, accidental write if the client SDK is fatRead-only credentials; no send/refund/create in the judge role
Shadow traffic to the real endpointCustomers see two agents; idempotency keys collideShadow against a sandbox tenant
Fixture that “checks the email arrived” by sending oneYou sent the emailAssert on the stored outbound payload
Eval job using the worker’s API keySame blast radius as the agentSeparate principal, deny-by-default writes
Sampling that retries failed tools “to be fair”You just became the retry stormScore the original ledger

OWASP’s LLM06 Excessive Agency is about workers. It applies to graders that inherited the same tools. Least privilege is not a slogan for the agent only.

Hard rules I now write on the eval service account before anyone adds a “helpful” tool:

  1. The evaluator does not call write tools. Ever. Same line as evaluators before agents.
  2. Replay uses recorded tool results. Missing results are a harness fail, not a license to hit production.
  3. If you need world-state, read a replica or an anonymized snapshot, not the live writable tenant.
  4. Idempotency keys on the worker do not protect you from a second actor. The eval is a second actor.
  5. A “dry run” flag in a prompt is not isolation. Isolation is credentials and network policy.
  • Eval principal cannot create, send, refund, delete, or update on production
  • Replay fixtures stub HTTP; CI has no CRM/SMTP/billing credentials
  • Shadow jobs have a tenant allowlist that is not a customer
  • Judge evidence package is a digest, not a live tool loop
  • Side-effect diffs are computed from the stored ledger, not from a second write

If the eval can email, you do not have an eval. You have a quieter intern with a cron.

Decision tree before any “replay” button exists:

  1. Is the original tool ledger complete? If no, fail the harness. Do not fill gaps live.
  2. Would this call create, update, send, refund, or delete? If yes, stub or deny.
  3. Do you need world-state? Read a snapshot keyed by the original run id.
  4. Still tempted to hit production “just this once”? That is the incident. Stop.
TemptationWhat people sayWhat to do
“Confirm the ticket closed”Outcome evalRead the stored ticket.status on the snapshot
“See if the email looks right”Qualitative checkGrade the stored outbound MIME / payload
“The sandbox is stale”Need fresh dataRefresh the snapshot on a schedule; do not punch through
“Shadow 1% of traffic”Safe canaryShadow a non-customer tenant, or do not shadow writes
“Judge should browse the order”EvidencePut the redacted order JSON in the package

Replay is a read of history. The moment it becomes a write, you left eval and entered production with worse observability.

Worked example: the support-draft eval that sent twice

Job: draft a support reply, optionally send. Worker had tickets.reply behind an approval flag. I added an online judge “to watch quality after send.” The judge got a tickets.get tool so it could “see the thread.” The SDK’s default client was the same bot user as the worker. One helper on that client was reply. Someone wired “fetch thread” as a generic run_tool(name, args) and the judge, looking for evidence, posted a second reply that quoted its own rubric.

No customer name. No invented ticket volume. The shape is enough: eval principal inherited write methods, and a tool-loop judge treated production as a retrieval index.

LayerWhat I shippedWhat I should have shipped
CredentialsWorker bot token reusedEval token, tickets.get only, no reply
Judge package“Go look at the ticket”Redacted thread digest already on the trace
ReplayRe-run worker on sampled idsScore stored artifacts + ledger
LabelsNone the first week40 stratified drafts, two raters on the “send” slice
SampleLast 200 sent ticketsSent + unsent + escalate, seven days, by queue
DashboardPass rate on sent ticketsFalse-pass vs humans; mix histogram; gold slice

What broke, mapped to the four failures:

  1. No gold. The judge learned that long, careful-sounding replies were “good.” Humans still rewrote short correct ones that cited the order id.
  2. Sampling. Unsent drafts and escalations never entered the sample, so the send-path bugs were the whole chart and still looked fine.
  3. Drift. A “clarified” rubric added “be thorough.” Verbosity bias got a permission slip. Pass rate climbed.
  4. Side effects. The second reply was the eval. Customers saw a duplicate. Idempotency on the worker did not apply to a second principal.

Fix that actually held:

  • Eval token scoped to read
  • Judge package is a digest; zero tools on the judge
  • Gold slice includes short-correct and long-wrong exemplars
  • Sample includes unsent and escalate
  • Duplicate-outbound detector on thread_id + eval_run_id

The send flag stayed behind approval until those boxes were checked. That is the opposite of “we’ll eval in production and then open the gate.”

Why did the dashboard stay green?

Because every lying number had a friend. Sampling bias fed the judge easy traces. The uncalibrated judge passed essays. Drift made the pass rate climb. Side-effect replay created “successful” world states that were actually duplicates. The chart was consistent. It was consistently wrong.

This is the production costume of the pass-rate lie. The metric panel — revision rate, trajectory, coverage, cost per success — is owned elsewhere. Here I only need the eval-side version:

Green numberHow it was manufacturedWhat to show instead
Online pass rateEasy-stratum sample + lenient judgeHard-stratum pass + false-pass vs humans
“N traces scored”Volume, not coverageFilled cells in the mix × tool matrix
Judge latency downTruncated tracesEvidence-quote rate on full ledgers
Cost per eval downSame-family judge, no goldCost per labeled disagreement
Week-over-week liftRubric loosened or mix shiftedGold-slice delta with pins frozen

OpenAI’s best-practices line still holds: combine metrics with human judgment so you are answering the right questions. A single online percentage answers a question nobody asked: “did the judge like this week’s easy completions?”

I now refuse to put online pass rate on a ship/no-ship toggle unless the gold-slice false-pass bound is next to it. An advisory score drawn as a gate is how teams get surprised.

Offline vs online: the split I skipped

I tried to make one loop do both jobs. CI wanted freeze. Production wanted mix. The compromise was a dataset that thawed every night and a “golden” that was just yesterday’s sample with the labels still missing.

LangChain’s improvement-loop writeup draws the split I needed: online evaluators watch live behavior and flag traces; offline evaluations are controlled experiments on curated datasets before a change ships. Online tells you what is going wrong. Offline tells you whether the fix actually addressed it.

Offline goldenOnline sample
SourceHarvested failures + signed fixturesStratified live traces
ToolsStubbedAlready happened; do not re-run writes
LabelsRequiredSparse; overlay judge + human queue
CadenceEvery behavior changeRolling window
DecisionBlock a mergeAlert, harvest, do not auto-widen writes
Failure if skippedYou ship regressions you already paid forYou miss mix the fixtures never covered

Procedure that finally held:

  1. Offline suite is append-only except for documented deletions. Pass-rate jumps after you deleted hard cases are rot.
  2. Online sample never mutates the golden set automatically. A human promotes a trace into a fixture.
  3. A PR can fail on offline gold. A PR cannot “pass” because online pass ticked up.
  4. Disagreements (judge-pass / human-fail) enter the annotation queue the same day, not the next quarter.
  • dataset_id for gold is different from sample_id for online
  • Promotion to gold requires a stubbed tool world and an expected terminal
  • Online alerts page a human; they do not flip a write allowlist
  • The last three harvested misses have case_ids in git or a signed waiver

Harvest — how a bad run becomes a fixture — is a different spoke. This page only insists the two loops do not share a write key.

How do I evaluate this in production without lying?

Target query, operational version: score a stratified sample of stored traces, against criteria you already froze, with a judge you have calibrated on gold, with a principal that cannot write.

StepOwnerDone when
1. Freeze criteria + rubric_idJob owner + engBuyer signed pass/fail lines
2. Mechanical graders on the ledgerEngSchema, allowlist, ids, budgets never hit the judge
3. Label a gold sliceJob ownerConfusion matrix exists; false-pass bound written
4. Sample policyEngStrata, window, reweight rule in config, not a comment
5. Score stored tracesEval serviceJudge pin + rubric on every span
6. Annotation queueJob ownerDisagreements SLA in hours, not “when we can”
7. HarvestEngFixture stubs, not live replay

Guardrails that belong on that loop:

  • Separate eval principal. Deny writes at the token.
  • Cap judge package size. A 50k-token chat dump grades prose. A one-page ledger with ids the judge must quote grades actions.
  • Independent scoring for revisions. Position bias is not a chatbot-only problem.
  • Evidence quotes tied to tool ids. “Looks careful” is not a criterion.
  • Pin worker, judge, tools, and sample policy. Log the identity line on every run.

If step 3 is missing, stop. You can still collect traces. You cannot trust scores.

Evidence package for an online row — minimum, not a chat dump:

FieldWhy
Job goal + acceptance linesThe question the run was supposed to answer
rubric_id + criterion listSame text the humans used
Final artifactsWhat would ship
Tool ledger digestNames, redacted args, result ids, timestamps
Harness terminal reasonpass / revise / escalate / budget
Sample metadatajob_type, locale, channel, window, stratum
Identity lineWorker pin, judge pin, tools version, sample policy

Leave out worker chain-of-thought, secrets, full RAG dumps, and prior judge rationales. Default no on giving the judge live tools. If a field is missing, the row is unscorable — that is a harness incident, not a 0.0 quality score.

What guardrails do I need on the eval itself?

The worker needs a sandbox. The eval needs one too. Why agent demos fail production owns schemas, auth bleed, and kill switches for the agent. Copy the furniture onto the grader.

GuardrailWorkerEval harness
Tool allowlistJob-scoped writesRead-only or none
Tenant bindCustomer tenantSnapshot / sandbox / stubs
Kill switchFreeze writes without a prompt deployDisable scoring and replay without a prompt deploy
Identity logmodel_id, tools_versionThose plus judge_pin, rubric_id, sample policy
BudgetTurn cap, dollar capJudge-token cap so you stop truncating the hard traces
AuthPer-tenant credentialsEval credentials that cannot mint worker tokens

Checklist before the first live sample:

  • Eval service account reviewed like a production actor
  • Network policy: no SMTP, no billing write, no CRM mutate
  • Replay path unit-tested with a fixture that would have written, and did not
  • Kill switch is a config flag, not a judge-prompt paragraph
  • PII redaction on packages that leave the tenant boundary (judge vendor, logs, gold store)

A prompt that says “you are only evaluating” is not a guardrail. Credentials are.

When is a workflow enough instead of an agent?

When you cannot isolate the eval, you cannot earn the loop.

If the job is a known path with known tools, a workflow already gives you deterministic tests: assert the payload, assert the idempotency key, assert the destination. You do not need a live judge to tell you a mapping job worked. You need a fixture.

You still want an agentYou want a workflow
The next action depends on messy evidenceThe path is known before the run starts
You can name pass/fail criteria and label goldYou cannot get two humans to agree after seeing the ledger
Eval can score stored traces with stubsThe only “eval” you can imagine is re-running writes
Blast radius is gated (draft-only, approval on sends)Irreversible tools with no replica and no stubs
Mix is wide enough that fixtures will miss itMix is a handful of shapes you can enumerate

Anthropic’s Building effective agents is blunt: the evaluator-refiner pattern only fits when you have clear evaluation criteria and iterative refinement actually improves the artifact. No criteria, no loop. No isolated eval, same conclusion.

After 20,000+ hours on agentic systems and 500+ automations, the honest fork is: if production eval requires a second copy of the blast radius, you are not ready for an agent. Keep the workflow. Collect traces in draft mode. Build gold. Then argue about autonomy.

A demo that cannot be evaluated without sending a real email was never a demo of an agent. It was a demo of a send button.

One-week honesty test I run with the buyer:

  1. Can we stub the irreversible tools today? If no, the agent stays draft-only.
  2. Can two humans label twenty hard traces without splitting on the rubric? If no, workshop criteria; do not hire a judge.
  3. Can the eval principal be denied writes in IAM, not in a prompt? If no, you do not have a harness.
  4. Is the job a lookup-plus-map that a workflow already does? If yes, stop the agent ticket.
  • Draft-only worker is acceptable for a week of harvest
  • “We’ll be careful in production” is not a stub
  • A workflow spike is a valid outcome of an agentic pilot, not a failed sale

If those four answers are no, the week should produce a workflow sketch and a criteria workshop, not an online leaderboard.

What I freeze before the next live sample

This is the packing list I use on a Spurlock Studios agentic pilot. Five days will not finish an academic inter-annotator study. Five days will tell you whether the measurement is usable or dangerous.

DayFreezeDo not do
1Criteria split: mechanical vs judge. Eval principal with writes denied.Give the judge the worker’s tools “just for outcome checks”
2Sample policy: strata, 7-day window, mix histogram. Pull 30–50 traces.Score last night’s completions and call it a baseline
3Human labels on the hard slice. First confusion matrix.Ship the online judge because the prompt “looks strict”
4Stubbed replay on three known-bad traces. Prove no live writes.Re-run the worker against production to “verify the note landed”
5Gold-slice ids frozen. Judge span on run id. False-pass bound written next to blast radius.Put online pass rate on a ship toggle

What “done” means on day five:

  • Mechanical checks no longer go through the judge
  • Eval credentials cannot write
  • Gold-slice ids exist and are not used for prompt fiddling
  • Sample mix is queryable next to the score
  • A named owner for the next re-calibration trigger
  • Three harvested misses have a path into fixtures, not into folklore

That is the same discipline as the rest of the control plane. The scorer is a component with a version. The sample is a component with a policy. The eval principal is a component with a deny list. None of those are a paragraph in a system prompt.

Anti-patterns that recreate the original incident:

Anti-patternWhat it costs
Online pass rate as the ship toggleYou widen writes on a lenient week
Gold slice = last week’s happy pathsEasy-case accuracy in a new costume
Judge tools “for evidence”Side effects; you are running a second agent
Sample = completed chats after 10pmMix lie
Rubric edits with no version bumpSilent drift
Same token for worker and evalShared blast radius
Deleting hard fixtures to “clean the suite”Pass-rate rot

If an anti-pattern is already in the dashboard, turn writes down before you argue about models.

How this spoke sits next to the cousins

Do not collapse four posts into one dashboard.

QuestionThis spokeSomewhere else
What broke when live eval hit production?Labels, sampling, drift, eval side effects—
What do I build first?Assume you already believe thisEvaluators before agents
Is the judge calibrated?I needed it; I did not have itLLM-as-judge reliability
Why did the demo die?Not this pageWhy agent demos fail production
What else has to exist?Eval as a second sandboxed systemOperating manual

If you skip the first column and only add an online judge, you have a second agent with a clipboard. The clipboard will write if you let it.

The only question here is: when production eval says pass, did the measurement survive contact with live tools — or did you just run the agent twice?

FAQ

What broke when I tried to evaluate an AI agent in production?

The measurement: no golden labels, a biased sample of easy completions, a drifting judge, and an eval harness that called write tools. The judge graded into a void while humans still rewrote. Scoring mutated the world it was supposed to observe. Fix those four before you trust a live score.

How do I measure whether broke when I tried to evaluate an AI agent in production is working?

Look at gold-slice false-pass versus humans, mix histograms next to the online number, and whether replay can mutate production. If judge-pass / human-fail is unqueryable, the eval is not working. A rising online pass rate with a frozen gold slice and a stable mix is the only lift I believe.

What usually fails first when teams try this?

Golden labels fail first: teams ship a judge prompt in an afternoon and skip 30–50 stratified human labels. Sampling bias is second — overnight completions, no escalations. Eval side effects show up the first time someone “replays” a write, and drift shows up later, which is why it gets blamed on the model.

How long does this take to show results?

A five-day pilot can tell you whether the harness is usable or dangerous: criteria split, denied writes, a labeled slice, stubbed replay, frozen gold ids. Trustworthy online numbers take longer — you need a week of stratified mix and an annotation SLA, not a weekend of green charts. I will not invent a universal week-three accuracy number.

What should I skip if I only have a week?

Skip widening writes, a same-family judge as the only gate, and re-running the worker to “check outcomes.” Do not skip labels, sample strata, or eval-credential deny lists. A thin week that freezes those three beats a theatrical week that scores five thousand unlabeled chats.

When is this not worth doing yet?

When two trained humans cannot agree on the ledger, when you have no replica or stubs for irreversible tools, or when the job is a known path a workflow already covers. Online eval of an agent you cannot isolate is how you buy a second copy of the blast radius. Keep writes off, or stay on a workflow, until gold and isolation exist.

CTA

Sandbox the measurement before you trust the live score.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

What broke when I tried to evaluate an AI agent in production?
The measurement: no golden labels, a biased sample of easy completions, a drifting judge, and an eval harness that called write tools. The judge graded into a void while humans still rewrote. Scoring mutated the world it was supposed to observe. Fix those four before you trust a live score.
How do I measure whether broke when I tried to evaluate an AI agent in production is working?
Look at gold-slice false-pass versus humans, mix histograms next to the online number, and whether replay can mutate production. If judge-pass / human-fail is unqueryable, the eval is not working. A rising online pass rate with a frozen gold slice and a stable mix is the only lift I believe.
What usually fails first when teams try this?
Golden labels fail first: teams ship a judge prompt in an afternoon and skip 30–50 stratified human labels. Sampling bias is second — overnight completions, no escalations. Eval side effects show up the first time someone “replays” a write, and drift shows up later, which is why it gets blamed on the model.
How long does this take to show results?
A five-day pilot can tell you whether the harness is usable or dangerous: criteria split, denied writes, a labeled slice, stubbed replay, frozen gold ids. Trustworthy online numbers take longer — you need a week of stratified mix and an annotation SLA, not a weekend of green charts. I will not invent a universal week-three accuracy number.
What should I skip if I only have a week?
Skip widening writes, a same-family judge as the only gate, and re-running the worker to “check outcomes.” Do not skip labels, sample strata, or eval-credential deny lists. A thin week that freezes those three beats a theatrical week that scores five thousand unlabeled chats.
When is this not worth doing yet?
When two trained humans cannot agree on the ledger, when you have no replica or stubs for irreversible tools, or when the job is a known path a workflow already covers. Online eval of an agent you cannot isolate is how you buy a second copy of the blast radius. Keep writes off, or stay on a workflow, until gold and isolation exist.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot