Spurlock Studios
Contact
Share LinkedIn X
A violet scale. Thesis: LLM AS JUDGE RELIABILITY CALIBRATE.

You cannot trust an LLM-as-judge out of the box — it will flatter bad agent runs if you never calibrate it. Treat the judge as a noisy instrument: measure agreement with humans, kill known biases, and re-validate after every model or rubric change. A high judge score with unmeasured calibration is a vanity metric with better fonts.

This is meta-eval. How to build the evaluator stack lives in evaluators before agents. This spoke asks whether the scorer itself is lying. Parent context: Agentic Systems Operating Manual. The cousin that owns the metric panel — pass rate, revision rate, coverage — is why pass rate lies.

The short answer

  • Mechanical checks first; judges only where humans would disagree on soft criteria.
  • Calibrate against a labeled set before the judge gates CI or production.
  • Agent trajectories break naive “grade the final answer” judges — score tools and intermediate claims too.
  • Watch position bias, verbosity bias, and same-model circular scoring.
  • Re-run calibration when the worker model, judge model, or rubric changes — judge drift is real.

What LLM-as-judge means in an agent harness

In an agent eval harness, the judge is a second model call (or panel) that scores a run against a rubric and returns structured verdicts: pass/fail, criterion codes, short evidence quotes.

It is not:

  • A replacement for schema validation
  • Ground truth by virtue of being “smarter”
  • Safe to share weights casually with the worker without measuring circularity

Typical placement:

  1. Worker run completes (or hits a checkpoint)
  2. Mechanical checks run (schema, allowlist, required artifacts)
  3. Judge scores remaining soft criteria
  4. Harness maps scores → pass / revise / escalate / fail

OpenAI’s own agent-eval docs treat this split as the default: deterministic graders for structure, model graders for open-ended quality (trace grading, agent evals). If step 2 is empty, you are paying a judge to do string checks. Stop.

RoleOwnsMust not own
Mechanical graderSchema, allowlists, budgets, idsTone, completeness vs a messy brief
LLM judgeSoft criteria with named evidenceBinary facts already on the ledger
HumanIrreversible tools until calibration holdsSpot-checking only the easy passes
WorkerThe artifact and the tool callsThe pass/fail that ships the run

Capability language beats fashion: pick a judge that follows rubrics and returns structured JSON reliably. Pin the id. Re-calibrate on change.

Why agent trajectories break naive judges

Chatbot judges often see: prompt, final answer, rubric. Agent runs add tools, multi-turn state, and side effects. Failure modes unique to agents:

FailureWhat a naive judge misses
Wrong tool, right-looking final textGrades the essay, ignores the CRM write
Hallucinated tool successBelieves the worker’s narration over the tool ledger
Criterion satisfied mid-trace then undoneScores the last message only
Policy near-miss“Sounds careful” while an irreversible tool nearly fired

RuVerBench (Peng et al., 2026) is the first public meta-eval aimed at this gap: 2,458 rubric-verification instances across deep research and agentic coding. Deep-research reports averaged 7.1K tokens; coding trajectories averaged 49.4K. Even frontier judges stayed strong-but-noisy. Weaker judges swung harder when the prompt changed. Batched verification traded accuracy for speed. Majority voting helped, then flattened.

JudgeBench makes the same point from the other side: many judges that look aligned on preference chats collapse toward chance on hard knowledge, math, and coding pairs. Human-preference agreement is not factual reliability.

Practical takeaway: feed the judge a structured trace digest, not a chat dump — tool names, redacted args/results, state transitions, and the final artifacts.

Do not invent a single published “accuracy %” for all judges. Your calibration numbers are the only ones that count for your rubric.

How to calibrate against human labels

Procedure that actually moves reliability:

  1. Sample 50–100 runs covering pass, fail, and escalate (stratify by job_type and difficulty). LangChain’s agent-eval note starts teams at 20+ labeled examples and grows from there; humans review on the order of 50–100 traces per hour, so the bottleneck is attention, not tooling (agent evals).
  2. Blind-label with humans using the same rubric the judge will see. Two raters when stakes are high; resolve disagreements explicitly and keep the disagreement codes.
  3. Run the judge on the same packages; store scores, evidence quotes, and the judge model pin.
  4. Compute agreement: per-criterion accuracy / F1, overall pass-fail agreement, and confusion pairs (judge pass / human fail is the dangerous cell).
  5. Tune rubric wording, evidence requirements, and which criteria stay mechanical.
  6. Freeze a calibration report with judge model id, rubric version, and date.
MetricWhy it matters
Human–judge pass/fail agreementCI gate sanity
False pass rate (judge pass, human fail)Customer risk
False fail rateCost / latency from over-refusal
Per-criterion agreementFinds broken rubric lines
Chance-corrected agreement (κ)Stops you celebrating easy-set accuracy

Raw agreement flatters. A 2026 large-scale audit of 21 judges across MT-Bench, JudgeBench, and RewardBench found exact-match agreement overstated chance-corrected discrimination by 33–41 percentage points on MT-Bench (Reliability without Validity). You do not need their exact delta. You need a metric that does not treat “both said pass on an obvious pass” as proof the judge can see a hard fail.

Hedge, not folklore: “good enough” for a soft-launch gate is often in the ballpark of strong majority agreement on pass/fail for your risk class — but you set the threshold from blast radius, not from a blog’s lucky number. The 2023 MT-Bench paper reported a strong judge matching human prefs in the same band as human–human agreement on their chat tasks, after bias mitigations (Zheng et al.). That is a historical receipt, not a CI threshold. Irreversible tools demand tighter false-pass bounds than draft-only jobs.

Calibration artifactStore it
rubric_id + semverOn every judge span
Judge model pin / provider snapshotNext to the score
Gold-slice idsNever used for prompt tuning
Confusion matrix by criterionWeekly drift check
Human override codesNext calibration set

If you cannot point at that report, you do not have a calibrated judge. You have a prompt.

What if the humans disagree with each other?

A judge cannot outrun its labelers. If two humans split on the same package, “judge–human agreement” is a coin flip dressed as a metric.

Do this before you blame the model:

  1. Sit both raters on the same five disagreements with the rubric on the table.
  2. Split each fight into rubric hole (the line is vague) vs attention miss (the evidence was there).
  3. Rewrite the hole as a mechanical check or add a positive and negative exemplar.
  4. Keep a third label, human_split, on packages you will not force into pass/fail.
  5. Report judge agreement against the adjudicated label, and separately against each rater.
Human patternWhat it means for the judge
High rater agreement, judge missesJudge or package problem — fix those
Low rater agreement, judge matches oneRubric problem — do not tune the judge to a coin flip
Both raters pass, later ops rejectOnline criterion missing from the rubric
Both raters fail, judge passesThe dangerous cell — raise this before any CI gate

Zheng et al. treated human–human agreement as the ceiling on their chat tasks (MT-Bench). Copy the method, not their percentage. Your ceiling is the adjudicated agreement on your hard stratum. If that ceiling is 70% because the rubric is mush, a 90% judge score means the judge learned one rater’s taste.

Irreversible tools do not get a mushy rubric. If two trained humans cannot agree after seeing the ledger, the job is not ready for a judge gate. Workshop the criterion. Then calibrate.

How do position and verbosity bias show up in agent traces?

These show up differently than in pairwise chatbot evals.

Position bias: When the judge sees multiple candidate revisions or tool results in a list, earlier or later items can win unfairly. Shi et al. measured this across pairwise and list-wise settings (15 judges, 22 tasks, 150,000+ instances): position bias is not random, it varies by judge and task, and it gets worse when the quality gap between candidates is small. Shuffle or score candidates independently when you compare revisions.

Verbosity bias: Long, confident worker narrations score higher than terse correct tool use. Zheng et al. documented the “repetitive list” attack on chat judges — padded answers winning over shorter correct ones (MT-Bench). On agent traces the costume is a three-paragraph “I verified in the CRM” speech with no matching ledger span.

Countermeasures:

  • Require evidence quotes tied to tool ledger ids, not vibes
  • Cap narrative length in the judge package
  • Score “correctness of actions” separately from “quality of prose”
  • Penalize unsupported claims explicitly in the rubric
BiasSymptom in tracesMitigation
PositionRevision A always wins when listed firstIndependent scoring / shuffle
VerbosityWordy fails beat short passesEvidence-first rubric
AuthorityJudge trusts “I verified via CRM” without tool spanLedger required
LeniencySoft criteria always “mostly met”Binary criteria + examples
Self-preferenceSame-family worker prose scores highDifferent-family judge, or measure it

If your judge prefers essays, your agent will learn to write essays instead of calling tools correctly.

The 2026 Reliability-without-Validity audit reported much smaller verbosity effects under a single pairwise rubric than the 2023 literature. Do not take that as a hall pass. Your agent traces are not MT-Bench pairs. Measure verbosity on your packages: correlate judge score with token count on the hard stratum. If the correlation is the story, the rubric is still grading prose.

Same model as worker — ever OK?

Sometimes, for low-stakes draft scoring in staging. Rarely for production gates on irreversible work.

Risks:

  • Shared blind spots (both miss the same policy hole)
  • Style favoritism (worker prose matches judge priors)
  • Correlated drift when the provider updates the family

Self-preference is measured, not vibes. Wataoka, Takahashi, and Ri found judges scoring lower-perplexity text higher than humans do — including text they were not told was their own. Play Favorites then isolated self-bias and family-bias: some judges systematically score their own completions, and other models in the same family, above an independent reference. G-Eval’s authors flagged the same tilt toward model-generated text when they introduced CoT + form-filling judges (Liu et al.).

SetupUse when
Same model family, same pinCheap staging smoke only
Same family, different pin / sizeAcceptable if calibrated; still watch circularity
Different vendor for judgePrefer for production gates when cost allows
Panel (2 judges + tie-break rule)High blast radius criteria, after single-judge work fails

A panel is not a personality. Independent scores, a predefined tie-break, and a recorded disagreement rate. Skip multi-agent debate theater as the SMB default — RuVerBench already saw majority voting flatten. If one calibrated judge plus mechanical checks is inside your false-pass band, stop adding voters.

How do you detect judge drift?

Judge drift is a silent production bug: worker prompts unchanged, online “pass rate” climbs or collapses, humans still rewrite.

Triggers that force a re-calibration run:

  • Judge model pin or provider snapshot changed
  • Rubric version bumped (even “clarifications”)
  • Worker model upgraded (distribution of traces changes)
  • New tool or side-effect class added
  • Human override rate diverges from judge pass rate for two weeks
  • Judge latency, JSON-schema fail rate, or empty-evidence rate jumps

Drift checks:

  1. Hold out a frozen gold slice (never used for prompt tuning).
  2. Weekly or on deploy: score the slice; alert if agreement or false-pass rate moves past your band.
  3. Sample online disagreements (human reject after judge pass) into the next calibration set.
  4. Keep the judge span on the same run id as the tool calls — same weekly ritual as the rest of the control plane.
SignalRead as
Gold-slice κ holds, online rewrite rate upWorker or input mix moved; judge may still be fine
Gold-slice false-pass up, rewrite rate upJudge got lenient, or rubric got vague
Gold-slice false-fail up, escalate rate upJudge got harsh; check pin and prompt
Judge JSON invalid / missing evidenceHarness bug, not a quality win

LangChain’s note is blunt: recalibrate regularly, because judges drift just like the agents they evaluate (agent evals). Anthropic’s multi-agent research writeup keeps humans in the loop for the same reason — people still catch hallucinations, system failures, and source misses that automation grades past (engineering note).

Judge spans belong on the trace beside tool calls. If you cannot query “show me last week’s judge-pass / human-fail,” you will learn about drift from a customer.

Which checks should never be a judge call?

Move these out of the LLM judge entirely:

CheckWhy mechanical
JSON / schema validityBinary, cheap
Required fields presentBinary
Tool allowlist / deny listPolicy, not taste
Max turns / budget exceededHarness facts
Forbidden strings / PII patternsRegex or classifiers
Idempotency key present on writesLedger fact
Tool result exists for every cited claim idLedger join
Terminal reason is a known enumHarness fact

Judges earn their tokens on: tone, completeness vs a messy brief, “did the research address the question,” soft brand constraints. If a criterion can be a unit test, make it a unit test.

A useful split when you are arguing with yourself:

  1. Write the criterion as a sentence.
  2. Name the observable (field, tool span, quote).
  3. If a junior engineer could assert it in ten lines of code, it is not a judge criterion.
  4. If two humans would still disagree after seeing the same evidence, it may be a judge criterion — and it needs exemplars.

OpenAI’s grader docs make the same cut: string and code graders for deterministic properties, model graders for the rest, and an explicit warning about grader hacking — the policy scores high on the grader and low with experts (graders). That is judge unreliability with an optimization loop attached. If the worker can see the rubric and is trained or prompted against it, your next calibration set must include the new failure costume, not last month’s.

Failure mode: correlated easy-case accuracy

What breaks: your calibration set is 80% obvious passes. Judge–human agreement looks excellent. Production is the hard 20%. The judge rubber-stamps fluent wrongness.

What it costs: CI stays green while revision rate stays ugly — the pass-rate lie with a judge costume. Why pass rate lies owns that panel. This post owns the scorer that inflated the first number.

What you do instead:

  1. Stratify the labeled set by difficulty and failure code.
  2. Track agreement on the hard stratum separately.
  3. Keep a rising share of production disagreements in the set.
  4. Never celebrate aggregate agreement alone.
  5. Include near-miss policy cases, hallucinated tool success, and undone mid-trace criteria — the rows naive judges miss.
StratumIf you skip itWhat the dashboard hides
Obvious passNothing usefulInflated κ
Obvious failJudge looks strictMisses leniency on the real risk
Hard / ambiguousThe job you hired the judge forFluent wrongness ships
New tool classYesterday’s calibrationSilent false passes

Easy cases are where judges look smart. Hard cases are why you hired them.

JudgeBench exists because preference-aligned judges can still lose on objective pairs. RuVerBench exists because long agentic traces make that worse. Your gold slice should look more like those than like a demo reel.

What rubric design survives agent traces?

Rules of thumb for judge-ready rubrics:

  1. One criterion, one failure code.
  2. Each criterion names observable evidence (artifact field, tool result, quote).
  3. Include 2–3 positive and negative exemplars per soft criterion.
  4. Separate “process” criteria (allowed tools, no speculative writes) from “outcome” criteria (user gets value).
  5. Version the rubric (rubric_id, semver). Store it on every judge span.
  6. Ban “mostly” and 1–5 vibes. Binary or a three-way pass / revise / fail with a named reason.

Bad criterion: “Be helpful and accurate.” Better: “Every numeric claim in the customer email appears in tool:billing.get result or is marked uncertain.”

Rubric smellWhat the judge will doFix
One “quality” scoreUn-actionable leniencySplit into codes ops can fix
No evidence slotGrades vibesRequire ledger-linked quotes
Process mixed with outcomePasses a polite wrong writeTwo criteria, two codes
Unversioned prose in a promptSilent driftrubric_id on the span
Exemplars all long essaysVerbosity biasInclude short correct traces

G-Eval’s useful inheritance is the form, not the brand: decompose the criterion, force a structured fill, then score (Liu et al.). The failure mode they already flagged — bias toward model-shaped text — is why the form still needs human labels on your traces.

What belongs in the judge evidence package?

If you dump the raw chat, you are grading creative writing. If you dump the whole RAG corpus, you are grading the judge’s patience.

Minimum package:

FieldWhy
Job goal + acceptance linesThe question the run was supposed to answer
rubric_id + criterion listSame text the humans used
Final artifactsWhat would ship
Tool ledger digestNames, redacted args, result ids, timestamps
Harness terminal reasonpass / revise / escalate / budget
Candidate list (if any)Independently scored, position shuffled

Leave out:

  • Worker chain-of-thought, unless you A/B’d it on the labeled set and false-pass improved
  • Secrets, raw PII, full document dumps
  • Prior judge rationales (the next judge will agree with the last one)
  • Unrelated memory / RAG hits the worker never used

Default no on CoT-in-the-package. It invites style grading and leaked rationalizations. Prefer actions + artifacts. If you experiment, keep it only when the hard-stratum false-pass rate moves the right way.

Cap the digest. RuVerBench’s coding traces averaged tens of thousands of tokens; that is the regime where even frontier judges get noisy. A one-page ledger with ids the judge must quote will beat a 50K-token dump you never measured.

How do you gate CI without false comfort?

Suggested promotion ladder:

GateJudge role
PR / prompt changeMechanical + judge on golden set; block on false-pass regressions vs baseline
Staging soakOnline sample; compare human spot-checks
Soft-launchJudge advisory or dual-run; humans still own irreversible tools
Autonomy expandJudge gate only after calibration report signed off

Agreement thresholds are a product decision. Document them next to blast radius. Do not copy a research paper’s headline number into your runbook without re-measuring on your traces.

You may block a PR when…You may not
Gold-slice false-pass exceeds bandAggregate pass rate ticked up
A new criterion has no human labelsThe flagship model is “the judge now”
Judge JSON / evidence fail rate spikesOne anecdotal “it looked better”
Hard-stratum κ dropsEasy-set agreement holds

OpenAI’s grader-hacking warning is the CI version of this: a worker that learns the scorer will green the eval and still fail experts (graders). Pair the judge gate with why pass rate lies — revision rate, trajectory, coverage, cost per success — so a lenient scorer cannot be the only veto.

If the judge is advisory, say so in the dashboard. An advisory score drawn as a ship/no-ship toggle is how teams get surprised.

What this post does not replace

SpokeOwns
Evaluators before agentsBuild order: criteria → mechanical → judge → online
Why pass rate liesThe metric panel a lying scorer can inflate
Agentic Systems Operating ManualThe rest of the stack around the scorer
This postMeta-eval: is the judge calibrated and stable?

If you skip the first two and only add a judge prompt, you have cosplay.

This spoke stays on the instrument. It does not tell you which jobs deserve an agent, how to sandbox tools, or how to price a fleet. Those are other pages. The only question here is: when the scorer says pass, should you believe it?

Pilot slice: calibration in five days

A Spurlock Studios agentic pilot can include a thin meta-eval pass when the job already has soft criteria. You will not finish academic-grade inter-annotator studies in five days. You will know whether the judge is roughly usable or actively dangerous.

DayJudge work
1Split mechanical vs judge criteria; write the evidence-package schema
2Label 30–50 runs (or dense fixtures), stratified, two raters on the hard slice
3First judge pass + confusion matrix; compute false-pass on the hard stratum
4Rubric surgery; kill verbosity and position loopholes; drop criteria that should be code
5Freeze rubric_id + pin; wire judge span to traces; hold out the gold slice

What “done” means on day five:

  • Mechanical checks no longer go through the judge
  • A written false-pass bound sits next to blast radius
  • Gold-slice ids are frozen and not used for prompt fiddling
  • Judge spans queryable by run id
  • A named owner for the next re-calibration trigger

That is the same discipline I use across 500+ automations and 20,000+ hours on agentic systems: the scorer is a component with a version, not a vibe in a system prompt. Book via /agentic.

Worked example: support draft agent

CriterionJudge or mechanical?
Contains order id from ticketMechanical
No refund promise unless tool says eligibleMechanical on tool + regex
Tone matches brand examplesJudge
Answers all explicit customer questionsJudge with checklist from ticket

Illustrative pattern (not a universal stat): judge passes “tone” on long drafts; humans fail short correct ones. Fix: verbosity penalty + exemplar shorts; require factual claims to cite tool:orders.get. That bias fix often beats swapping judge vendors.

Evidence package minimum: job goal, rubric version, final artifacts, tool ledger digest, harness terminal reason. No secrets, no giant RAG dumps. No ledger → you are grading creative writing.

Same-family judge on this job: fine for staging drafts after the day-3 matrix. Not fine as the only gate once the draft can trigger a refund tool.

Panel: only when blast radius is high and single-judge false-pass stays above band after rubric work. Independent scores, predefined tie-break.

Anti-patterns that keep this example lying:

“The flagship model is the judge, so we’re fine.” Capability helps; calibration decides. JudgeBench already showed strong judges losing on hard pairs.

Judge sees full chain-of-thought and grades style. Prefer actions + artifacts.

One giant “quality 1–5” score. Un-actionable. Prefer criterion codes ops can fix.

Recalibrating never. Then your dashboard is a fiction that ages.

Gold slice is last week’s happy paths. Then you re-discovered easy-case accuracy.

FAQ

When should mechanical checks replace a judge entirely?

Whenever the criterion is binary and observable without taste: schemas, allowlists, budgets, required ids, forbidden actions. Judges are for soft criteria. If your entire rubric is mechanical, delete the judge call and celebrate the latency win.

Should the judge see the worker’s chain of thought?

Default no for scoring. CoT invites style grading and leaked rationalizations. Prefer tool ledgers and artifacts. If you experiment with CoT-in-the-judge-package, A/B it on your labeled set — keep it only if false-pass rate improves on the hard stratum.

Same model as worker — ever OK?

For low-stakes staging or draft-only jobs after calibration, sometimes. For production gates on irreversible tools, prefer a different model family or a panel, and always measure circular agreement on hard cases. Self-preference and family-bias are documented; do not assume your pin is the exception.

What agreement rate with humans is “good enough” to gate CI?

Whatever bound matches your blast radius — documented, measured on a stratified labeled set, with special attention to false passes. There is no universal published percentage that absolves you from measuring on your rubric and traces.

How do position and verbosity bias show up in agent traces?

Position bias skews revision tournaments; verbosity bias rewards long narrations over correct short tool use. Mitigate with independent scoring, evidence-first rubrics, and ledger-linked claims. Measure both on your packages, not from a chat-benchmark abstract.

How does this relate to evaluators-before-agents without replacing it?

Evaluators before agents tells you to build the eval stack before autonomy. This post assumes that stack exists and asks whether the LLM judge component is calibrated. You need both: a real evaluator, and a scorer you have meta-evaluated.

CTA

Calibrate the scorer before you trust the score.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

When should mechanical checks replace a judge entirely?
Whenever the criterion is binary and observable without taste: schemas, allowlists, budgets, required ids, forbidden actions. Judges are for soft criteria. If your entire rubric is mechanical, delete the judge call and celebrate the latency win.
Should the judge see the worker’s chain of thought?
Default no for scoring. CoT invites style grading and leaked rationalizations. Prefer tool ledgers and artifacts. If you experiment with CoT-in-the-judge-package, A/B it on your labeled set — keep it only if false-pass rate improves on the hard stratum.
Same model as worker — ever OK?
For low-stakes staging or draft-only jobs after calibration, sometimes. For production gates on irreversible tools, prefer a different model family or a panel, and always measure circular agreement on hard cases. Self-preference and family-bias are documented; do not assume your pin is the exception.
What agreement rate with humans is “good enough” to gate CI?
Whatever bound matches your blast radius — documented, measured on a stratified labeled set, with special attention to false passes. There is no universal published percentage that absolves you from measuring on your rubric and traces.
How do position and verbosity bias show up in agent traces?
Position bias skews revision tournaments; verbosity bias rewards long narrations over correct short tool use. Mitigate with independent scoring, evidence-first rubrics, and ledger-linked claims. Measure both on your packages, not from a chat-benchmark abstract.
How does this relate to evaluators-before-agents without replacing it?
[Evaluators before agents](/blog/evaluators-before-agents) tells you to build the eval stack before autonomy. This post assumes that stack exists and asks whether the LLM judge component is calibrated. You need both: a real evaluator, and a scorer you have meta-evaluated.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot