Build the Evaluator Before the Agent
If the judge shares the worker's context, you are grading your own homework. Build criteria, evidence, and ceilings first — then grant real tool autonomy.
William Spurlock Founder — Spurlock Studios Updated 22 MIN
If you ship the agent before the evaluator, you are guessing. Accuracy jumps I trust never came from a warmer system prompt. They came from a second component whose only job is to disagree, using structured evidence the worker did not write.
This spoke sits under the Agentic Systems Operating Manual. The parent maps the stack. Here I own build order: independent criteria, evidence, and revision ceilings — then autonomy. How you harvest a bad run into a fixture lives in golden sets from agent failures. How Spurlock Studios installs the first harness is on /agentic.
The short answer
- Write pass/fail criteria with the buyer before anyone opens a tool schema.
- The evaluator sees criteria + artifact + retrieval evidence. Not the worker’s private reasoning.
- Code owns schema, enums, totals, allowlists, and citation presence. A model judge owns the residue.
- Failures name a criterion and a quote. Three revisions is the pilot default; the harness enforces the ceiling.
- Widen write tools only when golden-set pass rate, cost per pass, and escalate rate hold on sampled production.
Why does self-grading fail?
A model asked to check its own output is continuing a story in which it already “finished.” The verdict is correlated with the work. In production that looks like a green run where the tool errored, the summary soft-pedaled the error, and nobody opened the trace.
Self-grading is a hint inside a worker. It is not your metric. Zheng et al. documented the same family of failure in Judging LLM-as-a-Judge with MT-Bench: position bias, verbosity bias, and self-enhancement — models prefer their own prose. That paper is about chat assistants. Agents make it worse, because the worker also chose the tools and wrote the story of why those tools were fine.
| Setup | What the judge sees | What you actually measured |
|---|---|---|
| Worker says “looks good” | Its own draft + its own rationale | Confidence, not correctness |
| Same model, same thread | The conversation that produced the artifact | Continuity of a story |
| Independent evaluator | Criteria + artifact + evidence package | The contract you wrote |
| Human override | Same package plus the trace | The last veto, not the first grader |
- The worker cannot set
status: done - The judge does not receive chain-of-thought from the worker
- Every blocking fail names a criterion id
- Known-bad fixtures still fail after a prompt tweak
- A human can replay the evidence package without the chat
If any box is unchecked, you are still grading homework.
What is an independent evaluator?
An evaluator is a component — code, model, or both — that returns a structured verdict. It is not a vibes score. It is not a thumbs-up in Slack. Anthropic’s Demystifying evals for AI agents (Jan 2026) splits the same object into a task, a grader, a transcript, and an outcome. The outcome is the state of the world, not the last sentence the model uttered.
{
"verdict": "fail",
"evaluator_version": "refund-policy-v3",
"failures": [
{
"criterion": "severity_present_or_explicit_no_match",
"severity": "blocking",
"evidence": "summary claims policy X with zero citation URLs"
}
],
"next": "re-retrieve with query focused on refund policy; rewrite summary"
}
Rules that make it real:
- Independent context. Criteria, artifact, and evidence. No worker monologue.
- Mechanical first. Schema, enums, regex allowlists, arithmetic, required ids — assert in code.
- Model only for judgment. Tone, omission of material risk, “did this answer the question asked.”
- Evidence, not vibes. “Not good enough” is useless. “Criterion 3 failed because…” is a next action.
- Revision ceiling. Usually three. Then escalate with the full package.
- Read-only judge. The evaluator does not call write tools. Ever.
OpenAI’s agent evals and trace grading treat the trace as the unit you grade — model calls, tool calls, guardrails, handoffs — then tell you to move graded traces into a dataset when you need repeatability. That is the same split: one run is a clue; a fixture is a gate.
Microsoft Foundry’s agent evaluators make the same cut in vendor language: system evaluation (did the job finish with a usable deliverable) versus process evaluation (did the steps stay inside the contract). Binary pass/fail beats a 0–100 “quality” slider that hides which contract broke.
What criteria must exist before you write tools?
Sit with the buyer. Translate “good” into pass/fail lines. If you cannot, you are not ready to build. Ambiguous taste is a product workshop, not an agent ticket. Anthropic’s Building effective agents is blunt: the evaluator-refiner pattern only fits when you have clear evaluation criteria and iterative refinement actually improves the artifact. No criteria, no loop.
| Criterion shape | Example | Owner |
|---|---|---|
| Schema | JSON matches the contract; required keys present | Code |
| Arithmetic | Line totals equal header total within one cent | Code |
| Policy | Refund promise requires a citation URL or no_match | Code + light model |
| Completeness | Every asked field is answered or explicitly declined | Model judge |
| Safety | No PII in outbound email; no “we will sue them” | Code allowlists |
| Process | Forbidden tool never appears in the trace | Code over the trace |
Write criteria in this order:
- Name the job in one sentence the buyer will repeat.
- List irreversible side effects (money, legal, public send, delete).
- For each side effect, write the blocking fail that must stop the write.
- Add soft fails that may pass the run but open a ticket.
- Freeze the list as
evaluator_version. Tools come after.
- Every blocking criterion is testable on a fixture without a human in the room
- Soft vs blocking is labeled, not implied
- “Be helpful” and “be professional” have been deleted or split into observables
- The buyer signed the list, not just the demo script
A criterion that needs a page of interpretation is two criteria, or it is not a criterion.
Workshop script I run on day one of a pilot. Forty-five minutes. Whiteboard only.
- “What write, if wrong, do you have to unwind by hand?” Circle those. They are blocking.
- “Show me three tickets you already lost sleep over.” Those become never-auto-pass fixtures.
- “What would you accept as proof the agent was right?” If the answer is a feeling, keep workshopping.
- Read the list back. If two people disagree, the criterion is still mush. Split or delete.
- Freeze v0. Tools stay closed until v0 has an owner and a version string.
If step 3 dies, you do not have an agent job yet. You have a process argument wearing a model.
What counts as evidence, not vibes?
A verdict without evidence trains the worker to argue. A verdict with a quote trains the worker to fix a named hole. LangSmith’s evaluation concepts split graders into code, LLM-as-judge, human, and pairwise — and they still expect you to review scores and tune the judge prompt. Few-shot examples belong on the evaluator, never as a paste of golden answers into the worker.
| Evidence type | Good | Useless |
|---|---|---|
| Quote | Exact span from the artifact or tool payload | “The tone felt off” |
| Id | invoice_id, policy_url, tool_call_id | “It used the right tool, I think” |
| Counterfactual | “Totals differ by $4.12” | “Numbers seem high” |
| Absence | “Zero citation URLs; criterion requires one or no_match” | “Sources were weak” |
| Trace fact | email.send appeared after a fail | “It probably sent” |
Minimum evidence package the judge returns:
criterionid from the versioned rubric.severity:blockingorsoft.evidence: a quote or a computed delta, ≤200 characters.next: one action the worker may take, orescalate.
If the judge cannot fill those four fields, the criterion is not ready. Tighten it until a human could fill the same form from the artifact alone.
NIST’s AI 600-1 Generative AI Profile puts this in governance language: measure before you manage, and put output filters and human-review thresholds in front of downstream use. An evidence package is that filter in a shape an engineer can ship.
How do revision ceilings stop infinite hope?
A loop without a ceiling is a cost center with a chat UI. The worker will keep “trying” because trying is cheaper than admitting it is stuck. The evaluator + harness own the stop.
| Job class | Pilot ceiling | Why |
|---|---|---|
| High-cost side effect (refund, wire, public post) | 1 | Second try is a second blast radius |
| Draft + retrieve (support summary, research brief) | 3 | Enough to re-query; not enough to ramble |
| Cheap extract (OCR, field pull) | 1–2, then different path | A third fantasy pass invents lines |
| Exploratory research with no writes | 5 | Cost is tokens; still cap it |
Default for a Spurlock Studios pilot: three. The number is a constant in the state machine. It is not a suggestion in the system prompt.
Ceiling behavior that works:
- Worker produces a non-final artifact.
- Evaluator returns pass, fail+next, or escalate.
- Fail increments
revision_countand is the only reason the worker runs again. revision_count >= ceiling→ escalate with trace, cost, failures, and the last artifact.- Pass is the only path to a write tool or
done.
Never let the worker mark the run done. Terminal success is an evaluator privilege, or a human override with a name on it.
OWASP’s LLM06:2025 Excessive Agency names the failure this ceiling exists to stop: excessive functionality, excessive permissions, and excessive autonomy. A missing ceiling is excessive autonomy with a progress spinner.
Mechanical checks vs model judges — who owns the verdict?
Bias toward code. Models are for the residue. Anthropic’s evals post says it in one line: deterministic graders where possible, LLM graders where necessary, humans judiciously. OpenAI’s evals guide is the same stack — string_check and code graders for exact contracts, model graders for open-ended quality.
| Check type | Examples | Owner | Gates the write? |
|---|---|---|---|
| Schema | JSON matches Zod / JSON Schema | Code | Yes |
| Business rules | Totals equal line items; status in enum | Code | Yes |
| Retrieval contract | Citation required or explicit no_match | Code + light model | Yes if missing |
| Process | Forbidden tool; step budget; duplicate write | Code over the trace | Yes |
| Quality judgment | Summary completeness; email tone | Model evaluator | Only after code is green |
| Safety | Banned promises; PII patterns | Code + allowlists | Yes |
A model judge that re-checks “is this valid JSON?” is a bill you should not pay. A code check that tries to score “did we answer the customer’s actual question?” will rubber-stamp fluent misses.
Independence of context matters more than logo diversity. Same vendor, different prompt, no shared thread — fine, if you measure it. Same thread, “please confirm you did a good job” — not an evaluator.
Calibrating that model judge against human labels is a different job. This spoke stops at: do not turn the judge on until mechanical checks exist, and do not let the judge see the worker’s diary.
How do you wire evaluation into the agent loop?
Control flow that survives contact with a real ticket:
- Ingress validates the payload (required fields, size caps, auth).
- Worker produces an artifact in a non-final state. Writes are denied.
- Mechanical graders run on artifact + trace.
- If any blocking mechanical fail → attach evidence → revise or escalate.
- Model judge runs only on remaining soft/blocking judgment criteria.
- Pass → allow the specific write-back the job named. Not “whatever tool it wants.”
- Fail → attach evidence → worker revises with
nextonly. - Ceiling hit → escalate to a human with trace, cost, failures, and the last artifact.
| Stage | Allowed | Forbidden |
|---|---|---|
| Worker draft | Read tools, retrieve, compute | CRM write, refund, SMTP |
| Mechanical grade | Read artifact + trace | Mutate production |
| Model grade | Read evidence package | Mutate production |
| Pass gate | The one write the job named | A new tool the worker invented |
| Escalate | Human + full package | Silent “best effort” write |
-
doneis set by the harness after a pass, never by the model - Write tools are gated on evaluator pass, not on “the agent felt finished”
- Escalate payload includes
evaluator_versionandrevision_count - Kill switch exists outside the model (feature flag, queue pause)
Evaluators without sandboxes still let a bad action through between grades. Evaluators without a state machine still loop. Evaluators without a cost cap still burn money failing. Evaluation is first in design order, not the only layer. The operating manual maps the rest.
How do you test AI agents without fooling yourself?
A practical harness, in the order I actually run it after 500+ automations and 20,000+ hours on agentic systems:
1. Criteria before tools. If the buyer cannot sign pass/fail lines, stop. You are in a workshop.
2. A thin golden set, then grow it from failures.
Twenty real jobs beat a thousand synthetic toys. Include traps: empty retrieval, contradictory docs, hostile inputs, partial tool failures. Label expected severity (hard fail vs soft fail). Label three cases “must never auto-pass.” Those are regression anchors.
Do not boil the ocean. Anthropic’s evals piece is explicit: start small, encode expected behavior early, and grow. How you turn a production miss into a row — anonymize, stub tools, write the expected terminal — is the golden-set spoke. This page only requires that the set exists before write tools go live.
3. Split offline and online.
LangSmith’s table is the one I use in pilots: offline runs on a dataset (inputs, outputs, reference outputs) for regression; online runs on live traces (inputs and outputs only) for drift. Offline without online is academic. Online without offline is firefighting.
| Mode | Runs on | Has a reference? | Job |
|---|---|---|---|
| Offline | Frozen fixtures | Yes | Gate deploys |
| Online | Sampled production | Usually no | Catch new shapes |
| Harvest | A bad online run | Written after the fact | Feeds offline |
4. Track the numbers that earn autonomy.
- Pass rate on the golden set (by stratum: happy, trap, never-auto-pass)
- Average revisions to pass
- Cost per passing run
- Escalate rate
- Silent-fail rate (pass online, human later marks wrong) — expensive, sample it
Do not celebrate latency alone. Fast wrong is still wrong.
5. Version the evaluator.
When criteria change, bump evaluator_version. A jump in pass rate after you loosened criteria is not a model win. Treat evaluator diffs like worker diffs: same PR discipline, same rollback.
Who labels? A domain reviewer who will feel the pain of a wrong output — not only the engineer who owns the prompts. Disagreement between reviewers is useful. It means the criteria are still mush. Resolve the mush in the document before you ask a model to guess.
Which numbers earn a wider tool allowlist?
Autonomy is a promotion. The allowlist does not grow because the demo was charming.
| Signal | Widen | Hold | Shrink |
|---|---|---|---|
| Golden-set pass (blocking) | At or above the bar you set with the buyer | Within noise | Drop > agreed epsilon |
| Cost per passing run | Inside the band | At the ceiling | Above band for a week |
| Escalate rate | In the expected band | Spiking on one job type | Humans drowning |
| Silent-fail sample | Near zero on money/legal paths | Soft jobs only | Any money path |
| Never-auto-pass fixtures | Still fail | One flake under quarantine | Any of them pass |
Release rule I will put in a statement of work:
- Worker changes cannot ship if golden-set blocking pass rate drops more than the agreed epsilon.
- Cost per pass cannot leave the band without a named exception.
- A hotfix that ships without a new case has 48 hours to add one.
- Write-tool additions require a green run on every never-auto-pass fixture.
I have watched teams celebrate a 12-point pass-rate jump that was a deleted criterion. Version the rubric or you will lie to yourself with better charts.
Online samples hold for a defined window — a week of production, not an afternoon of dogfooding — before you hand the agent a refund tool. OWASP’s excessive-autonomy root cause is exactly this skip.
Why version the evaluator like production code?
The evaluator is a product artifact. It has a version, an owner, and a changelog. Shadow edits in a private doc recreate self-grading at the organizational layer.
| Change | Bump | Ship with |
|---|---|---|
| Typo in a criterion description | Patch | Note in the PR |
| New blocking criterion | Minor | At least one new fixture |
| Soft → blocking (or reverse) | Minor | Re-score the gold slice |
| Deleted criterion | Minor, and say so | Diff of pass-rate impact |
| New model judge prompt | Minor | Calibration note, even if thin |
| New mechanical check | Patch or minor | Unit tests for the check |
Store evaluator_version on every run. When a human later marks a silent fail, you need to know which contract they were graded against.
Publish the rubric where domain reviewers can edit it through a change process. If legal can rewrite “refund” in a Google Doc and the harness never sees it, you do not have an evaluator. You have folklore.
What breaks when you ship the agent first?
The failure mode is not “the demo was ugly.” The failure mode is a confident write you cannot reconstruct.
What I see after the first week of an agent-first build:
| Symptom | What actually broke | Cost |
|---|---|---|
| Green dashboard, angry buyer | Self-grade or a 0–100 slider hid the miss | Trust, then a rollback |
| Infinite “one more try” | No ceiling; the model is optimistic | Token burn, duplicate writes |
| “It used to work” | Prompt change, no golden set | A week of archaeology |
| Judge passes known-bad | Judge shares worker context, or criteria loosened | Silent-fail inventory |
| Refund went out twice | Write tools were live before pass/fail existed | Money + a lawyer |
| Team argues about “good” | Criteria were never signed | The agent becomes the scapegoat |
Fix is not a better model. Fix is to stop, write the contract, put writes behind a pass, and only then turn the loop back on. 35,000+ hours saved for clients came from deleting busywork — not from letting an ungraded loop create more of it.
If leadership rewards demos over scores, evaluation dies socially before it dies technically. Put pass rate and cost per pass in the same meeting as the feature launch. Refuse to widen the allowlist when scores regress.
Worked example: invoice line extraction
Job: extract line items from supplier PDFs into JSON for AP.
Mechanical criteria (write these first):
| Id | Check | Blocking? |
|---|---|---|
schema_valid | JSON matches the AP contract | Yes |
qty_positive | Every quantity > 0 | Yes |
currency_allowlist | Currency in {USD, EUR, GBP} | Yes |
totals_match | Sum of lines equals stated total within one cent | Yes |
vendor_known | vendor_id in ERP or flagged unknown_vendor | Yes |
Model criteria (narrow residue):
| Id | Check | Blocking? |
|---|---|---|
desc_not_boilerplate | Description is not empty boilerplate | Soft |
unreadable_source | Handwritten / unreadable PDF → fail with that code, do not invent lines | Yes |
Revision policy: one re-extract with a different OCR path, then escalate. No third fantasy pass.
Evidence on a totals fail: totals_match: header 1840.00 vs lines 1835.88, delta 4.12. The worker’s next action is “re-read the tax line,” not “try harder.”
That evaluator can be half code by lunch. The agent — or a simpler pipeline — becomes measurable the same day. Most “we need a bigger model” requests on this job die when the totals check is enforced.
Traps worth putting in the first twenty fixtures:
- Empty or whitespace-only PDF text layer
- Totals that include tax in the header and exclude it on the lines
- Mixed currencies on one invoice
- Partial OCR (one of three pages blank)
- Instructions in the PDF footer that try to override the rubric
- A near-duplicate invoice that should be idempotent
If your suite is only sunny-path invoices, your pass rate is a vanity metric.
What does evaluator-first look like in a five-day pilot?
The $1,500 · 5-day agentic pilot starts with criteria, not with a toy chatbot.
| Day | What ships | What does not |
|---|---|---|
| 1 | Job contract, irreversible list, evaluator shape, evaluator_version v0 | Tool shopping |
| 2 | Mechanical graders + evidence schema; first 15–20 fixtures | Write tools |
| 3 | Worker in a sandbox against the set; revision ceiling wired | Production CRM writes |
| 4 | Model judge on residue only; escalate path with the full package | Allowlist expansion |
| 5 | One passing path on your data; quote for hardening; scores on the table | “Autonomous” as a slide |
You keep the agent either way. Details: /agentic. Start: /contact?intent=agentic-pilot.
If a prospect wants agents without criteria, I push them toward a workflow or a scoping workshop. A loop with no grader is a demo you will pay for twice.
How this spoke sits next to golden sets and the rest of the stack
This page owns build order. It does not own fixture harvest, and it does not own judge meta-eval.
| Question | This spoke | Somewhere else |
|---|---|---|
| What do I build first? | Evaluator: criteria, evidence, ceiling | — |
| How do I turn a bad run into a CI row? | Require that you do it | Golden sets from agent failures |
| Is the model judge lying? | Do not turn it on until code checks exist | A calibration pass, not this page |
| What else has to exist? | Design-order first, not only-layer | Operating manual |
If you only remember one sentence: the evaluator is how you know whether anything else worked. Autonomy is a score you earn after that sentence is true.
FAQ
What is AI agent evaluation?
AI agent evaluation is judging agent outputs against explicit acceptance criteria with an independent process — mechanical checks plus a separate judge — and measuring pass rates on a frozen suite and on sampled production. It is not the agent saying “looks good.” The unit you grade is the trace plus the outcome in the world, not the last sentence in the chat.
How do you test AI agents before production?
Write criteria first, build a golden set of real jobs, implement an independent evaluator, run the suite on every meaningful change, and keep a human escalate path. Soft-launch with write gates until online scores match offline expectations. Harvest new production misses into fixtures instead of “fixing forward.”
Should the evaluator be another LLM?
Often only partly. Use code for anything assertable. Use a model evaluator when judgment is required — and still force structured output with per-criterion verdicts and evidence. Same-vendor or cross-vendor both work; independence of context matters more than logo diversity. Do not feed the worker’s private reasoning to the judge.
How many revision attempts should an agent get?
Three is a strong default for pilots. Some jobs warrant one (high-cost side effects). Some warrant five (cheap drafts with no writes). The number must be explicit and enforced by the harness, not left to the model’s optimism. When the ceiling hits, escalate with the evidence package.
When is evaluation “good enough” to widen autonomy?
When golden-set pass rate, cost per pass, and escalate rate meet the bar you set with the buyer — and online samples hold for a defined window. Never-auto-pass fixtures must still fail. Autonomy is earned by those scores, not by demo applause or a new model card.
How does this fit Spurlock Studios engagements?
Every agentic pilot and build assumes an evaluator harness. Day one is the job contract and the grader, not a chatbot. If a prospect wants agents without criteria, we push them toward automation or a scoping workshop first. See the operating manual and /agentic.
CTA
Criteria, evidence, ceiling — then tools. Install that order on your job in five days: /agentic · /contact?intent=agentic-pilot.
What questions does this article answer?
- What is AI agent evaluation?
- AI agent evaluation is judging agent outputs against explicit acceptance criteria with an independent process — mechanical checks plus a separate judge — and measuring pass rates on a frozen suite and on sampled production. It is not the agent saying “looks good.” The unit you grade is the trace plus the outcome in the world, not the last sentence in the chat.
- How do you test AI agents before production?
- Write criteria first, build a golden set of real jobs, implement an independent evaluator, run the suite on every meaningful change, and keep a human escalate path. Soft-launch with write gates until online scores match offline expectations. Harvest new production misses into fixtures instead of “fixing forward.”
- Should the evaluator be another LLM?
- Often only partly. Use code for anything assertable. Use a model evaluator when judgment is required — and still force structured output with per-criterion verdicts and evidence. Same-vendor or cross-vendor both work; independence of context matters more than logo diversity. Do not feed the worker’s private reasoning to the judge.
- How many revision attempts should an agent get?
- Three is a strong default for pilots. Some jobs warrant one (high-cost side effects). Some warrant five (cheap drafts with no writes). The number must be explicit and enforced by the harness, not left to the model’s optimism. When the ceiling hits, escalate with the evidence package.
- When is evaluation “good enough” to widen autonomy?
- When golden-set pass rate, cost per pass, and escalate rate meet the bar you set with the buyer — and online samples hold for a defined window. Never-auto-pass fixtures must still fail. Autonomy is earned by those scores, not by demo applause or a new model card.
- How does this fit Spurlock Studios engagements?
- Every agentic pilot and build assumes an evaluator harness. Day one is the job contract and the grader, not a chatbot. If a prospect wants agents without criteria, we push them toward automation or a scoping workshop first. See the [operating manual](/blog/agentic-systems-operating-manual) and [/agentic](/agentic).
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.