Spurlock Studios
Contact
Share LinkedIn X
A violet ring. Thesis: BUILD EVALUATOR BEFORE AGENT.

If you ship the agent before the evaluator, you are guessing. Accuracy jumps I trust never came from a warmer system prompt. They came from a second component whose only job is to disagree, using structured evidence the worker did not write.

This spoke sits under the Agentic Systems Operating Manual. The parent maps the stack. Here I own build order: independent criteria, evidence, and revision ceilings — then autonomy. How you harvest a bad run into a fixture lives in golden sets from agent failures. How Spurlock Studios installs the first harness is on /agentic.

The short answer

  • Write pass/fail criteria with the buyer before anyone opens a tool schema.
  • The evaluator sees criteria + artifact + retrieval evidence. Not the worker’s private reasoning.
  • Code owns schema, enums, totals, allowlists, and citation presence. A model judge owns the residue.
  • Failures name a criterion and a quote. Three revisions is the pilot default; the harness enforces the ceiling.
  • Widen write tools only when golden-set pass rate, cost per pass, and escalate rate hold on sampled production.

Why does self-grading fail?

A model asked to check its own output is continuing a story in which it already “finished.” The verdict is correlated with the work. In production that looks like a green run where the tool errored, the summary soft-pedaled the error, and nobody opened the trace.

Self-grading is a hint inside a worker. It is not your metric. Zheng et al. documented the same family of failure in Judging LLM-as-a-Judge with MT-Bench: position bias, verbosity bias, and self-enhancement — models prefer their own prose. That paper is about chat assistants. Agents make it worse, because the worker also chose the tools and wrote the story of why those tools were fine.

SetupWhat the judge seesWhat you actually measured
Worker says “looks good”Its own draft + its own rationaleConfidence, not correctness
Same model, same threadThe conversation that produced the artifactContinuity of a story
Independent evaluatorCriteria + artifact + evidence packageThe contract you wrote
Human overrideSame package plus the traceThe last veto, not the first grader
  • The worker cannot set status: done
  • The judge does not receive chain-of-thought from the worker
  • Every blocking fail names a criterion id
  • Known-bad fixtures still fail after a prompt tweak
  • A human can replay the evidence package without the chat

If any box is unchecked, you are still grading homework.

What is an independent evaluator?

An evaluator is a component — code, model, or both — that returns a structured verdict. It is not a vibes score. It is not a thumbs-up in Slack. Anthropic’s Demystifying evals for AI agents (Jan 2026) splits the same object into a task, a grader, a transcript, and an outcome. The outcome is the state of the world, not the last sentence the model uttered.

{
  "verdict": "fail",
  "evaluator_version": "refund-policy-v3",
  "failures": [
    {
      "criterion": "severity_present_or_explicit_no_match",
      "severity": "blocking",
      "evidence": "summary claims policy X with zero citation URLs"
    }
  ],
  "next": "re-retrieve with query focused on refund policy; rewrite summary"
}

Rules that make it real:

  1. Independent context. Criteria, artifact, and evidence. No worker monologue.
  2. Mechanical first. Schema, enums, regex allowlists, arithmetic, required ids — assert in code.
  3. Model only for judgment. Tone, omission of material risk, “did this answer the question asked.”
  4. Evidence, not vibes. “Not good enough” is useless. “Criterion 3 failed because…” is a next action.
  5. Revision ceiling. Usually three. Then escalate with the full package.
  6. Read-only judge. The evaluator does not call write tools. Ever.

OpenAI’s agent evals and trace grading treat the trace as the unit you grade — model calls, tool calls, guardrails, handoffs — then tell you to move graded traces into a dataset when you need repeatability. That is the same split: one run is a clue; a fixture is a gate.

Microsoft Foundry’s agent evaluators make the same cut in vendor language: system evaluation (did the job finish with a usable deliverable) versus process evaluation (did the steps stay inside the contract). Binary pass/fail beats a 0–100 “quality” slider that hides which contract broke.

What criteria must exist before you write tools?

Sit with the buyer. Translate “good” into pass/fail lines. If you cannot, you are not ready to build. Ambiguous taste is a product workshop, not an agent ticket. Anthropic’s Building effective agents is blunt: the evaluator-refiner pattern only fits when you have clear evaluation criteria and iterative refinement actually improves the artifact. No criteria, no loop.

Criterion shapeExampleOwner
SchemaJSON matches the contract; required keys presentCode
ArithmeticLine totals equal header total within one centCode
PolicyRefund promise requires a citation URL or no_matchCode + light model
CompletenessEvery asked field is answered or explicitly declinedModel judge
SafetyNo PII in outbound email; no “we will sue them”Code allowlists
ProcessForbidden tool never appears in the traceCode over the trace

Write criteria in this order:

  1. Name the job in one sentence the buyer will repeat.
  2. List irreversible side effects (money, legal, public send, delete).
  3. For each side effect, write the blocking fail that must stop the write.
  4. Add soft fails that may pass the run but open a ticket.
  5. Freeze the list as evaluator_version. Tools come after.
  • Every blocking criterion is testable on a fixture without a human in the room
  • Soft vs blocking is labeled, not implied
  • “Be helpful” and “be professional” have been deleted or split into observables
  • The buyer signed the list, not just the demo script

A criterion that needs a page of interpretation is two criteria, or it is not a criterion.

Workshop script I run on day one of a pilot. Forty-five minutes. Whiteboard only.

  1. “What write, if wrong, do you have to unwind by hand?” Circle those. They are blocking.
  2. “Show me three tickets you already lost sleep over.” Those become never-auto-pass fixtures.
  3. “What would you accept as proof the agent was right?” If the answer is a feeling, keep workshopping.
  4. Read the list back. If two people disagree, the criterion is still mush. Split or delete.
  5. Freeze v0. Tools stay closed until v0 has an owner and a version string.

If step 3 dies, you do not have an agent job yet. You have a process argument wearing a model.

What counts as evidence, not vibes?

A verdict without evidence trains the worker to argue. A verdict with a quote trains the worker to fix a named hole. LangSmith’s evaluation concepts split graders into code, LLM-as-judge, human, and pairwise — and they still expect you to review scores and tune the judge prompt. Few-shot examples belong on the evaluator, never as a paste of golden answers into the worker.

Evidence typeGoodUseless
QuoteExact span from the artifact or tool payload“The tone felt off”
Idinvoice_id, policy_url, tool_call_id“It used the right tool, I think”
Counterfactual“Totals differ by $4.12”“Numbers seem high”
Absence“Zero citation URLs; criterion requires one or no_match”“Sources were weak”
Trace factemail.send appeared after a fail“It probably sent”

Minimum evidence package the judge returns:

  1. criterion id from the versioned rubric.
  2. severity: blocking or soft.
  3. evidence: a quote or a computed delta, ≤200 characters.
  4. next: one action the worker may take, or escalate.

If the judge cannot fill those four fields, the criterion is not ready. Tighten it until a human could fill the same form from the artifact alone.

NIST’s AI 600-1 Generative AI Profile puts this in governance language: measure before you manage, and put output filters and human-review thresholds in front of downstream use. An evidence package is that filter in a shape an engineer can ship.

How do revision ceilings stop infinite hope?

A loop without a ceiling is a cost center with a chat UI. The worker will keep “trying” because trying is cheaper than admitting it is stuck. The evaluator + harness own the stop.

Job classPilot ceilingWhy
High-cost side effect (refund, wire, public post)1Second try is a second blast radius
Draft + retrieve (support summary, research brief)3Enough to re-query; not enough to ramble
Cheap extract (OCR, field pull)1–2, then different pathA third fantasy pass invents lines
Exploratory research with no writes5Cost is tokens; still cap it

Default for a Spurlock Studios pilot: three. The number is a constant in the state machine. It is not a suggestion in the system prompt.

Ceiling behavior that works:

  1. Worker produces a non-final artifact.
  2. Evaluator returns pass, fail+next, or escalate.
  3. Fail increments revision_count and is the only reason the worker runs again.
  4. revision_count >= ceiling → escalate with trace, cost, failures, and the last artifact.
  5. Pass is the only path to a write tool or done.

Never let the worker mark the run done. Terminal success is an evaluator privilege, or a human override with a name on it.

OWASP’s LLM06:2025 Excessive Agency names the failure this ceiling exists to stop: excessive functionality, excessive permissions, and excessive autonomy. A missing ceiling is excessive autonomy with a progress spinner.

Mechanical checks vs model judges — who owns the verdict?

Bias toward code. Models are for the residue. Anthropic’s evals post says it in one line: deterministic graders where possible, LLM graders where necessary, humans judiciously. OpenAI’s evals guide is the same stack — string_check and code graders for exact contracts, model graders for open-ended quality.

Check typeExamplesOwnerGates the write?
SchemaJSON matches Zod / JSON SchemaCodeYes
Business rulesTotals equal line items; status in enumCodeYes
Retrieval contractCitation required or explicit no_matchCode + light modelYes if missing
ProcessForbidden tool; step budget; duplicate writeCode over the traceYes
Quality judgmentSummary completeness; email toneModel evaluatorOnly after code is green
SafetyBanned promises; PII patternsCode + allowlistsYes

A model judge that re-checks “is this valid JSON?” is a bill you should not pay. A code check that tries to score “did we answer the customer’s actual question?” will rubber-stamp fluent misses.

Independence of context matters more than logo diversity. Same vendor, different prompt, no shared thread — fine, if you measure it. Same thread, “please confirm you did a good job” — not an evaluator.

Calibrating that model judge against human labels is a different job. This spoke stops at: do not turn the judge on until mechanical checks exist, and do not let the judge see the worker’s diary.

How do you wire evaluation into the agent loop?

Control flow that survives contact with a real ticket:

  1. Ingress validates the payload (required fields, size caps, auth).
  2. Worker produces an artifact in a non-final state. Writes are denied.
  3. Mechanical graders run on artifact + trace.
  4. If any blocking mechanical fail → attach evidence → revise or escalate.
  5. Model judge runs only on remaining soft/blocking judgment criteria.
  6. Pass → allow the specific write-back the job named. Not “whatever tool it wants.”
  7. Fail → attach evidence → worker revises with next only.
  8. Ceiling hit → escalate to a human with trace, cost, failures, and the last artifact.
StageAllowedForbidden
Worker draftRead tools, retrieve, computeCRM write, refund, SMTP
Mechanical gradeRead artifact + traceMutate production
Model gradeRead evidence packageMutate production
Pass gateThe one write the job namedA new tool the worker invented
EscalateHuman + full packageSilent “best effort” write
  • done is set by the harness after a pass, never by the model
  • Write tools are gated on evaluator pass, not on “the agent felt finished”
  • Escalate payload includes evaluator_version and revision_count
  • Kill switch exists outside the model (feature flag, queue pause)

Evaluators without sandboxes still let a bad action through between grades. Evaluators without a state machine still loop. Evaluators without a cost cap still burn money failing. Evaluation is first in design order, not the only layer. The operating manual maps the rest.

How do you test AI agents without fooling yourself?

A practical harness, in the order I actually run it after 500+ automations and 20,000+ hours on agentic systems:

1. Criteria before tools. If the buyer cannot sign pass/fail lines, stop. You are in a workshop.

2. A thin golden set, then grow it from failures.

Twenty real jobs beat a thousand synthetic toys. Include traps: empty retrieval, contradictory docs, hostile inputs, partial tool failures. Label expected severity (hard fail vs soft fail). Label three cases “must never auto-pass.” Those are regression anchors.

Do not boil the ocean. Anthropic’s evals piece is explicit: start small, encode expected behavior early, and grow. How you turn a production miss into a row — anonymize, stub tools, write the expected terminal — is the golden-set spoke. This page only requires that the set exists before write tools go live.

3. Split offline and online.

LangSmith’s table is the one I use in pilots: offline runs on a dataset (inputs, outputs, reference outputs) for regression; online runs on live traces (inputs and outputs only) for drift. Offline without online is academic. Online without offline is firefighting.

ModeRuns onHas a reference?Job
OfflineFrozen fixturesYesGate deploys
OnlineSampled productionUsually noCatch new shapes
HarvestA bad online runWritten after the factFeeds offline

4. Track the numbers that earn autonomy.

  • Pass rate on the golden set (by stratum: happy, trap, never-auto-pass)
  • Average revisions to pass
  • Cost per passing run
  • Escalate rate
  • Silent-fail rate (pass online, human later marks wrong) — expensive, sample it

Do not celebrate latency alone. Fast wrong is still wrong.

5. Version the evaluator.

When criteria change, bump evaluator_version. A jump in pass rate after you loosened criteria is not a model win. Treat evaluator diffs like worker diffs: same PR discipline, same rollback.

Who labels? A domain reviewer who will feel the pain of a wrong output — not only the engineer who owns the prompts. Disagreement between reviewers is useful. It means the criteria are still mush. Resolve the mush in the document before you ask a model to guess.

Which numbers earn a wider tool allowlist?

Autonomy is a promotion. The allowlist does not grow because the demo was charming.

SignalWidenHoldShrink
Golden-set pass (blocking)At or above the bar you set with the buyerWithin noiseDrop > agreed epsilon
Cost per passing runInside the bandAt the ceilingAbove band for a week
Escalate rateIn the expected bandSpiking on one job typeHumans drowning
Silent-fail sampleNear zero on money/legal pathsSoft jobs onlyAny money path
Never-auto-pass fixturesStill failOne flake under quarantineAny of them pass

Release rule I will put in a statement of work:

  1. Worker changes cannot ship if golden-set blocking pass rate drops more than the agreed epsilon.
  2. Cost per pass cannot leave the band without a named exception.
  3. A hotfix that ships without a new case has 48 hours to add one.
  4. Write-tool additions require a green run on every never-auto-pass fixture.

I have watched teams celebrate a 12-point pass-rate jump that was a deleted criterion. Version the rubric or you will lie to yourself with better charts.

Online samples hold for a defined window — a week of production, not an afternoon of dogfooding — before you hand the agent a refund tool. OWASP’s excessive-autonomy root cause is exactly this skip.

Why version the evaluator like production code?

The evaluator is a product artifact. It has a version, an owner, and a changelog. Shadow edits in a private doc recreate self-grading at the organizational layer.

ChangeBumpShip with
Typo in a criterion descriptionPatchNote in the PR
New blocking criterionMinorAt least one new fixture
Soft → blocking (or reverse)MinorRe-score the gold slice
Deleted criterionMinor, and say soDiff of pass-rate impact
New model judge promptMinorCalibration note, even if thin
New mechanical checkPatch or minorUnit tests for the check

Store evaluator_version on every run. When a human later marks a silent fail, you need to know which contract they were graded against.

Publish the rubric where domain reviewers can edit it through a change process. If legal can rewrite “refund” in a Google Doc and the harness never sees it, you do not have an evaluator. You have folklore.

What breaks when you ship the agent first?

The failure mode is not “the demo was ugly.” The failure mode is a confident write you cannot reconstruct.

What I see after the first week of an agent-first build:

SymptomWhat actually brokeCost
Green dashboard, angry buyerSelf-grade or a 0–100 slider hid the missTrust, then a rollback
Infinite “one more try”No ceiling; the model is optimisticToken burn, duplicate writes
“It used to work”Prompt change, no golden setA week of archaeology
Judge passes known-badJudge shares worker context, or criteria loosenedSilent-fail inventory
Refund went out twiceWrite tools were live before pass/fail existedMoney + a lawyer
Team argues about “good”Criteria were never signedThe agent becomes the scapegoat

Fix is not a better model. Fix is to stop, write the contract, put writes behind a pass, and only then turn the loop back on. 35,000+ hours saved for clients came from deleting busywork — not from letting an ungraded loop create more of it.

If leadership rewards demos over scores, evaluation dies socially before it dies technically. Put pass rate and cost per pass in the same meeting as the feature launch. Refuse to widen the allowlist when scores regress.

Worked example: invoice line extraction

Job: extract line items from supplier PDFs into JSON for AP.

Mechanical criteria (write these first):

IdCheckBlocking?
schema_validJSON matches the AP contractYes
qty_positiveEvery quantity > 0Yes
currency_allowlistCurrency in {USD, EUR, GBP}Yes
totals_matchSum of lines equals stated total within one centYes
vendor_knownvendor_id in ERP or flagged unknown_vendorYes

Model criteria (narrow residue):

IdCheckBlocking?
desc_not_boilerplateDescription is not empty boilerplateSoft
unreadable_sourceHandwritten / unreadable PDF → fail with that code, do not invent linesYes

Revision policy: one re-extract with a different OCR path, then escalate. No third fantasy pass.

Evidence on a totals fail: totals_match: header 1840.00 vs lines 1835.88, delta 4.12. The worker’s next action is “re-read the tax line,” not “try harder.”

That evaluator can be half code by lunch. The agent — or a simpler pipeline — becomes measurable the same day. Most “we need a bigger model” requests on this job die when the totals check is enforced.

Traps worth putting in the first twenty fixtures:

  • Empty or whitespace-only PDF text layer
  • Totals that include tax in the header and exclude it on the lines
  • Mixed currencies on one invoice
  • Partial OCR (one of three pages blank)
  • Instructions in the PDF footer that try to override the rubric
  • A near-duplicate invoice that should be idempotent

If your suite is only sunny-path invoices, your pass rate is a vanity metric.

What does evaluator-first look like in a five-day pilot?

The $1,500 · 5-day agentic pilot starts with criteria, not with a toy chatbot.

DayWhat shipsWhat does not
1Job contract, irreversible list, evaluator shape, evaluator_version v0Tool shopping
2Mechanical graders + evidence schema; first 15–20 fixturesWrite tools
3Worker in a sandbox against the set; revision ceiling wiredProduction CRM writes
4Model judge on residue only; escalate path with the full packageAllowlist expansion
5One passing path on your data; quote for hardening; scores on the table“Autonomous” as a slide

You keep the agent either way. Details: /agentic. Start: /contact?intent=agentic-pilot.

If a prospect wants agents without criteria, I push them toward a workflow or a scoping workshop. A loop with no grader is a demo you will pay for twice.

How this spoke sits next to golden sets and the rest of the stack

This page owns build order. It does not own fixture harvest, and it does not own judge meta-eval.

QuestionThis spokeSomewhere else
What do I build first?Evaluator: criteria, evidence, ceiling—
How do I turn a bad run into a CI row?Require that you do itGolden sets from agent failures
Is the model judge lying?Do not turn it on until code checks existA calibration pass, not this page
What else has to exist?Design-order first, not only-layerOperating manual

If you only remember one sentence: the evaluator is how you know whether anything else worked. Autonomy is a score you earn after that sentence is true.

FAQ

What is AI agent evaluation?

AI agent evaluation is judging agent outputs against explicit acceptance criteria with an independent process — mechanical checks plus a separate judge — and measuring pass rates on a frozen suite and on sampled production. It is not the agent saying “looks good.” The unit you grade is the trace plus the outcome in the world, not the last sentence in the chat.

How do you test AI agents before production?

Write criteria first, build a golden set of real jobs, implement an independent evaluator, run the suite on every meaningful change, and keep a human escalate path. Soft-launch with write gates until online scores match offline expectations. Harvest new production misses into fixtures instead of “fixing forward.”

Should the evaluator be another LLM?

Often only partly. Use code for anything assertable. Use a model evaluator when judgment is required — and still force structured output with per-criterion verdicts and evidence. Same-vendor or cross-vendor both work; independence of context matters more than logo diversity. Do not feed the worker’s private reasoning to the judge.

How many revision attempts should an agent get?

Three is a strong default for pilots. Some jobs warrant one (high-cost side effects). Some warrant five (cheap drafts with no writes). The number must be explicit and enforced by the harness, not left to the model’s optimism. When the ceiling hits, escalate with the evidence package.

When is evaluation “good enough” to widen autonomy?

When golden-set pass rate, cost per pass, and escalate rate meet the bar you set with the buyer — and online samples hold for a defined window. Never-auto-pass fixtures must still fail. Autonomy is earned by those scores, not by demo applause or a new model card.

How does this fit Spurlock Studios engagements?

Every agentic pilot and build assumes an evaluator harness. Day one is the job contract and the grader, not a chatbot. If a prospect wants agents without criteria, we push them toward automation or a scoping workshop first. See the operating manual and /agentic.

CTA

Criteria, evidence, ceiling — then tools. Install that order on your job in five days: /agentic · /contact?intent=agentic-pilot.

FAQ

What questions does this article answer?

What is AI agent evaluation?
AI agent evaluation is judging agent outputs against explicit acceptance criteria with an independent process — mechanical checks plus a separate judge — and measuring pass rates on a frozen suite and on sampled production. It is not the agent saying “looks good.” The unit you grade is the trace plus the outcome in the world, not the last sentence in the chat.
How do you test AI agents before production?
Write criteria first, build a golden set of real jobs, implement an independent evaluator, run the suite on every meaningful change, and keep a human escalate path. Soft-launch with write gates until online scores match offline expectations. Harvest new production misses into fixtures instead of “fixing forward.”
Should the evaluator be another LLM?
Often only partly. Use code for anything assertable. Use a model evaluator when judgment is required — and still force structured output with per-criterion verdicts and evidence. Same-vendor or cross-vendor both work; independence of context matters more than logo diversity. Do not feed the worker’s private reasoning to the judge.
How many revision attempts should an agent get?
Three is a strong default for pilots. Some jobs warrant one (high-cost side effects). Some warrant five (cheap drafts with no writes). The number must be explicit and enforced by the harness, not left to the model’s optimism. When the ceiling hits, escalate with the evidence package.
When is evaluation “good enough” to widen autonomy?
When golden-set pass rate, cost per pass, and escalate rate meet the bar you set with the buyer — and online samples hold for a defined window. Never-auto-pass fixtures must still fail. Autonomy is earned by those scores, not by demo applause or a new model card.
How does this fit Spurlock Studios engagements?
Every agentic pilot and build assumes an evaluator harness. Day one is the job contract and the grader, not a chatbot. If a prospect wants agents without criteria, we push them toward automation or a scoping workshop first. See the [operating manual](/blog/agentic-systems-operating-manual) and [/agentic](/agentic).
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot