The Evaluator Is the Product
Agent accuracy did not come from a better prompt or a bigger model. It came from separating the thing that does the work from the thing that judges it.
William Spurlock Founder — Spurlock Studios Updated 12 MIN
The single largest accuracy jump I have measured in an agentic system did not come from a model upgrade. It came from deleting a paragraph of a prompt and adding a second agent whose only job was to disagree.
This sits under the Agentic Systems Operating Manual. The parent maps the whole stack. Here I own one move: the evaluator is the product, and the worker is a thing the evaluator is allowed to reject. How you install that harness on a first job is on /agentic.
The short answer
- A model asked to check its own output is generating a continuation of a context that already contains the assertion that the work is done.
- Split the roles. The worker produces. A separate evaluator, with its own context, returns a verdict against criteria you wrote.
- Keep the evaluator off the worker’s reasoning. Give it the criteria and the artifact. Nothing else.
- Make criteria mechanical wherever you can. “Tests pass” beats “code is good.”
- Failures return evidence, not vibes. Cap retries, then escalate with the full trace.
Why does self-grading fail?
A model asked to check its own output is not checking anything. It is generating a continuation of a context that already contains the assertion that the work is done. The self-assessment is conditioned on the work, which is exactly the correlation you were trying to break.
That sounds abstract until you watch a run. The tool call errored. The model summarised the error as a minor issue. The run closed green. Nobody opened the trace because the last sentence said success. You did not measure correctness. You measured the model’s willingness to keep the story coherent.
Self-grading fails for a boring structural reason. The same hidden state that produced the artifact is still in the window when you ask “did you finish?” The model has already spent tokens arguing that it finished. Asking it to reverse that argument is asking it to contradict itself in public. Most of the time it will not. It will soften, rephrase, and call the miss a nuance.
In practice this shows up as agents that confidently report success on tasks they did not complete. The missed tool, the empty field, the partial write — all of it gets narrated into a passing grade. Raw single-pass accuracy on the internal benchmark I use sat around 72%. Most of the misses were not wrong answers — they were unnoticed failures. The worker was wrong and also in charge of noticing.
A self-check inside the worker prompt is a hint. It is not a metric. If the same component can set status: done, you are still grading homework.
Checklist for whether you are still self-grading:
- The worker can mark the run complete without an external verdict
- The “evaluator” is a paragraph at the end of the worker prompt
- The judge receives the worker’s chain-of-thought or tool-call narration
- A known-bad fixture still closes green after a prompt tweak
- Nobody can replay the evidence package without the chat transcript
If any of those are true, the accuracy number you are reporting is a confidence score.
How do you separate the worker from the evaluator?
The fix is structural. The worker produces output. A separate evaluator, with its own context containing only the acceptance criteria and the artifact, returns a verdict:
{
"verdict": "fail",
"criterion": "all tests pass",
"evidence": "3 failing in auth.test.ts",
"next": "fix the null guard in verifyToken"
}
That object is the product. Not the worker’s prose. Not a thumbs-up in Slack. A verdict with a named criterion, a quote from the world, and a next action the worker can actually take.
Three things make the split hold.
The evaluator never sees the worker’s reasoning. It sees the criteria and the artifact. Giving it the transcript reintroduces the correlation you just paid to remove. The moment the judge can read “I verified the tests and they look good,” you are back to self-grading with extra steps. Hide the inner monologue. Pass the file, the API response, the screenshot, the test output — the thing a human would inspect.
Criteria are mechanical wherever possible. “Tests pass” beats “code is good.” “Returns valid JSON matching this schema” beats “output is well-formed.” Anything you can assert in code should be asserted in code, and the model should only judge what genuinely needs judgement. Schema, enums, totals, allowlists, citation presence, required fields: those are functions. A model judge is for residue — tone against a brand brief, whether a summary omitted a material risk, whether two passages contradict. If you start with the model judge, you will never write the cheap checks.
Failure returns evidence, not vibes. The worker cannot act on “this is not right.” It can act on “three tests fail, here is the output.” A fail without a criterion id and a quote is a vibe. A vibe produces another speculative rewrite. Evidence produces a targeted patch, or an honest escalation.
With that loop in place and a ceiling of three correction attempts, the same benchmark runs at 99.4%. Same models. The difference is entirely architectural. I did not find a smarter worker. I stopped letting the worker grade the exam.
The sibling post on building the evaluator before the agent is the build-order version of this argument. This one is the product claim: if you only ship one new component, ship the judge.
What makes an evaluator verdict usable?
A usable verdict is small, typed, and boring. The JSON above is the shape. You can vary field names. You cannot vary the jobs.
verdict is pass or fail — or a third state that means “I cannot tell, escalate.” Fuzzy scores (“0.73 quality”) are how teams argue about thresholds instead of fixing the artifact. If you need a score for a dashboard, compute it from pass/fail over a set. Do not ask the judge for a vibe number and then treat 0.7 as a ship.
criterion names the rule that fired. “all tests pass.” “email field is present and parseable.” “no write to production.” Named criteria are how you debug the evaluator itself. When pass rate jumps after you loosened a rubric, that is not a model win. That is a judge that got friendlier. Version the criteria the way you version code.
evidence is a quote from the artifact or the tool output. Not a paraphrase. The worker — and later, a human — must be able to see the same bytes. If the evidence cannot be shown next to the criterion, the verdict is not replayable, and an unreplayable verdict will not survive a disagreement.
next is optional but it is how the loop closes. “Fix the null guard in verifyToken” is an instruction. “Try again” is how you burn the retry budget on the same miss.
A verdict that cannot be logged, replayed, and compared across runs is a comment. Comments are not a product.
Why does the retry ceiling matter?
Unbounded correction loops are how you spend $400 discovering that a task is impossible. Three attempts, then stop and escalate to a human with the full trace. An agent that knows how to give up is more useful than one that does not, because the failure arrives while someone can still do something about it.
The ceiling is a harness rule, not a prompt suggestion. If the stop lives in the system prompt, the model will negotiate it. If the stop lives in the loop that refuses to call the worker a fourth time, the model can plead all it wants. You still escalate.
Three is the number I use as a default, not a law of nature. The point is that there is a number, it is small, and hitting it is a first-class outcome: escalated, with the three artifacts, the three verdicts, and the original criteria attached. That package is what a human needs. A 40-turn transcript is what a human skims and then guesses about.
The same instinct shows up when agents loop on failed tools. Retrying without progress is not diligence. It is a missing stop condition. The evaluator ceiling and the tool-loop detector are the same idea applied to two different clocks: revisions of the artifact, and repeats of a dead call.
What the ceiling buys you:
- A bound on spend per task, so a stuck job cannot quietly become a bill.
- A bound on wall-clock, so a human still has the afternoon.
- A training signal: the cases that hit the ceiling are the cases you harvest into fixtures, not the cases you hand-wave as “the model having a bad day.”
Without a ceiling, you do not have an evaluator. You have a worker arguing with a mirror until the card declines.
What does this change about how you build agents?
Most teams shipping agents are optimising the wrong surface. They iterate on the worker prompt, upgrade the model, add tools. Meanwhile there is no independent judgement anywhere in the system, so nobody can tell whether any of it helped. A new tool that makes the worker faster at producing plausible wrong answers looks like a win until you measure against criteria the worker did not write.
Build the evaluator first. Write the pass/fail list with the person who will live with the output. Encode the mechanical checks in code. Only then give the worker tools. That order feels slow on day one and is the only reason day thirty is not a demo you are afraid to rerun.
It also changes what you ship to a customer. A demo is a worker that looked good once. A product is a worker that can be failed, revised a bounded number of times, and handed to a human with evidence when it still cannot pass. The evaluator is the thing that makes that sentence true. The model is a detail.
If you are already in production without this split, you do not need a rewrite of the worker. You need a second context, a criteria file, a verdict schema, and a loop that will not mark done until the judge says pass. That is the whole intervention. It is also the whole product.
Why can’t the worker grade itself?
Because the grade is generated from the same context that produced the work. The model is not inspecting an external result. It is continuing a story in which it already succeeded. That continuation is correlated with the artifact, which is the opposite of a check. Split the context or you do not have an evaluation — you have a more polite success message.
How many correction attempts should an agent get?
Enough to use the evidence, not enough to wander. I default to three, then escalate with the full trace. The number is a harness limit, not a suggestion in the prompt. Unbounded retries are how you spend real money confirming that the task, the tools, or the criteria were never going to work.
What questions does this article answer?
- Why can't the worker grade itself?
- Because the grade is generated from the same context that produced the work. The model is not inspecting an external result. It is continuing a story in which it already succeeded. That continuation is correlated with the artifact, which is the opposite of a check. Split the context or you do not have an evaluation — you have a more polite success message.
- How many correction attempts should an agent get?
- Enough to use the evidence, not enough to wander. I default to three, then escalate with the full trace. The number is a harness limit, not a suggestion in the prompt. Unbounded retries are how you spend real money confirming that the task, the tools, or the criteria were never going to work.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.