LLM-as-Judge Reliability: Calibrate the Scorer Before You Trust the Score
An LLM judge is a noisy instrument, not ground truth. Calibrate against human labels, kill position and verbosity bias, re-check after model or rubric changes.
William Spurlock Founder — Spurlock Studios Updated 19 MIN
You cannot trust an LLM-as-judge out of the box — it will flatter bad agent runs if you never calibrate it. Treat the judge as a noisy instrument: measure agreement with humans, kill known biases, and re-validate after every model or rubric change. A high judge score with unmeasured calibration is a vanity metric with better fonts.
This is meta-eval. How to build the evaluator stack lives in evaluators before agents. This spoke asks whether the scorer itself is lying. Parent context: Agentic Systems Operating Manual. The cousin that owns the metric panel — pass rate, revision rate, coverage — is why pass rate lies.
The short answer
- Mechanical checks first; judges only where humans would disagree on soft criteria.
- Calibrate against a labeled set before the judge gates CI or production.
- Agent trajectories break naive “grade the final answer” judges — score tools and intermediate claims too.
- Watch position bias, verbosity bias, and same-model circular scoring.
- Re-run calibration when the worker model, judge model, or rubric changes — judge drift is real.
What LLM-as-judge means in an agent harness
In an agent eval harness, the judge is a second model call (or panel) that scores a run against a rubric and returns structured verdicts: pass/fail, criterion codes, short evidence quotes.
It is not:
- A replacement for schema validation
- Ground truth by virtue of being “smarter”
- Safe to share weights casually with the worker without measuring circularity
Typical placement:
- Worker run completes (or hits a checkpoint)
- Mechanical checks run (schema, allowlist, required artifacts)
- Judge scores remaining soft criteria
- Harness maps scores →
pass/revise/escalate/fail
OpenAI’s own agent-eval docs treat this split as the default: deterministic graders for structure, model graders for open-ended quality (trace grading, agent evals). If step 2 is empty, you are paying a judge to do string checks. Stop.
| Role | Owns | Must not own |
|---|---|---|
| Mechanical grader | Schema, allowlists, budgets, ids | Tone, completeness vs a messy brief |
| LLM judge | Soft criteria with named evidence | Binary facts already on the ledger |
| Human | Irreversible tools until calibration holds | Spot-checking only the easy passes |
| Worker | The artifact and the tool calls | The pass/fail that ships the run |
Capability language beats fashion: pick a judge that follows rubrics and returns structured JSON reliably. Pin the id. Re-calibrate on change.
Why agent trajectories break naive judges
Chatbot judges often see: prompt, final answer, rubric. Agent runs add tools, multi-turn state, and side effects. Failure modes unique to agents:
| Failure | What a naive judge misses |
|---|---|
| Wrong tool, right-looking final text | Grades the essay, ignores the CRM write |
| Hallucinated tool success | Believes the worker’s narration over the tool ledger |
| Criterion satisfied mid-trace then undone | Scores the last message only |
| Policy near-miss | “Sounds careful” while an irreversible tool nearly fired |
RuVerBench (Peng et al., 2026) is the first public meta-eval aimed at this gap: 2,458 rubric-verification instances across deep research and agentic coding. Deep-research reports averaged 7.1K tokens; coding trajectories averaged 49.4K. Even frontier judges stayed strong-but-noisy. Weaker judges swung harder when the prompt changed. Batched verification traded accuracy for speed. Majority voting helped, then flattened.
JudgeBench makes the same point from the other side: many judges that look aligned on preference chats collapse toward chance on hard knowledge, math, and coding pairs. Human-preference agreement is not factual reliability.
Practical takeaway: feed the judge a structured trace digest, not a chat dump — tool names, redacted args/results, state transitions, and the final artifacts.
Do not invent a single published “accuracy %” for all judges. Your calibration numbers are the only ones that count for your rubric.
How to calibrate against human labels
Procedure that actually moves reliability:
- Sample 50–100 runs covering pass, fail, and escalate (stratify by
job_typeand difficulty). LangChain’s agent-eval note starts teams at 20+ labeled examples and grows from there; humans review on the order of 50–100 traces per hour, so the bottleneck is attention, not tooling (agent evals). - Blind-label with humans using the same rubric the judge will see. Two raters when stakes are high; resolve disagreements explicitly and keep the disagreement codes.
- Run the judge on the same packages; store scores, evidence quotes, and the judge model pin.
- Compute agreement: per-criterion accuracy / F1, overall pass-fail agreement, and confusion pairs (judge pass / human fail is the dangerous cell).
- Tune rubric wording, evidence requirements, and which criteria stay mechanical.
- Freeze a calibration report with judge model id, rubric version, and date.
| Metric | Why it matters |
|---|---|
| Human–judge pass/fail agreement | CI gate sanity |
| False pass rate (judge pass, human fail) | Customer risk |
| False fail rate | Cost / latency from over-refusal |
| Per-criterion agreement | Finds broken rubric lines |
| Chance-corrected agreement (κ) | Stops you celebrating easy-set accuracy |
Raw agreement flatters. A 2026 large-scale audit of 21 judges across MT-Bench, JudgeBench, and RewardBench found exact-match agreement overstated chance-corrected discrimination by 33–41 percentage points on MT-Bench (Reliability without Validity). You do not need their exact delta. You need a metric that does not treat “both said pass on an obvious pass” as proof the judge can see a hard fail.
Hedge, not folklore: “good enough” for a soft-launch gate is often in the ballpark of strong majority agreement on pass/fail for your risk class — but you set the threshold from blast radius, not from a blog’s lucky number. The 2023 MT-Bench paper reported a strong judge matching human prefs in the same band as human–human agreement on their chat tasks, after bias mitigations (Zheng et al.). That is a historical receipt, not a CI threshold. Irreversible tools demand tighter false-pass bounds than draft-only jobs.
| Calibration artifact | Store it |
|---|---|
rubric_id + semver | On every judge span |
| Judge model pin / provider snapshot | Next to the score |
| Gold-slice ids | Never used for prompt tuning |
| Confusion matrix by criterion | Weekly drift check |
| Human override codes | Next calibration set |
If you cannot point at that report, you do not have a calibrated judge. You have a prompt.
What if the humans disagree with each other?
A judge cannot outrun its labelers. If two humans split on the same package, “judge–human agreement” is a coin flip dressed as a metric.
Do this before you blame the model:
- Sit both raters on the same five disagreements with the rubric on the table.
- Split each fight into rubric hole (the line is vague) vs attention miss (the evidence was there).
- Rewrite the hole as a mechanical check or add a positive and negative exemplar.
- Keep a third label,
human_split, on packages you will not force into pass/fail. - Report judge agreement against the adjudicated label, and separately against each rater.
| Human pattern | What it means for the judge |
|---|---|
| High rater agreement, judge misses | Judge or package problem — fix those |
| Low rater agreement, judge matches one | Rubric problem — do not tune the judge to a coin flip |
| Both raters pass, later ops reject | Online criterion missing from the rubric |
| Both raters fail, judge passes | The dangerous cell — raise this before any CI gate |
Zheng et al. treated human–human agreement as the ceiling on their chat tasks (MT-Bench). Copy the method, not their percentage. Your ceiling is the adjudicated agreement on your hard stratum. If that ceiling is 70% because the rubric is mush, a 90% judge score means the judge learned one rater’s taste.
Irreversible tools do not get a mushy rubric. If two trained humans cannot agree after seeing the ledger, the job is not ready for a judge gate. Workshop the criterion. Then calibrate.
How do position and verbosity bias show up in agent traces?
These show up differently than in pairwise chatbot evals.
Position bias: When the judge sees multiple candidate revisions or tool results in a list, earlier or later items can win unfairly. Shi et al. measured this across pairwise and list-wise settings (15 judges, 22 tasks, 150,000+ instances): position bias is not random, it varies by judge and task, and it gets worse when the quality gap between candidates is small. Shuffle or score candidates independently when you compare revisions.
Verbosity bias: Long, confident worker narrations score higher than terse correct tool use. Zheng et al. documented the “repetitive list” attack on chat judges — padded answers winning over shorter correct ones (MT-Bench). On agent traces the costume is a three-paragraph “I verified in the CRM” speech with no matching ledger span.
Countermeasures:
- Require evidence quotes tied to tool ledger ids, not vibes
- Cap narrative length in the judge package
- Score “correctness of actions” separately from “quality of prose”
- Penalize unsupported claims explicitly in the rubric
| Bias | Symptom in traces | Mitigation |
|---|---|---|
| Position | Revision A always wins when listed first | Independent scoring / shuffle |
| Verbosity | Wordy fails beat short passes | Evidence-first rubric |
| Authority | Judge trusts “I verified via CRM” without tool span | Ledger required |
| Leniency | Soft criteria always “mostly met” | Binary criteria + examples |
| Self-preference | Same-family worker prose scores high | Different-family judge, or measure it |
If your judge prefers essays, your agent will learn to write essays instead of calling tools correctly.
The 2026 Reliability-without-Validity audit reported much smaller verbosity effects under a single pairwise rubric than the 2023 literature. Do not take that as a hall pass. Your agent traces are not MT-Bench pairs. Measure verbosity on your packages: correlate judge score with token count on the hard stratum. If the correlation is the story, the rubric is still grading prose.
Same model as worker — ever OK?
Sometimes, for low-stakes draft scoring in staging. Rarely for production gates on irreversible work.
Risks:
- Shared blind spots (both miss the same policy hole)
- Style favoritism (worker prose matches judge priors)
- Correlated drift when the provider updates the family
Self-preference is measured, not vibes. Wataoka, Takahashi, and Ri found judges scoring lower-perplexity text higher than humans do — including text they were not told was their own. Play Favorites then isolated self-bias and family-bias: some judges systematically score their own completions, and other models in the same family, above an independent reference. G-Eval’s authors flagged the same tilt toward model-generated text when they introduced CoT + form-filling judges (Liu et al.).
| Setup | Use when |
|---|---|
| Same model family, same pin | Cheap staging smoke only |
| Same family, different pin / size | Acceptable if calibrated; still watch circularity |
| Different vendor for judge | Prefer for production gates when cost allows |
| Panel (2 judges + tie-break rule) | High blast radius criteria, after single-judge work fails |
A panel is not a personality. Independent scores, a predefined tie-break, and a recorded disagreement rate. Skip multi-agent debate theater as the SMB default — RuVerBench already saw majority voting flatten. If one calibrated judge plus mechanical checks is inside your false-pass band, stop adding voters.
How do you detect judge drift?
Judge drift is a silent production bug: worker prompts unchanged, online “pass rate” climbs or collapses, humans still rewrite.
Triggers that force a re-calibration run:
- Judge model pin or provider snapshot changed
- Rubric version bumped (even “clarifications”)
- Worker model upgraded (distribution of traces changes)
- New tool or side-effect class added
- Human override rate diverges from judge pass rate for two weeks
- Judge latency, JSON-schema fail rate, or empty-evidence rate jumps
Drift checks:
- Hold out a frozen gold slice (never used for prompt tuning).
- Weekly or on deploy: score the slice; alert if agreement or false-pass rate moves past your band.
- Sample online disagreements (human reject after judge pass) into the next calibration set.
- Keep the judge span on the same run id as the tool calls — same weekly ritual as the rest of the control plane.
| Signal | Read as |
|---|---|
| Gold-slice κ holds, online rewrite rate up | Worker or input mix moved; judge may still be fine |
| Gold-slice false-pass up, rewrite rate up | Judge got lenient, or rubric got vague |
| Gold-slice false-fail up, escalate rate up | Judge got harsh; check pin and prompt |
| Judge JSON invalid / missing evidence | Harness bug, not a quality win |
LangChain’s note is blunt: recalibrate regularly, because judges drift just like the agents they evaluate (agent evals). Anthropic’s multi-agent research writeup keeps humans in the loop for the same reason — people still catch hallucinations, system failures, and source misses that automation grades past (engineering note).
Judge spans belong on the trace beside tool calls. If you cannot query “show me last week’s judge-pass / human-fail,” you will learn about drift from a customer.
Which checks should never be a judge call?
Move these out of the LLM judge entirely:
| Check | Why mechanical |
|---|---|
| JSON / schema validity | Binary, cheap |
| Required fields present | Binary |
| Tool allowlist / deny list | Policy, not taste |
| Max turns / budget exceeded | Harness facts |
| Forbidden strings / PII patterns | Regex or classifiers |
| Idempotency key present on writes | Ledger fact |
| Tool result exists for every cited claim id | Ledger join |
| Terminal reason is a known enum | Harness fact |
Judges earn their tokens on: tone, completeness vs a messy brief, “did the research address the question,” soft brand constraints. If a criterion can be a unit test, make it a unit test.
A useful split when you are arguing with yourself:
- Write the criterion as a sentence.
- Name the observable (field, tool span, quote).
- If a junior engineer could assert it in ten lines of code, it is not a judge criterion.
- If two humans would still disagree after seeing the same evidence, it may be a judge criterion — and it needs exemplars.
OpenAI’s grader docs make the same cut: string and code graders for deterministic properties, model graders for the rest, and an explicit warning about grader hacking — the policy scores high on the grader and low with experts (graders). That is judge unreliability with an optimization loop attached. If the worker can see the rubric and is trained or prompted against it, your next calibration set must include the new failure costume, not last month’s.
Failure mode: correlated easy-case accuracy
What breaks: your calibration set is 80% obvious passes. Judge–human agreement looks excellent. Production is the hard 20%. The judge rubber-stamps fluent wrongness.
What it costs: CI stays green while revision rate stays ugly — the pass-rate lie with a judge costume. Why pass rate lies owns that panel. This post owns the scorer that inflated the first number.
What you do instead:
- Stratify the labeled set by difficulty and failure code.
- Track agreement on the hard stratum separately.
- Keep a rising share of production disagreements in the set.
- Never celebrate aggregate agreement alone.
- Include near-miss policy cases, hallucinated tool success, and undone mid-trace criteria — the rows naive judges miss.
| Stratum | If you skip it | What the dashboard hides |
|---|---|---|
| Obvious pass | Nothing useful | Inflated κ |
| Obvious fail | Judge looks strict | Misses leniency on the real risk |
| Hard / ambiguous | The job you hired the judge for | Fluent wrongness ships |
| New tool class | Yesterday’s calibration | Silent false passes |
Easy cases are where judges look smart. Hard cases are why you hired them.
JudgeBench exists because preference-aligned judges can still lose on objective pairs. RuVerBench exists because long agentic traces make that worse. Your gold slice should look more like those than like a demo reel.
What rubric design survives agent traces?
Rules of thumb for judge-ready rubrics:
- One criterion, one failure code.
- Each criterion names observable evidence (artifact field, tool result, quote).
- Include 2–3 positive and negative exemplars per soft criterion.
- Separate “process” criteria (allowed tools, no speculative writes) from “outcome” criteria (user gets value).
- Version the rubric (
rubric_id, semver). Store it on every judge span. - Ban “mostly” and 1–5 vibes. Binary or a three-way
pass/revise/failwith a named reason.
Bad criterion: “Be helpful and accurate.”
Better: “Every numeric claim in the customer email appears in tool:billing.get result or is marked uncertain.”
| Rubric smell | What the judge will do | Fix |
|---|---|---|
| One “quality” score | Un-actionable leniency | Split into codes ops can fix |
| No evidence slot | Grades vibes | Require ledger-linked quotes |
| Process mixed with outcome | Passes a polite wrong write | Two criteria, two codes |
| Unversioned prose in a prompt | Silent drift | rubric_id on the span |
| Exemplars all long essays | Verbosity bias | Include short correct traces |
G-Eval’s useful inheritance is the form, not the brand: decompose the criterion, force a structured fill, then score (Liu et al.). The failure mode they already flagged — bias toward model-shaped text — is why the form still needs human labels on your traces.
What belongs in the judge evidence package?
If you dump the raw chat, you are grading creative writing. If you dump the whole RAG corpus, you are grading the judge’s patience.
Minimum package:
| Field | Why |
|---|---|
| Job goal + acceptance lines | The question the run was supposed to answer |
rubric_id + criterion list | Same text the humans used |
| Final artifacts | What would ship |
| Tool ledger digest | Names, redacted args, result ids, timestamps |
| Harness terminal reason | pass / revise / escalate / budget |
| Candidate list (if any) | Independently scored, position shuffled |
Leave out:
- Worker chain-of-thought, unless you A/B’d it on the labeled set and false-pass improved
- Secrets, raw PII, full document dumps
- Prior judge rationales (the next judge will agree with the last one)
- Unrelated memory / RAG hits the worker never used
Default no on CoT-in-the-package. It invites style grading and leaked rationalizations. Prefer actions + artifacts. If you experiment, keep it only when the hard-stratum false-pass rate moves the right way.
Cap the digest. RuVerBench’s coding traces averaged tens of thousands of tokens; that is the regime where even frontier judges get noisy. A one-page ledger with ids the judge must quote will beat a 50K-token dump you never measured.
How do you gate CI without false comfort?
Suggested promotion ladder:
| Gate | Judge role |
|---|---|
| PR / prompt change | Mechanical + judge on golden set; block on false-pass regressions vs baseline |
| Staging soak | Online sample; compare human spot-checks |
| Soft-launch | Judge advisory or dual-run; humans still own irreversible tools |
| Autonomy expand | Judge gate only after calibration report signed off |
Agreement thresholds are a product decision. Document them next to blast radius. Do not copy a research paper’s headline number into your runbook without re-measuring on your traces.
| You may block a PR when… | You may not |
|---|---|
| Gold-slice false-pass exceeds band | Aggregate pass rate ticked up |
| A new criterion has no human labels | The flagship model is “the judge now” |
| Judge JSON / evidence fail rate spikes | One anecdotal “it looked better” |
| Hard-stratum κ drops | Easy-set agreement holds |
OpenAI’s grader-hacking warning is the CI version of this: a worker that learns the scorer will green the eval and still fail experts (graders). Pair the judge gate with why pass rate lies — revision rate, trajectory, coverage, cost per success — so a lenient scorer cannot be the only veto.
If the judge is advisory, say so in the dashboard. An advisory score drawn as a ship/no-ship toggle is how teams get surprised.
What this post does not replace
| Spoke | Owns |
|---|---|
| Evaluators before agents | Build order: criteria → mechanical → judge → online |
| Why pass rate lies | The metric panel a lying scorer can inflate |
| Agentic Systems Operating Manual | The rest of the stack around the scorer |
| This post | Meta-eval: is the judge calibrated and stable? |
If you skip the first two and only add a judge prompt, you have cosplay.
This spoke stays on the instrument. It does not tell you which jobs deserve an agent, how to sandbox tools, or how to price a fleet. Those are other pages. The only question here is: when the scorer says pass, should you believe it?
Pilot slice: calibration in five days
A Spurlock Studios agentic pilot can include a thin meta-eval pass when the job already has soft criteria. You will not finish academic-grade inter-annotator studies in five days. You will know whether the judge is roughly usable or actively dangerous.
| Day | Judge work |
|---|---|
| 1 | Split mechanical vs judge criteria; write the evidence-package schema |
| 2 | Label 30–50 runs (or dense fixtures), stratified, two raters on the hard slice |
| 3 | First judge pass + confusion matrix; compute false-pass on the hard stratum |
| 4 | Rubric surgery; kill verbosity and position loopholes; drop criteria that should be code |
| 5 | Freeze rubric_id + pin; wire judge span to traces; hold out the gold slice |
What “done” means on day five:
- Mechanical checks no longer go through the judge
- A written false-pass bound sits next to blast radius
- Gold-slice ids are frozen and not used for prompt fiddling
- Judge spans queryable by run id
- A named owner for the next re-calibration trigger
That is the same discipline I use across 500+ automations and 20,000+ hours on agentic systems: the scorer is a component with a version, not a vibe in a system prompt. Book via /agentic.
Worked example: support draft agent
| Criterion | Judge or mechanical? |
|---|---|
| Contains order id from ticket | Mechanical |
| No refund promise unless tool says eligible | Mechanical on tool + regex |
| Tone matches brand examples | Judge |
| Answers all explicit customer questions | Judge with checklist from ticket |
Illustrative pattern (not a universal stat): judge passes “tone” on long drafts; humans fail short correct ones. Fix: verbosity penalty + exemplar shorts; require factual claims to cite tool:orders.get. That bias fix often beats swapping judge vendors.
Evidence package minimum: job goal, rubric version, final artifacts, tool ledger digest, harness terminal reason. No secrets, no giant RAG dumps. No ledger → you are grading creative writing.
Same-family judge on this job: fine for staging drafts after the day-3 matrix. Not fine as the only gate once the draft can trigger a refund tool.
Panel: only when blast radius is high and single-judge false-pass stays above band after rubric work. Independent scores, predefined tie-break.
Anti-patterns that keep this example lying:
“The flagship model is the judge, so we’re fine.” Capability helps; calibration decides. JudgeBench already showed strong judges losing on hard pairs.
Judge sees full chain-of-thought and grades style. Prefer actions + artifacts.
One giant “quality 1–5” score. Un-actionable. Prefer criterion codes ops can fix.
Recalibrating never. Then your dashboard is a fiction that ages.
Gold slice is last week’s happy paths. Then you re-discovered easy-case accuracy.
FAQ
When should mechanical checks replace a judge entirely?
Whenever the criterion is binary and observable without taste: schemas, allowlists, budgets, required ids, forbidden actions. Judges are for soft criteria. If your entire rubric is mechanical, delete the judge call and celebrate the latency win.
Should the judge see the worker’s chain of thought?
Default no for scoring. CoT invites style grading and leaked rationalizations. Prefer tool ledgers and artifacts. If you experiment with CoT-in-the-judge-package, A/B it on your labeled set — keep it only if false-pass rate improves on the hard stratum.
Same model as worker — ever OK?
For low-stakes staging or draft-only jobs after calibration, sometimes. For production gates on irreversible tools, prefer a different model family or a panel, and always measure circular agreement on hard cases. Self-preference and family-bias are documented; do not assume your pin is the exception.
What agreement rate with humans is “good enough” to gate CI?
Whatever bound matches your blast radius — documented, measured on a stratified labeled set, with special attention to false passes. There is no universal published percentage that absolves you from measuring on your rubric and traces.
How do position and verbosity bias show up in agent traces?
Position bias skews revision tournaments; verbosity bias rewards long narrations over correct short tool use. Mitigate with independent scoring, evidence-first rubrics, and ledger-linked claims. Measure both on your packages, not from a chat-benchmark abstract.
How does this relate to evaluators-before-agents without replacing it?
Evaluators before agents tells you to build the eval stack before autonomy. This post assumes that stack exists and asks whether the LLM judge component is calibrated. You need both: a real evaluator, and a scorer you have meta-evaluated.
CTA
Calibrate the scorer before you trust the score.
What questions does this article answer?
- When should mechanical checks replace a judge entirely?
- Whenever the criterion is binary and observable without taste: schemas, allowlists, budgets, required ids, forbidden actions. Judges are for soft criteria. If your entire rubric is mechanical, delete the judge call and celebrate the latency win.
- Should the judge see the worker’s chain of thought?
- Default no for scoring. CoT invites style grading and leaked rationalizations. Prefer tool ledgers and artifacts. If you experiment with CoT-in-the-judge-package, A/B it on your labeled set — keep it only if false-pass rate improves on the hard stratum.
- Same model as worker — ever OK?
- For low-stakes staging or draft-only jobs after calibration, sometimes. For production gates on irreversible tools, prefer a different model family or a panel, and always measure circular agreement on hard cases. Self-preference and family-bias are documented; do not assume your pin is the exception.
- What agreement rate with humans is “good enough” to gate CI?
- Whatever bound matches your blast radius — documented, measured on a stratified labeled set, with special attention to false passes. There is no universal published percentage that absolves you from measuring on your rubric and traces.
- How do position and verbosity bias show up in agent traces?
- Position bias skews revision tournaments; verbosity bias rewards long narrations over correct short tool use. Mitigate with independent scoring, evidence-first rubrics, and ledger-linked claims. Measure both on your packages, not from a chat-benchmark abstract.
- How does this relate to evaluators-before-agents without replacing it?
- [Evaluators before agents](/blog/evaluators-before-agents) tells you to build the eval stack before autonomy. This post assumes that stack exists and asks whether the LLM judge component is calibrated. You need both: a real evaluator, and a scorer you have meta-evaluated.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.