llm-as-judge
This is the llm-as-judge tag archive on Will's Journal: every published post that shares this tag, listed in one place.
It is a collection page, not a topic essay — scan the cluster here instead of filtering the full journal index.
Which posts are tagged llm-as-judge?
AI Agents What broke when I tried to evaluate an AI agent in production
Production eval failed on four fronts: no gold labels, biased samples, judge drift, eval tools that wrote. Isolate the harness before you trust the live score.
AI Agents LLM-as-Judge Reliability: Calibrate the Scorer Before You Trust the Score
An LLM judge is a noisy instrument, not ground truth. Calibrate against human labels, kill position and verbosity bias, re-check after model or rubric changes.
What should you know about this tag archive?
What is this tag archive?
This is the llm-as-judge tag archive on Will's Journal: every published post that shares this tag, listed in one place.
Is this a guide to the topic, or a list of posts?
A list of posts. This page is a collection, not a topic essay — the 2 posts below are the llm-as-judge cluster on Will's Journal.