golden set
This is the golden set tag archive on Will's Journal: every published post that shares this tag, listed in one place.
It is a collection page, not a topic essay — scan the cluster here instead of filtering the full journal index.
Which posts are tagged golden set?
AI Agents What broke when I tried to evaluate an AI agent in production
Production eval failed on four fronts: no gold labels, biased samples, judge drift, eval tools that wrote. Isolate the harness before you trust the live score.
AI Agents How do I test AI agents before they ship
Test AI agents before they ship on a golden set, staged tools, and shadow mode. Pass rate is not the gate — score cost, escalate, and known-bad fails too.
AI Agents How big should my golden set be before soft-launch
Before soft-launch, size the golden set by coverage — happy, auth fail, empty tool, policy deny, money, ambiguous — not a magic N or vanity pass rate.
AI Agents Golden Sets from Production Failures: Turn Bad Runs into Regression Fuel
Turn a bad production agent run into a regression test: harvest the trace, stub the tools, write the expected terminal, and gate deploys on that fixture.
What should you know about this tag archive?
What is this tag archive?
This is the golden set tag archive on Will's Journal: every published post that shares this tag, listed in one place.
Is this a guide to the topic, or a list of posts?
A list of posts. This page is a collection, not a topic essay — the 4 posts below are the golden set cluster on Will's Journal.