evals
This is the evals tag archive on Will's Journal: every published post that shares this tag, listed in one place.
It is a collection page, not a topic essay — scan the cluster here instead of filtering the full journal index.
Which posts are tagged evals?
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
Automation How long does it take to build AI workflow automation
AI workflow automation is a planning calendar: prompt eval, schema, and a human gate. Days-to-weeks bands are plans, not bids — a green model node is a demo.
AI Agents What broke when I tried to evaluate an AI agent in production
Production eval failed on four fronts: no gold labels, biased samples, judge drift, eval tools that wrote. Isolate the harness before you trust the live score.
AI Agents Build the Evaluator Before the Agent
If the judge shares the worker's context, you are grading your own homework. Build criteria, evidence, and ceilings first — then grant real tool autonomy.
AI Agents How do I know if my AI agent is actually working
You know an AI agent is working when rewrite rate, policy denials, cost, and time-to-done hold. Thumbs and CSAT hide unpaid human editors on real writes.
AI Agents How do I test AI agents before they ship
Test AI agents before they ship on a golden set, staged tools, and shadow mode. Pass rate is not the gate — score cost, escalate, and known-bad fails too.
AI Agents How do I manage multiple AI agents in production
Isolate each production agent: credentials, write surface, named owner, and evals. Handoff with typed contracts. Do not share one god-agent across jobs.
AI Agents Why did quality drop when we didn’t change our prompts
Quality dropped because a pin, tool schema, retrieval corpus, eval set, or traffic mix moved — not because you edited prompts. Isolate, then pin versions.
AI Agents How big should my golden set be before soft-launch
Before soft-launch, size the golden set by coverage — happy, auth fail, empty tool, policy deny, money, ambiguous — not a magic N or vanity pass rate.
AI Agents Golden Sets from Production Failures: Turn Bad Runs into Regression Fuel
Turn a bad production agent run into a regression test: harvest the trace, stub the tools, write the expected terminal, and gate deploys on that fixture.
AI Agents Pin the Model, Gate the Upgrade: Catch Agent Drift Before Customers Do
Yes—pin production agents to explicit model IDs. Floating aliases change behavior with no deploy. Upgrade only through a golden-set gate you actually run.
AI Agents LLM-as-Judge Reliability: Calibrate the Scorer Before You Trust the Score
An LLM judge is a noisy instrument, not ground truth. Calibrate against human labels, kill position and verbosity bias, re-check after model or rubric changes.
What should you know about this tag archive?
What is this tag archive?
This is the evals tag archive on Will's Journal: every published post that shares this tag, listed in one place.
Is this a guide to the topic, or a list of posts?
A list of posts. This page is a collection, not a topic essay — the 12 posts below are the evals cluster on Will's Journal.