Spurlock Studios
Contact
Tag

evals

This is the evals tag archive on Will's Journal: every published post that shares this tag, listed in one place.

It is a collection page, not a topic essay — scan the cluster here instead of filtering the full journal index.

Which posts are tagged evals?

A small stack of coins. Thesis: PASS RATE LIES REVISION RATE. AI Agents

Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage

Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.

25 MIN
Nested brass frames. Thesis: LONG TAKE BUILD AI WORKFLOW. Automation

How long does it take to build AI workflow automation

AI workflow automation is a planning calendar: prompt eval, schema, and a human gate. Days-to-weeks bands are plans, not bids — a green model node is a demo.

28 MIN
A cracked amber fuse. Thesis: BROKE TRIED EVALUATE AI AGENT. AI Agents

What broke when I tried to evaluate an AI agent in production

Production eval failed on four fronts: no gold labels, biased samples, judge drift, eval tools that wrote. Isolate the harness before you trust the live score.

29 MIN
A violet ring. Thesis: BUILD EVALUATOR BEFORE AGENT. AI Agents

Build the Evaluator Before the Agent

If the judge shares the worker's context, you are grading your own homework. Build criteria, evidence, and ceilings first — then grant real tool autonomy.

22 MIN
A small stack of coins. Thesis: KNOW IF AI AGENT ACTUALLY. AI Agents

How do I know if my AI agent is actually working

You know an AI agent is working when rewrite rate, policy denials, cost, and time-to-done hold. Thumbs and CSAT hide unpaid human editors on real writes.

25 MIN
A small stack of coins. Thesis: TEST AI AGENTS BEFORE THEY. AI Agents

How do I test AI agents before they ship

Test AI agents before they ship on a golden set, staged tools, and shadow mode. Pass rate is not the gate — score cost, escalate, and known-bad fails too.

26 MIN
An expired brass key. Thesis: MANAGE MULTIPLE AI AGENTS PRODUCTION. AI Agents

How do I manage multiple AI agents in production

Isolate each production agent: credentials, write surface, named owner, and evals. Handoff with typed contracts. Do not share one god-agent across jobs.

27 MIN
Nested brass frames. Thesis: DID QUALITY DROP WE DIDN. AI Agents

Why did quality drop when we didn’t change our prompts

Quality dropped because a pin, tool schema, retrieval corpus, eval set, or traffic mix moved — not because you edited prompts. Isolate, then pin versions.

29 MIN
A cracked amber fuse. Thesis: BIG SHOULD GOLDEN SET BEFORE. AI Agents

How big should my golden set be before soft-launch

Before soft-launch, size the golden set by coverage — happy, auth fail, empty tool, policy deny, money, ambiguous — not a magic N or vanity pass rate.

30 MIN
A folded lab sheet with no readable lines. Thesis: GOLDEN SETS PRODUCTION FAILURES TURN. AI Agents

Golden Sets from Production Failures: Turn Bad Runs into Regression Fuel

Turn a bad production agent run into a regression test: harvest the trace, stub the tools, write the expected terminal, and gate deploys on that fixture.

18 MIN
A violet ring. Thesis: PIN MODEL GATE UPGRADE CATCH. AI Agents

Pin the Model, Gate the Upgrade: Catch Agent Drift Before Customers Do

Yes—pin production agents to explicit model IDs. Floating aliases change behavior with no deploy. Upgrade only through a golden-set gate you actually run.

20 MIN
A violet scale. Thesis: LLM AS JUDGE RELIABILITY CALIBRATE. AI Agents

LLM-as-Judge Reliability: Calibrate the Scorer Before You Trust the Score

An LLM judge is a noisy instrument, not ground truth. Calibrate against human labels, kill position and verbosity bias, re-check after model or rubric changes.

19 MIN
FAQ

What should you know about this tag archive?

What is this tag archive?

This is the evals tag archive on Will's Journal: every published post that shares this tag, listed in one place.

Is this a guide to the topic, or a list of posts?

A list of posts. This page is a collection, not a topic essay — the 12 posts below are the evals cluster on Will's Journal.