How do I test AI agents before they ship
Test AI agents before they ship on a golden set, staged tools, and shadow mode. Pass rate is not the gate — score cost, escalate, and known-bad fails too.
William Spurlock Founder — Spurlock Studios 26 MIN
You test AI agents before they ship with three layers that can fail the release: a frozen golden set, tools that run in staging (stubs that can error), and shadow mode that never writes. A single pass rate on a happy-path demo is not a ship gate. Score cost per passing job, escalate rate, and whether the known-bad fixtures still fail — then compare shadow drafts to the live champion.
This spoke sits under the Agentic Systems Operating Manual. The independent judge — criteria, evidence, revision ceiling — lives in evaluators before agents. This page owns the stack that judge runs against before anyone gets a write tool. If the path is already a flowchart, stop and read when not to build an agent. How Spurlock Studios installs the first harness is on /agentic.
The short answer
- Freeze a golden set of real jobs, including traps and never-auto-pass fixtures, before you argue about a percentage.
- Run that set against staged tools that can empty, 429, timeout, and deny — not against a costume of production that always returns 200.
- Put the candidate in shadow mode on live traffic with write tools off. Champion serves. Challenger drafts. Harvest the diffs.
- Do not ship on pass rate alone. Require cost per pass, escalate rate, known-bad still fail, and process-plus-outcome grades.
- A workflow with one schema-checked model step is the default when you can still draw the graph.
What does “tested before they ship” mean?
Tested means a named candidate lost the right to write until three independent layers stayed green for a window you and the buyer signed. It does not mean the demo chat looked fluent. It does not mean a vendor dashboard showed a round number.
Anthropic’s Demystifying evals for AI agents (Jan 2026) splits the object you grade into a task, a grader, a transcript, and an outcome. The transcript is what the agent did. The outcome is the state of the world. “Ticket booked!” in the last sentence is not a booking. OpenAI’s agent evals make the same cut: the unit is the trace — model calls, tool calls, guardrails, handoffs — not a completion.
| Claim | What you actually ran | Ship? |
|---|---|---|
| “The demo was clean” | Happy path, live tools, one operator | No |
| “Pass rate is 94%” | Binary on the current set, tools unspecified | Not by itself |
| “We have evals” | A model asked itself if it did a good job | No |
| Golden set + staged tools | Frozen fixtures, stubs that can fail | Offline gate only |
| + shadow, writes off | Champion serves; challenger drafts on live traffic | Pre-ship, still no writes |
| All three + scoreboard | Pass, cost, escalate, known-bad, process and outcome | Then you may widen writes |
- Candidate has a version string (
agent_version,eval_set_version,evaluator_version) - Write tools default-deny on that candidate
- Golden set is frozen for the run, not edited mid-demo
- Staged tools can return empty, 429, timeout, auth fail, and policy deny
- Shadow path cannot SMTP, refund, deploy, or write CRM
- A human can paste a
run_idfrom an offline fail and from a shadow fail
If any box is empty, you ran a rehearsal. Rehearsals do not earn write tools.
Why is a pass rate not a ship gate?
A pass rate is one average over one set under one tool fiction. Agents are not unit tests. The same change that lifts the average can hide a money-path miss, a cost blow-up, or a known-bad fixture that started passing because you loosened the rubric.
Anthropic’s evals piece is explicit about the shape of the number. pass@k is the chance of at least one success in k trials. pass^k is the chance that all k trials succeed. At k=1 they are the same. By k=10 they tell opposite stories: pass@k drifts toward 100% while pass^k falls. Their illustration: a 75% per-trial success rate run three times is about 42% pass^3. Customer-facing agents live on consistency. A coding assistant you will retry can live on pass@k. Do not publish the generous one and ship the job that needed the strict one.
Microsoft Foundry’s agent evaluators split the same physics in vendor language: system evaluation (did the job finish with a usable deliverable) versus process evaluation (did the steps stay inside the contract). A 90% “task complete” with a forbidden tool in the trace is not a pass. It is a lucky outcome.
| Number | What it hides | What to pair it with |
|---|---|---|
| Overall pass % | Mixes happy path with traps | Pass by stratum; never-auto-pass must still fail |
| pass@k | One lucky trial | pass^k, or pass@1 if the product is one-shot |
| Task-complete % | Illegal path that still “finished” | Process grade on the trace |
| “Quality” 0–100 | Which contract broke | Binary per criterion |
| Latency | Fast wrong | Cost per passing job |
OpenAI’s evaluation best practices list the anti-pattern I see after 500+ automations and 20,000+ hours on agentic systems: biased datasets that do not look like production, and “it seems like it’s working” as a ship criterion. A green tile on a toy set is a press release.
Pick the pass family before you run the suite, not after the screenshot looks kind.
- One-shot product (refund, public send, deploy): score pass@1 and pass^k on a small k you can afford. Lucky retries are not a feature.
- Human-in-the-loop tool (draft the human will edit): pass@k can be honest. Still cap k. Still fail known-bad on the first try.
- Report both by stratum. Never average money-path with FAQ-path.
- Freeze k in the release ticket. Changing k to rescue a number is the same sin as editing the set mid-run.
A pass rate can ship when — and only when — it is one cell on a scoreboard the buyer signed, the set is frozen, the tools are honest, and the known-bad rows still fail. Otherwise it is a mood.
What belongs in the golden set for a pre-ship run?
The golden set is the frozen list of jobs you will replay on every behavior change. Inputs, expected terminal, tool stubs, and labels. It is not a folder of “pretty prompts.” Anthropic’s start line is 20–50 simple tasks drawn from real failures, not hundreds of paraphrases. That is a start, not a law, and it is not a substitute for coverage of irreversible tools.
This page does not own how big the set should be. It owns what has to be true of the rows before you treat a run as a pre-ship gate.
| Row type | Why it is in the set | Pre-ship rule |
|---|---|---|
| Happy path | Proves the job can finish | Required, never the majority |
| Empty tool / no retrieval | Catches fluent invention | Must hard-fail or explicit no_match |
| Auth fail | Catches retry-until-write | Must stop, not grind |
| Policy deny | Proves the gate fires | Deny must appear on the trace |
| Money / irreversible | Blast radius | Allow and deny coverage |
| Ambiguous / hostile | Criteria mush vs real traffic | Escalate or refuse, not “best effort” |
| Never-auto-pass | Regression anchors | If any of these pass, the suite is lying |
Every row needs:
case_idthat will still mean something in six months.- Input as the system will actually see it (payload, not a Slack anecdote).
- Stubbed tool world: what each tool returns on this case.
- Expected terminal:
pass,escalate, orabort— plus the blocking criteria that must fire. - Labels: stratum,
job_type, irreversible tools in play,must_never_auto_pass.
- At least one never-auto-pass fixture the buyer would fire someone over
- At least one row where the correct move is not to write
- Stubs checked into the repo next to the cases, not “whoever has staging up”
- Set version bumped when a row is added, edited, or deleted
- Worker few-shots are not a paste of golden answers (leakage)
A row without an expected terminal is a demo script. Demo scripts do not fail merges.
Shape I freeze for a pre-ship case. The numbers inside are labels, not a published bar.
{
"case_id": "crm-note-empty-search-001",
"job_type": "crm_note",
"stratum": "empty_tool",
"must_never_auto_pass": true,
"input": { "ticket_id": "T-1842", "question": "What is their refund window?" },
"stubs": {
"search.kb": { "status": 200, "records": [] }
},
"expected_terminal": "escalate",
"blocking": ["no_match_or_citation", "no_smtp_on_empty"]
}
If search.kb is empty and the candidate still writes a confident note — or worse, sends it — that row fails. If you cannot write this object, you do not have a case. You have a story.
How do I stage tools so a green suite is not a costume?
Staging for an agent is a tool facade that can fail the way production fails. It is not a second URL with the live Stripe key and a hopeful TEST_ prefix. If every stub returns 200 and a well-formed JSON body, you measured fluency on a greased floor.
LangSmith’s evaluation concepts split offline (dataset + reference outputs) from online (live traces, usually no reference). Offline is only as honest as the world you replay. If the replay cannot 429, you will not see the grind until a customer does.
| Staging lie | What you thought you tested | What actually ships |
|---|---|---|
| Always-200 stubs | Happy schema | First empty CRM payload invents fields |
| Production keys in “staging” | Isolation | Double send, double charge |
| Recorded 200s only | Replay | No timeout, no 429, no deny |
| Shared sandbox, dirty state | Idempotency | Order 2 sees order 1’s write |
| MCP to a real mailbox | “Just this once” | You shipped SMTP |
Minimum stub catalog for a pre-ship run:
- Happy — schema-valid payload, ids the case expects.
- Empty — 200 with zero records, or explicit not-found.
- Malformed — 200 with missing required keys.
- Auth — 401 / 403.
- Throttle — 429 with or without
Retry-After. - Timeout — hang past the tool budget, then error.
- Deny — policy or IAM refusal the agent must not retry-around.
- Write-shaped tools in staging are dry-run: validate args, log, return a fake id, do not mutate
- Fake ids cannot be copied into production
- Timeouts and 429s exist as first-class fixtures, not as “we’ll see in prod”
- The same
case_idcan pin which stub variant runs - A human can diff staging vs production allowlists and explain every extra scope
Stand up the facade in this order. Do not invert it.
- Inventory every tool the candidate may call. Writes get a dry-run twin with the same JSON schema.
- Implement the seven stub variants above as named fixtures, not as “toggle chaos.”
- Pin variants per
case_id. A happy-path case must not randomly 429 or you will chase flakes. - Run the golden set once with all-happy stubs. That run is a wiring check, not a ship signal.
- Run it again with the hostile pins the cases declare. That run is the gate.
- Diff the staging allowlist against production. Extra production scopes are a finding.
If staging cannot deny, you do not have a policy test. You have a costume.
What is shadow mode for an agent — and when is it fake?
Shadow mode means live traffic hits the champion, a copy of the request hits the candidate, and only the champion may change the world. The candidate produces an artifact and a trace. You grade both. You do not show the candidate to the user. You do not let it send.
LangSmith’s offline vs online table is the map: offline has references; online usually does not. Shadow sits in the crack. You have live inputs (online) and you still have the champion’s artifact as a comparison, plus the golden-set criteria as a grader. You still lack a single “correct” CRM note for a brand-new ticket. That is fine. You are hunting disagreement, cost, and process breaks — not a fantasy of labeled production.
| Setup | User sees | Candidate writes? | What you learn |
|---|---|---|---|
| Offline golden set | Nothing | No | Regression on frozen jobs |
| Dogfood with writes | Internal users | Yes | Incidents with a smaller audience |
| Canary with writes | A slice of customers | Yes | Real blast radius |
| Honest shadow | Champion only | No | Diffs, cost, process fails on live shape |
| Fake shadow | Champion only | Yes, by accident | Duplicate emails, duplicate refunds |
Honest shadow rules:
- Candidate tool runner is a dry-run facade. Same schemas, no side effects.
- Request copy is async. Champion latency does not wait on the challenger.
- Pair
champion_run_idandchallenger_run_idon onerequest_id. - Grade process (forbidden tool, step budget, schema) in code. Grade residue with the independent evaluator.
- Disagreements become golden-set candidates, anonymized, with a stubbed world.
- SMTP, refund, deploy, delete, CRM write are physically absent from the challenger allowlist
- Shadow errors cannot page the customer
- Shadow cost is capped per hour; a grind stops the challenger, not the champion
- Window length is a signed number of days, not “until we got bored”
- A kill switch exists outside the model
Fake shadow is the candidate with production tools and a comment that says “don’t actually send.” Comments are not gates. OWASP LLM06 Excessive Agency is exactly this skip: extra permissions, extra autonomy, extra blast radius.
Sampling is a signed number, not “whatever the queue did.” I will not dress a planning band as science. Agree the slice with the buyer, then hold it.
| Traffic | What I actually do in a pilot | What I do not do |
|---|---|---|
| Low, every request | Shadow 100% while cost cap holds | Shadow 100% with no cap |
| Medium | Sample by job_type, overweight money paths | Uniform 1% that misses refunds |
| High | Sample + replay buffer so champion never waits | Inline challenger on the hot path |
| New job type | 100% of that type until a week of pairs exist | “We’ll catch it in the average” |
If the challenger can write, you did not shadow. You dual-wrote.
What else do I score besides pass rate?
Score the numbers that would make you take the write tool away. I put five on one screen in a pilot. Everything else is a drill-down.
| Number | Definition | Lie if you skip it |
|---|---|---|
| Blocking pass (by stratum) | Share of frozen cases that meet every blocking criterion | Happy path inflates the average |
| Known-bad still fail | Never-auto-pass fixtures remain fail | Rubric was loosened or leaked |
| Cost per passing job | Dollars to a terminal pass, by job_type | Cheap fails look “efficient” |
| Escalate rate | Share that hit the ceiling or a human | Hidden by retries the user never saw |
| Process fail rate | Forbidden tool, over budget, schema miss — even if the note looks fine | Outcome-only grading |
Pair those with two sampling numbers you cannot get from the golden set alone:
- Shadow disagreement rate — champion vs challenger on live inputs, after process checks.
- Silent-fail sample — a human marks champion (or later, the candidate) wrong after a “pass.” Expensive. Sample it. Do not skip money paths.
NIST’s AI 600-1 Generative AI Profile puts this in governance language: measure before you manage, and put output filters and human-review thresholds in front of downstream use. A scoreboard is that filter in a shape an engineer can ship. A single percentage is not.
- Each number has an owner and a “hold / shrink / widen” rule
- Cost is dollars, flagged if estimated, not a weekly token dump
- Escalate includes
evaluator_versionandrevision_count - Process fails are counted even when the customer never saw a write
- No cell is “N/A because the vendor dashboard was blank”
Worked hold, with bars you and the buyer wrote — not a vendor default, not a number I invented for this page.
| Cell | Signed bar (example shape) | Hold | Widen writes |
|---|---|---|---|
| Blocking pass, happy | ≥ the bar you set | In band | In band and traps in band |
| Blocking pass, money | Own bar, usually stricter | Any miss | Zero blocking misses on money |
| Known-bad | 100% still fail | One flake quarantined | No flakes |
| Cost per pass | Band by job_type | At ceiling | Inside band for the window |
| Escalate | Expected band | Spike on one type | In band, humans not drowning |
| Process fail | Near zero on writes | Soft jobs only | Near zero on irreversible intent |
The numbers in the “bar” column are yours. Copying another team’s 91% is how you ship a costume with a nicer slide.
If the only chart you can screenshot is overall pass %, you are not ready to talk about shipping. You are ready to talk about adding columns.
How do I fail a ship even when the golden set is green?
The suite can be green and the candidate still stays denied. That is the point of a gate with more than one lock. Write the fail conditions down before the first run so nobody “interprets” a 91% on Friday.
| Green-looking fact | Still fail the ship if | Why |
|---|---|---|
| Overall pass ≥ bar | Any never-auto-pass fixture passed | The suite is now a cheerleader |
| Overall pass ≥ bar | Money/PII stratum below its own bar | Averages bury the expensive miss |
| Offline green | Staging cannot 429 or deny | You tested a costume |
| Offline green | Shadow disagreement on irreversible intent is high | Live shape is not the set |
| Offline green | Cost per pass left the signed band | You bought a grind |
| Offline green | Process fail rate spiked (forbidden tool, over budget) | Lucky outcomes |
| Offline green | Evaluator version changed without a bump | You moved the goalposts |
| Offline green | No run_id a human can open | You cannot debug the next miss |
Release rule I will put in a statement of work:
- Worker changes cannot ship if golden-set blocking pass drops more than the agreed epsilon, or if a never-auto-pass fixture passes.
- Write-tool additions require a green run on every never-auto-pass fixture and a dry-run shadow window.
- Cost per passing job cannot leave the band without a named exception and an expiry.
- A hotfix that ships without a new case has 48 hours to add one, or writes revert to deny.
- Evaluator diffs ship with the same discipline as worker diffs. A pass-rate jump after a looser rubric is not a model win.
- Fail conditions are in the ticket, not in Slack folklore
- The person who can override is named; “the team” is not a name
- Override writes a
waiver_idonto the release, with an expiry - Waivers cannot cover never-auto-pass fixtures
Bravery is not a test strategy. If you need a waiver to ship the refund tool, you do not have a refund tool. You have a hope.
How do I wire the evaluator into the pre-ship loop?
Do not rebuild the judge here. Build order, independence, evidence, and ceilings are the evaluators spoke. Pre-ship, the evaluator is a CI step and a shadow grader — not a vibe in the demo thread.
Control flow I actually run:
- Ingress of the fixture or the shadowed request. Schema, size cap, auth.
- Candidate produces a non-final artifact. Write tools denied.
- Mechanical graders on artifact + trace: schema, enums, forbidden tools, budgets, citation presence.
- Blocking mechanical fail → record evidence → count a fail. No “best effort” write.
- Model judge only on residue that needs judgment, with its own context. No worker diary.
- Scoreboard cells update. Known-bad fixtures must still fail.
- Shadow: store the pair, sample disagreements, harvest.
- Widen writes only if the signed cells hold for the signed window.
| Stage | Runs on | Write tools | Job |
|---|---|---|---|
| Offline | Frozen golden set | Denied | Regression gate |
| Staging soak | Same set + hostile stubs | Dry-run | Prove the facade can fail |
| Shadow | Live inputs, sampled | Denied | Shape vs champion |
| Soft-launch | Allowlisted tenants | Gated | Harvest, still a kill switch |
| Widen | Production | Named tools only | Promotion, not a vibe |
- CI fails the merge when blocking pass drops or a never-auto-pass passes
-
doneis set by the harness after a pass, never by the model - Shadow and offline share
evaluator_version - The judge cannot call write tools. Ever.
If the evaluator only runs when someone remembers to click “Grade” in a vendor UI, you do not have a pre-ship loop. You have a hobby.
What usually breaks in the first week of testing?
The first week does not fail because the model is “bad.” It fails because the test world is dishonest, the set is a pile of happy paths, or someone equated a percentage with a ship.
| Break | How it shows up | Fix this week |
|---|---|---|
| Costume staging | 100% pass, then production 429 grind | Add 429 / timeout / empty stubs as fixtures |
| Leakage | Pass jumps after you pasted goldens into the prompt | Strip few-shots; version the set |
| Self-grading | Worker says “looks good” | Independent evaluator; see the sibling spoke |
| Outcome-only | Fluent note, forbidden tool in the trace | Process graders on the trace |
| Fake shadow | Duplicate Slack, duplicate CRM | Pull write tools off the challenger |
| Unowned scoreboard | Pass % in a slide, no cost, no escalate | Five numbers, named owners |
| Set edited mid-run | Friday’s 91% is not Thursday’s 91% | Freeze eval_set_version for the candidate |
| No harvest | Same miss twice | Every shadow disagreement gets a ticket or a row |
OpenAI’s trace-grading guide exists because black-box “did the last sentence look okay?” cannot tell you where the agent went wrong. If week one cannot open a trace, stop adding cases. Instrument first.
- One incident from staging or shadow already became a golden row
- Someone other than the prompt author labeled at least three never-auto-pass cases
- Cost showed up as dollars, not as “tokens seemed fine”
- A deny or a 429 appeared on purpose, on a fixture, and the agent stopped
If week one produced only a pass-rate screenshot, you spent the week on theater. Theater does not get cheaper when you add a second model.
When should I skip the agent and test a workflow instead?
If you can draw the steps, do not buy an agent test harness to bless a loop you did not need. Test the workflow: schema in, deterministic route, one model step if extraction is messy, write behind a gate. That is cheaper to freeze, cheaper to stub, and honest about what failed.
The long form of this fork is when not to build an agent. Pre-ship, the test itself will tell you: if every golden row has a known tool sequence, your “agent” is a flowchart with extra latency.
| Signal | Test an agent loop | Test a workflow |
|---|---|---|
| Path | Branches on retrieved state you cannot enumerate | You can draw it on a whiteboard |
| Criteria | Signed, but the route varies | Signed, and the route is stable |
| Tools | Several, chosen at run time | Named nodes, named order |
| Failure | Needs a trace + outcome grade | Needs a node error + schema check |
| Volume | Enough to pay for evals and tokens | Tiny, or bursty but identical |
Decision list:
- Can you name the next tool without a model? Ship a workflow. Test the nodes.
- Is “good” still a feeling after a workshop? Do not test an agent. You do not have a job yet.
- Is recovery a restore-from-backup story? Do not give either shape write tools until that story is a runbook.
- Is the only reason for a loop “we wanted agents”? That is not a test plan.
- Job owner can say “workflow” or “agent” without looking at a vendor slide
- If workflow, golden set still exists — it is just cheaper: input → node → schema
- If agent, you accepted trace grading, staged tools, and shadow as the price of the loop
Anthropic’s Building effective agents is the same fork in vendor language: stay with the simpler shape when you can still draw the graph; the evaluator-refiner loop only fits when criteria are clear and iteration actually improves the artifact. No criteria, no loop, no agent test plan.
An untested agent is worse than an untested workflow. An unneeded agent that you did test is still a bill.
Failure mode: the demo that passed and the write that did not
The expensive miss looks like this. Sales runs five happy tickets on production tools. The CRM notes look like a sharp intern. Pass rate on a 12-row set of paraphrases is 94%. Someone turns on SMTP for “the rest of the company.” The first empty search result becomes a confident email. The first 429 becomes a retry loop that sends three times. The known-bad fixture was never in the set because it made the demo ugly.
| Moment | What they measured | What was true |
|---|---|---|
| Demo | Fluency on greased tools | Staging could not fail |
| “94%” | Binary on paraphrases | One stratum, leaked wording |
| Ship | A percentage on a slide | No cost, no escalate, no known-bad |
| Monday | Customer-visible send | Empty retrieval, no no_match |
| Tuesday | “The model got worse” | The test was always a costume |
What it costs: unwind by hand, a trust hit with the operator who already did this job, and a week of freeze while you build the set you skipped. I will not invent a dollar figure for a composite story. The pattern is the bill: you paid production to discover cases that staging was supposed to hold.
What you do instead:
- Stop writes. Kill switch, not a prompt that says “be careful.”
- Capture the miss as a golden row: input, stubbed empty (or 429), expected terminal
escalateorabort. - Add the stub variant if it did not exist.
- Re-run the suite. The new row must fail the old candidate.
- Shadow the fix with write tools off until disagreement and process fails are in band.
- Widen one named tool, not the whole allowlist.
- The miss has a
case_id - The old candidate fails that case
- SMTP is still deny until the shadow window holds
- The slide with “94%” is retired
If your response to a bad send is “we’ll add evals later,” you are still in the demo. Later is how the second send happens.
What does a one-week pre-ship test look like?
Five days. Writes stay deny. The output is a scoreboard and a freeze, not a launch tweet. Anthropic’s advice to start with a small set of real failures is the pace: do not spend the week generating 200 paraphrases.
| Day | Owner | Done when |
|---|---|---|
| 1 | Buyer + engineer | Blocking criteria signed; irreversible tools listed; never-auto-pass named |
| 2 | Engineer | Golden rows exist with stubs; set version frozen; CI can run the suite |
| 3 | Engineer | Stub catalog includes empty, 429, timeout, deny; costume staging deleted |
| 4 | Engineer + reviewer | Offline scoreboard filled; known-bad still fail; cost per pass visible |
| 5 | Engineer + buyer | Shadow on for a signed window; challenger cannot write; harvest queue staffed |
Skip if the week is short:
- Model bake-offs
- Pairwise “which prose is nicer”
- Expanding the set with LLM-generated paraphrases of the happy path
- Wiring a second vendor judge before mechanical checks exist
- Soft-launch emails, banners, or “limited beta” with writes on
Do not skip:
-
Independent evaluator (even if it is mostly code this week)
-
Never-auto-pass fixtures
-
Honest stubs
-
Shadow with writes off, even if the window starts on day five and continues
-
Kill switch outside the model
-
Day five artifact is a document: versions, cells, fails, waivers (hopefully none)
-
Buyer can say what would still fail the ship
-
Next week’s work is harvest + one named tool, not “turn it all on”
Daily stand-up for that week, fifteen minutes, same five questions:
- Which
eval_set_versionare we on, and did anyone edit a row? - Did a never-auto-pass fixture pass? If yes, stop the candidate.
- Did staging throw an empty, a 429, and a deny on purpose?
- Can we open yesterday’s worst
run_idin under a minute? - What harvest ticket moved a miss into a row?
If question 2 or 3 fails, the rest of the day is not “more cases.” It is fixing the world.
A week will not finish maturity. It will tell you whether you have a test or a costume. If you only have a costume, you do not ship. You fix the world the agent thinks it lives in.
FAQ
How do I test AI agents before they ship?
Run a frozen golden set against staged tools that can fail, then put the candidate in shadow mode on live traffic with write tools off. Grade transcript and outcome with an independent evaluator. Do not treat a demo pass rate as the gate. Ship writes only when pass-by-stratum, cost per pass, escalate rate, and known-bad fixtures all hold for a signed window.
How do I measure whether testing AI agents before they ship is working?
The test program is working when a bad candidate fails CI or shadow before a customer sees a write, and when each production miss becomes a golden row within an agreed SLA. Watch known-bad still fail, cost per passing job, escalate rate, process fails, and shadow disagreement — not a single average. If the only artifact is a weekly percentage, the program is not working.
What usually fails first when teams try this?
Costume staging: stubs that always return 200, or production keys behind a “staging” hostname. Second is a golden set of happy paraphrases with no never-auto-pass rows. Third is fake shadow, where the challenger can still send. Fix the tool facade and the known-bad fixtures before you buy another model.
How long does this take to show results?
A week is enough to freeze a small real-failure set, prove stubs can 429 and deny, and fail a merge on a known-bad fixture. Shadow needs a signed window on live shape after that; it is not an afternoon of dogfood. You will see results the first time a candidate dies in CI instead of in a customer thread. Broader coverage is a calendar.
What should I skip if I only have a week?
Skip bake-offs, paraphrase factories, and pairwise prose contests. Do not skip signed criteria, never-auto-pass fixtures, honest stubs, an independent mechanical grader, and shadow with writes off. If you cannot staff harvest, keep writes on deny. A prettier demo is not a test.
When is this not worth doing yet?
If you have no side-effect tools, you have a chatbot; ordinary evals on answers may be enough. If the path is a flowchart, test a workflow instead of an agent. If nobody will sign criteria or own the scoreboard, you are not testing — you are collecting screenshots. Do the workshop, or do not turn writes on.
CTA
Golden set, staged tools, shadow — then writes. Pass rate is one cell, not the gate.
What questions does this article answer?
- How do I test AI agents before they ship?
- Run a frozen golden set against staged tools that can fail, then put the candidate in shadow mode on live traffic with write tools off. Grade transcript and outcome with an independent evaluator. Do not treat a demo pass rate as the gate. Ship writes only when pass-by-stratum, cost per pass, escalate rate, and known-bad fixtures all hold for a signed window.
- How do I measure whether testing AI agents before they ship is working?
- The test program is working when a bad candidate fails CI or shadow *before* a customer sees a write, and when each production miss becomes a golden row within an agreed SLA. Watch known-bad still fail, cost per passing job, escalate rate, process fails, and shadow disagreement — not a single average. If the only artifact is a weekly percentage, the program is not working.
- What usually fails first when teams try this?
- Costume staging: stubs that always return 200, or production keys behind a “staging” hostname. Second is a golden set of happy paraphrases with no never-auto-pass rows. Third is fake shadow, where the challenger can still send. Fix the tool facade and the known-bad fixtures before you buy another model.
- How long does this take to show results?
- A week is enough to freeze a small real-failure set, prove stubs can 429 and deny, and fail a merge on a known-bad fixture. Shadow needs a signed window on live shape after that; it is not an afternoon of dogfood. You will see results the first time a candidate dies in CI instead of in a customer thread. Broader coverage is a calendar.
- What should I skip if I only have a week?
- Skip bake-offs, paraphrase factories, and pairwise prose contests. Do not skip signed criteria, never-auto-pass fixtures, honest stubs, an independent mechanical grader, and shadow with writes off. If you cannot staff harvest, keep writes on deny. A prettier demo is not a test.
- When is this not worth doing yet?
- If you have no side-effect tools, you have a chatbot; ordinary evals on answers may be enough. If the path is a flowchart, test a workflow instead of an agent. If nobody will sign criteria or own the scoreboard, you are not testing — you are collecting screenshots. Do the workshop, or do not turn writes on.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.