Spurlock Studios
Contact
Share LinkedIn X
A small stack of coins. Thesis: KNOW IF AI AGENT ACTUALLY.

You know an AI agent is actually working when four outcome numbers hold: human rewrite rate, policy denials with named codes, cost per shipped job, and time-to-done. Thumbs, CSAT, “messages sent,” and a green pass percentage answer a weaker question. The expensive failure is a run that looked busy, wrote somewhere, and still needed a human to finish the artifact.

This spoke sits under the Agentic Systems Operating Manual. It owns the weekly working verdict. Deploy gates that kill a flattering pass percentage live in why pass rate lies. If the path is still a graph you can draw, stop at when not to build an agent.

The short answer

  • Working means: the artifact ships without a rewrite, denials fire for the jobs that should be blocked, cost per clean ship is a number finance would defend, and trigger-to-shipped time stays in a band operators can live with.
  • Thumbs and CSAT measure vibe. Agents write. Measure the write.
  • A denial drought is not health. A correct deny is the agent doing its job.
  • Cost per run and time-to-first-token flatter grind. Cost per shipped-without-rewrite and time-to-done are the outcome clocks.
  • One screen, four tiles, a yes or no. If you need a speech, you do not know yet.

What does “actually working” mean for an agent?

Working is not “the model replied.” It is a bundle of outcomes on a real job: the artifact is usable, the path was allowed, the spend is inside a band, and the clock from trigger to shipped is inside a band. Microsoft Foundry splits this in their own catalog: Task Completion asks whether the agent produced a usable deliverable that meets the request. Their Customer Satisfaction evaluator is a separate 1–5 Likert. Keep them separate. For a CRM write, the Likert is not the verdict.

OpenAI is equally blunt about the unit of work. Their agent evals treat the trace — model calls, tool calls, guardrails, handoffs — as the thing you grade, not a single completion. LangSmith uses the same physics: a run is one unit of work; a trace is the collection of runs for one operation. If you cannot open that tree, you are grading a chatbot.

Claim people makeWhat it actually measuredWorking verdict?
“Users thumbs-up the replies”Chat vibe on a bubbleNo
“CSAT is 4.6”Survey / Likert on the sessionNo
“Pass rate is 91%”Evaluator on the final blobCousin — not this panel
“It is fast”Time to first tokenNo
“Tokens are down this week”Vendor invoice, not jobsNo
Rewrite rate in bandHumans did not finish the jobYes
Denials present, codes namedGate is wired and firingYes
Cost / shipped-without-rewrite in bandUnit economics of a real winYes
Time-to-done p50/p95 in bandOperators can wait that longYes

I have spent 20,000+ hours architecting agentic systems and 500+ automations are live. The fleets that survived a sales team all had those four outcome tiles. The ones that died had a thumbs widget and a Slack channel of “can you just fix this one.”

  • Outcome criteria written down independent of the worker
  • Rewrite flag on every done receipt
  • Denial code on every blocked tool proposal
  • Dollars and timestamps on the same run record
  • A weekly yes/no rule written before the first pretty dashboard

Empty boxes are a demo with a production URL.

Why don’t thumbs and CSAT tell you?

Thumbs measure whether someone felt okay about a reply. Agents are hired to finish jobs. A sales rep can smash 👍 on a note that still has the wrong account, then silently edit the CRM before the meeting. The thumb stayed green. The agent did not finish the work.

Microsoft Copilot Studio itself treats reactions and CSAT as separate tiles from conversation outcomes. Thumbs-up / thumbs-down sit under “Reactions.” Outcomes sit somewhere else. If a vendor that ships a thumbs widget still refuses to collapse those numbers, you should not collapse them either.

LangSmith will happily attach feedback scores to a run. That is useful as a comment. It is not a rewrite flag, a denial code, a dollar figure, or a clock.

SignalWhat it capturesHow it lies on a write agent
Thumbs / in-product 👍Momentary vibeSilent editors never click 👎
End-of-session CSATSurvey response, low nNon-responders are often the rewrite queue
“Messages sent”TrafficBusy is not done
Provider token chartInvoice shapeCheap failures look like a win
LLM-judge CSATA second model’s LikertGrade inflation wearing a lab coat
Rewrite rateHumans still touched the artifactHard to fake if the flag is on the receipt
Policy denialsGate fired or did notZero is a wiring bug until proven otherwise

Sources of flattering thumbs:

  1. The people who rewrite never vote.
  2. The UI asks for a thumb before the write lands.
  3. “Helpful” and “correct account” are different questions.
  4. A copilot that drafts most of a note still gets a 👍 from someone who likes typing less.
  5. You mixed chatbot sessions and write jobs on one tile.

If the only number an exec can quote is a star rating, the agent is a chatbot with extra tools.

How do you read rewrite rate as a working verdict?

Rewrite rate is the share of done runs a human still edited before the artifact shipped. It is the honesty metric. Pass can go up while rewrite stays ugly — that is judge inflation or grind, which why pass rate lies already owns. This page owns the operator read: if humans still finish the job, the agent is not working yet.

I will not invent a “typical” rewrite rate for your shop. Set the band from the human baseline: how often does the current operator redo their own first draft? If you do not know that number, you are not ready to call the agent a teammate.

Compute it from receipts, not from a vibe:

  1. Terminal status done is the denominator. escalate and abort are other tiles.
  2. A human sets rewrite_flag on the shipped artifact (or the write system’s audit log does it when a field changes after done).
  3. rewrite_rate = rewritten_done / done.
  4. Slice by job_type. A FAQ draft must not hide a CRM write.
  5. Pair with reject / escalate rate so “we just stop shipping the bad ones” cannot masquerade as a rewrite win.
PatternRewritePass / thumbsRead
Clean shipLowWhateverWorking on this job type
Silent editorHighHighCopilot. Do not widen writes
Judge got softHighClimbingDo not ship the prompt
Humans stopped lookingLowHighSampling bug — audit a slice
Escalate instead of rewriten/a (not done)LowerHonest control

Who may set the flag, and what counts:

EventCounts as rewrite?Who sets it
Human edits a required field after doneYesWrite-system audit, or the reviewer
Human fixes a typo in a non-criteria fieldYes, until you prove it is noise — start strictReviewer
Human rejects and the job becomes escalateNo — different tileRunner
Evaluator fails, runner retries, then doneNo — that is revision depth, not rewriteRunner (other spoke)
No human looks, artifact shipsNo — and you must sample or this tile goes quietSampling owner

Start strict. You can later split “criteria rewrite” from “cosmetic.” You cannot later invent history you never flagged.

  • Flag lives on the receipt, not in a spreadsheet
  • Default is “unreviewed,” not “clean”
  • Unreviewed done jobs are a third rate, not a silent zero
  • One named owner for the weekly rewrite tile

Gate idea: rewrite rate is a veto on the word “working.” A pretty thumb next to a climbing rewrite tile is an unpaid editor with a dashboard.

What should policy denials look like if the agent is working?

A working agent is denied sometimes. OWASP LLM01:2025 Prompt Injection exists because instructions and untrusted data share one channel. If the catalog never says no, you do not have a gate. You have a prompt that says “please don’t.”

Anthropic’s building-effective-agents note is the cost/latency warning: agents compound errors, so you test in a sandbox and put guardrails on the path. A denial is one of those guardrails firing. Treat a correct deny as a success of the control plane, not as the agent “failing.”

Denial shapeWhat it usually meansWorking?
Zero denials all weekGate not wired, or allowlist is “everything”No — drought
Named codes, writes that should die never landCatalog is liveYes
Flood of one code after a prompt changeYou loosened the model, not the gateInspect, do not celebrate volume
Flood of unknown_toolModel inventing verbsCap autonomy
Denials with no codeYou cannot debugNot working
Deny then a retry that writes anywayGate is advisoryIncident

Minimum denial catalog for a write agent:

  • policy_destination — wrong tenant / recipient
  • policy_amount — over cap
  • policy_unknown_tool — verb not on the allowlist
  • policy_schema — args failed the contract
  • policy_catalog_down — fail closed
  • Staging drill that must produce one expected code

A week of zeroes is the first smell. Drill it on purpose. If the drill is silent, the agent is not “working.” The gate is fiction.

deny and pending-approval are not the same tile. A rubber-stamp yes is how you get a named owner and still no control plane.

VerdictWhen to use itHealthy shapeWorking lie
allowSchema pass, destination pass, under capMajority of boring jobsAllow-all catalog
denyUnknown tool, wrong tenant, hard cap, catalog downSome, with codesZero, or deny-then-write
pending-approvalOver a soft cap, new recipient, irreversible firstSparse, payload-boundInstant yes without opening args

Staging drill you can run on Monday:

  1. Send a wrong-tenant payload. Expect policy_destination. Confirm no write.
  2. Send an over-cap amount. Expect policy_amount or pending-approval — whichever you wrote down.
  3. Propose a tool not on the allowlist. Expect policy_unknown_tool.
  4. Unplug the catalog. Expect policy_catalog_down and a stopped write.
  5. Paste the run_id from the write system’s audit log. If you cannot, the join is folklore.

Retries that ignore a deny and fire the write twice are a different failure: idempotent agent tool writes exist so a timeout cannot mint a second charge. Time-to-done and cost both lie if the same intent lands twice.

Which cost number proves the job is worth doing?

Cost per run flatters agents that fail cheap and pass expensive. Finance cares about cost per successful task, and the stricter form: cost per task that ships without rewrite.

Braintrust published the arithmetic in the open. On a large agent-trace study they showed cost per task and cost per success rank configs differently: cost per success = cost per task ÷ success rate. A config that succeeds one-in-six pays for roughly six attempts per win. Their follow-up on cost per resolved request is stricter: a request only counts if it clears quality gates. Those dollar figures are their experiment, not a Spurlock fleet statistic. Steal the formula. Do not steal the numbers as if they were yours.

MetricFlattersUse for the working verdict?
Tokens / weekQuiet weeks, cheap failuresNo
Cost / runFailures that die earlyCapacity only
Cost / passGrind that eventually goes greenWatch
Cost / successHonest attempt costYes
Cost / shipped-without-rewrite + human minutesIncludes the silent editorYes — this is the working number

Procedure on last week’s traces:

  1. Sum model + tool + evaluator spend for the job_type (cost_all).
  2. Count runs that met outcome criteria and were not rewritten (shipped_clean).
  3. cost_per_clean_ship = cost_all / shipped_clean.
  4. Add human_minutes * loaded_rate for the rewritten slice.
  5. Compare to the human-only baseline for the same job. If the agent is slower-and-dearer after human time, it is not working. It is a hobby.

Illustrative arithmetic — not a measured fleet statistic. 100 CRM-note jobs. $0.40 average model+tool spend per run ($40 all-in). 60 ship clean. 40 get rewritten at 8 human minutes each. Loaded rate $0.80 / minute → $256 human. Cost / run looks like $0.40. Cost / clean ship is $40 / 60 ≈ $0.67 before humans, and ($40 + $256) / 60 ≈ $4.93 after. If a human doing the job cold is cheaper than $4.93, the agent is not working on economics. The token chart will not tell you that.

SliceInclude in cost_all?Include in shipped_clean?
done, no rewriteYesYes
done, rewrittenYesNo
escalateYesNo
abort / denied before writeYes (cheap, usually)No
Evaluator / judge spendYesn/a
Human minutes on rewritesAdd for the honest numbern/a

Anthropic will tell you agentic systems trade latency and cost for task performance and that you should add the loop only when it demonstrably improves outcomes. “Demonstrably” is this ratio, not a token chart.

What is time-to-done, and why is first-token latency a lie?

Time-to-done is wall-clock from trigger to terminal done that a human shipped with no edit. Time-to-first-token is how fast the bubble started typing. Chat products optimize TTFT so the UI feels alive. Write agents are hired to finish. A snappy first token plus a long silent rewrite is a slow job with a costume.

LangSmith will show you trace start and end, and p50/p99 latency on a thread. That is necessary. It is not sufficient. Trace duration stops when the runner stops. Time-to-done keeps ticking through pending-approval, a human in the queue, and the edit before ship.

ClockStartsStopsLie if you use it as “working”
TTFTRequestFirst tokenIgnores tools, revisions, humans
Trace durationRunner startRunner endIgnores the rewrite after done
Time-to-terminalTriggerdone / escalate / abortCounts grind that still gets rewritten
Time-to-done (this page)TriggerShipped, no rewriteThis is the operator clock
Time-to-reverseBad writeRestoredAccountability, not this panel

Pilot bands are local. I will not paste a fake p95. Write yours from the human baseline: how long does the current operator take to finish the same job cold? The agent is working on latency when p50/p95 time-to-done is inside that band without pushing rewrite rate up.

Chart percentiles, not averages:

  • p50 and p95 time-to-done by job_type
  • Share of runs that miss the band
  • Split pending-approval wait vs model/tool wait vs human rewrite wait
  • Alert when p95 climbs while thumbs stay cute

Split the wait so you know which clock blew:

Wait stateWho owns the delayIf this is the p95 driver
Model + toolsLoop ownerCap revisions; check tool latency
pending-approval queueApprover named on the catalogApprovals are the product now
Human rewrite after doneDomain reviewerAgent is a copilot — do not widen
Retry / timeout / resumeRuntimeFix idempotency before you tune prompts
Queue in front of the runnerInfraNot an agent quality bug

Illustrative clock. Trigger at 09:02. First token 09:02.4. Runner prints done at 09:06. Sales edits the account link at 09:41 and ships. TTFT: 400ms. Trace duration: 4 minutes. Time-to-done: 39 minutes. The working clock is 39. The demo clock is 400ms.

A snappy demo and a late CRM note are different products.

Which run-record fields make the four tiles real?

You cannot veto what you did not emit. The four tiles are joins over a small set of fields on every run. Missing field, missing gate.

FieldTypeFeeds
run_idstringJoin to the write system
job_typestringEvery slice
terminalenum done / escalate / abortDenominator for rewrite
rewrite_flagbool / enum / unreviewedRewrite rate
denial_codestring | nullDenial tile
cost_usdnumberCost / clean ship
t_triggertimestampTime-to-done start
t_shippedtimestamp | nullTime-to-done stop
owner_idstringWho can widen writes

OpenAI’s trace grading model is the vendor version of the same idea: score the end-to-end record, not a black-box completion. You do not need their product. You need those fields on your record.

  • unreviewed is distinct from rewrite_flag=false
  • Denial code is set even when no write fired
  • t_shipped is the write system’s ship time, not the runner’s done
  • Cost flagged if estimated vs metered
  • A missing field fails the week’s “working?” question — it does not default to green

If a field is missing, the corresponding tile is theater.

How do you combine the four into a weekly yes or no?

Four tiles. Same screen. Named owner. The question is binary: would you widen write access this week?

RewriteDenialsCost / clean shipTime-to-doneVerdict
In bandCodes present, drill greenIn bandIn bandYes — widen one notch
ClimbingAnythingAnythingAnythingNo — unpaid editor
LowDrought (zero)PrettyPrettyNo — gate is fiction
In bandFlood, one new codeSpikeSpikeNo — inspect that code
In bandHealthySpikeIn bandNo — you bought the percentage
In bandHealthyIn bandp95 blownNo — operators will route around it
Low because nobody flagsDroughtTokens “down”TTFT “fine”You do not know

Ritual that fits a 15-minute standup:

  1. Read the four tiles. No slides.
  2. Open three rewritten receipts. Name the field humans still touch.
  3. Confirm last week’s denial drill still produces the expected code.
  4. Glance at cost / clean ship and p95 time-to-done vs the written bands.
  5. Say yes or no to widening writes. No speech.

Name an owner per tile so the standup has a human, not a vibe:

TileDefault ownerActs when
Rewrite rateDomain reviewer closest to the artifactRate climbs, or unreviewed share climbs
Policy denialsPolicy catalog ownerDrought, flood, or drill miss
Cost / clean shipWhoever pays the model + tool billSpike vs written band
Time-to-doneLoop ownerp95 crosses band
Yes/no on wideningRelease owner (not the model vendor)Any veto tile is red

If you cannot answer from that screen, you do not know if the agent is working. You are hoping.

How do you sample rewrites without asking for thumbs?

Thumbs are opt-in and biased. Sampling is a quota. You pull a slice of done jobs and a human marks rewrite / clean / escalate without a 👍 widget in the path.

SampleSize (pilot default)Purpose
Every write job100% until volume hurtsRewrite flag is the product
High-volume drafts10–20% after the first two weeksCatch silent editors
New job_type100% for two weeksYou have no baseline
After a prompt/tool change100% for the canary windowCanary is not a thumb
Denial drillOne forced case / weekProve the gate still exists

Procedure:

  1. Freeze the sample rule in writing (percent, job types, owner).
  2. Reviewers mark the receipt, not a chat bubble.
  3. Unreviewed counts as a separate rate. Do not fold it into “clean.”
  4. Promote repeated rewrite reasons into criteria or into a denial row.
  5. Never replace this sample with an LLM-judge CSAT. Microsoft will sell you a Customer Satisfaction evaluator as a 1–5 Likert. Useful as a cousin. It is not a rewrite flag.

If volume is a handful of jobs a week, skip the percent. Read all of them. Sampling is for when reading all of them is the bottleneck — not for when you would rather look at stars.

What does green thumbs plus angry ops look like?

Illustrative — not a measured fleet statistic. Product shows 4.6/5 thumbs on the agent UI. Offline pass is 90%. Rewrite rate on CRM notes is 40%. Policy denials are zero for two weeks because the catalog is still a paragraph in the system prompt. p95 time-to-done is 47 minutes because sales finishes the note after the runner prints done. Cost / run looks cheap. Cost / clean ship is three times the human baseline once you add those minutes.

What the dashboard showedWhat ops feltMissing outcome tile
4.6 thumbs“Can you just fix this one”Rewrite flag
90% passWrong account linksRewrite + destination deny
Token chart “fine”Humans still in the loopCost / clean ship
TTFT 600msNotes land after the standupTime-to-done
Zero denials“It never blocks anything”Denial drill

Fix, in order:

  1. Put rewrite_flag on the shipped artifact. Stop reading thumbs.
  2. Wire allow / deny / pending-approval before the tool. Drill it.
  3. Compute cost / clean ship from the same traces.
  4. Clock trigger → shipped, not first token.
  5. Only then ask whether pass rate is allowed to move. That question belongs to why pass rate lies.

This is the failure mode thumbs are built to hide: the bubble felt fine, the path was illegal or unfinished, and the humans quietly became the runtime.

How is this different from pass rate?

Pass rate asks whether an evaluator liked the final blob. This page asks whether the business outcome happened at an acceptable human load, control-plane pulse, price, and clock. You need both. They are not substitutes.

A 2026 coding-agent study asked the public version of the same split: Does Pass Rate Tell the Whole Story? Agents cleared a large share of benchmark tests while design-satisfaction sat in a much lower band. Tests went green. Maintainers still would not merge. That gap is rewrite rate wearing a lab coat.

QuestionPass-rate spokeThis spoke
Did criteria pass on the artifact?PrimaryInput, not verdict
How many evaluate→revise cycles?Revision depthShows up as time-to-done and cost
Was the path legal?Trajectory checklistShows up as denials + rewrite
Did a human still edit?Rewrite as a gate inputRewrite as the honesty tile
What did a win cost?Cost / successCost / clean ship
How long until shipped?Not the ownerTime-to-done
Thumbs / CSATIgnoreIgnore

Use pass rate to veto a merge. Use this panel to answer “is it actually working in the room where the work lands.” If pass is green and rewrite is ugly, the agent is not working. The judge is.

What should a five-day pilot leave you with?

A Spurlock Studios $1,500 · 5-day agentic pilot should leave you able to answer yes or no from four tiles — not a thumbs widget and a pass percentage. Fancy coverage math can wait. “We looked at stars” should already be dead.

DayOutcome you can point at
1Written definition of working: rewrite, denials, cost / clean ship, time-to-done
2rewrite_flag on every done receipt; human baseline named
3Denial catalog + staging drill that produces one expected code
4Cost / clean ship and p50/p95 time-to-done on last week’s traces (or the seed)
5One-screen yes/no rule; canary plan; rollback owner

/agentic · the parent map is the operating manual.

If day five still reports a star rating, the pilot failed the only test that matters.

What should you skip if you only have a week?

Skip anything that cannot move one of the four tiles.

TemptationWhy it waits
Thumbs widget on the agent UIContaminates the rewrite signal
CSAT survey at end of sessionLow n, wrong question
Vendor bake-off (three trace products)You do not yet emit the fields
Pass@k screenshotsResearch comfort, not a working verdict
Embedding / retrieval galleriesYou are debugging a write, not a demo
Copied SLO pack from another teamYou have no baseline
Per-model token leaderboardjob_type is the unit

Ship the flag, the drill, the two clocks, and the screen. Platforms come after the receipts exist.

When is this not worth doing yet?

If you should not have an agent, you should not have an agent working-verdict. When not to build an agent is the brake: known path, mushy criteria, tiny volume, or nobody owns the SOP. A workflow with retries and a dead-letter queue is the monitor. Do not buy four tiles for a graph you can still draw.

SituationMeasure this insteadThis four-tile panel?
Known path, rare exceptionsWorkflow success / fail / DLQSkip
One messy field, then deterministic routingSchema-check fail rateSkip
No evaluator, “we’ll know it when we see it”Write criteria firstSkip — nothing to score
Volume is a handful of jobs a weekA human reading the outputProbably skip
Writes are irreversible and unownedDo not ship the writeScoreboard will not save you
Loop already on live toolsThe four tilesDo this now

Decision list:

  1. Can you draw the path without a model in the room? Ship a workflow. Measure success, fail, DLQ.
  2. Is “good” still a taste argument? Stop. There is nothing to score but opinions.
  3. Will this loop attempt a write on real data this month? If no, wait.
  4. Can one person name the rewrite flag, the denial drill, and the two clocks? If no, you are not ready to call it working.
  5. If yes to a real write and named tiles — install them this week.

Anthropic’s line still holds: stay with the simplest solution until complexity pays for itself. A working-verdict panel is complexity. Pay it when the loop can attempt a side effect on real data.

Anti-patterns

Calling the agent working because thumbs are high. You measured vibe.

Treating a policy deny as the agent failing. A correct deny is the control plane working.

Celebrating a denial drought. The gate is probably unplugged.

Cost / run after a week of cheap failures. Failures are not a discount.

TTFT as the latency SLO. The operator waits on time-to-done.

No rewrite flag, then claiming rewrite rate is “low.” You did not measure it.

One blended tile across job types. A cheap FAQ hides a broken write.

Shipping on pass rate and calling this page done. Pass is a cousin. Read why pass rate lies for the merge veto. This page is the room-where-the-work-lands veto.

FAQ

How do I know if my AI agent is actually working?

Look at four outcome numbers, not thumbs: rewrite rate on done jobs, policy denials with named codes, cost per job that shipped without a human edit, and time-to-done from trigger to that clean ship. If any tile is missing, you do not know yet. A green pass percentage without those four is a weaker question wearing a stronger costume.

How do I measure whether my AI agent is actually working?

Put a rewrite flag on the shipped artifact, run a staging denial drill, compute cost / clean ship from the same traces, and chart p50/p95 trigger-to-shipped by job_type. Then apply a written yes/no rule: would you widen writes this week? If you need a speech to explain the dashboard, the measurement is not working.

What usually fails first when teams try this?

The rewrite flag and the denial drill. Teams ship a thumbs widget and a token chart, never count the human who still edits the note, and read a week of zero denials as health. Time-to-first-token is the next miss. Cost / run is the miss after that.

How long does this take to show results?

A thin four-tile panel can exist in five business days if the runner can emit a rewrite flag, a denial code, dollars, and timestamps. You will not have a stable baseline in five days. You will have the ability to answer “did this job ship clean.” Treat the first two weeks as instrumentation, then set bands from your human baseline — not from a percentage you copied.

What should I skip if I only have a week?

Skip thumbs widgets, CSAT surveys, vendor bake-offs, and Pass@k screenshots. Ship the rewrite flag, one denial drill, cost / clean ship, time-to-done percentiles, and a one-screen yes/no rule. That is the week. Platforms come after the receipts exist.

When is this not worth doing yet?

When you should not have an agent. If the path is known, measure the workflow. If criteria are mush, you have nothing to score. Build the loop — and this panel — only after a workflow, or a workflow plus one schema-checked LLM step, fails on real traffic.

CTA

Stop calling it working because someone clicked a thumb.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

How do I know if my AI agent is actually working?
Look at four outcome numbers, not thumbs: rewrite rate on `done` jobs, policy denials with named codes, cost per job that shipped without a human edit, and time-to-done from trigger to that clean ship. If any tile is missing, you do not know yet. A green pass percentage without those four is a weaker question wearing a stronger costume.
How do I measure whether my AI agent is actually working?
Put a rewrite flag on the shipped artifact, run a staging denial drill, compute cost / clean ship from the same traces, and chart p50/p95 trigger-to-shipped by `job_type`. Then apply a written yes/no rule: would you widen writes this week? If you need a speech to explain the dashboard, the measurement is not working.
What usually fails first when teams try this?
The rewrite flag and the denial drill. Teams ship a thumbs widget and a token chart, never count the human who still edits the note, and read a week of zero denials as health. Time-to-first-token is the next miss. Cost / run is the miss after that.
How long does this take to show results?
A thin four-tile panel can exist in five business days if the runner can emit a rewrite flag, a denial code, dollars, and timestamps. You will not have a stable baseline in five days. You will have the ability to answer “did this job ship clean.” Treat the first two weeks as instrumentation, then set bands from *your* human baseline — not from a percentage you copied.
What should I skip if I only have a week?
Skip thumbs widgets, CSAT surveys, vendor bake-offs, and Pass@k screenshots. Ship the rewrite flag, one denial drill, cost / clean ship, time-to-done percentiles, and a one-screen yes/no rule. That is the week. Platforms come after the receipts exist.
When is this not worth doing yet?
When you should not have an agent. If the path is known, measure the workflow. If criteria are mush, you have nothing to score. Build the loop — and this panel — only after a workflow, or a workflow plus one schema-checked LLM step, fails on real traffic.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot