How do I know if my AI agent is actually working
You know an AI agent is working when rewrite rate, policy denials, cost, and time-to-done hold. Thumbs and CSAT hide unpaid human editors on real writes.
William Spurlock Founder — Spurlock Studios 25 MIN
You know an AI agent is actually working when four outcome numbers hold: human rewrite rate, policy denials with named codes, cost per shipped job, and time-to-done. Thumbs, CSAT, “messages sent,” and a green pass percentage answer a weaker question. The expensive failure is a run that looked busy, wrote somewhere, and still needed a human to finish the artifact.
This spoke sits under the Agentic Systems Operating Manual. It owns the weekly working verdict. Deploy gates that kill a flattering pass percentage live in why pass rate lies. If the path is still a graph you can draw, stop at when not to build an agent.
The short answer
- Working means: the artifact ships without a rewrite, denials fire for the jobs that should be blocked, cost per clean ship is a number finance would defend, and trigger-to-shipped time stays in a band operators can live with.
- Thumbs and CSAT measure vibe. Agents write. Measure the write.
- A denial drought is not health. A correct deny is the agent doing its job.
- Cost per run and time-to-first-token flatter grind. Cost per shipped-without-rewrite and time-to-done are the outcome clocks.
- One screen, four tiles, a yes or no. If you need a speech, you do not know yet.
What does “actually working” mean for an agent?
Working is not “the model replied.” It is a bundle of outcomes on a real job: the artifact is usable, the path was allowed, the spend is inside a band, and the clock from trigger to shipped is inside a band. Microsoft Foundry splits this in their own catalog: Task Completion asks whether the agent produced a usable deliverable that meets the request. Their Customer Satisfaction evaluator is a separate 1–5 Likert. Keep them separate. For a CRM write, the Likert is not the verdict.
OpenAI is equally blunt about the unit of work. Their agent evals treat the trace — model calls, tool calls, guardrails, handoffs — as the thing you grade, not a single completion. LangSmith uses the same physics: a run is one unit of work; a trace is the collection of runs for one operation. If you cannot open that tree, you are grading a chatbot.
| Claim people make | What it actually measured | Working verdict? |
|---|---|---|
| “Users thumbs-up the replies” | Chat vibe on a bubble | No |
| “CSAT is 4.6” | Survey / Likert on the session | No |
| “Pass rate is 91%” | Evaluator on the final blob | Cousin — not this panel |
| “It is fast” | Time to first token | No |
| “Tokens are down this week” | Vendor invoice, not jobs | No |
| Rewrite rate in band | Humans did not finish the job | Yes |
| Denials present, codes named | Gate is wired and firing | Yes |
| Cost / shipped-without-rewrite in band | Unit economics of a real win | Yes |
| Time-to-done p50/p95 in band | Operators can wait that long | Yes |
I have spent 20,000+ hours architecting agentic systems and 500+ automations are live. The fleets that survived a sales team all had those four outcome tiles. The ones that died had a thumbs widget and a Slack channel of “can you just fix this one.”
- Outcome criteria written down independent of the worker
- Rewrite flag on every
donereceipt - Denial code on every blocked tool proposal
- Dollars and timestamps on the same run record
- A weekly yes/no rule written before the first pretty dashboard
Empty boxes are a demo with a production URL.
Why don’t thumbs and CSAT tell you?
Thumbs measure whether someone felt okay about a reply. Agents are hired to finish jobs. A sales rep can smash 👍 on a note that still has the wrong account, then silently edit the CRM before the meeting. The thumb stayed green. The agent did not finish the work.
Microsoft Copilot Studio itself treats reactions and CSAT as separate tiles from conversation outcomes. Thumbs-up / thumbs-down sit under “Reactions.” Outcomes sit somewhere else. If a vendor that ships a thumbs widget still refuses to collapse those numbers, you should not collapse them either.
LangSmith will happily attach feedback scores to a run. That is useful as a comment. It is not a rewrite flag, a denial code, a dollar figure, or a clock.
| Signal | What it captures | How it lies on a write agent |
|---|---|---|
| Thumbs / in-product 👍 | Momentary vibe | Silent editors never click 👎 |
| End-of-session CSAT | Survey response, low n | Non-responders are often the rewrite queue |
| “Messages sent” | Traffic | Busy is not done |
| Provider token chart | Invoice shape | Cheap failures look like a win |
| LLM-judge CSAT | A second model’s Likert | Grade inflation wearing a lab coat |
| Rewrite rate | Humans still touched the artifact | Hard to fake if the flag is on the receipt |
| Policy denials | Gate fired or did not | Zero is a wiring bug until proven otherwise |
Sources of flattering thumbs:
- The people who rewrite never vote.
- The UI asks for a thumb before the write lands.
- “Helpful” and “correct account” are different questions.
- A copilot that drafts most of a note still gets a 👍 from someone who likes typing less.
- You mixed chatbot sessions and write jobs on one tile.
If the only number an exec can quote is a star rating, the agent is a chatbot with extra tools.
How do you read rewrite rate as a working verdict?
Rewrite rate is the share of done runs a human still edited before the artifact shipped. It is the honesty metric. Pass can go up while rewrite stays ugly — that is judge inflation or grind, which why pass rate lies already owns. This page owns the operator read: if humans still finish the job, the agent is not working yet.
I will not invent a “typical” rewrite rate for your shop. Set the band from the human baseline: how often does the current operator redo their own first draft? If you do not know that number, you are not ready to call the agent a teammate.
Compute it from receipts, not from a vibe:
- Terminal status
doneis the denominator.escalateandabortare other tiles. - A human sets
rewrite_flagon the shipped artifact (or the write system’s audit log does it when a field changes afterdone). rewrite_rate = rewritten_done / done.- Slice by
job_type. A FAQ draft must not hide a CRM write. - Pair with reject / escalate rate so “we just stop shipping the bad ones” cannot masquerade as a rewrite win.
| Pattern | Rewrite | Pass / thumbs | Read |
|---|---|---|---|
| Clean ship | Low | Whatever | Working on this job type |
| Silent editor | High | High | Copilot. Do not widen writes |
| Judge got soft | High | Climbing | Do not ship the prompt |
| Humans stopped looking | Low | High | Sampling bug — audit a slice |
| Escalate instead of rewrite | n/a (not done) | Lower | Honest control |
Who may set the flag, and what counts:
| Event | Counts as rewrite? | Who sets it |
|---|---|---|
Human edits a required field after done | Yes | Write-system audit, or the reviewer |
| Human fixes a typo in a non-criteria field | Yes, until you prove it is noise — start strict | Reviewer |
Human rejects and the job becomes escalate | No — different tile | Runner |
Evaluator fails, runner retries, then done | No — that is revision depth, not rewrite | Runner (other spoke) |
| No human looks, artifact ships | No — and you must sample or this tile goes quiet | Sampling owner |
Start strict. You can later split “criteria rewrite” from “cosmetic.” You cannot later invent history you never flagged.
- Flag lives on the receipt, not in a spreadsheet
- Default is “unreviewed,” not “clean”
- Unreviewed
donejobs are a third rate, not a silent zero - One named owner for the weekly rewrite tile
Gate idea: rewrite rate is a veto on the word “working.” A pretty thumb next to a climbing rewrite tile is an unpaid editor with a dashboard.
What should policy denials look like if the agent is working?
A working agent is denied sometimes. OWASP LLM01:2025 Prompt Injection exists because instructions and untrusted data share one channel. If the catalog never says no, you do not have a gate. You have a prompt that says “please don’t.”
Anthropic’s building-effective-agents note is the cost/latency warning: agents compound errors, so you test in a sandbox and put guardrails on the path. A denial is one of those guardrails firing. Treat a correct deny as a success of the control plane, not as the agent “failing.”
| Denial shape | What it usually means | Working? |
|---|---|---|
| Zero denials all week | Gate not wired, or allowlist is “everything” | No — drought |
| Named codes, writes that should die never land | Catalog is live | Yes |
| Flood of one code after a prompt change | You loosened the model, not the gate | Inspect, do not celebrate volume |
Flood of unknown_tool | Model inventing verbs | Cap autonomy |
| Denials with no code | You cannot debug | Not working |
| Deny then a retry that writes anyway | Gate is advisory | Incident |
Minimum denial catalog for a write agent:
-
policy_destination— wrong tenant / recipient -
policy_amount— over cap -
policy_unknown_tool— verb not on the allowlist -
policy_schema— args failed the contract -
policy_catalog_down— fail closed - Staging drill that must produce one expected code
A week of zeroes is the first smell. Drill it on purpose. If the drill is silent, the agent is not “working.” The gate is fiction.
deny and pending-approval are not the same tile. A rubber-stamp yes is how you get a named owner and still no control plane.
| Verdict | When to use it | Healthy shape | Working lie |
|---|---|---|---|
allow | Schema pass, destination pass, under cap | Majority of boring jobs | Allow-all catalog |
deny | Unknown tool, wrong tenant, hard cap, catalog down | Some, with codes | Zero, or deny-then-write |
pending-approval | Over a soft cap, new recipient, irreversible first | Sparse, payload-bound | Instant yes without opening args |
Staging drill you can run on Monday:
- Send a wrong-tenant payload. Expect
policy_destination. Confirm no write. - Send an over-cap amount. Expect
policy_amountorpending-approval— whichever you wrote down. - Propose a tool not on the allowlist. Expect
policy_unknown_tool. - Unplug the catalog. Expect
policy_catalog_downand a stopped write. - Paste the
run_idfrom the write system’s audit log. If you cannot, the join is folklore.
Retries that ignore a deny and fire the write twice are a different failure: idempotent agent tool writes exist so a timeout cannot mint a second charge. Time-to-done and cost both lie if the same intent lands twice.
Which cost number proves the job is worth doing?
Cost per run flatters agents that fail cheap and pass expensive. Finance cares about cost per successful task, and the stricter form: cost per task that ships without rewrite.
Braintrust published the arithmetic in the open. On a large agent-trace study they showed cost per task and cost per success rank configs differently: cost per success = cost per task ÷ success rate. A config that succeeds one-in-six pays for roughly six attempts per win. Their follow-up on cost per resolved request is stricter: a request only counts if it clears quality gates. Those dollar figures are their experiment, not a Spurlock fleet statistic. Steal the formula. Do not steal the numbers as if they were yours.
| Metric | Flatters | Use for the working verdict? |
|---|---|---|
| Tokens / week | Quiet weeks, cheap failures | No |
| Cost / run | Failures that die early | Capacity only |
| Cost / pass | Grind that eventually goes green | Watch |
| Cost / success | Honest attempt cost | Yes |
| Cost / shipped-without-rewrite + human minutes | Includes the silent editor | Yes — this is the working number |
Procedure on last week’s traces:
- Sum model + tool + evaluator spend for the
job_type(cost_all). - Count runs that met outcome criteria and were not rewritten (
shipped_clean). cost_per_clean_ship = cost_all / shipped_clean.- Add
human_minutes * loaded_ratefor the rewritten slice. - Compare to the human-only baseline for the same job. If the agent is slower-and-dearer after human time, it is not working. It is a hobby.
Illustrative arithmetic — not a measured fleet statistic. 100 CRM-note jobs. $0.40 average model+tool spend per run ($40 all-in). 60 ship clean. 40 get rewritten at 8 human minutes each. Loaded rate $0.80 / minute → $256 human. Cost / run looks like $0.40. Cost / clean ship is $40 / 60 ≈ $0.67 before humans, and ($40 + $256) / 60 ≈ $4.93 after. If a human doing the job cold is cheaper than $4.93, the agent is not working on economics. The token chart will not tell you that.
| Slice | Include in cost_all? | Include in shipped_clean? |
|---|---|---|
done, no rewrite | Yes | Yes |
done, rewritten | Yes | No |
escalate | Yes | No |
abort / denied before write | Yes (cheap, usually) | No |
| Evaluator / judge spend | Yes | n/a |
| Human minutes on rewrites | Add for the honest number | n/a |
Anthropic will tell you agentic systems trade latency and cost for task performance and that you should add the loop only when it demonstrably improves outcomes. “Demonstrably” is this ratio, not a token chart.
What is time-to-done, and why is first-token latency a lie?
Time-to-done is wall-clock from trigger to terminal done that a human shipped with no edit. Time-to-first-token is how fast the bubble started typing. Chat products optimize TTFT so the UI feels alive. Write agents are hired to finish. A snappy first token plus a long silent rewrite is a slow job with a costume.
LangSmith will show you trace start and end, and p50/p99 latency on a thread. That is necessary. It is not sufficient. Trace duration stops when the runner stops. Time-to-done keeps ticking through pending-approval, a human in the queue, and the edit before ship.
| Clock | Starts | Stops | Lie if you use it as “working” |
|---|---|---|---|
| TTFT | Request | First token | Ignores tools, revisions, humans |
| Trace duration | Runner start | Runner end | Ignores the rewrite after done |
| Time-to-terminal | Trigger | done / escalate / abort | Counts grind that still gets rewritten |
| Time-to-done (this page) | Trigger | Shipped, no rewrite | This is the operator clock |
| Time-to-reverse | Bad write | Restored | Accountability, not this panel |
Pilot bands are local. I will not paste a fake p95. Write yours from the human baseline: how long does the current operator take to finish the same job cold? The agent is working on latency when p50/p95 time-to-done is inside that band without pushing rewrite rate up.
Chart percentiles, not averages:
- p50 and p95 time-to-done by
job_type - Share of runs that miss the band
- Split
pending-approvalwait vs model/tool wait vs human rewrite wait - Alert when p95 climbs while thumbs stay cute
Split the wait so you know which clock blew:
| Wait state | Who owns the delay | If this is the p95 driver |
|---|---|---|
| Model + tools | Loop owner | Cap revisions; check tool latency |
pending-approval queue | Approver named on the catalog | Approvals are the product now |
Human rewrite after done | Domain reviewer | Agent is a copilot — do not widen |
| Retry / timeout / resume | Runtime | Fix idempotency before you tune prompts |
| Queue in front of the runner | Infra | Not an agent quality bug |
Illustrative clock. Trigger at 09:02. First token 09:02.4. Runner prints done at 09:06. Sales edits the account link at 09:41 and ships. TTFT: 400ms. Trace duration: 4 minutes. Time-to-done: 39 minutes. The working clock is 39. The demo clock is 400ms.
A snappy demo and a late CRM note are different products.
Which run-record fields make the four tiles real?
You cannot veto what you did not emit. The four tiles are joins over a small set of fields on every run. Missing field, missing gate.
| Field | Type | Feeds |
|---|---|---|
run_id | string | Join to the write system |
job_type | string | Every slice |
terminal | enum done / escalate / abort | Denominator for rewrite |
rewrite_flag | bool / enum / unreviewed | Rewrite rate |
denial_code | string | null | Denial tile |
cost_usd | number | Cost / clean ship |
t_trigger | timestamp | Time-to-done start |
t_shipped | timestamp | null | Time-to-done stop |
owner_id | string | Who can widen writes |
OpenAI’s trace grading model is the vendor version of the same idea: score the end-to-end record, not a black-box completion. You do not need their product. You need those fields on your record.
-
unreviewedis distinct fromrewrite_flag=false - Denial code is set even when no write fired
-
t_shippedis the write system’s ship time, not the runner’sdone - Cost flagged if estimated vs metered
- A missing field fails the week’s “working?” question — it does not default to green
If a field is missing, the corresponding tile is theater.
How do you combine the four into a weekly yes or no?
Four tiles. Same screen. Named owner. The question is binary: would you widen write access this week?
| Rewrite | Denials | Cost / clean ship | Time-to-done | Verdict |
|---|---|---|---|---|
| In band | Codes present, drill green | In band | In band | Yes — widen one notch |
| Climbing | Anything | Anything | Anything | No — unpaid editor |
| Low | Drought (zero) | Pretty | Pretty | No — gate is fiction |
| In band | Flood, one new code | Spike | Spike | No — inspect that code |
| In band | Healthy | Spike | In band | No — you bought the percentage |
| In band | Healthy | In band | p95 blown | No — operators will route around it |
| Low because nobody flags | Drought | Tokens “down” | TTFT “fine” | You do not know |
Ritual that fits a 15-minute standup:
- Read the four tiles. No slides.
- Open three rewritten receipts. Name the field humans still touch.
- Confirm last week’s denial drill still produces the expected code.
- Glance at cost / clean ship and p95 time-to-done vs the written bands.
- Say yes or no to widening writes. No speech.
Name an owner per tile so the standup has a human, not a vibe:
| Tile | Default owner | Acts when |
|---|---|---|
| Rewrite rate | Domain reviewer closest to the artifact | Rate climbs, or unreviewed share climbs |
| Policy denials | Policy catalog owner | Drought, flood, or drill miss |
| Cost / clean ship | Whoever pays the model + tool bill | Spike vs written band |
| Time-to-done | Loop owner | p95 crosses band |
| Yes/no on widening | Release owner (not the model vendor) | Any veto tile is red |
If you cannot answer from that screen, you do not know if the agent is working. You are hoping.
How do you sample rewrites without asking for thumbs?
Thumbs are opt-in and biased. Sampling is a quota. You pull a slice of done jobs and a human marks rewrite / clean / escalate without a 👍 widget in the path.
| Sample | Size (pilot default) | Purpose |
|---|---|---|
| Every write job | 100% until volume hurts | Rewrite flag is the product |
| High-volume drafts | 10–20% after the first two weeks | Catch silent editors |
New job_type | 100% for two weeks | You have no baseline |
| After a prompt/tool change | 100% for the canary window | Canary is not a thumb |
| Denial drill | One forced case / week | Prove the gate still exists |
Procedure:
- Freeze the sample rule in writing (percent, job types, owner).
- Reviewers mark the receipt, not a chat bubble.
- Unreviewed counts as a separate rate. Do not fold it into “clean.”
- Promote repeated rewrite reasons into criteria or into a denial row.
- Never replace this sample with an LLM-judge CSAT. Microsoft will sell you a Customer Satisfaction evaluator as a 1–5 Likert. Useful as a cousin. It is not a rewrite flag.
If volume is a handful of jobs a week, skip the percent. Read all of them. Sampling is for when reading all of them is the bottleneck — not for when you would rather look at stars.
What does green thumbs plus angry ops look like?
Illustrative — not a measured fleet statistic. Product shows 4.6/5 thumbs on the agent UI. Offline pass is 90%. Rewrite rate on CRM notes is 40%. Policy denials are zero for two weeks because the catalog is still a paragraph in the system prompt. p95 time-to-done is 47 minutes because sales finishes the note after the runner prints done. Cost / run looks cheap. Cost / clean ship is three times the human baseline once you add those minutes.
| What the dashboard showed | What ops felt | Missing outcome tile |
|---|---|---|
| 4.6 thumbs | “Can you just fix this one” | Rewrite flag |
| 90% pass | Wrong account links | Rewrite + destination deny |
| Token chart “fine” | Humans still in the loop | Cost / clean ship |
| TTFT 600ms | Notes land after the standup | Time-to-done |
| Zero denials | “It never blocks anything” | Denial drill |
Fix, in order:
- Put
rewrite_flagon the shipped artifact. Stop reading thumbs. - Wire
allow/deny/pending-approvalbefore the tool. Drill it. - Compute cost / clean ship from the same traces.
- Clock trigger → shipped, not first token.
- Only then ask whether pass rate is allowed to move. That question belongs to why pass rate lies.
This is the failure mode thumbs are built to hide: the bubble felt fine, the path was illegal or unfinished, and the humans quietly became the runtime.
How is this different from pass rate?
Pass rate asks whether an evaluator liked the final blob. This page asks whether the business outcome happened at an acceptable human load, control-plane pulse, price, and clock. You need both. They are not substitutes.
A 2026 coding-agent study asked the public version of the same split: Does Pass Rate Tell the Whole Story? Agents cleared a large share of benchmark tests while design-satisfaction sat in a much lower band. Tests went green. Maintainers still would not merge. That gap is rewrite rate wearing a lab coat.
| Question | Pass-rate spoke | This spoke |
|---|---|---|
| Did criteria pass on the artifact? | Primary | Input, not verdict |
| How many evaluate→revise cycles? | Revision depth | Shows up as time-to-done and cost |
| Was the path legal? | Trajectory checklist | Shows up as denials + rewrite |
| Did a human still edit? | Rewrite as a gate input | Rewrite as the honesty tile |
| What did a win cost? | Cost / success | Cost / clean ship |
| How long until shipped? | Not the owner | Time-to-done |
| Thumbs / CSAT | Ignore | Ignore |
Use pass rate to veto a merge. Use this panel to answer “is it actually working in the room where the work lands.” If pass is green and rewrite is ugly, the agent is not working. The judge is.
What should a five-day pilot leave you with?
A Spurlock Studios $1,500 · 5-day agentic pilot should leave you able to answer yes or no from four tiles — not a thumbs widget and a pass percentage. Fancy coverage math can wait. “We looked at stars” should already be dead.
| Day | Outcome you can point at |
|---|---|
| 1 | Written definition of working: rewrite, denials, cost / clean ship, time-to-done |
| 2 | rewrite_flag on every done receipt; human baseline named |
| 3 | Denial catalog + staging drill that produces one expected code |
| 4 | Cost / clean ship and p50/p95 time-to-done on last week’s traces (or the seed) |
| 5 | One-screen yes/no rule; canary plan; rollback owner |
/agentic · the parent map is the operating manual.
If day five still reports a star rating, the pilot failed the only test that matters.
What should you skip if you only have a week?
Skip anything that cannot move one of the four tiles.
| Temptation | Why it waits |
|---|---|
| Thumbs widget on the agent UI | Contaminates the rewrite signal |
| CSAT survey at end of session | Low n, wrong question |
| Vendor bake-off (three trace products) | You do not yet emit the fields |
| Pass@k screenshots | Research comfort, not a working verdict |
| Embedding / retrieval galleries | You are debugging a write, not a demo |
| Copied SLO pack from another team | You have no baseline |
| Per-model token leaderboard | job_type is the unit |
Ship the flag, the drill, the two clocks, and the screen. Platforms come after the receipts exist.
When is this not worth doing yet?
If you should not have an agent, you should not have an agent working-verdict. When not to build an agent is the brake: known path, mushy criteria, tiny volume, or nobody owns the SOP. A workflow with retries and a dead-letter queue is the monitor. Do not buy four tiles for a graph you can still draw.
| Situation | Measure this instead | This four-tile panel? |
|---|---|---|
| Known path, rare exceptions | Workflow success / fail / DLQ | Skip |
| One messy field, then deterministic routing | Schema-check fail rate | Skip |
| No evaluator, “we’ll know it when we see it” | Write criteria first | Skip — nothing to score |
| Volume is a handful of jobs a week | A human reading the output | Probably skip |
| Writes are irreversible and unowned | Do not ship the write | Scoreboard will not save you |
| Loop already on live tools | The four tiles | Do this now |
Decision list:
- Can you draw the path without a model in the room? Ship a workflow. Measure success, fail, DLQ.
- Is “good” still a taste argument? Stop. There is nothing to score but opinions.
- Will this loop attempt a write on real data this month? If no, wait.
- Can one person name the rewrite flag, the denial drill, and the two clocks? If no, you are not ready to call it working.
- If yes to a real write and named tiles — install them this week.
Anthropic’s line still holds: stay with the simplest solution until complexity pays for itself. A working-verdict panel is complexity. Pay it when the loop can attempt a side effect on real data.
Anti-patterns
Calling the agent working because thumbs are high. You measured vibe.
Treating a policy deny as the agent failing. A correct deny is the control plane working.
Celebrating a denial drought. The gate is probably unplugged.
Cost / run after a week of cheap failures. Failures are not a discount.
TTFT as the latency SLO. The operator waits on time-to-done.
No rewrite flag, then claiming rewrite rate is “low.” You did not measure it.
One blended tile across job types. A cheap FAQ hides a broken write.
Shipping on pass rate and calling this page done. Pass is a cousin. Read why pass rate lies for the merge veto. This page is the room-where-the-work-lands veto.
FAQ
How do I know if my AI agent is actually working?
Look at four outcome numbers, not thumbs: rewrite rate on done jobs, policy denials with named codes, cost per job that shipped without a human edit, and time-to-done from trigger to that clean ship. If any tile is missing, you do not know yet. A green pass percentage without those four is a weaker question wearing a stronger costume.
How do I measure whether my AI agent is actually working?
Put a rewrite flag on the shipped artifact, run a staging denial drill, compute cost / clean ship from the same traces, and chart p50/p95 trigger-to-shipped by job_type. Then apply a written yes/no rule: would you widen writes this week? If you need a speech to explain the dashboard, the measurement is not working.
What usually fails first when teams try this?
The rewrite flag and the denial drill. Teams ship a thumbs widget and a token chart, never count the human who still edits the note, and read a week of zero denials as health. Time-to-first-token is the next miss. Cost / run is the miss after that.
How long does this take to show results?
A thin four-tile panel can exist in five business days if the runner can emit a rewrite flag, a denial code, dollars, and timestamps. You will not have a stable baseline in five days. You will have the ability to answer “did this job ship clean.” Treat the first two weeks as instrumentation, then set bands from your human baseline — not from a percentage you copied.
What should I skip if I only have a week?
Skip thumbs widgets, CSAT surveys, vendor bake-offs, and Pass@k screenshots. Ship the rewrite flag, one denial drill, cost / clean ship, time-to-done percentiles, and a one-screen yes/no rule. That is the week. Platforms come after the receipts exist.
When is this not worth doing yet?
When you should not have an agent. If the path is known, measure the workflow. If criteria are mush, you have nothing to score. Build the loop — and this panel — only after a workflow, or a workflow plus one schema-checked LLM step, fails on real traffic.
CTA
Stop calling it working because someone clicked a thumb.
What questions does this article answer?
- How do I know if my AI agent is actually working?
- Look at four outcome numbers, not thumbs: rewrite rate on `done` jobs, policy denials with named codes, cost per job that shipped without a human edit, and time-to-done from trigger to that clean ship. If any tile is missing, you do not know yet. A green pass percentage without those four is a weaker question wearing a stronger costume.
- How do I measure whether my AI agent is actually working?
- Put a rewrite flag on the shipped artifact, run a staging denial drill, compute cost / clean ship from the same traces, and chart p50/p95 trigger-to-shipped by `job_type`. Then apply a written yes/no rule: would you widen writes this week? If you need a speech to explain the dashboard, the measurement is not working.
- What usually fails first when teams try this?
- The rewrite flag and the denial drill. Teams ship a thumbs widget and a token chart, never count the human who still edits the note, and read a week of zero denials as health. Time-to-first-token is the next miss. Cost / run is the miss after that.
- How long does this take to show results?
- A thin four-tile panel can exist in five business days if the runner can emit a rewrite flag, a denial code, dollars, and timestamps. You will not have a stable baseline in five days. You will have the ability to answer “did this job ship clean.” Treat the first two weeks as instrumentation, then set bands from *your* human baseline — not from a percentage you copied.
- What should I skip if I only have a week?
- Skip thumbs widgets, CSAT surveys, vendor bake-offs, and Pass@k screenshots. Ship the rewrite flag, one denial drill, cost / clean ship, time-to-done percentiles, and a one-screen yes/no rule. That is the week. Platforms come after the receipts exist.
- When is this not worth doing yet?
- When you should not have an agent. If the path is known, measure the workflow. If criteria are mush, you have nothing to score. Build the loop — and this panel — only after a workflow, or a workflow plus one schema-checked LLM step, fails on real traffic.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.