How do I monitor AI agents in production
Monitor production agents on traces, tool errors, cost per turn, human rewrite rate, and policy denials — not chatbot thumbs. That is the ops scoreboard.
William Spurlock Founder — Spurlock Studios 28 MIN
You monitor AI agents in production with a scoreboard, not a chatbot survey. The numbers that matter are traces you can open, tool-error rates, cost per turn, human rewrite rate, and policy denials. Thumbs-up on a chat bubble hides the expensive failure: a run that spent money, wrote to the CRM, and still needed a human to fix the note.
This spoke sits under the Agentic Systems Operating Manual. If you still need a workflow instead of a loop, stop at when not to build an agent. If the loop only ever worked on staged tools, read why agent demos fail production before you trust a green tile.
The short answer
- Every production run needs a
run_ida human can paste, a tool-error code, a dollar figure for that turn, a rewrite flag, and a policy verdict. - Chatbot CSAT, “messages sent,” and provider token charts are traffic. They do not tell you whether the write was right.
- Tool errors and policy denials are the control-plane pulse. A week of zeroes usually means the gates are not wired, not that the agent is healthy.
- Cost belongs per turn and per passing job, by
job_type. Weekly token totals arrive too late to stop a grind. - Page on deltas from your baseline. Do not invent a 99% SLO because a vendor dashboard had a blank for one.
What does production monitoring mean for an agent?
Monitoring a chatbot asks “did the reply feel okay?” Monitoring an agent asks “did this job finish, with which tools, at what cost, under which policy, and did a human still have to rewrite the artifact?”
OpenAI’s eval guidance is explicit that the unit of work is the trace — model calls, tool calls, guardrails, and handoffs — not a single completion. LangSmith uses the same physics under different nouns: a run is one unit of work; a trace is the collection of runs for one operation. If you cannot open that tree, you are not monitoring the agent. You are monitoring the model vendor.
| Question | Chatbot view | Agent view |
|---|---|---|
| Did it work? | Thumbs / CSAT | Evaluator pass and no human rewrite |
| What broke? | “The model was weird” | Tool error code or policy denial code |
| What did it cost? | Tokens this week | Dollars per turn, by job_type |
| Can we debug it? | Paste the reply into Slack | Paste run_id, open the trace |
| Should we keep writing? | Vibes | Kill switch + denial rate |
I have spent 20,000+ hours on agentic systems and built 500+ automations. The agents that survived contact with a live CRM all had those five numbers on one screen. The ones that did not were demos with a production URL.
- One
run_idon every terminal (done/escalate/abort) - Tool spans with an error code, not a stack novel
- Dollars on the turn, flagged if estimated
- A rewrite / reject flag a human can set
- Policy denials counted even when the write never fired
If any box is empty, you have logs. Not monitoring.
What is the five-number scoreboard?
Five numbers. Named owners. Same screen every week. Everything else is a drill-down.
| Number | Definition | Healthy shape | Lie if you skip it |
|---|---|---|---|
| Traces opened | Share of incidents where ops opened the run_id in under a minute | You can actually debug | You are guessing from Slack screenshots |
| Tool-error rate | Tool spans with a non-empty error_code, per job_type | Spikes are investigated | You blame “the model” for a 401 |
| Cost per turn | Model $ + paid-tool $ for one loop iteration | Band by job_type | Cheap failures look efficient |
| Human rewrite rate | done runs a human still edited before the artifact shipped | Tracked beside pass | Pass rate greens a grind |
| Policy denials | Gate blocked a side effect, with a catalog code | Non-zero in staging; explained in prod | Gates are theater |
OpenTelemetry’s split is the one to steal: traces tell the story of one job; metrics are what you aggregate and page on. Do not page off a trace you might have sampled away. Do not debug an incident off a counter with no run_id.
| Scoreboard tile | Signal type | Owner |
|---|---|---|
| Tool-error rate | Metric | On-call |
| Cost per turn / cost per pass | Metric | Whoever holds the budget |
| Human rewrite rate | Metric + sample of traces | Domain reviewer |
| Policy denials | Metric | Policy owner |
| Trace jump (receipt → run) | Join table, not a chart | Eng |
How to compute each number without a data team:
- Traces opened: incidents this week where receipt →
run_idsucceeded, divided by incidents. If the denominator is zero, you did not have an incident — or you did not log one. - Tool-error rate: tool spans with
error_codeset / tool spans, bytool.nameandjob_type. Excludepolicy_*codes; those belong on the denial tile. - Cost per turn: sum of model $ and tool $ between
state:planenter and the next eval. If you cannot bound a turn, you cannot bound a grind. - Human rewrite rate:
human_rewrite=trueondonereceipts /donereceipts. Escalate packages that a reviewer rejects are a different tile — do not mix them. - Policy denials: spans with
policy_*/ attempted side effects. A denial on a read-only lookup is a mis-fired gate, not a win.
| Anti-pattern | What it does to the board |
|---|---|
| Mixing eval-fail into tool-error | You page the model for a connector outage |
| Mixing escalate-reject into rewrite | You punish the honest handoff |
Global average across job_type | One cheap job hides one expensive grind |
| Counting denials as errors | On-call fights the control plane |
Thumbs do not appear on this board. If someone wants a satisfaction widget, put it on the chatbot. Agents write to systems. Monitor the write.
Why are chatbot thumbs not monitoring?
A thumb is a feeling about a sentence. An agent’s damage is a side effect: the wrong account, the wrong refund, the public email that should have stayed a draft. Nobody thumbs-down a CRM note they never saw.
| Chatbot metric | What it measures | What it hides on an agent |
|---|---|---|
| Thumbs up / down | Tone of the last message | Wrong tool, wrong args, skipped eval |
| CSAT / “Was this helpful?” | Sentiment | Silent wrong writes |
| Messages sent | Activity | Terminal mix (done / escalate / abort) |
| Provider 200 / finish reason | The API accepted the call | 200 with an empty body |
| Token chart | Spend this hour | Cost per passing turn |
| “Agent said done” | Worker confidence | Evaluator never ran; human still rewrote |
The failure mode is familiar: ops watches a green CSAT tile, sales forwards a bad note three days later, and nobody can find the run. Correlation ids are cheaper than a dedicated “AI ops” hire.
OpenAI’s Agents SDK already records guardrail spans beside function spans. If your dashboard cannot show those events, you are not using the trace. You are using a chat log with extra fields.
- Remove thumbs from the agent ops screen
- Keep a “report wrong” control that captures
run_id - Route those reports into rewrite rate, not into a CSAT export
- Stop calling HTTP 200 a pass
Feelings are not a control loop.
How do you wire traces you can actually open?
A useful trace is not a JSON bucket. It is a tree a human can walk in one sitting: job, states, model calls, tools, eval, terminal. OpenTelemetry calls that path a trace and each unit of work a span. You still have to put agent fields on those spans — state name, side-effect class, reason code — because HTTP tracing never asked for them.
Procedure for a new job type:
- Mint
run_id/trace_idat intake. Never let a child agent start a second unrelated id. - Emit a span for every model call and every tool call. Name the tool; do not bury it in the prompt dump.
- Attach the write system’s receipt (
crm_note_id, ticket id, draft id) on success. - Store a catalog
reason_codeon the terminal. Free text cannot group. - Redact args and results at write time. OWASP still ranks Sensitive Information Disclosure as LLM02 — traces are a disclosure path.
- Prove the join: from a CRM note, open the run in under a minute. If you cannot, the trace is decoration.
| Must store | Must not store by default |
|---|---|
run_id, job_type, tenant | Full customer record |
Tool name, error_code, latency | Raw credentials |
| Side-effect class | Unredacted PII payloads |
| Evaluator verdict + criterion codes | A second copy of the artifact |
| Cost fields, revision count | Last month’s price list applied blindly |
| Terminal reason code | A paragraph |
LangSmith will happily keep a thread of many traces. Business agents need a third join the vendor will not invent for you: the receipt in the system of record.
- Receipt id →
run_idis a real lookup, not a folklore Slack search - Redaction policy applied at write, not “we will clean the bucket”
- Staging can dump more; production defaults to hashes plus a short preview
If the only way to debug is downloading a prompt dump, the UX failed.
What do tool errors tell you that model charts miss?
Most business damage is a tool call. Tracing only the model is how you miss the write. A provider latency chart will not tell you the CRM returned 200 with an empty body, or that the allowlisted “update note” tool started 401-ing at 2 a.m. when a token expired.
error_code | Means | Typical next action |
|---|---|---|
tool_timeout | Connector hung | Bound the wait; escalate the run |
tool_auth | 401 / 403 / scope miss | Rotate or shrink the allowlist |
tool_schema | Args failed the contract | Pin the schema; stop guessing fields |
tool_empty_200 | Success status, empty body | Treat as fail; do not eval an empty write |
tool_rate_limit | Vendor 429 | Backoff with a ceiling; do not spin |
tool_conflict | Idempotency or unique-key clash | Reuse the receipt; do not double-write |
tool_not_allowlisted | Model invented a tool | Denial, not a retry |
Count rates by tool.name and job_type. A 3% error rate on crm.get is a different incident than a 3% rate on crm.refund.
| Signal | Page? | Why |
|---|---|---|
| Spike vs trailing baseline for one tool | Yes | Connector or auth is on fire |
New error_code you have never seen | Yes | Catalog gap or vendor change |
Steady low-rate tool_empty_200 | Review this week | Silent wrongness factory |
| Model “sorry” text with no tool error | No | That is a prompt; look at eval fail instead |
OpenAI’s trace-grading questions start in the same place: did the agent pick the right tool, and did a guardrail fire when it should have? Grade the tool span. Do not grade the apology.
Procedure when you add a tool:
- Give it a stable
tool.name. Do not rename it every prompt tweak. - Classify the side effect (
read/draft/irreversible_write) before the first production call. - Map vendor failures onto the catalog above. Raw HTTP status is not a catalog.
- Emit
error_codeon fail and on 200-empty. Status is not truth. - Refuse to ship the tool if the span cannot name the side-effect class.
- New tool has a catalog mapping, not “see stack dump”
- Shadow tenant traffic hits the live connector shape
- On-call knows which tile moves when this tool dies
The miss I still see: teams alert on model 500s and ignore tool 200s. The 200s are where the money went.
How should you track cost per turn?
Cost that lands in a monthly export is archaeology. Cost that sits on the turn is a control.
A turn is one pass through plan → act → eval (and maybe revise). Cost per turn is model dollars plus paid-tool dollars for that pass. Cost per pass — dollars divided by evaluator-passing jobs — is the number finance can defend. Cheap failures make cost-per-run look efficient. Do not use that one.
Langfuse’s rule is the one to copy: ingested provider usage beats inferred token math, and inferred cost is computed at ingest against the price list you had then. Their docs are also blunt that only generation and embedding observations carry usage. Paid search, enrichment, and code-exec sandboxes will vanish unless you attach a dollar field to the tool span yourself.
| Cost field | Source of truth | Lie if you skip it |
|---|---|---|
| Input / output / cache tokens | Provider usage object | You invent a tokenizer |
| Model $ | Price list pinned at ingest | Last month’s list on this month’s run |
| Tool $ | Vendor meter or invoice map | “AI spend” misses the APIs |
| Revisions | Runner counter | Cheap first pass, expensive grind |
| Cost per turn | Sum for one loop iteration | You cannot see the grind |
| Cost per pass | Dollars / evaluator-passing jobs | Failures look cheap |
Chart bands by job_type, not by model nickname. The question is “what does a passing CRM note cost,” not “how many tokens did the flagship burn on Tuesday.”
- Provider usage copied onto the model span
- Estimates flagged
usage_estimated=true - Paid tools have a dollar field
- Cost per pass uses evaluator pass, not HTTP success
- Budget trips are events on the same board
If you cannot name this week’s cost per passing job, you are not monitoring unit economics. You are watching a token fireplace.
Why is human rewrite rate the honesty metric?
Pass rate can green while sales still rewrites half the notes. That is not a healthy agent. That is a vanity tile with a human in the loop you forgot to count.
Rewrite rate is simple: of the runs that ended done, what share did a human still edit, reject, or re-send before the artifact shipped? Capture it as a flag on the receipt, not as a feeling in standup.
| Capture method | Strength | Weakness |
|---|---|---|
“Report wrong” control with run_id | Fast, tied to the trace | Silent rewrites never click it |
| CRM edited-after-agent timestamp | Catches silent fixes | Need a join; timezone lies |
| Reviewer decision on an escalate package | Clean labels | Misses the done path |
| Sales forward of a bad note | High-signal | Slow; incomplete |
Procedure:
- Put
run_idon the artifact (hidden field, comment, custom property). - When a human edits it, set
human_rewrite=trueand keep a short reason code (wrong_account,tone,missing_field,policy). - Chart rewrite rate next to pass rate and cost per pass, by
job_type. - Sample rewritten traces into the golden set. That is how the evaluator learns.
- Do not celebrate a pass-rate jump that arrived with a rewrite-rate jump.
| Pattern | What it usually means | What you do |
|---|---|---|
| Pass up, rewrite up | Judge got lenient, or grind-to-green | Freeze deploys; inspect criteria |
| Pass flat, rewrite down | Humans gave up and stopped editing | That is not improvement |
| Pass up, rewrite down, cost flat | Actual progress | Widen autonomy a notch |
| Rewrite unmeasured | You are flying on thumbs | Stop. Wire the flag this week |
A high online pass with a rising rewrite rate is not a healthy agent. It is an unpaid editor.
What should policy denials look like in production?
Guardrails that never fire are decoration. Guardrails that always fire are a blocked product. You want a denial catalog, counted, with owners — the same way you count tool errors.
OpenAI’s SDK wraps each check in a guardrail_span. Use that event even if you never open their dashboard. Your runner should emit the same shape: gate name, tripwire yes/no, side-effect that was blocked.
| Denial code | Means | Staging expectation | Prod expectation |
|---|---|---|---|
policy_tool_not_allowed | Tool not on the allowlist | Forced tests hit this | Rare; investigate prompt or schema |
policy_schema_fail | Args failed JSON Schema | Common while pinning | Spike = vendor or prompt drift |
policy_side_effect | Write class not permitted in this state | Forced tests hit this | Should be near-zero if states are right |
policy_destination | Wrong tenant, wrong inbox, wrong env | Forced tests hit this | Page if this appears |
policy_pii | Payload tripped a DLP rule | Sampled tests | Page |
policy_budget | Tokens, dollars, or steps exhausted | Load tests | Page on band break |
policy_human_required | Gate demanded a reviewer | Normal for irreversible writes | Track as load, not as error |
A week of zero denials in production is a smell if you claim to have gates. Either the agent never attempted a write, or the gates are not on the path. Prove the opposite in staging: a deliberate bad tool call must show up as a denial on the scoreboard before you enable the write.
| Check | Pass when |
|---|---|
| Staging denial drill | Known-bad call produces the expected code |
Prod policy_destination | Stays at zero, or you are in an incident |
| Denial drought after a deploy | You re-run the staging drill the same day |
| Human-required rate | Has an owner; not treated as “errors” |
Anthropic’s effective-agents note still starts with the simplest solution that works. If your “agent” is a known path with a policy novel sitting in a prompt, you do not need more monitoring. You need a workflow. The monitoring question only pays once the loop can actually attempt a side effect.
Denials are a heartbeat. Silence is not health.
What do you page on versus review in a meeting?
Do not copy a percentage out of a blog and call it an SLO. I will not hand you a fake 99% / 15-minute / five-point band. Those numbers only mean something after your job_type has a trailing baseline.
Build the page list from deltas, not from folklore.
- Collect two weeks of the five scoreboard numbers in staging, then in a write-limited prod slice.
- Mark the median and a band you can live with — by
job_type, not globally. - Page when a metric leaves that band, or when a never-before-seen
error_code/ denial code appears. - Review the rest in a weekly thirty-minute meeting. If a tile has no owner, delete it.
- Revisit the band after you change tools, models, or criteria. Do not freeze a week-one number forever.
| Event | Page now | Review this week |
|---|---|---|
| Tool-error spike vs baseline | Yes | — |
policy_destination or policy_pii in prod | Yes | — |
| Cost-per-turn band break | Yes if writes are live | Yes if still in shadow mode |
| Rewrite-rate jump with pass-rate jump | After the first day of data | Always |
| New reason code | Yes (catalog gap) | Backfill the catalog |
| Trace store lag / join table down | Yes | — |
| Pretty latency heatmap | No | No — delete or assign |
OpenTelemetry’s sampling note is the right physics: if most requests finish clean, you do not need every trace. Keep 100% of writes, escalates, aborts, budget trips, and human-“wrong” reports. Sample ordinary passes. Never sample “successes only” to save storage — that is how you hide the story.
I am not going to print a keep-rate percentage for your volume. Start with 100% until the bill hurts, then sample passes only.
- Baseline exists before you page
- Bands are per
job_type - Kill switch is in the trace when thrown, not only in Slack
- No copied SLO from another company
Who pages, and when:
| Class | On-call, nights | Business hours |
|---|---|---|
| Live irreversible writes failing | Kill switch + page | Patch tool or schema |
| Cost-band break with writes live | Freeze new runs | Tighten budget or job scope |
| Rewrite-rate jump | No page | Domain reviewer + evaluator owner |
| Prompt wording | Never | Never at 2 a.m. |
| Credential rotation | Page if writes are failing | Rotate on a clock, not a vibe |
Write that split into the runbook before launch. Model-quality tweaks wait for daylight unless an active incident is burning money or customers.
A number you cannot defend in an incident review is not a control. It is a poster.
Failure mode: you monitored the demo
The demo had traces. The demo had a dashboard. Production still shipped nonsense. Monitoring the happy path is how that happens.
Why agent demos fail production owns the harness gaps — schemas, auth, evaluators, kill switches. This section owns the monitoring version of the same lie: you instrumented the staged run and called it ops.
| What you monitored | What broke | What it cost | What you do instead |
|---|---|---|---|
| Staged tools that always return clean JSON | Real vendor fields rename; 200-empty | Bad CRM writes, no error code | Point traces at the live connector in a shadow tenant |
Token chart, no job_type | One job type grinds; the average looks fine | Surprise invoice | Band cost per turn by job |
| Thumbs / CSAT | Silent wrong notes | Sales rewrite tax | Rewrite flag on the receipt |
| Zero denials, “gates exist in the prompt” | Prompt is not a gate | Irreversible send | Forced denial drill in staging |
| Traces without receipt ids | Cannot find the run | Hour-long incident | Join table the same day you enable writes |
| Sampling only successes | Incidents vanish from storage | Recurring “we have no trace” | Keep 100% of writes and aborts |
Soft-launch without a forced failure drill is still a demo with a production URL. The drill is simple: break auth, send a wrong-tenant payload, blow a tiny budget. The scoreboard must move. If it does not, you are not monitoring production. You are monitoring the brochure.
Do not widen autonomy until last week’s near-miss has a reason code you can name.
What can you install in a week?
You do not need a platform bake-off to start. You need the five numbers and a join table. Spurlock Studios puts the thin version on the board during the $1,500 · 5-day pilot on /agentic: structured run logs, tool errors, cost, rewrite capture, denials, and a screen someone actually reads. Full vendor suites can wait. Blindness should not.
Day-one workbook if you are standing this up yourself:
| Day | Ship | Done when |
|---|---|---|
| 1 | run_id on every terminal; redaction at write | You can grep a run without opening a dump |
| 2 | Tool spans + error_code catalog | A forced 401 shows up as tool_auth |
| 3 | Cost per turn (model + paid tools), estimates flagged | You can name yesterday’s dollars for one job_type |
| 4 | Rewrite flag + “report wrong” with run_id | A human edit flips the flag |
| 5 | Denial catalog + staging drill + one ops screen | Known-bad call produces the expected denial |
Checklist that kills the week if any box is empty:
- Receipt id →
run_idlookup works - Five scoreboard tiles have named owners
- Staging denial drill is recorded, not planned
- Kill switch leaves an
abortreason on the trace - Thumbs are not on the ops screen
What not to install in that week:
| Temptation | Why it waits |
|---|---|
| Vendor bake-off (Langfuse vs LangSmith vs “just OTel”) | You do not yet know which fields you emit |
| Copied SLO pack from another team | You have no baseline |
| Embedding / retrieval heatmaps | You are not debugging a write |
| Per-model token leaderboard | job_type is the unit, not the nickname |
| Chat thumbs on the agent UI | Contaminates the rewrite signal |
| Full prompt capture in prod | Privacy review first; hashes until then |
Skip vendor comparison. Skip embedding heatmaps. Skip a twelve-tab “AI hub.” If you only have a week, you are buying a scoreboard, not a museum.
How do you know the monitoring is working?
Monitoring is working when a bad write becomes a named span in minutes, not a mystery in standup. You will not get that from a thumbs widget.
| Test | Pass | Fail |
|---|---|---|
| Sales forwards a bad note | You open the trace in under a minute | You ask them to paste the reply |
| Auth token expires | tool_auth pages | Token chart looks “a bit high” |
Human rewrites a done note | Rewrite rate moves the same day | Pass rate stays green, nobody knows |
| Staging denial drill | Expected code on the board | Zero denials, “must be fine” |
| Budget trip | policy_budget + abort on the trace | Run hangs in act until someone notices |
| Weekly meeting | Five tiles, owners, one escalate package reviewed | Twelve tabs, no decisions |
Run those tests on purpose. Hope is not a monitor.
Rehearse once in staging before you call it production monitoring:
- Deploy a known-bad tool mapping (wrong field name, or expired token).
- Confirm
tool_schemaortool_authmoves the error tile. - Send a wrong-tenant payload. Confirm
policy_destinationdenies it and no write lands. - Mark a
doneartifact as rewritten. Confirm rewrite rate moves. - Throw the kill switch. Confirm the next run terminals as
abortwith a reason code, not a hungact. - From the write system’s audit log, open the
run_idwithout asking Slack.
| Rehearsal miss | What it means |
|---|---|
| Error tile did not move | You are still monitoring the model |
| Denial did not fire | The gate is in a prompt, not on the path |
| Rewrite did not move | The flag is not on the receipt |
| Kill switch left a hung span | The switch is a rumor |
| Could not open the run | The join table is folklore |
If pass rate is the only green tile you can point at, you do not have production monitoring yet. You have a demo metric that survived the meeting.
When is this not worth doing yet?
If you should not have an agent, you should not have an agent dashboard. When not to build an agent is the brake: known path, mushy criteria, tiny volume, or nobody owns the SOP. A workflow with retries and a dead-letter queue is the monitor. Do not buy traces for a graph you can still draw.
| Situation | Monitor this instead | Skip the agent scoreboard |
|---|---|---|
| Known path, rare exceptions | Workflow success / fail / DLQ | Yes |
| One messy field, then deterministic routing | Schema-check fail rate on that step | Yes |
| No evaluator, “we’ll know it when we see it” | Fix the process first | Yes — you have nothing to score |
| Volume is a handful of jobs a week | A human reading the output | Probably — the join table still helps if you insist on a loop |
| Writes are irreversible and unowned | Do not ship the write | Scoreboard will not save you |
| You already have the loop on live tools | The five-number board | No. Do this now |
Decision list before you buy the dashboard:
- Can you draw the path on a whiteboard without a model in the room? Ship a workflow. Monitor success, fail, and the dead-letter queue.
- Is “good” still a taste argument? Stop. Write criteria. There is nothing to monitor but opinions.
- Will this loop attempt a write on real data this month? If no, wait. Shadow-mode traces without a future write are homework.
- Can one person name the evaluator, the sandbox, and the kill switch? If no, you are not ready for an agent scoreboard.
- If yes to a real write, a named evaluator, and a kill switch — install the five numbers this week.
Anthropic will tell you to stay with the simplest solution until complexity pays for itself. Monitoring is complexity. Pay it when the loop can attempt a side effect on real data. Do not pay it to decorate a chatbot.
The parent map is still the operating manual. This page is the operator answer: five numbers, a join table, and a staging drill. Thumbs are a chatbot metric. Agents need a scoreboard.
FAQ
How do I monitor AI agents in production?
With a scoreboard, not a survey. Track traces you can open from a receipt id, tool-error rates, cost per turn, human rewrite rate, and policy denials — by job_type, with named owners. Provider token charts and chatbot thumbs are traffic. If a human cannot paste a run_id and see the tool that wrote, you are not monitoring the agent yet.
How do I measure whether agent monitoring is working?
Run the tests: a bad note opens a trace in under a minute, a forced 401 shows up as tool_auth, a human edit flips the rewrite flag, and a staging denial drill produces the expected code. If pass rate is the only tile that moves, the monitoring is not working. You are watching a vanity percentage.
What usually fails first when teams try this?
The join table and the rewrite flag. Teams emit pretty traces, then cannot get from a CRM note to a run_id, and they never count the human who still edits the artifact. Tool 200-empty and denial droughts are the next misses. Model 500s are usually the last thing that actually hurts.
How long does this take to show results?
A thin scoreboard can exist in five business days if the runner can emit ids, error codes, and cost fields. You will not have a stable baseline in five days. You will have the ability to debug the next bad write. Treat the first two weeks as instrumentation, then set bands from your medians — not from a percentage you copied.
What should I skip if I only have a week?
Skip vendor bake-offs, embedding galleries, copied SLOs, and thumbs widgets. Ship run_id, tool-error catalog, cost per turn, rewrite flag, denial drill, and one screen with owners. That is the week. Platforms come after the join table works.
When is this not worth doing yet?
When you should not have an agent. If the path is known, monitor the workflow. If criteria are mush, you have nothing to score. Build the loop — and this scoreboard — only after a workflow, or a workflow plus one schema-checked LLM step, fails on real traffic.
CTA
Install the five-number scoreboard. Then make someone read it every week.
What questions does this article answer?
- How do I monitor AI agents in production?
- With a scoreboard, not a survey. Track traces you can open from a receipt id, tool-error rates, cost per turn, human rewrite rate, and policy denials — by `job_type`, with named owners. Provider token charts and chatbot thumbs are traffic. If a human cannot paste a `run_id` and see the tool that wrote, you are not monitoring the agent yet.
- How do I measure whether agent monitoring is working?
- Run the tests: a bad note opens a trace in under a minute, a forced 401 shows up as `tool_auth`, a human edit flips the rewrite flag, and a staging denial drill produces the expected code. If pass rate is the only tile that moves, the monitoring is not working. You are watching a vanity percentage.
- What usually fails first when teams try this?
- The join table and the rewrite flag. Teams emit pretty traces, then cannot get from a CRM note to a `run_id`, and they never count the human who still edits the artifact. Tool 200-empty and denial droughts are the next misses. Model 500s are usually the last thing that actually hurts.
- How long does this take to show results?
- A thin scoreboard can exist in five business days if the runner can emit ids, error codes, and cost fields. You will not have a stable baseline in five days. You will have the ability to debug the next bad write. Treat the first two weeks as instrumentation, then set bands from *your* medians — not from a percentage you copied.
- What should I skip if I only have a week?
- Skip vendor bake-offs, embedding galleries, copied SLOs, and thumbs widgets. Ship `run_id`, tool-error catalog, cost per turn, rewrite flag, denial drill, and one screen with owners. That is the week. Platforms come after the join table works.
- When is this not worth doing yet?
- When you should not have an agent. If the path is known, monitor the workflow. If criteria are mush, you have nothing to score. Build the loop — and this scoreboard — only after a workflow, or a workflow plus one schema-checked LLM step, fails on real traffic.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.