Spurlock Studios
Contact
Share LinkedIn X
Two clipped paper packets. Thesis: MONITOR AI AGENTS PRODUCTION.

You monitor AI agents in production with a scoreboard, not a chatbot survey. The numbers that matter are traces you can open, tool-error rates, cost per turn, human rewrite rate, and policy denials. Thumbs-up on a chat bubble hides the expensive failure: a run that spent money, wrote to the CRM, and still needed a human to fix the note.

This spoke sits under the Agentic Systems Operating Manual. If you still need a workflow instead of a loop, stop at when not to build an agent. If the loop only ever worked on staged tools, read why agent demos fail production before you trust a green tile.

The short answer

  • Every production run needs a run_id a human can paste, a tool-error code, a dollar figure for that turn, a rewrite flag, and a policy verdict.
  • Chatbot CSAT, “messages sent,” and provider token charts are traffic. They do not tell you whether the write was right.
  • Tool errors and policy denials are the control-plane pulse. A week of zeroes usually means the gates are not wired, not that the agent is healthy.
  • Cost belongs per turn and per passing job, by job_type. Weekly token totals arrive too late to stop a grind.
  • Page on deltas from your baseline. Do not invent a 99% SLO because a vendor dashboard had a blank for one.

What does production monitoring mean for an agent?

Monitoring a chatbot asks “did the reply feel okay?” Monitoring an agent asks “did this job finish, with which tools, at what cost, under which policy, and did a human still have to rewrite the artifact?”

OpenAI’s eval guidance is explicit that the unit of work is the trace — model calls, tool calls, guardrails, and handoffs — not a single completion. LangSmith uses the same physics under different nouns: a run is one unit of work; a trace is the collection of runs for one operation. If you cannot open that tree, you are not monitoring the agent. You are monitoring the model vendor.

QuestionChatbot viewAgent view
Did it work?Thumbs / CSATEvaluator pass and no human rewrite
What broke?“The model was weird”Tool error code or policy denial code
What did it cost?Tokens this weekDollars per turn, by job_type
Can we debug it?Paste the reply into SlackPaste run_id, open the trace
Should we keep writing?VibesKill switch + denial rate

I have spent 20,000+ hours on agentic systems and built 500+ automations. The agents that survived contact with a live CRM all had those five numbers on one screen. The ones that did not were demos with a production URL.

  • One run_id on every terminal (done / escalate / abort)
  • Tool spans with an error code, not a stack novel
  • Dollars on the turn, flagged if estimated
  • A rewrite / reject flag a human can set
  • Policy denials counted even when the write never fired

If any box is empty, you have logs. Not monitoring.

What is the five-number scoreboard?

Five numbers. Named owners. Same screen every week. Everything else is a drill-down.

NumberDefinitionHealthy shapeLie if you skip it
Traces openedShare of incidents where ops opened the run_id in under a minuteYou can actually debugYou are guessing from Slack screenshots
Tool-error rateTool spans with a non-empty error_code, per job_typeSpikes are investigatedYou blame “the model” for a 401
Cost per turnModel $ + paid-tool $ for one loop iterationBand by job_typeCheap failures look efficient
Human rewrite ratedone runs a human still edited before the artifact shippedTracked beside passPass rate greens a grind
Policy denialsGate blocked a side effect, with a catalog codeNon-zero in staging; explained in prodGates are theater

OpenTelemetry’s split is the one to steal: traces tell the story of one job; metrics are what you aggregate and page on. Do not page off a trace you might have sampled away. Do not debug an incident off a counter with no run_id.

Scoreboard tileSignal typeOwner
Tool-error rateMetricOn-call
Cost per turn / cost per passMetricWhoever holds the budget
Human rewrite rateMetric + sample of tracesDomain reviewer
Policy denialsMetricPolicy owner
Trace jump (receipt → run)Join table, not a chartEng

How to compute each number without a data team:

  1. Traces opened: incidents this week where receipt → run_id succeeded, divided by incidents. If the denominator is zero, you did not have an incident — or you did not log one.
  2. Tool-error rate: tool spans with error_code set / tool spans, by tool.name and job_type. Exclude policy_* codes; those belong on the denial tile.
  3. Cost per turn: sum of model $ and tool $ between state:plan enter and the next eval. If you cannot bound a turn, you cannot bound a grind.
  4. Human rewrite rate: human_rewrite=true on done receipts / done receipts. Escalate packages that a reviewer rejects are a different tile — do not mix them.
  5. Policy denials: spans with policy_* / attempted side effects. A denial on a read-only lookup is a mis-fired gate, not a win.
Anti-patternWhat it does to the board
Mixing eval-fail into tool-errorYou page the model for a connector outage
Mixing escalate-reject into rewriteYou punish the honest handoff
Global average across job_typeOne cheap job hides one expensive grind
Counting denials as errorsOn-call fights the control plane

Thumbs do not appear on this board. If someone wants a satisfaction widget, put it on the chatbot. Agents write to systems. Monitor the write.

Why are chatbot thumbs not monitoring?

A thumb is a feeling about a sentence. An agent’s damage is a side effect: the wrong account, the wrong refund, the public email that should have stayed a draft. Nobody thumbs-down a CRM note they never saw.

Chatbot metricWhat it measuresWhat it hides on an agent
Thumbs up / downTone of the last messageWrong tool, wrong args, skipped eval
CSAT / “Was this helpful?”SentimentSilent wrong writes
Messages sentActivityTerminal mix (done / escalate / abort)
Provider 200 / finish reasonThe API accepted the call200 with an empty body
Token chartSpend this hourCost per passing turn
“Agent said done”Worker confidenceEvaluator never ran; human still rewrote

The failure mode is familiar: ops watches a green CSAT tile, sales forwards a bad note three days later, and nobody can find the run. Correlation ids are cheaper than a dedicated “AI ops” hire.

OpenAI’s Agents SDK already records guardrail spans beside function spans. If your dashboard cannot show those events, you are not using the trace. You are using a chat log with extra fields.

  • Remove thumbs from the agent ops screen
  • Keep a “report wrong” control that captures run_id
  • Route those reports into rewrite rate, not into a CSAT export
  • Stop calling HTTP 200 a pass

Feelings are not a control loop.

How do you wire traces you can actually open?

A useful trace is not a JSON bucket. It is a tree a human can walk in one sitting: job, states, model calls, tools, eval, terminal. OpenTelemetry calls that path a trace and each unit of work a span. You still have to put agent fields on those spans — state name, side-effect class, reason code — because HTTP tracing never asked for them.

Procedure for a new job type:

  1. Mint run_id / trace_id at intake. Never let a child agent start a second unrelated id.
  2. Emit a span for every model call and every tool call. Name the tool; do not bury it in the prompt dump.
  3. Attach the write system’s receipt (crm_note_id, ticket id, draft id) on success.
  4. Store a catalog reason_code on the terminal. Free text cannot group.
  5. Redact args and results at write time. OWASP still ranks Sensitive Information Disclosure as LLM02 — traces are a disclosure path.
  6. Prove the join: from a CRM note, open the run in under a minute. If you cannot, the trace is decoration.
Must storeMust not store by default
run_id, job_type, tenantFull customer record
Tool name, error_code, latencyRaw credentials
Side-effect classUnredacted PII payloads
Evaluator verdict + criterion codesA second copy of the artifact
Cost fields, revision countLast month’s price list applied blindly
Terminal reason codeA paragraph

LangSmith will happily keep a thread of many traces. Business agents need a third join the vendor will not invent for you: the receipt in the system of record.

  • Receipt id → run_id is a real lookup, not a folklore Slack search
  • Redaction policy applied at write, not “we will clean the bucket”
  • Staging can dump more; production defaults to hashes plus a short preview

If the only way to debug is downloading a prompt dump, the UX failed.

What do tool errors tell you that model charts miss?

Most business damage is a tool call. Tracing only the model is how you miss the write. A provider latency chart will not tell you the CRM returned 200 with an empty body, or that the allowlisted “update note” tool started 401-ing at 2 a.m. when a token expired.

error_codeMeansTypical next action
tool_timeoutConnector hungBound the wait; escalate the run
tool_auth401 / 403 / scope missRotate or shrink the allowlist
tool_schemaArgs failed the contractPin the schema; stop guessing fields
tool_empty_200Success status, empty bodyTreat as fail; do not eval an empty write
tool_rate_limitVendor 429Backoff with a ceiling; do not spin
tool_conflictIdempotency or unique-key clashReuse the receipt; do not double-write
tool_not_allowlistedModel invented a toolDenial, not a retry

Count rates by tool.name and job_type. A 3% error rate on crm.get is a different incident than a 3% rate on crm.refund.

SignalPage?Why
Spike vs trailing baseline for one toolYesConnector or auth is on fire
New error_code you have never seenYesCatalog gap or vendor change
Steady low-rate tool_empty_200Review this weekSilent wrongness factory
Model “sorry” text with no tool errorNoThat is a prompt; look at eval fail instead

OpenAI’s trace-grading questions start in the same place: did the agent pick the right tool, and did a guardrail fire when it should have? Grade the tool span. Do not grade the apology.

Procedure when you add a tool:

  1. Give it a stable tool.name. Do not rename it every prompt tweak.
  2. Classify the side effect (read / draft / irreversible_write) before the first production call.
  3. Map vendor failures onto the catalog above. Raw HTTP status is not a catalog.
  4. Emit error_code on fail and on 200-empty. Status is not truth.
  5. Refuse to ship the tool if the span cannot name the side-effect class.
  • New tool has a catalog mapping, not “see stack dump”
  • Shadow tenant traffic hits the live connector shape
  • On-call knows which tile moves when this tool dies

The miss I still see: teams alert on model 500s and ignore tool 200s. The 200s are where the money went.

How should you track cost per turn?

Cost that lands in a monthly export is archaeology. Cost that sits on the turn is a control.

A turn is one pass through plan → act → eval (and maybe revise). Cost per turn is model dollars plus paid-tool dollars for that pass. Cost per pass — dollars divided by evaluator-passing jobs — is the number finance can defend. Cheap failures make cost-per-run look efficient. Do not use that one.

Langfuse’s rule is the one to copy: ingested provider usage beats inferred token math, and inferred cost is computed at ingest against the price list you had then. Their docs are also blunt that only generation and embedding observations carry usage. Paid search, enrichment, and code-exec sandboxes will vanish unless you attach a dollar field to the tool span yourself.

Cost fieldSource of truthLie if you skip it
Input / output / cache tokensProvider usage objectYou invent a tokenizer
Model $Price list pinned at ingestLast month’s list on this month’s run
Tool $Vendor meter or invoice map“AI spend” misses the APIs
RevisionsRunner counterCheap first pass, expensive grind
Cost per turnSum for one loop iterationYou cannot see the grind
Cost per passDollars / evaluator-passing jobsFailures look cheap

Chart bands by job_type, not by model nickname. The question is “what does a passing CRM note cost,” not “how many tokens did the flagship burn on Tuesday.”

  • Provider usage copied onto the model span
  • Estimates flagged usage_estimated=true
  • Paid tools have a dollar field
  • Cost per pass uses evaluator pass, not HTTP success
  • Budget trips are events on the same board

If you cannot name this week’s cost per passing job, you are not monitoring unit economics. You are watching a token fireplace.

Why is human rewrite rate the honesty metric?

Pass rate can green while sales still rewrites half the notes. That is not a healthy agent. That is a vanity tile with a human in the loop you forgot to count.

Rewrite rate is simple: of the runs that ended done, what share did a human still edit, reject, or re-send before the artifact shipped? Capture it as a flag on the receipt, not as a feeling in standup.

Capture methodStrengthWeakness
“Report wrong” control with run_idFast, tied to the traceSilent rewrites never click it
CRM edited-after-agent timestampCatches silent fixesNeed a join; timezone lies
Reviewer decision on an escalate packageClean labelsMisses the done path
Sales forward of a bad noteHigh-signalSlow; incomplete

Procedure:

  1. Put run_id on the artifact (hidden field, comment, custom property).
  2. When a human edits it, set human_rewrite=true and keep a short reason code (wrong_account, tone, missing_field, policy).
  3. Chart rewrite rate next to pass rate and cost per pass, by job_type.
  4. Sample rewritten traces into the golden set. That is how the evaluator learns.
  5. Do not celebrate a pass-rate jump that arrived with a rewrite-rate jump.
PatternWhat it usually meansWhat you do
Pass up, rewrite upJudge got lenient, or grind-to-greenFreeze deploys; inspect criteria
Pass flat, rewrite downHumans gave up and stopped editingThat is not improvement
Pass up, rewrite down, cost flatActual progressWiden autonomy a notch
Rewrite unmeasuredYou are flying on thumbsStop. Wire the flag this week

A high online pass with a rising rewrite rate is not a healthy agent. It is an unpaid editor.

What should policy denials look like in production?

Guardrails that never fire are decoration. Guardrails that always fire are a blocked product. You want a denial catalog, counted, with owners — the same way you count tool errors.

OpenAI’s SDK wraps each check in a guardrail_span. Use that event even if you never open their dashboard. Your runner should emit the same shape: gate name, tripwire yes/no, side-effect that was blocked.

Denial codeMeansStaging expectationProd expectation
policy_tool_not_allowedTool not on the allowlistForced tests hit thisRare; investigate prompt or schema
policy_schema_failArgs failed JSON SchemaCommon while pinningSpike = vendor or prompt drift
policy_side_effectWrite class not permitted in this stateForced tests hit thisShould be near-zero if states are right
policy_destinationWrong tenant, wrong inbox, wrong envForced tests hit thisPage if this appears
policy_piiPayload tripped a DLP ruleSampled testsPage
policy_budgetTokens, dollars, or steps exhaustedLoad testsPage on band break
policy_human_requiredGate demanded a reviewerNormal for irreversible writesTrack as load, not as error

A week of zero denials in production is a smell if you claim to have gates. Either the agent never attempted a write, or the gates are not on the path. Prove the opposite in staging: a deliberate bad tool call must show up as a denial on the scoreboard before you enable the write.

CheckPass when
Staging denial drillKnown-bad call produces the expected code
Prod policy_destinationStays at zero, or you are in an incident
Denial drought after a deployYou re-run the staging drill the same day
Human-required rateHas an owner; not treated as “errors”

Anthropic’s effective-agents note still starts with the simplest solution that works. If your “agent” is a known path with a policy novel sitting in a prompt, you do not need more monitoring. You need a workflow. The monitoring question only pays once the loop can actually attempt a side effect.

Denials are a heartbeat. Silence is not health.

What do you page on versus review in a meeting?

Do not copy a percentage out of a blog and call it an SLO. I will not hand you a fake 99% / 15-minute / five-point band. Those numbers only mean something after your job_type has a trailing baseline.

Build the page list from deltas, not from folklore.

  1. Collect two weeks of the five scoreboard numbers in staging, then in a write-limited prod slice.
  2. Mark the median and a band you can live with — by job_type, not globally.
  3. Page when a metric leaves that band, or when a never-before-seen error_code / denial code appears.
  4. Review the rest in a weekly thirty-minute meeting. If a tile has no owner, delete it.
  5. Revisit the band after you change tools, models, or criteria. Do not freeze a week-one number forever.
EventPage nowReview this week
Tool-error spike vs baselineYes—
policy_destination or policy_pii in prodYes—
Cost-per-turn band breakYes if writes are liveYes if still in shadow mode
Rewrite-rate jump with pass-rate jumpAfter the first day of dataAlways
New reason codeYes (catalog gap)Backfill the catalog
Trace store lag / join table downYes—
Pretty latency heatmapNoNo — delete or assign

OpenTelemetry’s sampling note is the right physics: if most requests finish clean, you do not need every trace. Keep 100% of writes, escalates, aborts, budget trips, and human-“wrong” reports. Sample ordinary passes. Never sample “successes only” to save storage — that is how you hide the story.

I am not going to print a keep-rate percentage for your volume. Start with 100% until the bill hurts, then sample passes only.

  • Baseline exists before you page
  • Bands are per job_type
  • Kill switch is in the trace when thrown, not only in Slack
  • No copied SLO from another company

Who pages, and when:

ClassOn-call, nightsBusiness hours
Live irreversible writes failingKill switch + pagePatch tool or schema
Cost-band break with writes liveFreeze new runsTighten budget or job scope
Rewrite-rate jumpNo pageDomain reviewer + evaluator owner
Prompt wordingNeverNever at 2 a.m.
Credential rotationPage if writes are failingRotate on a clock, not a vibe

Write that split into the runbook before launch. Model-quality tweaks wait for daylight unless an active incident is burning money or customers.

A number you cannot defend in an incident review is not a control. It is a poster.

Failure mode: you monitored the demo

The demo had traces. The demo had a dashboard. Production still shipped nonsense. Monitoring the happy path is how that happens.

Why agent demos fail production owns the harness gaps — schemas, auth, evaluators, kill switches. This section owns the monitoring version of the same lie: you instrumented the staged run and called it ops.

What you monitoredWhat brokeWhat it costWhat you do instead
Staged tools that always return clean JSONReal vendor fields rename; 200-emptyBad CRM writes, no error codePoint traces at the live connector in a shadow tenant
Token chart, no job_typeOne job type grinds; the average looks fineSurprise invoiceBand cost per turn by job
Thumbs / CSATSilent wrong notesSales rewrite taxRewrite flag on the receipt
Zero denials, “gates exist in the prompt”Prompt is not a gateIrreversible sendForced denial drill in staging
Traces without receipt idsCannot find the runHour-long incidentJoin table the same day you enable writes
Sampling only successesIncidents vanish from storageRecurring “we have no trace”Keep 100% of writes and aborts

Soft-launch without a forced failure drill is still a demo with a production URL. The drill is simple: break auth, send a wrong-tenant payload, blow a tiny budget. The scoreboard must move. If it does not, you are not monitoring production. You are monitoring the brochure.

Do not widen autonomy until last week’s near-miss has a reason code you can name.

What can you install in a week?

You do not need a platform bake-off to start. You need the five numbers and a join table. Spurlock Studios puts the thin version on the board during the $1,500 · 5-day pilot on /agentic: structured run logs, tool errors, cost, rewrite capture, denials, and a screen someone actually reads. Full vendor suites can wait. Blindness should not.

Day-one workbook if you are standing this up yourself:

DayShipDone when
1run_id on every terminal; redaction at writeYou can grep a run without opening a dump
2Tool spans + error_code catalogA forced 401 shows up as tool_auth
3Cost per turn (model + paid tools), estimates flaggedYou can name yesterday’s dollars for one job_type
4Rewrite flag + “report wrong” with run_idA human edit flips the flag
5Denial catalog + staging drill + one ops screenKnown-bad call produces the expected denial

Checklist that kills the week if any box is empty:

  • Receipt id → run_id lookup works
  • Five scoreboard tiles have named owners
  • Staging denial drill is recorded, not planned
  • Kill switch leaves an abort reason on the trace
  • Thumbs are not on the ops screen

What not to install in that week:

TemptationWhy it waits
Vendor bake-off (Langfuse vs LangSmith vs “just OTel”)You do not yet know which fields you emit
Copied SLO pack from another teamYou have no baseline
Embedding / retrieval heatmapsYou are not debugging a write
Per-model token leaderboardjob_type is the unit, not the nickname
Chat thumbs on the agent UIContaminates the rewrite signal
Full prompt capture in prodPrivacy review first; hashes until then

Skip vendor comparison. Skip embedding heatmaps. Skip a twelve-tab “AI hub.” If you only have a week, you are buying a scoreboard, not a museum.

How do you know the monitoring is working?

Monitoring is working when a bad write becomes a named span in minutes, not a mystery in standup. You will not get that from a thumbs widget.

TestPassFail
Sales forwards a bad noteYou open the trace in under a minuteYou ask them to paste the reply
Auth token expirestool_auth pagesToken chart looks “a bit high”
Human rewrites a done noteRewrite rate moves the same dayPass rate stays green, nobody knows
Staging denial drillExpected code on the boardZero denials, “must be fine”
Budget trippolicy_budget + abort on the traceRun hangs in act until someone notices
Weekly meetingFive tiles, owners, one escalate package reviewedTwelve tabs, no decisions

Run those tests on purpose. Hope is not a monitor.

Rehearse once in staging before you call it production monitoring:

  1. Deploy a known-bad tool mapping (wrong field name, or expired token).
  2. Confirm tool_schema or tool_auth moves the error tile.
  3. Send a wrong-tenant payload. Confirm policy_destination denies it and no write lands.
  4. Mark a done artifact as rewritten. Confirm rewrite rate moves.
  5. Throw the kill switch. Confirm the next run terminals as abort with a reason code, not a hung act.
  6. From the write system’s audit log, open the run_id without asking Slack.
Rehearsal missWhat it means
Error tile did not moveYou are still monitoring the model
Denial did not fireThe gate is in a prompt, not on the path
Rewrite did not moveThe flag is not on the receipt
Kill switch left a hung spanThe switch is a rumor
Could not open the runThe join table is folklore

If pass rate is the only green tile you can point at, you do not have production monitoring yet. You have a demo metric that survived the meeting.

When is this not worth doing yet?

If you should not have an agent, you should not have an agent dashboard. When not to build an agent is the brake: known path, mushy criteria, tiny volume, or nobody owns the SOP. A workflow with retries and a dead-letter queue is the monitor. Do not buy traces for a graph you can still draw.

SituationMonitor this insteadSkip the agent scoreboard
Known path, rare exceptionsWorkflow success / fail / DLQYes
One messy field, then deterministic routingSchema-check fail rate on that stepYes
No evaluator, “we’ll know it when we see it”Fix the process firstYes — you have nothing to score
Volume is a handful of jobs a weekA human reading the outputProbably — the join table still helps if you insist on a loop
Writes are irreversible and unownedDo not ship the writeScoreboard will not save you
You already have the loop on live toolsThe five-number boardNo. Do this now

Decision list before you buy the dashboard:

  1. Can you draw the path on a whiteboard without a model in the room? Ship a workflow. Monitor success, fail, and the dead-letter queue.
  2. Is “good” still a taste argument? Stop. Write criteria. There is nothing to monitor but opinions.
  3. Will this loop attempt a write on real data this month? If no, wait. Shadow-mode traces without a future write are homework.
  4. Can one person name the evaluator, the sandbox, and the kill switch? If no, you are not ready for an agent scoreboard.
  5. If yes to a real write, a named evaluator, and a kill switch — install the five numbers this week.

Anthropic will tell you to stay with the simplest solution until complexity pays for itself. Monitoring is complexity. Pay it when the loop can attempt a side effect on real data. Do not pay it to decorate a chatbot.

The parent map is still the operating manual. This page is the operator answer: five numbers, a join table, and a staging drill. Thumbs are a chatbot metric. Agents need a scoreboard.

FAQ

How do I monitor AI agents in production?

With a scoreboard, not a survey. Track traces you can open from a receipt id, tool-error rates, cost per turn, human rewrite rate, and policy denials — by job_type, with named owners. Provider token charts and chatbot thumbs are traffic. If a human cannot paste a run_id and see the tool that wrote, you are not monitoring the agent yet.

How do I measure whether agent monitoring is working?

Run the tests: a bad note opens a trace in under a minute, a forced 401 shows up as tool_auth, a human edit flips the rewrite flag, and a staging denial drill produces the expected code. If pass rate is the only tile that moves, the monitoring is not working. You are watching a vanity percentage.

What usually fails first when teams try this?

The join table and the rewrite flag. Teams emit pretty traces, then cannot get from a CRM note to a run_id, and they never count the human who still edits the artifact. Tool 200-empty and denial droughts are the next misses. Model 500s are usually the last thing that actually hurts.

How long does this take to show results?

A thin scoreboard can exist in five business days if the runner can emit ids, error codes, and cost fields. You will not have a stable baseline in five days. You will have the ability to debug the next bad write. Treat the first two weeks as instrumentation, then set bands from your medians — not from a percentage you copied.

What should I skip if I only have a week?

Skip vendor bake-offs, embedding galleries, copied SLOs, and thumbs widgets. Ship run_id, tool-error catalog, cost per turn, rewrite flag, denial drill, and one screen with owners. That is the week. Platforms come after the join table works.

When is this not worth doing yet?

When you should not have an agent. If the path is known, monitor the workflow. If criteria are mush, you have nothing to score. Build the loop — and this scoreboard — only after a workflow, or a workflow plus one schema-checked LLM step, fails on real traffic.

CTA

Install the five-number scoreboard. Then make someone read it every week.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

How do I monitor AI agents in production?
With a scoreboard, not a survey. Track traces you can open from a receipt id, tool-error rates, cost per turn, human rewrite rate, and policy denials — by `job_type`, with named owners. Provider token charts and chatbot thumbs are traffic. If a human cannot paste a `run_id` and see the tool that wrote, you are not monitoring the agent yet.
How do I measure whether agent monitoring is working?
Run the tests: a bad note opens a trace in under a minute, a forced 401 shows up as `tool_auth`, a human edit flips the rewrite flag, and a staging denial drill produces the expected code. If pass rate is the only tile that moves, the monitoring is not working. You are watching a vanity percentage.
What usually fails first when teams try this?
The join table and the rewrite flag. Teams emit pretty traces, then cannot get from a CRM note to a `run_id`, and they never count the human who still edits the artifact. Tool 200-empty and denial droughts are the next misses. Model 500s are usually the last thing that actually hurts.
How long does this take to show results?
A thin scoreboard can exist in five business days if the runner can emit ids, error codes, and cost fields. You will not have a stable baseline in five days. You will have the ability to debug the next bad write. Treat the first two weeks as instrumentation, then set bands from *your* medians — not from a percentage you copied.
What should I skip if I only have a week?
Skip vendor bake-offs, embedding galleries, copied SLOs, and thumbs widgets. Ship `run_id`, tool-error catalog, cost per turn, rewrite flag, denial drill, and one screen with owners. That is the week. Platforms come after the join table works.
When is this not worth doing yet?
When you should not have an agent. If the path is known, monitor the workflow. If criteria are mush, you have nothing to score. Build the loop — and this scoreboard — only after a workflow, or a workflow plus one schema-checked LLM step, fails on real traffic.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot