Spurlock Studios
Contact
Share LinkedIn X
Two clipped paper packets. Thesis: OBSERVABILITY AGENTS TRACES SCORES DASHBOARD.

If your only view into an agent is a model-provider chart, you will miss the expensive failure: runs that look healthy, spend money, and ship wrong work. Observability for agents is traces with states and tool I/O, evaluator scores on the same timeline, and a weekly screen a human actually reads — not a graveyard of JSON in a bucket.

This spoke sits under the Agentic Systems Operating Manual. It owns the dashboard and the trace. Why pass rate lies owns which numbers may veto a deploy. Do not confuse the two.

The short answer

  • For each run, a human must answer six questions without spelunking: job, dying state, tools and redacted I/O, evaluator verdict, cost and revisions, terminal reason.
  • A useful trace is a tree of spans — run, state, model, tool, eval, terminal — with stable ids and a redaction policy applied at write time.
  • Cost is a first-class span field, not a later spreadsheet. Ingest provider usage when you have it; mark estimates when you do not.
  • Online sampling through the same evaluator is how you catch silent wrongness. HTTP 200 is not a pass.
  • One ops screen, named owners, thirty minutes a week. If a panel has no owner, delete it.

What AI agent observability must show

OpenTelemetry treats a trace as the path of a request and a span as one unit of work with a name, parent, timestamps, attributes, and a status. Agents need that shape plus fields HTTP tracing never asked for: state name, tool side-effect class, evaluator evidence, and budget events.

Microsoft Foundry’s agent-tracing overview is blunt about why chat logs fail here: many steps, order that changes with the input, long payloads, and nesting — a tool that calls another process that calls another tool. Their capture list is the right minimum: inputs and outputs, tool usage, token consumption, duration.

QuestionField that answers itFailure if missing
What job was this?job_id, job_type, tenantYou cannot group or page
Which state did it die in?State span nameYou debug the model, not the loop
Which tools ran?Tool spans + side-effect classDamage hides in a 200
What did they see and write?Redacted args/results or hashesYou cannot reconstruct the write
What did the evaluator say?Verdict + criterion codesYou are tracing activity, not correctness
How much did it cost?Tokens + dollars + revision countFinance finds out in week three
How did it end?done / escalate / abort + reason codeTrends are free text

If any answer requires downloading a raw prompt dump by default, the UX failed.

  • Job and tenant on the root span
  • State names match the runner, not marketing labels
  • Every tool span has a side-effect class
  • Evaluator verdict sits on the same timeline as the last write
  • Terminal reason is a catalog code, not a paragraph

Why provider dashboards miss silent wrongness

Provider dashboards answer “did the API accept the request?” They do not answer “did the CRM note name the right account?” Token burn, latency, and HTTP status are traffic. Silent wrongness is a plausible write that fails a soft criterion nobody watches.

OpenAI’s own agent-eval guidance starts in the same place: grade the trace — model calls, tool calls, guardrails, handoffs — when the unit of work is a loop, not a single completion. That is the opposite of staring at a spend chart.

SignalWhat it provesWhat it hides
Provider 200 / finish reasonThe call completedWrong tool, wrong args, ignored criteria
Token chartSpend this hourCost per passing job
p95 model latencyModel SLOTool timeouts and revision grind
“Agent said done”The worker is confidentThe evaluator never ran
Vanity pass rateA green tileRevision rate, coverage, rewrite by humans — see why pass rate lies

The failure mode: ops trusts the green tile, sales forwards a bad note, and nobody can open the run. Correlation ids are cheaper than a dedicated “AI ops” hire.

What a useful trace looks like

LangSmith’s model is the same physics under different nouns: a run is one unit of work (think span); a trace is the collection of runs for one operation, bound by a trace id. Their docs also cap a trace at 25,000 runs — a reminder that unbounded loops are a telemetry problem, not only a cost problem.

OpenAI’s Agents SDK records the tree you actually need: task, agent, turn, generation, function (tool), guardrail, handoff. Use that as a shopping list even if you never open their dashboard.

run (job_id, tenant, job_type)
  state:intake
  state:plan → model span (model id, tokens, latency)
  state:act  → tool:crm.get → tool:enrichment.lookup
  state:evaluate → mechanical checks + judge spans
  state:revise …
  terminal (reason_code, cost_usd, revision_count)
SpanMust storeMust not store by default
runjob_id, tenant, job_type, trace_idFull customer record
state:*Name, enter/exit, transition reasonChat transcript paste
modelProvider, model id, tokens, latency, cache hitUnredacted prompt forever
toolName, side-effect class, latency, error, arg/result hashRaw credentials, full PII payload
evalVerdict, criterion codes, evidence pointersAnother copy of the artifact
terminaldone / escalate / abort, reason codeA novel

Include stable ids for run, parent, and tool call. Point at artifacts with URIs. Apply redaction at write time, not in a later “we will clean the bucket” ticket.

Do not mark a run successful because the HTTP layer returned 200 if the evaluator failed. Terminal status is evaluator or human authority.

How do you trace LLM tool calls without leaking I/O?

Most business damage is a tool call. Tracing only the model is how you miss the write.

OpenTelemetry’s GenAI conventions now live in a dedicated repository (the old opentelemetry.io GenAI pages are stubs). They remain in Development as of August 2026, which is a reason to own an internal span schema and map outward — not a reason to skip tool spans. The convention’s tool-execution shape is the one to copy: an internal span named around the tool, with tool name and call id, and arguments/results only when policy allows.

Langfuse splits the same idea into observation types: a tool observation holds which tool ran, the arguments, and the return value. That is the I/O you debug. Cost, though, is tracked on generation and embedding observations unless you attach spend to the tool on purpose. If a paid search API or a code-exec sandbox has a meter, record it or it will vanish from unit economics.

OpenAI’s SDK is explicit that generation_span and function_span store inputs and outputs, and that this may be sensitive. Their default is to capture it (trace_include_sensitive_data defaults true). Production should flip that unless legal signed off.

Tool-span fieldWhy it exists
tool.nameGroup failures
side_effect_classread / draft / irreversible_write
duration_msFind the hung connector
error_codeLow-cardinality, not a stack dump
arg_hash / result_hashProve two runs touched the same shape without keeping PII
redacted_previewA few hundred characters for humans
external_receipt_idCRM note id, ticket comment id

Procedure for a new tool:

  1. Classify the side effect before you emit a span.
  2. Hash args and results. Store a redacted preview, not the raw body.
  3. Attach the write system’s receipt id on success.
  4. Set span status from the tool error and the evaluator, not from HTTP alone.
  5. Refuse to ship the tool if the span cannot name the side-effect class.

Hashes let you prove two runs touched the same payload shape without retaining PII forever. Previews let humans debug without opening cold storage.

side_effect_classExamplesTrace keep
readCRM get, search, lookupMay sample
draftInternal note, ticket comment in stagingKeep if later promoted
reversible_writeDraft email unsent, status flip with undoKeep on fail
irreversible_writePublic send, refund, delete, production CRM100%
policy_deniedGate blocked the call100%

If the runner cannot name the class, the tool is not ready for production. “It only posts a note” is still a write.

Where does cost belong on the timeline?

Cost that lands in a monthly export is archaeology. Cost that sits on the run is a control.

Langfuse’s rule is the one to steal: ingested provider usage beats inferred token math, and inferred cost is computed at ingest against the price list you had then. If the provider omits usage, estimate with a documented formula and mark usage_estimated=true so finance does not treat it as gospel.

Paid tools are the hole. Token dashboards never see a $0.04 enrichment call that ran twelve times on a revise loop. Attach that spend to the tool span or you will under-price the job.

Cost fieldSource of truthLie if you skip it
Input / output / cache tokensProvider usage objectYou invent a tokenizer
Model $Price list pinned at ingestLast month’s list on this month’s run
Tool $Vendor meter or your invoice map“AI spend” misses the APIs
RevisionsRunner counterCheap first pass, expensive grind
Cost per passDollars / evaluator-passing runsCheap failures look efficient
Budget eventsControl planeYou cannot explain the abort

Chart cost bands by job_type, not by model name. The question is “what does a passing CRM note cost,” not “how many tokens did Sol use on Tuesday.”

  • Provider usage copied onto the model span
  • Estimates flagged
  • Paid tools have a dollar field
  • Cost per pass uses evaluator pass, not HTTP success
  • Budget trips are first-class events on the same dashboard

Offline scores and online sampling

Offline: golden-set pass rate, cost per pass, revision depth — on every change that could move behavior.

Online: sample production runs through the same evaluator. Chart pass / fail / escalate, silent-fail samples (human overrides), cost bands, and drift after deploys.

OpenAI’s trace-grading workflow is the online half: inspect a representative trace, attach a grader, use the result to change prompts, tools, routing, or guardrails. Then they tell you to move to datasets when you need repeatability. That split is correct. Online sampling catches the live miss. Offline suites make the fix stay fixed.

Alert when online pass rate drops versus the trailing baseline, or when cost per pass spikes. That is how you catch prompt regressions and a rotten retrieval index.

ScoreOfflineOnline
Pass / fail / escalateFull golden setStratified sample
Revision depthEvery suite runSample + 100% of writes
Cost per passSuite batchRolling by job_type
Human overrideLabeler disagreement“Report wrong” + sales forwards
CoverageJob shapes in the setJob shapes seen this week

A high online pass with a rising human-rewrite rate is not a healthy agent. It is a vanity tile. The veto panel lives in why pass rate lies.

Which dashboard will ops actually read?

Keep one screen for the weekly meeting. Anything that needs a data scientist to interpret will not get read. Link out to full traces for incidents.

PanelPurposeOwner
Runs by terminal stateAre we escalating more?On-call lead
Pass rate (online sample)Quality, not vibesEvaluator owner
Cost per passUnit economicsWhoever holds the budget
Top evaluator failure codesWhere to fixDomain reviewer + eng
Kill-switch / budget eventsControl plane healthOn-call
p95 end-to-endSLO, secondaryEng

Label the tiles in business language: “Jobs finished cleanly,” “Jobs sent to a human,” “Jobs stopped for budget or policy,” “Average cost when successful,” “Top reasons for handoff.” Model jargon stays on the drill-down.

Vanity panelWhy it diesReplacement
Tokens by model, no job_typeNobody actsCost per passing job
“AI messages sent”Activity, not outcomesTerminal-state mix
Unowned latency heatmapsPretty, mutep95 with a page
Prompt gallery with no scoresTourismTrace + verdict
Twelve tabs of embeddingsNobody opens tab threeOne screen, six panels

If a panel has no owner, delete it. Orphan metrics create false comfort and burn attention.

Sampling that does not lie

Tracing 100% of runs is fine at low volume and expensive at high volume. OpenTelemetry’s sampling note is the right starting physics: if most requests finish clean, you do not need every trace. Head sampling cannot promise you will keep the errors — that needs tail sampling, or an explicit keep-list. Cost totals and terminal-state counts belong on metrics, which are built to aggregate, not on traces you might drop.

When you sample traces:

  1. Keep 100% of escalate, abort, budget trips, and irreversible writes.
  2. Sample passes, stratified by job_type and tenant size.
  3. Upsample for 48 hours after a deploy.
  4. Always keep a run that a human marked wrong.
  5. Never sample “successes only” to save money.
ClassTrace keep rateWhy
Irreversible write100%Disputes and rollbacks
Escalate / abort / budget100%That is the incident stream
Human “wrong” report100%Golden-set candidates
Ordinary pass5–20%, stratifiedBaseline without drowning storage
Post-deploy window100% for 48hRegressions cluster here

Sampling only successes hides the story. Sampling only errors hides the baseline you need to see drift.

How do you keep one story across handoffs and writes?

If Agent B starts a new unrelated id, you will never reconstruct the story. Propagate one trace_id / run_id through every hop.

The W3C Trace Context recommendation exists because vendor-private ids break at the boundary. traceparent is the portable position in the graph; tracestate is the vendor bag. Foundry’s multi-agent notes are built on that plus OpenTelemetry, and they call out the exact failure: tool arguments and results on execute_tool, plus evaluation events, only help if the parent id survived the hop.

LangSmith’s thread_id is the conversation-shaped version of the same rule: one turn is a trace; many turns are a thread. Business agents need a third join: the write system’s receipt.

HopWhat you propagateWhat you add
Runner → modeltrace_id, run_idModel span
Runner → toolSame ids + tool_call_idSide-effect class, receipt id
Agent A → Agent BSame trace_id; new child spansAgent name, handoff reason
Runner → CRM / ticketsrun_id on the noteexternal_receipt_id on the span
Human escalateSame ids on the packageReviewer id, decision

Store the CRM note id, ticket comment id, or email draft id on the trace. When a salesperson says “the agent wrote nonsense,” you jump to the run in seconds. Without that join, observability is a museum.

OpenAI’s SDK lets you wrap multiple run() calls in one parent trace() so they stay one story. Do the same in your runner even if you are not on that SDK.

Logging hygiene is a security control

Observability that leaks customer data is a security incident with charts.

OWASP’s current GenAI LLM Top 10 (2026, published August 4) still ranks Sensitive Information Disclosure as LLM02. Treat traces as a disclosure path: prompts, tool args, retrieved chunks, and error strings. Foundry’s own guidance matches: do not store secrets in prompts or span attributes; redact before telemetry; apply the same access controls you use for production logs.

LangSmith SaaS retains traces 180 days from ingest unless you copy a case into a dataset. That is a contract fact, not a default you should assume is right for your counsel. Prefer artifact URIs and hashes over duplicating sensitive payloads into a third-party SaaS.

RuleDefaultBreak-glass
Secrets / credentialsNever logRotate, do not “temporarily” print
PII in args/resultsRedact at writeAudit-logged unredact, time-boxed
Full promptsOff in prodOn in staging, or sampled with legal
Debug verbositySeparate flagHours, not weeks
Third-party SaaSHashes + URIsPayload only after a privacy review
RetentionDocumented, job-typedDispute hold, then delete

Before enabling full prompt capture in production, run a privacy review: what PII appears, who can access, retention, export paths. Tenant exports must not dump another tenant’s traces.

How do you run an incident from a trace?

When something bad ships:

  1. Find run ids from the write system’s audit (CRM, email, tickets) — not from a Slack screenshot of the model reply.
  2. Open the trace. Identify the first bad tool call or the first failed criterion that was ignored.
  3. Freeze writes if the pattern is broad (kill switch). The trace should show abort and a reason code, not a hung act.
  4. Patch the evaluator or the sandbox. Add a golden case from this run.
  5. Replay the suite before you re-enable autonomy.
StepDone whenCommon miss
LocateReceipt id → run id < 60sNo join table
DiagnoseFirst bad span namedBlaming “the model”
ContainKill switch in the traceSlack “please stop”
FixCriterion or tool patch + casePrompt-only vibes
ProveSuite green on the new caseRe-enable on hope

Blameless for humans. Ruthless for missing criteria.

On-call owns kill switches, credential rotations, and freeze-writes. Model-quality tweaks wait for business hours unless an active incident is burning money or customers. Write that sentence into the runbook before launch.

  • Kill switch location and who may throw it
  • How to freeze writes without killing reads
  • Where receipt id → run id is joined
  • Redaction break-glass, time-boxed
  • Who pages for cost-band breaches vs quality drops

Reason codes must be a catalog. Dashboards group on these. Free-text reasons make trends impossible.

CodeMeansTypical next action
eval_passEvaluator acceptedCount as a pass; sample the trace
eval_fail_exhaustedRevisions hit the ceilingAdd a golden case; inspect last tool
budget_exhaustedTokens, dollars, or stepsTighten the cage or the job
tool_auth_errorCredential or scope missRotate or shrink the allowlist
policy_violationGate denied a side effectCriteria or tool args
human_rejectReviewer said noTeach the evaluator
timeoutNo terminal in the SLOFind the hung span
out_of_scopeJob was not this jobIntake filter

When you compare prompt A vs B, run the same golden set, same tool stubs, same budget. Report pass rate, cost, latency. Store the batch under an experiment id on the root span. Intuition-only prompt merges are how regressions ship.

What vendor LLM observability still gets wrong

Many tools stop at prompt/response capture. Necessary. Insufficient.

Business agents need state names, tool side-effect classes, evaluator verdicts, and budget events in the same timeline. Buy or build toward that model. Do not confuse token charts with ops readiness.

Vendor surfaceUseful forNot enough because
Provider token / latency chartsSpend and SLONo tools, no evaluator, no job_type
SDK traces (OpenAI, others)Fast tree of generations and functionsDefaults often keep raw I/O; your states may be missing
LangSmith / Langfuse style appsRun trees, online evals, cost on generationsYou still have to emit side-effect class and reason codes
Foundry / App InsightsOTel-shaped agent tracesYou still own redaction and the weekly ritual
Homegrown JSON in a bucketCheap at ten runs a dayNobody opens it during an incident

Normalize provider-specific usage fields into your span schema. You will switch models. Your dashboards should not require a rewrite each time. Store raw provider payloads as optional debug attachments with stricter retention.

A “report wrong output” control that captures run_id is worth more than three vanity charts. Route those reports into the weekly review and into the golden set.

Weekly ritual and the vanity-dashboard failure

Ritual beats platform. Thirty minutes, named owners, same six panels.

  1. Glance terminal-state mix.
  2. Open the top three failure codes. Assign a fix owner or accept the rate.
  3. Check cost per pass against the band.
  4. Review one escalate package end to end — criteria, tool I/O, receipt id.
  5. Note kill-switch or budget events. If none ever fire, the switch is a rumor.
FailureWhat breaksWhat it costsWhat you do instead
Logs only on errorNo baselineYou cannot see driftKeep stratified passes
Full prompts in SlackRetention and accessA disclosure plus a useless archiveTrace backend, redacted
Metrics with no ownerMute tilesAttention and false comfortDelete or assign
Tracing only the modelMissed writesCustomer-facing nonsenseTool spans + receipts
Sampling only successesHidden failsRecurring incidents100% writes and aborts
Pass-rate-only screenGreen lieBad deploysPair with why pass rate lies

Example SLOs you can actually page: 99% of runs reach a terminal state within 15 minutes; under 1% abort for unknown errors; online pass rate stays within five points of offline. SLOs make observability a control, not a gallery.

Spurlock Studios installs the thin version during the $1,500 · 5-day pilot on /agentic: structured run logs, evaluator verdicts, cost, terminal reason, and the six-panel screen. Full platforms can wait. Blindness should not. The parent map is the operating manual.

Day-one workbook if you are standing this up yourself:

DayShip
1Span schema: run, state, model, tool, eval, terminal
2Emit spans from the runner with redaction and receipt ids
3One dashboard, six panels, named owners
4Alerts for budget trips and pass-rate drops
5Rehearse an incident with a deliberate bad deploy in staging

You do not need a perfect platform to start. You need correlated run ids and evaluator scores beside tool calls.

Add a “report wrong output” control that captures run_id on day one. That button is worth more than three vanity charts. Route reports into the weekly review and into golden-set candidates.

Correlate CRM write ids to run ids the same day — even before pretty charts. When sales forwards a bad note, you should open the trace in under a minute. If you cannot, the dashboard is decoration.

Traces, tool I/O, and cost. Not vanity.

FAQ

What is AI agent observability?

It is recording and reviewing each run with enough structure — states, model calls, redacted tool I/O, evaluator scores, cost, and terminal reasons — to debug, govern, and improve agents in production. Provider token charts are traffic. Observability is the story of one job from intake to done, escalate, or abort. If a human cannot answer those questions without a raw dump, you do not have it yet.

How do you trace LLM tool calls well?

Create a span for each model and tool invocation under a stable run id. Record tool name, side-effect class, latency, error codes, hashes, and a redacted preview — not raw credentials or full PII. Align terminal status with the evaluator, not with HTTP 200. Paid tools need their own dollar field or they disappear from unit economics.

Which metrics matter most week to week?

Online pass rate, escalate rate, cost per passing run, top evaluator failure codes, and budget or kill-switch events. Latency matters after correctness and cost. A green pass rate with a rising human-rewrite rate is a vanity tile; pair the screen with the veto metrics in why pass rate lies.

Do we need a special vendor on day one?

Not always. Structured logs plus a six-panel dashboard can cover a pilot. As volume grows, specialized tracing tools help if they ingest tool events, evaluator verdicts, and budget trips — not only prompts. Own the span schema either way. Vendor conventions move; your incident questions do not.

How does Spurlock Studios handle observability?

Thin traces and scores ship in the $1,500 · 5-day pilot: run ids, redacted tool I/O, evaluator verdicts, cost, and terminal reason, plus the weekly screen. Richer operator platforms land in fuller builds. The stack is described in the operating manual. Packaging lives on /agentic.

How does observability relate to evaluators?

Evaluators produce the quality signal. Observability stores that signal beside cost and tool I/O so a human can act. Without evaluators you are tracing activity, not correctness. Without traces, evaluator scores are a tile with no story. You need both on one timeline.

CTA

Install the minimum signal. Then make someone read it every week.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

What is AI agent observability?
It is recording and reviewing each run with enough structure — states, model calls, redacted tool I/O, evaluator scores, cost, and terminal reasons — to debug, govern, and improve agents in production. Provider token charts are traffic. Observability is the story of one job from intake to `done`, `escalate`, or `abort`. If a human cannot answer those questions without a raw dump, you do not have it yet.
How do you trace LLM tool calls well?
Create a span for each model and tool invocation under a stable run id. Record tool name, side-effect class, latency, error codes, hashes, and a redacted preview — not raw credentials or full PII. Align terminal status with the evaluator, not with HTTP 200. Paid tools need their own dollar field or they disappear from unit economics.
Which metrics matter most week to week?
Online pass rate, escalate rate, cost per passing run, top evaluator failure codes, and budget or kill-switch events. Latency matters after correctness and cost. A green pass rate with a rising human-rewrite rate is a vanity tile; pair the screen with the veto metrics in [why pass rate lies](/blog/why-pass-rate-lies).
Do we need a special vendor on day one?
Not always. Structured logs plus a six-panel dashboard can cover a pilot. As volume grows, specialized tracing tools help if they ingest tool events, evaluator verdicts, and budget trips — not only prompts. Own the span schema either way. Vendor conventions move; your incident questions do not.
How does Spurlock Studios handle observability?
Thin traces and scores ship in the **$1,500 · 5-day** pilot: run ids, redacted tool I/O, evaluator verdicts, cost, and terminal reason, plus the weekly screen. Richer operator platforms land in fuller builds. The stack is described in the [operating manual](/blog/agentic-systems-operating-manual). Packaging lives on [/agentic](/agentic).
How does observability relate to evaluators?
Evaluators produce the quality signal. Observability stores that signal beside cost and tool I/O so a human can act. Without evaluators you are tracing activity, not correctness. Without traces, evaluator scores are a tile with no story. You need both on one timeline.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot