Observability for Agents: Traces, Scores, and the Dashboard Ops Actually Reads
Provider dashboards miss silent wrongness. Agent observability is traces with redacted tool I/O, evaluator scores, and cost — the screen ops actually reads.
William Spurlock Founder — Spurlock Studios Updated 18 MIN
If your only view into an agent is a model-provider chart, you will miss the expensive failure: runs that look healthy, spend money, and ship wrong work. Observability for agents is traces with states and tool I/O, evaluator scores on the same timeline, and a weekly screen a human actually reads — not a graveyard of JSON in a bucket.
This spoke sits under the Agentic Systems Operating Manual. It owns the dashboard and the trace. Why pass rate lies owns which numbers may veto a deploy. Do not confuse the two.
The short answer
- For each run, a human must answer six questions without spelunking: job, dying state, tools and redacted I/O, evaluator verdict, cost and revisions, terminal reason.
- A useful trace is a tree of spans — run, state, model, tool, eval, terminal — with stable ids and a redaction policy applied at write time.
- Cost is a first-class span field, not a later spreadsheet. Ingest provider usage when you have it; mark estimates when you do not.
- Online sampling through the same evaluator is how you catch silent wrongness. HTTP 200 is not a pass.
- One ops screen, named owners, thirty minutes a week. If a panel has no owner, delete it.
What AI agent observability must show
OpenTelemetry treats a trace as the path of a request and a span as one unit of work with a name, parent, timestamps, attributes, and a status. Agents need that shape plus fields HTTP tracing never asked for: state name, tool side-effect class, evaluator evidence, and budget events.
Microsoft Foundry’s agent-tracing overview is blunt about why chat logs fail here: many steps, order that changes with the input, long payloads, and nesting — a tool that calls another process that calls another tool. Their capture list is the right minimum: inputs and outputs, tool usage, token consumption, duration.
| Question | Field that answers it | Failure if missing |
|---|---|---|
| What job was this? | job_id, job_type, tenant | You cannot group or page |
| Which state did it die in? | State span name | You debug the model, not the loop |
| Which tools ran? | Tool spans + side-effect class | Damage hides in a 200 |
| What did they see and write? | Redacted args/results or hashes | You cannot reconstruct the write |
| What did the evaluator say? | Verdict + criterion codes | You are tracing activity, not correctness |
| How much did it cost? | Tokens + dollars + revision count | Finance finds out in week three |
| How did it end? | done / escalate / abort + reason code | Trends are free text |
If any answer requires downloading a raw prompt dump by default, the UX failed.
- Job and tenant on the root span
- State names match the runner, not marketing labels
- Every tool span has a side-effect class
- Evaluator verdict sits on the same timeline as the last write
- Terminal reason is a catalog code, not a paragraph
Why provider dashboards miss silent wrongness
Provider dashboards answer “did the API accept the request?” They do not answer “did the CRM note name the right account?” Token burn, latency, and HTTP status are traffic. Silent wrongness is a plausible write that fails a soft criterion nobody watches.
OpenAI’s own agent-eval guidance starts in the same place: grade the trace — model calls, tool calls, guardrails, handoffs — when the unit of work is a loop, not a single completion. That is the opposite of staring at a spend chart.
| Signal | What it proves | What it hides |
|---|---|---|
| Provider 200 / finish reason | The call completed | Wrong tool, wrong args, ignored criteria |
| Token chart | Spend this hour | Cost per passing job |
| p95 model latency | Model SLO | Tool timeouts and revision grind |
| “Agent said done” | The worker is confident | The evaluator never ran |
| Vanity pass rate | A green tile | Revision rate, coverage, rewrite by humans — see why pass rate lies |
The failure mode: ops trusts the green tile, sales forwards a bad note, and nobody can open the run. Correlation ids are cheaper than a dedicated “AI ops” hire.
What a useful trace looks like
LangSmith’s model is the same physics under different nouns: a run is one unit of work (think span); a trace is the collection of runs for one operation, bound by a trace id. Their docs also cap a trace at 25,000 runs — a reminder that unbounded loops are a telemetry problem, not only a cost problem.
OpenAI’s Agents SDK records the tree you actually need: task, agent, turn, generation, function (tool), guardrail, handoff. Use that as a shopping list even if you never open their dashboard.
run (job_id, tenant, job_type)
state:intake
state:plan → model span (model id, tokens, latency)
state:act → tool:crm.get → tool:enrichment.lookup
state:evaluate → mechanical checks + judge spans
state:revise …
terminal (reason_code, cost_usd, revision_count)
| Span | Must store | Must not store by default |
|---|---|---|
run | job_id, tenant, job_type, trace_id | Full customer record |
state:* | Name, enter/exit, transition reason | Chat transcript paste |
model | Provider, model id, tokens, latency, cache hit | Unredacted prompt forever |
tool | Name, side-effect class, latency, error, arg/result hash | Raw credentials, full PII payload |
eval | Verdict, criterion codes, evidence pointers | Another copy of the artifact |
terminal | done / escalate / abort, reason code | A novel |
Include stable ids for run, parent, and tool call. Point at artifacts with URIs. Apply redaction at write time, not in a later “we will clean the bucket” ticket.
Do not mark a run successful because the HTTP layer returned 200 if the evaluator failed. Terminal status is evaluator or human authority.
How do you trace LLM tool calls without leaking I/O?
Most business damage is a tool call. Tracing only the model is how you miss the write.
OpenTelemetry’s GenAI conventions now live in a dedicated repository (the old opentelemetry.io GenAI pages are stubs). They remain in Development as of August 2026, which is a reason to own an internal span schema and map outward — not a reason to skip tool spans. The convention’s tool-execution shape is the one to copy: an internal span named around the tool, with tool name and call id, and arguments/results only when policy allows.
Langfuse splits the same idea into observation types: a tool observation holds which tool ran, the arguments, and the return value. That is the I/O you debug. Cost, though, is tracked on generation and embedding observations unless you attach spend to the tool on purpose. If a paid search API or a code-exec sandbox has a meter, record it or it will vanish from unit economics.
OpenAI’s SDK is explicit that generation_span and function_span store inputs and outputs, and that this may be sensitive. Their default is to capture it (trace_include_sensitive_data defaults true). Production should flip that unless legal signed off.
| Tool-span field | Why it exists |
|---|---|
tool.name | Group failures |
side_effect_class | read / draft / irreversible_write |
duration_ms | Find the hung connector |
error_code | Low-cardinality, not a stack dump |
arg_hash / result_hash | Prove two runs touched the same shape without keeping PII |
redacted_preview | A few hundred characters for humans |
external_receipt_id | CRM note id, ticket comment id |
Procedure for a new tool:
- Classify the side effect before you emit a span.
- Hash args and results. Store a redacted preview, not the raw body.
- Attach the write system’s receipt id on success.
- Set span status from the tool error and the evaluator, not from HTTP alone.
- Refuse to ship the tool if the span cannot name the side-effect class.
Hashes let you prove two runs touched the same payload shape without retaining PII forever. Previews let humans debug without opening cold storage.
side_effect_class | Examples | Trace keep |
|---|---|---|
read | CRM get, search, lookup | May sample |
draft | Internal note, ticket comment in staging | Keep if later promoted |
reversible_write | Draft email unsent, status flip with undo | Keep on fail |
irreversible_write | Public send, refund, delete, production CRM | 100% |
policy_denied | Gate blocked the call | 100% |
If the runner cannot name the class, the tool is not ready for production. “It only posts a note” is still a write.
Where does cost belong on the timeline?
Cost that lands in a monthly export is archaeology. Cost that sits on the run is a control.
Langfuse’s rule is the one to steal: ingested provider usage beats inferred token math, and inferred cost is computed at ingest against the price list you had then. If the provider omits usage, estimate with a documented formula and mark usage_estimated=true so finance does not treat it as gospel.
Paid tools are the hole. Token dashboards never see a $0.04 enrichment call that ran twelve times on a revise loop. Attach that spend to the tool span or you will under-price the job.
| Cost field | Source of truth | Lie if you skip it |
|---|---|---|
| Input / output / cache tokens | Provider usage object | You invent a tokenizer |
| Model $ | Price list pinned at ingest | Last month’s list on this month’s run |
| Tool $ | Vendor meter or your invoice map | “AI spend” misses the APIs |
| Revisions | Runner counter | Cheap first pass, expensive grind |
| Cost per pass | Dollars / evaluator-passing runs | Cheap failures look efficient |
| Budget events | Control plane | You cannot explain the abort |
Chart cost bands by job_type, not by model name. The question is “what does a passing CRM note cost,” not “how many tokens did Sol use on Tuesday.”
- Provider usage copied onto the model span
- Estimates flagged
- Paid tools have a dollar field
- Cost per pass uses evaluator pass, not HTTP success
- Budget trips are first-class events on the same dashboard
Offline scores and online sampling
Offline: golden-set pass rate, cost per pass, revision depth — on every change that could move behavior.
Online: sample production runs through the same evaluator. Chart pass / fail / escalate, silent-fail samples (human overrides), cost bands, and drift after deploys.
OpenAI’s trace-grading workflow is the online half: inspect a representative trace, attach a grader, use the result to change prompts, tools, routing, or guardrails. Then they tell you to move to datasets when you need repeatability. That split is correct. Online sampling catches the live miss. Offline suites make the fix stay fixed.
Alert when online pass rate drops versus the trailing baseline, or when cost per pass spikes. That is how you catch prompt regressions and a rotten retrieval index.
| Score | Offline | Online |
|---|---|---|
| Pass / fail / escalate | Full golden set | Stratified sample |
| Revision depth | Every suite run | Sample + 100% of writes |
| Cost per pass | Suite batch | Rolling by job_type |
| Human override | Labeler disagreement | “Report wrong” + sales forwards |
| Coverage | Job shapes in the set | Job shapes seen this week |
A high online pass with a rising human-rewrite rate is not a healthy agent. It is a vanity tile. The veto panel lives in why pass rate lies.
Which dashboard will ops actually read?
Keep one screen for the weekly meeting. Anything that needs a data scientist to interpret will not get read. Link out to full traces for incidents.
| Panel | Purpose | Owner |
|---|---|---|
| Runs by terminal state | Are we escalating more? | On-call lead |
| Pass rate (online sample) | Quality, not vibes | Evaluator owner |
| Cost per pass | Unit economics | Whoever holds the budget |
| Top evaluator failure codes | Where to fix | Domain reviewer + eng |
| Kill-switch / budget events | Control plane health | On-call |
| p95 end-to-end | SLO, secondary | Eng |
Label the tiles in business language: “Jobs finished cleanly,” “Jobs sent to a human,” “Jobs stopped for budget or policy,” “Average cost when successful,” “Top reasons for handoff.” Model jargon stays on the drill-down.
| Vanity panel | Why it dies | Replacement |
|---|---|---|
| Tokens by model, no job_type | Nobody acts | Cost per passing job |
| “AI messages sent” | Activity, not outcomes | Terminal-state mix |
| Unowned latency heatmaps | Pretty, mute | p95 with a page |
| Prompt gallery with no scores | Tourism | Trace + verdict |
| Twelve tabs of embeddings | Nobody opens tab three | One screen, six panels |
If a panel has no owner, delete it. Orphan metrics create false comfort and burn attention.
Sampling that does not lie
Tracing 100% of runs is fine at low volume and expensive at high volume. OpenTelemetry’s sampling note is the right starting physics: if most requests finish clean, you do not need every trace. Head sampling cannot promise you will keep the errors — that needs tail sampling, or an explicit keep-list. Cost totals and terminal-state counts belong on metrics, which are built to aggregate, not on traces you might drop.
When you sample traces:
- Keep 100% of
escalate,abort, budget trips, and irreversible writes. - Sample passes, stratified by
job_typeand tenant size. - Upsample for 48 hours after a deploy.
- Always keep a run that a human marked wrong.
- Never sample “successes only” to save money.
| Class | Trace keep rate | Why |
|---|---|---|
| Irreversible write | 100% | Disputes and rollbacks |
| Escalate / abort / budget | 100% | That is the incident stream |
| Human “wrong” report | 100% | Golden-set candidates |
| Ordinary pass | 5–20%, stratified | Baseline without drowning storage |
| Post-deploy window | 100% for 48h | Regressions cluster here |
Sampling only successes hides the story. Sampling only errors hides the baseline you need to see drift.
How do you keep one story across handoffs and writes?
If Agent B starts a new unrelated id, you will never reconstruct the story. Propagate one trace_id / run_id through every hop.
The W3C Trace Context recommendation exists because vendor-private ids break at the boundary. traceparent is the portable position in the graph; tracestate is the vendor bag. Foundry’s multi-agent notes are built on that plus OpenTelemetry, and they call out the exact failure: tool arguments and results on execute_tool, plus evaluation events, only help if the parent id survived the hop.
LangSmith’s thread_id is the conversation-shaped version of the same rule: one turn is a trace; many turns are a thread. Business agents need a third join: the write system’s receipt.
| Hop | What you propagate | What you add |
|---|---|---|
| Runner → model | trace_id, run_id | Model span |
| Runner → tool | Same ids + tool_call_id | Side-effect class, receipt id |
| Agent A → Agent B | Same trace_id; new child spans | Agent name, handoff reason |
| Runner → CRM / tickets | run_id on the note | external_receipt_id on the span |
| Human escalate | Same ids on the package | Reviewer id, decision |
Store the CRM note id, ticket comment id, or email draft id on the trace. When a salesperson says “the agent wrote nonsense,” you jump to the run in seconds. Without that join, observability is a museum.
OpenAI’s SDK lets you wrap multiple run() calls in one parent trace() so they stay one story. Do the same in your runner even if you are not on that SDK.
Logging hygiene is a security control
Observability that leaks customer data is a security incident with charts.
OWASP’s current GenAI LLM Top 10 (2026, published August 4) still ranks Sensitive Information Disclosure as LLM02. Treat traces as a disclosure path: prompts, tool args, retrieved chunks, and error strings. Foundry’s own guidance matches: do not store secrets in prompts or span attributes; redact before telemetry; apply the same access controls you use for production logs.
LangSmith SaaS retains traces 180 days from ingest unless you copy a case into a dataset. That is a contract fact, not a default you should assume is right for your counsel. Prefer artifact URIs and hashes over duplicating sensitive payloads into a third-party SaaS.
| Rule | Default | Break-glass |
|---|---|---|
| Secrets / credentials | Never log | Rotate, do not “temporarily” print |
| PII in args/results | Redact at write | Audit-logged unredact, time-boxed |
| Full prompts | Off in prod | On in staging, or sampled with legal |
| Debug verbosity | Separate flag | Hours, not weeks |
| Third-party SaaS | Hashes + URIs | Payload only after a privacy review |
| Retention | Documented, job-typed | Dispute hold, then delete |
Before enabling full prompt capture in production, run a privacy review: what PII appears, who can access, retention, export paths. Tenant exports must not dump another tenant’s traces.
How do you run an incident from a trace?
When something bad ships:
- Find run ids from the write system’s audit (CRM, email, tickets) — not from a Slack screenshot of the model reply.
- Open the trace. Identify the first bad tool call or the first failed criterion that was ignored.
- Freeze writes if the pattern is broad (kill switch). The trace should show
abortand a reason code, not a hungact. - Patch the evaluator or the sandbox. Add a golden case from this run.
- Replay the suite before you re-enable autonomy.
| Step | Done when | Common miss |
|---|---|---|
| Locate | Receipt id → run id < 60s | No join table |
| Diagnose | First bad span named | Blaming “the model” |
| Contain | Kill switch in the trace | Slack “please stop” |
| Fix | Criterion or tool patch + case | Prompt-only vibes |
| Prove | Suite green on the new case | Re-enable on hope |
Blameless for humans. Ruthless for missing criteria.
On-call owns kill switches, credential rotations, and freeze-writes. Model-quality tweaks wait for business hours unless an active incident is burning money or customers. Write that sentence into the runbook before launch.
- Kill switch location and who may throw it
- How to freeze writes without killing reads
- Where receipt id → run id is joined
- Redaction break-glass, time-boxed
- Who pages for cost-band breaches vs quality drops
Reason codes must be a catalog. Dashboards group on these. Free-text reasons make trends impossible.
| Code | Means | Typical next action |
|---|---|---|
eval_pass | Evaluator accepted | Count as a pass; sample the trace |
eval_fail_exhausted | Revisions hit the ceiling | Add a golden case; inspect last tool |
budget_exhausted | Tokens, dollars, or steps | Tighten the cage or the job |
tool_auth_error | Credential or scope miss | Rotate or shrink the allowlist |
policy_violation | Gate denied a side effect | Criteria or tool args |
human_reject | Reviewer said no | Teach the evaluator |
timeout | No terminal in the SLO | Find the hung span |
out_of_scope | Job was not this job | Intake filter |
When you compare prompt A vs B, run the same golden set, same tool stubs, same budget. Report pass rate, cost, latency. Store the batch under an experiment id on the root span. Intuition-only prompt merges are how regressions ship.
What vendor LLM observability still gets wrong
Many tools stop at prompt/response capture. Necessary. Insufficient.
Business agents need state names, tool side-effect classes, evaluator verdicts, and budget events in the same timeline. Buy or build toward that model. Do not confuse token charts with ops readiness.
| Vendor surface | Useful for | Not enough because |
|---|---|---|
| Provider token / latency charts | Spend and SLO | No tools, no evaluator, no job_type |
| SDK traces (OpenAI, others) | Fast tree of generations and functions | Defaults often keep raw I/O; your states may be missing |
| LangSmith / Langfuse style apps | Run trees, online evals, cost on generations | You still have to emit side-effect class and reason codes |
| Foundry / App Insights | OTel-shaped agent traces | You still own redaction and the weekly ritual |
| Homegrown JSON in a bucket | Cheap at ten runs a day | Nobody opens it during an incident |
Normalize provider-specific usage fields into your span schema. You will switch models. Your dashboards should not require a rewrite each time. Store raw provider payloads as optional debug attachments with stricter retention.
A “report wrong output” control that captures run_id is worth more than three vanity charts. Route those reports into the weekly review and into the golden set.
Weekly ritual and the vanity-dashboard failure
Ritual beats platform. Thirty minutes, named owners, same six panels.
- Glance terminal-state mix.
- Open the top three failure codes. Assign a fix owner or accept the rate.
- Check cost per pass against the band.
- Review one escalate package end to end — criteria, tool I/O, receipt id.
- Note kill-switch or budget events. If none ever fire, the switch is a rumor.
| Failure | What breaks | What it costs | What you do instead |
|---|---|---|---|
| Logs only on error | No baseline | You cannot see drift | Keep stratified passes |
| Full prompts in Slack | Retention and access | A disclosure plus a useless archive | Trace backend, redacted |
| Metrics with no owner | Mute tiles | Attention and false comfort | Delete or assign |
| Tracing only the model | Missed writes | Customer-facing nonsense | Tool spans + receipts |
| Sampling only successes | Hidden fails | Recurring incidents | 100% writes and aborts |
| Pass-rate-only screen | Green lie | Bad deploys | Pair with why pass rate lies |
Example SLOs you can actually page: 99% of runs reach a terminal state within 15 minutes; under 1% abort for unknown errors; online pass rate stays within five points of offline. SLOs make observability a control, not a gallery.
Spurlock Studios installs the thin version during the $1,500 · 5-day pilot on /agentic: structured run logs, evaluator verdicts, cost, terminal reason, and the six-panel screen. Full platforms can wait. Blindness should not. The parent map is the operating manual.
Day-one workbook if you are standing this up yourself:
| Day | Ship |
|---|---|
| 1 | Span schema: run, state, model, tool, eval, terminal |
| 2 | Emit spans from the runner with redaction and receipt ids |
| 3 | One dashboard, six panels, named owners |
| 4 | Alerts for budget trips and pass-rate drops |
| 5 | Rehearse an incident with a deliberate bad deploy in staging |
You do not need a perfect platform to start. You need correlated run ids and evaluator scores beside tool calls.
Add a “report wrong output” control that captures run_id on day one. That button is worth more than three vanity charts. Route reports into the weekly review and into golden-set candidates.
Correlate CRM write ids to run ids the same day — even before pretty charts. When sales forwards a bad note, you should open the trace in under a minute. If you cannot, the dashboard is decoration.
Traces, tool I/O, and cost. Not vanity.
FAQ
What is AI agent observability?
It is recording and reviewing each run with enough structure — states, model calls, redacted tool I/O, evaluator scores, cost, and terminal reasons — to debug, govern, and improve agents in production. Provider token charts are traffic. Observability is the story of one job from intake to done, escalate, or abort. If a human cannot answer those questions without a raw dump, you do not have it yet.
How do you trace LLM tool calls well?
Create a span for each model and tool invocation under a stable run id. Record tool name, side-effect class, latency, error codes, hashes, and a redacted preview — not raw credentials or full PII. Align terminal status with the evaluator, not with HTTP 200. Paid tools need their own dollar field or they disappear from unit economics.
Which metrics matter most week to week?
Online pass rate, escalate rate, cost per passing run, top evaluator failure codes, and budget or kill-switch events. Latency matters after correctness and cost. A green pass rate with a rising human-rewrite rate is a vanity tile; pair the screen with the veto metrics in why pass rate lies.
Do we need a special vendor on day one?
Not always. Structured logs plus a six-panel dashboard can cover a pilot. As volume grows, specialized tracing tools help if they ingest tool events, evaluator verdicts, and budget trips — not only prompts. Own the span schema either way. Vendor conventions move; your incident questions do not.
How does Spurlock Studios handle observability?
Thin traces and scores ship in the $1,500 · 5-day pilot: run ids, redacted tool I/O, evaluator verdicts, cost, and terminal reason, plus the weekly screen. Richer operator platforms land in fuller builds. The stack is described in the operating manual. Packaging lives on /agentic.
How does observability relate to evaluators?
Evaluators produce the quality signal. Observability stores that signal beside cost and tool I/O so a human can act. Without evaluators you are tracing activity, not correctness. Without traces, evaluator scores are a tile with no story. You need both on one timeline.
CTA
Install the minimum signal. Then make someone read it every week.
What questions does this article answer?
- What is AI agent observability?
- It is recording and reviewing each run with enough structure — states, model calls, redacted tool I/O, evaluator scores, cost, and terminal reasons — to debug, govern, and improve agents in production. Provider token charts are traffic. Observability is the story of one job from intake to `done`, `escalate`, or `abort`. If a human cannot answer those questions without a raw dump, you do not have it yet.
- How do you trace LLM tool calls well?
- Create a span for each model and tool invocation under a stable run id. Record tool name, side-effect class, latency, error codes, hashes, and a redacted preview — not raw credentials or full PII. Align terminal status with the evaluator, not with HTTP 200. Paid tools need their own dollar field or they disappear from unit economics.
- Which metrics matter most week to week?
- Online pass rate, escalate rate, cost per passing run, top evaluator failure codes, and budget or kill-switch events. Latency matters after correctness and cost. A green pass rate with a rising human-rewrite rate is a vanity tile; pair the screen with the veto metrics in [why pass rate lies](/blog/why-pass-rate-lies).
- Do we need a special vendor on day one?
- Not always. Structured logs plus a six-panel dashboard can cover a pilot. As volume grows, specialized tracing tools help if they ingest tool events, evaluator verdicts, and budget trips — not only prompts. Own the span schema either way. Vendor conventions move; your incident questions do not.
- How does Spurlock Studios handle observability?
- Thin traces and scores ship in the **$1,500 · 5-day** pilot: run ids, redacted tool I/O, evaluator verdicts, cost, and terminal reason, plus the weekly screen. Richer operator platforms land in fuller builds. The stack is described in the [operating manual](/blog/agentic-systems-operating-manual). Packaging lives on [/agentic](/agentic).
- How does observability relate to evaluators?
- Evaluators produce the quality signal. Observability stores that signal beside cost and tool I/O so a human can act. Without evaluators you are tracing activity, not correctness. Without traces, evaluator scores are a tile with no story. You need both on one timeline.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.