Why did quality drop when we didn’t change our prompts
Quality dropped because a pin, tool schema, retrieval corpus, eval set, or traffic mix moved — not because you edited prompts. Isolate, then pin versions.
William Spurlock Founder — Spurlock Studios 29 MIN
Quality dropped because something else moved. The prompt file is the last place I look. After 20,000+ hours on agentic systems and 500+ automations, the incidents that present as “we didn’t change the prompts” almost always resolve to a silent model swap, a tool schema change, retrieval corpus drift, eval set rot, or a traffic mix shift. Prompts staying still does not freeze an agent.
This spoke sits under the Agentic Systems Operating Manual. It owns the differential diagnosis when copy did not change and scores did. The evaluator that makes the diagnosis honest is built before the agent. Tool-retry storms that look like “worse quality” are often a missing loop detector — that harness lives in why agents loop on failed tools.
The short answer
- Prompts are one axis. Quality is the product of pin, tools, corpus, evals, and mix.
- Five causes, in the order I walk them: silent model swap, tool schema change, retrieval corpus drift, eval set rot, traffic mix.
- Pin versions for
model_id,prompt_version,tools_version, index hash, and eval-suite membership. Log them on every run. - Isolate with diffs across the drop window. Do not start by rewriting copy.
- A prompt edit that “fixes” a moved pin trains the next incident. Roll the identity back, then decide if copy should move.
What does a quality drop look like when prompts did not change?
Chatbots hide the drop in tone. Tool agents deposit it in CRM fields, ticket notes, and payment payloads. Support says “it got weird.” Finance says cost per pass left the band. Nobody opened the prompt PR.
| Signal you actually have | What it is not | First cause to check |
|---|---|---|
| Human revision rate up, prompt hash unchanged | “The model got dumber” | Pin, schema, mix |
| Offline golden still green; online pass down | A bad eval vendor | Traffic mix or corpus |
| Argument errors / retries up | A wording problem | Tool schema, then loop detector |
| Confident wrong citations | A “knowledge” gap in the prompt | Retrieval corpus / index hash |
| Offline pass rate up while humans still rewrite | A model win | Eval set rot |
| Cost per pass up, same job volume | A prompt that got wordier | Pin, sampling, reasoning effort |
| Same fingerprint retried until the turn cap | A quality complaint | Failed-tool loop, not copy |
I will not invent a vendor score like “provider X dropped 12%.” Public model cards do not publish your job’s argument accuracy. Your golden set does. If you cannot name last Tuesday’s identity line, you are grading folklore.
- Prompt hash for the drop window matches the week before
- I can name
model_id/resolved_idfor both weeks - I can name
tools_versionand the retrieval index hash - I can name which eval cases were added or removed
- I can show online mix (
job_type, locale, channel) vs the golden mix
If the first box is the only one you can check, you have an alibi, not a diagnosis.
Why are unchanged prompts a weak alibi?
A prompt is a file. An agent run is a request contract plus tools plus retrieved text plus the tickets that showed up. Holding one file still while the rest floats is how you ship a behavior change with a clean git log on prompts/.
| What people freeze | What still moves | Why the freeze lies |
|---|---|---|
| System prompt markdown | Alias behind gpt-5.6 / latest / gemini-flash-latest | Weights and defaults can move with no deploy (OpenAI model guidance, Gemini version patterns) |
| Prompt text | Tool JSON: required fields, enums, descriptions | The model is filling a different form |
| Prompt text | Chunker, embedder, index snapshot, allowlisted sources | Citations come from the index |
| Prompt text | Golden membership and judge rubric | You changed the test, not the student |
| Prompt text | Who files tickets, from where, in what language | Offline fixtures are last month’s mix |
| Prompt text | Sampling, thinking, reasoning.effort | Same ID, different traces |
Dateless Claude IDs from the 4.6 generation on are pinned snapshots, not evergreen pointers (Claude model IDs). OpenAI family aliases and Gemini -latest are the opposite. Read the ID rules before you treat “we didn’t touch prompts” as “nothing moved.”
- Production forbids aliases in CI (
latest,gpt-5.6,gemini-flash-latest, baresonnet/opus) - Prompt PRs are not the only diffs that gate the golden set
- On-call can paste an identity line from last Tuesday in thirty seconds
Bravery is not a restore strategy. An unchanged prompt is a clue. It is not a root cause.
Did a silent model swap move under you?
Yes, if production sent a different identity string — or the same alias resolved to different weights — and you never ran the golden set on the new identity. The prompt file staying still is expected. That is the point of an alias.
| Swap type | What changed | How it shows up on a tool agent |
|---|---|---|
| Alias float | Provider pointed gpt-5.6 / latest / -latest at new weights or defaults | Tool-choice shifts; verbosity jumps; cost per pass leaves the band; no git deploy |
| Sampling / thinking default | reasoning.effort, thinking-on-by-default, temperature semantics | Longer traces, new 400s, sudden brevity on the same ID |
| Serving-infra tweak | Router, safety classifier, sampler under a still-valid ID | Rare tone or refusal shifts. Anthropic is explicit that infra can still move on a pinned ID (model IDs) |
| Partner-cloud skew | Bedrock / Vertex clock ≠ Claude API / OpenAI API clock | Same marketing name, different retirement and routing |
| “Friday upgrade” | Someone flipped a dashboard model picker | Identity moved; prompts did not |
I do not have a public percentage for “how much quality fell.” I have incidents where the repo had no prompt diff and the identity string did. That is enough to treat aliases as a drift channel.
Check the window like this:
- Pull
model_idandresolved_idfor seven days before the complaint and seven days after. - Diff
reasoning.effort/ thinking / sampling knobs on the same rows. - If either column moved, freeze prompts and tools and run the golden set against both identities.
- Canary is not a vibe. Hold if write-tool argument accuracy drops, even when pass rate is flat.
- Roll back to the last scored pin. Then open an upgrade PR. Do not “fix it in the prompt.”
- Prod
model_idis a snapshot or documented stable ID, not an alias - Staging may float only if it logs
resolved_idand someone reads the nightly score - Pin changes go through the same golden-set gate as prompt PRs
The model did not owe you a press release. Your config owed you an ID.
Did a tool schema change break argument accuracy?
Yes, when the form the model fills changed and the prompt still describes last month’s form. Required fields appear. Enums rename. Descriptions that used to hint “leave this blank” now imply a write. The golden set still passes if it stubs the old schema.
| Schema move | Symptom | Prompt-only “fix” that makes it worse |
|---|---|---|
| New required field | Argument errors, then retries on the same fingerprint | “Always fill every field” → junk writes |
| Enum renamed / added | Wrong-tool or invalid-enum retries | Extra examples that fight the live schema |
| Description drift | Tool-choice flip; extra calls per turn | Longer system prompt, same 400s |
| Handler behavior change, schema text same | Side effects the eval never stubbed | You scored the JSON, not the write |
| Auth / idempotency change | Duplicate writes that look like “the agent is sloppy” | Copy about “be careful” |
This is why agents loop on failed tools when quality “drops.” The harness treats every turn as progress. The schema started returning permanent errors. Copy cannot honor retryable: false.
Walk the schema before you touch a sentence:
- Diff
tools_versionand the JSON Schema hash across the drop window. - Replay golden write-tool cases against the new schema with tools stubbed.
- Score argument accuracy separately from pass rate. A higher pass with worse args is a lie.
- If the live handler changed, add a contract test on the side effect, not only the model’s JSON.
- Bump
tools_version. Do not bury the schema inside a prompt SemVer.
# Incident grep — schema, not copy
tools_version: crm-tools@1.6.2 → crm-tools@1.7.0
schema_hash: 9f3a… → c21b…
arg_accuracy: hold
retryable_false_honored: false # now you also have a loop
- Tool PRs run the golden set even when no prompt file changed
- Write tools have argument-accuracy as a hard gate
- Permanent tool errors are
retryable: falsein the harness, not in a paragraph
If on-call cannot tell “new required field” from “model got vague,” you will rewrite the prompt every time sales ships a CRM field.
Did the retrieval corpus drift?
Yes, when the index, chunker, embedder, or allowlisted sources moved and the prompt still says “cite the knowledge base.” The model is not remembering your wiki. It is quoting whatever the retriever dumped into the window.
| Corpus move | Symptom | Why the prompt looks innocent |
|---|---|---|
| Re-embed / reindex without a snapshot | Confident wrong sources; citation URLs that 404 | Prompt still says “use retrieved context” |
| Chunker or overlap change | Answers skip the paragraph that used to be in chunk 0 | Same query, different window |
| Embedder pin change | Near-duplicate chunks crowd out the right one | You did not edit instructions |
| Allowlist / ACL change | Missing policy docs; the model fills from pretraining | Looks like a refusal or a hallucination |
| Stale memory / thread store | Yesterday’s ticket facts leak into today’s write | Prompt never mentioned memory |
Citation cases belong in the golden set the way evaluators before agents already demand evidence. If the evaluator does not see the retrieved chunks, you are grading vibes.
- Index hash / corpus snapshot ID is in the run log next to
model_id - Embedder ID and chunker version are pinned, not “whatever the pipeline used”
- Golden citation slice fails when the source is wrong, even if the answer sounds fine
- Reindex is a scored change: freeze prompts, run citation cases, then flip the hash
# Pin the corpus the way you pin the model
retrieval:
index_hash: kb-support@2026-06-10
embedder_id: # the ID you evaluated, not “default”
chunker_version: markdown-v3.1
allowlist: support-kb
A prompt that says “don’t hallucinate” is not a retriever. If online wrong-source rate rose and the prompt hash did not, pull the index hash before you add another sentence about being careful.
Did the eval set rot?
Yes, when the suite that tells you quality is “fine” no longer represents the job. Pass rate is a function of the cases you kept. Delete the hard rows, loosen the rubric, or let the judge share the worker’s context, and the dashboard goes green while humans still rewrite.
| Rot type | What changed | Dashboard lie |
|---|---|---|
| Hard cases deleted | Someone “cleaned up flakes” | Pass rate up; production still angry |
| Easy cases added | Demos and happy paths padded the suite | Pass rate up; write tools still miss fields |
| Rubric loosened | Evaluator version bumped without a note | Same artifacts, new passes |
| Judge prompt / model floated | The scorer moved, the worker did not | You measured judge drift |
| Fixtures frozen while mix moved | Suite is last quarter’s tickets | Offline green, online red |
| Vendor eval dashboard only | Suite lives in a product that can shut down | You cannot replay the incident |
Own the suite in git. OpenAI announced deprecation of its Evals platform on 2026-06-03, with a read-only window starting 2026-10-31 and a scheduled shutdown on 2026-11-30 (OpenAI deprecations). If your only regression memory lives in a vendor UI, the vendor can retire the UI.
Rot check:
- Diff eval-suite membership (added / removed / renamed cases) across the drop window.
- Diff evaluator version: criteria, ceilings, judge identity. A jump in pass rate after a looser rubric is not a model win.
- Re-run the old suite against current prod identity. If old-suite fails and new-suite passes, you aged the test.
- Harvest one real failure from this week into git before you celebrate any recovery.
- Version the evaluator the way you version prompts. The evaluator spoke already treats a self-grading worker as homework.
| Signal | Treat as | Do not do |
|---|---|---|
| Pass ↑, revision rate ↑ | Rot or mix | Ship a “quality improved” note |
| Pass ↑, hard cases missing | Rot | Add more easy demos |
| Pass flat, arg accuracy down | Schema or pin | Call it a wash |
| Judge pass ↑, human pass down | Judge drift | Trust the model-as-judge alone |
- Eval suite membership is reviewed in PRs like code
- Deleted cases need a reason, not a cleanup commit
- Evaluator version is in the identity line
- A vendor dashboard is a view, not the source of truth
If you cannot replay last month’s failing ticket as a fixture, you do not have evals. You have a mood.
Did the traffic mix change?
Yes, when the tickets, locales, channels, or job types in production diverged from the golden set while prompts and pins stayed put. Offline-only teams learn this from angry humans. Online sampling with the same evaluator catches it first.
| Mix shift | What you see | Offline golden |
|---|---|---|
| New locale / language | Online pass down; citation style breaks | Still last language |
| New channel (chat → email, or the reverse) | Length and tone leave the band | Fixtures are the old channel |
| Seasonal ticket types | New job_type values the tools were not built for | Suite never had those rows |
| Power users vs first-timers | Ambiguous asks; more escalate | Goldens were clean demos |
| New product / SKU / venue | Tools missing fields the new object needs | Schema looks “unchanged” in git because nobody added the object |
| Burst volume | Timeouts, partial writes, retry storms | Load was never in the suite |
This is the cause that most often gets blamed on the model. The pin is unchanged. The prompt is unchanged. The customers are not.
| Compare | How | Hold if |
|---|---|---|
job_type histogram | Online 7-day vs golden labels | A new label is > a small slice and untested |
| Locale / language | Online vs golden | Untested locale is now material |
| Channel | Chat / email / form | Channel that is untested is now the majority |
| Ticket length | Token in vs golden | Median length left the band you scored |
| First-time vs repeat | Account age | First-time share jumped and goldens were all repeats |
- Online sample uses the same evaluator as the offline suite
- Mix dashboards sit next to pass rate, not in a different product
- A new
job_typeis a scored change, same as a tool PR - You do not swap the pin to “fix” a seasonal mix
If online pass or revision rate diverges from offline while the pin and prompt hash are unchanged, investigate mix and corpus first. Swapping the pin to chase seasonality is how you get two problems.
Walk a mix incident without touching copy:
- Freeze
prompt_version,model_id, andtools_versionfor the week. - Dump seven-day online histograms for
job_type, locale, channel, and median input tokens. - Diff those histograms against golden labels. Name the slice that appeared.
- Harvest five tickets from that slice into git with stubbed tools and expected verdicts.
- Score current prod identity on the new rows. If they fail, you found the drop. If they pass, the humans are scoring a criterion you never wrote down — that is evaluator work, not a prompt emergency.
Which versions must you pin so the next drop is diagnosable?
Pin the request contract, not the prompt file alone. An identity line that on-call can grep is the difference between a one-hour isolation and a week of prompt archaeology.
| Axis | Example | Bumps when | Drop symptom if it floats |
|---|---|---|---|
prompt_version | support-triage@3.2.0 | Wording, tool descriptions inside the prompt, rubrics you stuffed into copy | This post’s alibi. Still pin it. |
model_id | claude-sonnet-5 or gpt-5.6-terra | Pin change | Silent swap |
request_contract | reasoning.effort=medium | Sampling / thinking / effort | Same ID, new traces |
tools_version | crm-tools@1.7.0 | Schema / handler contract | Argument errors |
index_hash | kb-support@2026-06-10 | Reindex, chunker, embedder, allowlist | Wrong citations |
eval_suite | golden@2026-06-17 | Cases added or removed | Rot |
eval_version | judge@2.1.0 | Criteria, ceilings, judge identity | Dashboard lie |
mix_snapshot | online@2026-06-10 | What you sampled for the weekly score | You cannot prove mix |
Prefer the most specific model ID the provider documents as a pinned snapshot or stable ID. Re-check the provider page before you copy strings months later. As of August 2026 docs, that means OpenAI gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna by role (not the gpt-5.6 alias), Anthropic claude-sonnet-5 / claude-opus-5 / claude-fable-5, Gemini gemini-3.6-flash as a documented stable ID — not gemini-flash-latest (OpenAI, Claude IDs, Gemini 3.6 Flash).
# models.yml — the line you grep during the drop
agents:
support_triage:
model_id: gpt-5.6-terra
reasoning_effort: medium
prompt_version: support-triage@3.2.0
tools_version: crm-tools@1.7.0
index_hash: kb-support@2026-06-10
eval_suite: golden@2026-06-17
eval_version: judge@2.1.0
last_eval_passed: 2026-06-16
rollback_model_id: gpt-5.6-terra
- CI denylists aliases
- CI fails if
last_eval_passedis missing or older than your SLA - Prompt bump and pin bump are separate axes
- A tools bump scores argument accuracy before anyone celebrates pass rate
If on-call cannot paste model=… | prompt=… | tools=… | index=… | golden=… from logs, you will keep holding a prompt file still and calling it an investigation.
How do you isolate which cause moved?
Treat it as an incident, not a writing workshop. Confirm the alibi, then walk the five causes in order. Stop at the first diff that explains the signals. You can have two causes. You still isolate one at a time.
- Confirm the alibi. Hash
prompt_versionacross the drop window. If it did change, this is not this post. Score that prompt PR. - Diff the pin.
model_id,resolved_id,reasoning.effort, thinking, sampling. If any moved, replay golden write-tool cases on both identities with tools stubbed. - Diff the schema.
tools_versionand schema hash. Replay argument-accuracy cases on the new schema. Check whetherretryable: falseis honored. - Diff the corpus. Index hash, embedder, chunker, allowlist. Run the citation slice. Pull three wrong-source traces and quote the chunks.
- Diff the suite. Membership and
eval_version. Re-run the old suite on current prod identity. - Diff the mix. Online
job_type/ locale / channel vs golden labels. If online and offline diverged, harvest the new mix into fixtures before you touch copy. - Only then open a prompt PR. It must still pass the suite on the restored identity.
| Isolation result | Ship | Hold |
|---|---|---|
| Pin moved, golden worse on candidate | Rollback pin | Prompt edit on the new pin |
| Schema moved, arg accuracy down | Schema revert or scored tools bump | “Be more complete” copy |
| Index hash moved, citations wrong | Roll index or scored reindex | “Don’t hallucinate” sentence |
| Old suite fails, new suite passes | Restore cases / rubric | Celebrate the pass-rate jump |
| Mix diverged, pin+prompt+tools still | Harvest fixtures; maybe a new job | Swap pin to chase seasonality |
| Nothing in identity moved | Serving-infra or an unlogged knob | Guess in Slack |
Gate the diagnosis the same way you gate an upgrade: pass rate, argument accuracy, cost per pass, new failure codes. A recommended replacement on a vendor deprecation page is a candidate, not a pass (OpenAI deprecations).
Two diffs at once are common. Sequence the rollback so you do not “fix” the wrong axis:
- Restore the last scored pin first if
model_idorresolved_idmoved. Replay write-tool cases. - If args are still wrong, restore or rescore tools_version. Do not stack a prompt edit on a new schema.
- If citations are still wrong, restore index_hash. The model is quoting the window, not your new paragraph.
- Only after identity is still and online/offline still diverge, harvest mix into the suite.
- Prompt PR last, on the restored identity, with the isolation note in the description.
- One owner for the isolation, not “the team”
- Every step produces a diff, not a theory
- Prompt PRs are locked until step 6 is written down
- Rollback is an identity flip, not a new paragraph
- Two-cause incidents still get one rollback at a time
If you cannot run steps 1–6 because the fields were never logged, that is the finding. Instrument, then come back. Guessing the model is cheaper emotionally. It is more expensive in production.
What breaks if you rewrite the prompt to “fix” a moved pin?
You paper over the wrong axis. The next provider notice, schema field, or reindex puts you back in the same thread — plus a prompt that only works on the broken identity.
What broke: Prod used a floating alias “so we always get the best.” A routing or default change shifted tool verbosity and argument discipline. Pass rate dipped. Cost per pass jumped. Git showed no prompt diff. Someone “fixed quality” by adding two pages of examples overnight. The alias kept floating.
Cost: The examples encoded the new (worse) tool-choice. A later rollback of the pin made the new prompt fail the old golden set. You now had two moving parts and one weekend. Support ate the writes. I will not dress that week up as a benchmark. The receipt is the missing identity line and the prompt PR that should not have shipped.
Instead: Restore the last scored pin. Freeze copy. Run the golden set. Canary. Then decide whether copy deserves a bump.
| Prompt-rewrite “fix” | What it actually did | What to do instead |
|---|---|---|
| Added examples for the new enum | Hid a schema bump inside copy | Bump tools_version; score args |
| “Always cite sources” paragraph | Hid a bad index behind wording | Pin index_hash; fail wrong sources |
| “Be concise” after cost jumped | Hid a reasoning-effort default | Pin request_contract |
| Deleted “flaky” golden cases | Hid mix or pin drift behind rot | Restore cases; harvest new failures |
| Swapped the alias again | Hid the first swap with a second | Pin a snapshot ID; gate the upgrade |
- Prompt PRs during a quality incident require an isolation note (which cause, which diff)
- Rollback map exists before anyone edits copy
- Golden set is run on the restored identity, not only on the “fixed” prompt
The prompt is not the product. The scored identity is. Copy that only works on an unpinned alias is a time bomb with nice formatting.
What must every run log so you can prove it next week?
If you cannot grep last Tuesday, you cannot prove what drifted. Folklore in Slack is not an identity line. Emit one line per production run, and keep it boring.
| Field | Example | Proves |
|---|---|---|
run_id | 01J… | The job |
model_id | gpt-5.6-terra | Weights you think you bought |
resolved_id | gpt-5.6-terra | Same as model_id on a pin; different if staging floated |
prompt_version | support-triage@3.2.0 | The alibi |
tools_version | crm-tools@1.7.0 | The form |
request_contract | reasoning.effort=medium | The knobs |
index_hash | kb-support@2026-06-10 | The corpus |
eval_suite | golden@2026-06-17 | What last scored this pin |
eval_version | judge@2.1.0 | Who graded |
job_type / locale / channel | refund / en-US / email | Mix |
| Terminal reason | ok / wrong_tool / no_progress | Whether you had a loop, not a vibe |
support_triage | model=gpt-5.6-terra | resolved=gpt-5.6-terra | prompt=support-triage@3.2.0 | tools=crm-tools@1.7.0 | index=kb-support@2026-06-10 | golden=2026-06-17 | effort=medium | job=refund | locale=en-US
| Role | Owns | Does not own |
|---|---|---|
| Job owner | Pass criteria, “ship / hold” on the weekly score | Quiet alias swaps |
| Eng | models.yml, CI denylist, log fields | Rewriting copy during an incident without an isolation note |
| On-call | Rollback to rollback_model_id from logs | Inventing a new pin at 2 a.m. |
- Identity line is in the trace, not only in the repo
- Online sample and offline suite share evaluator identity
- Provider changelog URLs are in the runbook (OpenAI, Anthropic, Gemini)
- Someone is named on the weekly score
A pin that never gets a candidate bake-off is a pin you will rip out under a retirement deadline. OpenAI documents at least six months’ notice for GA model retirement after announcement; Anthropic documents at least sixty days before retirement for public models; Gemini publishes shutdown dates per ID and about two weeks’ notice on preview and -latest breaking changes (OpenAI, Anthropic deprecations, Gemini deprecations). Those are notice windows, not quality SLAs. Do not treat them as a promised pass rate.
What does a one-week triage look like?
If you only have a week, do not rewrite the prompt library. Instrument identity, split online vs offline, and walk the five diffs. A Spurlock $1,500 · 5-day agentic pilot is this shape: one job, a scored identity, a gate, an online sample.
Day-by-day, one job only:
- Day 1 — alibi and logs. Confirm prompt hash. Emit the identity line. Denylist aliases in CI. If you cannot log it today, you are not triaging quality. You are guessing.
- Day 2 — pin and schema. Diff
model_id/resolved_id/ schema hash for two weeks. Roll back any unscored pin. Stub tools and run write-tool cases. - Day 3 — corpus and citations. Record
index_hash. Run the citation slice. Freeze reindex until the slice is green. - Day 4 — suite hygiene. Diff membership. Restore deleted hard cases. Harvest three live failures into git. Version the evaluator.
- Day 5 — mix and hold/ship. Compare online mix to golden labels. Canary only if identity is scored. Write the isolation note. Prompt edits, if any, are a separate PR.
| If you skip | What you will still not know on Friday |
|---|---|
| Identity logs | Whether the pin moved |
| Argument-accuracy gate | Whether the schema moved |
| Citation slice | Whether the corpus moved |
| Membership diff | Whether the suite rotted |
| Online vs offline | Whether the customers changed |
| Prompt rewrite | Anything useful |
- One job, not five agents
- Prompt freeze until isolation is written
- Rollback pin named
- Three harvested failures in git
- Weekly score owner named
You do not need multi-provider routing this week. You need to name what ran. If you cannot, do not scale the loop.
When is a workflow enough instead of hunting agent drift?
When the job is a known graph with known side effects, a workflow will not present as “quality dropped and we didn’t change prompts,” because it does not have an alias and a retriever to float. If your pain is a missing field in a CRM mapping, you do not have an agent-quality incident. You have an unversioned mapping.
| Situation | Hunt the five causes | Use a workflow instead |
|---|---|---|
| Model chooses tools, writes CRM, cites a corpus | Yes | No — you bought an agent, so buy identity |
| Fixed mapping, retries, webhooks | Maybe the mapping, not the model | Yes — automation shape |
You cannot name model_id for last Tuesday | Instrument first | Do not add more agent autonomy |
| No evaluator, no golden set | Build the evaluator first | Do not widen the loop |
| Same fingerprint retried until cap | Loop detector, then maybe quality | Fix the harness (failed-tool loops) |
| One tool, one field, one destination | You probably never needed an agent | Workflow |
Decision list:
- If there is no identity line, stop. Log, then return.
- If there is no evaluator, stop. You cannot see a drop you cannot score.
- If the job is a mapping, delete the agent costume.
- If the job needs judgment plus tools, hunt the five causes before you edit copy.
- If two weeks of isolation still show “nothing moved,” look at serving infra and unlogged knobs — not a new system prompt.
- Job owner can say “agent” or “workflow” without a slide
- Autonomy widens only when golden pass, cost per pass, and escalate rate hold on an online sample
- Prompt edits are a last step, not a first
Unchanged prompts are a true statement and a useless finding. Pin the versions. Diff the five causes. Then, and only then, touch a sentence.
FAQ
Why did quality drop when we didn’t change our prompts?
Because prompts are one axis. The drop almost always comes from a silent model swap, a tool schema change, retrieval corpus drift, eval set rot, or a traffic mix shift — often more than one. Confirm the prompt hash, then diff model_id, tools_version, index hash, eval membership, and online mix before you rewrite copy.
How do I measure whether the diagnosis is working?
You have a working diagnosis when the identity line explains the signals and a rollback or scored bump moves revision rate, argument accuracy, or wrong-source rate back toward the pre-drop band. Offline pass rate alone is not the measure — it can rise from eval rot. Use the same evaluator online and offline, and treat argument accuracy and citation correctness as hard gates.
What usually fails first when teams try this?
Logs. Teams rewrite the prompt because they can, then discover they cannot grep last Tuesday’s model_id. The next failure is treating pass rate as quality while hard cases quietly left the suite. Fix identity and membership before you hold a writing sprint.
How long does this take to show results?
A week of identity logs plus golden-vs-online split is enough to name the cause for one job. Repair time depends on which axis moved: a pin rollback can be a config PR the same day; a rotten suite or a bad reindex takes as long as restoring fixtures and snapshots. I will not invent a recovery percentage.
What should I skip if I only have a week?
Skip prompt rewrites, multi-provider routing, and a second agent. Instrument the identity line, denylist aliases, run write-tool and citation slices, diff eval membership, and compare online mix to golden labels. Harvest three live failures into git. That is the week.
When is this not worth doing yet?
When you have no evaluator, no production logs, and no pin — build those first, or you will be guessing. When the job is a fixed mapping, use a workflow instead of an agent and stop looking for model drift. When you already know the prompt did change, this diagnosis does not apply; score that prompt PR.
CTA
Want the five-cause isolation, pinned identity, and a golden-set gate on one real job this week? Start at /agentic or /contact?intent=agentic-pilot.
What questions does this article answer?
- Why did quality drop when we didn’t change our prompts?
- Because prompts are one axis. The drop almost always comes from a silent model swap, a tool schema change, retrieval corpus drift, eval set rot, or a traffic mix shift — often more than one. Confirm the prompt hash, then diff `model_id`, `tools_version`, index hash, eval membership, and online mix before you rewrite copy.
- How do I measure whether the diagnosis is working?
- You have a working diagnosis when the identity line explains the signals and a rollback or scored bump moves revision rate, argument accuracy, or wrong-source rate back toward the pre-drop band. Offline pass rate alone is not the measure — it can rise from eval rot. Use the same evaluator online and offline, and treat argument accuracy and citation correctness as hard gates.
- What usually fails first when teams try this?
- Logs. Teams rewrite the prompt because they can, then discover they cannot grep last Tuesday’s `model_id`. The next failure is treating pass rate as quality while hard cases quietly left the suite. Fix identity and membership before you hold a writing sprint.
- How long does this take to show results?
- A week of identity logs plus golden-vs-online split is enough to name the cause for one job. Repair time depends on which axis moved: a pin rollback can be a config PR the same day; a rotten suite or a bad reindex takes as long as restoring fixtures and snapshots. I will not invent a recovery percentage.
- What should I skip if I only have a week?
- Skip prompt rewrites, multi-provider routing, and a second agent. Instrument the identity line, denylist aliases, run write-tool and citation slices, diff eval membership, and compare online mix to golden labels. Harvest three live failures into git. That is the week.
- When is this not worth doing yet?
- When you have no evaluator, no production logs, and no pin — build those first, or you will be guessing. When the job is a fixed mapping, use a workflow instead of an agent and stop looking for model drift. When you already know the prompt *did* change, this diagnosis does not apply; score that prompt PR.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.