Spurlock Studios
Contact
Share LinkedIn X
Nested brass frames. Thesis: DID QUALITY DROP WE DIDN.

Quality dropped because something else moved. The prompt file is the last place I look. After 20,000+ hours on agentic systems and 500+ automations, the incidents that present as “we didn’t change the prompts” almost always resolve to a silent model swap, a tool schema change, retrieval corpus drift, eval set rot, or a traffic mix shift. Prompts staying still does not freeze an agent.

This spoke sits under the Agentic Systems Operating Manual. It owns the differential diagnosis when copy did not change and scores did. The evaluator that makes the diagnosis honest is built before the agent. Tool-retry storms that look like “worse quality” are often a missing loop detector — that harness lives in why agents loop on failed tools.

The short answer

  • Prompts are one axis. Quality is the product of pin, tools, corpus, evals, and mix.
  • Five causes, in the order I walk them: silent model swap, tool schema change, retrieval corpus drift, eval set rot, traffic mix.
  • Pin versions for model_id, prompt_version, tools_version, index hash, and eval-suite membership. Log them on every run.
  • Isolate with diffs across the drop window. Do not start by rewriting copy.
  • A prompt edit that “fixes” a moved pin trains the next incident. Roll the identity back, then decide if copy should move.

What does a quality drop look like when prompts did not change?

Chatbots hide the drop in tone. Tool agents deposit it in CRM fields, ticket notes, and payment payloads. Support says “it got weird.” Finance says cost per pass left the band. Nobody opened the prompt PR.

Signal you actually haveWhat it is notFirst cause to check
Human revision rate up, prompt hash unchanged“The model got dumber”Pin, schema, mix
Offline golden still green; online pass downA bad eval vendorTraffic mix or corpus
Argument errors / retries upA wording problemTool schema, then loop detector
Confident wrong citationsA “knowledge” gap in the promptRetrieval corpus / index hash
Offline pass rate up while humans still rewriteA model winEval set rot
Cost per pass up, same job volumeA prompt that got wordierPin, sampling, reasoning effort
Same fingerprint retried until the turn capA quality complaintFailed-tool loop, not copy

I will not invent a vendor score like “provider X dropped 12%.” Public model cards do not publish your job’s argument accuracy. Your golden set does. If you cannot name last Tuesday’s identity line, you are grading folklore.

  • Prompt hash for the drop window matches the week before
  • I can name model_id / resolved_id for both weeks
  • I can name tools_version and the retrieval index hash
  • I can name which eval cases were added or removed
  • I can show online mix (job_type, locale, channel) vs the golden mix

If the first box is the only one you can check, you have an alibi, not a diagnosis.

Why are unchanged prompts a weak alibi?

A prompt is a file. An agent run is a request contract plus tools plus retrieved text plus the tickets that showed up. Holding one file still while the rest floats is how you ship a behavior change with a clean git log on prompts/.

What people freezeWhat still movesWhy the freeze lies
System prompt markdownAlias behind gpt-5.6 / latest / gemini-flash-latestWeights and defaults can move with no deploy (OpenAI model guidance, Gemini version patterns)
Prompt textTool JSON: required fields, enums, descriptionsThe model is filling a different form
Prompt textChunker, embedder, index snapshot, allowlisted sourcesCitations come from the index
Prompt textGolden membership and judge rubricYou changed the test, not the student
Prompt textWho files tickets, from where, in what languageOffline fixtures are last month’s mix
Prompt textSampling, thinking, reasoning.effortSame ID, different traces

Dateless Claude IDs from the 4.6 generation on are pinned snapshots, not evergreen pointers (Claude model IDs). OpenAI family aliases and Gemini -latest are the opposite. Read the ID rules before you treat “we didn’t touch prompts” as “nothing moved.”

  • Production forbids aliases in CI (latest, gpt-5.6, gemini-flash-latest, bare sonnet / opus)
  • Prompt PRs are not the only diffs that gate the golden set
  • On-call can paste an identity line from last Tuesday in thirty seconds

Bravery is not a restore strategy. An unchanged prompt is a clue. It is not a root cause.

Did a silent model swap move under you?

Yes, if production sent a different identity string — or the same alias resolved to different weights — and you never ran the golden set on the new identity. The prompt file staying still is expected. That is the point of an alias.

Swap typeWhat changedHow it shows up on a tool agent
Alias floatProvider pointed gpt-5.6 / latest / -latest at new weights or defaultsTool-choice shifts; verbosity jumps; cost per pass leaves the band; no git deploy
Sampling / thinking defaultreasoning.effort, thinking-on-by-default, temperature semanticsLonger traces, new 400s, sudden brevity on the same ID
Serving-infra tweakRouter, safety classifier, sampler under a still-valid IDRare tone or refusal shifts. Anthropic is explicit that infra can still move on a pinned ID (model IDs)
Partner-cloud skewBedrock / Vertex clock ≠ Claude API / OpenAI API clockSame marketing name, different retirement and routing
“Friday upgrade”Someone flipped a dashboard model pickerIdentity moved; prompts did not

I do not have a public percentage for “how much quality fell.” I have incidents where the repo had no prompt diff and the identity string did. That is enough to treat aliases as a drift channel.

Check the window like this:

  1. Pull model_id and resolved_id for seven days before the complaint and seven days after.
  2. Diff reasoning.effort / thinking / sampling knobs on the same rows.
  3. If either column moved, freeze prompts and tools and run the golden set against both identities.
  4. Canary is not a vibe. Hold if write-tool argument accuracy drops, even when pass rate is flat.
  5. Roll back to the last scored pin. Then open an upgrade PR. Do not “fix it in the prompt.”
  • Prod model_id is a snapshot or documented stable ID, not an alias
  • Staging may float only if it logs resolved_id and someone reads the nightly score
  • Pin changes go through the same golden-set gate as prompt PRs

The model did not owe you a press release. Your config owed you an ID.

Did a tool schema change break argument accuracy?

Yes, when the form the model fills changed and the prompt still describes last month’s form. Required fields appear. Enums rename. Descriptions that used to hint “leave this blank” now imply a write. The golden set still passes if it stubs the old schema.

Schema moveSymptomPrompt-only “fix” that makes it worse
New required fieldArgument errors, then retries on the same fingerprint“Always fill every field” → junk writes
Enum renamed / addedWrong-tool or invalid-enum retriesExtra examples that fight the live schema
Description driftTool-choice flip; extra calls per turnLonger system prompt, same 400s
Handler behavior change, schema text sameSide effects the eval never stubbedYou scored the JSON, not the write
Auth / idempotency changeDuplicate writes that look like “the agent is sloppy”Copy about “be careful”

This is why agents loop on failed tools when quality “drops.” The harness treats every turn as progress. The schema started returning permanent errors. Copy cannot honor retryable: false.

Walk the schema before you touch a sentence:

  1. Diff tools_version and the JSON Schema hash across the drop window.
  2. Replay golden write-tool cases against the new schema with tools stubbed.
  3. Score argument accuracy separately from pass rate. A higher pass with worse args is a lie.
  4. If the live handler changed, add a contract test on the side effect, not only the model’s JSON.
  5. Bump tools_version. Do not bury the schema inside a prompt SemVer.
# Incident grep — schema, not copy
tools_version: crm-tools@1.6.2 → crm-tools@1.7.0
schema_hash: 9f3a… → c21b…
arg_accuracy: hold
retryable_false_honored: false   # now you also have a loop
  • Tool PRs run the golden set even when no prompt file changed
  • Write tools have argument-accuracy as a hard gate
  • Permanent tool errors are retryable: false in the harness, not in a paragraph

If on-call cannot tell “new required field” from “model got vague,” you will rewrite the prompt every time sales ships a CRM field.

Did the retrieval corpus drift?

Yes, when the index, chunker, embedder, or allowlisted sources moved and the prompt still says “cite the knowledge base.” The model is not remembering your wiki. It is quoting whatever the retriever dumped into the window.

Corpus moveSymptomWhy the prompt looks innocent
Re-embed / reindex without a snapshotConfident wrong sources; citation URLs that 404Prompt still says “use retrieved context”
Chunker or overlap changeAnswers skip the paragraph that used to be in chunk 0Same query, different window
Embedder pin changeNear-duplicate chunks crowd out the right oneYou did not edit instructions
Allowlist / ACL changeMissing policy docs; the model fills from pretrainingLooks like a refusal or a hallucination
Stale memory / thread storeYesterday’s ticket facts leak into today’s writePrompt never mentioned memory

Citation cases belong in the golden set the way evaluators before agents already demand evidence. If the evaluator does not see the retrieved chunks, you are grading vibes.

  • Index hash / corpus snapshot ID is in the run log next to model_id
  • Embedder ID and chunker version are pinned, not “whatever the pipeline used”
  • Golden citation slice fails when the source is wrong, even if the answer sounds fine
  • Reindex is a scored change: freeze prompts, run citation cases, then flip the hash
# Pin the corpus the way you pin the model
retrieval:
  index_hash: kb-support@2026-06-10
  embedder_id:  # the ID you evaluated, not “default”
  chunker_version: markdown-v3.1
  allowlist: support-kb

A prompt that says “don’t hallucinate” is not a retriever. If online wrong-source rate rose and the prompt hash did not, pull the index hash before you add another sentence about being careful.

Did the eval set rot?

Yes, when the suite that tells you quality is “fine” no longer represents the job. Pass rate is a function of the cases you kept. Delete the hard rows, loosen the rubric, or let the judge share the worker’s context, and the dashboard goes green while humans still rewrite.

Rot typeWhat changedDashboard lie
Hard cases deletedSomeone “cleaned up flakes”Pass rate up; production still angry
Easy cases addedDemos and happy paths padded the suitePass rate up; write tools still miss fields
Rubric loosenedEvaluator version bumped without a noteSame artifacts, new passes
Judge prompt / model floatedThe scorer moved, the worker did notYou measured judge drift
Fixtures frozen while mix movedSuite is last quarter’s ticketsOffline green, online red
Vendor eval dashboard onlySuite lives in a product that can shut downYou cannot replay the incident

Own the suite in git. OpenAI announced deprecation of its Evals platform on 2026-06-03, with a read-only window starting 2026-10-31 and a scheduled shutdown on 2026-11-30 (OpenAI deprecations). If your only regression memory lives in a vendor UI, the vendor can retire the UI.

Rot check:

  1. Diff eval-suite membership (added / removed / renamed cases) across the drop window.
  2. Diff evaluator version: criteria, ceilings, judge identity. A jump in pass rate after a looser rubric is not a model win.
  3. Re-run the old suite against current prod identity. If old-suite fails and new-suite passes, you aged the test.
  4. Harvest one real failure from this week into git before you celebrate any recovery.
  5. Version the evaluator the way you version prompts. The evaluator spoke already treats a self-grading worker as homework.
SignalTreat asDo not do
Pass ↑, revision rate ↑Rot or mixShip a “quality improved” note
Pass ↑, hard cases missingRotAdd more easy demos
Pass flat, arg accuracy downSchema or pinCall it a wash
Judge pass ↑, human pass downJudge driftTrust the model-as-judge alone
  • Eval suite membership is reviewed in PRs like code
  • Deleted cases need a reason, not a cleanup commit
  • Evaluator version is in the identity line
  • A vendor dashboard is a view, not the source of truth

If you cannot replay last month’s failing ticket as a fixture, you do not have evals. You have a mood.

Did the traffic mix change?

Yes, when the tickets, locales, channels, or job types in production diverged from the golden set while prompts and pins stayed put. Offline-only teams learn this from angry humans. Online sampling with the same evaluator catches it first.

Mix shiftWhat you seeOffline golden
New locale / languageOnline pass down; citation style breaksStill last language
New channel (chat → email, or the reverse)Length and tone leave the bandFixtures are the old channel
Seasonal ticket typesNew job_type values the tools were not built forSuite never had those rows
Power users vs first-timersAmbiguous asks; more escalateGoldens were clean demos
New product / SKU / venueTools missing fields the new object needsSchema looks “unchanged” in git because nobody added the object
Burst volumeTimeouts, partial writes, retry stormsLoad was never in the suite

This is the cause that most often gets blamed on the model. The pin is unchanged. The prompt is unchanged. The customers are not.

CompareHowHold if
job_type histogramOnline 7-day vs golden labelsA new label is > a small slice and untested
Locale / languageOnline vs goldenUntested locale is now material
ChannelChat / email / formChannel that is untested is now the majority
Ticket lengthToken in vs goldenMedian length left the band you scored
First-time vs repeatAccount ageFirst-time share jumped and goldens were all repeats
  • Online sample uses the same evaluator as the offline suite
  • Mix dashboards sit next to pass rate, not in a different product
  • A new job_type is a scored change, same as a tool PR
  • You do not swap the pin to “fix” a seasonal mix

If online pass or revision rate diverges from offline while the pin and prompt hash are unchanged, investigate mix and corpus first. Swapping the pin to chase seasonality is how you get two problems.

Walk a mix incident without touching copy:

  1. Freeze prompt_version, model_id, and tools_version for the week.
  2. Dump seven-day online histograms for job_type, locale, channel, and median input tokens.
  3. Diff those histograms against golden labels. Name the slice that appeared.
  4. Harvest five tickets from that slice into git with stubbed tools and expected verdicts.
  5. Score current prod identity on the new rows. If they fail, you found the drop. If they pass, the humans are scoring a criterion you never wrote down — that is evaluator work, not a prompt emergency.

Which versions must you pin so the next drop is diagnosable?

Pin the request contract, not the prompt file alone. An identity line that on-call can grep is the difference between a one-hour isolation and a week of prompt archaeology.

AxisExampleBumps whenDrop symptom if it floats
prompt_versionsupport-triage@3.2.0Wording, tool descriptions inside the prompt, rubrics you stuffed into copyThis post’s alibi. Still pin it.
model_idclaude-sonnet-5 or gpt-5.6-terraPin changeSilent swap
request_contractreasoning.effort=mediumSampling / thinking / effortSame ID, new traces
tools_versioncrm-tools@1.7.0Schema / handler contractArgument errors
index_hashkb-support@2026-06-10Reindex, chunker, embedder, allowlistWrong citations
eval_suitegolden@2026-06-17Cases added or removedRot
eval_versionjudge@2.1.0Criteria, ceilings, judge identityDashboard lie
mix_snapshotonline@2026-06-10What you sampled for the weekly scoreYou cannot prove mix

Prefer the most specific model ID the provider documents as a pinned snapshot or stable ID. Re-check the provider page before you copy strings months later. As of August 2026 docs, that means OpenAI gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna by role (not the gpt-5.6 alias), Anthropic claude-sonnet-5 / claude-opus-5 / claude-fable-5, Gemini gemini-3.6-flash as a documented stable ID — not gemini-flash-latest (OpenAI, Claude IDs, Gemini 3.6 Flash).

# models.yml — the line you grep during the drop
agents:
  support_triage:
    model_id: gpt-5.6-terra
    reasoning_effort: medium
    prompt_version: support-triage@3.2.0
    tools_version: crm-tools@1.7.0
    index_hash: kb-support@2026-06-10
    eval_suite: golden@2026-06-17
    eval_version: judge@2.1.0
    last_eval_passed: 2026-06-16
    rollback_model_id: gpt-5.6-terra
  • CI denylists aliases
  • CI fails if last_eval_passed is missing or older than your SLA
  • Prompt bump and pin bump are separate axes
  • A tools bump scores argument accuracy before anyone celebrates pass rate

If on-call cannot paste model=… | prompt=… | tools=… | index=… | golden=… from logs, you will keep holding a prompt file still and calling it an investigation.

How do you isolate which cause moved?

Treat it as an incident, not a writing workshop. Confirm the alibi, then walk the five causes in order. Stop at the first diff that explains the signals. You can have two causes. You still isolate one at a time.

  1. Confirm the alibi. Hash prompt_version across the drop window. If it did change, this is not this post. Score that prompt PR.
  2. Diff the pin. model_id, resolved_id, reasoning.effort, thinking, sampling. If any moved, replay golden write-tool cases on both identities with tools stubbed.
  3. Diff the schema. tools_version and schema hash. Replay argument-accuracy cases on the new schema. Check whether retryable: false is honored.
  4. Diff the corpus. Index hash, embedder, chunker, allowlist. Run the citation slice. Pull three wrong-source traces and quote the chunks.
  5. Diff the suite. Membership and eval_version. Re-run the old suite on current prod identity.
  6. Diff the mix. Online job_type / locale / channel vs golden labels. If online and offline diverged, harvest the new mix into fixtures before you touch copy.
  7. Only then open a prompt PR. It must still pass the suite on the restored identity.
Isolation resultShipHold
Pin moved, golden worse on candidateRollback pinPrompt edit on the new pin
Schema moved, arg accuracy downSchema revert or scored tools bump“Be more complete” copy
Index hash moved, citations wrongRoll index or scored reindex“Don’t hallucinate” sentence
Old suite fails, new suite passesRestore cases / rubricCelebrate the pass-rate jump
Mix diverged, pin+prompt+tools stillHarvest fixtures; maybe a new jobSwap pin to chase seasonality
Nothing in identity movedServing-infra or an unlogged knobGuess in Slack

Gate the diagnosis the same way you gate an upgrade: pass rate, argument accuracy, cost per pass, new failure codes. A recommended replacement on a vendor deprecation page is a candidate, not a pass (OpenAI deprecations).

Two diffs at once are common. Sequence the rollback so you do not “fix” the wrong axis:

  1. Restore the last scored pin first if model_id or resolved_id moved. Replay write-tool cases.
  2. If args are still wrong, restore or rescore tools_version. Do not stack a prompt edit on a new schema.
  3. If citations are still wrong, restore index_hash. The model is quoting the window, not your new paragraph.
  4. Only after identity is still and online/offline still diverge, harvest mix into the suite.
  5. Prompt PR last, on the restored identity, with the isolation note in the description.
  • One owner for the isolation, not “the team”
  • Every step produces a diff, not a theory
  • Prompt PRs are locked until step 6 is written down
  • Rollback is an identity flip, not a new paragraph
  • Two-cause incidents still get one rollback at a time

If you cannot run steps 1–6 because the fields were never logged, that is the finding. Instrument, then come back. Guessing the model is cheaper emotionally. It is more expensive in production.

What breaks if you rewrite the prompt to “fix” a moved pin?

You paper over the wrong axis. The next provider notice, schema field, or reindex puts you back in the same thread — plus a prompt that only works on the broken identity.

What broke: Prod used a floating alias “so we always get the best.” A routing or default change shifted tool verbosity and argument discipline. Pass rate dipped. Cost per pass jumped. Git showed no prompt diff. Someone “fixed quality” by adding two pages of examples overnight. The alias kept floating.

Cost: The examples encoded the new (worse) tool-choice. A later rollback of the pin made the new prompt fail the old golden set. You now had two moving parts and one weekend. Support ate the writes. I will not dress that week up as a benchmark. The receipt is the missing identity line and the prompt PR that should not have shipped.

Instead: Restore the last scored pin. Freeze copy. Run the golden set. Canary. Then decide whether copy deserves a bump.

Prompt-rewrite “fix”What it actually didWhat to do instead
Added examples for the new enumHid a schema bump inside copyBump tools_version; score args
“Always cite sources” paragraphHid a bad index behind wordingPin index_hash; fail wrong sources
“Be concise” after cost jumpedHid a reasoning-effort defaultPin request_contract
Deleted “flaky” golden casesHid mix or pin drift behind rotRestore cases; harvest new failures
Swapped the alias againHid the first swap with a secondPin a snapshot ID; gate the upgrade
  • Prompt PRs during a quality incident require an isolation note (which cause, which diff)
  • Rollback map exists before anyone edits copy
  • Golden set is run on the restored identity, not only on the “fixed” prompt

The prompt is not the product. The scored identity is. Copy that only works on an unpinned alias is a time bomb with nice formatting.

What must every run log so you can prove it next week?

If you cannot grep last Tuesday, you cannot prove what drifted. Folklore in Slack is not an identity line. Emit one line per production run, and keep it boring.

FieldExampleProves
run_id01J…The job
model_idgpt-5.6-terraWeights you think you bought
resolved_idgpt-5.6-terraSame as model_id on a pin; different if staging floated
prompt_versionsupport-triage@3.2.0The alibi
tools_versioncrm-tools@1.7.0The form
request_contractreasoning.effort=mediumThe knobs
index_hashkb-support@2026-06-10The corpus
eval_suitegolden@2026-06-17What last scored this pin
eval_versionjudge@2.1.0Who graded
job_type / locale / channelrefund / en-US / emailMix
Terminal reasonok / wrong_tool / no_progressWhether you had a loop, not a vibe
support_triage | model=gpt-5.6-terra | resolved=gpt-5.6-terra | prompt=support-triage@3.2.0 | tools=crm-tools@1.7.0 | index=kb-support@2026-06-10 | golden=2026-06-17 | effort=medium | job=refund | locale=en-US
RoleOwnsDoes not own
Job ownerPass criteria, “ship / hold” on the weekly scoreQuiet alias swaps
Engmodels.yml, CI denylist, log fieldsRewriting copy during an incident without an isolation note
On-callRollback to rollback_model_id from logsInventing a new pin at 2 a.m.
  • Identity line is in the trace, not only in the repo
  • Online sample and offline suite share evaluator identity
  • Provider changelog URLs are in the runbook (OpenAI, Anthropic, Gemini)
  • Someone is named on the weekly score

A pin that never gets a candidate bake-off is a pin you will rip out under a retirement deadline. OpenAI documents at least six months’ notice for GA model retirement after announcement; Anthropic documents at least sixty days before retirement for public models; Gemini publishes shutdown dates per ID and about two weeks’ notice on preview and -latest breaking changes (OpenAI, Anthropic deprecations, Gemini deprecations). Those are notice windows, not quality SLAs. Do not treat them as a promised pass rate.

What does a one-week triage look like?

If you only have a week, do not rewrite the prompt library. Instrument identity, split online vs offline, and walk the five diffs. A Spurlock $1,500 · 5-day agentic pilot is this shape: one job, a scored identity, a gate, an online sample.

Day-by-day, one job only:

  1. Day 1 — alibi and logs. Confirm prompt hash. Emit the identity line. Denylist aliases in CI. If you cannot log it today, you are not triaging quality. You are guessing.
  2. Day 2 — pin and schema. Diff model_id / resolved_id / schema hash for two weeks. Roll back any unscored pin. Stub tools and run write-tool cases.
  3. Day 3 — corpus and citations. Record index_hash. Run the citation slice. Freeze reindex until the slice is green.
  4. Day 4 — suite hygiene. Diff membership. Restore deleted hard cases. Harvest three live failures into git. Version the evaluator.
  5. Day 5 — mix and hold/ship. Compare online mix to golden labels. Canary only if identity is scored. Write the isolation note. Prompt edits, if any, are a separate PR.
If you skipWhat you will still not know on Friday
Identity logsWhether the pin moved
Argument-accuracy gateWhether the schema moved
Citation sliceWhether the corpus moved
Membership diffWhether the suite rotted
Online vs offlineWhether the customers changed
Prompt rewriteAnything useful
  • One job, not five agents
  • Prompt freeze until isolation is written
  • Rollback pin named
  • Three harvested failures in git
  • Weekly score owner named

You do not need multi-provider routing this week. You need to name what ran. If you cannot, do not scale the loop.

When is a workflow enough instead of hunting agent drift?

When the job is a known graph with known side effects, a workflow will not present as “quality dropped and we didn’t change prompts,” because it does not have an alias and a retriever to float. If your pain is a missing field in a CRM mapping, you do not have an agent-quality incident. You have an unversioned mapping.

SituationHunt the five causesUse a workflow instead
Model chooses tools, writes CRM, cites a corpusYesNo — you bought an agent, so buy identity
Fixed mapping, retries, webhooksMaybe the mapping, not the modelYes — automation shape
You cannot name model_id for last TuesdayInstrument firstDo not add more agent autonomy
No evaluator, no golden setBuild the evaluator firstDo not widen the loop
Same fingerprint retried until capLoop detector, then maybe qualityFix the harness (failed-tool loops)
One tool, one field, one destinationYou probably never needed an agentWorkflow

Decision list:

  1. If there is no identity line, stop. Log, then return.
  2. If there is no evaluator, stop. You cannot see a drop you cannot score.
  3. If the job is a mapping, delete the agent costume.
  4. If the job needs judgment plus tools, hunt the five causes before you edit copy.
  5. If two weeks of isolation still show “nothing moved,” look at serving infra and unlogged knobs — not a new system prompt.
  • Job owner can say “agent” or “workflow” without a slide
  • Autonomy widens only when golden pass, cost per pass, and escalate rate hold on an online sample
  • Prompt edits are a last step, not a first

Unchanged prompts are a true statement and a useless finding. Pin the versions. Diff the five causes. Then, and only then, touch a sentence.

FAQ

Why did quality drop when we didn’t change our prompts?

Because prompts are one axis. The drop almost always comes from a silent model swap, a tool schema change, retrieval corpus drift, eval set rot, or a traffic mix shift — often more than one. Confirm the prompt hash, then diff model_id, tools_version, index hash, eval membership, and online mix before you rewrite copy.

How do I measure whether the diagnosis is working?

You have a working diagnosis when the identity line explains the signals and a rollback or scored bump moves revision rate, argument accuracy, or wrong-source rate back toward the pre-drop band. Offline pass rate alone is not the measure — it can rise from eval rot. Use the same evaluator online and offline, and treat argument accuracy and citation correctness as hard gates.

What usually fails first when teams try this?

Logs. Teams rewrite the prompt because they can, then discover they cannot grep last Tuesday’s model_id. The next failure is treating pass rate as quality while hard cases quietly left the suite. Fix identity and membership before you hold a writing sprint.

How long does this take to show results?

A week of identity logs plus golden-vs-online split is enough to name the cause for one job. Repair time depends on which axis moved: a pin rollback can be a config PR the same day; a rotten suite or a bad reindex takes as long as restoring fixtures and snapshots. I will not invent a recovery percentage.

What should I skip if I only have a week?

Skip prompt rewrites, multi-provider routing, and a second agent. Instrument the identity line, denylist aliases, run write-tool and citation slices, diff eval membership, and compare online mix to golden labels. Harvest three live failures into git. That is the week.

When is this not worth doing yet?

When you have no evaluator, no production logs, and no pin — build those first, or you will be guessing. When the job is a fixed mapping, use a workflow instead of an agent and stop looking for model drift. When you already know the prompt did change, this diagnosis does not apply; score that prompt PR.

CTA

Want the five-cause isolation, pinned identity, and a golden-set gate on one real job this week? Start at /agentic or /contact?intent=agentic-pilot.

FAQ

What questions does this article answer?

Why did quality drop when we didn’t change our prompts?
Because prompts are one axis. The drop almost always comes from a silent model swap, a tool schema change, retrieval corpus drift, eval set rot, or a traffic mix shift — often more than one. Confirm the prompt hash, then diff `model_id`, `tools_version`, index hash, eval membership, and online mix before you rewrite copy.
How do I measure whether the diagnosis is working?
You have a working diagnosis when the identity line explains the signals and a rollback or scored bump moves revision rate, argument accuracy, or wrong-source rate back toward the pre-drop band. Offline pass rate alone is not the measure — it can rise from eval rot. Use the same evaluator online and offline, and treat argument accuracy and citation correctness as hard gates.
What usually fails first when teams try this?
Logs. Teams rewrite the prompt because they can, then discover they cannot grep last Tuesday’s `model_id`. The next failure is treating pass rate as quality while hard cases quietly left the suite. Fix identity and membership before you hold a writing sprint.
How long does this take to show results?
A week of identity logs plus golden-vs-online split is enough to name the cause for one job. Repair time depends on which axis moved: a pin rollback can be a config PR the same day; a rotten suite or a bad reindex takes as long as restoring fixtures and snapshots. I will not invent a recovery percentage.
What should I skip if I only have a week?
Skip prompt rewrites, multi-provider routing, and a second agent. Instrument the identity line, denylist aliases, run write-tool and citation slices, diff eval membership, and compare online mix to golden labels. Harvest three live failures into git. That is the week.
When is this not worth doing yet?
When you have no evaluator, no production logs, and no pin — build those first, or you will be guessing. When the job is a fixed mapping, use a workflow instead of an agent and stop looking for model drift. When you already know the prompt *did* change, this diagnosis does not apply; score that prompt PR.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot