Spurlock Studios
Contact
Share LinkedIn X
A violet ring. Thesis: PIN MODEL GATE UPGRADE CATCH.

Yes — pin the model version for production agents, and upgrade only through a golden-set gate. Floating aliases (gpt-5.6, latest, bare family nicknames) can change weights and defaults under you with no deploy. Customers notice first as “the agent got weird.” You notice last as a support spike.

This spoke sits under the Agentic Systems Operating Manual. The gate that makes a pin change a scored PR is golden sets from agent failures. Doctrine lives in the manual. Identity and the upgrade ritual live here.

The short answer

  • Pin explicit model IDs in prod. Store them in config, not string literals scattered across services.
  • Floating aliases are a drift channel. Behavior changes without a PR.
  • Upgrade = config PR + golden-set gate + canary. Never “swap the alias on Friday.”
  • Drift is not only model weights. Prompts, tools, retrieval indexes, sampling defaults, serving infra, and input mix move too — schedule checks even when nothing deployed.
  • Providers do not keep snapshots forever. Pinning buys stability until deprecation or retirement. Watch their notices.

What does model and prompt drift look like on a tool agent?

Chatbots hide drift in prose. Tool agents deposit it in CRM notes, ticket fields, and payment payloads. After 500+ automations and 20,000+ hours on agentic systems, the incidents I see most are not “the model got dumber.” They are identity strings that moved, or prompts that moved while the pin stayed put.

Drift typeWhat changedSymptom on a tool agent
Model alias floatProvider pointed the alias at new weights or a new defaultTool-choice shifts; verbosity jumps; cost per pass leaves the band
Sampling / reasoning defaultreasoning.effort, thinking, or temperature semantics changedSame ID, longer traces, new 400s, or sudden brevity
Prompt editSystem prompt or tool descriptions editedSame model, new failure codes
Tool schema driftEnums or required fields changedArgument errors, retry loops
Retrieval / memory driftIndex or memory policy changedConfident wrong citations
Input-distribution driftNew ticket types, seasons, localesOffline golden set still green; online revision rate up
Serving-infra driftRouter, safety classifier, or sampler changed under a pinned IDRare tone or refusal shifts with no config PR

Agents amplify small drifts. One extra tool call per turn compounds cost. One wrong enum compounds retries. If the only score you watch is “did the chat sound fine,” you will ship the wrong write.

  • I can name the exact model_id string production sent last Tuesday
  • I can name prompt_version and tools_version for that same run
  • I can show the last golden-set score for that pin
  • I can show online pass rate and cost per pass against a trailing baseline

If any box is empty, customers are already running the A/B test.

Why do floating aliases change behavior with no deploy?

Aliases exist for convenience. Production needs identity. If the model string in prod is not the string you evaluated, you are testing on customers.

Examples of float risk, checked against vendor docs in August 2026:

Alias / pointerWhat it actually doesProd risk
OpenAI gpt-5.6Routes to gpt-5.6-sol today (model guidance, Sol model page)Routing policy or the snapshot behind the alias can move; your evals never ran
OpenAI chat-latestPoints at the latest ChatGPT Instant snapshot; OpenAI says the underlying snapshot updates regularly (changelog)Chat-shaped, not agent-shaped. Do not put it on a write path
Gemini gemini-flash-latestHot-swapped on every new Flash release; breaking changes get about two weeks of email notice (model version patterns)The string stays still. The weights do not
Claude convenience aliases (opus, sonnet, older dateless pointers)Resolve to a recommended or latest dated snapshot for that line (model IDs and versions)Fine for a spike. Illegal in prod

Dateless is not the same as floating. On Claude, from the 4.6 generation on, a dateless ID such as claude-sonnet-5 is a pinned snapshot, not an evergreen pointer. Anthropic does not update the weights of an existing ID; a new release ships a new ID. That is the opposite of gpt-5.6 and gemini-flash-latest. Read the ID rules before you treat “no date in the string” as a warning or a blessing.

# Illegal in prod — CI should fail these
model: gpt-5.6
model: latest
model: gemini-flash-latest
model: sonnet
model: opus

Which model IDs should you pin in August 2026?

Prefer the most specific ID your provider documents as a pinned snapshot or stable model ID. Re-check the provider model page before you copy these into a deploy months later. Training data and this table both go stale.

OpenAI — GPT-5.6 family

Sources: OpenAI model guidance, GPT-5.6 Sol, models list (checked 2026-08-16).

RolePin this IDNotes
Flagship / hard agent reasoninggpt-5.6-solgpt-5.6 alias routes here — do not use the alias in prod
Balanced workergpt-5.6-terraCost and quality middle
High-volume / latency-sensitivegpt-5.6-lunaVolume tier

If the model page lists a more specific snapshot ID than the tier name, prefer the snapshot for regulated or high-stakes agents. Do not invent dated strings that are not on the page. Sol’s public card prices text at $5 / $30 per 1M input / output tokens as of this check — another reason volume classify should not sit on the flagship pin.

Pin reasoning.effort with the ID. GPT-5.6 accepts none, low, medium (default), high, xhigh, and max. An upgrade that keeps gpt-5.6-sol but changes effort is still an upgrade. GPT-5.6 also defaults persisted reasoning to all_turns; earlier families defaulted to current_turn. That default shift changes traces even when the family name looks familiar.

Anthropic — Claude 5-era pins

Sources: model IDs and versions, models overview, model deprecations (checked 2026-08-16). From the 4.6 generation onward, dateless IDs are pinned snapshots, not evergreen pointers. Convenience aliases can move — avoid them in prod.

RolePin this IDTentative retirement floor (Anthropic table)
Long-horizon / highest capabilityclaude-fable-5Not sooner than 2027-06-09
Complex agent / codingclaude-opus-5Not sooner than 2027-07-24
General productionclaude-sonnet-5Not sooner than 2027-06-30
Fast classify / extractclaude-haiku-4-5-20251001Not sooner than 2026-10-15

Haiku 4.5’s floor is weeks away from this review, not years. A “we pinned it, we are done” stance on that ID is a Q4 outage on the calendar. Older dated pins such as claude-sonnet-4-5-20250929 sit even closer (not sooner than 2026-09-29). If you still hold a pre-4.6 dated ID, open the migration ticket this week.

Opus 5 is not a drop-in string swap. Anthropic’s Opus 5 note: thinking is on by default, and disabling thinking with effort xhigh or max returns a 400. That belongs in the golden set as a contract test, not a surprise in prod.

Google — Gemini Flash workhorse

Sources: Gemini 3.6 Flash model card (GA 2026-07-21), model version patterns, Gemini deprecations, Gemini changelog.

RolePin this IDNotes
Agentic / coding workhorsegemini-3.6-flashDocumented stable ID; no separate dated suffix on the public card as of Jul 2026
Cost / latency Lite classgemini-3.5-flash-liteUse when the job earned the cheaper tier (GA the same day as 3.6 Flash)

Honest limit: Gemini’s public card for 3.6 Flash does not expose a YYYY-MM-DD snapshot the way older stacks did. Pin the stable ID you evaluated, track the deprecations table and changelog, and re-run the golden set when they announce a replacement. Do not pretend a dated pin exists if the docs do not list one. Do not use gemini-flash-latest in prod.

Google’s own version page: stable IDs “usually don’t change”; -latest aliases get hot-swapped; preview IDs may be retired with at least two weeks’ notice. As of this check, gemini-3.6-flash has no shutdown date announced. That is not a forever SLA. It is “no date yet.”

How long do providers keep pinned snapshots?

Do not invent retention SLAs. Use the published notice policies and watch the deprecation tables.

ProviderWhat they commit (as of Aug 2026 docs)Practical meaning
OpenAIFor GA models, at least 6 months notice before retirement after deprecation announcement; specialized variants ≥3 months; previews may be ~2 weeks (deprecations)Pinning ≠ forever. Calendar the shutdown date when it appears. Safety or compliance can shorten the window
AnthropicAt least 60 days notice before retirement for publicly released models; retired IDs fail (model deprecations)Weights may be preserved internally (deprecation commitments); API access is not. Migrate before the retirement date
Google GeminiShutdown dates live on the deprecations page; preview / -latest breaking changes: about 2 weeks (version patterns)Treat changelog + golden-set canaries as the control, not an assumed multi-year pin

Hedged on purpose: third-party blogs sometimes claim “12 months” or other windows. Prefer the provider deprecation page over folklore. Partner clouds (Bedrock, Vertex) set their own retirement schedules; a Claude ID that is Active on the Claude API can already be on a different clock in a partner catalog.

OpenAI’s June 11, 2026 notice is the shape you should expect: dated GPT-5 and o3 snapshots were given a December 11, 2026 shutdown, with replacements named as gpt-5.6-sol / terra / luna. That is six months on the calendar. The replacement still has to pass your golden set. A recommended ID is not a passed eval.

What else must you pin besides the model ID?

A pinned ID with floating knobs is a floating model. Upgrade gates fail in boring ways when the string matches and the request body does not.

KnobWhy it belongs next to model_idFailure if it floats
reasoning.effort (OpenAI GPT-5.6)Same ID, six effort levels, different tracesCost and latency jump; tool loops lengthen
Thinking / effort (Claude Opus 5)Thinking on by default; some disable combos 400Prod 400s after a “simple” pin flip
Sampling paramsClaude Opus 4.7+ returns 400 on non-default temperature / top_p / top_k; Gemini 3.6 Flash deprecates the same trio (changelog)Silent ignore or hard error, depending on vendor
prompt_versionWording is behaviorNew refusals, new tool choice
tools_versionSchema is behaviorArgument errors
Retrieval index hashCitations come from the index, not the pinConfident wrong sources
# Pin the request contract, not only the ID
agents:
  support_triage:
    model_id: gpt-5.6-terra
    reasoning_effort: medium
    prompt_version: support-triage@3.2.0
    tools_version: crm-tools@1.7.0
  • Model ID is a snapshot or stable ID, not an alias
  • Reasoning / thinking / sampling are explicit in config
  • Prompt and tool versions are explicit
  • CI diffs the whole request contract, not only model_id

Is staging on a float and prod on a pin a good pattern?

Yes — if staging evals run on a schedule and a human owns failures. Prod stays pinned. Staging float without alerts is how you learn about provider changes from a social feed instead of CI.

EnvironmentModel stringJob
Local / spikeAlias OKExplore
StagingCandidate pin or intentional float + nightly evalDetect provider moves
ProductionPinned ID onlyStable behavior

Staging-on-float only works if something reads the nightly eval. A floating staging env nobody watches is theater.

Good staging float jobs:

  1. Resolve the alias to the current underlying ID (log both strings).
  2. Run the same golden set you run on the prod pin.
  3. Diff failure codes and cost per pass, not only pass rate.
  4. Open a ticket the same day the alias moves or the score leaves the band.

If you cannot resolve an alias to an underlying ID, treat the alias as opaque and do not use it as a control. Gemini -latest is the extreme case: the string never changes, so your logs lie unless you also record provider changelog dates.

What is the upgrade procedure that does not lie?

Upgrade is a change to identity. Treat it like a schema migration, not like a dashboard toggle.

  1. Pick the candidate ID from provider docs (not a social post).
  2. Freeze prompts, tools, stubs, and sampling knobs.
  3. Run the full golden set plus the cost band against the current pin.
  4. Diff failure codes — not only pass rate.
  5. Canary a small online share with revision rate and cost per pass watched.
  6. Merge the config PR that changes the pin.
  7. Keep the old pin in a rollback map for 48–72 hours.

Gate rule example (tune to your risk):

SignalShipHold
Offline pass rate≥ baseline − 1 ptDrop > 1 pt
Argument accuracy≥ baselineAny drop on write tools
Cost per pass≤ baseline × 1.15Above
New failure codesNone criticalAny policy / wrong_tool surge
New HTTP 400s from sampling / thinkingZeroAny

A recommended replacement on a deprecation page is a candidate, not a pass. OpenAI can tell you to move gpt-5-2025-08-07 to gpt-5.6-sol. Your write-tool argument accuracy can still drop. That is the whole point of the gate.

Own the suite in git. OpenAI announced deprecation of its Evals platform on 2026-06-03, with a read-only window starting 2026-10-31 and a scheduled shutdown on 2026-11-30 (deprecations). If your only regression memory lives in a vendor dashboard, the vendor can retire the dashboard. Fixtures you harvested from production failures survive that.

What belongs in SemVer for prompts versus model pins?

Keep separate version axes. Incident response needs to answer “which weights?” in one query. An opaque “agent version 42” cannot.

AxisExampleBumps when
prompt_versionsupport-triage@3.2.0Wording, tool descriptions, rubrics
model_idclaude-sonnet-5Pin change
tools_versioncrm-tools@1.7.0Schema / handler contract
eval_suitegolden@2026-03-12Cases added or removed
request_contractreasoning.effort=mediumSampling / thinking / effort

Rules that keep the axes honest:

  1. A model pin change does not require a prompt SemVer bump — and must still pass the golden set.
  2. A prompt bump with the same pin still requires the suite. Same weights, new instructions, new failures.
  3. A tools bump is a contract change. Score argument accuracy before you celebrate a higher pass rate.
  4. Never bury model_id inside a Docker tag or a marketing “agent v4.”
# Readable in an incident channel
support_triage | model=claude-sonnet-5 | prompt=support-triage@3.2.0 | tools=crm-tools@1.7.0 | golden=2026-03-12

If on-call cannot paste that line from logs in thirty seconds, the version scheme is decoration.

How do you schedule drift checks when nothing deployed?

Calendar, not vibes. Input-distribution drift shows up online first. Offline-only teams learn from angry humans.

CadenceWhat you runWhat a miss costs
WeeklyOnline sample pass rate, revision rate, cost per pass vs trailing baselineQuiet cost creep; “it got wordy” tickets
On provider email / changelogOpen an upgrade ticket the same dayYou start the 60-day or 6-month clock late
MonthlyStaging float vs prod pin bake-off on the golden setYou discover the alias moved from a customer, not from CI
After any tool or prompt PRFull suite — pin unchangedYou ship prompt drift and blame the model
QuarterlyCandidate upgrade on the next Active IDForced retirement becomes a rewrite
  • Provider changelog URLs are in the runbook (OpenAI, Anthropic, Gemini)
  • Deprecation-table review is a recurring ticket, not a memory
  • Online sampling uses the same evaluator as the offline suite
  • Someone is named on the weekly score, not “the team”

A pin that never gets a candidate bake-off is a pin you will rip out under a retirement deadline. The operating manual already treats evaluators as load-bearing. This spoke adds the calendar that keeps those evaluators pointed at identity, not only at prompts.

What belongs in the golden-set gate for a pin change?

A pin change is the cheapest time to find that the new ID loves a different tool, a longer trace, or a sloppier enum. Harvested failures beat synthetic demos. Build the rows the way the golden-set spoke describes: inputs, stubbed tools, expected terminal verdict.

Minimum slices for a model-upgrade PR:

SliceWhy it existsHold if
Write-tool argument accuracyNew IDs miss required fields in new waysAny drop vs current pin
Policy / deny casesNew IDs get braver or more timidA deny case becomes an allow
Cost-band casesFlagship pins are expensive; effort defaults moveCost per pass > 1.15× baseline
HTTP contract casesThinking / sampling 400sAny new 400 on a fixture that passed
Citation / retrieval casesSame pin, different index — or a model that ignores the indexConfident wrong source
# CI shape — pin change is just another scored diff
golden --suite support-triage --baseline-model claude-sonnet-5 --candidate-model claude-opus-5
# fail on: pass_rate, arg_accuracy, cost_per_pass, new_codes

Do not let the candidate “win” on pass rate while losing on write-tool arguments. Pass rate is a vanity metric when the write is wrong. Stub the tools so CI does not call live CRM. Anonymize before the fixture lands in git.

What still moves when the model ID is pinned?

Pinning is necessary. It is not a freeze-frame of the whole system.

Anthropic is explicit: weights stay fixed for a given ID, but serving infrastructure — request router, safety classifiers, sampling logic — can still change, and you may see minor behavior shifts (model IDs and versions). Google says stable IDs “usually don’t change,” which is not “never.” OpenAI’s alias story is the louder version of the same industry habit: convenience names move; even some “stable” serving paths get ops patches.

So you still need online sampling. A pinned ID with a rotting index, a rewritten tool description, or a new season of tickets will drift. The pin removes the cheapest, dumbest channel. It does not retire the weekly score.

Still moves on a pinned IDHow you catch itWhat you do not do
Serving-infra tweaksOnline sample + weekly goldenBlame “the model got worse” and silently swap aliases
Prompt / tool editsSuite on every PRShip on Friday and score on Monday
Index / memory policyCitation slice + online wrong-source rateRe-embed production without a fixture
Input mixOnline vs offline divergenceRetrain the prompt on anecdotes

If online pass or revision rate diverges from offline while the pin is unchanged, investigate distribution and tools first. Swapping the pin to “fix” a seasonal ticket mix is how you get two problems.

Worked failure: the quiet alias weekend

What broke: Prod used gpt-5.6 “so we always get the best.” A routing or default change shifted tool verbosity. Average tool calls per ticket left the band. Pass rate dipped. Cost per pass jumped. No deploy in git. The identity string moved; the repo did not.

Cost: Budget alerts. Weekend rollback to an explicit gpt-5.6-terra pin. Emergency golden-set triage while support ate the weird writes.

Instead: Prod pin gpt-5.6-sol or gpt-5.6-terra by role. Staging tracks the alias and logs the resolved ID. Upgrade only through the gate. CI denylists gpt-5.6.

The model did not “get dumber.” Your identity string did. I will not dress that weekend up as a benchmark. The receipt is the missing deploy and the missing gate.

What breaks if you never unpin?

Pinning without a migration calendar delays a worse incident.

FailureHow it shows upPrevention
Retirement hard-failAnthropic retired IDs return errors; OpenAI shutdown dates are published; Gemini endpoints turn off on the deprecations tableCalendar the date the day the notice lands
Prompt debtPatches that only paper over old-model quirksCandidate upgrade quarterly so patches stay portable
Forced big-bangCompetitors already scored the next Active ID; you migrate under a deadlineKeep a rollback pin and a scored candidate
Partner-cloud skewBedrock or Vertex retires the same Claude name on a different clockTrack the catalog you actually call

Anthropic’s own history on the deprecations page is the warning label: dated 4.x IDs have already retired on the Claude API in 2026. If you still hold claude-haiku-4-5-20251001 as a volume pin, the tentative floor is 2026-10-15. That is a pin with an expiration date, not a personality.

Pair every prod pin with a written successor candidate and a last-eval date. If last_eval_passed is older than 90 days, the pin is a rumor.

What config shape survives a review?

# models.yml — reviewed in PRs
agents:
  support_triage:
    provider: anthropic
    model_id: claude-sonnet-5          # pinned snapshot ID
    reasoning_effort: null             # unused on this pin; keep the key
    prompt_version: support-triage@3.2.0
    tools_version: crm-tools@1.7.0
    eval_suite: golden@2026-03-12
    last_eval_passed: 2026-03-10
    rollback_model_id: claude-sonnet-4-6
    successor_candidate: claude-opus-5

CI fails if model_id matches a denylist of aliases (latest, gpt-5.6, gemini-flash-latest, sonnet, opus). CI also fails if last_eval_passed is missing or older than your SLA.

Review questions a stranger should be able to answer from this file:

  1. What string does production send?
  2. What string do we roll back to?
  3. When did that pin last pass the suite?
  4. What candidate are we scoring next?

If the answers live in a Slack thread, you do not have a pin. You have folklore.

Who owns the pin, and what must every run log?

A pin without an owner expires in silence. Split the work so the person who likes new models is not the only person who can flip prod.

RoleOwnsDoes not own
Job ownerPass criteria, deny cases, “ship / hold” on the scoreThe raw model string in a service
Engmodels.yml, CI denylist, rollback mapQuiet alias swaps “to see if it is better”
On-callRollback to rollback_model_id from logsInventing a new pin during an incident

Every production run should emit a single identity line. If you cannot grep last Tuesday’s writes by model_id, you cannot prove what drifted.

FieldExampleWhy
model_idgpt-5.6-terraWeights you think you bought
resolved_idgpt-5.6-terraSame as model_id on a pin; different if staging floated
prompt_versionsupport-triage@3.2.0Instructions
tools_versioncrm-tools@1.7.0Schema
request_contractreasoning.effort=mediumKnobs
eval_suitegolden@2026-03-12What last scored this pin
  • Job owner named on the weekly score
  • Eng named on the denylist and the rollback map
  • On-call can flip to rollback_model_id without a design review
  • Identity line is in the trace, not only in the repo

What is the pilot minimum?

A Spurlock $1,500 · 5-day agentic pilot ships:

  1. One pinned model_id per agent role in config, with a rollback ID
  2. A golden-set gate wired so a pin change is a scored PR
  3. An online sampling panel that would have caught the quiet alias weekend

You do not need multi-provider routing on day one. You need identity and a gate. The rest of the stack — evaluator, sandbox, kill switch — is in the operating manual. This week answers “can we name the weights and prove an upgrade?” If you cannot, do not scale the loop.

  • Prod strings are snapshot or stable IDs
  • Aliases are denylisted in CI
  • Golden set runs on pin diffs
  • Weekly online sample exists
  • Deprecation URLs are in the runbook

FAQ

Which OpenAI / Anthropic / Gemini IDs should I pin today?

As of August 2026 docs: OpenAI gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna by role (avoid the gpt-5.6 alias in prod); Anthropic claude-fable-5, claude-opus-5, claude-sonnet-5, or dated claude-haiku-4-5-20251001; Gemini gemini-3.6-flash (and gemini-3.5-flash-lite when the job earned the cheaper tier). Re-verify on provider model pages before hardcoding months later. Dateless Claude 4.6+ IDs are snapshots; OpenAI and Gemini family aliases are not.

How long do providers keep pinned snapshots?

Until they deprecate and retire them — not forever. OpenAI documents at least six months’ notice for GA model retirement after announcement; Anthropic documents at least sixty days before retirement for public models. Gemini publishes shutdown dates per ID and gives about two weeks’ notice on preview and -latest breaking changes. Track the deprecation tables. Do not assume multi-year API access.

Staging on floating vs prod on pinned — good pattern?

Yes, if staging evals run on a schedule and someone owns failures. Prod stays pinned. Staging should log the resolved underlying ID when the alias moves. Staging float without alerts is how you learn about provider changes from a social feed instead of CI.

What belongs in SemVer for prompts vs model pins?

Version prompts and rubrics (prompt_version), tool contracts (tools_version), and the eval suite separately from model_id. A model pin change is a config change that must pass the golden-set gate even when the prompt SemVer does not bump. Incident logs should print all four, plus any reasoning or thinking knob you pinned.

How does online sampling catch input-distribution drift?

Offline goldens freeze yesterday’s tickets. Online samples score today’s mix with the same evaluator. When online pass or revision rates diverge from offline, you are seeing distribution drift — not necessarily a bad pin. Investigate mix, tools, and indexes before you “fix” the model.

What breaks if I never unpin?

Forced retirement outages, prompt debt that only works on the old pin, and a painful big-bang migration. Dated Claude pins already have 2026 retirement floors measured in weeks. Pin for stability. Schedule candidate upgrades so unpinning is a controlled PR, not an incident.

CTA

Want pins, gates, and drift panels on a real agent job in five days? Start at /agentic or /contact?intent=agentic-pilot.

FAQ

What questions does this article answer?

Which OpenAI / Anthropic / Gemini IDs should I pin today?
As of August 2026 docs: OpenAI `gpt-5.6-sol` / `gpt-5.6-terra` / `gpt-5.6-luna` by role (avoid the `gpt-5.6` alias in prod); Anthropic `claude-fable-5`, `claude-opus-5`, `claude-sonnet-5`, or dated `claude-haiku-4-5-20251001`; Gemini `gemini-3.6-flash` (and `gemini-3.5-flash-lite` when the job earned the cheaper tier). Re-verify on provider model pages before hardcoding months later. Dateless Claude 4.6+ IDs are snapshots; OpenAI and Gemini family aliases are not.
How long do providers keep pinned snapshots?
Until they deprecate and retire them — not forever. OpenAI documents at least six months’ notice for GA model retirement after announcement; Anthropic documents at least sixty days before retirement for public models. Gemini publishes shutdown dates per ID and gives about two weeks’ notice on preview and `-latest` breaking changes. Track the deprecation tables. Do not assume multi-year API access.
Staging on floating vs prod on pinned — good pattern?
Yes, if staging evals run on a schedule and someone owns failures. Prod stays pinned. Staging should log the resolved underlying ID when the alias moves. Staging float without alerts is how you learn about provider changes from a social feed instead of CI.
What belongs in SemVer for prompts vs model pins?
Version prompts and rubrics (`prompt_version`), tool contracts (`tools_version`), and the eval suite separately from `model_id`. A model pin change is a config change that must pass the golden-set gate even when the prompt SemVer does not bump. Incident logs should print all four, plus any reasoning or thinking knob you pinned.
How does online sampling catch input-distribution drift?
Offline goldens freeze yesterday’s tickets. Online samples score today’s mix with the same evaluator. When online pass or revision rates diverge from offline, you are seeing distribution drift — not necessarily a bad pin. Investigate mix, tools, and indexes before you “fix” the model.
What breaks if I never unpin?
Forced retirement outages, prompt debt that only works on the old pin, and a painful big-bang migration. Dated Claude pins already have 2026 retirement floors measured in weeks. Pin for stability. Schedule candidate upgrades so unpinning is a controlled PR, not an incident.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot