Pin the Model, Gate the Upgrade: Catch Agent Drift Before Customers Do
Yes—pin production agents to explicit model IDs. Floating aliases change behavior with no deploy. Upgrade only through a golden-set gate you actually run.
William Spurlock Founder — Spurlock Studios Updated 20 MIN
Yes — pin the model version for production agents, and upgrade only through a golden-set gate. Floating aliases (gpt-5.6, latest, bare family nicknames) can change weights and defaults under you with no deploy. Customers notice first as “the agent got weird.” You notice last as a support spike.
This spoke sits under the Agentic Systems Operating Manual. The gate that makes a pin change a scored PR is golden sets from agent failures. Doctrine lives in the manual. Identity and the upgrade ritual live here.
The short answer
- Pin explicit model IDs in prod. Store them in config, not string literals scattered across services.
- Floating aliases are a drift channel. Behavior changes without a PR.
- Upgrade = config PR + golden-set gate + canary. Never “swap the alias on Friday.”
- Drift is not only model weights. Prompts, tools, retrieval indexes, sampling defaults, serving infra, and input mix move too — schedule checks even when nothing deployed.
- Providers do not keep snapshots forever. Pinning buys stability until deprecation or retirement. Watch their notices.
What does model and prompt drift look like on a tool agent?
Chatbots hide drift in prose. Tool agents deposit it in CRM notes, ticket fields, and payment payloads. After 500+ automations and 20,000+ hours on agentic systems, the incidents I see most are not “the model got dumber.” They are identity strings that moved, or prompts that moved while the pin stayed put.
| Drift type | What changed | Symptom on a tool agent |
|---|---|---|
| Model alias float | Provider pointed the alias at new weights or a new default | Tool-choice shifts; verbosity jumps; cost per pass leaves the band |
| Sampling / reasoning default | reasoning.effort, thinking, or temperature semantics changed | Same ID, longer traces, new 400s, or sudden brevity |
| Prompt edit | System prompt or tool descriptions edited | Same model, new failure codes |
| Tool schema drift | Enums or required fields changed | Argument errors, retry loops |
| Retrieval / memory drift | Index or memory policy changed | Confident wrong citations |
| Input-distribution drift | New ticket types, seasons, locales | Offline golden set still green; online revision rate up |
| Serving-infra drift | Router, safety classifier, or sampler changed under a pinned ID | Rare tone or refusal shifts with no config PR |
Agents amplify small drifts. One extra tool call per turn compounds cost. One wrong enum compounds retries. If the only score you watch is “did the chat sound fine,” you will ship the wrong write.
- I can name the exact
model_idstring production sent last Tuesday - I can name
prompt_versionandtools_versionfor that same run - I can show the last golden-set score for that pin
- I can show online pass rate and cost per pass against a trailing baseline
If any box is empty, customers are already running the A/B test.
Why do floating aliases change behavior with no deploy?
Aliases exist for convenience. Production needs identity. If the model string in prod is not the string you evaluated, you are testing on customers.
Examples of float risk, checked against vendor docs in August 2026:
| Alias / pointer | What it actually does | Prod risk |
|---|---|---|
OpenAI gpt-5.6 | Routes to gpt-5.6-sol today (model guidance, Sol model page) | Routing policy or the snapshot behind the alias can move; your evals never ran |
OpenAI chat-latest | Points at the latest ChatGPT Instant snapshot; OpenAI says the underlying snapshot updates regularly (changelog) | Chat-shaped, not agent-shaped. Do not put it on a write path |
Gemini gemini-flash-latest | Hot-swapped on every new Flash release; breaking changes get about two weeks of email notice (model version patterns) | The string stays still. The weights do not |
Claude convenience aliases (opus, sonnet, older dateless pointers) | Resolve to a recommended or latest dated snapshot for that line (model IDs and versions) | Fine for a spike. Illegal in prod |
Dateless is not the same as floating. On Claude, from the 4.6 generation on, a dateless ID such as claude-sonnet-5 is a pinned snapshot, not an evergreen pointer. Anthropic does not update the weights of an existing ID; a new release ships a new ID. That is the opposite of gpt-5.6 and gemini-flash-latest. Read the ID rules before you treat “no date in the string” as a warning or a blessing.
# Illegal in prod — CI should fail these
model: gpt-5.6
model: latest
model: gemini-flash-latest
model: sonnet
model: opus
Which model IDs should you pin in August 2026?
Prefer the most specific ID your provider documents as a pinned snapshot or stable model ID. Re-check the provider model page before you copy these into a deploy months later. Training data and this table both go stale.
OpenAI — GPT-5.6 family
Sources: OpenAI model guidance, GPT-5.6 Sol, models list (checked 2026-08-16).
| Role | Pin this ID | Notes |
|---|---|---|
| Flagship / hard agent reasoning | gpt-5.6-sol | gpt-5.6 alias routes here — do not use the alias in prod |
| Balanced worker | gpt-5.6-terra | Cost and quality middle |
| High-volume / latency-sensitive | gpt-5.6-luna | Volume tier |
If the model page lists a more specific snapshot ID than the tier name, prefer the snapshot for regulated or high-stakes agents. Do not invent dated strings that are not on the page. Sol’s public card prices text at $5 / $30 per 1M input / output tokens as of this check — another reason volume classify should not sit on the flagship pin.
Pin reasoning.effort with the ID. GPT-5.6 accepts none, low, medium (default), high, xhigh, and max. An upgrade that keeps gpt-5.6-sol but changes effort is still an upgrade. GPT-5.6 also defaults persisted reasoning to all_turns; earlier families defaulted to current_turn. That default shift changes traces even when the family name looks familiar.
Anthropic — Claude 5-era pins
Sources: model IDs and versions, models overview, model deprecations (checked 2026-08-16). From the 4.6 generation onward, dateless IDs are pinned snapshots, not evergreen pointers. Convenience aliases can move — avoid them in prod.
| Role | Pin this ID | Tentative retirement floor (Anthropic table) |
|---|---|---|
| Long-horizon / highest capability | claude-fable-5 | Not sooner than 2027-06-09 |
| Complex agent / coding | claude-opus-5 | Not sooner than 2027-07-24 |
| General production | claude-sonnet-5 | Not sooner than 2027-06-30 |
| Fast classify / extract | claude-haiku-4-5-20251001 | Not sooner than 2026-10-15 |
Haiku 4.5’s floor is weeks away from this review, not years. A “we pinned it, we are done” stance on that ID is a Q4 outage on the calendar. Older dated pins such as claude-sonnet-4-5-20250929 sit even closer (not sooner than 2026-09-29). If you still hold a pre-4.6 dated ID, open the migration ticket this week.
Opus 5 is not a drop-in string swap. Anthropic’s Opus 5 note: thinking is on by default, and disabling thinking with effort xhigh or max returns a 400. That belongs in the golden set as a contract test, not a surprise in prod.
Google — Gemini Flash workhorse
Sources: Gemini 3.6 Flash model card (GA 2026-07-21), model version patterns, Gemini deprecations, Gemini changelog.
| Role | Pin this ID | Notes |
|---|---|---|
| Agentic / coding workhorse | gemini-3.6-flash | Documented stable ID; no separate dated suffix on the public card as of Jul 2026 |
| Cost / latency Lite class | gemini-3.5-flash-lite | Use when the job earned the cheaper tier (GA the same day as 3.6 Flash) |
Honest limit: Gemini’s public card for 3.6 Flash does not expose a YYYY-MM-DD snapshot the way older stacks did. Pin the stable ID you evaluated, track the deprecations table and changelog, and re-run the golden set when they announce a replacement. Do not pretend a dated pin exists if the docs do not list one. Do not use gemini-flash-latest in prod.
Google’s own version page: stable IDs “usually don’t change”; -latest aliases get hot-swapped; preview IDs may be retired with at least two weeks’ notice. As of this check, gemini-3.6-flash has no shutdown date announced. That is not a forever SLA. It is “no date yet.”
How long do providers keep pinned snapshots?
Do not invent retention SLAs. Use the published notice policies and watch the deprecation tables.
| Provider | What they commit (as of Aug 2026 docs) | Practical meaning |
|---|---|---|
| OpenAI | For GA models, at least 6 months notice before retirement after deprecation announcement; specialized variants ≥3 months; previews may be ~2 weeks (deprecations) | Pinning ≠ forever. Calendar the shutdown date when it appears. Safety or compliance can shorten the window |
| Anthropic | At least 60 days notice before retirement for publicly released models; retired IDs fail (model deprecations) | Weights may be preserved internally (deprecation commitments); API access is not. Migrate before the retirement date |
| Google Gemini | Shutdown dates live on the deprecations page; preview / -latest breaking changes: about 2 weeks (version patterns) | Treat changelog + golden-set canaries as the control, not an assumed multi-year pin |
Hedged on purpose: third-party blogs sometimes claim “12 months” or other windows. Prefer the provider deprecation page over folklore. Partner clouds (Bedrock, Vertex) set their own retirement schedules; a Claude ID that is Active on the Claude API can already be on a different clock in a partner catalog.
OpenAI’s June 11, 2026 notice is the shape you should expect: dated GPT-5 and o3 snapshots were given a December 11, 2026 shutdown, with replacements named as gpt-5.6-sol / terra / luna. That is six months on the calendar. The replacement still has to pass your golden set. A recommended ID is not a passed eval.
What else must you pin besides the model ID?
A pinned ID with floating knobs is a floating model. Upgrade gates fail in boring ways when the string matches and the request body does not.
| Knob | Why it belongs next to model_id | Failure if it floats |
|---|---|---|
reasoning.effort (OpenAI GPT-5.6) | Same ID, six effort levels, different traces | Cost and latency jump; tool loops lengthen |
| Thinking / effort (Claude Opus 5) | Thinking on by default; some disable combos 400 | Prod 400s after a “simple” pin flip |
| Sampling params | Claude Opus 4.7+ returns 400 on non-default temperature / top_p / top_k; Gemini 3.6 Flash deprecates the same trio (changelog) | Silent ignore or hard error, depending on vendor |
prompt_version | Wording is behavior | New refusals, new tool choice |
tools_version | Schema is behavior | Argument errors |
| Retrieval index hash | Citations come from the index, not the pin | Confident wrong sources |
# Pin the request contract, not only the ID
agents:
support_triage:
model_id: gpt-5.6-terra
reasoning_effort: medium
prompt_version: support-triage@3.2.0
tools_version: crm-tools@1.7.0
- Model ID is a snapshot or stable ID, not an alias
- Reasoning / thinking / sampling are explicit in config
- Prompt and tool versions are explicit
- CI diffs the whole request contract, not only
model_id
Is staging on a float and prod on a pin a good pattern?
Yes — if staging evals run on a schedule and a human owns failures. Prod stays pinned. Staging float without alerts is how you learn about provider changes from a social feed instead of CI.
| Environment | Model string | Job |
|---|---|---|
| Local / spike | Alias OK | Explore |
| Staging | Candidate pin or intentional float + nightly eval | Detect provider moves |
| Production | Pinned ID only | Stable behavior |
Staging-on-float only works if something reads the nightly eval. A floating staging env nobody watches is theater.
Good staging float jobs:
- Resolve the alias to the current underlying ID (log both strings).
- Run the same golden set you run on the prod pin.
- Diff failure codes and cost per pass, not only pass rate.
- Open a ticket the same day the alias moves or the score leaves the band.
If you cannot resolve an alias to an underlying ID, treat the alias as opaque and do not use it as a control. Gemini -latest is the extreme case: the string never changes, so your logs lie unless you also record provider changelog dates.
What is the upgrade procedure that does not lie?
Upgrade is a change to identity. Treat it like a schema migration, not like a dashboard toggle.
- Pick the candidate ID from provider docs (not a social post).
- Freeze prompts, tools, stubs, and sampling knobs.
- Run the full golden set plus the cost band against the current pin.
- Diff failure codes — not only pass rate.
- Canary a small online share with revision rate and cost per pass watched.
- Merge the config PR that changes the pin.
- Keep the old pin in a rollback map for 48–72 hours.
Gate rule example (tune to your risk):
| Signal | Ship | Hold |
|---|---|---|
| Offline pass rate | ≥ baseline − 1 pt | Drop > 1 pt |
| Argument accuracy | ≥ baseline | Any drop on write tools |
| Cost per pass | ≤ baseline × 1.15 | Above |
| New failure codes | None critical | Any policy / wrong_tool surge |
| New HTTP 400s from sampling / thinking | Zero | Any |
A recommended replacement on a deprecation page is a candidate, not a pass. OpenAI can tell you to move gpt-5-2025-08-07 to gpt-5.6-sol. Your write-tool argument accuracy can still drop. That is the whole point of the gate.
Own the suite in git. OpenAI announced deprecation of its Evals platform on 2026-06-03, with a read-only window starting 2026-10-31 and a scheduled shutdown on 2026-11-30 (deprecations). If your only regression memory lives in a vendor dashboard, the vendor can retire the dashboard. Fixtures you harvested from production failures survive that.
What belongs in SemVer for prompts versus model pins?
Keep separate version axes. Incident response needs to answer “which weights?” in one query. An opaque “agent version 42” cannot.
| Axis | Example | Bumps when |
|---|---|---|
prompt_version | support-triage@3.2.0 | Wording, tool descriptions, rubrics |
model_id | claude-sonnet-5 | Pin change |
tools_version | crm-tools@1.7.0 | Schema / handler contract |
eval_suite | golden@2026-03-12 | Cases added or removed |
request_contract | reasoning.effort=medium | Sampling / thinking / effort |
Rules that keep the axes honest:
- A model pin change does not require a prompt SemVer bump — and must still pass the golden set.
- A prompt bump with the same pin still requires the suite. Same weights, new instructions, new failures.
- A tools bump is a contract change. Score argument accuracy before you celebrate a higher pass rate.
- Never bury
model_idinside a Docker tag or a marketing “agent v4.”
# Readable in an incident channel
support_triage | model=claude-sonnet-5 | prompt=support-triage@3.2.0 | tools=crm-tools@1.7.0 | golden=2026-03-12
If on-call cannot paste that line from logs in thirty seconds, the version scheme is decoration.
How do you schedule drift checks when nothing deployed?
Calendar, not vibes. Input-distribution drift shows up online first. Offline-only teams learn from angry humans.
| Cadence | What you run | What a miss costs |
|---|---|---|
| Weekly | Online sample pass rate, revision rate, cost per pass vs trailing baseline | Quiet cost creep; “it got wordy” tickets |
| On provider email / changelog | Open an upgrade ticket the same day | You start the 60-day or 6-month clock late |
| Monthly | Staging float vs prod pin bake-off on the golden set | You discover the alias moved from a customer, not from CI |
| After any tool or prompt PR | Full suite — pin unchanged | You ship prompt drift and blame the model |
| Quarterly | Candidate upgrade on the next Active ID | Forced retirement becomes a rewrite |
- Provider changelog URLs are in the runbook (OpenAI, Anthropic, Gemini)
- Deprecation-table review is a recurring ticket, not a memory
- Online sampling uses the same evaluator as the offline suite
- Someone is named on the weekly score, not “the team”
A pin that never gets a candidate bake-off is a pin you will rip out under a retirement deadline. The operating manual already treats evaluators as load-bearing. This spoke adds the calendar that keeps those evaluators pointed at identity, not only at prompts.
What belongs in the golden-set gate for a pin change?
A pin change is the cheapest time to find that the new ID loves a different tool, a longer trace, or a sloppier enum. Harvested failures beat synthetic demos. Build the rows the way the golden-set spoke describes: inputs, stubbed tools, expected terminal verdict.
Minimum slices for a model-upgrade PR:
| Slice | Why it exists | Hold if |
|---|---|---|
| Write-tool argument accuracy | New IDs miss required fields in new ways | Any drop vs current pin |
| Policy / deny cases | New IDs get braver or more timid | A deny case becomes an allow |
| Cost-band cases | Flagship pins are expensive; effort defaults move | Cost per pass > 1.15× baseline |
| HTTP contract cases | Thinking / sampling 400s | Any new 400 on a fixture that passed |
| Citation / retrieval cases | Same pin, different index — or a model that ignores the index | Confident wrong source |
# CI shape — pin change is just another scored diff
golden --suite support-triage --baseline-model claude-sonnet-5 --candidate-model claude-opus-5
# fail on: pass_rate, arg_accuracy, cost_per_pass, new_codes
Do not let the candidate “win” on pass rate while losing on write-tool arguments. Pass rate is a vanity metric when the write is wrong. Stub the tools so CI does not call live CRM. Anonymize before the fixture lands in git.
What still moves when the model ID is pinned?
Pinning is necessary. It is not a freeze-frame of the whole system.
Anthropic is explicit: weights stay fixed for a given ID, but serving infrastructure — request router, safety classifiers, sampling logic — can still change, and you may see minor behavior shifts (model IDs and versions). Google says stable IDs “usually don’t change,” which is not “never.” OpenAI’s alias story is the louder version of the same industry habit: convenience names move; even some “stable” serving paths get ops patches.
So you still need online sampling. A pinned ID with a rotting index, a rewritten tool description, or a new season of tickets will drift. The pin removes the cheapest, dumbest channel. It does not retire the weekly score.
| Still moves on a pinned ID | How you catch it | What you do not do |
|---|---|---|
| Serving-infra tweaks | Online sample + weekly golden | Blame “the model got worse” and silently swap aliases |
| Prompt / tool edits | Suite on every PR | Ship on Friday and score on Monday |
| Index / memory policy | Citation slice + online wrong-source rate | Re-embed production without a fixture |
| Input mix | Online vs offline divergence | Retrain the prompt on anecdotes |
If online pass or revision rate diverges from offline while the pin is unchanged, investigate distribution and tools first. Swapping the pin to “fix” a seasonal ticket mix is how you get two problems.
Worked failure: the quiet alias weekend
What broke: Prod used gpt-5.6 “so we always get the best.” A routing or default change shifted tool verbosity. Average tool calls per ticket left the band. Pass rate dipped. Cost per pass jumped. No deploy in git. The identity string moved; the repo did not.
Cost: Budget alerts. Weekend rollback to an explicit gpt-5.6-terra pin. Emergency golden-set triage while support ate the weird writes.
Instead: Prod pin gpt-5.6-sol or gpt-5.6-terra by role. Staging tracks the alias and logs the resolved ID. Upgrade only through the gate. CI denylists gpt-5.6.
The model did not “get dumber.” Your identity string did. I will not dress that weekend up as a benchmark. The receipt is the missing deploy and the missing gate.
What breaks if you never unpin?
Pinning without a migration calendar delays a worse incident.
| Failure | How it shows up | Prevention |
|---|---|---|
| Retirement hard-fail | Anthropic retired IDs return errors; OpenAI shutdown dates are published; Gemini endpoints turn off on the deprecations table | Calendar the date the day the notice lands |
| Prompt debt | Patches that only paper over old-model quirks | Candidate upgrade quarterly so patches stay portable |
| Forced big-bang | Competitors already scored the next Active ID; you migrate under a deadline | Keep a rollback pin and a scored candidate |
| Partner-cloud skew | Bedrock or Vertex retires the same Claude name on a different clock | Track the catalog you actually call |
Anthropic’s own history on the deprecations page is the warning label: dated 4.x IDs have already retired on the Claude API in 2026. If you still hold claude-haiku-4-5-20251001 as a volume pin, the tentative floor is 2026-10-15. That is a pin with an expiration date, not a personality.
Pair every prod pin with a written successor candidate and a last-eval date. If last_eval_passed is older than 90 days, the pin is a rumor.
What config shape survives a review?
# models.yml — reviewed in PRs
agents:
support_triage:
provider: anthropic
model_id: claude-sonnet-5 # pinned snapshot ID
reasoning_effort: null # unused on this pin; keep the key
prompt_version: support-triage@3.2.0
tools_version: crm-tools@1.7.0
eval_suite: golden@2026-03-12
last_eval_passed: 2026-03-10
rollback_model_id: claude-sonnet-4-6
successor_candidate: claude-opus-5
CI fails if model_id matches a denylist of aliases (latest, gpt-5.6, gemini-flash-latest, sonnet, opus). CI also fails if last_eval_passed is missing or older than your SLA.
Review questions a stranger should be able to answer from this file:
- What string does production send?
- What string do we roll back to?
- When did that pin last pass the suite?
- What candidate are we scoring next?
If the answers live in a Slack thread, you do not have a pin. You have folklore.
Who owns the pin, and what must every run log?
A pin without an owner expires in silence. Split the work so the person who likes new models is not the only person who can flip prod.
| Role | Owns | Does not own |
|---|---|---|
| Job owner | Pass criteria, deny cases, “ship / hold” on the score | The raw model string in a service |
| Eng | models.yml, CI denylist, rollback map | Quiet alias swaps “to see if it is better” |
| On-call | Rollback to rollback_model_id from logs | Inventing a new pin during an incident |
Every production run should emit a single identity line. If you cannot grep last Tuesday’s writes by model_id, you cannot prove what drifted.
| Field | Example | Why |
|---|---|---|
model_id | gpt-5.6-terra | Weights you think you bought |
resolved_id | gpt-5.6-terra | Same as model_id on a pin; different if staging floated |
prompt_version | support-triage@3.2.0 | Instructions |
tools_version | crm-tools@1.7.0 | Schema |
request_contract | reasoning.effort=medium | Knobs |
eval_suite | golden@2026-03-12 | What last scored this pin |
- Job owner named on the weekly score
- Eng named on the denylist and the rollback map
- On-call can flip to
rollback_model_idwithout a design review - Identity line is in the trace, not only in the repo
What is the pilot minimum?
A Spurlock $1,500 · 5-day agentic pilot ships:
- One pinned
model_idper agent role in config, with a rollback ID - A golden-set gate wired so a pin change is a scored PR
- An online sampling panel that would have caught the quiet alias weekend
You do not need multi-provider routing on day one. You need identity and a gate. The rest of the stack — evaluator, sandbox, kill switch — is in the operating manual. This week answers “can we name the weights and prove an upgrade?” If you cannot, do not scale the loop.
- Prod strings are snapshot or stable IDs
- Aliases are denylisted in CI
- Golden set runs on pin diffs
- Weekly online sample exists
- Deprecation URLs are in the runbook
FAQ
Which OpenAI / Anthropic / Gemini IDs should I pin today?
As of August 2026 docs: OpenAI gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna by role (avoid the gpt-5.6 alias in prod); Anthropic claude-fable-5, claude-opus-5, claude-sonnet-5, or dated claude-haiku-4-5-20251001; Gemini gemini-3.6-flash (and gemini-3.5-flash-lite when the job earned the cheaper tier). Re-verify on provider model pages before hardcoding months later. Dateless Claude 4.6+ IDs are snapshots; OpenAI and Gemini family aliases are not.
How long do providers keep pinned snapshots?
Until they deprecate and retire them — not forever. OpenAI documents at least six months’ notice for GA model retirement after announcement; Anthropic documents at least sixty days before retirement for public models. Gemini publishes shutdown dates per ID and gives about two weeks’ notice on preview and -latest breaking changes. Track the deprecation tables. Do not assume multi-year API access.
Staging on floating vs prod on pinned — good pattern?
Yes, if staging evals run on a schedule and someone owns failures. Prod stays pinned. Staging should log the resolved underlying ID when the alias moves. Staging float without alerts is how you learn about provider changes from a social feed instead of CI.
What belongs in SemVer for prompts vs model pins?
Version prompts and rubrics (prompt_version), tool contracts (tools_version), and the eval suite separately from model_id. A model pin change is a config change that must pass the golden-set gate even when the prompt SemVer does not bump. Incident logs should print all four, plus any reasoning or thinking knob you pinned.
How does online sampling catch input-distribution drift?
Offline goldens freeze yesterday’s tickets. Online samples score today’s mix with the same evaluator. When online pass or revision rates diverge from offline, you are seeing distribution drift — not necessarily a bad pin. Investigate mix, tools, and indexes before you “fix” the model.
What breaks if I never unpin?
Forced retirement outages, prompt debt that only works on the old pin, and a painful big-bang migration. Dated Claude pins already have 2026 retirement floors measured in weeks. Pin for stability. Schedule candidate upgrades so unpinning is a controlled PR, not an incident.
CTA
Want pins, gates, and drift panels on a real agent job in five days? Start at /agentic or /contact?intent=agentic-pilot.
What questions does this article answer?
- Which OpenAI / Anthropic / Gemini IDs should I pin today?
- As of August 2026 docs: OpenAI `gpt-5.6-sol` / `gpt-5.6-terra` / `gpt-5.6-luna` by role (avoid the `gpt-5.6` alias in prod); Anthropic `claude-fable-5`, `claude-opus-5`, `claude-sonnet-5`, or dated `claude-haiku-4-5-20251001`; Gemini `gemini-3.6-flash` (and `gemini-3.5-flash-lite` when the job earned the cheaper tier). Re-verify on provider model pages before hardcoding months later. Dateless Claude 4.6+ IDs are snapshots; OpenAI and Gemini family aliases are not.
- How long do providers keep pinned snapshots?
- Until they deprecate and retire them — not forever. OpenAI documents at least six months’ notice for GA model retirement after announcement; Anthropic documents at least sixty days before retirement for public models. Gemini publishes shutdown dates per ID and gives about two weeks’ notice on preview and `-latest` breaking changes. Track the deprecation tables. Do not assume multi-year API access.
- Staging on floating vs prod on pinned — good pattern?
- Yes, if staging evals run on a schedule and someone owns failures. Prod stays pinned. Staging should log the resolved underlying ID when the alias moves. Staging float without alerts is how you learn about provider changes from a social feed instead of CI.
- What belongs in SemVer for prompts vs model pins?
- Version prompts and rubrics (`prompt_version`), tool contracts (`tools_version`), and the eval suite separately from `model_id`. A model pin change is a config change that must pass the golden-set gate even when the prompt SemVer does not bump. Incident logs should print all four, plus any reasoning or thinking knob you pinned.
- How does online sampling catch input-distribution drift?
- Offline goldens freeze yesterday’s tickets. Online samples score today’s mix with the same evaluator. When online pass or revision rates diverge from offline, you are seeing distribution drift — not necessarily a bad pin. Investigate mix, tools, and indexes before you “fix” the model.
- What breaks if I never unpin?
- Forced retirement outages, prompt debt that only works on the old pin, and a painful big-bang migration. Dated Claude pins already have 2026 retirement floors measured in weeks. Pin for stability. Schedule candidate upgrades so unpinning is a controlled PR, not an incident.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.