Your Model Has a Retirement Date. Your Team Doesn’t Have a Switch Rule.

On 30 September, Anthropic told developers that Claude Sonnet 4.5 retires from the Claude API on 30 November. On 23 October, OpenAI switches off GPT-4, GPT-4 Turbo, o1 and o3-mini. If your agent runs on one of them, your next model migration already has a deadline, and you didn’t choose it.

The last time a model deprecation notice landed on my team, we put one ticket on the planning board: “Update model endpoint string.” Two story points. We marked it done in two days and rolled it to staging, then spent the next three weeks finding out why three separate code-refactoring workflows were failing without an error. The model had stopped returning the structured schema they expected. Changing the string took ten minutes. Understanding what it broke took weeks.

Most migrations happen after the model is gone

Hyungjin Lukas Kim, a researcher at Myongji University in Seoul, went through 22,555 commits in 17,703 GitHub repositories to see what open-source applications do when a provider retires their model. He estimates that 82% of migrations off a retired model were committed after the shutdown date. The median fix landed 39 days after it. About 4% of the migration commits mention an evaluation suite, and 93% swap one model ID for another where it’s written in the code.

Some of those applications didn’t crash. In 8% of the migrations that the study’s two reviewers checked by hand, an error handler had caught the retired model’s error and returned an apologetic chat message, a null result, a fallback to demo mode or, in one case, a constant answer.

The constant answer came from an app that identifies dog breeds. After its model was retired, it showed no error and told every user the same thing. The commit that fixed it says breed identification “always returned ‘Golden Retriever’”.

Kim studied open-source projects through their commit messages, so a company with a release process might do better. Mine had a release process, and it produced a two-point ticket.

The switch rule in four parts

A switch rule is four decisions you write down before you need them.

Put the retirement date on the roadmap

The dates are public before the deprecation email arrives. Anthropic promises “at least 60 days’ notice” and lists a tentative retirement date against each active model. As I write this, Claude Haiku 4.5’s reads “not sooner than October 15, 2026”. OpenAI promises at least six months’ notice for “generally available models”. Preview models get far less.

Projects met a long deadline and missed a short one. In Kim’s data, 89% of migrations came after the shutdown under Anthropic’s 60 to 114 day notices, against 13% under the one-year notice OpenAI gave before shutting down its Assistants API. An abstraction layer such as LiteLLM or OpenRouter didn’t help: applications that already routed calls through one were as likely to migrate late as the rest. Experience didn’t help either. Among repositories hit by two or more retirements, 73% of migrations came after the shutdown at the first retirement, and 73% at the later ones.

A 60-day notice lands inside a quarter you’ve already committed, so someone has to hold back capacity for it, and deciding that is the PM’s job.

Write the rule before the release

In the first week of September, Anthropic, Meta, Google and OpenAI each shipped a model within three days of each other. Zhen Lu, CEO of Runpod, told CNBC: “I feel like model fatigue is a real thing.” Suresh Vasudevan, who runs Clockwork Systems, said that when his startup wants to evaluate ten models for a task, it may test only five. A team without a written rule turns each of those releases into a debate, or skips it.

Write the rule per task type. The one I’d start from switches to the candidate model when:

  • all critical test cases pass, because no cost saving buys back a failure there
  • quality on each task type stays within N points of today
  • cost per successful task falls by at least X%, or p95 latency by Y%
  • the migration work fits inside Z days

Cost per successful task is what a run costs divided by how often runs succeed. Arize and LangWatch both wrote it up this summer, and LangWatch’s worked example has a $0.80 run that succeeds 78% of the time costing $1.03 per successful task. If you price per outcome, it sets your margin.

The PM sets those thresholds and makes the call; engineering owns the eval harness that produces the numbers.

On an enterprise SQL-generation agent, we made one gate non-negotiable: zero regressions on schema-mapping safety cases. A candidate model came in x% cheaper and y% faster, and it failed n of the m complex multi-join edge cases the old model passed. Our rule already said cost savings couldn’t buy back a critical failure, so it was an immediate no-go until we re-tuned the prompt. Without a rule like that you get what Steve James calls “eval theatre”: results that arrive after the ship decision, on a dashboard no one owns.

Test model, prompt and settings together

Changing the model ID is the smallest part of a swap. Your prompt was tuned to the old model’s habits, and some API settings may not carry over. Run the candidate on your test cases twice, once with today’s prompt and once with a re-tuned one. The gap between those two runs is the migration work you’re signing up for.

Run each case more than once, too. The same agent can pass a task on Monday and fail it on Tuesday, and Anthropic’s guide to agent evals shows how fast that adds up: at a 75% success rate per attempt, the chance of three passes in a row is about 42%.

The test set can be small. The same guide reports that teams “delay building evals because they think they need hundreds of tasks”, when “20-50 simple tasks drawn from real failures is a great start.” Report the results by task type, because one average can hide a drop on the task your biggest customer runs.

Intercom works this way on Fin, its customer-support agent. Pedro Tabacof, a principal ML scientist on the Fin team, described the process in one sentence: “Every change, from a one-line prompt tweak to swapping the underlying model, has to earn its way to production through backtests (replaying past conversations offline), then an A/B test with real customers, then a monitored rollout.” They score latency and costs next to resolution rate, the share of conversations Fin resolves.

Keep the test cases, and the checks that score them, in files you own. OpenAI’s own Evals platform goes read-only on 31 October and shuts down on 30 November, so a provider can retire your eval tool as well as your model.

Time the verdict, then keep watching

Time to verdict is the number of days between a new model’s release, or a deprecation notice, and your team’s decision. With the first three parts in place, you count it in days. Without them you count it in weeks. Anthropic’s guide to agent evals describes the gap: “When more powerful models come out, teams without evals face weeks of testing while competitors with evals can quickly determine the model’s strengths, tune their prompts, and upgrade in days.”

After you switch, keep a small fixed set of test cases running on a schedule, because a model’s behaviour can change while its name stays the same, and the provider may not be the first to notice. In Anthropic’s September 2025 postmortem, a routing bug sent Sonnet 4 requests to the wrong servers, and the answers got worse. At its worst hour it hit 16% of requests, and the company wrote that “the evaluations we ran simply didn’t capture the degradation users were reporting.”

A worked example: the Claude Sonnet 4.5 deprecation

Say your agent runs on Claude Sonnet 4.5 today. Your date is 30 November, 61 days after the notice, with Q4 already planned. Sonnet 5.5 is the replacement Anthropic recommends, a newer model in the same family. Write the rule now, for each task type that runs on Sonnet 4.5, before anyone on the team has tried the new model. The thresholds are yours to set; mine wouldn’t fit your product.

The test is where the 61 days go. Anthropic’s migration guide lists “five settings that return a 400 error” on Sonnet 5.5: thinking budgets, non-default sampling parameters such as temperature, assistant prefill, forced tool choice, and turning thinking off. Sonnet 4.5 accepted all five. Thinking also runs by default on the new model, and it’s billed in full as output tokens even though, by default, the response shows none of it.

Sonnet 4.5 costs $3 per million input tokens and $15 per million output tokens; Sonnet 5.5 costs $2 and $10. That reads as a 33% cut. The migration guide also says the same text counts as about 30% more tokens on Sonnet 5.5, so by my arithmetic the same prompt costs about 13% less (two-thirds of the price, on 1.3 times the tokens), before thinking adds anything. The amount thinking adds depends on the effort level, a setting Sonnet 4.5 doesn’t have. The guide’s checklist covers both in one line: “Re-run your effort sweep, and re-baseline cost.”

Give the verdict a target, two weeks say, and keep the rest of the 61 days for shadow traffic and a small canary before 30 November.

Where the rule doesn’t fit

A prototype or a low-volume agent may not need all of this, because a full harness can cost more than the switch saves. I don’t know where that line sits, and I haven’t found anyone who’s measured it. Even for a prototype, keep 20 to 50 real test cases in a spreadsheet.

Preview models don’t give you time to run the rule. OpenAI’s deprecation page says they “may be retired with much shorter notice, such as 2 weeks”, and adds: “We don’t recommend using preview models for business-critical production workloads unless you can migrate on short notice.”

Fine-tuned models need more lead time. In Kim’s data, an application that only sends prompts changed a median of 6 lines of code to migrate, and one built on a fine-tuned model nearly 700, because the tuning doesn’t carry over to the replacement model.

And if no one will act on the result, settle who owns the decision before you build the dashboard.

Your first step

Open a spreadsheet with one row for each model ID your product calls. In each row, note where the ID is set (Kim found it hard-coded in 94% of the applications that had to migrate) and the model’s retirement date from the provider’s deprecation page, or its “not sooner than” date if that’s all the page gives. Then add one more column: the result that would make you switch. If your team can’t fill that column, book the meeting where you fill it.

A new model will ship before you finish that spreadsheet. Will your team reach a verdict on it in days, or in weeks?

Share via
Copy link