In this article 5 sections

OpenAI’s deprecation table schedules four model IDs to shut down today: gpt-3.5-turbo-instruct, babbage-002, davinci-002, and gpt-3.5-turbo-1106. Every row recommends gpt-5.6-terra. That makes the change look like one model-name replacement. It is at least three different migrations. (OpenAI deprecations)

gpt-3.5-turbo-instruct is an instruction-following model that only uses the legacy Completions endpoint. babbage-002 and davinci-002 are base models, not instruction-tuned assistants, and also use legacy Completions. gpt-3.5-turbo-1106 is a dated chat snapshot. Their prompts, response parsers, quality expectations, and cost baselines are not interchangeable. (OpenAI model pages) (Babbage 002) (Davinci 002) (GPT-3.5 Turbo)

The replacement is much broader. GPT-5.6 Terra supports legacy Completions, Chat Completions, and Responses. It also defaults to medium reasoning effort and publishes a 1.05 million-token context window with a 128,000-token maximum output. Compatibility gives teams a way to stage the move, but it does not prove behavioral equivalence. (GPT-5.6 Terra)

One target model, three contracts

The safest plan begins by classifying each production caller by its old contract.

  1. Instruction completion: preserve the existing prompt and legacy endpoint first. Change only the model, cap the output, and measure whether the new model follows the same delimiters and stopping rules.
  2. Base-model continuation: decide whether the workload actually needs raw continuation semantics. A base model may have been used for classification, ranking, or synthetic text through hand-built prefixes. Moving that workload to an instruction model changes the abstraction, not just the weights.
  3. Chat snapshot: keep the Chat Completions message structure for the first comparison, or migrate deliberately to Responses as a separate change. Do not attribute parser or state-management failures to the model if the endpoint changed at the same time.
Conceptual migration map. Three legacy caller contracts remain separate while each establishes a GPT-5.6 Terra baseline behind its current interface. Only after that checkpoint does an optional second release move the interface to Responses.
Conceptual migration map synthesized from OpenAI’s model pages, GPT-5.6 Terra documentation, and the Responses migration guide. It is an editorial rollout pattern, not an official OpenAI migration sequence.

This separation matters because OpenAI’s Responses guide names several API-level changes. Responses returns typed output items instead of a choices array, moves Structured Outputs definitions from response_format to text.format, uses a different function-calling shape, and offers different state handling. Those are useful capabilities, but they create a second migration axis. (OpenAI Responses migration guide)

If a team changes the model and API together, a failed result has too many plausible causes. The model may have produced different text. The parser may have read the wrong object. A stream consumer may have ignored a typed event. State may have been stored or replayed differently. The clean cutover is therefore sequential: stabilize Terra behind the old interface where supported, then move the interface with a new baseline.

The replacement can invert the cost profile

Published token prices make the same point from another direction. Terra is listed at $2 per million input tokens and $12 per million output tokens. The retiring davinci-002 is listed at $2 and $2. babbage-002 is listed at $0.40 and $0.40. gpt-3.5-turbo-instruct is listed at $1.50 and $2. (GPT-5.6 Terra) (Babbage 002) (Davinci 002) (GPT-3.5 Turbo Instruct)

For the same token counts, Terra’s output rate is six times the published rate of Davinci 002 and GPT-3.5 Turbo Instruct, and 30 times Babbage 002’s. Its input rate is unchanged from Davinci 002, about one-third higher than GPT-3.5 Turbo Instruct, and five times Babbage 002’s.

Those ratios are not a bill forecast. A stronger model may finish with fewer retries, produce more useful work, or generate a different number of tokens. Terra’s default reasoning effort also means a latency-sensitive migration should set and test effort intentionally rather than inherit the default accidentally. The operational metric is cost per accepted task, with output-token rate, output length, retry count, latency, and acceptance measured separately.

A migration test should therefore keep four distributions, not one average:

  • input and output tokens per attempt;
  • wall-clock latency by task class;
  • accepted, rejected, retried, and manually corrected outcomes;
  • cost per accepted result.

An output cap belongs in the first release even if the new model permits much longer responses. A larger ceiling is capacity, not a safe default.

Treat the deadline as a routing event

The deprecation notice identifies a replacement, not a rollout design. A thin routing adapter can make the remaining decisions explicit:

workload class -> endpoint -> model -> reasoning effort -> output cap -> parser version

Version that tuple. Log it with each attempt. That record lets an operator separate a model regression from an API migration, configuration change, or parser defect.

Run representative fixtures through the old and new routes while the old route still responds. For deterministic classifiers or extractors, compare task-specific invariants rather than prose similarity. For free-form generation, sample failures and manual corrections by workload. For chat, include long conversations and boundary cases around message roles. A global pass rate can hide a broken low-volume route, so preserve results by caller and task class.

If the old model is already unavailable, the same discipline still applies. Keep the last known outputs and production distributions as the control. Route Terra through the old endpoint first where possible, then change the endpoint only after the model baseline is stable. The absence of live dual-running makes stored fixtures and parser-level tests more important, not less.

OpenAI’s table also lists October 23 shutdowns for additional snapshots and fine-tuned models. That is a separate migration wave. Fine-tuned babbage-002 and davinci-002 workloads should not inherit today’s acceptance thresholds merely because their base names look familiar. (OpenAI deprecations)

Stay Sharp: canary a model as a release

Google’s SRE Workbook defines a canary as a partial, time-limited deployment evaluated against a control. It requires a way to send only part of the population to the candidate, an evaluation process, and integration of the result into the release decision. The technique applies to configuration and data changes as well as binaries. (Google SRE Workbook)

For a model migration, select the canary with a stable rule such as a tenant or request hash. Random per-attempt routing can send retries to different models and contaminate both user experience and measurements. Keep the control and canary active during the same period so traffic mix and external conditions are comparable.

Define the release gate before traffic moves. A useful gate combines hard failures, task acceptance, p95 latency, output tokens, and cost per accepted result. Give low-volume but consequential workloads their own guardrails. Then make rollback independent of the application deployment, so changing the route does not require shipping new code.

The lesson is narrower than “test your prompts.” A model ID, endpoint, reasoning setting, output cap, and parser form one release unit. Canary that unit, and change one axis at a time when attribution matters.

What to watch

  • Whether OpenAI documents the exact retirement behavior and error responses for the four September 28 model IDs.
  • Whether workloads moved from base models preserve their original ranking or continuation semantics after instruction tuning enters the path.
  • Whether output length, reasoning effort, or retries cause cost per accepted task to diverge from simple token-price estimates.
  • Whether teams facing the October 23 wave separate fine-tuned-model validation from today’s base and chat migrations.