September 28, 2026

OpenAI legacy completion models shut down today: what breaks

OpenAI shut down gpt-3.5-turbo-instruct, babbage-002, davinci-002, and gpt-3.5-turbo-1106 today. What breaks, the gpt-5.6-terra migration, and cost.

News

Four OpenAI models stopped taking API calls today. gpt-3.5-turbo-instruct, babbage-002, davinci-002 and gpt-3.5-turbo-1106 all hit their shutdown date on September 28, 2026, exactly as OpenAI told developers they would a year ago. OpenAI's deprecations page announced the retirement on September 26, 2025, and put it plainly: "we are deprecating a set of older OpenAI models with declining usage over the next six to twelve months. Access to these models will be shut down on the dates below." If a production app is still calling one of these four models, the notice window closed a long time ago. The work left is fixing it.

Which models are gone, and what replaces them

Three of the four ran on OpenAI's legacy Completions endpoint, the plain-text-in, plain-text-out API that predates Chat Completions. gpt-3.5-turbo-instruct billed at $1.5 per million input tokens and $2 per million output tokens, and its own documentation described it as "an older model only compatible with the legacy Completions endpoint." davinci-002 charged $2 per million tokens for input and output alike, and OpenAI's page for that model states plainly: "[m]ost customers should use GPT-3.5 or GPT-4." babbage-002, the cheapest of the three at $0.4 per million tokens either direction, supported only that same legacy Completions endpoint and nothing else.

gpt-3.5-turbo-1106 is the fourth name on the same shutdown table. OpenAI's GPT-3.5 Turbo model page lists it as a snapshot of a model that supports Chat Completions and Responses, so it never shared the other three's endpoint problem. Its OpenAI model page no longer resolves, so no price appears here.

OpenAI's stated replacement for all four is gpt-5.6-terra. Its model page lists $2 per million input tokens, $0.2 per million cached input tokens, and $12 per million output tokens, a 1,050,000-token context window, a 128,000-token output cap, and a knowledge cutoff of February 16, 2026. It supports Chat Completions, Responses, and Batch, but it does not support the legacy Completions API, and that's the detail that turns three of these four migrations into an actual rewrite rather than a config change. This kind of forced swap, an old model's fixed shutdown date paired with a pricier replacement, is exactly what AI model releases and pricing: what changes for engineering teams tracks across every vendor.

How to migrate off gpt-3.5-turbo-instruct, babbage-002, davinci-002 and gpt-3.5-turbo-1106

1. Find every remaining call to the four retiring models

Search every service, script, notebook, and test suite for the four model strings directly. Check beyond the obvious inference path, too: legacy Completions calls tend to hide in older internal tools, a classification script, a batch-labeling job, a fine-tuning pipeline, that nobody's touched since it was written, which is exactly why they're still pointed at models built for that endpoint in the first place.

2. Rewrite legacy Completions-format prompts for the API gpt-5.6-terra actually supports

Because gpt-5.6-terra doesn't run on legacy Completions, a straight model-ID swap fails immediately for gpt-3.5-turbo-instruct, davinci-002, and babbage-002. Those three calls have to move to Chat Completions or Responses. OpenAI's own migration guide doesn't start from legacy Completions, though. It walks teams from Chat Completions to Responses, stating "While Chat Completions remains supported, Responses is recommended for all new projects," and describes the format shift by saying Chat Completions "is an array of Messages" while "Responses API uses Items." None of that covers a team starting from a plain-text Completions prompt with no message structure at all. Budget time to write that first conversion by hand, since OpenAI hasn't published a guide for it. gpt-3.5-turbo-1106 skips this problem: it already ran on Chat Completions, so pointing its calls at gpt-5.6-terra is a model-ID change rather than a format rewrite, though it still needs the re-testing and re-pricing covered in the next two steps.

3. Rebudget: gpt-5.6-terra is not priced like the models it replaces

gpt-5.6-terra's output price of $12 per million tokens runs 6 times what gpt-3.5-turbo-instruct or davinci-002 charged for output, and 30 times what babbage-002 charged. Input pricing moves less sharply, about a third higher than gpt-3.5-turbo-instruct's $1.5 rate, but a workload that leaned on babbage-002 for high-volume, low-cost completions is the one that feels this cutover hardest. gpt-3.5-turbo-1106's own pricing isn't available since its model page no longer resolves, so that call needs its own check against gpt-5.6-terra's published rates rather than an assumption the swap comes out even. Run the numbers against last month's actual token volume for every one of the four before trusting any of these migrations is cost-neutral. It probably isn't.

4. Re-test the calls that are failing right now

If step 1 turned up any live traffic, it's already erroring today. Confirm the error paths your app takes when a model call fails, since a silent retry loop against a dead model ID burns latency and budget without ever succeeding. Once the rewritten prompt is live on gpt-5.6-terra, re-run whatever evaluation set covers that feature. A Completions-to-Chat-format rewrite changes how the model reads the prompt, and that alone can shift output quality enough to notice.

OpenAI forced a similar migration this same stretch with GPT-6 Astra release: pricing, context, API changes: tool calls there now require the Responses API, the same endpoint the three legacy Completions models above are moving to. Claude Opus 5.5: what to re-price and re-test now walks through the same re-pricing and re-testing checklist for a different vendor, worth checking if the same pipeline calls both providers. DeepSeek V4.1 Flash: pricing, the v4-pro reroute and EU data covers another case of calls getting silently rerouted to a different model, exactly what step 1's search is meant to catch before it happens unnoticed.

FAQ

What happens if my app still calls gpt-3.5-turbo-instruct, babbage-002, davinci-002, or gpt-3.5-turbo-1106 after September 28, 2026?

The calls fail. OpenAI's deprecations page states access to these models "will be shut down" on their listed date, and September 28 is that date for all four. There's no grace window described anywhere in the source, so migrating the call is the only way to restore it.

Is gpt-5.6-terra more expensive to run than gpt-3.5-turbo-instruct?

Yes, on output tokens specifically. gpt-5.6-terra charges $12 per million output tokens against gpt-3.5-turbo-instruct's $2, a 6x jump. Input pricing is closer: $2 per million against gpt-3.5-turbo-instruct's $1.5, roughly a third more. Any workload with a heavy output-to-input ratio should expect its per-request cost to move more than the headline numbers suggest.

Does gpt-5.6-terra support the legacy Completions API?

No. Its model page lists Chat Completions, Responses, and Batch as supported, and doesn't include legacy Completions among them. That's why swapping the model ID alone doesn't work for gpt-3.5-turbo-instruct, davinci-002, or babbage-002: the endpoint those three ran on isn't one gpt-5.6-terra accepts calls through. gpt-3.5-turbo-1106 already sat on Chat Completions, so its swap is simpler, though it still needs the re-test and re-pricing check the other three get.

Does OpenAI provide a migration guide for moving off the legacy Completions API?

Not for this specific jump. OpenAI's migration guide covers moving from Chat Completions to Responses; it doesn't start from legacy Completions at all. A team rewriting gpt-3.5-turbo-instruct, babbage-002, or davinci-002 prompts into message or item format is working from the general shape of that guide rather than a document written for its exact starting point.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now