October 5, 2026

OpenAI deprecations: gpt-5.1 and TTS models removal dates

OpenAI removes tts-1 on 6 Jan 2027 and gpt-5.1, gpt-5.3-codex and gpt-5.4-nano on 1 Apr. Dates, replacements, price changes and what to test.

News

OpenAI's latest model deprecations put three removal dates on the calendar, two announced on 1 October 2026 and one on 26 August: 6 January 2027 for the text-to-speech models, 26 February 2027 for the transcription models, and 1 April 2027 for gpt-5.3-codex, gpt-5.1 and gpt-5.4-nano. Text models have a drop-in-looking replacement. The TTS models don't, and that's the part to plan around first. Every date and replacement below comes from OpenAI's deprecations page.

Which OpenAI models go away, and when

The page announces the text models this way: "The following models are deprecated and will be removed from the API on April 1, 2027, with six months' notice." For text-to-speech, the page says: "The following text-to-speech models are deprecated and will be removed from the API on January 6, 2027, with at least three months' notice." Its table lists gpt-realtime-2.1-mini as the replacement for all four.

ModelRemoval dateOpenAI's replacement
gpt-5.3-codex1 April 2027gpt-6-sol
gpt-5.11 April 2027gpt-6-sol
gpt-5.4-nano1 April 2027gpt-6-luna
tts-16 January 2027gpt-realtime-2.1-mini
tts-1-hd6 January 2027gpt-realtime-2.1-mini
gpt-4o-mini-tts-2025-03-206 January 2027gpt-realtime-2.1-mini
gpt-4o-mini-tts-2025-12-156 January 2027gpt-realtime-2.1-mini

An older entry on the same page, dated 26 August 2026, covers transcription models. It says OpenAI notified developers using "whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize of their deprecation and removal from the API on February 26, 2027." All four list gpt-live-transcribe or gpt-transcribe as the replacement.

The notice periods differ. The text models get six months. The TTS models get "at least three months," which makes January the nearest deadline and the one with the least slack.

If you've been through this before, the pattern is familiar. OpenAI shut down four legacy models on 28 September, covered in OpenAI's September shutdown of four legacy models, and Anthropic ran its own version in the Sonnet 4.5 retirement walkthrough. The difference this time is that the notice hasn't expired yet.

What a gpt-5.1 or gpt-5.3-codex to gpt-6-sol move changes

Prices below are per 1M tokens on the Standard tier, checked 2026-10-05 against OpenAI's pricing page.

ModelInputCached inputOutput
gpt-5.1$1.25$0.125$10.00
gpt-5.3-codex$1.75$0.175$14.00
gpt-6-sol$2.00$0.20$10.00
gpt-5.4-nano$0.20$0.02$1.25
gpt-6-luna$0.10$0.01$0.50

The comparisons that follow are our arithmetic on those list prices, not OpenAI's claims. gpt-6-sol's input price is 1.6 times gpt-5.1's ($2.00 against $1.25), while output is the same $10.00. Against gpt-5.3-codex, Sol's input costs about 14% more ($2.00 against $1.75) and its output about 29% less ($10.00 against $14.00). On the small-model side, gpt-6-luna's input is half of gpt-5.4-nano's ($0.10 against $0.20) and its output is 0.4 times as much ($0.50 against $1.25).

Those ratios don't add up to a saving or a surcharge. Your bill depends on your input-to-output mix, how much of the prompt gets cached, and how many reasoning tokens a given model spends, none of which a price table shows. Replay a week of real traffic against the new model before you forecast anything.

On specs, OpenAI's gpt-6-sol model page lists a 1,050,000-token context window, with up to 922,000 tokens of input and 128,000 of output, and a knowledge cutoff of April 20, 2026. It describes Sol as "Built for complex coding and agentic workflows." gpt-5.1 has a 400,000-token context and 128,000 max output, so the window grows by more than 2.5 times while the output cap stays put. gpt-6-luna matches Sol's context figures, with a cutoff of May 18, 2026, and OpenAI calls it "Our most efficient model for focused, high-volume tasks."

Sol also has a newer sibling. If you're choosing a target rather than following OpenAI's suggestion, GPT-6.1 Sol pricing is worth a look before you commit.

TTS: the replacement is not a drop-in

Judgment first: swapping tts-1 for gpt-realtime-2.1-mini is not a like-for-like change. The facts behind that come from OpenAI's gpt-realtime-2.1-mini model page. It lists the Realtime endpoint (v1/realtime) as its only supported endpoint, so Chat Completions, Responses and /audio/speech are not listed. You connect over WebRTC, WebSocket or SIP. The model has a 128,000-token context and 32,000 max output, supports function calling and prompt caching, and the page lists text, audio and image as inputs and text and audio as outputs.

If your server POSTs text to /audio/speech and gets audio back, that code path doesn't carry over. Moving to gpt-realtime-2.1-mini means building around a Realtime session. For many teams that's a feature rebuild, not a config change. That's the real decision behind the January date: rebuild on Realtime, or move speech to a different vendor.

There's also a documentation inconsistency. OpenAI's text-to-speech guide still calls gpt-4o-mini-tts "our newest and most reliable text-to-speech model" and shows no deprecation notice in its page text as of 5 October 2026, even though the deprecations page lists both gpt-4o-mini-tts snapshots for removal. Confirm with OpenAI or your account contact before you commit engineering time either way. If you do rebuild, the Responses API build tutorial is the place to start for the text side of the work.

Pricing changes shape, not just size

The units differ, which makes a clean comparison impossible. Checked 2026-10-05:

ModelPricing
tts-1$15.00 per 1M characters
tts-1-hd$30.00 per 1M characters
gpt-4o-mini-tts$12.00 per 1M audio output tokens, plus $0.60 per 1M text input tokens
gpt-realtime-2.1-miniAudio $10.00 in and $20.00 out per 1M tokens; text $0.60 in and $2.40 out per 1M tokens

Characters and tokens don't convert without knowing your text, your voice and how long each reply runs. Don't derive a per-minute figure from this table. Re-estimate from real traffic instead.

What this means for your AI roadmap and budget

Three dates, three owners. Put 6 January, 26 February and 1 April into Q4 and Q1 planning now, and name one person per date who is responsible for finding every call to the affected models. The text migration is mostly an eval pass: run your prompts against gpt-6-sol and gpt-6-luna, compare quality, then re-forecast token spend on your own traffic mix. The voice migration is a decision, and as a judgment call, it's the one that needs a senior engineer's time this quarter: rebuild on a Realtime session, or switch vendors. Make that call before the holidays, because January leaves little room to change course.

For the budget side, AI model releases and pricing: what changes for engineering teams tracks the other forced moves and price changes as they land, which is where to look before you lock a migration target.

FAQ

What happens to calls to gpt-5.1 after 1 April 2027?

OpenAI says gpt-5.1 will be removed from the API on that date. A removed model isn't available to call, so any service still pointing at it needs to be on gpt-6-sol, OpenAI's listed replacement, before then. gpt-5.3-codex goes the same day, also to gpt-6-sol.

When does tts-1 stop working?

On 6 January 2027, with tts-1-hd and both gpt-4o-mini-tts snapshots (2025-03-20 and 2025-12-15). OpenAI gave "at least three months' notice."

Is gpt-realtime-2.1-mini a text-to-speech replacement?

OpenAI names it as the migration target, but its model page lists the Realtime endpoint as the only supported endpoint, so Chat Completions, Responses and /audio/speech are not listed. Existing text-in, audio-out calls to the speech endpoint need to be rebuilt around a Realtime session.

Is gpt-6-sol cheaper than gpt-5.1?

Not on input: $2.00 against $1.25 per 1M tokens. Output is $10.00 for both. Whether your total bill goes up depends on your token mix and reasoning tokens, so test on real traffic.

Which model replaces gpt-5.4-nano?

gpt-6-luna. Per the pricing page, its input is $0.10 against nano's $0.20, and its output is $0.50 against $1.25 per 1M tokens (checked 2026-10-05).

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now