September 25, 2026

DeepSeek V4.1 Flash: pricing, the v4-pro reroute and EU data

DeepSeek retired V4-Flash and rerouted v4-pro to V4.1-Flash on Sept 14. What changed in pricing, what to re-check, and the EU data question.

News

Insight

DeepSeek V4.1 Flash launched on September 10, 2026, and if your team has deepseek-v4-flash or deepseek-v4-pro hardcoded anywhere in a config file, that name doesn't point where it used to. DeepSeek retired the old V4-Flash and V4-Flash-Vision-Exp models outright, and four days later it started routing every deepseek-v4-pro request to the new flash model too, billed at flash prices. Nothing about your integration changed. The model answering it did.

What DeepSeek actually shipped on September 10

DeepSeek's launch notes describe the new model as "smarter, faster, more efficient" than what it replaces, and the architecture backs that up on paper. V4.1-Flash is a mixture-of-experts model with a 552-billion-parameter backbone that activates 8 billion parameters for input processing and 16 billion for output, according to the same notes. Its Hugging Face card lists the total model size at 763 billion parameters. The same source puts the memory savings in concrete terms too: a quarter of the HBM and an eighth of the SSD storage that V4-Flash needed for its KV cache. V4.1-Flash is live on the API now, with native multimodal support built in from launch.

The part that matters more than the model: your pinned name now points somewhere else

DeepSeek retired V4-Flash and V4-Flash-Vision-Exp outright. For compatibility, any code still calling deepseek-v4-flash or deepseek-v4-flash-vision-exp now routes to V4.1-Flash, per the same launch notes. That's the smaller change.

The bigger one lands four days later. Starting at 04:00 UTC on September 14, 2026, every request to deepseek-v4-pro also routes to V4.1-Flash, billed at V4.1-Flash rates, and the notes say that routing continues until V4.1-Pro launches. If your production code has "v4-pro" pinned somewhere, in an environment variable, a config file, a hardcoded string in a service nobody's touched since it shipped, that string now resolves to a different, cheaper, architecturally distinct model.

That's the same kind of silent default that changes which file quietly starts governing an agent: a vendor treating a name as an interface rather than a fixed artifact. The name stays stable. What's behind it doesn't have to.

What V4.1-Flash actually costs, off-peak vs. peak

New pricing took effect at 04:00 UTC on September 10, 2026, the same day the model launched, per DeepSeek's pricing page. The page lists rates for two model IDs, deepseek-flash and deepseek-v4-pro, each split into off-peak and peak windows:

ModelMetricOff-peakPeak
deepseek-flashInput, cache hit$0.003 / 1M tokens$0.006 / 1M tokens
deepseek-flashInput, cache miss$0.15 / 1M tokens$0.3 / 1M tokens
deepseek-flashOutput$0.6 / 1M tokens$1.2 / 1M tokens
deepseek-v4-proInput, cache hit$0.022 / 1M tokens$0.044 / 1M tokens
deepseek-v4-proInput, cache miss$0.66 / 1M tokens$1.32 / 1M tokens
deepseek-v4-proOutput$1.98 / 1M tokens$3.96 / 1M tokens

Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays, per the same page. That's seven hours a weekday where every one of those rates exactly doubles. A team that models its budget off the off-peak column and never checks what fraction of its traffic actually lands in those windows will miss the real number by a wide margin.

Here's the part worth flagging directly to whoever owns the DeepSeek integration: the announcement says deepseek-v4-pro requests now bill at V4.1-Flash rates, but the pricing page still lists a separate, higher rate card under the deepseek-v4-pro model ID itself. Which number actually lands on your invoice is a question only your own billing dashboard can answer. Pull an actual invoice from after September 14 before trusting either table as the final word.

Context length is listed at 1M tokens for both model IDs, with a maximum output of 384K tokens, on the same pricing page.

Self-hosting instead of calling the API

V4.1-Flash's weights are released under the MIT license, according to DeepSeek's Hugging Face model card, which also confirms the 552-billion-parameter backbone, the 8-billion/16-billion active-parameter split described in the launch notes, and context windows of up to one million tokens. MIT is about as permissive as a license gets: it allows commercial use and doesn't require open-sourcing anything built on top of it.

That's a real alternative to calling a China-hosted endpoint, if a team has the hardware and the ops capacity to run a mixture-of-experts model with 763 billion total parameters. That's not a small if. Serving that model in production takes GPU infrastructure most teams don't have sitting around, and self-hosting still carries real operational cost of its own. For a team already uneasy about a vendor silently changing which model answers a pinned name, owning the weights outright removes that specific risk.

What to re-check before you keep routing production traffic through a pinned name

A team running production traffic against a pinned DeepSeek model name has three things worth checking before the next invoice arrives.

1. Re-run your evals

Whatever eval suite qualified deepseek-v4-pro for production, run it again against whatever's actually answering that name today. No published benchmark score for V4.1-Flash against your specific workload exists anywhere. The only score that tells you anything is the one you generate yourself, against the model actually serving your traffic now.

2. Re-price for the peak window

Pull the last week of request logs and check what fraction landed inside the Monday-to-Friday peak windows, 01:00-04:00 and 06:00-10:00 UTC. Every rate in the pricing table above exactly doubles during those hours. A monthly estimate built entirely on the off-peak column is a best case rather than a forecast.

3. Confirm where the data goes

DeepSeek is a China-based company running a China-hosted API. If your workload includes anything covered by GDPR, that's a cross-border data question worth answering before the next contract renewal.

Copilot buyers hit a similar surprise on their own bill this quarter, when GitHub changed the timing of seat charges rather than the price itself, and Google's own Flash line is set to double its price on a fixed date is the same genre of story from the calendar side, a scheduled increase a team can plan around months out. DeepSeek's version has no calendar. The name looks identical in every dashboard and every changelog; only the invoice tells you something moved. A team that picked DeepSeek this quarter and now has to weigh sticking with V4.1-Flash against evaluating alternatives is looking at the renewal math that applies whenever the tool market shifts underneath a Q1 pick.

The EU question: is calling a China-hosted API a GDPR problem?

China doesn't appear anywhere on the European Commission's list of adequacy decisions, the countries and territories the EU has formally recognized as offering an adequate level of data protection. That list runs from Andorra and Argentina through Japan, South Korea, Switzerland, the UK, and the US's Data Privacy Framework for commercial organizations. China isn't on it, in any category.

That absence is what creates a cross-border transfer question in the first place. It doesn't answer what to do about it. Which transfer mechanism applies, whether the data involved even triggers GDPR for your use case, whether existing contractual terms already cover it, is a legal question for counsel. Check the actual mechanism with counsel before routing anything regulated through a China-hosted endpoint.

Separately, if a system built on V4.1-Flash reaches EU users, a separate set of EU disclosure duties that apply regardless of GDPR status come into play under the EU AI Act's transparency rules. That's a different law with a different deadline. Checking one doesn't cover the other.

Frequently asked questions

What model actually handles deepseek-v4-pro requests now?

V4.1-Flash. As of 04:00 UTC on September 14, 2026, every request sent to deepseek-v4-pro routes to V4.1-Flash, billed at V4.1-Flash rates, according to DeepSeek's launch notes, until V4.1-Pro launches. The v4-pro name still works in your code. It's not the model you originally integrated.

Does deepseek-v4-flash still work in my code?

Yes, for now. DeepSeek retired the original V4-Flash and V4-Flash-Vision-Exp models, but calls to deepseek-v4-flash and deepseek-v4-flash-vision-exp route to V4.1-Flash for compatibility. The launch notes don't give a date for when that compatibility routing ends.

Is DeepSeek V4.1 Flash open-weights?

Yes. The weights are released under the MIT license on Hugging Face. A team with the infrastructure to run a 763-billion-parameter mixture-of-experts model can self-host it instead of calling DeepSeek's API.

Does China have an EU GDPR adequacy decision?

No. China doesn't appear on the European Commission's list of countries and territories with an adequacy decision. What that means for a specific transfer depends on the data involved, the safeguards in place, the mechanism your contract relies on. Check with counsel for what applies to your data flows.

How much more does peak pricing cost than off-peak?

Exactly double, across every line item DeepSeek publishes. Input cache-hit pricing for deepseek-flash goes from $0.003 to $0.006 per million tokens between off-peak and peak, and output goes from $0.6 to $1.2 per million tokens. Peak hours run 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now