September 22, 2026

Gemini 3.8 Flash: what its price doubling means for coding agents

Gemini 3.8 Flash's introductory price holds through 2026, then doubles January 1, 2027. What that does to a coding-agent budget, and what the benchmarks show.

Insight

Google shipped Gemini 3.8 Flash on September 2, 2026, alongside a restricted sibling called Gemini 3.8 Flash Cyber that most teams will never get to use. The GA model is the one that matters for anyone running coding agents on Google's stack, and Gemini 3.8 Flash pricing is why: introductory rates hold through the end of 2026, then the price doubles on January 1, 2027. If a team has standardized on Flash for agent workloads, that date belongs on next year's budget, not in a footnote.

What Gemini 3.8 Flash costs, and when the price changes

The numbers: $0.75 per million input tokens and $3.75 per million output tokens, in effect through December 31, 2026. From January 1, 2027, those rates become $1.50 per million input tokens and $7.50 per million output tokens, according to Google's launch post announcing 3.8 Flash and 3.8 Flash Cyber. That's not a modest adjustment. Both the input price and the output price exactly double, on the same day, for every customer still on the model past the new year.

DeepMind's model card for the GA model shows the same two price points but labels them differently: $0.75/$3.75 is "no caching" pricing, and $1.50/$7.50 is "regular" pricing, with no calendar dates attached to either. The launch post is the only place Google frames these as an introductory rate expiring December 31 and a standard rate taking over January 1. Use the model card to confirm the figures; use the launch post to know when they change.

What the price doubling actually does to a monthly agent bill

Take a hypothetical team running a coding agent against Gemini 3.8 Flash at 50 million input tokens and 10 million output tokens a month. That volume is a round number picked to make the arithmetic legible, not a measured or reported usage figure from any real account. At introductory pricing, input runs $37.50 (50 x $0.75) and output runs $37.50 (10 x $3.75): $75 a month combined. At standard pricing from January 1, 2027, the same volume costs $75 for input (50 x $1.50) and $75 for output (10 x $7.50): $150 a month.

Same usage, double the bill, because both sides of the rate sheet doubled at once. Scale the hypothetical up to a team burning ten times that volume and the gap goes from $75 to $750 a month, all arithmetic on Google's own published rates and nothing else. The real number for any given team depends entirely on actual token volume, which only that team's own usage dashboard knows. The part that's certain: whatever a team pays under introductory pricing today, budget for roughly double it from January.

Coding benchmarks: DeepSWE, and two Terminal-bench scores that aren't the same test

Gemini 3.8 Flash scores 73.7% on DeepSWE v1.1, according to DeepMind's model card for the model. On Terminal-bench, the same card lists two numbers that look contradictory until you check the fine print: 89.4% on Terminal-bench 2.1 and 19.1% on Terminal-bench 4.0.

Those aren't the same test. Terminal-bench 2.1 and Terminal-bench 4.0 are different benchmark versions, listed on the model card as separate entries rather than a before-and-after on one suite. Reading 89.4% next to 19.1% as a collapse in capability is the wrong read. It's evidence that version 4.0 is a harder benchmark, not that the model got worse at anything between the two rows. Whether either number maps to a real team's delivery speed is a separate question, and what the research actually shows about AI-assisted coding speed is a reasonable check to run before a benchmark score goes into a purchasing decision.

Context window, knowledge cutoff, and where to run it

The same model card puts Gemini 3.8 Flash's context window at up to 1 million input tokens and 64K output tokens, with a knowledge cutoff of March 2026 (some domains are limited to January 2025). That's enough context to hold a sizeable codebase in a single prompt, which matters more for agent workloads than a benchmark table does.

The GA model runs in Google AI Studio, Android Studio, Gemini Enterprise, the Gemini app, the Gemini API, Google AI Mode, and Google Antigravity, per the model card, with the launch post adding Google Sheets and Search AI Mode to that list. That's a wide surface for a model that's four months from a price increase.

Gemini 3.8 Flash Cyber: a gated sibling, not something most teams can use

Gemini 3.8 Flash Cyber shipped the same day as the GA model, and it's worth naming for one reason: to be clear it isn't available to whoever's reading this. Access is limited to the Fairwind Program eligibility rules Google set out in the same post, which cover government authorities, critical infrastructure operators, and software maintainers. There's no public pricing, no self-serve signup, and no path for a typical engineering team to try it this year.

Everything in this article about pricing, benchmarks, and availability applies to the GA Flash model, not to Cyber. If a team doesn't sit in one of those three categories, Cyber simply isn't a purchasing option right now.

What this means if you're already running a coding agent

A price doubling four months out is a renewal-planning input, not an emergency. A team with Gemini 3.8 Flash baked into an agent pipeline has from now until January 1, 2027 to either budget for double the per-token cost or requalify the workload against a cheaper option before the standard rate takes over.

That's the same category of decision GitHub's own changes to Copilot Business and Enterprise billing this quarter forced on teams paying for coding-agent seats, just on the model-pricing side of the stack instead of the seat side. And the renewal math that applies when a tool market shifts underneath a Q1 pick is exactly the position a team that adopted Flash under introductory pricing this September will be in come January. Whichever way a team leans, the decision now has a hard date attached to it, not a vague sense that prices might move eventually.

Frequently asked questions

When does Gemini 3.8 Flash's price go up?

Introductory pricing, $0.75 per million input tokens and $3.75 per million output tokens, holds through December 31, 2026. Standard pricing, $1.50 per million input tokens and $7.50 per million output tokens, starts January 1, 2027, per Google's launch post.

Is Gemini 3.8 Flash Cyber available to everyone?

No. Cyber is restricted to the Fairwind Program: government authorities, critical infrastructure operators, and software maintainers. There's no general-availability or self-serve path for anyone outside those categories.

Do the two Terminal-bench scores mean the model got much worse on harder tasks?

No. The 89.4% score is on Terminal-bench 2.1 and the 19.1% score is on Terminal-bench 4.0, two different benchmark versions listed separately on the model card. They aren't directly comparable, and the gap reflects benchmark difficulty rather than a drop in capability.

Where can I use Gemini 3.8 Flash?

Google AI Studio, Android Studio, Gemini Enterprise, the Gemini app, the Gemini API, Google AI Mode, Google Antigravity, and Google Sheets.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now