September 26, 2026

Grok 4.7 release: pricing, context window, availability

Grok 4.7 shipped Sept 21, 2026 at $2/$6 per M tokens, GA in Cursor and the Grok API. What its pricing tiers mean for a Q4 2026 model pick.

News

Insight

xAI shipped Grok 4.7 on September 21, 2026, priced at $2 per million input tokens and $6 per million output tokens, the same rate the vendor quotes for Grok 4.6. That flat headline is only half the bill. The numbers worth checking before rolling Grok 4.7 into production live in xAI's separate pricing documentation: a prompt that crosses 200,000 tokens bills at exactly double the headline rate, and routing traffic through the US endpoint adds another 10% on top.

What xAI shipped on September 21, 2026

Grok 4.7 shipped as a generally available release. xAI's Grok 4.7 launch post states the model is available today in Cursor and Grok Build, plus through the Grok API, third-party coding tools, and model routers and cloud platforms. That's a wide release surface for a day-one launch.

xAI frames the release as an improvement over Grok 4.6 on both benchmarks and general capability, comparable to other frontier models. The launch post backs that with a benchmark table: 46.3% on CursorBench 4.0, 71.0% on DeepSWE v1.1 at high reasoning effort, 64.0% on EEBench, a score of 1,657 on AA Briefcase v1.1, 37.6% on Terminal-Bench 4.0, 19.6% on the Harvey Legal Agent Benchmark, 56.7% on HealthBench Professional, 62.4% on LatchBio's biosafety benchmark, and a 3.3% risky-prompt rate on HackerBench v0.3. None of those suites overlap with what OpenAI, Google, or Anthropic report for their own Q4 2026 releases, so a direct benchmark comparison across vendors isn't available regardless of price.

What Grok 4.7 costs through the API, and where the price doubles

Standard pricing starts at $2 per million input tokens and $6 per million output tokens. xAI also serves a fast variant at twice the price with twice the output speed, aimed at teams that need lower latency more than the lowest bill. The same launch post adds that Grok 4.7 is "served at the same price and speed as Grok 4.6," a rare case where a new flagship model doesn't force a repricing decision on top of a capability one.

That flat headline rate only holds inside a threshold, and the threshold doesn't appear on the launch post at all.

The 200K-token pricing cliff

xAI's API pricing page sets a separate tier once a single prompt crosses 200,000 tokens; the threshold applies per prompt, and a long session made up of shorter prompts stays under it. Below that line, input runs $2 per million tokens, output $6 per million, and cached input $0.50 per million. Cross it, and the whole request bills at the higher tier: input doubles to $4 per million, output doubles to $12 per million, and cached input doubles to $1 per million. A team running large codebases or long documents through Grok 4.7 should check its average prompt size against that line before estimating spend off the $2/$6 headline.

xAI isn't the only vendor building a usage threshold into this quarter's pricing. DeepSeek's own peak-pricing tiers split the day into off-peak and peak windows instead of prompt length, but the effect lands the same way: the sticker price only covers part of the traffic.

The US regional endpoint premium

The same pricing page also sets a 10% surcharge on the US regional endpoint: input, output, and cached input all bill at 1.1x the global rate once traffic is pinned to US infrastructure. For a team routing to the US endpoint for data-residency reasons, that surcharge is a real, quantifiable line item worth building into a spend forecast.

Context window, knowledge cutoff, and where it runs

xAI's model documentation for Grok 4.7 lists a 500,000-token context window, more than double the Grok 4.6 default most teams have tuned prompts around, and a knowledge cutoff of May 2026. xAI calls it "the most capable model we've built," a framing every vendor applies to its newest release, so check it against your own eval suite before taking it at face value.

A team that's already wired Grok into a router or a coding tool doesn't need a new integration path to try 4.7. Swap the model string, and re-run the eval suite before pinning it in production.

Grok 4.7 vs. GPT-6 Astra vs. Gemini 3.8 Flash vs. Opus 5.5: the Q4 2026 price comparison

ModelInput $/MOutput $/MContext windowStatus
Grok 4.7$2$6500,000 tokensGA, Sept 21, 2026
GPT-6 Astra$10$501,050,000 tokensGA, Sept 3, 2026
Gemini 3.8 Flash$0.75 (rising to $1.50 on Jan 1, 2027)$3.75 (rising to $7.50 on Jan 1, 2027)1,000,000 input tokens (64K output)GA
Opus 5.5$4$201,000,000 tokensGA, Sept 22, 2026

Price alone doesn't settle a Q4 2026 model pick, since none of the vendor-stated benchmark suites overlap, but it narrows the field fast.

At $2 input and $6 output, Grok 4.7 undercuts GPT-6 Astra's own pricing and context window by a wide margin: Astra runs $10 input and $50 output, with a much larger 1,050,000-token context window attached to that bill. Gemini 3.8 Flash's scheduled price increase still undercuts Grok 4.7 on raw price, at $0.75 input and $3.75 output through the end of 2026, though that rate doubles on January 1, 2027.

Anthropic's Opus 5.5, released the day after Grok 4.7 on September 22, 2026, sits between the two on price: $4 per million input tokens and $20 per million output, per Anthropic's own announcement. Anthropic states the model "costs 40% less than Opus 5 on typical workloads." That comparison is against Anthropic's own prior model, so it doesn't change where Opus 5.5 lands against Grok 4.7 here.

What changes for a team already running Grok 4.6

A team already on Grok 4.6 can treat this as a low-friction upgrade. Neither the launch post nor the pricing docs mention a forced API or parameter migration, unlike GPT-6 Astra's move to the Responses API for tool calls. Still, don't just flip a pinned model name in production. xAI says Grok 4.7 "uses a new, larger base model compared to Grok 4.6," and a larger base model can shift behavior even when price and latency stay put. Re-run evals before the switch goes live.

That eval work matters more the longer a model pick sits unexamined. A Q1 tool pick already two quarters stale was made on data that's aged out by the time most teams get around to revisiting it. Four flagship models landed within the same few weeks this quarter: Grok 4.7, GPT-6 Astra, Gemini 3.8 Flash, and Opus 5.5, each with its own pricing structure and its own migration cost. Whatever a team picks now is worth calendaring a review for next quarter.

FAQ

What does Grok 4.7 cost through the API?

$2 per million input tokens and $6 per million output tokens, the same rate xAI charged for Grok 4.6. Prompts over 200,000 tokens bill at double that rate on both sides, and traffic routed to the US regional endpoint carries a 10% surcharge on every tier.

Is Grok 4.7 generally available?

Yes. It's available now in Cursor and Grok Build, plus through the Grok API, third-party coding tools, and model routers and cloud platforms. It shipped September 21, 2026 without a preview or beta phase.

What is Grok 4.7's context window?

500,000 tokens, more than double Grok 4.6's default, with a knowledge cutoff of May 2026.

Do I need to change my code to move from Grok 4.6 to Grok 4.7?

No forced migration is stated on any of xAI's primary pages, unlike GPT-6 Astra's required move to the Responses API for tool calls. Price and speed stay the same. xAI does say Grok 4.7 runs on a new, larger base model, so re-run your evals before pinning the new model name in production.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now