GPT-6 Astra reached the API around September 3, 2026, the date OpenAI's own system card for the model states as its publication point. If your team is weighing a move off GPT-5.6, the benchmark scores aren't the part to look at first. Three things in your integration code will break or narrow the moment you switch models: how you call tools, how you set reasoning effort, and where your EU traffic is allowed to land.
What GPT-6 Astra costs and how big its context window is
The core numbers, per OpenAI's model specs page for gpt-6-astra:
| Spec | Value |
|---|---|
| Input | $10 per million tokens |
| Cached input | $1 per million tokens |
| Output | $50 per million tokens |
| Context window | 1,050,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | April 30, 2026 |
| Tier 1 rate limits | 500 requests/minute, 500,000 tokens/minute |
The specs page attaches no end date to that rate card. Gemini 3.8 Flash's scheduled price increase is the counterexample: its launch price holds through the end of 2026 and doubles on January 1, 2027.
OpenAI's migration guide for Astra (linked in the next section) makes a cost claim worth budgeting around: the model tends to reach a stronger result while spending far fewer output tokens to get there, which can lower what a given task actually costs even though the per-token rate hasn't changed. If your team tracks spend per token rather than spend per completed task, this release is a reason to switch which number you watch.
How to migrate tool calling to GPT-6 Astra
1. Switch tool calling to the Responses API
Astra still answers plain chat requests through Chat Completions. Its tool calling doesn't: "GPT-6 Astra supports Chat Completions, but its tool calling requires Responses," per OpenAI's own migration guidance for the model. Any function-calling code still pointed at Chat Completions needs to move to the Responses API before Astra's tools work at all.
2. Replace reasoning_effort: none with low
Astra drops support for the none reasoning-effort setting entirely. Anywhere your pipeline currently sends none, switch it to low.
3. Drop temperature, top_p, and top_logprobs
Once reasoning effort isn't set to none, Astra doesn't accept temperature, top_p, or top_logprobs in the same call. Any code that hardcodes those parameters alongside a reasoning-effort value needs a conditional for this model, or the request will fail.
EU data residency and what it does and doesn't cover
Astra's EU residency guarantee comes with one condition attached. The same guidance states it plainly: "EU data residency is available only with Standard processing." A GDPR-conscious team routing Astra traffic through the EU needs to confirm which processing mode its actual integration uses, because the residency guarantee doesn't extend automatically to every way of calling the model. Confirm it before compliance signs off on where Astra traffic lives.
Why OpenAI rates Astra "Critical" for cybersecurity, and what that means operationally
Astra is the first OpenAI model to reach the "Critical" level of cybersecurity capability under OpenAI's Preparedness Framework, according to that same system card. The system card says that rating means "the model can identify previously unknown security vulnerabilities and develop novel exploitation methods across well-protected systems with minimal human guidance." The same system card also states that monitorability has decreased relative to GPT-5.6 Sol, a separate problem from raw capability: it's harder to observe what Astra is doing while it's doing it.
A team running Astra inside Codex or any agentic pipeline should treat that combination, higher capability alongside lower visibility, as a reason to apply real controls. The isolation-tier and kill-switch controls the NCSC recommends for agentic AI are the closest match: credential scoping, defined oversight modes, and folding agent activity into whatever monitoring your security team already runs.
Don't reduce any of this to a single capability score either. What Devin's RSA-260 result does and doesn't show is a good example of reading past a capability headline before it goes in a deck.
Behavioral changes worth testing before rollout
Two habits shift with Astra, according to OpenAI's own notes on how Astra's behavior differs from GPT-5.6. It's "more likely to ask the user a question when additional input could materially change the result," and it's "more sensitive to instructions contained in skills and other files" than its predecessor was. Audit AGENTS.md and skills files for stale or conflicting guidance before rollout: Astra will follow what's written there more literally than GPT-5.6 did.
That shift matters more the further your pipeline leans on written instructions instead of a person typing in the moment, which is the same territory as what happens to CI once agents write most of the code. The review burden moves from the code to the instructions that produced it.
The wider GPT-6 family: Sol and Luna in GitHub Copilot
Astra isn't the only new OpenAI model to land this month. GPT-6 Sol and GPT-6 Luna went into GitHub Copilot on September 22, 2026, per GitHub's own changelog announcing the two models, which ties them directly back to "the previously released GPT-6 Astra."
GitHub describes Sol as "a balanced model for interactive and agentic coding," a "strong all-round choice for development tasks that benefit from careful, multistep validation," available on Copilot Pro+, Max, Business, and Enterprise plans. Luna is "a lightweight, cost-efficient model for smaller, faster tasks and the lowest-cost option in the GPT-6 family," available more broadly, on Pro, Pro+, Max, Business, and Enterprise. Both are billed under usage-based billing, and GitHub says rollout will be gradual, so don't expect every seat to see either model on day one.
That billing detail is worth checking against GitHub's own Copilot billing changes this quarter: a new usage-based model landed in Copilot on September 22, weeks before seats start getting charged upfront on October 1, so review the two changes together.
FAQ
What does GPT-6 Astra cost through the API?
$10 per million input tokens, $1 per million cached input tokens, and $50 per million output tokens, per OpenAI's model specs page. The model supports a 1,050,000-token context window and up to 128,000 tokens of output, with a knowledge cutoff of April 30, 2026.
Do I need to change my tool-calling code to use GPT-6 Astra?
Yes, if it's currently built on Chat Completions. Astra's tool calling only works through the Responses API. You'll also need to replace any reasoning_effort: none setting with low, and drop temperature, top_p, and top_logprobs from calls that aren't using none effort.
What does OpenAI's "Critical" cybersecurity rating for Astra mean?
Astra is the first OpenAI model to reach the Critical level of cybersecurity capability under OpenAI's Preparedness Framework, meaning it can discover unknown security flaws and develop exploitation methods with minimal human guidance. OpenAI's system card also notes that monitorability has decreased relative to GPT-5.6 Sol, a separate concern from the capability rating itself.
Are GPT-6 Sol and GPT-6 Luna the same as GPT-6 Astra?
No. They're separate models in the same family. Both landed in GitHub Copilot on September 22, 2026, weeks after Astra shipped. Sol is positioned as a balanced, all-round coding model; Luna is the lightweight, lowest-cost option.
