September 27, 2026

Claude API refusal billing: what changed and what to fix

Anthropic now bills some refusals before any output. Here's which categories are billed, what stays free, and what to fix in retry and cost-monitoring code.

News

Claude API refusal billing changed on September 24, 2026, and it can move your bill even though you sent no new traffic. Anthropic started charging for some refusals that arrive before the model produces a single token of output, a response type that used to cost nothing. The change hits Claude Fable 5.1, Claude Fable 5, Claude Opus 5.5, and Claude Opus 5, on every platform that runs them: Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry.

What changed on September 24, 2026

Anthropic's own release note spells out the mechanism. According to the September 24, 2026 entry in Anthropic's release notes, "we're resuming billing for refusals that arrive before any output when stop_details.category is "bio", "frontier_llm", or "reasoning_extraction", the categories where we measure low volumes of false positives." The note goes on to confirm two things stay put: mid-stream refusals were already billed, and fallback credit is unchanged.

The word "resuming" matters. Anthropic ran this billing rule before, turned it off, and has now switched it back on for three specific categories. The other categories stay free before output, and the models affected are limited to the four named above, the ones running Anthropic's current safety classifiers.

Which refusal categories are billed now, and which stay free

A refusal on the Messages API is a normal HTTP 200 with stop_reason set to "refusal". content is empty, and the detail lives in stop_details, a field that's null for every other stop reason, per Anthropic's refusals and fallback docs. stop_details carries type, category, explanation, and recommended_model (the last one populated only when fallbacks is set).

Six category values exist, and only three of them trigger a charge when the refusal lands before any output:

CategoryBilled before output?
bioYes
frontier_llmYes
reasoning_extractionYes
cyberNo
general_harmsNo
nullNo

Anthropic frames the billed three as the categories where it currently measures low false-positive rates, and it says directly that the billed list may change as it keeps refining those measurements. Treat this table as the state of play in September 2026. Anything that hardcodes these three category names as a fixed list will need a second look whenever Anthropic updates the docs.

How much a billed refusal costs, and how mid-stream refusals differ

A pre-output refusal in a billed category costs the same as any other request, charged at the rates of whichever model ran it. So the model you pin also sets the price of a billed refusal, whichever platform serves it.

Mid-stream refusals, where the model starts streaming and then declines partway through, were already billed before September 24 and stay that way: input tokens plus whatever output already streamed, at normal rates. Nothing about that path moved. And regardless of category or timing, a refused request still counts against your rate limits and still reports token counts in usage, billed or not.

What this means for fallback and retry logic

Fallback credit itself hasn't changed. A refusal can still carry stop_details.fallback_credit_token and stop_details.fallback_has_prefill_claim, and redeeming that token on a retry still avoids re-paying the cache-write cost of warming the prompt on a fallback model, per Anthropic's fallback-credit docs. Both fields are null when no credit is available, exactly as before.

What did change is the cost that arrives ahead of the fallback. When a billed-category refusal triggers a retry on a different model, a team now pays for the refusal itself, on top of the fallback request that follows it. Before September 24, that first leg was free; now it isn't. That combined cost, refusal plus fallback, is the number to model going forward. The fallback credit mechanism itself is doing exactly what it did last month.

If you're running server-side fallback, each attempt bills independently under the same before-output or mid-stream rule, and each one's tokens land on its own usage.iterations entry. A three-attempt fallback chain that hits bio twice before succeeding on the third model now shows two billed refusals and one successful completion in that array, where it used to show two free refusals followed by a charge.

What to update in cost monitoring and budget alerts

Any dashboard or alert built on the assumption that a refusal with no output is always free now under-counts spend for bio, frontier_llm, and reasoning_extraction traffic. That's a one-line fix: branch on stop_details.category instead of on whether output was empty.

If your logging already captures stop_reason for observability but never reads stop_details.category, run a one-time breakdown of refusals by category before your next invoice lands, so you know the exposure rather than finding it on the bill. And if your team built its own retry ladder outside the SDK's middleware or the server-side fallbacks parameter, check that it accounts for billed and unbilled refusals per attempt. The built-in paths already handle both cases correctly; a hand-rolled one might not.

Billing shifts like this one keep landing on our AI model releases and pricing: what changes for engineering teams coverage, and this one has close company. Teams already re-pricing after the same week's Opus 5.5 pricing and breaking-change list have a reason to fold this into the same pass, since Opus 5.5 is one of the four models this billing rule touches. It's also worth comparing notes with a similar billing-rule change from GitHub Copilot, where GitHub moved seat charges from arrears to upfront with about five weeks' notice. Both change the invoice without touching your code.

FAQ

What is stop_details.category, and where do I read it?

It's a field on the Messages API response, populated only when stop_reason is "refusal". stop_details itself is null for every other stop reason. Alongside category, it also carries type, explanation, and recommended_model (present only when fallbacks is set).

Does this billing change apply to every Claude model?

No. It's scoped to the models running Anthropic's current safety classifiers: Claude Fable 5.1, Claude Fable 5, Claude Opus 5.5, and Claude Opus 5. Older or unrelated models aren't part of this update.

Will Anthropic add more billed categories later?

Possibly. Anthropic's own framing treats the billed list as something it will keep adjusting as it measures false-positive rates per category, not a fixed set. Check the field reference in Anthropic's refusals and fallback docs before hardcoding today's three categories anywhere.

Does a refused request still count against my rate limit?

Yes. That's true whether the refusal is billed or not, and whether it arrives before any output or mid-stream. Rate limits track the request itself, not whether Anthropic charged for it.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now