Anthropic's CI job volume grew 25x in six months. That number describes one company's own codebase, reported by that company, about a tool it sells. No control group, no comparison against another engineering org, no independent audit. Keep that in mind for every figure below, because all of them trace back to the same place: Sachin Malhotra's September 14, 2026 account of how agentic coding strained CI at Anthropic, a single post with a single named author. What makes it worth reading anyway is the sequence it walks through: three fixes that each worked for a while and then didn't, followed by a rebuild that treated the real problem instead of the symptom.
The numbers behind Anthropic's CI crunch, and why they're not a benchmark
Anthropic says its engineers now ship 8x as much code per quarter as they did across 2021-2025, and that Claude authors 80% of it. That's the upstream cause. The downstream effect is CI job volume up 25x over six months, and total tests in the codebase growing 10x over the same window. More code, written faster, by a system that doesn't get tired, produces more pull requests, and every pull request runs the full test suite until something forces a change.
Read those numbers as what they are: one company's internal telemetry about its own dogfooding of its own product, published on its own blog. Anthropic doesn't disclose cost figures, headcount, or how any of this compares to a team not building the AI in question. Its own generalizability claim is a prediction, not a benchmark: "Scaling CI is a challenge more engineering teams are likely to soon face as agents continue to accelerate code generation and review." That's an interested party forecasting a trend it has a commercial reason to want true, not a finding about the industry. It happens to line up with HighCircl's reading of where DORA's 2024 findings point: testing and deployment pipelines as the likely place AI-driven throughput gains turn into delivery instability, an inference drawn from DORA's data rather than a conclusion DORA itself states. Anthropic's post is the concrete case that reading was gesturing at, not proof it generalizes.
Three stopgaps, three failures
Before the rebuild, Anthropic tried three fixes, in order, each buying less time than the one before it. The pattern is the useful part: a team recognizing a symptom, patching it, watching the patch stop working, and repeating that twice more before addressing the actual cause.
A bigger machine (70 days)
The first move was doubling CPU cores. That addressed runtime on any single job, which is a reasonable first guess when tests are running slow. It lasted 70 days before job volume outgrew the extra capacity. Doubling compute doesn't help when the number of jobs, not the length of any one job, is the thing compounding.
Sharding by package (29 days)
Next came splitting the suite by package and running shards in parallel. This bought less time than the first fix, 29 days, because it also treated the wrong variable. Sharding parallelizes execution of a fixed suite; it does nothing when the suite itself, and the number of pull requests demanding a full run of it, keeps growing underneath the fix.
Daily restarts (under a day)
The last stopgap was restarting the CI service daily to clear accumulated state. It held for less than a day. By this point, according to the post, tests that were "already super flaky" were compounding fast enough that flaky reds started blocking merges almost as soon as a restart finished. State was piling back up faster than a daily reset could clear it.
Every one of these fixes treated CI runtime as the problem. The actual problem was structural: the full suite ran on every pull request regardless of what that pull request actually touched. No amount of faster hardware or more frequent restarts changes that math.
The fix: a deterministic test-impact-analysis service
Anthropic's rebuild has two deterministic components. A listener records test results from every CI run. A selector reads that history and decides which tests actually need to run on a given pull request, instead of running everything every time. That's the structural change the three stopgaps never touched.
The redesign also added an in-memory data store to offload processing, which made the service, in the post's own words, "stateless and hence, horizontally scalable." That's the direct answer to why the earlier patches kept degrading under load: a stateful single-instance service accumulates exactly the kind of pressure that daily restarts were trying, and failing, to relieve.
One engineer built it in three weeks. Anthropic's own estimate is that the same project would have taken closer to a quarter a year earlier. The post doesn't explain that gap, it doesn't credit Claude Code with writing the service, so there's nothing here to attribute beyond what's stated. What is stated, and worth taking at face value, is the pressure that made the rebuild necessary in the first place: "Writing code is no longer the constraint, and once PR review gets accelerated, CI starts feeling the pressure." That line, more than the multipliers around it, is the part likely to hold true regardless of whose codebase you're running.
What this doesn't tell you if you don't have a spare senior engineer for three weeks
Anthropic's post names a coming problem for other teams and then stops. Its one line about broader relevance, that more engineering teams are likely to soon face this, isn't followed by any guidance. That's a real gap, not an oversight worth forgiving: a company that just solved this for itself had every reason to publish a checklist for teams without its bench, and didn't.
Most engineering orgs don't have a senior engineer they can pull off other work for three uninterrupted weeks to rebuild CI infrastructure from scratch. That's worth saying plainly instead of skating past it. Before reaching for a custom listener and selector, the lower-effort move is diagnosis: is your CI pain runtime-bound or volume-bound? Anthropic's own sequence shows that question is exactly what its first two stopgaps got wrong. A bigger machine and package sharding both assumed the problem was how long each job took. It wasn't. If your team is burning time on faster hardware or smarter parallelization and the pain keeps coming back on a shorter cycle each time, that's the same signal Anthropic got twice, first after 70 days, then again after 29 more. Worth checking before committing to a rebuild is whether your CI provider already offers some form of test selection or impact analysis you haven't turned on, rather than assuming a from-scratch build is the only option.
There's also a code-provenance angle Anthropic's post doesn't raise, because it isn't the post's subject. When 80% of a codebase's new code comes from an AI system, what your CI needs to catch changes along with what your CI needs to run, and a faster, better-targeted test suite doesn't automatically catch the failure modes that matter most in AI-authored code. Test-impact analysis solves throughput. It doesn't solve what the tests are checking for. And as agents move from writing code to triggering CI runs and merges on their own, the same question that governs an agent deciding when to fix CI and merge without being asked applies here too: faster CI is only good news if the merges it clears are ones a human, or a deliberately scoped policy, actually wanted cleared. That's a governance question, and it sits next to the security side of wiring agents into CI/CD rather than replacing it. Solving throughput and solving who gets to act on that throughput are two separate projects, and Anthropic's post is only about the first one.
Frequently asked questions
How much did Anthropic's CI job volume grow?
25x over six months, according to Anthropic's own September 2026 blog post. The company attributes the growth to shipping roughly 8x more code per engineer per quarter than its 2021-2025 average, with Claude authoring about 80% of that code.
What is a test-impact-analysis service, and how does Anthropic's version work?
It's a service that decides which tests actually need to run on a given pull request instead of running the full suite every time. Anthropic's version has two deterministic parts: a listener that records results from every CI run, and a selector that uses that history to choose which tests run on each new pull request. An in-memory data store made the service stateless and horizontally scalable.
Why did Anthropic's first three fixes fail?
Each one treated CI runtime as the bottleneck when the actual problem was volume: the full suite running on every pull request regardless of what changed. Doubling CPU cores lasted 70 days, sharding tests by package lasted 29 days, and daily service restarts bought less than a day of relief before flaky tests and accumulated state caught back up.
Is the 25x CI growth figure a benchmark for other engineering teams?
No. It's Anthropic's self-reported figure about its own codebase, with no control group, no cross-company comparison, and no independent audit. Anthropic's own forward-looking claim, that more engineering teams are likely to face similar pressure, is a prediction from a company that sells the AI coding tool in question, not a measured industry benchmark.
How long did it take Anthropic to build the fix?
One engineer, three weeks. Anthropic says the same project would have taken closer to a quarter a year earlier, though the post doesn't explain what closed that gap and doesn't attribute the build itself to Claude Code.
