Eric Lu started the clock on Devin's RSA-260 project at 00:11:58 Pacific time on August 13, 2026, with a single prompt to Cognition's coding agent. Three weeks later, on September 3, that project had helped factor RSA-260, a number that had stood as the largest RSA modulus ever publicly factored. That's the headline most outside coverage will run with. The steering behind it is the more interesting number, and it's the one that headline is likely to drop.
What Devin actually built, and how fast
Lu's opening prompt asked Devin to build "glas," a GPU drop-in replacement for las, the CPU lattice siever inside CADO-NFS, the open-source number field sieve implementation the project would end up relying on throughout. He gave Devin access to Modal's GPU infrastructure and let it start working.
About two hours in, before going to bed, Lu raised the bar: he added a stretch goal for glas to handle RSA-250-scale parameters, not just the smaller test cases it had been working with. Devin didn't clear that bar in those first two hours. It took "another 7 hours of iteration" after Lu set it, roughly nine hours from the initial prompt, before Devin had a version of glas that handled RSA-250-scale parameters and beat the CPU baseline it was replacing. Nine hours to a working, faster GPU siever from one prompt is still a fast start. It's a different claim than two hours, and worth stating precisely rather than rounding down to the number that sounds more impressive.
Does this mean RSA encryption is broken?
No, and Cognition's own post is direct about it. RSA-2048, the key size actually used to secure real systems, "remains roughly a billion times harder than RSA-1024 and does not appear to be meaningfully affected by this work," according to the write-up Cognition published on factoring RSA-260. The post adds that the efficiency gains from glas "have little impact on the feasibility of factoring RSA-2048-sized numbers with GNFS," the general number field sieve algorithm both RSA-260 and RSA-2048 factorizations would use. Nothing in the source material goes further than that, and neither does this article. If your key sizes were adequate last week, this result doesn't change that assessment.
The three-week, $400,000 factorization
Getting glas working was the fast part. Using it to actually factor RSA-260 took from August 13 to September 3, 2026, about three weeks, and consumed 4,900 GPU-days, roughly 13.5 GPU-years of compute. That split across three stages: 643 GPU-days on polynomial selection, 3,813 on lattice sieving, and 467 on solving the resulting sparse linear system.
At market GPU rates, Cognition estimates the run at around $400,000. That's roughly 10 times cheaper than the previous public state of the art, RSA-250, factored in February 2020. Some of that gap is glas itself; some of it is five years of GPU price and performance improvement that would have shown up in any GPU-based sieving effort, agent-built or not. The post doesn't separate the two, so neither does this article.
How much of this was Devin, and how much was Eric Lu
The session-by-session breakdown Cognition published counts 233 total Devin sessions across the project. Lu sent 3,328 messages totaling 82,702 words into 192 of those sessions used for factoring; separately, 36 sessions received no intervention at all. He kept an average of three sessions going at once, peaking at 18 concurrent, across the three-week run.
Lu doesn't dress this up as a hands-off result. In his own words: "I cannot claim that Devin iterated autonomously on the entire end-to-end pipeline." He describes his own contributions as "setting a hierarchy of goals and keeping Devin properly scoped," "recognizing when Devin was doing something unproductive and redirecting," "recognizing repeated inefficiencies in Devin's workflows," and "organizing experimental frameworks and results." That's a domain expert running near-continuous project management across nearly three weeks, not a single prompt left to unfold on its own. An agent that produces a record-setting GPU siever under 82,702 words of directed human steering is still a genuinely interesting result. It's a different result from one an agent produced alone, and the gap between those two claims is the whole story.
What made this task unusually easy to hand to an agent
Three conditions lined up here that don't line up for most engineering work.
First, Lu reports "essentially no algorithmic advancements" in the project. Implementing lattice sieving and sparse linear system solving on GPUs, in his words, required only "good old performance engineering," not new mathematics. Lu says he doesn't fully understand the mathematics his own project depended on: "I do not understand much of the underlying mathematics... I cannot tell you how the polynomial is used in the siever." The task was porting known techniques to new hardware, not inventing them.
Second, CADO-NFS already existed as an open-source reference. It "provided all the relevant techniques, the pipeline stages and their interfaces, and a reference CPU implementation," per Cognition's account. Devin had a working blueprint to translate, not a blank page.
Third, a factorization either produces the correct prime factors or it doesn't. That's a binary, machine-checkable success criterion, the kind almost no ordinary feature work gets. Most software changes don't self-verify; they need a human to judge whether the behavior is actually right, which is a slower and fuzzier loop than "did the number factor correctly."
One more detail belongs here: the compute ran "at no marginal cost on spare or fragmented compute" inside Cognition's existing training clusters, not a dedicated deployment built for the project. The $400,000 figure describes what that compute would cost at market rates, not what Cognition actually spent to run it.
What this does and doesn't tell you about AI agents
Take the pieces together and you get a real result on a narrow task: a performance-engineering port of known algorithms onto new hardware, with a self-verifying output, run under close and constant supervision from someone who understood the engineering even where he didn't understand the mathematics. That's worth taking seriously as a data point about what a coding agent can do when a domain expert stays in the loop for weeks.
What it isn't is evidence about unattended agent operation, and it isn't a productivity study of ordinary feature work. Most engineering tasks don't ship with a binary correctness check, an existing reference implementation to port, or an expert available to intervene in 192 of 233 sessions. Nor did Devin start from unfamiliar code; CADO-NFS handed it a working blueprint, not a blank page, as noted above. That matters because familiarity isn't something AI reliably substitutes for: a study on AI-assisted engineer ramp time found experienced developers working in codebases they already knew well were 19% slower with AI enabled, and that AI assistance didn't reliably close the gap for new hires unfamiliar with the code. The same caution applies whenever a single narrow result gets generalized past what it measured. A separate lab study on AI code suggestions shows how easily a controlled finding gets stretched into a claim about production work it never tested. This is that same pattern, run in reverse: an unusually favorable task, honestly disclosed, that outside coverage will likely compress into "an AI agent did expert cryptography research alone." It didn't. Cognition's own numbers say so, and that candor is worth more than the headline it complicates. It's also worth contrasting with vendors that publish a capability number and stop there. The same scrutiny applies to a vendor's own reported PR-merge numbers when there's no session-level detail behind them; Cognition published the detail, which is exactly why the RSA-260 result holds up to a closer look instead of falling apart under one.
Frequently asked questions
Does Devin's RSA-260 factorization mean RSA encryption is broken?
No. Cognition states that RSA-2048, the key size actually used to secure systems today, "remains roughly a billion times harder than RSA-1024 and does not appear to be meaningfully affected by this work."
How autonomous was Devin during the RSA-260 project?
Not fully, and Eric Lu says so directly: "I cannot claim that Devin iterated autonomously on the entire end-to-end pipeline." He sent 3,328 messages across 192 of the 233 sessions the project used. Only 36 sessions ran without any intervention from him.
How long and how much did the full factorization take?
About three weeks, from August 13 to September 3, 2026, and 4,900 GPU-days of compute, estimated at roughly $400,000 at market GPU rates. That's about 10 times cheaper than the prior public record, RSA-250, factored in February 2020.
Does this result predict how AI agents will perform on ordinary software engineering work?
No, and Cognition's post doesn't claim it does. The task had a binary, machine-checkable success criterion, an existing open-source reference implementation to port, and near-continuous supervision from a domain expert across most of its sessions. Most engineering work has ambiguous specs, no automatic correctness check, and nowhere near that level of hands-on direction. This result says nothing about how agents perform on merged pull requests or everyday feature work, because the source material never measures that.
