September 26, 2026

AI coding agents and what the research actually shows

AI coding agents keep making capability claims. Here's how to tell a vendor benchmark from a controlled study before you act on either.

Tech

News

AI coding agents don't fail the way junior engineers fail, and treating them like an underperforming hire is the first mistake most engineering leaders make. An agent that writes a clean pull request on a scoped bug fix can spin for hours on a multi-file refactor, produce code that passes review but ships a security hole, or start a task nobody asked it to start. Deciding what work to hand an agent, and how to interview an engineer who works alongside one, means understanding the specific failure modes rather than the marketing copy that sits on top of them.

The material here spans benchmark write-ups, including long-horizon Android tasks that test whether an agent can hold a multi-step goal across an unfamiliar codebase, and AGENTS.md, the convention that tells an agent how a repository expects work to get done. Coordinator agents that start work unprompted, before anyone assigns a ticket, get covered too, alongside what agentic coding does to CI load once an agent can open a pull request as fast as it can think of one.

There's coverage of the security of AI-suggested code, of ramp time and speed studies on developers working alongside an agent, and of capability signals like the RSA factoring result that vendors point to as proof an agent has crossed some meaningful line. Each of those deserves a different kind of scrutiny, because a benchmark score, a security audit, and a productivity study measure entirely different things and get quoted as if they're interchangeable.

A vendor's capability claim and a controlled study read differently once you check what each one actually measured, even when they cite the same benchmark. Anthropic, OpenAI, and Google all publish results that look like research and read like marketing, because the model that produced the result is also the model they're selling. Before quoting any of it to a hiring committee or a board, check the task set, the baseline it was compared against, and whether a human graded the output or another model did. Most of the interesting failures live in that fine print, not in the headline number a press release leads with.

Every guide on AI coding agents

FAQ

Should a coding agent get the same code review as a human engineer?

Yes, and arguably a stricter one. An agent won't push back when a reviewer asks it to justify a decision the way a human engineer might, and it won't flag its own uncertainty unless it's specifically prompted to. Treat an agent's pull request as untrusted output from a fast contractor: read the diff, not just the test result.

Who's accountable when an AI coding agent ships a bug?

The engineer who approved the merge, the same as with any other pull request. An agent doesn't carry accountability, it executes a task and returns an artifact, so the review discipline around it has to be at least as strong as it would be for a junior hire, arguably stronger given how confidently an agent can present broken code as finished work.

How do you interview an engineer for AI-assisted work?

Ask them to describe a time an agent's output looked right and wasn't, and how they caught it. Engineers who've only ever accepted agent suggestions without pushback usually can't answer that with any specificity. The ones worth hiring have a habit of treating agent output the way they'd treat a pull request from someone they don't fully trust yet.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now