AI coding agents don't fail the way junior engineers fail, and treating them like an underperforming hire is the first mistake most engineering leaders make. An agent that writes a clean pull request on a scoped bug fix can spin for hours on a multi-file refactor, produce code that passes review but ships a security hole, or start a task nobody asked it to start. Deciding what work to hand an agent, and how to interview an engineer who works alongside one, means understanding the specific failure modes rather than the marketing copy that sits on top of them.
The material here spans benchmark write-ups, including long-horizon Android tasks that test whether an agent can hold a multi-step goal across an unfamiliar codebase, and AGENTS.md, the convention that tells an agent how a repository expects work to get done. Coordinator agents that start work unprompted, before anyone assigns a ticket, get covered too, alongside what agentic coding does to CI load once an agent can open a pull request as fast as it can think of one.
There's coverage of the security of AI-suggested code, of ramp time and speed studies on developers working alongside an agent, and of capability signals like the RSA factoring result that vendors point to as proof an agent has crossed some meaningful line. Each of those deserves a different kind of scrutiny, because a benchmark score, a security audit, and a productivity study measure entirely different things and get quoted as if they're interchangeable.
A vendor's capability claim and a controlled study read differently once you check what each one actually measured, even when they cite the same benchmark. Anthropic, OpenAI, and Google all publish results that look like research and read like marketing, because the model that produced the result is also the model they're selling. Before quoting any of it to a hiring committee or a board, check the task set, the baseline it was compared against, and whether a human graded the output or another model did. Most of the interesting failures live in that fine print, not in the headline number a press release leads with.
Every guide on AI coding agents
- Android Bench 2.0: what it says about AI vs senior Android engineers: Google's Android Bench 2.0 shows AI agents pass 28% of long-horizon tasks. What that means for which Android work you can hand to AI today.
- What is AGENTS.md, and why Claude Code now reads it: AGENTS.md is the open coding-agent standard. Claude Code reads it as a CLAUDE.md fallback (Sept 18, 2026): what to put in it, and the Bedrock/Vertex gap.
- What agentic coding does to CI, and how Anthropic fixed it: Anthropic's CI jobs grew 25x in six months once Claude began writing most of the code. Three quick fixes failed before a rebuild worked. What it means for you.
- Devin's RSA-260 siever: a capability signal, not a security one: Devin helped factor RSA-260 over 3 weeks and 4,900 GPU-days, with Eric Lu steering closely. The authors say RSA-2048 isn't affected. What the result shows.
- Cursor Projects: what changes when an agent decides when work starts: Cursor's coordinator starts work unprompted, including on PR events. What that changes, why its 30% claim is thin, and the gates to set first.
- What a 100-developer study found when AI code suggestions varied in security: Only 22% of final code submissions were fully secure in a 100-developer study testing AI code suggestions. Here's what that means for review policy.
- How long does it take an AI-assisted engineer to become productive?: DX's Q4 2025 cohort comparison found AI-assisted developers reaching their 10th pull request in 49 days versus 91 for non-AI users, drawn from a dataset of roughly 400 companies. Here's what that gap does and doesn't tell you about ramp time.
- How much faster are AI-assisted developers? What the studies actually show: Microsoft measured a 55.8% speedup on a controlled task. METR measured 19% slower on real work. Here's what determines which number applies to your team.
FAQ
Should a coding agent get the same code review as a human engineer?
Yes, and arguably a stricter one. An agent won't push back when a reviewer asks it to justify a decision the way a human engineer might, and it won't flag its own uncertainty unless it's specifically prompted to. Treat an agent's pull request as untrusted output from a fast contractor: read the diff, not just the test result.
Who's accountable when an AI coding agent ships a bug?
The engineer who approved the merge, the same as with any other pull request. An agent doesn't carry accountability, it executes a task and returns an artifact, so the review discipline around it has to be at least as strong as it would be for a junior hire, arguably stronger given how confidently an agent can present broken code as finished work.
How do you interview an engineer for AI-assisted work?
Ask them to describe a time an agent's output looked right and wasn't, and how they caught it. Engineers who've only ever accepted agent suggestions without pushback usually can't answer that with any specificity. The ones worth hiring have a habit of treating agent output the way they'd treat a pull request from someone they don't fully trust yet.
