October 6, 2026

How senior engineers use AI coding agents in 2026

How senior engineers run AI coding agents: the loop, the checks, the review habits, and what to test for when you hire. Sources from Anthropic and a 2025 study.

Guide

Tech

Senior engineers who get useful work out of AI coding agents don't prompt harder. They build a loop around the agent. That's what engineers who write about their workflow describe, and it's what the one observational study of experienced developers found. That loop has a check the agent can run, a small step, a clean context, and a review that isn't the agent's own. Our guides on AI in the engineering workflow: tools, permissions and policy cover governance. This one covers practice.

What an AI coding agent loop actually is

Between your prompt and the diff, the agent is doing something specific. Simon Willison's short definition is that an agent runs tools in a loop to achieve a goal, and he adds that loops suit problems with clear success criteria. Anthropic's post on building effective agents draws the line differently: workflows follow predefined code paths, while agents direct their own process and tool use. The same post says it's "crucial for the agents to gain 'ground truth' from the environment at each step". Tool results are what keep the agent from guessing.

Claude Code's own docs describe three phases: gather context, take action, verify results. Each tool result feeds the next decision, and you can interrupt at any point. Matt Pocock makes the point that matters for control in What Is An Agent?: the agent, not your code, decides which tool to call and when to stop.

That last part is where the loop gets its human gates. Here's the loop as a text diagram.

Illustrative loop, not a vendor spec:

text
goal
  |
  v
plan  <--- human gate: approve or cut the plan
  |
  v
act (edit files, run commands)
  |
  v
observe tool result
  |
  v
verify against a check (tests, build, types, lint)
  |
  +-- fails --> back to act (or /clear and re-plan)
  |
  v
passes
  |
  v
review in a fresh session  <--- human gate: read the diff
  |
  v
merge or reject

If you'd rather see the loop in code, we build an agent loop in code with approval and cost limits in a separate tutorial. The difference is that the tutorial builds an agent, and this article is about using one.

The diagram has one consequence worth stating plainly. Without a check, the agent decides it's finished when the work looks finished. The Claude Code docs say as much: "Claude stops when the work looks done." The practices in the sources replace that judgment with a check.

Six practices behind useful agent output

Give the agent a check it can run

The Claude Code best practices page puts it in one line: "Give Claude a check it can run: tests, a build, a screenshot to compare. It's the difference between a session you watch and one you walk away from." Anthropic's agents post makes the same case from the other side: code suits agents because solutions are verifiable through automated tests, and the agent can iterate on the results.

Addy Osmani calls testing "the single biggest differentiator between agentic engineering and vibe coding", in his post on agentic engineering. Pocock's tips for AI coding with Ralph Wiggum list the usual feedback sources (types, tests, lint, hooks) and say more of them means higher quality. His line on speed: "Never outrun your headlights."

In our view, the first question on any task is how you'll know it's done. If the answer is "I'll read it and see", the task isn't ready for an agent yet.

Plan first, then execute in small steps

The Claude Code docs describe four phases: explore, plan, implement, commit, with plan mode separating exploration from execution. They also give the exception that stops this becoming ceremony: "If you could describe the diff in one sentence, skip the plan."

The most relevant behavioural evidence we found is a December 2025 paper from Huang, Reyna, Lerner, Xia and Hempel, a study of experienced developers working with agents. It observed 13 developers and surveyed 99, where "experienced" meant at least three years of professional work (median nine years in the field observations; the paper says "experienced", not "senior"). The authors conclude that "professional developers do not vibe code. Instead, they carefully control the agents through planning and supervision." The detail is more useful than the headline. The largest plans ran past 10 steps, and two exceeded 70, but they were executed in small chunks. The median was 1.8 steps per prompt, and two participants never ran more than six at once.

Our reading: a step is small. Think one prompt, a couple of edits, one check. The authors suggest, without having tested it, that explicit multi-step plans in separate files work well when executed in manageable chunks.

One task per iteration

Geoffrey Huntley's original Ralph loop is a shell one-liner, while :; do cat PROMPT.md | claude-code ; done, built on one thing per loop and tests as backpressure. Pocock's Getting Started with Ralph describes it as "a technique for running AI coding agents in a loop". In his version a PRD plus a progress file carry state between iterations, and each iteration takes exactly one task.

Why small? Because a failure in a small step is cheap to diagnose and cheap to throw away. A failure in a forty-file change is neither. Our judgment: the "one task" rule is the cheapest practice on this list and an easy one to skip, because a big prompt feels like progress.

Keep context short, current and clean

Anthropic's best practices page calls the context window "the most important resource to manage." Its post on effective context engineering explains why: as the number of tokens grows, the model's ability to recall information from that context decreases. The remedies it names are compaction, structured note-taking and sub-agent architectures with clean context windows.

The habits that follow are small. After two failed corrections on the same issue, the docs say to /clear and rewrite the prompt, rather than keep arguing in a polluted session. The project context file stays short and gets pruned, and the docs suggest asking of each line whether removing it would cause mistakes. Pocock's guide to AGENTS.md agrees: "The ideal AGENTS.md is small, focused, and points elsewhere." If you're deciding between file formats, we cover how Claude Code reads AGENTS.md separately. Step 2 below shows what a short one looks like.

Shape the codebase so the agent can be checked

Pocock's article on codebases agents love argues that codebase shape affects output more than the prompt does, with deep modules and tested interfaces as the model. His framing is that AI is a new starter with no memory, and a new starter does well in a repo with clear seams and fast tests, and badly in one where behaviour is spread across everything.

Our judgment, and an implication leaders tend to miss: this turns technical debt into an agent-productivity cost. A module with no tests can't be given a check, so the agent can't be walked away from there. The cleanup work seniors used to argue for on maintainability grounds now has a direct payoff you can point to.

Isolate, parallelise, and review with fresh eyes

Worktrees give parallel sessions isolated git checkouts, so they don't collide, according to the best practices page. Willison recommends sandboxes, credential limits and budget limits for loops that run unattended. The governance side of that sits in our policy pieces, not here.

Review is where the sources are strictest. The same docs page recommends a writer and a reviewer in separate sessions, because a fresh context won't be biased toward code it just wrote. It also recommends asking the agent to show test output or the commands it ran, rather than asserting success. And it has the rule that should hang on a wall: "If you can't verify it, don't ship it."

The study found the same habit in people. Nine of the 13 observed developers carefully reviewed every agentic change, and three who let agents drive on unfamiliar tasks still monitored decisions and outputs closely. Osmani's post on loop engineering says the verification burden stays human and warns of comprehension debt, the gap between the code you've merged and the code you understand. His closing line is the best summary of the senior stance: "Build the loop. But build it like someone who intends to stay the engineer, not just the person who presses go."

How to run an agent loop on a real task

This is a pilot a lead can hand to a team, built on a real Ralph loop. The scripts and prompts below are Matt Pocock's, from AI Hero, reproduced with his credit. Where something is ours, it says so. Pick a task where "done" is a command that passes (a failing test, a bug with a reproduction) and keep the first one boring on purpose. The PATH fix below edits ~/.bashrc; on macOS, where zsh is the default shell, use ~/.zshrc instead.

1. Install Claude Code and start it in a sandbox

Install the native binary with the command from Matt Pocock's Ralph starter guide on AI Hero:

Bash
curl -fsSL https://claude.ai/install.sh | bash

If your shell says "command not found: claude", his fix is to add the install location to your PATH:

Bash
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bashrc
source ~/.bashrc

Then install Docker Desktop (his article says 4.50 or later) and start Claude Code inside its sandbox from your project folder:

Bash
docker sandbox run claude

You authenticate on first run, and per the article the credentials are stored in a Docker volume.

One caveat as of October 2026: Docker has since removed the docker sandbox command. On Docker 29.8.1 it prints ""docker sandbox" is deprecated and has been removed" and points to Docker Sandboxes, whose CLI starts the same agent with sbx run claude. If the command above errors on your machine, install sbx, swap it in here and in the unattended script in step 6, and check sbx run --help for how it passes flags through to Claude. The loop itself doesn't change.

2. Write a short AGENTS.md

Keep the project rules file to what applies to every task. In his AGENTS.md guide, Pocock says the root file needs only a one-sentence project description, the package manager if it isn't npm, and build or typecheck commands if they're non-standard. His example of what to cut is a list of language rules like this:

markdown
Always use const instead of let.
Never use var.
Use interface instead of type when possible.
Use strict null checks.

The guide's list runs longer. Replace the lot with one pointer to a separate file, which he calls a light touch, with no "always" and no all-caps:

markdown
For TypeScript conventions, see docs/TYPESCRIPT.md

Claude Code reads CLAUDE.md, so the guide symlinks it to AGENTS.md:

Bash
# Create a symlink from AGENTS.md to CLAUDE.md
ln -s AGENTS.md CLAUDE.md

Commit both so the team shares one version.

3. Turn requirements into PRD items

Start Claude Code and press shift-tab to enter plan mode, which Pocock's Ralph article uses to iterate on the plan before anything is written. Then, from his Ralph tips, give it this prompt to convert your feature requirements into checkable items:

text
Convert my feature requirements into structured PRD items.
Each item should have: category, description, steps to verify, and passes: false.
Format as JSON. Be specific about acceptance criteria.

Each item has the shape Pocock gives, with a passes flag the loop flips to true when the item is done:

JSON
{
  "category": "functional",
  "description": "New chat button creates a fresh conversation",
  "steps": [
    "Click the 'New Chat' button",
    "Verify a new conversation is created",
    "Check that chat area shows welcome state"
  ],
  "passes": false
}

Tell Claude to save the result to PRD.md, as the Ralph article does, and cut any item you couldn't check in one iteration. Then create the progress file the loop appends to:

Bash
touch progress.txt

4. Add feedback loops

The loop is only as good as its checks. Pocock's tip 5 gives this prompt, which makes typecheck, tests and lint all pass before a commit:

text
Before committing, run ALL feedback loops:
1. TypeScript: npm run typecheck (must pass with no errors)
2. Tests: npm run test (must pass)
3. Lint: npm run lint (must pass)
Do NOT commit if any feedback loop fails. Fix issues first.

Swap in your own commands if your project isn't a Node one. We'd keep these lines in AGENTS.md so every iteration reads them, which is our choice, not his.

5. Run one watched iteration

This is Matt Pocock's ralph-once.sh from the Ralph article: one iteration, human in the loop.

Bash
#!/bin/bash

claude --permission-mode acceptEdits "@PRD.md @progress.txt \
1. Read the PRD and progress file. \
2. Find the next incomplete task and implement it. \
3. Commit your changes. \
4. Update progress.txt with what you did. \
ONLY DO ONE TASK AT A TIME."

Make it executable and run it:

Bash
chmod +x ralph-once.sh
./ralph-once.sh

--permission-mode acceptEdits auto-accepts file edits so the loop doesn't stall, and @PRD.md @progress.txt pulls the plan and progress files into the prompt. Note that this script calls claude directly, with no sandbox, so it runs on your machine. Watch what the agent does when a check fails, because that tells you whether your check is any good. Pocock's tips pair this mode with learning and prompt refinement. Our rule: don't move to step 6 until a few watched runs in a row behave.

6. Run unattended with an iteration cap

Go away from keyboard only once the watched runs behave. This is Matt Pocock's afk-ralph.sh from the same article:

Bash
#!/bin/bash
set -e

if [ -z "$1" ]; then
  echo "Usage: $0 <iterations>"
  exit 1
fi

for ((i=1; i<=$1; i++)); do
  result=$(docker sandbox run claude --permission-mode acceptEdits -p "@PRD.md @progress.txt \
  1. Find the highest-priority task and implement it. \
  2. Run your tests and type checks. \
  3. Update the PRD with what was done. \
  4. Append your progress to progress.txt. \
  5. Commit your changes. \
  ONLY WORK ON A SINGLE TASK. \
  If the PRD is complete, output <promise>COMPLETE</promise>.")

  echo "$result"

  if [[ "$result" == *"<promise>COMPLETE</promise>"* ]]; then
    echo "PRD complete after $i iterations."
    exit 0
  fi
done

It differs from the watched script in three ways. It runs inside the Docker sandbox, -p switches Claude to non-interactive print mode so the script can capture the output, and the prompt tells the agent to run tests and type checks. The argument is the iteration cap ($1), which Pocock says prevents runaway costs. The loop also exits early when the agent prints the <promise>COMPLETE</promise> sigil, which the script greps for. His tips suggest 5-10 iterations for small tasks and 30-50 for larger ones. His example run uses 20:

Bash
./afk-ralph.sh 20

His article's instruction after that is "Go make coffee. Come back to commits." His tips say loops usually take 30-45 minutes. One tradeoff he notes: inside the sandbox, your global AGENTS.md and user skills don't load.

7. Review the diff in a fresh session

Don't let the session that wrote the code grade it. Open a new session (or run /clear) and give it this prompt. It's ours, not Pocock's:

text
Review the commits made by the loop against PRD.md, as a skeptic.
List any PRD item marked passes: true that you can't confirm, any test that would pass without testing the behaviour, and any change outside the PRD.
Run the test suite and show the commands you ran and their output.

Asking for commands and output instead of "it works" is the habit the Claude Code docs recommend. Then read the diff yourself. The fresh session catches what the writing session rationalised, and your read catches what neither understands. If a correction fails twice in the same session, /clear and rewrite the prompt.

8. Record what the loop cost and caught

Note the time, the tokens or spend, the number of times a check failed, and what the review rejected. After three or four tasks you have a real picture of where agents help in your codebase. Those notes are the raw material for any later rollout measure.

What the evidence says about results

Do these practices make experienced engineers faster? We found no study that measures these practices against outcomes, and the Huang paper itself calls for that validation.

The nearest evidence is mixed. METR's 2025 study of early-2025 AI tools gave 16 experienced open-source developers 246 issues. They expected a 24% speedup and afterwards believed they'd gained 20%, but "When developers are allowed to use AI tools, they take 19% longer to complete issues". In its February 2026 update, METR called the new data "an unreliable signal of the current productivity effect", since 30% to 50% of developers said they held back tasks they didn't want to do without AI. DORA's 2025 report found AI adoption positive for delivery throughput and negatively related to delivery stability, and its summary line is "AI doesn't fix a team; it amplifies what's already there." For the studies side by side, read what the speed studies actually show.

Trust follows the same pattern. The Stack Overflow 2025 survey found 46% of developers distrust AI accuracy against 33% who trust it, and 66% cite "almost right, but not quite" as a frustration. That's the failure mode the loop exists to catch.

Our read: the data supports caution about speed claims, and it supports the claim that the surrounding practice matters more than the model. It doesn't prove any one habit above pays for itself.

When senior engineers do not use agents

Where do the best engineers keep their hands on the keyboard? The Huang survey has tallies. Among open-ended survey answers coded by task, mentions of small, simple or straightforward tasks as suitable outnumbered mentions as unsuitable 33 to 1 (following well-defined plans 28 to 2, complex tasks 3 to 16, high-level plans 13 to 23). The survey recruited from AI-focused repositories, so it leans positive. The mean suitability rating respondents gave their own task was 4.73 of 6, so experienced developers do find real uses for agents. The pattern is well-specified work in; for high-level planning and design, experienced developers disagree.

Anthropic's advice is the same in spirit: add multi-step agentic systems "only when simpler solutions fall short". The best practices page lets you skip the plan when the diff fits in a sentence, and our extension of that logic: you can skip the agent when typing is faster than explaining.

Adoption sits in the same place. In the Stack Overflow survey, 14.1% of respondents say they use agents daily at work, 9% weekly, and 37.9% say they don't use them and don't plan to.

Our judgment: restraint is a signal, not a lag. An engineer who says "I wouldn't give an agent this migration, the design is still moving" is applying the same test they apply to a junior's pull request. Be wary of the candidate who has never found a task to keep.

What to look for when you hire a senior now

Which interview signals predict someone who gets good output from agents? We found no study that validates any. What follows is our judgment, derived from the practices above, not a tested predictor.

SignalProbeRed flag
Defines done as a runnable check before prompting"Walk me through the last task you gave an agent. What told you it was finished?""I read the diff and it looked right"
Decomposes work and can show a planAsk for a plan file or notes from a recent taskPlans exist only in their head, or steps are huge
Keeps a context file and explains it"What's in your project rules file, and what did you remove?"No file, or a file that's pages long
Reviews agent diffs and shows what they rejected"What did the agent get wrong last week?"Can't name a single case
Names tasks they wouldn't delegate"Where do you keep the agent out?"Delegates everything, or nothing
Knows cost and sandbox limits"What can your agent touch, and what can it spend?"Broad credentials, no budget cap

Two exercises work well. Ask for a walk-through of a real recent agent session, with the plan and the review. Or give them a small repo with one failing test, let them use their own tools, and watch how they set up the check before they start. This is a different job from the one in our piece on running a live coding interview when candidates use AI, which is about catching hidden AI use. Here the AI is allowed, and what you're testing is the loop.

What this means for your team shape and hiring bar

The decision is what to require from a senior now that the job includes supervising an agent loop, and whether your seniority mix changes.

Our judgment: loop skills compound with seniority. The checks, plans and reviews all need someone who knows what breaks in production, and a loop built by someone who doesn't will pass its own tests and still ship the wrong thing. The evidence is thin: the one controlled study that tested seniority (Google's, covered in our speed-studies article) found no significant difference.

So change the interview loop before you change the team. Add the session walk-through and the failing-repo exercise, and score the six signals above. Then pilot the eight steps on one team for a month and look at what the loop cost and caught.

Don't read this as "no juniors". The evidence here is about experienced developers and says nothing about how people become one. For that decision, read whether a startup should still hire juniors.

FAQ

What is an agent loop in AI coding?

It's the cycle an AI coding agent runs between your request and the finished diff: gather context, act with tools such as file edits and commands, observe the result, and verify against a check, repeating until the check passes or a human stops it. Willison defines an agent as something that runs tools in a loop to achieve a goal. Seniors add the check and the review gates.

Do experienced developers trust agents to write production code?

Not without supervision. In the 2025 observational study, nine of 13 developers carefully reviewed every agentic change, and the authors say professionals plan and supervise rather than vibe code. In Stack Overflow's 2025 survey, 46% of developers distrust AI accuracy and 33% trust it. Trust is conditional on checks and review.

Are AI coding agents faster for senior engineers?

The evidence doesn't settle it. METR's 2025 study of early-2025 AI tools found experienced open-source developers took 19% longer while believing they were faster, and its 2026 follow-up was too affected by selection to be a reliable signal. DORA's 2025 report found higher throughput alongside lower delivery stability. No study measures the practices in this article against outcomes.

How do you interview a senior engineer for agent skill?

Ask for a walk-through of a recent agent session, including the plan, the check and what they rejected in review. Or give them a small repo with a failing test and let them use their own tools. Look for a runnable definition of done, small steps, a short context file, named tasks they'd keep from the agent, and awareness of cost and sandbox limits. These are our judgment, not validated predictors.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now