October 7, 2026

How to run a coding interview where candidates can use AI

How to run an AI-assisted coding interview: pick the task, set up tools, watch the right signals, score verifying and overriding, and know when to ban AI.

Guide

Recruiting

To run an AI-assisted coding interview, allow the tool, change the task so the tool can't finish it alone, and score how the candidate directs, verifies and overrides the output. This is our design, drawn from what Coinbase, Canva and the interviewing vendor Karat have published. No study shows it predicts job performance, and we won't pretend otherwise.

Why does a coding interview need to change when candidates use AI?

Because the standard questions stop discriminating. Coinbase, in its July 2026 write-up of rebuilding its engineering interviews, says that when it turned on full AI assistance for its standard frontend coding questions, "the AI solved most of them directly." It names two failure modes in the old format: false positives, where a candidate passes on preparation rather than capability, and false negatives, where a strong engineer is marked down for lacking recall. Coinbase also reported "an 84% correlation between two different interview rounds"; our read is that the two rounds were measuring much the same thing. Its engineers now review code that AI wrote, and AI crossed 50% of merged code in Q4 2025. Our read: if engineers now mostly review AI-written code, a clock-and-memory test measures less of the job than it used to.

Canva reached the same conclusion from the other side. In its post on letting candidates use AI, Simon Newton says that when Canva tested its Computer Science Fundamentals questions with AI tools, they produced correct, well-documented solutions in seconds. Canva now expects backend, ML and frontend candidates to use tools such as Copilot, Cursor or Claude in technical interviews.

Both companies kept AI in the room and changed what they ask. That's the pattern the rest of this article follows. AI as the candidate's tool is a different question from AI as the interviewer, which where AI screens fit and where an engineer is needed covers.

What should you test in an AI-allowed coding interview?

Three things: how the candidate directs the tool, how they evaluate what it returns, and whether they know where it fails.

Coinbase calls its version AI Fluency and splits it into Usage, Application and Understanding Limits, and says the definition applies to junior and senior engineers alike. Karat's published rubric for human-AI interviews lands in similar territory with code productivity, product sense and problem solving, and AI proficiency, which covers prompting and evaluating generated code for correctness and risk. Karat is a vendor selling interviewing, so read it as one design, not a standard.

The honest limit: nobody has shown these signals predict performance on the job. Coinbase itself is running 45- and 90-day pulse surveys and says it will publish what it learns. Treat everything below as a reasonable hypothesis you test on your own hires.

How to run an AI-allowed coding interview

1. Decide where AI is allowed in your loop

Set a policy per stage instead of one rule for the whole process. Our read: AI on for the repository exercise, and off for a short fundamentals or reasoning check if you want that signal separately. Meta shows one version of a controlled setup. Its hiring page says select roles include "an authorized AI assistant within CoderPad," and that no outside AI tools are authorized. Coinbase says it doesn't add rounds, and a small team shouldn't either. Change one round of the technical screening process you already run. For reference, HighCircl's own vetting has four stages run by senior engineers, and about 1 in 10 applicants pass.

2. Design a task where AI gets you started but not finished

Use a small repository with a planted fault, an ambiguous spec, or a constraint a model is likely to miss. Coinbase's backend questions are repository-based: debugging, code review for correctness and performance, and rollback scenarios. Canva says it uses realistic product problems. KORE1, a recruiter, suggests replacing LeetCode-style take-homes with code review on code that contains planted bugs.

Before you use any task, run it through a current model yourself. If one prompt solves it, it won't separate candidates. That's our recommendation, not a published method. A generic example: a payments service with a retry bug that only appears when a timeout and a duplicate request overlap, plus a one-line spec that leaves the idempotency behaviour open.

3. Set up the environment and tell candidates the rules

Let candidates use their own editor and AI tool, or provide one. If you provide one, give them a practice session first, as Meta does. Send the rules and the format in advance: what's allowed, how long it runs, what you'll be watching for. Ask candidates to say their prompts aloud as they type them. Whether to record is a policy choice, so decide it once and state it.

Confirm who is actually on the call before any technical round. For spotting hidden or undisclosed AI use, see running a live session and catching hidden AI use. In an AI-allowed format you want the tool in the open, so Canva's stance is a good default: transparency over detection.

4. Observe the same behaviours every time

Write the list down and watch for it in every interview. How does the candidate frame the task? What do they ask before they prompt? Do they read the output, or paste it on? What do they run? What do they reject? How do they recover when the first answer is wrong?

Canva's account of the good and the bad is useful here. Candidates with little AI experience often struggled, with poor judgment about how to guide the tools. Strong candidates asked clarifying questions, used AI for defined subtasks and reviewed what came back. For the on-the-job version of these habits, see what senior engineers do with an agent loop.

5. Score process and outcome separately

Karat's principles are to separate outcomes from process, treat AI as an engineering resource, and score observable behaviour only. Use the rubric below. A candidate can reach a working fix by pasting blindly, and another can run out of time after excellent verification. Those need different scores.

Have two interviewers score independently before they compare notes. That's our recommendation, and it matters more here than in a classic interview because the behaviours are easier to over-read.

6. Review and rotate the task every quarter

Questions have a half-life. Coinbase runs quarterly review cycles and retires questions that stop producing signal. Anthropic had the same experience with a take-home test: successive Claude models beat it, and the team redesigned it more than once. Keep two or three variants of your task and swap one each quarter. If you can, log each candidate's scores next to their performance at 45 and 90 days, so your rubric gets tested against your own hires. Coinbase uses 45- and 90-day pulse surveys for a similar purpose.

How to score prompting, verifying and overriding AI output

This table is our design. It's built from the criteria in Karat's, Coinbase's and Canva's published material, and it's not a validated instrument. Score each dimension from 1 to 5, using the anchors at 1, 3 and 5.

DimensionWeak (1)Solid (3)Strong (5)
Framing and promptingPastes the task in whole and accepts the replyBreaks the task into parts and gives the tool the relevant codeAsks clarifying questions first, states constraints, uses the tool for subtasks
VerificationAccepts output without running itRuns the code and the existing testsReads the diff, writes a test that targets the risky path, checks edge cases
OverridingKeeps a wrong suggestion, or rewrites everything by handSpots a wrong suggestion and fixes itRejects it for a stated reason and keeps the design decision they already made
CommunicationSilent, or narrates after the factExplains choices when askedExplains what they asked for, what they expected and why they changed course
OutcomeFault not foundFault fixed with a gapFault fixed, tested, and the trade-off named

If this starts to look like a checklist of the same agent-loop signals described elsewhere, cut dimensions. In our experience, five dimensions is about as many as a two-person panel can score reliably.

When should you not allow AI in a coding interview?

In our view, switch it off when you need a fundamentals floor for the role: work close to the metal, or on algorithms, where an engineer who can't reason without a tool is a risk. Use a short AI-off segment for that, and keep the rest open.

Be careful with very junior hires, where you're testing how someone learns. Coinbase says its fluency definition applies to junior and senior engineers alike, so don't treat juniors as a blanket exception. An AI-off reasoning segment can be enough.

Switch it off in regulated environments where the tools can't see your code, since the candidate can't work in a setup that mirrors the job.

And skip any task a current model solves without real help from the candidate. Anthropic's take-home is a cautionary example: successive Claude models beat it and it had to be redesigned. If AI is on and the task is trivial, you're measuring typing.

What this means for your next technical hire

If you hire senior engineers, you're deciding whether the interview should test the work they do now. Coinbase's finding that AI wrote more than half of its merged code in Q4 2025, with humans reviewing roughly all of it, is the case for testing review and judgment over recall. Canva's observation that candidates with minimal AI experience often struggled is the case for including some AI use in the format.

Our recommendation: change one round, not the whole loop. Pilot the AI-allowed format on your next two senior hires, score them with the rubric above, and compare those scores with how they perform at 90 days. If your team has 10-40 engineers and no interview platform, a candidate's own editor and a shared screen is enough to start.

The next hiring-bar decision is seniority mix, and whether to hire juniors now that AI writes the first draft is where to take it.

FAQ

Should you let candidates use AI in a coding interview?

Yes, if the job allows it. Canva and Coinbase both allow it and changed the task to suit. Decide per stage, as in step 1, and tell candidates in advance.

What should an AI-allowed coding interview test?

How the candidate directs the tool, verifies its output and overrides it when it's wrong. It shouldn't test typing speed or recall of algorithms. The rubric above turns that into scores.

Do you need special tooling to run an AI-allowed interview?

No. A candidate's own editor, their own AI tool and a shared screen are enough. Meta's provided assistant inside CoderPad and Coinbase's setup are examples of what larger teams build, not requirements. That's our judgment, not a published finding.

How do you stop candidates from simply pasting the prompt into a model?

You don't try to. You want them to use the model. Pick a task that needs follow-up questions and judgment, and the pasted answer won't be enough on its own. For detecting undisclosed use, the live coding article covers it.

Can junior engineers be interviewed with AI too?

Yes. Coinbase says its AI Fluency definition applies to junior and senior engineers equally. Use the same dimensions. Our read: expect less on the outcome and on independence, but still expect them to run and check what the tool gives them.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now