October 2, 2026

AI vs human technical screening: where each fits

What the main AI interview study measured, how Arc and Turing describe their vetting, and which screening stages still need an engineer in the room.

Guide

Recruiting

In AI vs human technical screening, our view is to automate the top of the funnel and keep an engineer on any stage that judges reasoning. AI handles volume, scheduling and consistent first-pass questions. The widely cited study on AI interviews didn't test engineers.

That study is worth reading closely, because the headline number travels without its setting. So are the vendor pages that sell "AI vetting", which often describe something narrower than the phrase suggests.

What counts as AI technical screening

Three different things get filed under one label, and mixing them up is where most bad comparisons start.

The first is AI matching. A model reads profiles and job descriptions and proposes candidates. Nobody is being assessed on anything yet, so it's closer to search than to screening.

The second is automated testing: multiple-choice knowledge checks, auto-graded coding challenges, and take-homes scored by a model. The format is fixed and the grading is mechanical, which is why it scales.

The third is an AI-run interview, where a voice or video agent asks questions and scores the answers. This is the one people usually mean when they ask whether AI can replace a human interviewer.

What a major AI-interview field experiment measured

A widely cited study is Jabarian and Henkel's natural field experiment on automated job interviews. It's not about engineers.

The setting was PSG Global Solutions, a Teleperformance subsidiary, hiring for customer-service roles in the Philippines that paid roughly $280-435 a month. The authors report 70,884 applications, of which 67,056 were randomized, handled by 131 recruiters. AI voice agents ran the interviews. The headline result is that applicants interviewed by an AI voice agent were 12% more likely to get an offer than those interviewed by a human recruiter, with gains in starts and retention too.

Two details matter more than the headline. Every hiring decision stayed with a human recruiter, so the experiment tested AI as an interviewer, not AI as a decider. In about 5% of AI interviews the applicant said they were unwilling to keep talking to the AI, and in 7% the AI voice agent ran into technical difficulties.

None of the roles were software engineering. Every role in the sample was entry-level customer service. A senior engineer defending a design trade-off is a different task. If a page uses the 12% figure to argue that AI interviews work for engineers, it's borrowing a result from a different job.

What Arc and Turing say their vetting is

"AI vetting" sounds like one thing. The vendors' own pages describe different things.

Arc's HireAI page calls HireAI "our GPT-4-powered AI recruiter" and promotes it for instant candidate matches. That's the first category above: matching. On the same page, Arc says developers must pass a profile screening, complete a behavioral interview, and pass a technical interview or pair programming, and that it accepts the top 2% of applicants. The page doesn't say who conducts those interviews. What it does show is that HireAI is a matching tool, not an AI screening interview.

Turing's developer test preparation article, dated February 21, 2025, walks through its stages: profile setup, a work experience survey (multiple choice), tech stack tests (multiple choice), and a live coding challenge. It also describes an "AI Matching Engine," AIME. The article doesn't mention a human interview. That's what the article describes, not proof that Turing has none. This is based on a single Turing post from February 2025, so it may not reflect Turing's current process. (HighCircl's own comparison found Turing's current pages don't describe a vetting process, so treat this post as dated.) Turing claims the top 1%.

PlatformAI component, in their wordsHuman component, in their words
ArcHireAI, "our GPT-4-powered AI recruiter" for candidate matchesLists a behavioral interview and a technical interview or pair programming; doesn't say who conducts them
TuringAIME, an "AI Matching Engine"; multiple-choice survey and tech stack testsNot stated. The article lists a live coding challenge but doesn't say who runs or scores it; no human interview is mentioned

If you want the vendor-level comparison with pricing and fit, how Toptal and Turing compare for EU teams and Arc's HireAI against human-led alternatives cover it. This article is about the method.

Where AI-run screens break for engineers

The format breaks before the AI does, and cheating is the clearest example.

Karat, which sells human-led technical interviews and so has a stake here, reports a tech leader's claim that 80% of candidates used a large language model on a top-of-funnel code test despite a ban on it. That's one leader's number, not a measured rate, and it comes from a vendor with a view. But the second point in the same post holds up on its own logic: Karat says telling AI-written code from human-written code by the output alone is "nearly impossible," and that the difference becomes clear when you watch the process.

That's the real problem for any unobserved test. A take-home or auto-graded challenge hands you finished output. If the output can't tell you who wrote it, the score tells you less than it appears to. Adding an AI grader on top doesn't fix this. It grades the same artifact faster.

The second break is follow-up. A good engineering interview turns on the moment you ask "why did you pick that?" and the answer either holds or doesn't. A scripted agent can ask a follow-up, but in our view an engineer who has shipped and maintained systems that failed is better placed to judge the answer. Senior signal lives in how someone handles a constraint they didn't expect.

What we'd rely on is observation. Watching someone reason live, with questions that adapt to what they just said, is harder to fake than finished output (Karat's point: the difference becomes clear when you watch the process), and running a live coding interview and catching AI cheating covers how to do it. Identity fraud is a separate problem from skill cheating, and screening for fake candidates before the first call deals with that one.

Where AI helps

AI earns its place where the task is repetitive, the criteria are explicit, and a wrong call is cheap to reverse.

Scheduling is the obvious one. Nobody's judgment improves by coordinating calendars by hand. Consistency is the second: an automated first screen asks every applicant the same questions in the same order, and the paper describes AI-led interviews as more structured and consistent. Volume is the third. If you have thousands of applicants for a role, no team of engineers will interview them all.

Is it enough on its own? Depends on how many interviews you run. HackerEarth, which sells assessment tooling, says that below roughly 50 interviews a month, a well-designed coding test plus a human interviewer usually outperforms its listed platforms on cost and candidate experience. Treat it as vendor opinion with no data behind it on that page.

How to split the stages

This is a placement table, not a procedure. It reflects the evidence above.

StageAutomate?Why
Sourcing and matchingYesSearch over profiles; a human reviews what comes out
Scheduling and logisticsYesNo judgment involved
Knowledge checks (multiple choice)Yes, with low weightCheap and consistent, but answers can be looked up or put to a model
Unobserved coding test or take-homePartlyOutput can't show who wrote it, so a human must discuss it with the candidate afterwards
Behavioral and communication interviewKeep humanThe study covered entry-level customer-service calls; this isn't one
Review of a take-homeKeep humanThe question is why they made each choice, which only a conversation reveals
Live architectural sessionKeep humanJudging reasoning needs follow-up questions from someone who has built similar systems

The pattern is simple. Automate stages where you're sorting. Keep humans where you're judging.

What engineer-led vetting means

HighCircl's approach is the human end of this table: four stages run by senior engineers, covering background and experience verification, a communication and product-thinking assessment, a take-home project that mirrors real work, and a live technical session on architectural reasoning. About 1 in 10 applicants pass. If you're building your own version, a technical screening process that senior engineers run lays out the steps.

FAQ

Is AI screening good enough for senior engineers?

For the top of the funnel, yes: matching, scheduling and consistent first-pass checks. For judging reasoning, we'd keep a human, because the cited study didn't test it. A large field experiment covered customer-service roles in the Philippines, with humans making every hiring decision, and included no engineers.

Can AI detect AI cheating?

Karat says that telling AI-written code from human code by output alone is "nearly impossible," and that the difference shows when you watch the process. That argues for observed sessions over a smarter grader. Karat sells human-led interviews, so weigh it accordingly.

Does Arc use AI to vet developers?

Arc's HireAI page describes HireAI as a GPT-4-powered AI recruiter that produces instant candidate matches. The same page lists a profile screening, a behavioral interview, and a technical interview or pair programming, but doesn't say who conducts them.

Does Turing use human interviews?

The Turing article describes a profile setup, multiple-choice work experience and tech stack tests, a live coding challenge, and AI matching through AIME. It doesn't mention a human interview. That's a statement about one article, not about everything Turing does.

What does engineer-led vetting mean?

It means senior engineers run the vetting stages themselves, including the live session where reasoning gets judged. HighCircl uses four stages run this way, and about 1 in 10 applicants pass.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now