You can hire a strong backend engineer and still ship an AI feature that quietly gets worse every week. Nothing crashes. The answers just get less right, nobody has a test that says so, and the first person to notice is a customer. That's the real hiring risk.
It covers how to hire AI engineers, which in practice means the same person as an LLM developer: an engineer who builds features on top of language models, with retrieval, evals, cost control and fallbacks. It doesn't cover candidates who use Copilot or Cursor to write code faster. It sits in our series Hire developers by tech stack: rates, vetting and interview guides. There's no salary table: we couldn't find a primary, dated source for European AI-engineer pay.
What is an AI engineer, and how is it different from an ML engineer?
"AI engineer" gets used as a catch-all, which is how companies end up interviewing researchers for a product job. There are three different people behind the title.
The AI or LLM application engineer builds product features on top of models that someone else trained. Their work is prompts, retrieval, tool calls, evaluation, latency and cost. The ML engineer trains and fine-tunes models and runs the training infrastructure. The data engineer builds the pipelines that feed both, and for a retrieval system that means ingestion, cleaning and keeping the index fresh.
Most companies shipping an LLM feature need the first profile. Andrej Karpathy, quoted in Latent Space's 2023 essay on the role, put it this way: "One can be quite successful in this role without ever training anything." The essay itself says that for shipping AI products, you want engineers, not researchers. That's the filter for your job spec: strong software engineering first, model knowledge second.
Python is the language most of this work is written in, but the language isn't the job. A great Python developer who has never had to measure whether an answer got worse after a model upgrade is a backend hire, not an AI hire.
What do you need the first hire to own?
Don't write the spec from a skills list. Write it from the failure you already have, or the one you're most afraid of. Three cover most teams.
Prototype-to-production gap
The demo works. Production doesn't: timeouts, malformed outputs, a provider outage, a prompt that behaves differently on real user input. You need someone who has been paged for an LLM feature and can describe what they changed afterwards.
Retrieval and data
Answers sound confident and are wrong because the wrong documents were retrieved, or the index is stale, or chunking split a table in half. This hire needs to think about ingestion and freshness as much as about the model. If your data pipeline is also a mess, you may need a data engineer first, and our guide to hiring data engineers covers that profile.
Evals and quality
Nobody can say whether the feature got better or worse after the last change. Here you need someone who builds the measurement before the fix.
What to put in the job description
Write requirements you can test in an interview. Model API experience in production, not a side project. Retrieval design, including how they'd debug it. An evaluation approach. Observability for prompts and outputs. Latency handling, retries and fallbacks. Per-request cost awareness.
Leave out "passionate about AI" and any list of framework names. Frameworks change every quarter, and a candidate who can explain why they chose one is worth more than one who lists five.
How to vet an AI engineer
Each step below is a short exercise you can run in 30 to 60 minutes.
1. Give a small eval-writing exercise
Hand the candidate about 20 example inputs from your domain and a feature that's failing on some of them. Ask them to design the test set and the grading.
OpenAI's docs say that writing evals to see how your LLM application performs, especially when upgrading or trying new models, is an essential part of building reliable applications. Anthropic's testing guidance adds the design rules worth scoring against: mirror your real task distribution and include edge cases, and structure the questions so grading can be automated, by string match, code or another model. The same page notes that most use cases need evaluation along several criteria, and latency and cost are among them.
Score the candidate on three things. Are the cases specific to your task rather than generic? Did they include the ugly edge cases? Can the grading run without a human reading every output? A candidate who starts by asking what "good" means for your users is already ahead.
2. Review a broken RAG pipeline
Show a pipeline where the retrieved chunks are wrong but the final answers sound right. Ask what they'd check first.
You want to see them separate retrieval from generation. Good candidates look at what was retrieved before touching the prompt. They ask about chunking, embeddings, metadata filters and freshness. Weak candidates rewrite the prompt, because prompts are the only lever they've ever pulled.
3. Ask about latency, retries and fallbacks
Ask how they'd keep a feature usable when the model is slow or down. In a Latent Space interview, Elicit's head of engineering James Brady says they often see a 10x variation in P90 latency within half an hour to an hour when prompting models. A candidate who's shipped one knows that already.
Listen for timeouts, retries with backoff, a fallback model or a degraded non-AI path, streaming to hide wait time, and how they'd know it was happening (logs, traces, alerts). "We'd just retry" isn't an answer.
4. Ask them to cut the cost of a request path
Give them a prompt that carries a large, mostly unchanged context and gets called thousands of times a day. Ask how they'd reduce the cost per request without hurting quality.
Good answers name levers: prompt caching, batching work that isn't latency-sensitive, a smaller model for easy cases, shorter prompts, caching whole responses. Anthropic's docs give the shape of the savings. As of 2026-09-30, prompt cache reads cost 0.1 times the base input price for most models, and batch processing cuts costs by 50% for work that can wait. Prices change, so don't test whether they memorised a number. Test whether they know the levers exist and would measure the effect.
A candidate who can't say how they'd make a request cheaper is missing a skill your finance team will eventually notice.
5. Run the real-work interview
Skip the algorithm puzzle. Give them a small feature or bug in a real codebase and watch how they work. The Latent Space interview above describes hiring for this role that way, on tasks inside a real code base. You're checking whether they can read unfamiliar code, run it, and reason about a failure, which is most of the job. If you're setting up this stage, our piece on running a live coding interview covers the format.
Interview questions that reveal production experience
Each question has a tell in the answer.
- Tell me about an LLM feature that got worse after a change. How did you find out? A good answer names a test set or monitoring. A weak one says a user complained.
- How do you decide a new model is safe to switch to? Look for a fixed eval set, comparison on your own cases, and a rollout plan.
- A RAG answer is wrong. Walk me through debugging it. Retrieval first, then context assembly, then generation.
- What do you log for each model call? Inputs, outputs, model version, latency, token counts, and retrieved sources.
- How do you handle output that isn't valid JSON? Validation, constrained output where available, retry with the error, and a fallback.
- How would you stop a prompt injection in retrieved content? Treating retrieved text as untrusted, limiting tool permissions, and not trusting the model as the security boundary.
- When would you not use an LLM for this? Seniors have an answer. Rules, classic search or a small classifier are often better and cheaper.
- What's the cost per request on your last project, and what moved it? They don't need the exact figure, but they should know it existed.
Red flags
Only demo experience: everything they describe ended at "it worked in the notebook". No eval story, or "we checked it by eye". They can't explain a failure they personally debugged, only what the framework docs say. They treat the prompt as the whole product. They talk about the model and never about the data, the fallbacks or the bill.
Where do AI engineers come from, and how should you hire?
There are three routes: a direct hire, a freelancer, or an agency or staff augmentation partner. None is wrong. The vetting above works the same on all three, and any partner that can't tell you how it screens for evals and production experience is screening for keywords.
A direct hire fits when the AI feature is core and you'll be iterating on it for years. Freelancers suit a bounded piece of work with a clear finish line, such as an eval suite or a cost audit, as long as you can review their output yourself.
When staff augmentation fits an AI feature
Staff augmentation fits spike work: you need an engineer for a quarter to get a feature from prototype to production, and you don't want a permanent seat before you know the shape of the work. It fits badly when nobody in-house can judge the work, because then there's no one to run the exercises above.
Data residency and GDPR when prompts contain customer data
Before any engineer touches customer data, get answers to three questions. Does the model provider's data processing agreement cover your use, and can you sign it? What does the provider retain from prompts and outputs, for how long, and is it used for training? Where do your retrieval data and embeddings live, and who can access them?
A good AI engineer asks these before you do. Ask candidates how they'd keep personal data out of prompts, logs and vector stores unless it's needed, and how they'd delete it on request. Someone who has never thought about it will have a long pause.
What HighCircl offers
HighCircl puts a vetted shortlist of 3-5 engineers in front of you within 72 hours. Four stages, all run by senior engineers, sit behind it: background and experience verification, a communication and product-thinking assessment, a take-home project that mirrors real work, and a live session on architectural reasoning. Roughly 1 in 10 applicants pass. Engineers work from seven European countries. There's a 20% margin on top of what the engineer earns, capped, with no minimum hour commitment, no subscription and no recruitment fee.
FAQ
How is an AI engineer different from an ML engineer?
An AI engineer builds product features on top of existing models: retrieval, evals, latency, cost and fallbacks. An ML engineer trains and fine-tunes models. Andrej Karpathy, quoted in Latent Space's 2023 essay on the role, says you can succeed as an AI engineer without ever training anything. The essay adds that shipping AI products calls for engineers, not researchers.
What should I test in an AI engineer interview?
Test the things that break in production: writing an eval for a failing feature, debugging a RAG pipeline that retrieves the wrong documents, handling slow or failing model calls, and cutting the cost of a request path. Then run a small task in a real codebase. Trivia about model architectures tells you little about whether they can ship.
How long does it take to hire an AI engineer?
It depends on seniority, location and how specific your spec is, and we haven't found a primary source for a reliable figure, so we won't quote one. What shortens it is a spec built around one failure mode you need fixed, and an interview loop you've written before the first candidate arrives.
Do I need a PhD?
Usually not. For engineers who build on top of models, software engineering skill and curiosity about how models behave matter more than research credentials. You'd want a PhD-level hire for training new models, which is a different role.
Should I hire freelance, agency, or full-time?
Full-time when the AI feature is core and long-lived, freelance for a bounded project you can review yourself, and an agency or staff augmentation partner when you need a senior engineer for a quarter or two. Whichever route you take, run the same eval, RAG and cost exercises on the person who'll do the work.
