AI Engineer interview questions and how to answer them

AI engineer interviews test two things at once: can you write solid production code, and do you know how model features fail. Most loops spend more time on the second than candidates expect. Come ready to talk about a system you built, how you knew it worked, and what you'd do differently.

The process

What happens in each round

  1. 1

    Recruiter screen

    What happens

    Whether you've shipped a model feature to real users or only built demos, and whether your background fits a product team or a research team.

  2. 2

    Coding round

    What happens

    Plain Python and data handling. Expect parsing, batching API calls, retries with backoff, or cleaning a messy text dataset. They want readable code and sensible error handling, not clever tricks.

  3. 3

    System design for an LLM feature

    What happens

    How you'd build something like a support assistant or document search end to end: retrieval, prompting, caching, evaluation, monitoring and cost. They listen for tradeoffs you name without being prompted.

  4. 4

    Project walkthrough

    What happens

    Depth on one thing you built. They'll keep asking why until they hit the edge of what you actually know.

  5. 5

    Behavioral and product sense

    What happens

    How you handle ambiguity, push back on a launch that isn't ready, and explain model limits to people who aren't engineers.

Questions you're likely to get

1.Walk me through how you'd build a question-answering assistant over our internal docs.

Why they ask

It's the most common thing AI engineers get asked to build, and it touches every part of the job in one answer.

How to answer

  • Start with the user and the failure that hurts most, such as a confident wrong answer about policy
  • Cover ingestion and chunking, and say why chunk size depends on how the docs are written
  • Explain retrieval choices: embeddings, a vector store like pgvector, and hybrid keyword search for names and codes
  • Describe the prompt, citations back to source passages, and what the assistant says when retrieval comes back empty
  • Close with an evaluation set built from real questions and how you'd monitor it after launch
2.How do you evaluate a feature whose output is free text?

Why they ask

Evaluation is where demos and products split. Interviewers use this question to sort people who've shipped from people who've prototyped.

How to answer

  • Build a graded example set from real or realistic inputs, including the ugly ones
  • Use exact checks where you can, like whether the right document was retrieved or the JSON parses
  • Use a model as a grader for fuzzier qualities, and check the grader against human labels first
  • Run the suite on every prompt or model change, the way you'd run unit tests
3.Your assistant is making up answers. How do you track down why?

Why they ask

Hallucination complaints land on the AI engineer's desk constantly. They want a debugging order, not a buzzword.

How to answer

  • Pull the failing conversations and check whether the right source was retrieved at all
  • If retrieval missed, look at chunking, embeddings and query rewriting before touching the prompt
  • If retrieval hit, tighten the prompt to answer only from context and to say when it doesn't know
  • Add the failures to the eval set so the fix stays fixed
4.When would you fine-tune a model instead of improving the prompt or retrieval?

Why they ask

Fine-tuning is expensive and often unnecessary. They want to see judgment, not enthusiasm.

How to answer

  • Default to prompting and retrieval because they're cheaper to change and easier to debug
  • Fine-tune for consistent format or tone, a narrow task at high volume, or to move to a smaller, cheaper model
  • Mention parameter-efficient methods like LoRA and the need for clean labeled data
  • Say how you'd compare the tuned model against the baseline on the same eval set
5.Our model bill is too high. What would you look at first?

Why they ask

Cost surprises are common with model features, and the engineer who can cut them without hurting quality is worth a lot.

How to answer

  • Break the spend down by feature and by step, since one chatty agent loop is often most of it
  • Trim context: fewer retrieved chunks, shorter system prompts, summarized history
  • Cache repeated requests and route easy cases to a smaller model
  • Prove each change holds quality on the eval set before shipping it
6.How would you keep response times low for a chat feature?

Why they ask

Users leave slow chat boxes. They want to know you think about latency as part of the design.

How to answer

  • Stream tokens so the user sees text right away
  • Run retrieval and other lookups in parallel instead of in sequence
  • Keep prompts short and pick the smallest model that passes your evals
  • Track the slow tail in your dashboards, not only the average
7.How do you protect a model feature from prompt injection or leaking data?

Why they ask

Any feature that reads user input or outside documents can be tricked. Security teams will ask you this before launch anyway.

How to answer

  • Treat model output as untrusted input, especially before it calls a tool or touches a database
  • Limit what tools the model can call and require confirmation for anything destructive
  • Enforce access control at retrieval time so users only see documents they're allowed to see
  • Test with known injection examples and log suspicious requests
8.Tell me about an agent or tool-calling system you've built. What went wrong?

Why they ask

Agents fail in loops, wrong tool picks and runaway costs. They want to hear that you've been burned and learned from it.

How to answer

  • Describe the task and why it needed tools instead of a single prompt
  • Name a specific failure, like the agent looping on a search or passing a malformed argument
  • Explain the fix: step limits, stricter schemas, fewer tools, or a fixed workflow for the common path
  • Say what you'd do differently from the start
9.A provider is retiring the model version you depend on. How do you handle the switch?

Why they ask

This happens often, and it tests whether your system is built to survive change.

How to answer

  • Keep the model name in config, behind a thin wrapper, so switching is a small change
  • Run the full eval set against the new model and compare side by side
  • Retune prompts where scores drop, then roll out to a small slice of traffic first
  • Watch user feedback and error rates before moving everyone over
10.Tell me about a time you told a product manager a model feature wasn't ready.

Why they ask

Pressure to launch is constant in this role. They want someone who can say no with evidence and still keep the relationship.

How to answer

  • Set up what was being pushed and why the pressure was real
  • Show the evidence you brought, like eval results or examples of harmful answers
  • Offer the alternative you proposed, such as a narrower launch or a human review step
  • Share how it ended and what you'd repeat
11.What's a recent change in model tooling you tried, and did you keep it?

Why they ask

The tools move fast. They want to know you experiment with judgment and don't chase every release.

How to answer

  • Name something concrete you actually used, like a new structured output feature or an eval framework
  • Say what problem you hoped it would solve
  • Explain how you tested it and whether it earned a place in your stack

Mistakes that sink good candidates

Talking only about prompts and never about how you measured results

Claiming experience with every framework on the posting when you've used one

Describing a demo as if it ran in production with real users

Dismissing cost, latency or security as someone else's problem

Need more AI Engineer interviews to prep for?

HeroApply applies to AI Engineer jobs that match you, every day. 4,114 jobs are open today.

Find AI Engineer jobs