Why do trivia questions fail for AI engineers?
Most AI interview question lists test recall: define attention, explain gradient descent, name three vector databases. Strong candidates can answer those, and so can weak candidates with a search tab open. What predicts success in an applied AI or ML role is judgment: choosing how to measure quality, deciding when a model is good enough, and knowing what to do when it fails in production. The questions below are built to surface that judgment. Each one asks about a decision, and each comes with what a strong answer tends to include.
How should you run the interview?
Pick five or six questions from the groups below and ask them in the same order to every candidate. Score each answer against the same rubric before you discuss the candidate with anyone else. That is what structured interviewing means, and it is one of the most reliable ways to compare candidates fairly. Use follow-up questions to dig into the candidate's own example: what they chose, what they rejected and why.
What questions test evaluation judgment?
- Tell me about an AI feature you shipped. How did you decide it was good enough to launch?
What a strong answer includes: A defined evaluation set, a quality bar agreed with the product owner, and honesty about what the evaluation did not cover. - How would you build an evaluation set for a feature with no labelled data?
What a strong answer includes: Starting small with real user inputs, using human review to label them, and a clear view on when an LLM-as-judge is acceptable and how to check it against human ratings. - Your offline metrics improved but users say the product got worse. What do you do?
What a strong answer includes: Questioning whether the evaluation set reflects real usage, looking at actual failing cases, and treating user feedback as data rather than noise. - When would you not trust an LLM-as-judge score?
What a strong answer includes: When the judge shares blind spots with the model being tested, on subjective or domain-specific tasks, and whenever it has not been calibrated against human labels recently.
What questions test retrieval and system design?
- A retrieval-augmented system gives confident wrong answers. Walk me through how you would debug it.
What a strong answer includes: Separating retrieval failures from generation failures, inspecting the retrieved chunks, and checking chunking, embeddings and ranking before touching the prompt. - When would you fine-tune a model instead of improving prompts or retrieval?
What a strong answer includes: Fine-tuning for consistent format, style or narrow tasks at volume, retrieval for knowledge that changes, and a clear view of the cost of maintaining a fine-tuned model. - How would you design an agent that takes actions in a customer's account?
What a strong answer includes: Limiting the actions available, asking for confirmation before irreversible steps, logging every action, and planning for failure and rollback. - You have two weeks to prove an AI feature is worth pursuing. What do you build?
What a strong answer includes: The thinnest version real users can try, with a way to measure whether it works.
What questions test cost, latency and production trade-offs?
- Your feature works but costs too much per request. What are your options?
What a strong answer includes: Smaller or cheaper models for easier cases, caching, shorter context and routing, with the quality impact of each change measured. - How do you decide between a larger model and a faster one?
What a strong answer includes: Starting from the user's latency tolerance and the quality bar, then testing both on the evaluation set rather than guessing. - What do you monitor after an AI feature goes live?
What a strong answer includes: Quality signals, not just uptime: sampled output review, user feedback, changes in inputs, cost per request, and error and refusal rates. - A model provider updates a model and your outputs shift. How do you protect against that?
What a strong answer includes: Pinned model versions where possible, regression evaluations before switching, and alerts when outputs change.
What questions test ML engineering depth?
- Tell me about a model that performed well in training and badly in production. What happened?
What a strong answer includes: A specific cause, such as data leakage, distribution shift or a mismatch between offline and online features. - How do you decide when to retrain a model?
What a strong answer includes: Monitoring for drift and performance decline, a defined trigger, and the cost of retraining weighed against its impact. - How would you handle a heavily imbalanced dataset for a classification problem?
What a strong answer includes: Choosing metrics that reflect the imbalance before choosing techniques, and tying the choice to the business cost of each type of error. - What does a reproducible training pipeline look like to you?
What a strong answer includes: Versioned data, code and configuration, tracked experiments, and the ability to rebuild a past model.
What questions test ownership and working style?
- Tell me about a time you pushed back on a product request because of an AI limitation.
What a strong answer includes: Explaining the limitation in plain language, offering an alternative, and following through. - Describe a production incident you owned. What did you change afterwards?
What a strong answer includes: Specifics, ownership of their part, and a lasting fix rather than a patch. - How do you explain model uncertainty to a non-technical stakeholder?
What a strong answer includes: Concrete examples instead of jargon, and expectations set before launch. - What is something in AI engineering you changed your mind about in the last year?
What a strong answer includes: A real update backed by experience. It shows they keep learning in a field that moves quickly.
What questions test safety and responsible use?
- How would you stop an LLM feature from leaking data between customers?
What a strong answer includes: Keeping each customer's data separate at retrieval, enforcing access checks outside the model, and testing for prompt injection. - What is prompt injection, and how have you defended against it?
What a strong answer includes: Treating model output and retrieved content as untrusted, limiting which tools the model can call, and having tested attacks themselves. - If you built an AI tool that screens job applicants, what would you check before launch?
What a strong answer includes: Bias testing across groups, a human reviewing every decision, and candidate notice and local rules such as New York City's Local Law 144. - When should a human stay in the loop on an AI decision?
What a strong answer includes: When errors are costly or hard to reverse, when decisions affect people's opportunities, and when the system cannot explain itself well enough to audit.
How do you score the answers?
Use one simple rubric for every question. Score 1 for a generic or textbook answer, 2 for a real example with limited reasoning, 3 for a real example with clear trade-offs, and 4 for a real example with trade-offs, a measured outcome and what they would do differently next time. Write down the evidence behind each score. Compare candidates on that evidence, not on how confident they sounded.
Which questions fit which role?
For ML engineers lean on the evaluation and ML depth groups. For applied AI and LLM engineers lean on evaluation, retrieval and production trade-offs. For founding engineers add the ownership questions and the two-week build question. For AI sales and solutions engineers adapt the cost and safety questions and ask the candidate to explain those trade-offs to a buyer. For the wider process, read our guide on how to hire AI engineers.