By Daniel Whitmore, MSc in Industrial and Organizational Psychology
In short
An AI interview is a structured conversation an AI conducts on the employer's behalf: it asks role-relevant questions, adapts follow-ups to probe the candidate's answers, and captures the exchange as a transcript. In a defensible setup, that transcript is the output — evidence a named evaluator scores against a rubric — not a black-box "fit" number. AI interviews add the most value for structured verbal reasoning, communication, and situational judgement; they do not replace a hands-on work sample, and they should never infer personality or emotion from a candidate's face or tone of voice.
What an AI interview actually is (and what it isn't)
Strip away the marketing and an AI interview is a conversation with structure. An AI system asks a candidate a set of role-relevant questions, listens to or reads the answers, and asks follow-ups to go deeper — the way a good interviewer would when a reply is thin or a claim needs a concrete example. The candidate can respond by voice or in writing. What comes out the other end is a record of what was actually said: a transcript, tied to the questions that prompted it.
That record is the point. In a well-built AI interview, the transcript is the deliverable — evidence a person reads and scores against a rubric. The AI ran the conversation so a human didn't have to schedule and sit through every first-round call; it did not decide anything.
Contrast that with the version that earned AI interviews their bad reputation: systems that analyze facial expressions, micro-gestures, or vocal tone and infer personality traits or a "culture fit" score. That approach has weak scientific support and a large fairness problem — it can penalize an accent, a flat affect, a poor webcam, or a disability that has nothing to do with the job. The useful kind of AI interview produces evidence a human can check. The harmful kind produces a number no one can defend. Everything below assumes the first kind.
| Evidence-gathering AI interview | Black-box AI "analysis" |
|---|---|
| Asks structured, job-relevant questions | Scores face, tone, or word choice |
| Output is a transcript a human reads | Output is an opaque trait/fit score |
| A named evaluator scores against a rubric | The tool ranks or screens automatically |
| Auditable: you can see what was said and why it scored | Unauditable: the reasoning is hidden |
How the conversation works: structure first, follow-ups within bounds
The single biggest driver of whether an interview predicts job performance is structure — the same core questions, tied to the competencies the role needs, scored the same way for everyone. Decades of selection research point the same direction: structured interviews substantially outperform unstructured ones, and the structure is what does the work. An AI interviewer's real advantage is that it applies that structure consistently. It doesn't get tired by the tenth candidate, doesn't drift into rapport-based small talk, and doesn't ask the strong-résumé candidate easier questions.
But structure alone can feel like a form. The reason to use AI rather than a static questionnaire is the adaptive follow-up: when a candidate gives a vague answer to a situational question, the AI asks for the specific action they took; when they mention a decision, it asks what they weighed. Done well, this probes the same competencies more deeply without turning into a different interview for each person.
The design risk is a follow-up that wanders — off the role, into personal territory, or into a rabbit hole that advantages whoever is chattier. So the follow-ups have to be bounded: on SkillCort, an AI interview's follow-ups stay inside the role's competency map. The system digs deeper on what the role actually needs and doesn't improvise its way into questions you'd never approve a human asking. Structured spine, adaptive depth, fixed boundaries.
How AI interviews are scored — and how they should be
Here is where most of the trust is won or lost. In a defensible AI interview, scoring runs off the transcript against a rubric a human wrote: each competency has descriptors, and each answer is judged against them. The AI can help — summarize a long transcript, point to the passage where a candidate acknowledged the customer or weighed a trade-off, and even draft a rubric-anchored score per competency. What it cannot do is decide. Every criterion score waits in a pending state until a named evaluator confirms or overrides it, and overriding is as easy as confirming.
This matters more for interviews than for almost anything else, because a conversation is exactly the kind of rich, ambiguous evidence where a model reads competently but not completely. It can miss that a candidate's "unpolished" answer described precisely the right judgement, or that a fluent answer said nothing. The human reading the transcript catches that — and owns the call.
The anti-pattern is the hidden "interview score": a single number the tool assigns and passes downstream, with the transcript buried or discarded. If you cannot open a candidate's interview, read what they said, and see which passage earned which criterion score, you do not have an assessment — you have a verdict with no evidence behind it. And no candidate should carry an AI interview score from one employer to the next; each hiring decision is its own, judged on its own evidence.
What an AI interview can and can't measure
AI interviews are strong where the evidence is verbal and structured. They are a good fit for situational judgement ("a customer says X — what do you do and why?"), communication quality, reasoning you can hear a candidate work through, and role knowledge probed with follow-ups. For high-volume, early-stage screening — customer support, contact centre, inside sales — a structured AI interview can gather far more comparable evidence than a rushed human phone screen, without the scheduling bottleneck.
They are weak, or simply the wrong tool, elsewhere. An interview asks a candidate to describe or reason about work; it does not ask them to do it. For anything hands-on — writing the actual reply to the angry customer, working the ticket, building the thing — a work sample beats any interview, because it measures the real output rather than a candidate's account of it. And no AI interview should try to read personality, emotion, or "fit" from a face or a voice; that is the discredited approach from the first section, not a bonus feature.
The practical answer is rarely "interview or work sample." It's both, in sequence: a structured AI interview to gather broad, comparable evidence efficiently, then a work sample for the finalists to see them do the job. The interview tells you how someone thinks and communicates; the work sample tells you what they can actually produce.
Fairness, bias, and the law
AI does not remove bias from interviewing; it changes where bias can enter and how visible it is. The real risks in an evidence-gathering AI interview are concrete: speech-to-text that transcribes some accents less accurately than others, questions that aren't genuinely job-relevant, and — worst of all — automated scoring that screens people out with no human reading the evidence. Each has a mitigation: job-relevant structured questions, transcripts a human reviews, and a confirm-or-override checkpoint on every score so no one is rejected by a machine.
Regulation is converging on exactly this discipline. New York City's Local Law 144 requires an independent bias audit of automated employment decision tools and candidate notice before use. Illinois's Artificial Intelligence Video Interview Act requires notifying candidates, explaining how the AI works, obtaining consent, and deleting recordings on request. The EU AI Act classifies AI used in hiring as high-risk, with obligations around human oversight, transparency, and record-keeping. The specifics differ by jurisdiction and change over time — treat this as orientation, not legal advice, and check the rules that apply to you — but the direction is unmistakable: notice, human oversight, auditability.
The reassuring part is that the compliant design and the effective design are the same design. Job-relevant structured questions, a transcript you can show, human-owned scoring, and an audit trail are what makes an AI interview both fair to defend and good at predicting performance. A tool whose pitch is "our AI ranks and screens candidates for you" is selling you the architecture the law is moving against — and the one least likely to survive a candidate's, or a regulator's, question of "why."
A checklist for a defensible AI interview
If you are introducing AI interviews, or auditing one you already run, these are the questions that separate an evidence-gathering interview from a black box. If you can answer yes to all of them, you have a process you can stand behind.
- Structured spine: does every candidate get the same core, job-relevant questions, scored against the same rubric?
- Bounded follow-ups: do adaptive probes stay inside the role's competencies rather than wandering off-role?
- Transcript as evidence: can you open any interview and read exactly what was said?
- Human-owned scoring: does a named evaluator confirm or override every criterion score, with overriding as easy as accepting?
- No face/voice inference: is the system scoring answers, never facial expression, emotion, or "fit"?
- Candidate transparency: are candidates told an AI is interviewing them, how it works, and how their data is handled?
- Paired with a work sample: do finalists also do the actual job, not just describe it?
- Audited for impact: are you checking outcomes for adverse impact across groups, and fixing questions that don't hold up?
Key takeaways
- An AI interview should gather evidence — structured questions, adaptive follow-ups, a transcript a human scores — not analyze a candidate's face or voice for a hidden "fit" score.
- Structure is what makes an interview predictive; an AI interviewer's real value is applying that structure consistently, with follow-ups bounded to the role.
- Scoring must run off the transcript against a human-written rubric, with a named evaluator confirming or overriding every criterion score.
- AI interviews measure verbal reasoning and communication well but can't replace a hands-on work sample — use both, in sequence.
- Fair-by-design and lawful-by-design are the same design: job-relevant questions, a readable transcript, human-owned decisions, and an audit trail (NYC LL144, Illinois AIVIA, EU AI Act).
Sources & further reading
- Levashina, Hartwell, Morgeson & Campion (2014), "The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature," Personnel Psychology
- McDaniel, Whetzel, Schmidt & Maurer (1994), "The Validity of Employment Interviews: A Comprehensive Review and Meta-Analysis," Journal of Applied Psychology
- Regulation (EU) 2024/1689 (EU AI Act) — Official Journal of the European Union
- NYC Local Law 144 — Automated Employment Decision Tools (NYC DCWP)
