What you'll need
- A role where spoken communication or verbal reasoning genuinely matters
- One well-formed opening question you would ask every candidate anyway
- Evaluation criteria for what a strong spoken answer contains
- A named evaluator who will confirm or override every AI score suggestion
Step 1: Understand the turn-based format before you build
The interview is a sequence of turns, not a free-flowing call. The candidate hears a question and speaks their answer — up to around three minutes per turn. Speech-to-text produces a transcript that is stored as evidence, and from that transcript an adaptive follow-up question is generated and spoken back to the candidate, with subtitles on screen.
This structure is the point: every candidate starts from the same opening question, follow-ups stay anchored to what the candidate actually said, and everything is captured as reviewable text. It behaves like a structured interview that never gets tired, runs at whatever hour suits the candidate, and leaves a complete written record.
Step 2: Write the opening question
The opening question sets the ceiling for the whole interview, because adaptive follow-ups can only dig into what the candidate offers. Ask for a specific, recent, first-person account: "Tell me about a time you had to deliver bad news to a customer — what was the situation and what did you do?" beats "How do you handle difficult conversations?" every time.
Aim for a question a strong candidate can answer substantively in two to three minutes of speech. Avoid multi-part questions — the follow-up mechanism handles the parts you would have bundled in. If you have several independent things to probe, they are separate opening questions, not one long one.
- Ask for a concrete situation, not a philosophy
- One question, one topic — follow-ups handle depth
- Answerable in two to three minutes of natural speech
- Neutral wording that does not telegraph the ideal answer
Step 3: Define the evaluation criteria
Attach a rubric with the criteria a strong spoken answer should show — for example structure, ownership of the outcome, concrete detail, and communication clarity. Use anchored level descriptions so evaluators know what "strong" concretely sounds like in a transcript. As with any SkillCort rubric, criterion weights are explicit and the rubric freezes into a snapshot when attached.
Write criteria that can be evidenced from a transcript, since the transcript is what gets scored. "Explains their own role rather than the team's" is checkable in text; "seems confident" is not, and criteria that lean on vocal impressions rather than content invite bias against accents and speech styles.
Step 4: Configure the typing fallback and accessibility
A typing fallback is always available: candidates can type their answers instead of speaking, and no microphone is ever required to complete the interview. This is not a degraded mode — typed answers flow into the same evidence record and the same rubric — so a candidate with no quiet room, no working microphone, or a speech difference is not filtered out by their equipment.
Make the fallback visible in your candidate instructions rather than leaving it to be discovered. Questions are spoken back with subtitles, so candidates who process text better than audio are covered on the listening side too. Score typed and spoken transcripts against the same criteria and be conscious that fluency of delivery was never one of them.
Step 5: Add the interview to a delivery and invite candidates
Place the voice interview in your assessment and publish once the checklist clears, then create a delivery with a name and time window and send email invitations. Keep integrity settings light — the interview format needs microphone consent, not lockdown, and heavy proctoring on a conversational exercise reads as hostile.
The consent and preflight gate tells candidates exactly what happens: their speech is transcribed, the transcript is stored as evidence, follow-up questions are AI-generated, and a typing fallback exists. The preflight microphone check happens before the first question, so equipment problems surface as setup fixes, not lost attempts.
Step 6: Review transcripts and confirm every score
As interviews complete, each turn's transcript sits in the candidate's evidence record. The AI drafts criterion-level score suggestions from the transcripts, with the model and version logged. A named human evaluator reads the transcript, then confirms or overrides each suggestion — the confirmed score is the score, and the evaluator's name is on it.
Read the transcript before looking at the suggestions when you can; it keeps your own judgment primary. Where you disagree with a suggestion, override it and record why. Those overrides are useful signal about where the criteria need sharper anchors, and they are exactly the human-in-the-loop record that makes the process defensible.
Step 7: Decide where voice AI interviews fit — and where they do not
The format shines for asynchronous screening at scale: fifty candidates each get the same structured opening, adaptive depth, and a scored transcript, without fifty calendar slots. It works as a consistent first conversation before humans spend live time, and pairs naturally with work-sample tasks in the same assessment.
It is the wrong tool for final-round judgment, for roles where spoken interaction is peripheral, or as a substitute for ever speaking to the person you hire. It screens; it does not select. The Decision Board shows its scores alongside all other evidence, and a human weighs the whole picture — no candidate is rejected by a model.
Pro tips
- Record yourself answering your own opening question — if you cannot give a substantive answer in three minutes, candidates cannot either.
- Two or three opening questions with adaptive follow-ups beat six shallow ones; depth is what the format buys you.
- Read a few full transcripts before trusting any score suggestion pattern — it calibrates you to what the AI sees and misses.
- State the typing fallback in the invitation email, not just at the gate; some candidates will not start if they think a microphone is mandatory.
- Do not score delivery fluency unless the job genuinely requires spoken polish — transcripts are for judging content.