By Daniel Whitmore, MSc in Industrial and Organizational Psychology
In short
In hiring, AI should operate as decision support: it can draft tasks and rubrics, summarize responses and interview transcripts, and propose criterion-level scores — but a named person confirms or overrides every score, and every decision belongs to a human who can explain it. Each AI-assisted output should carry provenance (which model, which version), so the role AI played in a decision stays auditable after the fact.
"The AI picked the candidate" is two failures in one sentence
The first failure is accountability. A hiring decision affects a person's livelihood, and someone must be able to own it — to explain it to the candidate, defend it to a stakeholder, and answer for it if it is challenged. "The model scored them lower" is not an explanation; it is an abdication. When no named human stands behind a rejection, the organization has not automated a decision. It has orphaned one.
The second failure is quality, and it is less discussed. A model reading a support candidate's reply to an angry customer can assess structure and tone competently, but it does not know that your team is drowning in refund escalations this quarter, that the role backfills your best de-escalator, or that the hiring manager can coach writing but not judgement. The decision-relevant context lives with people. A model that decides alone decides with less information than the panel had.
So the popular framing — human oversight as a safety brake bolted onto an AI decision — gets it backwards. The human is not reviewing the AI's decision; the AI is preparing the human's. The inversion matters in practice, too. A reviewer who rubber-stamps machine output drifts toward the machine's blind spots, while an evaluator who owns the call uses the machine to see more, not to think less.
What good AI assistance actually looks like
Rejecting AI-as-decider is not rejecting AI — that would waste the most reviewer-friendly technology assessment has seen. Work-sample evaluation produces exactly the kind of evidence AI is genuinely good with: long written responses, multi-step task outputs, whole sessions of activity that take real time to read cold. Used as an assistant, it makes human evaluation faster and more consistent without displacing the judgement it exists to serve.
The test for every feature is the same: does it prepare the evaluator's judgement, or preempt it? Assistance that passes hands a human better material and leaves the conclusion open; assistance that fails quietly narrows the human's role to a yes-or-no on the machine's opinion. In a customer-support or inside-sales hire, the assistance that passes looks like this:
- Summaries: condense a candidate's full ticket-queue exercise or discovery-call write-up into a brief an evaluator can verify against the raw output in one click.
- Evidence pointers: highlight where a reply acknowledged the customer, where a policy call was made, where an objection was actually answered versus deflected.
- Draft criterion scores: propose a rubric-anchored score per criterion, with the passages that justify it — a starting point, never a verdict.
- Consistency flags: notice when an evaluator's score diverges sharply from the rubric descriptor they cited, and surface it for calibration.
- Integrity signals: flag session anomalies for proportional human review — never an automatic rejection.
| AI drafts and surfaces | A named human decides |
|---|---|
| Task, question, and rubric drafts | What actually measures the role |
| Summaries of long responses and transcripts | How to weigh the evidence |
| Criterion-level score suggestions | The confirmed score — and any override |
| Evidence pointers and integrity flags | Whether a candidate advances |
Confirm or override: the checkpoint that keeps humans in charge
The design detail that separates decision support from decision laundering is the explicit checkpoint. Every AI-drafted score in SkillCort sits in a pending state until a named evaluator confirms it or overrides it — and an override requires nothing more than the evaluator's judgement. No justification tax, no friction that quietly punishes disagreeing with the machine. If overriding is harder than accepting, you have built an AI decision-maker with extra steps.
Overrides are also signal, not noise. When evaluators consistently override the draft score on a judgement-within-policy criterion, that says the rubric descriptor is ambiguous or the task needs redesign — feedback a silent rubber-stamp workflow would never surface. The checkpoint keeps humans in charge and, as a side effect, keeps the assistance honest about where it helps and where it does not.
This is why we insist every confirmation carries a name. "An evaluator approved it" is accountability; "the system processed it" is not. The unit of decision-making in hiring is a person with a rationale, and the tooling should make that person visible rather than abstract them away. When a candidate asks why they were not advanced, the answer should begin with who decided — and the evidence file should be able to say.
Provenance: every assisted output carries its receipts
If AI touches an evaluation, the record must say so. Every AI-assisted output in a candidate's evidence file carries provenance: which model produced it, which version, when it ran, what it was asked to do, and who confirmed or overrode the result. Not as a compliance ornament — as the thing that makes an assisted decision auditable at all.
Provenance answers the questions that otherwise become unanswerable months later. Was this summary generated before or after the rubric changed? Did the draft scores for these two candidates come from the same model version, so their starting points are comparable? Which human turned this suggestion into a score? Without logging, "AI-assisted" degrades into "AI-decided, unverifiably." With it, an auditor, a candidate, or your own team can reconstruct exactly where the machine's contribution ended and the human's judgement began.
Regulation is converging on the same line
This is not only our philosophy; it is the direction of the law. The EU AI Act treats AI systems used in hiring and employment decisions as high-risk, with obligations that point squarely at human oversight, transparency, and record-keeping — the same triad as confirm-or-override checkpoints, labeled assistance, and provenance logs. Regulators elsewhere are scrutinizing automated employment decisions along similar lines. The precise obligations vary; the direction does not.
Teams that build the human-decision discipline now are not making a compliance bet — they are building the process regulators are describing, because it was the right process anyway. A vendor whose pitch is "the AI ranks and rejects for you" is selling you the exact architecture the regulatory trajectory runs against. Decision support with named accountability is not the cautious option. It is the durable one.
Our position, stated plainly
AI in assessment should make human judgement faster, more consistent, and better evidenced — and should never replace it. In practice that means: AI summarizes and points at evidence; AI drafts criterion-level scores against a rubric humans wrote; a named person confirms or overrides every score; every assisted output logs its model, version, and reviewer; integrity signals go to human review, never auto-rejection; and no candidate ever carries a hidden AI score across employers.
This is what work-sample-first assessment demands. The whole point of asking a candidate to do the actual job — answer the angry customer, work the pipeline — is that a person who understands the role reads the work and decides. Anything that removes the person removes the point. Score to story, machine to margin, decision to a name: that is the standard we build to, and the one we think the industry ends up at.
Key takeaways
- "The AI picked the candidate" fails twice: no one owns the decision, and the decision was made with less context than the humans had.
- Good AI assistance is concrete: summaries, evidence pointers, rubric-anchored draft scores, and calibration flags — never selection or rejection.
- Every AI-drafted score waits for a named evaluator to confirm or override it, and overriding must be as easy as accepting.
- Provenance logging — model, version, timestamp, task, reviewer — is what makes AI-assisted evaluation auditable instead of unverifiable.
- The EU AI Act treats hiring as high-risk, and its emphasis on oversight, transparency, and records points the same direction as human-decided assessment.
