By Daniel Whitmore, MSc in Industrial and Organizational Psychology
In short
A hiring decision file is a single record that pairs every candidate's score with the evidence that produced it: the work-sample outputs, the rubric scores each evaluator gave, and an audit trail of who decided what and why. Unlike a standalone test score, it can be re-opened and defended months later — to a hiring manager, a rejected candidate, or an auditor. Teams build one by scoring real work against shared rubrics and recording the rationale for every advance.
The problem with a single number
A single score is seductive because it makes comparison feel easy: 82 beats 76, done. But the number hides everything that matters. Which competencies drove it? Was the 76 a candidate who wrote a brilliant reply to an angry customer but stumbled on a queue-triage item — exactly the trade-off your team lead would want to weigh? A leaderboard flattens judgment into arithmetic, and arithmetic cannot answer the question you will eventually be asked: why this person and not that one.
The failure shows up at the worst moments. A hiring manager asks why their favorite interviewee was passed over. A candidate asks for feedback. An executive asks why two offices hire to visibly different standards. If your only artifact is a ranked list, every answer is a reconstruction from memory — which is to say, an assumption. Decades of selection research point the same way: structured, job-related evidence beats unstructured impressions. The number is not the evidence; it is a summary of evidence you may no longer be able to produce.
| Standalone score | Decision file | |
|---|---|---|
| What it tells you | A rank: 82 beats 76 | Why: which competencies drove the result |
| Answering a challenge | “The test said so” | Evidence, scores, and recorded rationale |
| Six months later | Unexplainable | Reconstructable end to end |
| Panel discussion | Impressions compete | Everyone reads the same evidence |
What a decision file actually contains
A decision file is the complete record of how one candidate was evaluated for one role. It is not a report generated after the fact; it accumulates as the assessment runs, so nothing has to be reconstructed. When a decision is questioned, you open the file instead of your memory.
For an inside-sales req, that means you can show the exact discovery-call scenario every candidate received, the email each one wrote to revive a stalled deal, the rubric levels those responses earned, which evaluator scored them, and the sentence of rationale behind each borderline call. The decision stops being an opinion with a number attached and becomes a conclusion you can walk anyone through.
- The task, exactly as the candidate saw it — prompt, materials, and time allowed.
- The candidate's actual output: the reply they wrote, the queue they triaged, the call plan they drafted.
- Rubric scores per competency, with the level descriptors they were scored against.
- Evaluator identity and a short written rationale for each judgment call.
- Any integrity signals, plus the human review that resolved them — never an auto-verdict.
- The final decision and the stated reason, timestamped.
Work-sample outputs: the center of gravity
The heart of the file is the candidate's real work product, because it is the one artifact nobody can argue with. A résumé claims skill; an interview performs it; a work sample demonstrates it. When a support candidate has actually de-escalated a simulated billing dispute in writing, the debate shifts from "did you get a good vibe" to "read this reply — does it meet our bar?" That is a conversation a panel can have productively.
Work-sample outputs also make disagreement useful instead of political. If two evaluators split on a sales candidate, they are not trading impressions; they are pointing at the same objection-handling email and arguing about a specific sentence against a specific rubric line. The evidence anchors the debate. Without it, the loudest or most senior voice wins by default — which is exactly the failure mode a defensible process exists to prevent.
This is why SkillCort is work-sample-first rather than another test engine with a question bank. Multiple-choice results compress into a number naturally, and the number is all you keep. A work product stays reviewable forever: the panel today, the hiring manager next week, and the auditor next year are all looking at the same thing the candidate actually made.
Rubric scores: judgment made comparable
Raw work products alone are not enough, because two evaluators can read the same reply and weigh it differently. The rubric is what converts individual judgment into comparable data. Written before anyone reviews a response, it describes in concrete terms what strong, acceptable, and weak look like for each competency — "acknowledged the customer's frustration before proposing a fix" rather than "good empathy."
Scored against that standard, judgment items deserve proportional credit, not pass/fail gates. Most real support and sales scenarios have a best answer, an acceptable one, and a poor one; a rubric that only knows right and wrong throws away the distinction your team actually cares about. Record the level and the reason. "Acceptable — resolved the issue but overpromised a refund timeline" is evidence; a bare 3/5 is trivia.
The scores then earn their place in the decision file precisely because they are traceable end to end. Anyone reviewing the hire can follow the chain: this competency, this level descriptor, this specific sentence in the candidate's output, this evaluator, this written rationale. A single composite number breaks that chain at every link — it tells you the destination while erasing the route.
The audit trail: decisions that survive scrutiny
The third layer is the trail of who did what, when, and why. It sounds bureaucratic until the day you need it: a rejected candidate questions the outcome, a new leader asks how the last three support hires were chosen, or two panelists remember a calibration discussion differently. An audit trail replaces recollection with record — which evaluator scored which submission, what was flagged, how the flag was resolved, and what reasoning closed the decision.
Integrity signals belong in this trail with special care. Proportional checks — appropriate to the stakes of the role, disclosed to candidates — sometimes surface anomalies. The trail should show that a human reviewed each one and what they concluded, never that a system auto-rejected anyone. And the record stays inside this hiring decision: no score or suspicion follows a candidate to another employer. A file you would be uncomfortable showing the candidate is a file built wrong.
Where AI fits: support, never verdict
A rich decision file creates real reading load, and this is where AI genuinely helps. It can summarize a long written response, map where an answer touched each rubric competency, draft first-pass notes for an evaluator to confirm or correct, and surface anomalies for human review. Used this way, AI makes the panel faster and more consistent without touching the decision itself.
The line to hold is bright: AI never selects or rejects a candidate. The moment a model's output becomes the verdict, your decision file stops being a record of human judgment and becomes a record of deference — unexplainable in exactly the situations the file exists for. In a defensible file, every AI summary is labeled as a summary, and every conclusion has a human's name on it.
Assembling a file your whole panel trusts
Panel trust is earned by process, not asserted. Evaluators trust a file when they helped set its standard, when their scoring aligns with colleagues on shared samples, and when the file demonstrably changes outcomes — the strong-work-sample, quiet-interview candidate who gets a fair hearing because the evidence spoke. Build that trust deliberately, before the first real candidate arrives.
The payoff compounds with every req. Questions that used to trigger an awkward reconstruction — why him, why not her, why did the two panels land on different standards — now share one calm answer: open the file. New evaluators onboard faster because the standard is written down, and disputes shrink because the evidence is already assembled. That is the difference between defending a decision and merely having made one.
- Write the rubric with the panel, from a role blueprint, before reviewing any response.
- Calibrate on two or three sample responses until scores converge on the same level for the same reasons.
- Require a one-sentence rationale on every judgment score — it takes seconds and doubles the file's value.
- Route borderline and high-stakes calls through a second reviewer, inside the same file.
- Close every req by recording the decision and its stated reason while the context is fresh.
Key takeaways
- A single score cannot answer the question every hiring decision eventually faces: why this candidate and not that one.
- A decision file pairs three layers — real work-sample outputs, rubric scores with rationale, and an audit trail of who decided what and why.
- Work products anchor panel disagreement to evidence instead of seniority or volume; rubrics make individual judgment comparable.
- AI can summarize, map, and flag for human review, but every verdict in the file carries a human's name — and nothing follows the candidate to other employers.
- Trust is built before candidates arrive: co-written rubrics, calibration on samples, and recorded rationale on every judgment call.
