By Daniel Whitmore, MSc in Industrial and Organizational Psychology
In short
A candidate evidence wallet is the full record behind an assessment score: the task outputs the candidate actually produced, the criterion-level rubric scores with each evaluator's notes, and the integrity timeline of the attempt. Reading it means starting from the role-fit number and drilling down one layer at a time — so a hiring panel debates the same evidence instead of trading impressions.
Start with the number, then refuse to stop there
Picture a shortlist for a customer support role. One candidate sits at the top with a strong role-fit score; another trails a few points behind. If the panel treats those numbers as the decision, the meeting is over in five minutes — and it will be the wrong five minutes. A single score compresses judgement, tone, troubleshooting, and follow-through into one digit, and compression always loses information.
The score's real job is to tell you where to look, not what to conclude. Two candidates can land near-identical totals through completely different routes: one wrote warm, precise replies but stumbled on a refund-policy judgement call; the other handled policy flawlessly but opened every message with the fix and no acknowledgement. Same number, different hires. The evidence wallet exists so the panel can see which story the number is summarizing.
So the walkthrough starts by opening the wallet, not by ranking the totals. Everything that follows is the same evidence every panel member sees, in the same order: the work itself, the criterion-level scoring with evaluator notes, and the integrity timeline. No layer asks you to take anyone's word for anything.
Layer one: the task outputs — the work itself
The first layer is the candidate's actual work, presented next to the prompt they received. For the support candidate, that means the full reply they wrote to a frustrated customer whose order arrived damaged, the order in which they triaged a small queue, and the response they chose in a judgement scenario — with their reasoning. For an inside-sales candidate, it is the discovery email they drafted from a lead brief, their objection-handling response, and how they prioritized a pipeline snapshot.
Reading raw output changes the texture of the discussion immediately. "Scored well on written clarity" is abstract; the reply that opens with "I can see why this is frustrating, and here is exactly what I will do" is not. Panel members notice things a score cannot carry: a candidate who quoted the policy accurately but buried the apology, or one who promised a delivery date nobody could guarantee.
This is also where work-sample assessment earns its keep over a quiz. A multiple-choice result gives you a right-or-wrong tally; a work sample gives you an artifact you can hold up in the room and discuss. The output is the evidence — everything else in the wallet is structure around it.
Layer two: criterion scores and the notes behind them
The second layer breaks the role-fit number into the criteria the role blueprint defined — and shows how each evaluator scored each one, with their written rationale. For the support role, that might be empathy and tone, written clarity, judgement within policy, and accuracy. For inside sales: discovery quality, objection handling, written persuasion, and pipeline judgement. Each criterion carries the rubric descriptor it was scored against, so "3 out of 4 on judgement" is anchored to concrete language, not a reviewer's mood.
The evaluator notes are the most underrated part of the wallet. A note like "acknowledged the customer before the fix, but committed to a refund outside policy without flagging it" tells the panel precisely what to weigh. When two evaluators diverge on the same criterion, the wallet shows both scores and both rationales side by side — disagreement becomes visible and discussable instead of silently averaged away.
Reading this layer, the panel can ask sharper questions. Is the policy stumble a training gap or a judgement gap? Did the second evaluator read the tone differently, and does their note explain why? These are questions about evidence, and they have answerable forms — unlike "I just got a better feeling from the other one."
- Every criterion links back to the role blueprint, so the panel debates what the role needs, not what interviews rewarded.
- Scores are anchored to rubric descriptors written before anyone reviewed a response.
- Evaluator notes record the why behind each score, in the evaluator's own words.
- Divergent scores are shown side by side, turning disagreement into a calibration conversation.
- AI-drafted summaries or suggested scores, where used, are labeled as such — a human confirmed or overrode every one.
Layer three: the integrity timeline
The third layer is the integrity timeline: a factual record of how the session unfolded. When the candidate started and finished each task, how time was distributed, and any signals worth a human look — a long idle gap, a large block of pasted text, a pattern that reads differently from the rest of the session. It is a timeline, not a verdict.
Reading it well means reading it proportionally. A tab switch during a support task that explicitly allowed reference material is not a finding; it is a candidate doing what the instructions permitted. A response pasted in wholesale seconds after the prompt opened is worth discussing — and the discussion belongs to the panel, in context, alongside the quality of the work itself. Nothing in the timeline rejects anyone automatically, and no signal follows a candidate to another employer.
In practice, the timeline mostly builds confidence rather than suspicion. Seeing that a candidate spent nine minutes drafting and revising the hard reply — deleting a curt first attempt — is evidence of exactly the deliberation the role needs. The integrity layer is there to make the evidence trustworthy, not to put candidates on trial.
What changes when the panel reads the same wallet
Now run the hiring meeting again — same two candidates, but this time everyone has read the same wallet. The conversation changes shape. Instead of trading impressions ("she seemed sharper on the phone screen"), the panel points at artifacts: this reply, this criterion, this note, this moment in the timeline. Selection research has long favored structured, job-related evidence over unstructured impressions, and a shared wallet is what makes structure survive contact with a group discussion.
Disagreements get better, too. When the sales manager and the recruiter split on a candidate, the wallet localizes the split: they agree on written persuasion, diverge on pipeline judgement, and the divergence traces to one prioritization call the candidate explained in their rationale. That is a resolvable disagreement. Without shared evidence, the same split becomes a status contest, and the loudest reader of the room wins.
The wallet also protects the decision after it is made. If a rejected candidate or an internal stakeholder asks how the call went, the answer is not a reconstruction from memory — it is the file: task, output, rubric, scores, notes, timeline, and the named people who decided. That is what a defensible hiring decision looks like.
How to run the walkthrough with your own panel
None of this requires a heavyweight process — it requires an order of operations. Ask every panel member to read the wallet before the meeting, and open the meeting at the evidence rather than the ranking. The score is the table of contents; the layers are the book. For a customer-support or inside-sales shortlist, fifteen minutes of pre-reading per candidate is usually enough, and it repays itself before the first meeting ends.
A few simple disciplines keep the discussion honest and short. They are not rules about the tool; they are rules about how people read together, and they work in any panel that agrees to hold them. Most teams find that reading the same evidence makes meetings faster rather than slower, because the debate starts where the actual disagreement lives instead of spending half the hour discovering it.
- Read outputs before scores: form a view of the work, then check it against the rubric layer.
- Discuss criteria one at a time; do not let a strong criterion halo the weak ones.
- Treat evaluator-note divergence as the agenda — that is where the meeting earns its time.
- Review integrity signals in context and as a group; never let a flag stand in for a judgement.
- Record the panel's final rationale in the wallet, so the decision file is complete the moment the decision is made.
Key takeaways
- A role-fit score tells you where to look, not what to decide — two identical totals can summarize very different candidates.
- Read the wallet in layers: raw task outputs first, then per-criterion scores with evaluator notes, then the integrity timeline.
- Integrity signals are context for human review, never automatic verdicts — and they never follow a candidate across employers.
- Shared evidence changes panel dynamics: disagreements localize to specific criteria and artifacts instead of competing impressions.
- The finished wallet doubles as the decision file — the record that makes the hire explainable and defensible later.
