What a work sample actually is
A work sample is a short, realistic task that mirrors what the person will do on the job, performed under conditions close enough to real work that the result is meaningful. It is not a knowledge quiz and it is not a personality questionnaire; it is a slice of the actual work, scored against a standard.
The appeal is simple: the best predictor of whether someone can do a task is watching them do a representative version of it. Done well, a work sample gives you evidence you can compare across candidates, rather than an impression you have to defend from memory. Done badly — too long, ambiguous, or scored on gut feel — it just adds cost, so the examples below pair each task with how to keep it fair and proportional.
Example: customer support ticket
Give the candidate a realistic inbound message — an upset customer whose order arrived damaged, say — plus the relevant policy, and ask them to write the reply. You are looking at how they open (acknowledgement before policy), whether they stay accurate (no invented ship dates), and whether they end with a clear next step.
A lighter variant is a 'choose the best reply' work sample: present two or three drafted responses and ask which is strongest and why. It removes writing-speed noise and grades pure judgement about tone, clarity, and accuracy — useful for high-volume screening. Candidates can practise the format on our free customer-support and email-writing practice sets so the task itself is no surprise.
- Task: reply to a realistic customer message using the given policy.
- Measures: empathy and tone, written clarity, accuracy, clear next step.
- Fair scoring: rubric with strong/acceptable/weak descriptors per dimension.
Example: sales role-play or written follow-up
For inside sales, a short role-play (live or recorded) around discovery or objection handling reveals whether a candidate listens before pitching, handles a price objection honestly, and advances the deal without pressure. A lower-effort written version asks them to draft the follow-up email after a described call, or to plan the questions they would ask in discovery.
The signal you want is judgement, not aggression: does the candidate understand the buyer's need, tell the truth about fit, and propose a clear next step? Score against a rubric so 'persuasive' does not quietly become 'pushy'.
- Task: role-play or written follow-up for a described sales situation.
- Measures: discovery, honest objection handling, follow-through, integrity.
- Fair scoring: reward customer-first judgement, not pressure tactics.
Example: business email / writing task
Written communication is most of the job in many support, sales, and coordination roles, yet it rarely gets assessed directly. A writing work sample asks the candidate to respond to a realistic scenario — delivering bad news about a deadline, correcting a mistake, or answering a comparison question honestly — and scores tone, clarity, accuracy, and structure.
To keep scoring objective and fast, a strong option is a multiple-choice 'which drafted reply is best' format rather than free-text grading: it isolates the writing judgement, removes subjective marking, and gives instant, consistent results — which is exactly the format of our free business email writing practice.
- Task: choose or draft the strongest reply to a realistic work situation.
- Measures: tone, clarity, accuracy, structure, judgement about what to include.
- Fair scoring: 'which reply is best' isolates judgement and avoids subjective grading.
Example: data accuracy / attention to detail
For roles where a small error is expensive — order processing, data entry, account setup — a data-accuracy work sample asks the candidate to reconcile figures, spot inconsistencies between a form and a record, or catch an error in a report. Numerical and logical reasoning practice tests are a candidate-side warm-up, but the work sample itself should use the kind of data the role really touches.
Score on accuracy and on the process the candidate used, not just the final answer — a candidate who documents how they checked their work is showing a habit you want. Keep the task short; attention to detail is visible in ten focused minutes.
- Task: reconcile or verify realistic figures and flag discrepancies.
- Measures: accuracy, thoroughness, a sensible checking process.
- Fair scoring: credit method and documented checks, not only the final number.
How to score work samples fairly with a shared rubric
A work sample is only as fair as its rubric. Before anyone reviews a response, write down what a strong, acceptable, and weak answer looks like for each competency, in concrete language. This is what makes two reviewers reach the same score and what turns a subjective reaction into comparable evidence.
Score proportionally where tasks have more than one reasonable answer — best response full marks, acceptable partial, poor none — and record the reasoning, not just the number. Use a second reviewer on borderline or high-stakes cases, and keep any integrity monitoring proportional to the stakes. Fair, transparent scoring is also the most defensible if a decision is ever challenged.
- Write concrete strong/acceptable/weak descriptors per competency, in advance.
- Score proportionally; record the rationale alongside the score.
- Calibrate with a second reviewer on borderline and high-stakes cases.
- Keep integrity checks proportional and fair — never a surveillance dragnet.
How AI supports evaluation — and how samples become a decision file
AI earns its place in evaluation by making human review faster and more consistent: summarising a long response, mapping an answer to the competencies it touched, drafting first-pass notes, and flagging a possible integrity signal for a person to check. That reduces reviewer fatigue and helps evaluators stay aligned to the rubric.
The line that must not move is who decides. AI does not select, reject, rank, or carry a hidden score about a candidate across employers — a human reviews the evidence and makes the call. Assembled together, the task, the response, the rubric scores, the reviewers, and the reasoning form a decision file: comparable across candidates, explainable to stakeholders, and defensible after the fact. That is what turns a set of work samples into an evidence-based hiring decision.
Key takeaways
- A work sample is a short, realistic slice of the job — the closest thing to watching someone do it.
- Match the sample to the role: support tickets, sales role-plays or follow-ups, writing tasks, data-accuracy checks.
- A 'which reply is best' format gives fast, objective writing signal without subjective free-text grading.
- Score with a rubric written in advance, proportionally, with recorded rationale and a second reviewer on close calls.
- AI supports evaluation (summaries, flags) but never decides; the assembled evidence becomes a defensible decision file.