Step 1: Pick the slice of real work worth simulating
Start from your role blueprint and choose the moments where the competencies you weighted most heavily actually show up. The best slices are decision-dense: a short stretch of work where the candidate must notice something, choose between defensible options, and produce an artifact. A support agent's whole shift is not a task; the ten minutes where a refund request sits just outside policy and the customer is already frustrated is.
Go back to the artifacts you gathered for the blueprint and pick a real case as your seed. A genuine ticket, anonymized and lightly edited, beats an invented one every time — real cases carry the awkward details that make judgment visible, like a customer who is partly right, or a prospect whose objection hides a budget problem. Invented scenarios drift toward tidy puzzles with one right answer, which is exactly what a work sample is not.
One slice per task. If you find yourself bundling a triage exercise, a written reply, and a process question into one prompt, split them into separate tasks. Each task should observe one primary moment of work, with at most two or three blueprint competencies surfacing naturally inside it — a task built to check everything produces evidence about nothing in particular.
Step 2: Choose the format that matches the skill
The format is not a style choice — it is a validity choice. Every format captures some behaviors and hides others, so match it to how the skill is actually exercised on the job. Written judgment belongs in a ticket simulation. Spoken composure belongs in an audio or video response. If you assess a phone-heavy inside sales role entirely in writing, you have measured a different job.
When two formats could work, prefer the one closer to the job's real medium, and only trade down for practical reasons you can name. A live role-play is the highest-fidelity way to observe negotiation, but an asynchronous audio response to a recorded objection scales further and still captures tone, structure, and recovery. Write the trade-off into the task notes so future reviewers know it was deliberate.
The mapping below is a reliable default for support and inside sales roles. Treat it as a starting point rather than a rulebook: if your support team handles half its volume on live chat, a timed chat simulation may fit better than an email-style ticket, and if your reps sell over video, a video response beats audio alone. The blueprint tells you which behaviors matter; the format decides whether you will actually see them.
- Ticket simulation → written judgment, tone, and policy application: reply to a live-feeling support ticket or email thread.
- Case study → prioritization and analysis: triage a queue of five tickets, or pick which three of eight stalled deals to work first, with reasoning.
- File upload → produced artifacts: a post-call follow-up email, a short proposal, a corrected data sheet.
- Audio or video response → spoken communication: respond to a recorded angry customer or a prospect's pricing objection.
- Role-play → live interaction and adaptability: a discovery call or escalation handled with a trained interviewer, scored on the same rubric.
Step 3: Write context that feels like the job
A realistic task gives the candidate what a new hire would actually have: a company one-pager, the relevant policy excerpt, the customer's history, the prospect's last two emails. Write these as in-world documents, not exam preamble. A support candidate should read a refund policy formatted like a help-center article, not a bullet list labeled 'rules for this test.'
Include the friction that makes the work real. The customer's second message contradicts their first. The CRM notes are thin. The prospect asked a question your one-pager only half answers. These are not tricks — they are the texture of the job, and they are where your blueprint behaviors become visible. A candidate who asks for the missing information, or names the ambiguity in their reply, is showing you exactly what you came to see.
Keep the reading load honest. Everything you include should be plausibly consulted while doing the task; if a document exists only to bury a gotcha, cut it. Aim for a context pack a candidate can absorb in a few minutes, because the skill you are measuring is the work, not speed-reading.
Step 4: Set constraints that reveal judgment
Constraints are what turn an open prompt into a work sample. State the candidate's role and authority explicitly: you can refund up to a stated amount without approval; you cannot promise a ship date; discounting beyond a threshold needs a manager. Judgment is only observable against boundaries, and unstated boundaries just measure who happens to guess your policies.
Define the deliverable precisely — one reply the customer will receive, or one follow-up email plus a two-line CRM note — and say who will read it. 'Write to the customer, not to us' changes what candidates produce, and it makes outputs comparable. Leave the approach open: the constraint is the situation and the deliverable, never the path. If your prompt implies the answer, you have written a comprehension check, not a work sample.
Step 5: Timebox fairly
Set the time limit from evidence, not instinct. Have two or three people on the current team do the task cold and note their times; a fair candidate limit is comfortably above what a competent insider needs, because candidates lack context and are working under pressure. The limit should make the task honest — no outsourcing an afternoon of polishing — without turning it into a typing race.
Keep the total sequence proportional to the role. For support and inside sales hiring, a focused work-sample sequence of roughly thirty to forty-five minutes end to end respects candidates and still yields rich evidence; a multi-hour take-home filters for free time, not skill. If a single task needs an hour, question the task before questioning the limit — you probably picked too wide a slice in step one.
Tell candidates the timing rules up front: how long each task allows, whether they can pause between tasks, and what happens if they run out. Time pressure that surprises people measures their anxiety, not their skill, and accommodations should be straightforward to grant because the limits were never load-bearing traps.
Step 6: Standardize inputs so outputs are comparable
Comparability is the whole point of a structured work sample: every candidate faces the same situation, so differences in output reflect differences in skill. Audit the task for anything that could vary between candidates — instructions phrased differently by different recruiters, a scenario updated mid-req, an attachment some candidates receive and others do not — and lock it all down in one canonical task definition.
Standardize the response envelope too. If one candidate answers in a paragraph and another in a formatted document because the prompt never said, your evaluators will score presentation instead of the competency. Specify the medium, the rough length, and the fields to fill. Then version the task exactly like the blueprint: candidates within a req always see the same version, and any change ships as a new version with a note.
- One canonical prompt — no recruiter paraphrasing, no per-candidate edits.
- Identical context pack and attachments for every candidate in the req.
- Specified deliverable format and length, so evaluators compare substance.
- Version-locked per req; changes create a new version with a changelog note.
- Consistent tooling: same editor, same upload flow, same recording setup.
Step 7: Pilot internally before any candidate sees it
Run the finished task on three kinds of insiders: a top performer, a solid-but-average performer, and someone adjacent to the role, such as a teammate from another queue. You are testing two things. Clarity: did anyone misread the instructions, miss an attachment, or run out of time for the wrong reasons? And discrimination: did the top performer's output actually look different from the average one's? If everyone produces the same answer, the task is too easy or too closed — widen the gray area.
Debrief each pilot participant while the task is fresh: what felt artificial, what information they wanted and could not find, where they guessed at your intent. Fix the prompt, not the person. Then keep the pilot outputs — they become your first anchor examples when you build the rubric, and the range between the strong and average responses tells you what a meaningful score difference will look like. Only when the checklist below is clean does the task go live.
- Instructions survived contact: no pilot participant asked a question the prompt should have answered.
- The strong performer's output was visibly better in ways your blueprint behaviors describe.
- Time limit left the average insider a reasonable margin.
- The seed scenario is fully anonymized — no real customer names, amounts, or identifiers.
- An answer-guide draft exists: the range of strong, acceptable, and weak responses the pilot revealed, ready to feed the rubric.
Key takeaways
- Simulate a decision-dense slice of real work, seeded from an anonymized real case — one slice per task.
- Choose format by skill: ticket simulations for written judgment, audio or video for spoken skills, case studies for prioritization, role-plays for live interaction.
- Write in-world context with realistic friction, and state role, authority, and deliverable explicitly.
- Timebox from insider trials, keep the full sequence proportional, and standardize every input so outputs compare cleanly.
- Pilot on strong, average, and adjacent insiders — the task must separate them before it meets a candidate.