What you'll need
- A SkillCort workspace (or a demo)
- A real example of the work the role does — an actual ticket, brief, or task
- The two or three competencies you want the task to reveal
- Someone who does the job today, to sanity-check realism
Step 1: Pick the slice of real work
Start with the job, not the test. List what the role produces in a normal week, then pick one task that is frequent, important, and observable in 20 to 40 minutes. For a support role, that is almost always a ticket: a frustrated customer, a partial refund request that sits just outside policy, and a reply that has to be written.
Choose a task where quality varies between people who hold the same title. Replying to a routine password reset separates nobody; replying to an angry customer with a policy trade-off separates a strong hire from an average one. That variance is the signal you are paying for.
Finally, strip anything a new hire would not plausibly know — internal jargon, product trivia, tribal shortcuts that only make sense after six months in the seat. Provide whatever context is genuinely needed inside the task itself. The test should measure skill, not tenure at your company.
- Frequent: the role does this weekly, not once a year
- Important: getting it wrong has a real cost
- Timeboxed: a fair attempt fits in 20–40 minutes
- Self-contained: no insider knowledge required
Step 2: Set up a folder in the items library
Open the items library and create a folder for the role — for example, Support Specialist. Folders keep every role's tasks in one place, and each folder carries its own tags, so you can label items by competency (tone, policy judgment, troubleshooting) or difficulty and filter on those labels later.
Do this before writing anything. When you assemble the assessment, you will pull items in through a picker that filters by folder, tag, and type — a few minutes of structure now makes that step instant, and it makes the next role's assessment faster to build too.
Step 3: Build the task item and choose the right item type
Create the item manually and match the response format to the skill. A support reply belongs in a rich-text response item; a spreadsheet task fits table input; a spoken skill fits audio or video recording; design work fits file upload. The rule: the candidate should produce the same artifact the job produces.
For the ticket task, use the side-by-side stimulus layout: the customer's message, order details, and the relevant policy excerpt on one side, the reply editor beside it. The candidate reads the case and writes in one view, exactly as they would in a helpdesk.
Write the customer message from a real (anonymized) ticket rather than inventing one. Real tickets have the messy detail — a half-relevant complaint, an unclear timeline — that makes judgment visible. You can also generate variations with AI, uploading the real ticket as a grounding document, then edit the drafts by hand.
Step 4: Attach a rubric before anyone answers
Open-ended work needs a written standard, defined before you see a single response. Build a level-based matrix rubric: criteria as rows (tone, accuracy, policy judgment, clarity of next step), levels as columns, and a concrete behavior description anchoring every cell — what a strong, acceptable, and weak reply actually looks like.
Weight the criteria by what the role rewards. If de-escalation matters more than grammar, say so in the weights, not in evaluators' heads. You can optionally make the rubric visible to candidates, which is honest and tends to produce more focused responses.
When you attach the rubric, SkillCort freezes it into a snapshot. Later edits to the rubric bank never change a live assessment, so every candidate in a run is scored against exactly the same standard — the first applicant and the fortieth are measured by the identical yardstick, which is what makes their scores comparable.
Step 5: Assemble the assessment in the builder
Create the assessment and lay it out as sections, pages, and blocks. A simple work-sample test might be two sections: a short warm-up (a few multiple-choice judgment questions, quick-pasted into the library) and the ticket task itself. Import items through the picker, filtering by your folder and tags and multi-selecting.
Set timing deliberately. A total time limit keeps the experience predictable; a per-page limit on the ticket page stops one task from eating the whole attempt. Give the work sample room — a rushed reply measures typing speed, not judgment.
Set section scoring weights so the work sample carries the decision — for example, ×3 on the ticket section against ×1 on the warm-up. Those weights flow into role-fit on the Decision Board later. Finish the candidate instructions field: what the task is, how long it takes, and how it will be scored.
Step 6: Preview as a candidate and clear the publish checklist
Click Preview as candidate in the builder header and take the test yourself, end to end. Read the ticket cold. Is anything ambiguous? Does the timing feel fair? Can you write a strong reply with only what is on screen? Fix everything that trips you — a candidate under time pressure has no chance to ask.
Then work through the publish checklist. It blocks Publish until required setup is complete — scoring attached, timing set, instructions written — so a half-configured test cannot reach a real candidate. Treat any remaining checklist item as a genuine gap, not a formality.
Step 7: Publish and configure delivery
Publish, then create a delivery: give it a name, set the time window candidates can start in, and choose security toggles. Keep integrity proportional to the stakes — for a standard support hire, copy/paste blocking and fullscreen lockdown with a reasonable violation limit are usually plenty. Webcam or screen recording belongs in higher-stakes runs, not everywhere by default.
Candidates pass a consent and preflight gate before starting, so every measure you enable is disclosed up front and their setup is checked before the clock starts. Nothing runs silently in the background, which protects both the candidate's trust and your process's defensibility. Send invitations by email and watch attempts arrive within the window you set.
Step 8: Evaluate against the rubric and decide on evidence
As attempts come in, evaluators score each reply against the shared rubric — in matrix scoring, clicking the level card that matches the response. AI can draft a summary and suggest criterion-level scores, with the model and version logged, but a named person confirms or overrides every score. AI supports the read; it never makes the call.
When scoring is done, open the Decision Board to compare candidates on weighted role-fit and drill into the actual replies behind each number. The decision rests on what candidates did, with the evidence one click away — which is the entire point of a work-sample test.
Pro tips
- One great task beats five mediocre ones — depth of signal, not breadth of trivia.
- Write the rubric before you collect a single response, never after.
- Pilot the test on a current employee; their score sanity-checks both task and rubric.
- Tell candidates it is a realistic work sample — framing honestly improves completion.
- Reuse the folder: tagged items make the next role's assessment an afternoon's work.