What you'll need
- A role to hire for, with input from someone who does or manages the work
- A SkillCort workspace with permission to build and publish
- The evaluators who will score, with time reserved for a calibration session
- Your candidate list, or the pipeline that will produce it
- Agreement on who owns the final decision
Step 1: Start from a role blueprint, not a question list
Before touching the platform, write down what the role actually requires: the handful of competencies that separate someone who thrives from someone who struggles, and roughly how they rank. This blueprint is the spine of everything downstream — tasks demonstrate these competencies, rubrics score them, weights encode their ranking, and role-fit reports against them.
Keep it observable. "Handles ambiguous requests with sound judgment" can become a task and a rubric criterion; "team player" cannot. If you cannot imagine the work sample that would reveal a competency, it does not belong in the blueprint yet.
Step 2: Build the tasks and rubrics in the items library
In the items library, create the tasks that let candidates demonstrate each competency — organized in folders with folder-scoped tags so the next assessment can reuse them. Lead with work samples: a realistic ticket to answer, a case to analyze, a file to produce. SkillCort's item types cover text and rich-text responses, file uploads, code, table inputs, audio and video recording, and more, with side-by-side stimulus layouts for document-based tasks.
You can author manually, AI-generate items in bulk — up to 100 per run, optionally grounded in your own source documents — or quick-paste multiple-choice questions. Treat AI drafts as drafts: review and edit every generated item before it faces a candidate.
Give every judgment-scored task a rubric. Derive criteria from the blueprint, anchor levels in observable behaviors, and weight the criteria that matter most. Rubric design is a craft with its own tutorial; the non-negotiable here is that rubrics exist before responses do.
Step 3: Assemble the assessment in the builder
In the builder, structure the assessment as sections, pages, and blocks — typically a section per competency area, pages that pace the work, and blocks holding the items you import from your library, filtered by folder, tag, or type. Set a total time limit and per-page limits that reflect realistic working pace with margin for first-time readers.
Set section scoring weights (×N) to mirror the blueprint's ranking — these weights flow directly into role-fit, so this is where your priorities become arithmetic. If you use shuffling for fairness, the builder offers it at three levels; keep work-sample sequences in a fixed, sensible order where the flow matters.
Step 4: Preview, QA, and clear the publish checklist
Use "Preview as candidate" from the builder header and walk every page on desktop and on a phone, answering for real. Time yourself against your limits, verify every rubric and answer key against the task as finally written, and remember that a task without an answer key auto-scores as null — flagged for manual review, not zero.
The publish checklist blocks Publish until required setup is complete; clear it, then run a short pilot with one or two colleagues under real conditions before any candidate is invited. The full QA routine has its own tutorial — do not skip it, since a bug found by a candidate costs signal you cannot recover.
Step 5: Configure delivery with proportional security
Publish, then create the delivery: a name, the time window candidates can attempt in, and security proportional to the stakes. SkillCort offers lockdown with optional violation limits, copy/paste blocking, second-screen detection, webcam or screen recording, and ID check — toggles, not a package deal.
Proportional is the operative word. A first-round work sample rarely justifies webcam recording; a certification exam may justify most of the toggles. Candidates see a consent and preflight gate before starting, so whatever you enable is disclosed and checked up front — every control you add is friction for honest candidates, so add only what the stakes earn.
Step 6: Invite candidates and monitor the window
Send email invitations from the delivery. Around the invitation, set expectations in plain language: what the assessment involves, how long it takes, that it works best on a desktop if any task needs one, and that responses are scored against a consistent rubric by human evaluators.
During the window, watch completion rather than performance. Nudge non-starters before the window closes, and note any integrity events for structured review later — during delivery they are observations on a timeline, not judgments. Resist peeking at early responses to form impressions; scoring starts after calibration, against the rubric, or the rubric was theater.
Step 7: Evaluate with a calibrated panel
Before scoring begins, run a calibration session: evaluators independently score two or three reference responses against the shared rubric, compare per-criterion deltas, and document agreed interpretations. It takes an hour and is the difference between scores that are comparable and scores that reflect who happened to review whom. The session format has its own tutorial.
Then score. In the level-based matrix, evaluators click the level card whose anchored description matches the response, attaching notes as evidence. AI can draft summaries and criterion-level suggestions — from responses or interview transcripts, with model and version logged — but a named person confirms or overrides every score. Multiple evaluators are averaged within each criterion before weights combine anything.
- Calibrate before the review window opens, not midway
- Attach an evidence note to every criterion score
- Route null auto-scores (no answer key) to manual review promptly
- Run a blind re-score drift check if the window spans weeks
Step 8: Decide on the board and export the audit pack
On the Decision Board, read role-fit as a weighted starting point: it reflects your criterion, skill, and section weights, and it orders your attention. Compare finalists per competency, drill from any score into the underlying evidence — task outputs, rubric notes, the integrity timeline — and resolve risk flags as context for human review, never as auto-rejects. Reading the board well is its own tutorial.
Record the decision with its reasoning, then export the artifacts: a white-label candidate evidence report for stakeholder conversations, and the audit pack as the durable record — one self-contained file with the assessment structure and weights, rubric versions, every score with its named evaluator and timestamp, plain-language methodology, AI provenance, integrity resolutions, and the decision log.
That file is the point of the whole exercise. You did not just rank candidates; you produced a decision you can explain — to a stakeholder, to an auditor, or to a candidate who asks how the call was made.
Pro tips
- Timebox the first run: a focused assessment shipped this week beats a comprehensive one shipped next month, and the library makes the second run dramatically faster.
- Let the blueprint discipline scope — every task, criterion, and weight should trace to a competency, or it is measuring something you did not choose.
- Reuse deliberately: folders and tags in the items library turn this assessment into the starting point for the next role in the family.
- Keep total candidate time proportional to the stage — a long assessment filters for free time, not skill.
- Do a retrospective after the first cohort: which tasks separated candidates, which rubric anchors caused hesitation, which security toggles earned their friction.