What you'll need
- A SkillCort workspace (or a demo)
- Three to five real incidents where judgment made the difference
- A view on what good judgment means for this role
- A colleague who handled those incidents, for a realism check
Step 1: Collect real incidents, not hypotheticals
The best SJT scenarios are not invented — they are remembered. Ask the team for moments in the last quarter where two reasonable people might have acted differently: the discount request that stretched policy, the deadline that collided with a quality problem, the colleague who quietly missed a handoff.
Real incidents carry the texture that makes judgment visible: incomplete information, mild time pressure, and no option that is perfect. Invented scenarios drift toward one obviously right answer, and an SJT where every strong candidate spots the intended answer instantly is measuring reading comprehension, not judgment under real constraints.
Anonymize as you collect. Keep the structure of the dilemma — the competing pressures, the missing information, the deadline — and drop the names, account details, and anything traceable to a real customer or colleague. The predictive value lives in the shape of the situation, not in its private specifics.
Step 2: Define what good judgment looks like for the role
Before writing a single scenario, write down the judgment dimensions you are scoring — for a support role that might be de-escalation, fair use of policy, and knowing when to escalate; for a manager, prioritization and how they handle ambiguity. Two to four dimensions is plenty.
For each dimension, note what better and worse choices look like in behavioral terms. This list becomes your answer key or rubric later, and it stops the test from drifting into personality guessing. You are scoring decisions and their trade-offs, not temperament.
Step 3: Draft the scenario stems in the items library
Create a folder for the role in the items library and tag items by judgment dimension, so you can filter and pull the right mix later. Write each scenario stem as a short, concrete situation: who is involved, what just happened, what constraint applies, and what the candidate must decide.
Keep stems to 80–150 words. Long enough that the trade-off is real, short enough that reading speed is not the test. If you have many scenarios drafted elsewhere, AI generation can produce variations — upload a source document such as an incident log as grounding, state a count, and edit the results by hand.
- One decision per scenario — never two dilemmas in one stem
- Include the constraint that creates the tension (policy, time, resource)
- Write in second person: "A customer tells you..."
- No insider jargon a new hire could not know
Step 4: Write response options with genuine trade-offs
For the multiple-choice format, write four or five options where every option is something a real person might do. The classic pattern: one best response, one or two acceptable responses with a visible downside, and one or two poor responses that a plausible but weaker candidate would pick.
The discipline is in the wrong answers. A poor option should not be absurd — it should be the tempting shortcut: quoting policy verbatim at an upset customer, escalating something the candidate should own, promising a fix they cannot guarantee. If your team debates which option is best, the scenario is working.
If you already have scenarios and options drafted in a document or spreadsheet, quick-paste brings multiple-choice questions into the items library in bulk, ready to edit. That is often faster than authoring each item by hand, and it keeps the writing work in whatever tool your team drafted in.
Step 5: Choose the scoring model: answer key or rubric
Multiple-choice SJTs score against a key, and the key should be proportional: full credit for the best response, partial credit for acceptable ones, none for poor ones. Judgment is rarely binary, and pass/fail scoring on a trade-off question punishes reasonable thinking. Record the rationale for each option's credit so evaluators and stakeholders can see why.
For richer signal, add one or two written-response scenarios using a text or rich-text item: "What would you do, and why?" Score these against a rubric with criteria such as reasoning quality, awareness of trade-offs, and chosen action — the rubric freezes into a snapshot when attached, so the standard cannot drift mid-run.
Step 6: Assemble, timebox, and shuffle in the builder
Create the assessment and import your scenarios through the picker, filtering by folder and dimension tags. A balanced SJT covers each judgment dimension with at least two scenarios, so one oddly worded item cannot decide a dimension on its own.
Set a total time limit that allows careful reading — roughly two to three minutes per multiple-choice scenario, more for written responses, enforced per page if you want steady pacing. Enable shuffle at the block level so scenario order varies per attempt (seeded, so each attempt is reproducible). Write candidate instructions that say plainly: choose the response closest to what you would actually do.
Step 7: Preview, publish, and deliver
Use Preview as candidate from the builder header and take the test cold. Watch for the two classic SJT failures: an option that is accidentally obvious, and a stem that needs context only insiders have. Fix both before the publish checklist, which blocks Publish until scoring, timing, and instructions are complete.
Publish, create a delivery with a name and time window, and keep security light — an SJT is usually a screening step, so copy/paste blocking is typically enough. Candidates pass the consent and preflight gate, invitations go out by email, and results flow into evaluation and the Decision Board alongside your other evidence.
Pro tips
- If every strong performer on your team picks the same option instantly, the item is too easy — sharpen the trade-off.
- Pilot scenarios on current employees and compare their picks to your key before any candidate sees them.
- Mix formats: multiple-choice for coverage, one written scenario for depth.
- Partial credit is not generosity — it is what makes judgment scoring honest.
- Refresh scenarios from new incidents each cycle; the library folder makes swaps painless.