What you'll need
- A finished rubric attached to the assessment (a frozen snapshot)
- 2–3 reference responses spanning different quality levels
- All evaluators who will score candidates, with 45–60 minutes together
- A place to record agreed interpretations, shared with the panel
Step 1: Pick two or three reference responses
Choose responses that will make disagreement visible: one clearly strong, one clearly weak, and — most importantly — one genuinely borderline. Pilot-run responses from colleagues work well; so do early candidate responses if your review window allows scoring to begin after calibration.
The borderline response is where calibration earns its keep. Everyone agrees on excellence and failure; panels split on the response that is strong on one criterion and shaky on another. If none of your references produces that tension, draft one that does.
Step 2: Brief the panel on the rubric before anyone scores
Walk the panel through the rubric together: each criterion, its anchored level descriptions, and its weight. The goal is shared understanding of the standard, not of any response — do not discuss the references yet. Invite questions about ambiguous anchor language now, while no score is at stake.
Remind the panel how matrix scoring works in SkillCort: for each criterion, you click the level card whose description matches the response, and that level's score is recorded. The instruction that matters most: score against the written anchors, not against your private sense of what a good answer looks like.
Step 3: Score the references independently
Each evaluator scores every reference response alone, without discussion, using the shared rubric. Independence is the whole point — the session measures how the panel naturally diverges, and any conferring before scores are committed destroys that measurement. Give everyone the same quiet window to work in, and hold all comparisons until every score is in.
Ask evaluators to attach a short note to each criterion score citing the evidence in the response that justified the level they chose. In SkillCort, notes attach to scores as evidence, and in this session they are what turns "we disagree" into "we are reading this anchor differently" — the productive kind of disagreement.
Step 4: Compare scores per criterion, not per total
Bring the scores together and compare them criterion by criterion. Totals hide the story: two evaluators can land on similar overall scores through opposite readings of two criteria, which means their agreement is luck, not calibration. The per-criterion deltas show exactly where interpretations diverge.
For each criterion with a spread, have the differing evaluators read their evidence notes aloud — the notes show what each person was actually looking at when they chose a level, which is far more useful than debating the numbers. Almost every gap resolves into one of a few patterns, and naming the pattern tells you what to fix.
- An ambiguous anchor that supports two readings — tighten the language interpretation
- One evaluator scoring an adjacent criterion's evidence — restate the boundary
- A private standard imported from past experience — return to the written anchor
- Genuine judgment differences on borderline work — agree on a convention
Step 5: Agree on anchors and document every interpretation
Resolve each divergence into an explicit agreement: which level does this kind of response earn, and why. Write the agreements down in a shared calibration note — anchor clarifications, boundary rules between criteria, conventions for recurring edge cases. Undocumented agreements evaporate within a week and never reach an evaluator who joins later.
Remember that the attached rubric is a frozen snapshot, so mid-assessment you are documenting interpretations of the anchors, not rewriting them — which is what keeps scoring consistent for candidates already evaluated. Feed genuine wording fixes back into the rubric in your bank so the next assessment attaches a sharper version.
Step 6: Re-score a reference to confirm alignment
Close the session by having everyone independently score one more response — a fresh reference if you have one, or the borderline case again — applying the documented agreements. If the per-criterion spread has visibly narrowed, the panel is calibrated enough to start scoring candidates.
If a criterion still splits the panel, do not average the problem away. A persistent split means the anchor or the agreement is still ambiguous; spend ten more minutes on that one criterion rather than sending an unresolved disagreement into every candidate's score.
Step 7: Schedule drift checks during the review window
Calibration decays. Over a long review window, evaluators drift — standards creep up as strong responses reset expectations, or soften with fatigue. Plan brief drift checks: periodically pick a recently scored response and have a second evaluator score it blind, then compare per-criterion deltas exactly as in the session.
This is also where AI assistance stays in its lane. SkillCort's AI can draft summaries and criterion-level score suggestions, which speeds review — but every score is confirmed or overridden by a named person, and drift checks verify the humans, since it is their judgment the decision will rest on. Reconvene briefly if a drift check surfaces a new divergence.
Pro tips
- Calibrate before the review window opens, not after half the candidates are scored — recalibrating midway creates two populations of scores.
- Keep the session to 45–60 minutes; a tight session on three responses beats a long one that dissolves into rubric redesign.
- Never let the most senior person score first or speak first in comparisons — independence is the mechanism, hierarchy is its enemy.
- Store the calibration note next to the assessment so every future evaluator inherits the agreed interpretations.
- Watch for one evaluator being consistently harsher across all criteria — that is a personal baseline to discuss, not a rubric problem.