What you'll need
- A SkillCort workspace (or a demo)
- One or two real inbound emails the role answers — anonymized
- The facts and policy the reply must respect
- Two or three replies from your team, from strong to weak, for calibration
Step 1: Pick a real inbound message worth replying to
Choose an email where the reply requires judgment, not just information. For support hiring, that is the frustrated customer whose request sits half inside policy — a refund on day 35 of a 30-day window. For sales, an inbound lead with a vague ask and a buried objection about price.
Anonymize the message but keep its texture: the irritation, the half-relevant aside, the question the sender did not quite ask. Sanitized, tidy prompts produce sanitized, tidy replies, and you learn nothing about how the candidate handles the real inbox. Gather the supporting facts too — the order record, the policy line, the pricing detail — because the reply will be scored on accuracy against them.
Step 2: Build the item: message on one side, reply beside it
Create a folder in the items library for the role and add an item with the side-by-side stimulus layout: the inbound email and any needed context — order details, the policy excerpt, the pricing page — on one side, the reply editor beside it. That mirrors a real inbox, where the message stays visible while you write.
Choose a text or rich-text response for the answer. Rich text lets candidates use structure — a short greeting, grouped points, a clear next step — which is itself part of what you are scoring. In the prompt, state whom the reply goes to and what a sent-quality reply means: this is the email the customer receives, not a draft.
Step 3: Anchor a rubric on tone, clarity, accuracy, and structure
Build a level-based matrix rubric with those four criteria and write an observable anchor in every cell. Tone: acknowledges the customer's frustration before the fix, versus opens with policy. Clarity: the next step is unmissable, versus buried in a paragraph. Accuracy: every stated fact and promise is true to the brief. Structure: scannable in ten seconds, versus a wall of text.
Weight for the role — tone usually leads in support, clarity and accuracy in sales. Keep grammar and typos as a minor input to clarity rather than a criterion of their own; you are hiring communicators, not proofreaders. Attaching the rubric freezes a snapshot, so the standard is identical for every candidate in the run.
- Tone: matches warmth to the situation, acknowledges before fixing
- Clarity: the point and next step are unmissable
- Accuracy: nothing promised that the policy cannot deliver
- Structure: scannable — short paragraphs, one idea each
Step 4: Assemble and timebox in the builder
Create the assessment and import the item through the picker. One strong email task can carry a short screen on its own; for more coverage, add a second inbound message of a different type — one angry customer, one confused prospect — each in its own page and block.
Set a total time limit of 20–30 minutes and a per-page limit of about 10–15 minutes per email, matching the pace of a real queue. In the candidate instructions, say exactly what they are being scored on — tone, clarity, accuracy, structure — and that they should send the reply they would genuinely send. Honest framing produces representative writing.
Step 5: Calibrate evaluators on sample replies
Before real responses arrive, have every evaluator independently score your two or three team-written samples — one strong, one middling, one weak — against the rubric. Compare scores criterion by criterion and argue out the gaps: one evaluator's "warm" is another's "unprofessional" until the anchors settle the question.
Tighten any anchor that produced disagreement, then keep the calibrated samples as reference points during live scoring. Twenty minutes of calibration is the difference between a rubric on paper and a shared standard in practice — especially for tone, the most subjective criterion on the sheet.
Step 6: Preview, publish, deliver, and evaluate
Run Preview as candidate from the builder header and write a reply yourself under the clock; fix anything ambiguous in the message or the context you provided. Clear the publish checklist — it blocks Publish until scoring, timing, and instructions are complete — then publish and create a delivery with a name, a time window, and proportional security. For a writing screen, copy/paste blocking is typically all you need.
Invite candidates by email; they pass the consent and preflight gate and write their replies. In evaluation, your calibrated panel scores each reply on the shared matrix by clicking level cards. AI can draft a summary and suggest criterion scores, model and version logged, and a named evaluator confirms or overrides each one — then the Decision Board lines up candidates on weighted role-fit, with every actual reply one click away.
Pro tips
- Use a real anonymized email — invented prompts are always too tidy to test judgment.
- Score the reply a customer would receive, not an essay about customer service.
- Keep grammar a minor input to clarity, not a criterion of its own.
- Calibrate on samples before live scoring; tone is where evaluators drift most.
- Two contrasting emails (angry customer, confused prospect) double the signal for ten extra minutes.