Step 1: Classify the stakes of the assessment
Before choosing a single signal, name what the assessment decides. A 20-minute screening task for a customer support role that leads to an interview is low stakes: a false negative costs the candidate an interview slot, and there are later checkpoints. A final-stage sales assessment that directly gates an offer, or a certification a person will carry to clients, is high stakes — the outcome is consequential and there may be no later correction.
Write the classification down as one sentence: 'This assessment decides whether a support candidate advances to a structured interview — low stakes.' That sentence becomes the justification for every control you add or refuse to add. If you cannot justify a control against the stakes, it does not belong in the assessment.
Stakes can also differ within one hiring funnel. An early inside-sales writing screen and a final role-play for the same role deserve different controls, so classify each stage on its own rather than setting one policy for the whole pipeline. A funnel that tightens monitoring as stakes rise feels fair to candidates and keeps early stages friction-light.
Step 2: Pick signals proportional to the stakes
Integrity signals escalate in intrusiveness, and each step up must be earned by the stakes. At the light end sit passive environment flags: tab switches, copy-paste events, unusually fast completion. In the middle sit webcam snapshots or screen recording. At the heavy end sits full lockdown browsing. Start at the lightest tier that protects the decision, and move up only when the classification from step one demands it.
For a low-stakes support screening, tab-switch and paste flags are usually enough — a candidate pasting a polished paragraph into a live-chat simulation is worth a look, and nothing more intrusive is warranted. A high-stakes certification might justify screen recording. Lockdown should be rare and deliberate, because every tier you add filters out some honest candidates who are uncomfortable being watched, not just dishonest ones.
Also decide what you will deliberately not collect, and write that down too. Keystroke-level surveillance, always-on audio, or any signal you have no plan to review adds risk and candidate discomfort without adding decision value. A signal nobody will look at is pure cost — it cannot protect the decision, but it can still damage the experience and your reputation with candidates.
- Light: tab-switch flags, copy-paste events, timing anomalies — fits most screening.
- Medium: webcam snapshots or screen recording — reserve for high-stakes, offer-gating stages.
- Heavy: lockdown environments — rare, deliberate, and justified in writing.
- Never: signals you cannot name a reviewer and a review moment for.
Step 3: Tell candidates exactly what is monitored
Transparency is not a courtesy; it is part of the control. Before the assessment starts, tell candidates precisely which signals are collected, why, and how they are used — 'we log tab switches and pasted text; a reviewer sees these alongside your work; nothing is decided automatically.' Vague warnings like 'this session may be monitored' create anxiety without deterring anything.
Honest disclosure changes behavior in the right direction. A support candidate who knows paste events are logged will still paste their own drafted reply from a notes app — and can explain it, because you invited explanation rather than ambush. The candidates you lose to clear disclosure are mostly the ones planning to outsource the work; the ones you keep perform with less stress and rate the experience better.
Put the disclosure where it cannot be missed: on the assessment invitation and again on the pre-start screen, in plain language rather than legal boilerplate. Two sentences are enough — what is collected, and who reviews it. If a candidate later asks about a flag, that disclosure is what makes the conversation feel fair instead of adversarial.
Step 4: Review flags on a timeline, as context
A flag is a timestamp, not a finding. Review integrity signals on a timeline against what the candidate was doing at that moment, and most flags explain themselves. A tab switch during an inside-sales research task — where the instructions said to look up the prospect's company — is the candidate doing the job. Ten tab switches in a two-minute window right before a suspiciously polished essay answer is a different picture.
Read the timeline the way you read the work sample: as evidence to interpret. Ask what the flag coincides with, whether the surrounding work is consistent with the candidate's other responses, and whether the instructions made the flagged behavior reasonable. Then record your interpretation next to the flag, so the decision file shows a human looked and what they concluded.
Resist scoring integrity numerically. A single 'risk score' invites exactly the shortcut this playbook exists to prevent: treating an aggregate as a verdict instead of reading the events underneath it. Keep the flags as a list of moments with human interpretations attached, and let the decision board show the interpretation — never a number that pretends to settle the question.
- Ask what the candidate was doing when the flag fired — the task context often explains it.
- Compare the flagged answer's quality and voice against the candidate's other responses.
- Check whether the instructions invited the behavior (research tasks invite tab switches).
- Write down the interpretation, not just the flag, in the decision file.
Step 5: Define escalation before you need it
Decide now what happens when a reviewer finds a pattern that genuinely concerns them, because improvising this under time pressure produces bad outcomes. The escalation path is human at every step: a reviewer documents the concern, a second person looks at the same timeline, and if the concern holds, someone talks to the candidate. 'We noticed several long pastes during the written scenario — can you walk us through how you worked?' is a fair question, and honest candidates answer it easily.
Give the candidate a real chance to explain, and offer a supervised re-assessment when the explanation is plausible but the evidence is compromised. What must never exist in the path is an automatic reject: no threshold of flags, no risk score, no model output that removes a candidate without a named person deciding and recording why. An integrity signal ends careers only when humans stop reading it.
Step 6: Keep the controls inside your walls
Whatever a timeline shows, it describes one candidate in one assessment for one employer. Integrity findings must never follow a candidate across companies as a reputation score — a paste flag from someone else's screening process is not evidence in yours, and treating it as such would punish people for context you cannot see. Each assessment starts clean.
Close the loop by auditing the controls themselves once per hiring round. Count how many flags fired, how many a reviewer actually consulted, and how many changed an interpretation. Signals that never influenced a decision are candidates for removal; a stage where reviewers wished for more context is a candidate for one tier up. Proportionality is a setting you revisit, not a value you set once.
Key takeaways
- Classify the stakes first — a screening task and an offer-gating assessment deserve different controls, and the classification justifies every signal you collect.
- Escalate intrusiveness deliberately: passive flags for screening, recording only for high stakes, lockdown rarely — and never collect a signal nobody will review.
- Tell candidates exactly what is monitored and how it is used; clear disclosure deters misconduct better than vague warnings and treats honest candidates fairly.
- Flags are context on a timeline, not verdicts — a human interprets each one against the task and records the interpretation.
- Escalation means a conversation or a re-assessment decided by a named person; an automatic reject must not exist anywhere in the path.