What you'll need
- A documented certification standard: what competencies it attests and at what level
- A pass criterion agreed before any candidate sits the exam
- An item pool large enough to support variation between attempts
- Sign-off from your program owner on which proctoring measures the stakes justify
Step 1: Define the standard and the pass criteria first
Write down what the certificate attests before building anything: the competencies covered, the level of performance required, and the pass criterion. Decide now whether passing means an overall threshold, a minimum in every section, or both. A pass line drawn after seeing results is indefensible; one documented in advance is a standard.
Map each competency to a section with an explicit scoring weight. The weights flow into the final score, so they are part of the standard itself — a certificate where "safety procedures" carries 10% of the score attests something different from one where it carries 40%. Record the rationale alongside the weights.
Step 2: Build a banked item pool in the library
Create a dedicated folder for the certification with folder-scoped tags per competency and difficulty tier. Author items manually for anything nuanced, quick-paste multiple-choice questions you already have, and use AI generation — with your syllabus or course material uploaded as a grounding document — to draft breadth, up to 100 items per run.
Treat AI-generated items as drafts, not finished questions. Review every one for accuracy, a single defensible correct answer, and alignment with the standard before it enters the pool. For a certification, a bad item is not just noise — it is a result someone can legitimately appeal.
Step 3: Assemble the exam with seeded shuffle as parallel forms
Structure the exam as sections, pages, and blocks matching your competency map, importing pool items through the picker modal with folder, tag, and type filters. Then enable shuffle at the section, page, and block levels as your design allows — each attempt gets its own seeded order, so no two candidates sit an identical sequence.
This is how you get parallel forms without maintaining separate exam documents: the content and difficulty are constant, the ordering varies per attempt, and the seed makes each candidate's experience internally consistent. Answer keys passed between candidates lose most of their value when position no longer identifies the question.
Step 4: Set strict, disclosed time limits
Set a total time limit derived from timed pilot runs, not guesswork — use "Preview as candidate" to time a full attempt yourself and add measured headroom. For a certification, the limit must be identical for every candidate in a cohort; comparability is part of what makes the result meaningful.
Add per-page time limits where later material could reveal earlier answers or where timed performance is part of the standard. Remember the mechanics: when a page's time expires the candidate is moved on and the page locks. State the total limit, the per-page limits, and the locking behavior explicitly in the candidate instructions.
Step 5: Configure proctoring proportional to the stakes
In the delivery settings, escalate security to match what the certificate is worth. For a genuine credential, that typically means fullscreen lockdown with violation limits, copy/paste blocking, second-screen detection — which blocks the start if an extra display is detected and flags mid-exam connections — webcam recording, screen recording, and an ID check. Safe Exam Browser is available where you control the machines.
Configure the lockdown violation limits deliberately: maximum violations, seconds allowed per leave, and total seconds outside fullscreen. Exceeding a cap force-submits the attempt with an audit event — it is a configured limit the candidate was told about, not a verdict. A human still reviews the attempt and its evidence before any decision about the result is made.
- Fullscreen lockdown with explicit violation caps (count, seconds per leave, total seconds)
- Copy/paste blocking and second-screen detection
- Webcam and screen recording for review, not automated judgment
- ID check to bind the attempt to the person
- Safe Exam Browser for fully managed environments
Step 6: Be transparent at the consent gate
Every candidate passes a consent and preflight gate before starting, and it states exactly what is monitored: recordings, lockdown rules, violation limits, and identity verification. Do not soften or bury this. Candidates who know the rules can consent meaningfully, set up their environment properly, and cannot later claim they were monitored covertly.
The preflight also protects your data quality — camera, microphone, and display checks happen before the clock starts, so technical failures surface as fixable setup problems rather than mid-exam integrity flags. A transparent gate is not a courtesy on top of a defensible exam; it is part of what makes the exam defensible.
Step 7: Score every attempt against the same standard
For open-response items, attach rubrics before delivery — each rubric freezes into a snapshot when attached, so every candidate in the cohort is measured against literally the same criteria even if you refine the rubric for the next cycle. Use anchored level descriptions so different evaluators land on the same score for the same performance.
AI can draft summaries and criterion-level score suggestions with the model and version logged, but a named human confirms every score. Apply the pass criterion you documented in step one mechanically and identically to every attempt. Any integrity flags from the delivery are reviewed by a human alongside the substantive work — a flag is context for a decision, never the decision itself.
Step 8: Export the audit pack as the record
For each certified — or denied — candidate, the audit pack export assembles the defensible record: what they were asked, what they submitted, rubric snapshots and confirmed scores, evaluator identities, integrity events with their configured thresholds, and AI model and version logs for any suggestions used. This is what you produce when a result is appealed or a regulator asks how the certificate was earned.
Pair it with the candidate evidence report — printable, PDF-ready, and white-label — for the candidate-facing side. Archive the audit pack per cohort as a matter of routine, not only when challenged; a record assembled at decision time is worth far more than one reconstructed a year later.
Pro tips
- Pilot the full exam with internal staff under real delivery settings before the first cohort — proctoring misconfigurations only show up under proctoring.
- Size violation limits for humans, not ideals: a notification banner or an accidental swipe should consume budget, not end an attempt.
- Keep the item pool growing between cohorts and retire items that leak or underperform; a static pool decays.
- Document the pass criterion and section weights in your program records, not just in the platform — appeals are answered from documents.
- Never let a force-submit stand as a rejection on its own; a human reviews the attempt, the recordings, and the audit events before any outcome is final.