By Daniel Whitmore, MSc in Industrial and Organizational Psychology
In short
Item quality signals are statistics computed from real candidate responses that tell you whether a question is doing its job. The two that matter most are difficulty — the share of candidates who answer correctly — and discrimination — how well the question separates strong candidates from weak ones. A question almost everyone gets right or wrong carries little information, and one with low discrimination fails to sort candidates at all. Both are signals to check the answer key and revise the wording before the question shapes another decision.
A question that looks fine can still measure nothing
Most assessment questions are written once, reviewed for wording, and then trusted forever. That trust is misplaced — not because authors are careless, but because the quality of a question is not visible in its text. Two questions can read equally well and behave completely differently the moment candidates answer them: one cleanly separates the people who know the material from the people who do not, the other splits the room at random.
Item analysis is the discipline of judging questions by their behavior instead of their appearance. It comes from classical test theory, and for a work-sample bank you only need two of its measures to catch the questions that quietly hurt you: how hard the question is, and how well it discriminates. Both are computed from the answers your own candidates have already given, which is what makes them trustworthy — they describe your test, on your population, not a textbook ideal.
The point is not to chase a perfect number on every item. It is to surface the small set of questions that are actively distorting your results, so a human can look at them and decide whether to fix, keep, or retire each one. A good bank is not one where every question is flawless; it is one where the flawed questions are known.
Difficulty: too easy and too hard both cost you information
Difficulty, confusingly named, is just the proportion of candidates who answer a question correctly — a question answered right by 70% of people has a difficulty of 0.70. Its job is not to be high or low but to be informative, and the extremes carry almost no information. When 95% or more of candidates get a question right, it no longer separates anyone: strong and weak candidates both clear it, so the item adds length without adding signal. It may still belong at the start of an assessment as a warm-up, but it is not doing measurement work.
The opposite extreme is more dangerous. When only 20% or fewer answer correctly, the honest first suspicion is not that your candidates are weak — it is that the question is broken. A miskeyed answer, an ambiguous stem, two defensible options, or a task that tests something the role never needs will all drive the correct rate into the floor. A genuinely hard question that a fifth of a strong pool can still solve is legitimate and valuable; a question almost no one solves usually has a problem you can fix in a minute once you go looking.
So difficulty is best read as a pair of guardrails, not a target. Very easy questions are candidates for retirement or repositioning; very hard questions are candidates for a key-and-wording check. Everything in between is doing its job, and does not need your attention.
Discrimination: does the question tell strong from weak?
Discrimination is the more powerful of the two signals and the one teams most often miss. It asks a single question: do the candidates who do well on this item also tend to do well on the assessment overall? When the answer is yes, the item is pulling in the same direction as the rest of the test — it discriminates. When there is no relationship, the item is measuring something else, or nothing, no matter how sensible it reads.
Concretely, discrimination is the correlation between getting one item right and the candidate's total score. High values mean strong candidates get it and weak candidates miss it, which is exactly what you want. A value near zero means the item is not sorting candidates at all. A negative value is the alarm worth acting on immediately: it means your stronger candidates are more likely to get the item wrong — the classic fingerprint of a miskeyed answer or a trick that punishes the people who understand the material most deeply.
A widely used rule of thumb treats discrimination below about 0.20 as poor — the item is barely separating strong from weak and is a candidate for revision. But discrimination is only meaningful once enough people have answered; on three responses it is noise. That is why a responsible flag waits for a minimum sample before it trusts the number, and why a fresh question is never condemned on its first few candidates.
These signals are earned, not guessed
The honest limitation of item analysis is that it cannot tell you anything the day you write a question. Difficulty and discrimination are properties of responses, so a brand-new item has no health at all until candidates have answered it enough times to say something real. This is a feature, not a gap: it keeps the signals grounded in evidence instead of opinion, and it protects a new question from being judged before it has had a fair run.
It also means item quality is a maintenance habit, not a one-time gate. A bank that was clean a year ago drifts as the role changes, as the candidate pool shifts, and as questions get reused across assessments. The questions worth watching are the ones with real usage and enough responses to trust — which, in practice, is a small, changing subset of the bank rather than the whole thing. You are not auditing everything; you are reading the handful of items the data has flagged.
How SkillCort surfaces this: the 'Needs attention' flag
SkillCort computes item health from real candidate responses across every assessment a question appears in, and turns the two signals into a single, plain filter on your item bank: Needs attention. An item lands there when the data shows one of three concrete problems — low discrimination once at least five people have answered, a difficulty at or above 95% (very easy), or a difficulty at or below 20% (very hard). Nothing is flagged on a hunch, and nothing is flagged before it has the responses to justify it.
Each flagged item carries its reason inline, as a chip next to the question: Low discrimination, Very easy, or Very hard, with the underlying numbers on hover. So the bank does not just tell you a question is weak — it tells you why, which is what turns a warning into a next action. An item bank analytics view rolls the same signals up across the whole bank, adding which items have enough data to judge and which have never been used, so you can see the health of your measurement at a glance rather than one question at a time.
The design rule behind all of this is the same one that governs the rest of the platform: the system surfaces the signal and the reason, and a human makes the call. SkillCort will tell you a question is very hard and show you that only 12% answered it correctly; it will not silently delete the question or overrule your judgment about whether that difficulty is a problem or the point.
What to do when a question is flagged
Start with the answer key, especially for very-hard and negative-discrimination items. A surprising share of alarming numbers resolve the moment you re-read the keyed answer and realize a second option is also correct, or that the intended answer is wrong. This single check fixes more flagged items than any amount of rewriting.
If the key is right, read the item as a confused candidate would. Very hard and low-discrimination questions often share a root cause: ambiguous wording, two defensible answers, or a stem that quietly tests something outside the skill you meant to measure. Tightening the language so only one reading survives is usually enough to move the numbers. Very easy items are the gentler case — keep them as warm-ups if they serve that purpose, or retire them so the assessment spends its time where measurement actually happens.
Finally, resist the urge to over-react to small samples. A question flagged on the strength of a handful of responses is a question to watch, not to condemn; give it more candidates before you act. Item quality is a slow, cumulative read, and the goal is a bank whose weak questions are known and managed — not a bank scrubbed of every imperfect number.
Key takeaways
- A question's quality lives in how candidates answer it, not in how it reads — two equally clean questions can behave completely differently.
- Difficulty is a pair of guardrails: very easy items (≥95% correct) add no signal; very hard items (≤20% correct) usually point to a miskeyed or ambiguous question.
- Discrimination — whether strong candidates get the item and weak ones miss it — is the sharpest signal; below ~0.20 is weak, and a negative value almost always means a miskeyed answer.
- These statistics are earned from real responses, so new questions have no health until enough candidates answer; a responsible flag waits for a minimum sample.
- SkillCort's 'Needs attention' filter surfaces flagged items with the reason inline (Low discrimination / Very easy / Very hard) — the system shows the signal, a human decides the fix.