By Daniel Whitmore, MSc in Industrial and Organizational Psychology
In short
Weighting competencies means declaring that some parts of a role count for more than others when a candidate's score is combined. Decades of research show that equal weights rank candidates about as well as carefully derived differential weights, so weighting is not a way to make a score more predictive. Its value is job-relatedness: a weighted score reflects what the role actually is, which is what makes a hiring decision explainable and defensible when someone asks why one candidate was preferred.
Counting how many bars a candidate clears hides which ones
The simplest way to summarize a candidate against a role is to count: they met four of the seven competencies the role requires. It reads cleanly and it is easy to explain, which is exactly why so many scorecards stop there. The problem is what the count throws away — it treats every competency as interchangeable, so it cannot tell you whether the four they cleared were the four that matter.
Picture a support role where the work is overwhelmingly written: handling customers in writing, with empathy, all day. The role also lists product knowledge, prioritization and documentation, which are real but secondary. One candidate clears written communication and empathy and misses the other three. Another clears the three secondary ones and misses the two that define the job. By the count, the second candidate looks stronger — three beats two. By the job, the first candidate is the only one who can actually do the work.
Nothing about that outcome is a bug in the arithmetic; the arithmetic did exactly what it was told. It is a modeling failure: the model was never given the information that some parts of this role carry the job and others support it. Weighting is how you supply that information.
What the research actually says about weights
Here is where most vendor material gets uncomfortable, because the honest finding is not the flattering one. A long line of research shows that unit weights — treating every component equally — rank people about as well as weights derived with far more effort. Wainer titled his 1976 paper on estimating coefficients in linear models with the phrase that has followed the finding ever since: it don't make no nevermind. Dawes made the companion argument in 1979 about what he called improper linear models, showing that simple equal-weight composites are remarkably robust and often beat more elaborate schemes on new data.
Bobko, Roth and Buster returned to the question in 2007 specifically for the composites we build in selection, reviewing the literature and concluding that unit weights are a highly appropriate approach under a wide range of conditions. The reason is not mysterious. Components of a well-built assessment tend to correlate with each other, and once they do, shifting weight between them moves the resulting rank order far less than intuition suggests. Differential weights also carry their own estimation error, which can quietly cancel out whatever precision they were supposed to add.
So if a platform tells you that weighting your competencies will make its scores more accurate or more predictive, treat that claim carefully. On the evidence, that is not what weighting reliably buys. Saying so plainly costs a vendor a nice-sounding sentence, and it is the only version of the sentence that survives contact with the literature.
So why weight at all? Because the score has to match the job
The case for weighting is not prediction. It is content validity — the requirement that an assessment, and the score it produces, represent the job it claims to measure. A number that says a candidate is a strong fit for a writing-heavy role while quietly counting writing as one-sixth of the result is not describing that role. It may still rank candidates acceptably; it just cannot be explained without embarrassment.
That explainability is the practical payoff. Selection decisions get questioned — by a hiring manager who disagrees, by a candidate who asks why, sometimes by a lawyer. In those moments the useful artifact is a documented chain: this is the job analysis, these are the competencies it produced, this is how much each one counts and why, and here is the evidence for this candidate against them. A weighted composite makes that chain complete. An unweighted one leaves a gap where the reasoning about relative importance should be, and 'we counted them all the same because that was the default' is a weak answer when the job plainly does not work that way.
There is a second, quieter benefit. Setting weights forces a conversation that teams usually skip. Deciding that written communication counts three times what documentation counts requires the hiring manager and the recruiter to agree on what the role is — before candidates arrive, not while arguing about one. Most of the value of a weighting exercise is captured before a single score is computed.
Weights answer a different question than must-haves
Weighting is often confused with marking something essential, and the two are genuinely different instruments. A weight says how much a competency counts in a total where strength in one area can offset weakness in another — a compensatory model. A must-have says no amount of strength elsewhere makes up for a shortfall here — a hurdle. The literature on selection models treats these as complementary designs rather than rivals, and comparisons of compensatory and multiple-hurdle approaches turn on cost, sequencing and utility rather than one being correct.
The failure mode is using one where you need the other. If you express a genuine must-have as a heavy weight, a candidate can still clear the bar overall while being unable to do the non-negotiable part of the job. If you express a mere priority as a must-have, you start rejecting people for shortfalls the role can absorb — and once several competencies are flagged essential, almost everyone fails one, and the flag stops carrying information.
A practical rule: reach for a must-have when the honest answer to 'could we hire someone who is weak here?' is no, and reach for weight when the answer is 'yes, but it would cost us'. Keep must-haves scarce. Their power comes from being rare.
- Weight — how much this counts in a score where strengths can offset weaknesses.
- Must-have — a shortfall reported on its own, never averaged away by strength elsewhere.
- Use both: a weighted composite for the overall picture, hurdles for the non-negotiables.
What a missing measurement must never mean
Weighted models introduce a failure mode worth naming, because it punishes candidates for something they did not do. If a role lists a heavily weighted competency that the assessment never actually measured, a naive composite treats the absence as a zero — and the candidate is marked down for a question nobody asked them.
The correct handling is to exclude what was not assessed from both the numerator and the denominator, so the score describes the ground actually covered. A candidate should be measured against the evidence that exists, and the report should be explicit about how much of the role that evidence spans. That transparency also does something useful for the employer: a competency that keeps showing up as unmeasured is a gap in the assessment, and seeing it stated is what prompts someone to add the task that closes it.
This is the same principle that governs the rest of a defensible scoring model. Absence of evidence is not evidence of absence, and a scoring system that silently converts one into the other will eventually produce a decision nobody can justify.
Setting weights you can actually defend
Start at equal. Given what the research says about unit weights, equal weighting is a defensible default rather than a lazy one, and it should be what your blueprint does until someone can say why a competency deserves more. Deviating is the decision that needs a reason, not staying put.
Derive the differences from the job, not from the assessment. The question is how central a competency is to performing the role — how much of the work touches it, how visible the consequences are when it is weak — and that question is answered by the people who know the job, ideally with the same job-analysis input that produced the competency list. Weights invented to make a favored candidate look better are not weights; they are a result working backwards.
Keep the scale coarse and the reasoning written down. Fine-grained weights imply a precision the underlying judgments do not have, and the difference between 1.7 and 1.8 will not survive scrutiny — a small set of steps is enough to express what a team actually believes. Record why each non-default weight exists, so the model can be explained months later by someone who was not in the room.
- Default to equal weights; treat any deviation as a decision that needs a justification.
- Derive weights from the job analysis, before you see candidates.
- Use a coarse scale — a few steps, not decimals implying false precision.
- Write down the reason for every non-default weight.
- Revisit weights when the role changes; a stale blueprint measures a job you no longer have.
How SkillCort implements this
A role blueprint in SkillCort lists the competencies a role requires, each with a target level, a weight and an optional critical flag. Candidate evidence rolls up from the tasks they completed into those competencies, and the report shows two numbers side by side: how many of the assessed competencies cleared their target, and — only when the blueprint actually carries non-default weights — what share of the role's assessed weight was met. With equal weights the two are identical by construction, so nothing redundant is displayed and no existing number silently shifts when the feature arrives.
Critical competencies stay outside the composite. A shortfall there is reported as its own gap rather than being smoothed into an average, which keeps the compensatory summary and the non-negotiables in separate columns where they belong. Competencies with no evidence are excluded from both figures, so an unmeasured requirement can never drag a candidate down.
The claim attached to all of this is deliberately narrow. SkillCort does not tell you the weighted figure predicts job performance better than the unweighted one, because the evidence does not support that. It tells you the number reflects the role you defined — and pairs it with the blueprint, the evidence and the audit trail that let you show your reasoning when someone asks.
Key takeaways
- Counting how many competencies a candidate clears treats every part of a role as interchangeable, which can rank the wrong person first.
- Equal weights rank candidates about as well as differential weights (Wainer 1976; Dawes 1979; Bobko, Roth & Buster 2007) — weighting is not an accuracy upgrade.
- The real case for weighting is content validity: the score should describe the job, which is what makes a decision explainable.
- Weights and must-haves answer different questions — use a weighted composite for the overall picture and keep hurdles scarce for the non-negotiables.
- Never let an unmeasured competency count as a zero; exclude it and disclose the coverage.
Sources & further reading
- Wainer (1976), Estimating coefficients in linear models: It don't make no nevermind — Psychological Bulletin
- Dawes (1979), The robust beauty of improper linear models in decision making — American Psychologist
- Bobko, Roth & Buster (2007), The Usefulness of Unit Weights in Creating Composite Scores — Organizational Research Methods
- Ock (2018), The Utility of Personnel Selection Decisions: Comparing Compensatory and Multiple-Hurdle Selection Models — Journal of Personnel Psychology