Reliability is whether a test gives the same answer twice. Validity is whether that answer means what you claim it means. A hiring assessment can be highly reliable and still predict nothing about job performance — consistency is a precondition for validity, never a substitute for it. Reliability is measured on the scores; validity is argued about the inferences you draw from them.
That distinction is the single most useful thing an HR team can learn about assessment, and it is routinely collapsed. Vendors quote a reliability coefficient because it is easy to compute from one administration. Validity requires evidence about the job, the candidates, and the decision — which is harder, slower, and the only thing a regulator or a tribunal will actually ask about.
The two questions, side by side
| | Reliability | Validity | |---|---|---| | Question answered | Would this score repeat? | Does this score support the decision I am making? | | What it is a property of | The test and the group taking it | The interpretation, not the test | | Typical evidence | Internal consistency (α, ω), test–retest, inter-rater agreement, standard error of measurement | Content, response process, internal structure, relations to other variables, consequences | | Can be high while the other is low? | Yes — a reliable test of the wrong thing | No — an unreliable test cannot be valid | | Who defines the standard | Classical test theory / IRT | AERA/APA/NCME *Standards* (2014); SIOP *Principles* (2018) |
A bathroom scale that reads 4 kg heavy every time is perfectly reliable and completely invalid for the inference "this is my weight." Most bad assessments fail exactly this way — not by producing noise, but by producing stable, confident, irrelevant numbers.
Reliability sets a mathematical ceiling on validity
The two are not independent. Under classical test theory, the maximum correlation you can observe between a predictor and a criterion is the square root of the product of their reliabilities — the correction for attenuation, a result that dates to Spearman in 1904.
Work the numbers. A test with internal consistency of 0.85, validated against supervisor performance ratings whose inter-rater reliability is 0.60, cannot correlate above 0.71 with that criterion — no matter how well designed it is. And 0.60 is not a pessimistic figure: it is the value Berry, Lievens, Zhang and Sackett (2024) use as the standard correction for supervisor ratings of overall job performance in the updated personnel selection meta-analytic matrix.
This is why published operational validities look modest. In that same updated matrix, structured interviews sit at .42, biodata at .38, general mental ability and integrity tests at .31, and situational judgment tests at .26. Those are corrected for range restriction and criterion unreliability but not for predictor unreliability. A vendor quoting a validity above .60 against a supervisor-rating criterion is either measuring something else or correcting more aggressively than the field now accepts.
Reliability is not a property of your test
This is the part that surprises practitioners: the same test can report wildly different reliability depending on who sits it.
Tighe, McManus, Dewhurst, Chis and Mucklow (2010), analysing MRCP(UK) examinations, documented the paradox directly. Part 2 of the exam showed lower reliability than Part 1 (mean 0.802 versus 0.907) while having a smaller standard error of measurement (2.77% versus 3.20%) — that is, it measured each candidate more precisely while scoring worse on the headline coefficient. The reason is range restriction: Part 2 candidates have already passed Part 1, so the spread of ability is narrower, and reliability coefficients shrink when spread shrinks.
Their Monte Carlo simulation makes the point unarguable. Restricting the sample to passing candidates only, reliability fell from 0.897 to 0.704 — while SEM stayed essentially flat at around 3.1%. The authors' conclusion: reliability "is not a property of an assessment, but a joint property" of the test and the candidate group.
The practical implication for hiring: a reliability figure quoted without the population it was computed on is close to meaningless. An assessment showing α = 0.91 on a general applicant pool may show 0.70 on your shortlisted final ten — the exact group where you need precision most.
Use the standard error of measurement, not the coefficient
SEM converts reliability into the unit that matters: how wrong an individual score might be. The formula is SEM = SD × √(1 − reliability).
Take a test with a candidate standard deviation of 12 points and internal consistency of 0.85. SEM is 4.65 points. A 95% confidence band is roughly ±9 points. So a candidate scoring 72 has a true score plausibly anywhere between 63 and 81 — and the "clear winner" who scored 76 is statistically indistinguishable from them.
Three rules follow:
- Never rank-order candidates on gaps smaller than one SEM. Treat them as tied and let structured interview or work-sample evidence break it.
- Set score bands, not point cut-offs. A single number pretending to three-digit precision invites decisions the measurement cannot support.
- Report the band to hiring managers. They will otherwise read 72 as meaningfully better than 70. On AssessAll, integrity signals are reported the same way — as bands rather than a binary cheat/no-cheat verdict — precisely because a probabilistic signal shown as a certainty gets misused.
One further caution: Cronbach's alpha assumes all items contribute equally to the construct (tau-equivalence). When they do not — and in mixed-format assessments they rarely do — alpha understates reliability, which is why McDonald's omega is increasingly preferred in psychometric reporting.
What validity evidence actually looks like
The 2014 Standards frame validity as a single unified concept supported by five sources of evidence: test content, response processes, internal structure, relations to other variables, and consequences of testing. The older shorthand of "three types of validity" is retired; the EEOC's Uniform Guidelines (29 CFR 1607) still recognise criterion-related, content and construct strategies, but they are routes to one argument, not three separate qualities.
For a hiring assessment, a defensible minimum looks like this:
- A documented job analysis linking each competency measured to work actually performed.
- Content-linkage evidence: subject-matter experts rating each item's job relevance.
- Adverse impact monitoring by group, with the four-fifths rule as a screening flag rather than a legal safe harbour.
- A criterion study where sample size permits, or documented content validity where it does not.
- A record of who set the cut score and by what method.
Custom-built assessments — including the AI-graded scenario sets we build for clients at AssessAll — are worth this documentation trail at design time, because retrofitting a validity argument after a rejected candidate complains is considerably harder than writing it down in advance.
When each check matters more
If you are screening high volume with a wide ability range, reliability coefficients will look flattering and tell you little; watch SEM and adverse impact instead. If you are assessing a narrow, pre-filtered pool — final-round candidates, promotion panels, certification retakes — expect reliability to look poor and judge the instrument on SEM and content evidence. And if the assessment drives a legally consequential decision, no reliability figure of any size substitutes for a documented job-relevance argument.
The takeaway: reliability tells you how much of a score is signal; validity tells you whether the signal is the one you need. Ask any vendor for both — the population the reliability was computed on, and the evidence supporting the specific inference you intend to make.