Predictive validity is the degree to which a hiring method's scores actually forecast how someone will perform on the job, expressed as a correlation between selection scores and later job performance. It is the single most useful number in hiring — and in 2022 the accepted numbers changed. A major re-analysis by Sackett and colleagues found that decades of meta-analytic estimates had been inflated by a statistical overcorrection, and the revised rankings quietly rewrote what a good selection process should look like. Most hiring funnels haven't caught up.
The old hierarchy — and why it broke
For nearly 25 years, the reference point was Schmidt and Hunter's landmark 1998 meta-analysis of 85 years of selection research. Its headline: general mental ability (GMA) tests were the king of predictors, with a validity of about .51, with work samples (.54) and structured interviews (.51) close by. That study shaped textbooks, vendor pitches, and thousands of hiring processes built around cognitive testing.
Then Sackett, Zhang, Berry, and Lievens re-examined the underlying corrections. Meta-analyses adjust raw correlations for range restriction — the fact that you only observe job performance for people you actually hired, who are a narrower slice than the applicant pool. The re-analysis showed those adjustments had been applied too aggressively and too uniformly, systematically overstating validity, with cognitive ability inflated most of all.
The revised rankings
The corrected estimates reorder the leaderboard. Approximate mean validities from the revision:
- Structured interviews: .42 — now at the top
- Job knowledge tests: .40
- Empirically keyed biodata: .38
- Work sample tests: .33
- Cognitive ability tests: .31 — down from .51
- Integrity tests: .31
- Conscientiousness: .19
- Unstructured interviews: .19
- Reference checks: .13
- Years of education: .10; years of experience: .09
Three shifts matter more than any single number.
1. Structure beats format
The gap between structured and unstructured interviews (.42 vs .19) is now the most dramatic contrast in the table. Same conversation, same candidate, same interviewer — but fixed job-related questions, anchored rating scales, and consistent scoring roughly double the predictive power. The lesson generalises beyond interviews: the value is not in the format (interview, test, exercise) but in the standardisation. An improvised "tell me about yourself" chat and a coin flip are closer than most hiring managers would like to believe.
2. Job-specific signal beats general traits
The methods that held their value — structured interviews, job knowledge tests, work samples — all measure something close to the job itself. The methods that fell furthest — GMA, broad personality traits — measure general characteristics and hope they transfer. That doesn't make cognitive ability useless; .31 is still meaningful, and the researchers stress wide variability across roles. But "smart and conscientious" is no longer a defensible substitute for "can demonstrably do the work."
3. No single method is enough
Even the best single predictor leaves most of the variance in job performance unexplained (a .42 correlation explains about 18% of it). The practical implication, which the revision's authors themselves emphasise, is combination: two or three moderately valid, weakly correlated methods — say a job knowledge or scenario assessment plus a structured interview — stacked in sequence, with scores combined mechanically rather than by gut feel.
What this means for your funnel
Put a structured assessment before the interview, not instead of it. Interviews are expensive — 30 to 60 minutes of a manager's time per candidate — so their .42 validity is only affordable deep in the funnel. Job knowledge tests, scenario-based judgment questions, and work-sample-style exercises deliver comparable signal upstream at a fraction of the cost per candidate. Screening on CV keywords and pedigree (education: .10, experience: .09) filters out exactly the candidates the evidence says you should be testing.
Standardise the interview you keep. Write questions from a job analysis, use behaviourally anchored rating scales, score independently before discussing. This is the cheapest validity upgrade available in all of hiring — it costs process discipline, not money.
Prefer demonstration over description. Wherever feasible, replace "have you done X?" with "do X now." Classic work samples were expensive to run at scale, which is why they were historically reserved for late stages. AI-graded scenario responses and simulations have largely dissolved that constraint — platforms like AssessAll can score open-ended, job-realistic scenario answers automatically and, with AI proctoring that assigns High/Medium/Low integrity bands, preserve the trust in results that remote testing otherwise loses.
Combine mechanically. Decades of evidence show that mechanically combining scores (weighted sums, fixed rules) beats holistic judgment. Decide the weights before you see candidates, then follow them.
The stakes: what a mis-hire actually costs
If the validity decimals feel academic, price the error instead. The US Department of Labor's long-standing conservative floor puts the cost of a bad hire at 30% of first-year earnings; SHRM's benchmarks run 50–75% of annual salary for entry-level roles and 100–150% for mid-level technical and managerial hires; and around three-quarters of employers admit to having made a bad hire. Moving from an unstructured process (~.19) to a stacked, structured one (combined validities in the .5+ range) doesn't shave percentage points off a metric — it materially changes how often you pay those bills.
For volume hiring, the arithmetic gets sharper. A BPO or campus drive screening thousands of applicants multiplies both the cost of weak methods and the return on strong ones. This is also where pay-as-you-go economics change what's practical: at roughly ₹30 (US$0.50) per assessment credit on AssessAll — with 100 free credits for individuals and 250 for companies to start — running a validated job knowledge or scenario assessment on every applicant costs less than the coffee served at one interview loop.
Reading the table honestly
Two caveats keep this evidence-based rather than evidence-flavoured. First, the revised figures are averages with wide credibility intervals — structured interview validity ranges from roughly .18 to .66 across studies. Context, role complexity, and execution quality matter enormously; a badly built "structured" interview inherits none of the .42. Second, rankings are relative, not gospel: the revision's authors urge designers to weigh validity alongside subgroup differences, candidate reactions, cost, and testing volume. A slightly less valid method that candidates complete willingly and that shrinks adverse impact can be the better system choice.
The takeaway
The best available evidence now says hiring accuracy comes from structure and job-relevance, layered in combination — not from any single silver-bullet test or a brilliant interviewer's instinct. Audit your funnel against the revised rankings: every unstructured conversation, CV filter, and pedigree proxy you replace with a standardised, job-specific measure is a measurable reduction in the mis-hires you'll be paying for next year.