What does Sackett et al. (2022) actually say about selection test validity?
A 2022 re-analysis found that meta-analyses of hiring tests had corrected for range restriction too aggressively, overstating how well those tests predict job performance. Correcting the error cut most published validity estimates by .10 to .20 and moved structured interviews to the top of the ranking, ahead of cognitive ability tests.
Sackett, P. R., Zhang, C., Berry, C. M. & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
Primary source opened and quotes confirmed on .
In its own words
“Key findings are that most of the same selection procedures that ranked high in prior summaries remain high in rank, but with mean validity estimates reduced by .10–.20 points. Structured interviews emerged as the top-ranked selection procedure.”
“We conclude that our selection procedures remain useful, but selection predictor–criterion relationships are considerably lower than previously thought.”
What it does not say
Each of these is a claim made in this market and attributed to the source above. None of them is supported by it.
Commonly claimed: That assessments do not work, or that testing has been debunked.
The abstract's own closing sentence is that selection procedures remain useful. The finding is about the size of the coefficients, not about whether the instruments predict anything.
Commonly claimed: That the ranking of predictors was overturned.
The same sentence says most procedures that ranked high before remain high in rank. What changed at the top is that structured interviews overtook cognitive ability; most of the order below that is stable.
Commonly claimed: That the revised figures are what you would observe in your own hiring data.
They are still corrected operational validities, corrected less aggressively. An observed correlation in one employer's pipeline will normally be lower again, because that pipeline has already been filtered.
The figures
Revised operational validity estimates, as reported by SIOP's own publication
| Predictor | Revised validity |
|---|---|
| Structured interviews | .42 |
| Job knowledge tests | .40 |
| Empirically keyed biodata | .38 |
| Work sample tests | .33 |
| Cognitive ability tests | .31 |
| Interests (fit-based) | .24, up from .10 in the 1998 table |
These per-predictor figures are not in the paper's abstract, and the article itself is paywalled. The values above are as reported in TIP (Trends and Issues in Psychology) 60(3), the magazine of the Society for Industrial and Organizational Psychology, in an article by Patrick Gavan O'Shea and Adrienne Fox Luscombe of HumRRO — read at source on 16 September 2026. Where a figure matters to a decision, buy or borrow the paper rather than trusting any secondary table, this one included.
What the error was
Validity coefficients are routinely corrected upwards because the people in a validation study are not a random slice of applicants — they have usually already been hired, which narrows the spread of scores and shrinks the correlation. The correction is legitimate. The size of the correction depends on an assumption about how much narrowing occurred, and meta-analyses carried that assumption in an artifact distribution.
The 2022 paper's contribution is to show that those distributions were built in ways that systematically overstated the narrowing, principally by applying restriction estimates derived from predictive designs to concurrent studies, where selection on the measure had not happened at all. The abstract describes and critiques five such approaches and concludes that each has significant issues that often result in substantial overcorrection.
Why it matters commercially rather than academically
Almost every validity number circulating in assessment marketing traces back to the meta-analytic tradition this paper corrects. A vendor quoting a headline coefficient from before 2022 is quoting a figure its own field has since revised downward, and the revision is not small: .10 to .20 on a scale where the whole useful range is roughly .1 to .5.
The practical consequence is not to stop testing. It is that a business case built on the older numbers is overstated by a knowable amount, and that the comparison between instruments has changed at the top — a well-run structured interview is now the benchmark a test has to beat rather than the soft complement to it.
Where this source is used here
These pages argue from the source above. If it is ever superseded, these are the pages that have to change.
Read next
- Schmidt & Hunter (1998) — the table everyone still quotes — The 1998 meta-analysis behind almost every validity figure in assessment marketing — and the reason quoting it in 2026 is a dated claim.
- Sackett et al. (2023) — what the revision means for a screen — The follow-up that prices the decision: dropping cognitive ability from a six-predictor composite costs .05 of validity, not .20.
Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.
Last reviewed