All articles
Assessment Science10 October 2026·6 min read

Work Sample Validity Fell From .54 to .33. The Test Didn't Change - the Correction Did.

Sackett et al. (2022) found range-restriction corrections had inflated selection validity estimates by .10-.20 points. What the correction does, where the overcorrection came from, the revised numbers, and six questions to ask before accepting a validity coefficient.

By AssessAll Editorial

A validity coefficient is not a measurement. It is an observed correlation that has been statistically corrected upward to estimate what the relationship would have been in an unrestricted applicant pool. The size of that correction is a methodological choice, and in 2022 a re-analysis of the selection literature concluded the field had been making it too generously for decades.

The headline numbers moved a long way. Sackett, Zhang, Berry and Lievens (2022) revisited the meta-analyses behind Schmidt and Hunter's canonical 1998 table and found that mean validity estimates for widely used predictors dropped by .10 to .20 points. Work sample tests fell from .54 to .33. Cognitive ability tests fell from .51 to .31. No test changed. The correction did.

What the correction is actually doing

Range restriction is the problem that you can only observe job performance for people you hired. If your applicants span the full ability distribution but your incumbents are the top slice of it, the correlation you compute inside that slice understates the relationship in the full pool. Correcting for it means estimating the applicant-pool standard deviation and scaling the observed correlation back up.

The scaling factor is called the U ratio: unrestricted predictor SD divided by restricted predictor SD. It is not a small knob. In the working paper version of the analysis, the authors show that selecting the top 50% directly on a predictor restricts its SD to .60, giving a U ratio of 1.67 — a 67% increase in the observed validity. Get the U ratio wrong and you have not made a rounding error. You have changed the conclusion.

Where the overcorrection came from

The 2022 argument is not that the formulas are wrong. It is that they were applied to the wrong studies.

Direct range restriction happens only when incumbents were selected on the predictor being validated, and nothing else. Sackett et al. call that case "exceedingly rare." This is the scenario where corrections are large and legitimate.

Indirect range restriction happens when people were hired on something else — an interview, a degree screen, a manager's judgement — and the predictor under study merely correlates with it. Here the effect is far smaller. The paper shows that when the predictor correlates .50 or less with whatever actually drove selection, and the selection ratio is anything above 10%, the restricted SD stays at .90 or higher. The correction factor is generally under 10%.

That matters because most validation studies are concurrent: the test is given to current employees who were never selected on it. The paper tallies the concurrent share of the meta-analyses that fed the 1998 table — 98% of the studies in Roth et al.'s (2005) work sample meta-analysis, 95% in McDaniel et al.'s (2007) situational judgment test meta-analysis, 76% in Ones et al.'s (1993) integrity test meta-analysis. U ratios estimated from the minority of predictive studies were then applied across the whole database. In the integrity case, the correction lifted the mean correlation from .29 to .34 across 655 studies, three-quarters of which had little restriction to correct for.

The revised estimates

Operational validity for predicting job performance, 1998 estimate first, 2022 estimate second:

  • Structured employment interviews — .51, now .42 (highest ranked)
  • Job knowledge tests — .48, now .40
  • Empirically keyed biodata — .35, now .38 (one of the few that rose)
  • Work sample tests — .54, now .33
  • Cognitive ability tests — .51, now .31
  • Integrity tests — .41, now .31
  • Assessment centres — .37, now .29
  • Situational judgment tests — .26 (knowledge and behavioural-tendency variants alike)
  • Unstructured employment interviews — .38, now .19
  • Conscientiousness, overall measures — .31, now .19
  • Years of job experience — .18, now .07

The ranking survived better than the levels. Most of what was near the top is still near the top. But the spread between methods narrowed, and cognitive ability — the field's long-standing focal predictor — is no longer at the front of the queue.

What did not change

Two things are easy to over-read here.

First, selection still works. In the 2023 follow-up, Revisiting the design of selection systems, the same authors report that for composites of one to five predictors, mean validity was .51 under the old estimates and .47 under the revised ones. Their own framing: it would be incorrect to read this work as challenging the value of selection practice.

Second, the weighting changed more than the toolkit. In a six-predictor composite, dropping cognitive ability to zero weight cost .20 of validity under the old numbers (.66 to .46) and .05 under the revised ones (.61 to .56). Cognitive ability's standardised weight falls from .40 to .23. Given that it also carries the largest Black–White mean difference in the table (d = .79, against .23 for structured interviews and .10 for integrity tests), that reweighting changes the validity–diversity arithmetic materially.

The argument is not settled

Treat this as a live methodological dispute, not a closed verdict. Oh and colleagues published a direct challenge in 2023 to the recommendation against correcting concurrent studies; Sackett and colleagues replied in A response to speculations about concurrent validities in selection. Anyone quoting either set of numbers as settled fact is overstating the state of the evidence.

Six questions to ask before you accept a validity number

  1. Is this corrected or uncorrected? If the vendor cannot say, you do not have a number you can interpret.
  2. Corrected for what? Criterion unreliability, range restriction, or both. Corrections compound.
  3. Where did the U ratio come from? An assumed value, a hiring rate, or an actual applicant-pool SD. Sackett et al. are explicit that treating a firm's overall hire rate as the selection ratio for one measure is wrong.
  4. Predictive or concurrent design? A concurrent study with a large restriction correction deserves scepticism.
  5. What is the credibility interval? Structured interviews average .42 with an 80% interval of .18 to .66. Empirically keyed biodata average .38 with a lower bound of .26 — a lower mean but a safer floor, which may matter more if you are risk-averse.
  6. What criterion was predicted? Supervisor ratings, training performance and turnover are not interchangeable outcomes.

What to do in your own funnel

The practical conclusion is unglamorous: your own data beats anyone's meta-analysis. The Uniform Guidelines and the Standards for Educational and Psychological Testing both ask for validity evidence appropriate to the inference you are making, in the context you are making it. An uncorrected local correlation between your assessment scores and your own performance data — even a modest one — tells you more about your funnel than a corrected field-wide average.

That means keeping scores and outcomes joinable from the start. Assessment data is worth more when it carries its conditions with it: which form the candidate took, when, and under what administration conditions. AssessAll records an integrity band alongside each score rather than folding it into the score, so a later local validation can filter on it instead of guessing. Where a role has no off-the-shelf analogue, a custom assessment built against your own job analysis gives you a cleaner criterion to validate against.

Takeaway: When someone quotes you a validity coefficient, the number is downstream of a correction decision you did not see. Ask what that decision was — and run an uncorrected local check against your own outcomes, because that is the one number nobody adjusted.

#validity#range-restriction#meta-analysis#psychometrics#selection-science#criterion-validity

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free