How much does a hiring test actually improve who you hire?
A validity coefficient on its own tells you nothing you can act on. What a selection procedure buys depends on three numbers together: its validity, your base rate (the share of applicants who would succeed anyway) and your selection ratio (the share you hire). The same test can lift success rates by twenty points or by six, depending on the other two.
1. How many of the people you hire will work out?
The Taylor–Russell model, published in 1939 and still the correct answer to the question. Nothing you type is sent anywhere — the whole calculation runs in your browser.
- Success rate among hires
- 71.5%up from 50% with no test
- Extra hires who work out
- +21.5of 100 hired, per round
- Share of the maximum gain
- 43%a perfect predictor reaches 100%
71.5% of your hires would work out, against 50% with no instrument at all. That is 21.5 more people per round of 100 hires, and 43% of everything a perfect predictor could achieve at this selection ratio.
This figure assumes you hire strictly top-down and that everyone you offer accepts. Every candidate above the cut who declines, drops out or is passed over pulls the realised success rate back toward 50%. The second assumption is the one nobody states: Scullen and Meyer showed in 2014 that because applicants apply to several employers at once, small pools lose their best candidates to other offers, and the expected quality of new hires is materially lower than a selection-ratio model predicts. Read 71.5% as a ceiling, not a forecast.
2. What is that worth?
The Brogden–Cronbach–Gleser model, using the applicants, hires and validity from panel 1. Currency is whatever you type — the arithmetic does not care.
- Gain per hire, per year
- 9,793r 0.31 × SDy 18,000 × 1.75
- Net over the cohort
- 2.93m2.94m gross less 8,000 testing
- Validity needed to break even
- 0.00you are using 0.31
- That gain, as a share of salary
- 22%per hire, every year — your credibility check
- Criterion gain (Brogden)
- 0.54 SD31% of the maximum — exactly r
Net 2.93m over 3.0 years. The break-even validity is 0.00; you are using 0.31. Because the cost falls on everyone tested and the gain only on those hired, this number is far more sensitive to your selection ratio than to the price per test.
Read that 22%-of-salary figure before you put this in a paper. The model says each hire is worth 9,793 more per year on a salary of 45,000. Utility estimates routinely come out this large, and the best-known experiment on what happens next found that showing managers a positive utility analysis made them less willing to adopt a valid selection procedure than the validity evidence alone did. Present the success-rate figure from panel 1 to a decision-maker; keep this panel for sizing your own options.
SDy is the weakest number in this calculation and it is doing the most work. Schmidt and Hunter give 40% of salary as a lower bound on the standard deviation of output value, not a typical figure, and every estimate of it in the literature is contested. Halve it and the gain halves. Treat the output as an order of magnitude, and never quote it to two significant figures in a business case.
A validity coefficient is not an answer to any business question
Ask an assessment vendor how good their instrument is and you will be given a number between zero and one. It is a real number, it is usually honestly derived, and on its own it tells you nothing you can act on. It is a correlation. Nobody hires a correlation.
The same coefficient produces wildly different outcomes depending on two facts about your situation that the vendor does not know. Take a validity of 0.31, the current published estimate for a cognitive ability test, against a base rate of 50% — half your applicants would work out if you hired at random:
| If you hire… | Selection ratio | Success rate among hires | Lift |
|---|---|---|---|
| 1 applicant in 20 | 0.05 | 74.8% | +24.8 pp |
| 1 applicant in 10 | 0.10 | 71.5% | +21.5 pp |
| 1 applicant in 2 | 0.50 | 60.0% | +10.0 pp |
| 7 applicants in 10 | 0.70 | 56.2% | +6.2 pp |
One instrument, one coefficient, four answers ranging from transformative to marginal. And note which rows are which. Vendor examples are habitually drawn at selection ratios of 0.05 and 0.10; Schmidt and Hunter observed in 1998 that actual selection ratios are typically in the 0.30 to 0.70 range, which is the bottom half of that table.
This is not an argument against testing. A ten-point lift on a thousand hires is a hundred people, and few interventions available to an HR function move a hundred people. It is an argument against quoting a coefficient as though it were a result, which is what almost every page in this market does.
Three numbers, and the one you can actually move
The base rate is the share of your applicants who would perform satisfactorily if you hired them at random. It sets the headroom. At a base rate of 80% there are at most twenty points for any instrument to capture, and a validity of 0.31 captures about thirteen of them at an aggressive selection ratio. If your base rate is already high, selection is not your problem and no assessment will make it your solution.
The selection ratio is how many you hire divided by how many apply. It is a multiplier on validity, because an instrument can only sort people you are in a position to reject. At a selection ratio of 1.0 the most valid test ever built changes nothing at all. This is also the number most organisations can move fastest — widening the top of the funnel usually beats improving the instrument, and it is rarely the option on the table because nobody sells it.
The validity coefficient is the one everybody argues about and the one with the least room in it. The published estimates for whole categories of instrument sit between 0.19 and 0.42. Choosing well inside that band matters; it matters less than either of the other two numbers.
A useful discipline: before evaluating any instrument, write down your base rate and your selection ratio and compute the ceiling. If the best possible predictor could only buy you four points, stop the procurement.
The validity numbers, and the two revisions the market has not caught up with
The presets in the calculator are the estimates from Sackett, Zhang, Berry and Lievens (2022), who re-analysed the evidence base and found that corrections for restriction of range had been applied systematically too aggressively for twenty-five years. The revision reordered the field.
| Predictor | Schmidt & Hunter 1998 | Sackett et al. 2022 | SD of ρ |
|---|---|---|---|
| Structured interviews | .51 | .42 | .19 |
| Job knowledge tests | .48 | .40 | .13 |
| Empirically keyed biodata | .35 | .38 | .09 |
| Work sample tests | .54 | .33 | .09 |
| Cognitive ability (GMA) | .51 | .31 | .14 |
| Integrity tests | .41 | .31 | .20 |
| Assessment centres | .37 | .29 | .09 |
| Conscientiousness | .31 | .21 | .15 |
| Unstructured interviews | .38 | .19 | .16 |
| Years of job experience | .18 | .07 | .11 |
The first thing the market has not absorbed is that cognitive ability is no longer top of the list. Structured interviewing predicts better than any test in the table. Any vendor page still quoting 0.51 for an ability test in 2026 is quoting a superseded figure, and several do.
The second is newer still and almost nobody has it. Sackett and colleagues published a follow-up in 2023 collecting estimates that had appeared since, and it revises cognitive ability again — downward, to 0.23, on the basis of a meta-analysis of 113 twenty-first-century studies, which drops it from fifth to twelfth among the predictors examined. Assessment centres move the other way, up to 0.33. Both figures are in the calculator’s preset list, marked. Run the model at 0.31 and then at 0.23: on a 1-in-10 selection ratio and a 50% base rate the success rate falls from 71.5% to 66.0%, which is the difference between five and a half extra good hires per hundred.
Third, and least discussed: read the SD column. These are means of distributions, not constants. Structured interviews have an SD of ρ of 0.19 and integrity tests 0.20, which means the honest way to use this calculator is to run it twice — once at the mean and once a standard deviation below — and to make the decision that survives both.
The money model, and the experiment that says be careful who you show it to
The second panel implements Brogden–Cronbach–Gleser, the standard utility model. The gain per person hired per year is the validity coefficient multiplied by SDy — the standard deviation of the value of an employee’s output — multiplied by the mean standardised predictor score of the people you actually hired. Costs are subtracted across everyone tested, not everyone hired, which is why the selection ratio drives the answer harder than the price per test does.
Two things about that model deserve to be on the page rather than in a footnote.
SDy is a guess wearing a decimal point. The conventional estimate is 40% of salary, from Schmidt and Hunter, and it is worth reading what they actually wrote: that figure is described as a minimum— “at minimum 40% of the mean salary of the job” — and “a lower bound value”, with actual values “typically considerably higher”. It is not a typical value and it is not measured. Halve it and every number in the panel halves.
The output is usually too large to be believed, and that is an empirical finding rather than an opinion. Latham and Whyte gave 143 experienced managers either a validity-based case for adopting a selection procedure, or the same case plus a utility analysis showing substantial net financial benefits. The utility analysis reduced managerial support for adopting the procedure. Whyte and Latham ran it again in 1997 with an internationally recognised expert explaining the logic on video and then answering questions live, and the effect survived that too.
The expert in question, Steven Cronshaw, published a reply in the same issue that is the most useful thing anyone has written about how to use these numbers. He accepts the effect is real and substantial, and argues the moderator is perceived persuasive intent: a utility analysis read as an attempt to sell an intervention backfires, where the same analysis used at arm’s length to inform an investment decision may not. That is the reason this page ships a calculator you drive with your own numbers rather than a return-on-investment claim with ours in it — and the reason the honest thing to take to a decision-maker is the success-rate figure from the first panel, not the currency figure from the second.
What the model assumes, and where it breaks
- The criterion is a line drawn through a continuum.Taylor–Russell needs performance sorted into satisfactory and unsatisfactory, which real performance is not. Where you draw that line changes the base rate, and the base rate changes the answer. Two people modelling the same job with different definitions of “worked out” will get different results and both will be right.
- Predictor and criterion are assumed jointly normal and linearly related. Reference implementations of the model, including Niels Waller’s TaylorRussell package for R, integrate a multivariate normal distribution to produce the tables. Where the relationship is not linear — a threshold effect, say — the tables mislead.
- You hire strictly top-down, and everyone you offer accepts.The second half of that is almost never stated and is the assumption most likely to fail. Scullen and Meyer argued in 2014 that the standard models rest on the “inherent, but dubious, assumption that all job seekers in a given applicant pool are pursuing that particular job and no other jobs”, and that once candidates apply to several employers at once, small pools lose their strongest people to competing offers — producing significant negative effects on expected new-hire quality that the selection-ratio models do not capture. Treat the first panel’s output as a ceiling.
- The two panels measure the gain differently and do not agree. The first reports the share of the maximum possible improvement in the success rate; the second reports the gain on a continuous criterion, where — this is Brogden’s 1946 result — the share of the maximum possible gain is exactly the validity coefficient itself. At a 10% selection ratio and a 50% base rate those are 43% and 31% of maximum respectively. Neither is wrong. Dichotomising a criterion changes what “the gain” means, and any page that reports one of these numbers without saying which is being imprecise.
- None of this is a fairness analysis.A procedure can be valid, utility-positive, and still produce a selection rate for one group below four-fifths of the highest group’s. The two questions are independent and both have to be answered. Run the adverse impact ratio on the same cut you modelled here.
What AssessAll can and cannot tell you here
The validity presets in this calculator are published meta-analytic estimates for categories of instrument. None of them is an AssessAll figure, and there is no AssessAll figure to offer. A criterion-validity coefficient requires scores collected before hire and performance measured afterwards, on a sample large enough and spread across enough employers to mean anything. AssessAll has not published such a study on any instrument, so no coefficient for an AssessAll assessment appears on this page or anywhere else on this site.
What the platform does publish is the specification, which is the part you can check before you have outcome data: the item count and time limit on every form, the competencies each one reports, and what the report does at the boundary. Those are the inputs to a sensible decision when a validity coefficient is not available — and asking a vendor for them is a fair test of whether they have them.
One practical consequence of the arithmetic above, and it applies to buying from us as much as from anyone: because cost falls on everyone tested and gain only on those hired, a single expensive instrument administered to a whole applicant pool is the worst shape a screen can have. A short, cheap stage that removes most of the pool before an expensive one runs changes the economics far more than a better instrument does. That is what a staged assessment blueprint is for, and it is the argument the calculator makes on its own if you halve the number of applicants tested at full price.
Frequently asked questions
What does a validity coefficient of 0.31 actually buy you?
It depends entirely on two other numbers. At a base rate of 50% — half your applicants would succeed if you hired at random — a test with a validity of 0.31 lifts the success rate among the people you hire to 71% if you hire 1 applicant in 10, to 60% if you hire 1 in 2, and to 56% if you hire 7 in 10. The same coefficient, the same test, three different answers. A coefficient quoted without a base rate and a selection ratio is not an answer to any business question.
What is the base rate in hiring?
The base rate is the proportion of your applicants who would perform satisfactorily if you hired them at random with no selection procedure at all. It is the floor any instrument has to beat, and it sets the ceiling on how much one can help: at a base rate of 80% there are at most 20 percentage points of headroom for the best possible predictor to capture, so a test bought to fix a hiring problem in that pool will disappoint however valid it is.
What is the selection ratio, and why does it matter more than validity?
The selection ratio is the number of people you hire divided by the number who apply. It matters more than validity because it multiplies it: a test can only sort people you are able to reject. At a selection ratio of 1.0 — you hire everyone who applies — no instrument on earth changes who you end up with. Schmidt and Hunter noted in 1998 that actual selection ratios are typically in the 0.30 to 0.70 range, which is far less favourable than the 0.05 and 0.10 columns that make vendor examples look impressive.
What are the Taylor–Russell tables?
A set of tables published by H. C. Taylor and J. T. Russell in the Journal of Applied Psychology in 1939 that convert a validity coefficient into the practical outcome a selection decision-maker cares about: the proportion of selected candidates who will be satisfactory, given the base rate and the selection ratio. They remain the standard answer to the question and are the model behind the first panel of this calculator.
Is cognitive ability still the best predictor of job performance?
No, and the correction is more recent than most vendor material. Schmidt and Hunter's 1998 figure of 0.51 for general mental ability was revised to 0.31 by Sackett, Zhang, Berry and Lievens in 2022, after they showed that corrections for range restriction had been applied systematically too aggressively. On those revised figures, structured interviews (0.42) and job knowledge tests (0.40) both predict better. A 2023 follow-up goes further, citing a meta-analysis of 113 twenty-first-century studies that puts cognitive ability at 0.23, which moves it from fifth to twelfth in the ranking of predictors.
How do you calculate the ROI of a pre-employment test?
The standard model is Brogden–Cronbach–Gleser: the gain per person hired, per year, is the validity coefficient multiplied by SDy (the standard deviation of the value of output, conventionally estimated at 40% of salary or more) multiplied by the mean standardised test score of the people you hired. Subtract the cost of testing everyone, not just those hired. The output is usually implausibly large, which is a known problem with the model rather than a reason to celebrate.
Why should you be careful showing a utility estimate to a manager?
Because the best-known experiment on the question found it backfires. Latham and Whyte gave 143 experienced managers either a validity argument for a selection procedure or the same argument plus a utility analysis showing substantial net financial benefits. The utility analysis reduced their support for adopting the procedure. Whyte and Latham repeated it in 1997 with an internationally recognised expert presenting the analysis on video and then live, and the effect held.
Does this calculator store the numbers I enter?
No. The entire calculation runs in your browser. Nothing you type is transmitted to AssessAll or to anyone else, there is no signup, and no result is saved.
Sources
Every item below was verified at a DOI record, a publisher-deposited abstract or the paper’s own text on 13 September 2026.
- Taylor, H. C., & Russell, J. T. (1939). The relationship of validity coefficients to the practical effectiveness of tests in selection: discussion and tables. Journal of Applied Psychology 23(5), 565–578. doi:10.1037/h0057079
- Brogden, H. E. (1946). On the interpretation of the correlation coefficient as a measure of predictive efficiency. Journal of Educational Psychology 37(2), 65–76 — the result that the validity coefficient is the share of the maximum possible gain on a continuous criterion.
- Brogden, H. E. (1949). When testing pays off. Personnel Psychology 2(2), 171–183 — the linear utility model.
- Cronbach, L. J., & Gleser, G. C. (1965). Psychological Tests and Personnel Decisions (2nd ed.). University of Illinois Press, Urbana.
- Schmidt, F. L., Hunter, J. E., McKenzie, R. C., & Muldrow, T. W. (1979). Impact of valid selection procedures on work-force productivity. Journal of Applied Psychology 64(6), 609–626.
- Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology. Psychological Bulletin 124(2), 262–274 — the 40%-of-salary lower bound on SDy and the 0.30–0.70 selection-ratio range, both p. 263.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: addressing systematic overcorrection for restriction of range. Journal of Applied Psychology 107(11), 2040–2068. doi:10.1037/apl0000994
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology 16, 283–300 (open access) — the source of the 0.23 and 0.33 updates. doi:10.1017/iop.2023.24
- Latham, G. P., & Whyte, G. (1994). The futility of utility analysis. Personnel Psychology 47(1), 31–46.
- Whyte, G., & Latham, G. (1997). The futility of utility analysis revisited: when even an expert fails. Personnel Psychology 50(3), 601–610.
- Cronshaw, S. F. (1997). Lo! The stimulus speaks: the insider’s view on Whyte and Latham’s “The futility of utility analysis”. Personnel Psychology 50(3), 611–615.
- Scullen, S. E., & Meyer, R. D. (2014). Pitfalls in modelling the effects of recruitment on selection outcomes. Journal of Management 40(6), 1675–1699 — on the offer-acceptance assumption.
- Waller, N. G. TaylorRussell(R package, University of Minnesota) — a reference implementation, used here to corroborate the model’s normality assumption and the 1939 pagination.
Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.
Last reviewed
This page describes measurement and decision models, not legal requirements. A procedure can be valid and utility-positive and still produce adverse impact; the two questions are independent. Where a selection decision has legal consequences, take advice from a qualified employment lawyer in your jurisdiction.
Related reading
- What is predictive validity? — and why 0.51 became 0.31
- What is validity?
- What is criterion validity?
- What is a cut score?
- Adverse impact ratio calculator
- How do you set a defensible pass mark?
- How many candidates before a score means anything?
- Skills assessment vs interviews — what the evidence says
- Assessment blueprints — what a whole screen should contain
- What a numerical reasoning score actually means
- All free tools