Free tools

How many candidates before a score means anything?

Sample size in assessment answers two separate questions. How many sittings a pass rate needs before it is stable is a question about a group, and it is governed by the width of a confidence interval. How much a single candidate's score can be trusted is a question about precision, and it is governed by the standard error of measurement. A large sample never makes an unreliable score precise.

1. How stable is this pass rate?

A question about the cohort. Nothing you type is sent anywhere — the whole calculation runs in your browser.

Observed pass rate
35.0%42 of 120
95% interval (Wilson)
27.143.9%± 8.4 percentage points
Sittings for ± 5.0 pp
350230 more than you have

Not yet precise enough for your target. You are at ± 8.4 pp and asked for ± 5.0 pp. The remedy is 230 more completed sittings, assuming the rate stays near 35.0%. If you cannot assume that, size for the worst case — 385 sittings, which is what a 50% rate needs.

2. How precise is one candidate’s score?

A question about a person, and the one the panel above cannot answer. Reliability and score SD come from your instrument’s technical manual; if the vendor will not give you both, that is itself the finding.

Standard error (SEM)
± 5.0912.0 × √(1 − 0.82)
95% band on this score
53.073.0observed 63.0, cut 60.0
Chance they are on the other side
27.8%truly below the cut

This score cannot separate pass from fail. The 95% band runs from 53.0 to 73.0 and the cut sits inside it, so roughly 28 in 100 candidates scoring 63.0 have a true score on the other side of 60.0. The remedy is not a second look at this score. It is more evidence — another instrument, a structured interview, a work sample — or a wider borderline band that routes these candidates to a human.

The second number, which stops the first being misread.Regressing for unreliability — Kelley’s estimate, mean + r × (score − mean) — puts this candidate’s best true-score estimate at 61.6, not 63.0. An unreliable score is systematically too far from the mean, and the correction always pulls toward it.

Group-level reliability, individual-level use. α = 0.82 is comfortably above the 0.70 that gets quoted as “acceptable”, but that figure was Nunnally’s bar for early-stage research. His bar for decisions about individuals was 0.90 minimum, 0.95 desirable — which is why the band above is as wide as it is.

The remedy, in items. By the Spearman–Brown prophecy formula, reaching α = 0.90 from α = 0.82 needs the test to be 1.98× longer — about 60 items instead of 30, assuming the added items are of similar quality. Lengthening only works if you are adding items that measure the same thing; adding a different construct raises the count and lowers the reliability.

What it would take to call this one. For a 3.0-point margin to clear the 95% band at this SD, the SEM would have to fall to 1.53 — a reliability of about 0.984. If that is out of reach, the honest move is to stop treating scores this close to the bar as decidable.

Everything above is computed in your browser. No signup, no analytics on these inputs, nothing stored, nothing transmitted.

Two questions that look like one

“How many do we need?” is asked in assessment programmes about two different things, and the answers move in different directions.

The group question. Your pass rate, your mean score, your time-to-hire — these are estimates of something about a population, and every one of them comes with an interval that narrows as the sample grows. At a 35% pass rate you need roughly 350 completed sittings for a 95% interval of ±5 percentage points, about 90 for ±10, and about 30 for ±17. Collect four times as many and the interval halves, because it scales with the square root of n.

The person question.How much a single candidate’s score can be trusted is not a sample-size question at all. It is set by the instrument’s reliability and the spread of scores, through the standard error of measurement, and it does not improve by one point when the thousandth candidate sits the test.

Conflating the two is how a team ends up with a beautifully precise cohort dashboard sitting on top of individual decisions the measurement cannot support. The second panel of the calculator exists to keep the two apart.

Why the interval here is a Wilson interval

The formula most people were taught — p ± 1.96 × √(p(1−p)/n), the Wald interval — is unreliable at exactly the sample sizes assessment teams actually have. Its true coverage drops well below the nominal 95% for small n and for rates near 0 or 1, it can return bounds outside 0–100%, and when nobody passes it collapses to zero width, reporting perfect certainty from the least informative data possible.

Agresti and Coull set out the case against it in The American Statistician in 1998, under the title Approximate is better than “exact” for interval estimation of binomial proportions. The alternative they favour is the score interval Edwin Wilson published in the Journal of the American Statistical Association in 1927, which is what this calculator uses: it stays inside the range, it keeps close to its nominal coverage at small n, and it costs one extra line of arithmetic.

This matters more than it sounds. A team reporting “a 35% pass rate” on 40 sittings with a Wald interval understates its own uncertainty, and the number then gets quoted onward without the interval at all.

The standard error of measurement, and the number it hides

SEM = SD × √(1 − reliability). With a score standard deviation of 12 and a reliability of 0.82 — which is a respectable figure for a well-built commercial instrument — the SEM is about 5.1 points. A 95% band around any single observed score is therefore about ±10 points wide.

Now put a cut score three points below a candidate’s result. On those numbers the probability that their true score is actually below the bar is around 28%. Nothing about that candidate is borderline in the report; the number reads as a pass. The uncertainty is invisible unless someone computes it, which is the entire argument for printing a band rather than a point.

The calculator also returns Kelley’s regressed true-score estimate — cohort mean + reliability × (observed score − cohort mean). Because measurement error inflates extreme scores in both directions, the best estimate of a true score always sits closer to the mean than the observed one. It is an uncomfortable output: it says the top of your shortlist is probably not as far ahead as the ranking implies, and the bottom is probably not as far behind.

When the band does straddle the bar, lengthening the test is the honest remedy and the Spearman–Brown prophecy formula prices it — going from α = 0.82 to α = 0.90 needs roughly twice the items, and reaching 0.95 needs about four times as many. That is a real cost, and it is the right thing to weigh against widening a borderline band instead.

The 0.70 rule is a misquotation

“Cronbach’s alpha above 0.70 is acceptable” is the most repeated sentence in assessment procurement, and it is a misreading of its own source. In Psychometric Theory (2nd ed., 1978, pp. 245–246) Nunnally recommended 0.70 for the early stages of research, where the interest is in correlations and group differences. For applied settings where important decisions turn on an individual’s score he was explicit that a reliability of 0.90 is the minimum that should be tolerated and 0.95 the desirable standard.

Lance, Butts and Michels traced how this and three other cutoffs became folklore in The Sources of Four Commonly Reported Cutoff Criteria: What Did They Really Say? (Organizational Research Methods 9(2), 2006, 202–220). The pattern is always the same: a conditional recommendation gets cited, the condition falls away, and the number survives on its own.

Two practical consequences. First, ask a vendor which decision their reliability figure was computed for — an alpha of 0.78 is fine evidence for a cohort report and thin evidence for a rejection. Second, alpha is a lower bound on reliability only under assumptions most real tests violate; Klaas Sijtsma set out the case against relying on it in Psychometrikain 2009 (74(1), 107–120), and McDonald’s ω is the better figure to ask for when it is available.

What to do with the numbers this returns

  1. Decide the precision before you look at the rate.“We need to know the pass rate to within five points” is a decision about how the number will be used. Choosing it after seeing the interval is how targets get set to whatever the data happened to support.
  2. Report counts, not rates, under about 30.“9 of 26” is honest and legible. “34.6%” on the same data implies a precision that is not there.
  3. Suppress cuts below 50 rather than caveating them. A footnote does not stop a thin subgroup figure being screenshotted into a board pack. If you cannot report a role, a site or a demographic group at n ≥ 50, do not break it out.
  4. Print a band on every individual score you act on. If the band straddles the cut, the score has not made the decision and something else has to.
  5. Route the borderline to a human, deliberately. Set the width from the SEM rather than from a round number, and write down what the human is supposed to look at.
  6. Re-check after any change to the instrument. A new item set, a new rubric, a re-trained grader or a new delivery mode moves the distribution against a cut score that nobody re-examined. Treat it as a new instrument and a new sample.

Where these thresholds sit inside AssessAll

The same arithmetic decides what the platform will and will not say, and the numbers are constants in the codebase rather than positions in a brochure.

  • Borderline results go to a person. On AssessAll Certified, a composite within 3 points of the nearest level cut is flagged borderline and escalated to human review rather than being resolved by the arithmetic. That is precisely the case the second panel above describes.
  • A credential’s validity report is stamped DRAFT below n = 200. Under that many scored sittings the report says INSUFFICIENT SAMPLE and every band claim in it is labelled a design prior — a defensible starting point set by design, not a cut score derived from data.
  • Item calibration will not run below 200 responses per item, and the adaptive engine stops a section when the standard error on the ability estimate reaches 0.30 — a precision rule, not an item-count rule.
  • No benchmark is published at all. The internal gate requires at least 400 completed sittings on the single assessment being reported, from at least 10 distinct organisations, with no organisation contributing more than a quarter of the sample, and suppresses any reported cut below n = 50.

What that means in practice, plainly: AssessAll has not passed that gate on any instrument, so there is no AssessAll score benchmark, industry norm or national percentile — and any that appear here later will carry their sample size and their reliability on the same page. Where the platform reports a percentile inside a hiring drive it is a cohort percentile: the share of scored candidates in that same pipeline at or below this one. Not a national norm, and never presented as one.

Frequently asked questions

How many candidates do you need before a pass rate means anything?

It depends entirely on how precise you need the answer to be. At a 35% pass rate, roughly 350 completed sittings gets you a 95% interval of about ±5 percentage points; about 90 gets you ±10; about 30 gets you ±17, which is wider than most differences anyone would act on. There is no universal number — decide the precision you need first, then read the sample size off it.

Does a bigger sample make an individual score more accurate?

No, and this is the most common and most expensive misunderstanding in assessment analytics. Sample size governs how precisely you know a group statistic such as a pass rate or a mean. The precision of one person's score is governed by the standard error of measurement, which is the score's standard deviation multiplied by the square root of one minus the reliability. It does not change however many other people sit the test.

What is the standard error of measurement?

The SEM is the standard deviation of the errors around an individual's observed score: SEM = SD × √(1 − reliability). With a score SD of 12 and a reliability of 0.82, the SEM is about 5.1 points, so a 95% band around an observed score is roughly ±10 points. A candidate three points above a cut score is, on those numbers, only about 72% likely to be truly above it.

Is a Cronbach's alpha of 0.70 good enough?

For early-stage research, which is what Nunnally actually recommended it for. For decisions about individuals he wrote in Psychometric Theory (2nd ed., 1978, pp. 245–246) that a reliability of 0.90 is the minimum that should be tolerated and 0.95 the desirable standard. Lance, Butts and Michels documented in Organizational Research Methods in 2006 how the 0.70 figure came to be quoted as a universal standard it was never intended to be.

Why does this calculator use a Wilson interval rather than the usual formula?

Because the textbook Wald interval — p ± 1.96√(p(1−p)/n) — behaves badly at exactly the sample sizes assessment teams have. Its actual coverage falls well below the nominal 95% for small n and for rates near 0 or 1, where it can produce bounds outside the range or collapse to zero width. Agresti and Coull made the case in The American Statistician in 1998; the Wilson score interval, published in 1927, is well behaved in both cases and costs nothing extra to compute.

How many sittings before you can publish a benchmark?

More than a pass rate needs, because a benchmark is used to judge people outside the sample. AssessAll's own internal gate, in scripts/qa/research-readiness.ts, requires at least 400 completed sittings on the single assessment being reported, from at least 10 distinct organisations, with no organisation contributing more than 25% of the sample, and suppresses any reported cut below n = 50. The platform has not passed that gate on any instrument, which is why no AssessAll score benchmark is published.

Does this calculator store the numbers I enter?

No. The entire calculation runs in your browser. Nothing you type is transmitted to AssessAll or to anyone else, there is no signup, and no result is saved.

Sources

  • Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association 22(158), 209–212.
  • Agresti, A., & Coull, B. A. (1998). Approximate is better than “exact” for interval estimation of binomial proportions. The American Statistician 52(2), 119–126.
  • Nunnally, J. C. (1978). Psychometric Theory (2nd ed.), pp. 245–246. McGraw-Hill.
  • Lance, C. E., Butts, M. M., & Michels, L. C. (2006). The sources of four commonly reported cutoff criteria: what did they really say? Organizational Research Methods 9(2), 202–220.
  • Sijtsma, K. (2009). On the use, the misuse, and the very limited usefulness of Cronbach’s alpha. Psychometrika 74(1), 107–120.
  • Spearman, C. (1910) and Brown, W. (1910), British Journal of Psychology 3 — the prophecy formula.
  • AERA, APA & NCME (2014). Standards for Educational and Psychological Testing, chapter 2 (reliability and errors of measurement).
SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed

This page describes measurement methods, not legal requirements. Where a selection decision has legal consequences, take advice from a qualified employment lawyer in your jurisdiction.

Related reading