All guides

How do you decide between two final candidates whose scores are close?

Two finalists, two close scores. The question is not who scored higher — it is whether the gap is bigger than the test's error. The standard error of the difference answers it, and here is the arithmetic.

Last updated

The short answer

When two final candidates' scores are close, the question is whether the gap exceeds the test's error, not who scored higher. The statistic is the standard error of the difference: √2 times the standard error of measurement. If the gap is under two of those, the test has not separated them and the decision belongs to other evidence.

The statistic almost nobody names, and why it is not the one you think

Most advice on this question offers tiebreaker questions — ask about a failure, ask what they would do in the first ninety days. That is interview advice given to a measurement problem. If the two finalists sat the same assessment and came out three points apart, there is a prior question: can this instrument tell those two people apart at all?

There are two different error statistics and they answer two different questions. The standard error of measurement puts an interval around one person's score. The standard error of the difference puts an interval around the gap between two people's scores. Maury Buster's paper on applied statistical banding, presented to the International Personnel Assessment Council in 2005 and read at source on 23 September 2026, states the split plainly: use the standard error of measurement to establish the "interval of likely/possible true scores around a given individual's score", and the standard error of the difference to "test the significance between two individuals' scores".

The relationship between them is fixed and simple. Both candidates' scores carry error, so the gap carries both errors: SED = √2 × SEM, which is about 1.41 times as wide. The comparison you actually want to make is therefore always a harder test to pass than the interval on either individual score, and that is the part that catches people out.

The arithmetic, on a worked example

The full band formula, as given in the IPAC banding material, is Band Width = C × SDx × (1 − rxx)^½ × 1.414, where rxx is the test's reliability, SDx is the standard deviation of scores, and C is the normal deviate for the confidence level you want — 1.96 for 95%. The 1.414 is the √2 that turns a one-person interval into a two-person one.

Take a scale with a standard deviation of 10 and a reported reliability of 0.85. AssessAll publishes no reliability coefficient for any of its own instruments, because none has yet passed its internal research-readiness gate; 0.85 is a figure typical of published full-length commercial forms, used here as an illustration and labelled as one. Then:

SEM = 10 × √(1 − 0.85) = 3.87 points. SED = 1.414 × 3.87 = 5.48 points. A 95% interval on the difference = 1.96 × 5.48 = 10.7 points.

So on that scale two finalists need to be more than about eleven points apart before the instrument has said anything. A candidate on 71 and a candidate on 64 are seven points apart and are not distinguishable. Sorting the column and taking the top name has manufactured a distinction the data cannot support — which is the argument banding exists to make.

Notice what that number is next to the scale it sits on: 10.7 is larger than the standard deviation of the scale itself. That is not an artefact of the example. It is the strongest published objection to this whole method, and it is in the next section rather than buried.

The mistake that looks careful: comparing the two confidence intervals

The sophisticated-looking move is to put a confidence interval around each candidate's score and check whether the two intervals overlap. It feels rigorous. It is the wrong test, and it fails in the direction that costs you a real distinction.

Run it on the same numbers. Each candidate's own 95% interval is ±1.96 × 3.87 = ±7.6 points, so the two intervals stop overlapping only when the candidates are more than 15.2 points apart. The correct test — the standard error of the difference — declares a real gap at 10.7. Two finalists 12 points apart are genuinely different on the instrument and the overlap check will tell you they are the same.

The overlap-of-intervals test is roughly 40% more conservative than the test you want. If you are going to do the arithmetic at all, do this one.

What to do when the test has not separated them

The finding is not "we cannot decide". It is "this instrument has finished contributing, and the decision is now being made on something else — so choose what that something else is deliberately."

Use a second, independent, job-related source of evidence: a work sample scored against a rubric written before anyone sat it, a structured interview with fixed questions and anchored ratings, or a reference on a specific named behaviour rather than a general impression. Independent is the operative word — a second rater watching the same interview is not a second source.

Be explicit that the rule you use inside the tie is itself a decision with consequences. In the United States, 42 U.S.C. §2000e-2(l) makes it unlawful to "adjust the scores of, use different cutoff scores for, or otherwise alter the results of, employment related tests on the basis of race, color, religion, sex, or national origin" — so a within-tie rule that refers to a protected characteristic is within the words of that provision whatever the intent. A rule that decides on independent job-related evidence is not doing the thing the provision names. This is United States federal law, it is not legal advice, and other jurisdictions treat the question differently.

And write the rule down before you look at the two names. A tie-break rule invented after the shortlist exists is an unexamined preference with arithmetic in front of it.

Four published objections to the method on this page

One. The band is enormous. Frank Schmidt, writing in the ten-question banding symposium in Personnel Psychology 54 (Campion et al., 2001), at p. 154: "If the test distribution is approximately normal and if we take the highest score as being the one at the 99.9 percentile, the resulting 95% SED band will contain 38% of the score range and about 25% of the job applicants." A rule that declares a quarter of your applicant pool mutually indistinguishable is not a small adjustment to how you shortlist.

Two. It can erase differences that are real. John Kehoe, in the same symposium at p. 158: "the 1.96 SDdiff bandwidth is frequently equal to or larger than the standard deviation of observed scores on the measurement procedure (SD). This means, for example, that the Cascio et al. (1991) procedure will frequently conclude that true score differences as large as one SD do not exist." The worked example above is exactly that case — 10.7 against an SD of 10 — which is why it was left in rather than tuned.

Three. Which reliability you pick changes the answer. Reliability is not one number. Buster's 2005 IPAC case study reports the same examination returning a test–retest reliability of 0.64 and an internal-consistency reliability of 0.84, producing materially different band widths on identical scores. An author who wants a wider band can get one by choosing the lower coefficient, and nothing in the formula stops them. Ask which coefficient was used and why, every time — see Cronbach's alpha and test–retest reliability for why the two measure different things.

Four. The underlying SEM formula has been challenged. "Corrections to Twenty Years of Banding: The Necessity of Precision" (East Carolina University) argues that "the current procedure for calculating the standard error of measurement (SEM), which is used when calculating SED, is erroneous", following Dudek (1979), and that regression to the mean across administrations has been overlooked in the banding literature. Applying the corrections changes selection outcomes. This is a live methodological dispute, not a settled recipe.

And one thing this does not buy you. Charles Igou's 2011 IPAC paper reports that "statistical banding without minority preference does not reduce adverse impact or increase diversity", and that because "candidates generally don't understand statistical banding", organisations adopting it should expect more appeals and challenges rather than fewer. Banding is a measurement-honesty tool. It is not a diversity intervention, and selling it as one is how it ends up in court. If adverse impact is the question you are actually asking, the adverse impact ratio and the four-fifths rule are where to start instead.

The three questions to ask before any of this works — including of AssessAll

Every number on this page needs two inputs from the vendor: the standard deviation of the scale, and a reliability coefficient with its type named. Without both, you cannot compute a standard error of the difference and there is no honest way to say whether two finalists differ.

So ask for three things in writing. What is the reliability coefficient for this specific form, and is it internal consistency, test–retest or interrater? What is the standard deviation of the scale? And what is the standard error of measurement in score points — which is the same information in the form you can actually use at a decision boundary.

Apply that to this site too. AssessAll publishes no reliability coefficient, no standard error of measurement and no norm group for any of its own instruments today, because none has yet cleared its internal research-readiness gate. That means the arithmetic on this page cannot currently be run on an AssessAll score either. It is stated here rather than omitted, because a page arguing that buyers should demand these numbers is worth nothing if it quietly exempts the vendor publishing it. What AssessAll does report on behavioural results is a confidence band rather than a bare point score, and the validity checks widen that band rather than auto-rejecting anyone — which is the right shape, and is not a substitute for the coefficient.

A vendor who will not give you a standard error is asking you to treat a three-point gap as real on their word. That is the whole question this page is about.

Sources

Campion, M. A., Outtz, J. L., Zedeck, S., Schmidt, F. L., Kehoe, J. F., Murphy, K. R., & Guion, R. M. (2001). The controversy over score banding in personnel selection: Answers to 10 key questions. Personnel Psychology, 54(1), 149–185. Quotations above are from Schmidt at p. 154 and Kehoe at p. 158; read in the open full text on 23 September 2026.

Cascio, W. F., Outtz, J., Zedeck, S., & Goldstein, I. L. (1991). Statistical implications of six methods of test score use in personnel selection. Human Performance, 4(4), 233–264 — the source of the two-SED band width referred to throughout the symposium above. Abstract read at the publisher; the full text was not opened, and nothing is claimed here about its contents beyond what the 2001 symposium quotes.

Buster, M. A. (2005). Applied issues in statistical banding. International Personnel Assessment Council conference paper, read at annex.ipacweb.org on 23 September 2026 — source of SEM = σ√(1 − r), SED = √2 SEM, the individual-versus-two-individuals distinction, and the 0.64 / 0.84 same-exam case study.

Igou, C. (2011). To band or not to band: Is that the question? International Personnel Assessment Council conference paper, read at annex.ipacweb.org on 23 September 2026 — source of the band-width formula with the 1.414 factor and of the adverse-impact and candidate-comprehension findings.

Corrections to Twenty Years of Banding: The Necessity of Precision, East Carolina University (thescholarship.ecu.edu), abstract read 23 September 2026 — source of the Dudek (1979) SEM objection and the regression-to-the-mean point.

42 U.S.C. §2000e-2(l), added by section 106 of the Civil Rights Act of 1991. Read at source on 16 September 2026.

Frequently asked questions

How big does the gap between two candidates have to be before it is real?

Larger than about two standard errors of the difference. The standard error of the difference is √2 times the standard error of measurement, and the standard error of measurement is the scale's standard deviation times the square root of one minus the reliability. On a scale with a standard deviation of 10 and a reliability of 0.85 that works out to about 11 points, so two finalists seven points apart are not distinguishable by that instrument.

Can I just check whether the two candidates' confidence intervals overlap?

No — that is the wrong test and it errs in the expensive direction. On a scale with a standard deviation of 10 and a reliability of 0.85, each candidate's own 95% interval is about ±7.6 points, so the intervals stop overlapping only at a gap of 15.2 points, while the correct test declares a real difference at 10.7. The overlap check will tell you that two genuinely different finalists are the same.

What should decide it when the test cannot separate two finalists?

A second, independent, job-related source of evidence chosen and written down before you look at the two names: a work sample scored against a rubric that existed beforehand, a structured interview with fixed questions and anchored ratings, or a reference on a specific named behaviour. A second rater watching the same interview is not an independent source. In the United States, 42 U.S.C. §2000e-2(l) makes a tie-break rule that refers to race, colour, religion, sex or national origin unlawful; this is not legal advice and other jurisdictions differ.

Does AssessAll publish the numbers needed to run this calculation?

No. AssessAll publishes no reliability coefficient, no standard error of measurement and no norm group for any of its own instruments, because none has yet cleared its internal research-readiness gate. The arithmetic on this page therefore cannot currently be run on an AssessAll score. AssessAll's behavioural reports do carry a confidence band rather than a bare point score, and validity checks widen that band rather than rejecting anyone automatically, but a band is not a substitute for a published coefficient.

Is treating close scores as equivalent the same as banding, and is banding safe to use?

It is the same underlying reasoning, and banding is contested rather than settled. In the Personnel Psychology symposium on banding (Campion et al., 2001), Frank Schmidt notes at p. 154 that a 95% band can contain about 25% of applicants, and John Kehoe notes at p. 158 that the 1.96 band is frequently as wide as the whole score standard deviation, so real one-SD differences get declared non-existent. A 2011 IPAC paper also reports that statistical banding without minority preference does not reduce adverse impact or increase diversity. Use the arithmetic to know what your instrument can support; do not adopt formal banding as a fairness programme without advice.

Turn a validity coefficient into the number that matters
If the test cannot separate two finalists, the next question is how much your selection process is worth at all. The selection utility calculator answers it from your base rate and selection ratio.
Open the selection utility calculator