What is Standard error of the difference?
Also called Standard error of the difference between two scores, Standard error of a difference score, SEdiff, Sdiff
The standard error of the difference is the error attached to the gap between two scores rather than to either score alone. Because both scores carry measurement error, the gap carries both: on a single instrument it is √2 times the standard error of measurement, about 1.41 times wider. A gap under two of them is not a real difference.
Two error statistics, two different questions
The standard error of measurement answers a question about one person: how far this score would move if the same person sat the test again. The standard error of the difference answers a question about two people: how far the gap between them would move. They are not interchangeable, and the second is always the larger of the two.
The National Council on Measurement in Education's own instructional module on the standard error of measurement states the principle in one sentence — "An important principle to remember is that the difference between two test scores is less reliable than the two individual scores" — and gives the formula for two scores on the same test as √2 times the scale's standard deviation times the square root of one minus the reliability. That is exactly √2 times the standard error of measurement.
Maury Buster's paper on applied statistical banding for the International Personnel Assessment Council draws the same line in operational terms: use the standard error of measurement to establish the "interval of likely/possible true scores around a given individual's score", and the standard error of the difference to "test the significance between two individuals' scores". Running the individual interval on a two-candidate comparison is the most common misuse of either number.
When the two scores come from different tests, the formula is not √2 × SEM
The 1.41 multiplier holds only when both scores come from the same instrument, with the same reliability. Comparing a candidate's numerical reasoning score against their verbal reasoning score, or one candidate's result on test A against another's on test B, is a different calculation, and the NCME module gives it: the scale's standard deviation times the square root of two minus the first test's reliability minus the second's.
That difference matters because the weaker of the two reliabilities does most of the damage. On a scale with a standard deviation of 10, two scores each at reliability 0.85 give a standard error of the difference of 5.48. Swap one of them for a test at 0.70 and it rises to 6.71 — so the gap two people need before the comparison means anything grows from about 10.7 points to about 13.1. Comparing across two instruments is not the same test as comparing within one, and quoting a single "SEdiff" for a mixed battery is a category error.
Which reliability coefficient goes into it, and the field does not agree
Every version of the formula takes a reliability coefficient as its input, and almost every vendor publishes exactly one: Cronbach's alpha, an internal-consistency figure computed from a single sitting. Whether that is the right input is a live methodological question rather than a settled one, and the answer that is emerging is no.
Blampied's 2022 open-access review in The Cognitive Behaviour Therapist records the position: Jacobson and Truax, who introduced the statistic to clinical practice, "recommended that rxx should be used" — test–retest reliability — and "McAleavey (2021) has argued that only test–retest reliability estimated over short inter-test intervals should be used, because it is the measure of reliability matching the pre–post repeated measures aspect of the data being analysed, and coefficient α should not be used."
The consequence is arithmetical and it runs one way. The standard error of measurement is the scale's standard deviation times the square root of one minus the reliability, so a lower reliability produces a wider error band. Where a vendor publishes only alpha and alpha is the higher of their two figures, the standard error of the difference you compute from it is too narrow, and the gap you conclude is real may not be. Ask for the test–retest figure and the interval it was measured over, and run the arithmetic on whichever coefficient is lower.
Hiring has no name for this number. Clinical psychology has had one since 1991.
This is the same statistic that clinical and neuropsychology call Sdiff, the denominator of the Reliable Change Index. Jacobson and Truax defined it in the Journal of Consulting and Clinical Psychology in 1991 as the quantity that "describes the spread of the distribution of change scores that would be expected if no actual change had occurred", and attached a criterion to it: "An RC larger than 1.96 would be unlikely to occur (p < .05) without actual change."
That field has had a named statistic, a published threshold and thirty-five years of argument about how to estimate it. The hiring lane has the same arithmetic, the same instruments and no name for it — which is why score gaps between finalists get read as rankings. The threshold transfers directly: divide the gap by the standard error of the difference, and if the result is under 1.96 the instrument has not separated the two people.
It is worth being clear about what the criterion does and does not establish. Jacobson and Truax are explicit that it "tells us whether change reflects more than the fluctuations of an imprecise measuring instrument" — no more than that. A gap that clears 1.96 is a real difference on the instrument. Whether it is a difference that should decide anything is a question about the instrument's validity, not about its error.
The smallest gap that counts, at four reliabilities
- Scale with a standard deviation of 10 — the scaling most score reports use.
- Reliability 0.95 → SEM 2.24, SED 3.16 → smallest real gap (1.96 × SED): 6.2 points
- Reliability 0.90 → SEM 3.16, SED 4.47 → smallest real gap: 8.8 points
- Reliability 0.85 → SEM 3.87, SED 5.48 → smallest real gap: 10.7 points
- Reliability 0.70 → SEM 5.48, SED 7.75 → smallest real gap: 15.2 points
Dropping the reliability from 0.95 to 0.70 more than doubles the gap two candidates need before the instrument has said anything about which of them is stronger. Two figures decide whether a shortlist is a ranking or a tie, and a vendor who publishes neither the scale's standard deviation nor a reliability coefficient with its type named has made the calculation impossible.
Not the same as standard error of measurement
The standard error of measurement puts a band around one person's score; the standard error of the difference puts a band around the gap between two people's scores and is 1.41 times wider on a single instrument. Checking whether two candidates' individual bands overlap is a different and more conservative test: on a scale with a standard deviation of 10 at reliability 0.85 the overlap check only declares a difference at 15.2 points, where the correct test declares one at 10.7. It will call real differences ties.
Standard error of measurementSources
- Harvill, L. M. (1991). Standard Error of Measurement. NCME Instructional Topics in Educational Measurement, Module 9 — open PDF, read at source 25 September 2026; equations 9 and 10 give the same-test and different-test forms
- Jacobson, N. S., & Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12–19 — free full text at a university repository, read at source 25 September 2026
- Blampied, N. M. (2022). Reliable change and the reliable change index: still useful after all these years? The Cognitive Behaviour Therapist, 15, e50 — open access, read at source 25 September 2026; source of the McAleavey (2021) position on which reliability coefficient to use
- Buster, M. A. (2005). Applied Issues in Statistical Banding. International Personnel Assessment Council
Read next
Related terms
More terms beginning with S
Check a selection process against the four-fifths rule
Free, no signup, computed in your browser — with the remedy, not just the verdict.
Last reviewed