---
title: "Standard error of the difference"
description: "The standard error of the difference is the error attached to the gap between two scores rather than to either score alone. Because both scores carry measurement error, the gap carries both: on a single instrument it is √2 times the standard error of measurement, about 1.41 times wider. A gap under two of them is not a real difference."
canonical: https://www.assessall.com/guides/glossary/s/standard-error-of-the-difference
updated: 2026-09-25
source: AssessAll
---

# Standard error of the difference

_Also known as: Standard error of the difference between two scores, Standard error of a difference score, SEdiff, Sdiff._

The standard error of the difference is the error attached to the gap between two scores rather than to either score alone. Because both scores carry measurement error, the gap carries both: on a single instrument it is √2 times the standard error of measurement, about 1.41 times wider. A gap under two of them is not a real difference.

<!-- #two-error-statistics-two-different-questions -->
## Two error statistics, two different questions

The [standard error of measurement](https://www.assessall.com/guides/glossary/s/standard-error-of-measurement) answers a question about one person: how far this score would move if the same person sat the test again. The standard error of the difference answers a question about two people: how far the gap between them would move. They are not interchangeable, and the second is always the larger of the two.

The National Council on Measurement in Education's own instructional module on the standard error of measurement states the principle in one sentence — "An important principle to remember is that the difference between two test scores is less reliable than the two individual scores" — and gives the formula for two scores on the same test as √2 times the scale's standard deviation times the square root of one minus the reliability. That is exactly √2 times the standard error of measurement.

Maury Buster's paper on applied statistical banding for the International Personnel Assessment Council draws the same line in operational terms: use the standard error of measurement to establish the "interval of likely/possible true scores around a given individual's score", and the standard error of the difference to "test the significance between two individuals' scores". Running the individual interval on a two-candidate comparison is the most common misuse of either number.

<!-- #when-the-two-scores-come-from-different-tests-the-formula-is -->
## When the two scores come from different tests, the formula is not √2 × SEM

The 1.41 multiplier holds only when both scores come from the same instrument, with the same reliability. Comparing a candidate's numerical reasoning score against their verbal reasoning score, or one candidate's result on test A against another's on test B, is a different calculation, and the NCME module gives it: the scale's standard deviation times the square root of two minus the first test's reliability minus the second's.

That difference matters because the weaker of the two reliabilities does most of the damage. On a scale with a standard deviation of 10, two scores each at reliability 0.85 give a standard error of the difference of 5.48. Swap one of them for a test at 0.70 and it rises to 6.71 — so the gap two people need before the comparison means anything grows from about 10.7 points to about 13.1. **Comparing across two instruments is not the same test as comparing within one, and quoting a single "SEdiff" for a mixed battery is a category error.**

<!-- #which-reliability-coefficient-goes-into-it-and-the-field-doe -->
## Which reliability coefficient goes into it, and the field does not agree

Every version of the formula takes a reliability coefficient as its input, and almost every vendor publishes exactly one: [Cronbach's alpha](https://www.assessall.com/guides/glossary/c/cronbachs-alpha), an internal-consistency figure computed from a single sitting. Whether that is the right input is a live methodological question rather than a settled one, and the answer that is emerging is no.

Blampied's 2022 open-access review in *The Cognitive Behaviour Therapist* records the position: Jacobson and Truax, who introduced the statistic to clinical practice, "recommended that rxx should be used" — test–retest reliability — and **"McAleavey (2021) has argued that only test–retest reliability estimated over short inter-test intervals should be used, because it is the measure of reliability matching the pre–post repeated measures aspect of the data being analysed, and coefficient α should not be used."**

The consequence is arithmetical and it runs one way. The standard error of measurement is the scale's standard deviation times the square root of one minus the reliability, so **a lower reliability produces a wider error band**. Where a vendor publishes only alpha and alpha is the higher of their two figures, the standard error of the difference you compute from it is too narrow, and the gap you conclude is real may not be. Ask for the [test–retest](https://www.assessall.com/guides/glossary/t/test-retest-reliability) figure and the interval it was measured over, and run the arithmetic on whichever coefficient is lower.

<!-- #hiring-has-no-name-for-this-number-clinical-psychology-has-h -->
## Hiring has no name for this number. Clinical psychology has had one since 1991.

This is the same statistic that clinical and neuropsychology call Sdiff, the denominator of the Reliable Change Index. Jacobson and Truax defined it in the *Journal of Consulting and Clinical Psychology* in 1991 as the quantity that "describes the spread of the distribution of change scores that would be expected if no actual change had occurred", and attached a criterion to it: "An RC larger than 1.96 would be unlikely to occur (p < .05) without actual change."

That field has had a named statistic, a published threshold and thirty-five years of argument about how to estimate it. The hiring lane has the same arithmetic, the same instruments and no name for it — which is why score gaps between finalists get read as rankings. The threshold transfers directly: divide the gap by the standard error of the difference, and if the result is under 1.96 the instrument has not separated the two people.

It is worth being clear about what the criterion does and does not establish. Jacobson and Truax are explicit that it "tells us whether change reflects more than the fluctuations of an imprecise measuring instrument" — no more than that. A gap that clears 1.96 is a real difference on the instrument. Whether it is a difference that should decide anything is a question about the instrument's [validity](https://www.assessall.com/guides/glossary/v/validity), not about its error.

---

Source: https://www.assessall.com/guides/glossary/s/standard-error-of-the-difference
