What is Test–retest reliability?
Also called Temporal stability
Test–retest reliability is the correlation between the scores of the same people measured twice with an interval between sittings. It is the most intuitive form of reliability and the least forgiving, because it exposes both measurement noise and genuine change in the person, and the two are difficult to separate.
Why the interval has to be reported
A retest after two days is inflated by memory of the previous answers. A retest after two years is deflated by real development in the person. Neither is wrong, but a coefficient without its interval cannot be interpreted, and the interval is often omitted from marketing material precisely because the short ones look better.
Ask for the interval, the sample size, and the population. A stability figure from a student sample over three weeks says little about working adults over a year.
The type-flip problem
Instruments that convert a continuous score into a category — a type, a letter, a colour — have a reliability problem their continuous scores do not. If a scale is cut at its midpoint, someone who scores near that midpoint will flip categories on a retest even when their underlying score barely moves.
This is why a questionnaire can report respectable continuous stability and still hand a substantial share of people a different type a few weeks later. Both figures can be true at once, and only the second is what the person experiences.
Not the same as internal consistency
Internal consistency is measured inside one sitting and describes whether the items hang together. Test–retest is measured across two sittings and describes whether the result holds. A scale can score well on the first and poorly on the second.
Cronbach's alphaRead next
Related terms
Check a selection process against the four-fifths rule
Free, no signup, computed in your browser — with the remedy, not just the verdict.
Last reviewed