What is Reliability?
Reliability is the consistency of a measurement: how much of the variation in scores is signal rather than noise. It is reported as a coefficient between 0 and 1, and it comes in several distinct kinds — internal consistency, test–retest, parallel forms and inter-rater — which answer different questions and are not interchangeable.
Which kind is being quoted
Internal consistency asks whether the items agree with each other in a single sitting. Test–retest asks whether the same people get the same result on a second occasion. Parallel forms asks whether two versions of the test rank people the same way. Inter-rater asks whether two humans scoring the same response agree.
A vendor who says 'reliability of 0.89' without saying which kind has not answered the question. For a behavioural questionnaire used to advise someone about their working style, test–retest is the figure that matters, and it is usually the lowest and the least often published.
Reliability caps validity
A test cannot correlate with an outcome more strongly than its own consistency allows; the theoretical ceiling on a validity coefficient is the square root of the product of the two measures' reliabilities. Unreliable measurement is therefore not merely untidy, it puts a hard limit on how useful the instrument can be.
The reverse does not hold. A perfectly reliable measure of the wrong thing is perfectly consistent and perfectly useless, which is why reliability is a precondition for validity and never a substitute for it.
Not the same as validity
Reliability is about consistency; validity is about whether the inference is justified. A bathroom scale that reads four kilograms heavy every time is perfectly reliable and invalid for the use it is put to.
ValidityRelated terms
Check a selection process against the four-fifths rule
Free, no signup, computed in your browser — with the remedy, not just the verdict.
Last reviewed