What is Reliability?

Reliability is the consistency of a measurement: how much of the variation in scores is signal rather than noise. It is reported as a coefficient between 0 and 1, and it comes in several distinct kinds — internal consistency, test–retest, parallel forms and inter-rater — which answer different questions and are not interchangeable.

Which kind is being quoted

Internal consistency asks whether the items agree with each other in a single sitting. Test–retest asks whether the same people get the same result on a second occasion. Parallel forms asks whether two versions of the test rank people the same way. Inter-rater asks whether two humans scoring the same response agree.

A vendor who says 'reliability of 0.89' without saying which kind has not answered the question. For a behavioural questionnaire used to advise someone about their working style, test–retest is the figure that matters, and it is usually the lowest and the least often published.

Reliability caps validity

A test cannot correlate with an outcome more strongly than its own consistency allows; the theoretical ceiling on a validity coefficient is the square root of the product of the two measures' reliabilities. Unreliable measurement is therefore not merely untidy, it puts a hard limit on how useful the instrument can be.

The reverse does not hold. A perfectly reliable measure of the wrong thing is perfectly consistent and perfectly useless, which is why reliability is a precondition for validity and never a substitute for it.

Not the same as validity

Reliability is about consistency; validity is about whether the inference is justified. A bathroom scale that reads four kilograms heavy every time is perfectly reliable and invalid for the use it is put to.

Validity

Related terms

Check a selection process against the four-fifths rule

Free, no signup, computed in your browser — with the remedy, not just the verdict.

Open the calculator

Last reviewed