What is Interrater reliability?

Also called Inter-rater reliability, Rater agreement

Interrater reliability is the degree to which two independent raters judging the same person or the same piece of work arrive at the same score. It sets a ceiling on everything built from those ratings: a criterion no two raters agree on cannot be predicted well by any test, however good the test is.

The number, and the split inside it that matters

Supervisory ratings of job performance are the criterion most selection research is validated against, and their interrater reliability has been cumulated twice. Viswesvaran, Ones and Schmidt reported 0.52 for overall job performance in 1996. Salgado and Moscoso, cumulating 219 independent coefficients over a total sample of 43,203 people in Frontiers in Psychology in 2019, reported an observed 0.56, and 0.61 after correction for range restriction.

The finding worth acting on is not the headline figure but the split by why the rating was collected. Observed interrater reliability was 0.45 when the ratings were made for administrative purposes and 0.61 when they were made for research. After correction the research figure rises to 0.69 — and the administrative figure stays at 0.45. The correction that lifts one does nothing for the other.

Administrative ratings are the ones organisations actually run on: the appraisal that feeds a bonus, the talent-review placement, the readiness call before a promotion. So the ratings used to make decisions about people are, in the literature, the least reliable version of the rating that exists — and the gap is not a measurement artefact that a statistical correction removes.

How many raters it takes

The reliability of an average of several raters is higher than the reliability of any one of them, and the relationship is the Spearman-Brown formula applied to raters rather than to items: k times r, divided by one plus k minus one times r, where r is the reliability of a single rater and k is the number of raters.

Run it at 0.45, the administrative figure, and the practical answer falls out. One rater is 0.45. Two are 0.62. Three are 0.71. Five are 0.80. Nothing about the rating form changed across those four lines. This is why adding a second and third independent rater is almost always a better investment than redesigning the scale, and why a single-manager judgement should not be the sole basis for a promotion decision.

What it does to a validity coefficient

An unreliable criterion drags down any validity coefficient measured against it, which is why published validities are often corrected for criterion unreliability before being compared. That correction is legitimate and it is also where a great deal of the field's inflation has come from, because the size of the correction depends entirely on which reliability figure is fed into it.

So when a vendor quotes a validity coefficient, two questions decide whether the number means anything: what the criterion was, and which interrater reliability was used to correct it. A validity corrected using a research-grade 0.69 will look considerably better than the same study corrected using the 0.45 that describes the ratings the client actually keeps.

What a second and third rater buy you

  • Single rater, administrative purpose: r = 0.45
  • Two raters: (2 × 0.45) ÷ (1 + 1 × 0.45) = 0.90 ÷ 1.45 = 0.62
  • Three raters: (3 × 0.45) ÷ (1 + 2 × 0.45) = 1.35 ÷ 1.90 = 0.71
  • Five raters: (5 × 0.45) ÷ (1 + 4 × 0.45) = 2.25 ÷ 2.80 = 0.80

A one-manager readiness rating is a 0.45 instrument. If a promotion rests on it, the cheapest available improvement is not a better form — it is a second and a third independent rater, and the arithmetic says roughly where to stop.

Not the same as internal consistency

Internal consistency describes whether the items of one instrument hang together for one respondent. Interrater reliability describes whether two people scoring the same performance agree. A rating form can be internally consistent and still produce a different verdict depending on who filled it in.

Cronbach's alpha

How AssessAll handles it

The 360 report will not display a rater group's average until at least three people in that group have responded; below the floor the group is merged into a combined others, and if that is still short the figure is withheld rather than shown. The floor is set in code as at least three and cannot be configured lower. It was built for rater anonymity rather than for reliability, but three raters is also roughly where the arithmetic above turns a 0.45 single-rater judgement into something around 0.71, so the two reasons land in the same place.

Sources

Read next

Related terms

More terms beginning with I

Check a selection process against the four-fifths rule

Free, no signup, computed in your browser — with the remedy, not just the verdict.

Open the calculator

Last reviewed