Evidence index
Meta-analysis · 2019

How much do two managers agree when they rate the same employee?

Two supervisors rating the same employee's overall job performance agree far less than most hiring material implies. A 2019 meta-analysis of 224 samples put the observed interrater reliability at 0.61 where ratings were collected for research and 0.45 where they were collected for appraisal or pay decisions — the purpose of the rating, not the scale, moved the number most.

Citation

Salgado, J. F. & Moscoso, S. (2019). Meta-Analysis of Interrater Reliability of Supervisory Performance Ratings: Effects of Appraisal Purpose, Scale Type, and Range Restriction. Frontiers in Psychology, 10, 2281.

Primary source opened and quotes confirmed on .

In its own words

“The best estimates of the observed ryy for overall job performance are 0.61 for research-collected ratings and 0.45 for administrative-collected ratings.”
Abstract, Results
“When the ratings had been done for administrative purposes, the observed interrater reliability fell to 0.45, but the observed interrater reliability rose to 0.61 when the ratings had been done for research purposes.”
Results, overall job performance
“Appraisal purpose moderates ryy and researchers and practitioners should be aware of its effects before collecting ratings or using empirically-derived interrater reliability distributions”
Abstract, Conclusions, point 1
“The most relevant finding of Viswesvaran et al.'s meta-analysis was to show that the observed interrater reliability (sample size weighted) was 0.52 (K = 40, N = 14,650) for supervisory ratings of overall job performance.”
Introduction, on the figure the market still quotes
“Wilmot et al. (2014) affirmed that the 0.52 interrater reliability estimate found by Viswesvaran et al. (1996) is really the interrater reliability for single-scale measures of overall job performance.”
Introduction
“if the uncorrected reliability estimates are used for correcting validity coefficients for attenuation, there is a possibility of substantial overestimation of the validity”
Discussion, on correcting validity
“the unrestricted criterion reliability must be used for correcting validity coefficients”
Discussion
“the corrected interrater reliability is larger for the multi-item scales than for the mono-item scales”
Results, scale type
“The vast majority of the studies (94.5%) used two raters”
Method, description of the database
“We cannot conduct separate analyses for the administrative ratings because the 18 coefficients were obtained from published studies.”
Method, on the limits of the administrative subset

What it does not say

Each of these is a claim made in this market and attributed to the source above. None of them is supported by it.

Commonly claimed: That the interrater reliability of performance ratings is 0.52.

That figure is Viswesvaran, Ones and Schmidt's 1996 observed estimate, and this paper quotes it in order to revise it: 0.56 across all purposes, 0.61 for research-collected ratings, 0.45 for administrative ones. The paper also records Wilmot et al.'s finding that 0.52 was really the figure for single-scale measures. A newer meta-analysis by Zhou, Sackett, Shen and Beatty in 2024 reports 0.65 and explicitly advises against using any grand mean at all. Quoting 0.52 in 2026 quotes a number its own field has revised twice.

Commonly claimed: That 0.45 is the reliability of 360-degree feedback, or of peer or subordinate ratings.

This meta-analysis is about supervisory ratings only — one supervisor compared with another supervisor. It reports nothing about peers, direct reports or self-ratings, and nothing about a multisource composite. Any figure applied to a 360 instrument on the strength of this paper is applied outside the population it was estimated on, and this site says so on the pages where it uses the number.

Commonly claimed: That administrative ratings are unreliable because managers inflate them.

The paper establishes that appraisal purpose moderates the coefficient. It does not establish the mechanism, and it does not test leniency, motivated distortion or rater training as explanations. A moderator is a fact about where the number goes, not a theory about why.

Commonly claimed: That 0.45 is a safe figure to divide a validity coefficient by.

It is the observed administrative value, and the paper could not produce a range-restriction-corrected administrative estimate at all — the sentence saying so is quoted above. The Discussion says the unrestricted criterion reliability is what a validity correction requires, and that using uncorrected estimates risks substantial overestimation of validity. Because the correction divides by the square root of the coefficient, the lowest available figure produces the largest published validity, which is the direction this paper warns about.

Commonly claimed: That a higher coefficient here means a better rating form.

Interrater reliability as measured in this literature is a property of a study — its raters, its purpose, its scale length and how much the sample's performance varied. It is an input you borrow when you have no local estimate, not a score your instrument earns.

The figures

How many independent supervisors a composite rating needs to reach a reliability of 0.80, by which published coefficient you start from

Starting coefficient, and where it comes fromOne raterRaters needed for 0.80
0.45 — observed, administrative purpose (Salgado & Moscoso, 2019)0.455
0.52 — observed, single-scale overall (Viswesvaran et al., 1996)0.524
0.57 — managerial jobs (Zhou et al., 2024)0.574
0.61 — observed, research purpose (Salgado & Moscoso, 2019)0.613
0.65 — overall estimate (Zhou et al., 2024)0.653
0.68 — non-managerial jobs (Zhou et al., 2024)0.682

Computed with the Spearman-Brown formula, rk = kr / (1 + (k - 1)r) — the same formula the 2019 paper used in reverse, to reduce multi-rater studies back to a one-rater figure. Nothing in this table is quoted from either paper: it is arithmetic on their published coefficients and every row is reproducible in a spreadsheet. The 2024 figures are taken from that paper's abstract, which is free at the repository linked below; its full text is paywalled and was not opened.

Why this source carries more of this site than any other

Interrater reliability here means the correlation between two supervisors independently rating the same person on the same performance construct. It is the ceiling on how much of a rating is about the person rather than about the rater, and it is the denominator in the correction that turns an observed validity coefficient into a published one.

This paper is the most load-bearing citation on this site. Its administrative figure of 0.45 is the number behind the rater-count arithmetic on the interrater reliability entry, the three-rater floor recommended for a 360 cycle, the caution printed on the annotated 360 sample report, and the warning against reading a one-box difference on a 9-box grid as a difference. Four separate surfaces argue from one coefficient in one paper, which is exactly the dependency this index exists to make visible.

Two moderators produced it. Ratings gathered for a research study reach 0.61; ratings gathered because they will feed an appraisal, a pay decision or a promotion reach 0.45. Same instrument, same kind of rater, same construct, separated only by what the rating was for. Scale length is the second moderator and runs as you would expect — the paper reports corrected interrater reliability as larger for multi-item than for mono-item scales, and quotes the authors' own earlier figures of 0.45 for mono-item against 0.64 for multiple-item scales. A single overall rating box is the least reliable way to collect the thing and the most common.

The population limit, which is the part most easily lost

Every coefficient in this paper describes two supervisors. It reports nothing about peers, nothing about direct reports, nothing about self-ratings, and nothing about the composite a multi-rater instrument prints. Where this site uses 0.45 to argue for a rater floor in a 360 cycle, it is borrowing a supervisor-versus-supervisor figure and applying it to a different rater population, and the honest statement is that no better figure is being used because a comparably cumulated one for peer and subordinate ratings is not to hand.

There is a second boundary inside the method. The vast majority of the studies used two raters, and where more were used the estimates were reduced back to a one-rater figure with the Spearman-Brown formula. So every published coefficient here is a single-rater number, and every practical use of it has to build back up — which is legitimate, and is the arithmetic in the table above.

And the administrative subset is thinner than the headline suggests. The paper states plainly that it could not run separate analyses for the administrative ratings, because those 18 coefficients came from published studies. The 0.45 is an observed value with no range-restriction correction available behind it. Anyone quoting it as though it were the corrected, unrestricted figure — including anyone quoting it from this site — is quoting the one number in the paper that could not be corrected.

The 2024 update, and what it changes

Zhou, Sackett, Shen and Beatty published an updated meta-analysis in the Journal of Applied Psychology in 2024, over 132 independent samples, using a weighting method chosen so that large samples do not dominate the result. They report 0.65 overall — higher than the 0.56 in this paper and higher than the 0.52 the market has quoted since 1996 — and they report that it varies by job type: 0.57 for managerial positions against 0.68 for non-managerial ones. Their recommendation is the consequential part. They advise against using an overall grand mean of interrater reliability at all, and recommend job-specific or local reliabilities instead.

That recommendation is aimed squarely at what almost every vendor does, this one included, when it reaches for a single published coefficient to run an argument. It does not make the 2019 purpose split wrong — the 2024 paper examines rating purpose too — but it does mean that a site arguing from one grand mean while a newer cumulation says not to use one owes its readers the sentence saying so. This page is that sentence, and the pages listed below are the ones that have to change.

The job-type finding also has teeth for multi-rater feedback specifically. A 360 cycle is usually run on a manager, and the managerial figure is the lower of the two. Read against the table above, that is the difference between two raters and four for the same target reliability — on the target whose feedback is hardest to collect.

Which figure goes into a validity correction, and which direction the error runs

A test correlating 0.30 with supervisor ratings is published as 0.30 divided by the square root of the criterion reliability, so the choice of denominator moves the headline number without touching the test or the sample.

Take the same observed 0.30. Divide by the square root of 0.45, the administrative figure, and it is published as 0.45. Divide by the square root of 0.61, the research figure from the same paper, and it is published as 0.38. Divide by the square root of 0.65, the 2024 estimate, and it is published as 0.37. One correlation, three defensible denominators, a published coefficient anywhere from 0.37 to 0.45, and no arithmetic error anywhere in the chain.

The paper is explicit about which direction the error runs. It says the unrestricted criterion reliability must be used for correcting validity coefficients, and that using uncorrected estimates risks substantial overestimation of validity. Because the correction divides by a square root, the lowest available coefficient produces the largest published validity — so reaching for the administrative 0.45 because it is the most conservative-sounding number produces the least conservative result. That is the same overcorrection mechanism Sackett and colleagues identified across this literature in 2022.

Where this source is used here

These pages argue from the source above. If it is ever superseded, these are the pages that have to change.

Read next

SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed