All articles
L&D & Capability4 October 2026·6 min read

A 360 Moves Ratings by 0.15 of a Standard Deviation: Six Myths About Multisource Feedback

Meta-analysis puts improvement after 360-degree feedback at d = 0.05-0.15, and a third of feedback interventions make performance worse. Six claims about multisource feedback tested against the evidence: panel sizes, rater self-selection, cross-source disagreement, and development vs decision use.

By AssessAll Editorial

360-degree feedback, or multisource feedback, collects ratings of one person's behaviour from their manager, peers, direct reports and themselves. The best meta-analytic estimate of how much it improves those ratings over time is d = 0.05 to 0.15 — below Cohen's threshold for a small effect. It is a development input that can sharpen self-awareness. It is not a measurement instrument accurate enough to carry a promotion decision.

That gap between what a 360 is sold as and what it demonstrably does is wide enough that in May 2026 a Campbell Collaboration team registered a new protocol to redo the evidence synthesis, writing that there is still "no consensus on whether it predicts improvements in employee performance" (Campbell Systematic Reviews). Their background section puts adoption at 34% of UK HR leaders and up to half of medium and large US organisations, having grown from 27% in 2003 to 48% in 2013. A tool that widespread deserves a clear account of what the research actually supports.

Here are six claims made routinely about 360s, each tested against the published evidence.

Myth 1: A 360 improves performance

The reference point is Smither, London and Reilly's meta-analysis of 26 longitudinal studies in Personnel Psychology (2005). Corrected mean effect sizes for improvement in ratings over time, by rater source:

  • Direct reports — d = 0.15 across 21 studies, N = 7,705
  • Supervisors — d = 0.15 across 10 studies, N = 5,358
  • Peers — d = 0.05 across 7 studies, N = 5,331
  • Self-ratings — d = −0.04 across 11 studies, N = 3,684

Cohen's benchmarks put a small effect at roughly 0.20. The authors' own conclusion is blunt: "it is unrealistic for practitioners to expect large across-the-board performance improvement after people receive multisource feedback."

The CIPD's evidence review on performance feedback reaches the same reading: improvements in performance over time are small. None of this makes 360s worthless. It makes the expectation that running one causes capability to rise an unsupported one.

Myth 2: More rater sources mean better signal

The CIPD review is explicit that receiving feedback from more sources does not make a difference compared with receiving feedback from direct reports only. Smither and colleagues coded their studies for exactly this distinction — upward feedback only, versus feedback from multiple sources — and it is their data the CIPD is reading when it reports no difference.

The practical implication cuts against most 360 configurations. Adding a client panel and a dotted-line manager to an already long questionnaire adds respondent burden and adds very little that direct reports were not already telling you.

Myth 3: Feedback can only help

Kluger and DeNisi's meta-analysis in Psychological Bulletin (1996) covered 607 effect sizes and 23,663 observations. The average feedback intervention produced d = 0.41 — and in roughly a third of the studies analysed, performance went down.

The mechanism is reasonably well mapped. Brett and Atwater (2001) found that managers who rated themselves higher than others did reported significantly more negative reactions, and that low or lower-than-expected ratings "did not result in enlightenment or awareness but rather in negative reactions such as anger and discouragement." Nowack and Mashihi's review in the *Consulting Psychology Journal* (2012) catalogues the rest: poorly designed 360s can increase disengagement and depress individual and team performance.

A 360 is an intervention with a downside distribution. It should be designed like one.

Myth 4: Disagreement between rater groups means the data is bad

This is the most common misreading, and it is backwards. A 360 rests on the premise that different levels see genuinely different behaviour, so some cross-source disagreement is expected and informative rather than a defect. Scullen, Mount and Goff argued that rating differences often reflect real variation in how someone performs in front of different groups, not observer bias.

What the correlations actually look like: self-ratings correlate with other sources at roughly .3 to .6, with peer and supervisor ratings converging more closely with each other, and supervisors are the single most reliable source of job performance ratings (Conway and Huffcutt, 1997).

The correct inference from a wide self-versus-other gap is not "the instrument failed." It is a data point about calibration — and it is the one thing a 360 does better than any assessment.

Myth 5: Letting people pick their own raters biases the result

Plausible, and not what the research found. Nieman-Gonder and colleagues (2006) compared ratings from raters selected by the feedback recipient against raters they did not select, using multiple accuracy measures. Self-selected raters were as accurate, or more accurate, than non-selected raters.

What matters is not who chose the panel but whether the panel is large enough and has actually observed the behaviour being rated. A participative selection between the individual and their manager raises acceptance of the results without costing accuracy.

Myth 6: Five or six raters is enough

This is where most 360 programmes quietly fail. Greguras and Robie's analysis of within-source reliability concluded that reaching acceptable reliability — .70 or higher — takes at least four supervisors, eight peers and nine direct reports. 3D Group's research indicates that two or fewer respondents in a given group is inadequate for reliable measurement at all.

Set that against normal practice: a manager with three reports, two peers nominated, one supervisor. Every group in that report is below the threshold, and the sub-scores for each competency are noisier still, because reliability was estimated on overall scales rather than on the three-item competency blocks the report prints.

If the panel cannot be populated to those numbers, the honest move is to report fewer, broader scores — not to print a competency-by-competency bar chart that implies precision the data cannot support.

Where this leaves the design

The one finding in this literature with real leverage is about purpose, not instrumentation. In Smither and colleagues' data, the average uncorrected effect across rater sources was 0.25 in studies where the 360 was used for development, against 0.08 where it was used for administrative decisions such as promotion or performance review. The developmental use is roughly three times as effective — and it is also the use with no legal exposure.

That second point gets overlooked. Once 360 ratings feed a promotion decision, they are functioning as a selection procedure. The EEOC's guidance on employment tests and selection procedures lists performance appraisals among the tools covered, and applies to promotion as well as hiring. A rating instrument with interrater reliability nobody has estimated, panels below minimum size, and no job analysis behind the competency list is a weak thing to defend.

A workable split:

  1. Use the 360 for what it measures well — how a person is experienced by the people around them, and where their self-view diverges from that.
  2. Size the panels before you launch, and drop any rater group that cannot reach at least three respondents.
  3. Keep it out of the decision — development purpose, not administrative, with that stated to raters and recipients.
  4. Measure capability separately with an instrument scored against a fixed standard rather than against opinion. Pre- and post-programme scenario assessments graded on a defined rubric — the kind AssessAll runs as AI-graded scenarios, on pay-as-you-go credits rather than an annual platform licence — answer "did skill change?" The 360 answers "how is this person landing?" Those are two different questions.
  5. Expect a minority to react badly, and resource the debrief accordingly. Feedback delivered without a conversation is where the negative third of Kluger and DeNisi's distribution lives.

A 360 is a mirror, and a reasonably good one. The mistake is asking a mirror to certify a skill. Run the multisource feedback for self-awareness, run an assessment against a standard for capability, and stop reporting one as evidence of the other.

#360-degree-feedback#multisource-feedback#leadership-development#performance-ratings#rater-reliability#l-and-d

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free