---
title: "What should a 360 feedback report actually show you?"
description: "Nine pages rank for this question and none of them answers it. Here are the seven things a 360 report has to show before its averages mean anything — how many raters sit behind each number, how much they disagreed, what was marked not observed, and the two different anonymity thresholds nobody volunteers."
canonical: https://www.assessall.com/guides/what-a-360-feedback-report-should-show-you
updated: 2026-09-28
source: AssessAll
---

# What should a 360 feedback report actually show you?

Nine pages rank for this question and none of them answers it. Here are the seven things a 360 report has to show before its averages mean anything — how many raters sit behind each number, how much they disagreed, what was marked not observed, and the two different anonymity thresholds nobody volunteers.

_Last updated 2026-09-28._

<!-- #the-short-answer -->
## The short answer

A 360 feedback report should show you four things most do not: how many people sit behind each number, how much they disagreed, which behaviours were marked not observed, and the threshold at which a group's ratings were withheld. Without those four an average is unreadable — two raters and nine look identical.

Everything below is that claim unpacked into seven checks you can run against any vendor's sample report in about ten minutes, plus the one quantity no report in this market prints and the arithmetic that has to stand in for it.

<!-- #what-the-pages-ranking-for-this-question-actually-compare -->
## What the pages ranking for this question actually compare

Searched on 28 September 2026, the phrase *best 360 degree feedback tool* returns nine results, and the set has been stable across three measurements eleven days apart: Betterworks, AIHR, Qualaroo, Eletive, People Managing People, Workhuman, NanoGlobals, SelectSoftwareReviews and Launch-360. Two of them were opened in full and read on 28 September 2026 — AIHR's twelve-tool listicle and SelectSoftwareReviews' buyer guide.

**Neither mentions a reliability coefficient, a standard error, a confidence interval, rater-agreement or standard deviation, a minimum number of raters per group, an anonymity threshold, or a not-observed response option.** AIHR's stated criteria are automation, customisation, reporting functionality, user experience and business outcomes. SelectSoftwareReviews works from customisable surveys, anonymous feedback, integrated performance management and insightful analytics. Both are careful, useful pages. Both treat a 360 as a piece of performance-management software rather than as a measurement instrument, which is a reasonable thing for a software buyer's guide to do and a poor basis for deciding what a report has to contain.

So this page is not a better listicle. It is the question those pages leave on the table: once you have the tool, what does the artefact it produces have to show you before you can read it?

<!-- #the-seven-checks-in-the-order-they-matter -->
## The seven checks, in the order they matter

**1. The rater split, not the response rate.** A cover line reading *8 of 8 raters completed* is a response rate. It is not the number of people behind any figure inside. Ask for the split by group. A 360 with a 100% response rate and one peer is thinner evidence than a 360 with a 60% response rate and six, and the report will not tell you which you are holding unless it prints the split.

**2. An observed count on every row.** Each behaviour needs the number of people who actually answered it printed beside its average, and a floor below which the row does not appear at all. Without a floor, a behaviour one rater saw once can top or bottom a ranked table. With one, a ranking is a ranking rather than an anecdote.

**3. A not-observed option, and a count of how often it was used.** If the scale has no way for a rater to say they have never watched this person run a difficult conversation, every average in the report contains guesses from people with no exposure. That is not conservatism; it is noise recorded as data. The report should show how many ratings each competency is built from and how many people declined to rate it.

**4. A spread statistic per item, not only a mean.** This is the single most useful thing an artefact can carry and the one most likely to be missing. An item averaging 3.7 where the answers run from 1 to 6 describes nobody's experience of the person. A mean plus a standard deviation, or a banded agreement label derived from it, turns that from a mild development area into the most productive twenty minutes of the debrief. Ask what the bands are and what statistic they come from.

**5. Which groups are shown, and which are inside the combined column anyway.** Count the chart's bars against the rater split on the cover. A group below the anonymity floor is usually withheld from the charts while its ratings remain inside the combined others average — which is the correct behaviour, and which almost no report states on the page. Ask directly: are the withheld group's ratings still in the combined figure?

**6. Two anonymity thresholds, because there are two.** The minimum group size that governs *scores* and the rule that governs *written comments* are usually different numbers applied to different groupings. Comments are commonly released once the whole pool of others clears the floor, not once each group does — so a team of two direct reports can have its scores withheld and its comments published in the same document. Ask for both numbers separately. Nobody volunteers the second one.

**7. What the vendor says the report is for.** SHL's own Enterprise Leadership 360 sample states on its face that *It is recommended that you read through your report with a trained facilitator*, and its disclaimer limits use of the questionnaire to people trained in its interpretation — read at source on 15 September 2026. That is an honest statement about how these instruments are designed to be used. It is also the reason a line manager who downloads a sample PDF to decide whether to buy gets a document that answers none of their questions.

<!-- #the-one-number-no-360-report-in-this-market-prints -->
## The one number no 360 report in this market prints

Ten 360 reports were read at source on 15 September 2026 — Mercer Mettl, three SHL reports, Hogan via its UK distributor, DecisionWise, Appraisal360, Info-Tech, hr-survey.com and 360feedback.app. Several print the spread and one prints per-item standard deviations with a plain-English explanation of them. **Not one prints a reliability coefficient, a standard error of measurement or a confidence interval.** Neither does AssessAll's own 360 report, and saying so is the point: this is a gap in the category, not a shortcoming of one vendor, and a buyer who asks for it will be the first person that quarter to ask.

That leaves you doing the arithmetic the report should have done. The nearest published figure is the interrater reliability of performance ratings, and it moved recently in the unhelpful direction. Zhou, Sackett, Shen and Beatty cumulated 132 independent samples in 2024 and report 0.65 overall, **0.57 for managerial positions against 0.68 for non-managerial ones**, and they advise against using an overall grand mean at all. Everyone who is the subject of a leadership 360 is a manager, so 0.57 is the applicable figure — not the 0.45 administrative grand mean that circulated for years, and not the flattering 0.65.

Put 0.57 through the [Spearman-Brown formula](https://www.assessall.com/guides/glossary/s/spearman-brown-formula) and the picture is stark. One rater gives 0.57. Two give 0.73. Three give 0.80. Four give 0.84. Six give 0.89. A single manager's bar printed to one decimal place beside a six-person others bar, in the same typeface, at the same width, carries roughly two thirds of the latter's reliability — and nothing on the chart says so.

Two consequences worth carrying into a vendor conversation. **A minimum group size of three is not only a privacy rule; it is approximately where the number becomes worth printing**, because three raters is where this arithmetic first crosses the conventional 0.80 threshold and two does not reach 0.73. And a report that shows you a manager bar without a caveat is showing you the least reliable line on the chart in the position readers trust most.

<!-- #the-seam-to-find-before-the-debrief-in-whatever-tool-you-use -->
## The seam to find before the debrief, in whatever tool you use

Most 360 reports compute their headline summary boxes — confirmed strengths, blind spots, development areas — by one rule, and the per-competency badges by another, with different thresholds. The two rules are individually defensible and they are not mutually exclusive, so the same competency can appear as a confirmed strength on page one and carry a development-area badge on page two.

AssessAll's own sample does exactly this: Goal-Setting & Performance sits in confirmed strengths with an others average of 4.4, because the strengths rule needs 4.0 of 6, and carries a development-area badge on its competency card, because the quadrant band needs 4.5. Both statements are true under their own rule. To the person holding the report it reads as a contradiction, and they will ask.

The check is quick and it works on any vendor's sample: find the two rules, find their thresholds, and look for a competency that falls between them. Then rehearse the answer. Publishing the seam is cheaper than being asked about it in a debrief — which is why it is written on the sample page rather than left for a customer to discover.

<!-- #the-six-questions-to-send-a-vendor -->
## The six questions to send a vendor

Paste these into an email. Every one is answerable in a sentence, and a vendor who cannot answer them about their own artefact has told you something.

What is the minimum number of responses per rater group before that group is shown, and what happens to those ratings when the group is below it — withheld, merged, or dropped?

What is the separate threshold for releasing written comments, and is it applied per group or across the whole pool of others?

Does the scale carry a not-observed option, is it excluded from the averages, and does the report print how often it was used?

What spread statistic appears per item, and what are the bands for any agreement label derived from it?

What is the minimum observed count for a behaviour to appear in a ranked table?

Do you publish a reliability coefficient, standard error or confidence interval for this instrument — and if not, what figure should I use to interpret a single manager's rating?

<!-- #sources -->
## Sources

Zhou, S., Sackett, P. R., Shen, W., & Beatty, A. S. (2024). An updated meta-analysis of the interrater reliability of supervisory performance ratings. Journal of Applied Psychology. 132 independent samples. Abstract read at source; the full text is paywalled, and that limitation is recorded on the [evidence entry for this source](https://www.assessall.com/guides/evidence/salgado-moscoso-2019-interrater-reliability).

Salgado, J. F., & Moscoso, S. (2019), cumulating 224 samples, is the earlier source of the administrative-versus-research split — observed reliability of 0.45 for administratively-collected ratings against 0.61 for research ones. Set out in full, with what it does not establish, in the [evidence index](https://www.assessall.com/guides/evidence/salgado-moscoso-2019-interrater-reliability).

The ten competitor 360 reports were opened and read on 15 September 2026; the panel-by-panel walkthrough of AssessAll's own artefact, including the SHL facilitator disclaimer quoted above, is at [what a 360 feedback report looks like](https://www.assessall.com/guides/reports/360-feedback-report).

The search result set described in the second section was recorded on 28 September 2026 and the two pages named there were opened in full on the same date. Search results are re-ranked continuously and this set will change; the finding is dated deliberately.

<!-- #faq -->
## Frequently asked questions

### What should a 360 feedback report contain as a minimum?

The rater split by group rather than only a response rate; an observed count on every behaviour with a floor below which it does not appear; a not-observed option excluded from the averages and counted on the page; a spread statistic per item as well as a mean; a statement of which groups are shown and whether a withheld group's ratings remain inside the combined figure; and both anonymity thresholds, the one for scores and the separate one for comments.

### How many raters does a 360 need per group?

Three is the usual minimum and there are two independent reasons for it. It is roughly where anonymity becomes defensible, and it is approximately where the number becomes worth printing: applying the Spearman-Brown formula to the 0.57 interrater reliability that Zhou, Sackett, Shen and Beatty report for managerial positions in 2024, three raters reach 0.80 and two reach 0.73. Four raters reach 0.84. Below three you are reading one or two people's opinion presented as a group average.

### Why does a 360 report show a manager's score as a single bar?

Because the manager is usually told their view will be attributed, so it is shown alone rather than pooled. The consequence is that the manager bar is one rater printed to one decimal place beside a combined bar built from six, in the same typeface at the same width. On the nearest published figure for managerial ratings, one rater carries about 0.57 and six carry about 0.89. Read the manager's line as one person's view, which is what it is.

### Is a 360 feedback report anonymous?

Partly, and the honest answer is a set of specific rules rather than a yes. Peer, direct-report and external ratings should be anonymous by mechanism; a single named manager's view cannot be. The rules that matter are a minimum group size that merges rather than drops, comments rewritten to remove identifying detail, and a release gate — and the score threshold and the comment threshold are usually different numbers applied to different groupings. Ask for both.

### Should a 360 report be used for pay or promotion decisions?

The safer answer is no, for a mechanical reason rather than an ideological one. Ratings collected for administrative purposes show substantially lower interrater agreement than ratings collected for development: 0.45 against 0.61 observed in Salgado and Moscoso's 2019 cumulation of 224 samples. Attaching a consequence changes what raters write, and the gap does not close with statistical correction. Use a 360 to decide what a person works on, and instruments designed for decisions to decide about a person.

---

Source: https://www.assessall.com/guides/what-a-360-feedback-report-should-show-you
