All guides

How to read an assessment report (without misreading it)

Assessment reports are misread more often than they are wrong. Here's how to interpret scores, bands, percentiles and competency breakdowns — including the norm group, the standard error behind a single score, and the mistakes to avoid.

Last updated

The one rule that prevents most misreads

An assessment report is evidence about specific measured competencies, at a specific moment, on a specific instrument — not a verdict on a person. Read it as 'what this person demonstrated on these tasks', and most interpretation mistakes disappear before they start.

That framing matters because the most common misuse of assessment reports is over-extension: treating a reasoning score as general intelligence, a behavioural style as a fixed identity, or one sitting's result as permanent. Good reports scope their own claims; good readers keep them scoped.

Scores, bands, and percentiles — what each actually says

A raw or scaled score states how the person performed against the instrument's own scale — AssessAll's English suites, for example, score 10–90 and map that number to a CEFR band. A band groups scores into named ranges (B2, 'proficient', 'developing') so a decision-maker reads a meaning rather than a bare number. A percentile compares the person to a norm group: 78th percentile means better than 78% of that group — which is only as meaningful as the group itself, so check who the comparison is against.

The practical order to read them: band first for the headline, score for precision, percentile for context. A B2 at 58/90 and a B2 at 66/90 are the same band with different headroom — the score tells you which end of the band you are hiring.

Read the competency breakdown, not just the total

A single overall number hides the shape of the result, and the shape is usually the useful part. 'B2 overall — C1 reading, B1 pronunciation' describes a person who processes written English at a professional level but needs support in spoken clarity; the overall alone would have hidden exactly what a voice-process recruiter needed to know.

The same logic applies to judgement and leadership reports: two identical totals can decompose into opposite profiles — strong prioritisation with weak coaching versus the reverse — that suit different roles. On AssessAll, every report breaks results into competencies, and those competency scores accumulate into the person's Skill Passport across sittings.

Confidence, consistency, and response quality

Stronger reports also tell you how much to trust the result. Signals to look for: consistency checks (paired items that should agree), response-quality flags (rushed answers, straight-lining, an unusual completion time), and — where AI grading is involved — whether low-confidence answers were re-graded by a stronger model before the score was finalised.

Treat these as weighting instructions. A clean sitting with consistent responses deserves full weight; a flagged sitting deserves a follow-up conversation or a retest, not silent acceptance of the number.

The three lines that decide whether the rest of the report means anything

Before reading a single score, find three things. **Who the norm group is.** A percentile is a rank inside a specific comparison sample, and 'all test takers' and 'shortlisted graduates for this role' put the same raw score at wildly different percentiles. Where AssessAll reports a percentile inside a hiring drive it is a cohort percentile — the share of scored candidates in that same pipeline at or below this one, not a national or industry norm — and a report that leaves the group unnamed has not told you what its percentile means.

**The reliability, and the standard error that comes out of it.** Reliability is a property of a group; the standard error of measurement is what it says about one person, in score points: SEM = SD × √(1 − reliability). With a score standard deviation of 12 and a reliability of 0.82, the SEM is about 5.1, so a 95% band around any single score is roughly ±10 points. Our sample size and reliability calculator turns those two published numbers into the band and into the odds that a candidate near your bar is on the wrong side of it.

**Which version of the instrument produced it, and when.** A new item set, a re-trained AI grader, a re-worded rubric or a new delivery mode all move the score distribution against a cut score nobody re-examined. Two reports from the same product six months apart are not necessarily on the same scale, and a report that does not state its version cannot tell you whether they are.

Why the top of your shortlist is probably not as far ahead as it looks

Measurement error pushes extreme scores further from the mean than the underlying ability warrants — high scorers had luck on the day more often than low scorers did. So the best estimate of someone's true score always sits closer to the cohort mean than their observed score, by exactly the fraction the instrument is unreliable. Kelley's estimate makes it concrete: true score ≈ mean + reliability × (observed − mean). On a cohort mean of 55 and a reliability of 0.82, an observed 63 is best estimated at about 61.6.

That is a small correction and a large consequence. It says the gap between your first and third ranked candidates is systematically overstated, and it is the statistical reason a ranked shortlist should be treated as a band of comparable people rather than an order of merit. Nunnally's own standard for decisions about individuals — set out in Psychometric Theory in 1978 — was a reliability of 0.90 minimum and 0.95 desirable; the 0.70 that gets quoted as 'acceptable' was his bar for early-stage research, not for deciding about a person.

The five most common misreads

Treating a behavioural profile as pass/fail — DISC and similar instruments describe style, not ability, and have no failing result. Comparing scores across different instruments as if they shared a scale. Reading a percentile against the wrong norm group. Letting one strong competency halo the rest of the report. And deciding on a borderline score alone, when the report is one input that should be paired with a structured interview.

The fix for all five is the same discipline: know what each number measures, against whom, and pair the report with one other source of evidence before a consequential decision.

Frequently asked questions

What is a good score on an assessment?

It depends on the instrument's scale and the bar the role needs — there is no universal 'good'. Anchor on the band and the role benchmark: a voice-support role screening at CEFR B2 needs 53+ on AssessAll's 10–90 English scale, while a back-office chat role may hire well at strong B1. The report's bands exist precisely so you don't have to guess.

What does a percentile mean in an assessment report?

It ranks the person against a comparison group: 78th percentile means they scored higher than 78% of that group. Always check which group — 'all test takers' and 'shortlisted graduates for this role' can put the same raw score at very different percentiles.

Can a candidate fail a DISC or personality assessment?

No. Behavioural instruments describe how a person tends to operate — they have no passing score. Using them as pass/fail screens is the most common misuse; use them to inform interviews, onboarding, and team composition after ability has been measured with the right instruments.

How recent does an assessment result need to be?

It depends on what was measured. Reasoning ability is stable over years; language proficiency and technical skills move with practice, so results older than 12–18 months deserve re-verification for critical roles. A Skill Passport helps here: it timestamps every result and shows the trend, so you can see whether a competency is current, improving, or stale.

What should I do when a score sits right on the cut score?

Treat it as undecided by the measurement. If the 95% band around the score straddles the bar — and with a typical reliability of 0.82 that band is roughly ±10 points wide — then the instrument is not separating that candidate from the threshold, and whatever you decide is being decided by noise. The standard responses are a documented borderline band that routes those results to a human, a second independent source of evidence, or a re-sit. On AssessAll Certified the borderline band is built in: a composite within a few points of the nearest level cut is escalated to human review rather than resolved by arithmetic.

How big does the sample need to be before a pass rate in a report is trustworthy?

Larger than most dashboards imply. At a 35% pass rate, roughly 350 completed sittings gets a 95% interval of about ±5 percentage points; about 90 gets ±10; about 30 gets ±17, which is wider than most differences anyone would act on. Under about 30 sittings, report the counts rather than the rate. Note that this is a fact about the cohort figure only — no sample size makes an individual score more precise.

Should one assessment report decide a hire?

No — and well-designed reports say so themselves. Best practice pairs the report with one independent source of evidence, usually a structured interview focused on whatever the report left uncertain. The report's job is to make that conversation sharper, not to replace it.

See a report built for reading
Competency breakdowns, bands, response-quality signals, and results that accumulate into a verified Skill Passport.
Browse assessments