A norm-referenced score says where a candidate stands relative to other people: 78th percentile, second decile, above average. A criterion-referenced score says what the candidate can do relative to a defined standard: meets the proficiency requirement, or does not. They answer different questions, and using one for a decision that requires the other is a common and consequential scoring error.
The distinction decides whether your scores survive a year of market change, whether you can defend a rejection, and whether a result means anything outside the organisation that produced it.
The two questions a score can answer
The *Standards for Educational and Psychological Testing* — the joint AERA, APA and NCME document governing professional testing practice — puts the choice at the top of test design. Specification of intended uses, it states, "will include an indication of whether the test score interpretations will be primarily norm-referenced or criterion-referenced. When scores are norm-referenced, relative score interpretations are of primary interest."
Relative interpretation ranks; absolute interpretation compares performance against a defined level of competence. A test built well for one purpose may serve the other poorly: "Tests designed to facilitate one type of interpretation may function less effectively for the other type of interpretation."
Most organisations run both without noticing. Campus shortlisting is almost always norm-referenced — take the top N. Certification and compliance testing is almost always criterion-referenced — demonstrate the competency or repeat the module. Problems begin when the output of one gets quietly treated as the other.
Where norm-referenced scoring is the right choice
Ranking is correct when the decision itself is comparative and supply-constrained.
- Fixed-slot selection. With 40 trainee seats and 2,000 applicants, the question is not "who is competent?" but "who are the strongest 40?" The Standards acknowledge this: cut scores set "to select a specified number of examinees (e.g., to identify a fixed number of job applicants for further screening)" may need little further justification.
- Triage under volume. When the top of the funnel is enormous, a relative cut is the only tractable first filter.
- Competitive entrance examinations. India's CAT reports percentiles precisely because the prize is a limited number of seats, not a certificate of ability.
Where criterion-referenced scoring is the right choice
Absolute interpretation is correct when the decision depends on capability, not rank.
- Readiness decisions. Can this agent handle a live customer call unsupervised? A percentile cannot answer that. Twenty people can sit in the top decile of a weak cohort and none be ready.
- Pre/post training measurement. Score against the cohort and a group that improves uniformly shows no movement at all — everyone keeps their rank.
- Anything portable. A criterion statement travels; a percentile does not survive leaving its reference group. This is why competency frameworks and National Occupational Standards are written as criteria rather than ranks, and why a Skill Passport recording demonstrated competencies is more useful to a future employer than one recording that someone beat 71% of a cohort nobody can now describe.
- Fairness scrutiny. "Why was this person rejected?" is easier to answer against a documented standard than against a shifting distribution.
The failure mode: a percentile that quietly becomes a standard
The characteristic mistake is to build a norm-referenced cut once, then leave it in place as though it were a competence standard. It is not, because the reference group moves.
The Standards are blunt about this dependency: "The validity of norm-referenced interpretations depends in part on the appropriateness of the reference group to which test scores are compared," and for established tests, "periodic review is generally required to ensure the continued utility of their norms. Renorming may be required to maintain the validity of norm-referenced test score interpretations." Standard 5.9 places a continuing obligation on publishers: as long as the test remains in print, it is the publisher's responsibility "to renorm the test with sufficient frequency to permit continued accurate and appropriate score interpretations."
Applicant pools have moved unusually fast. LinkedIn reported an average of 11,000 job applications submitted per minute, a 45% increase in a year, with generative AI tools contributing to the volume, as reported by *The New York Times*. A funnel that grew that much is not the same distribution it was. A percentile cut calibrated on the older pool now selects a different absolute level of skill — possibly higher, possibly lower — and nothing in your dashboard will tell you which.
The reverse error also happens: ranking candidates by criterion-referenced results, as though a scale built to sort competent from not-yet-competent says who is best.
What regulators and professional standards actually require
29 CFR 1607.5 of the US Uniform Guidelines on Employee Selection Procedures treats ranking as the more demanding use: "Evidence which may be sufficient to support the use of a selection procedure on a pass/fail (screening) basis may be insufficient to support the use of the same procedure on a ranking basis under these guidelines."
The same section sets the expectation for absolute standards: "Where cutoff scores are used, they should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force." Adverse-impact analysis under the four-fifths convention in 1607.4 applies to either approach.
SIOP's *Principles for the Validation and Use of Personnel Selection Procedures* adds a documentation duty most organisations skip: the normative group "should be described in terms of its relevant demographic and occupational characteristics," and "the time frame in which the normative results were established should be stated." If you cannot say who your norm group was and when it was collected, your percentiles are decoration. Standard 5.21 of the Standards imposes the parallel duty on the other side: where interpretations involve cut scores, "the rationale and procedures used for establishing cut scores should be documented clearly."
Banding: the middle path, and its unresolved argument
Score banding treats scores within a range as statistically indistinguishable given measurement error, then lets other factors decide within the band — ranking's efficiency without claiming precision the instrument lacks.
It remains contested. Campion and colleagues' review of the banding controversy in *Personnel Psychology* set out the competing positions across ten questions without resolving them, and the SIOP Principles require only that "the basis for its development and the decision rules to be followed in its administration should be clearly documented." Treat banding as a defensible option that raises the documentation burden, not a way to dodge the choice.
A practical checklist
- Name the decision first. Fixed number of slots, or a capability threshold? That settles which scale you need before any statistics.
- Write the standard, if the decision needs one. A criterion cut needs a defined performance level and a documented rationale — not a round number inherited from school marking.
- Document your norm group. Who, how many, when. Put the collection date on the report.
- Diarise renorming. Pick a review interval — annually for high-volume roles — and check whether the distribution has shifted.
- Label the two scales separately on the report — "78th percentile" and "meets standard" are not adjacent numbers.
- Run adverse-impact analysis on the rule you actually use, including any within-band discretion.
- Re-examine thresholds inherited from a previous vendor or year. They encode a reference group you may no longer hire from.
When you need both
Often you do, and the Standards permit it: "Both criterion-referenced and norm-referenced scales may be developed and used with the same test scores if appropriate methods are used to validate each type of interpretation."
The sequence that works is criterion first, norm second: apply an absolute competence threshold to establish who can do the work, then rank the qualified group to allocate limited slots. That order catches what a pure ranking cannot — filling every seat from a cohort in which nobody met the standard. It does require both interpretations to be validated on the same instrument, which is where AI-graded scenario assessments help: a response can be scored against a defined behavioural standard and placed on a scale, provided both are documented. Organisations without an in-house psychometrics function can use AssessAll's custom assessment services for that standard-setting work rather than leaving the threshold to guesswork.
Takeaway
A percentile tells you who won a particular contest on a particular day against a particular crowd; a criterion tells you whether someone can do the job. Decide which claim your decision requires, document the reference group or standard behind it, and revisit both on a schedule — the crowd changes even when the job does not.