All guides
Glossary

Assessment and psychometrics glossary

This is a measurement glossary, not a general HR glossary. It covers the vocabulary that appears the moment you open a score report or challenge a vendor's claim — reliability, validity, norms, cut scores, item statistics and fairness analysis — because that is the vocabulary buyers are handed without explanation and the one that decides whether a number can carry a decision.

Every entry opens with a definition that stands on its own, then shows the arithmetic, the term it is most often confused with, and the primary source. 25 terms so far, across 13 letters.

  • Adaptive testing

    Adaptive testing is a test design in which each question is chosen from a calibrated bank on the basis of the answers already given: a correct answer leads to a harder question, an incorrect one to an easier question. It reaches a given level of precision in far fewer items than a fixed test.

  • Adverse impact ratio

    The adverse impact ratio is one group's selection rate divided by the selection rate of the group selected most often. A ratio below 0.80 is the threshold at which United States federal enforcement agencies generally treat a selection procedure as showing adverse impact and expect the employer to justify it.

  • Angoff method

    The Angoff method sets a passing score by expert judgement rather than by ranking candidates. A panel estimates, for every question, the probability that a just-barely-qualified candidate would answer it correctly. Those probabilities are averaged across judges and summed across questions, and the total becomes the recommended cut score.

  • Banding

    Banding reports a score as a range rather than a point, on the reasoning that differences smaller than the measurement error are not real differences. Candidates whose scores fall inside the same band are treated as equivalent, and any ordering among them is decided on other evidence rather than on the score.

  • Construct validity

    Construct validity is the evidence that a test measures the psychological attribute it claims to measure, rather than something merely correlated with it. It is built from convergent evidence — the score tracks other measures of the same attribute — and discriminant evidence, that it does not track measures of different attributes.

  • Criterion validity

    Criterion validity is the strength of the relationship between a test score and an independent outcome the test is meant to anticipate — typically job performance, training results or retention. It is reported as a correlation coefficient, and it is measured either at the same time as the outcome or in advance of it.

  • Cronbach's alpha

    Cronbach's alpha is an index of internal consistency: how far the items in a scale behave as though they measure one thing. It runs from 0 to about 1 and rises with both the average correlation between items and the sheer number of items, which is why a high alpha on its own proves very little.

  • Cut score

    A cut score is the point on a score scale at which a decision changes — pass or fail, shortlist or reject, certified or not. It is a policy choice made by the organisation using the test rather than a property of the test itself, and moving it changes who is selected without changing anyone's score.

  • Differential item functioning

    Differential item functioning occurs when two groups of equal ability on the attribute being measured have different chances of answering a particular question correctly. It is a property of the item rather than of the groups, and it is the statistical evidence used to identify questions that need review for bias.

  • Face validity

    Face validity is whether a test looks relevant to the people taking it. It is not evidence that the test measures anything, and an instrument can be highly valid with poor face validity or the reverse. It matters anyway, because it drives completion rates, candidate goodwill and whether a hiring manager trusts the result.

  • Four-fifths rule

    The four-fifths rule is the rule of thumb in the United States Uniform Guidelines on Employee Selection Procedures: if a group's selection rate is less than four-fifths of the highest group's rate, federal enforcement agencies will generally regard the procedure as having adverse impact. It is a trigger for scrutiny, not a legal verdict.

  • Item analysis

    Item analysis is the statistical review of how each question in a test performed. Two numbers carry most of it: difficulty, the proportion of candidates who answered correctly, and discrimination, the correlation between getting that question right and scoring well on the test overall. Weak items are revised or retired on this evidence.

  • Item response theory

    Item response theory models the probability of a given answer as a function of the candidate's underlying ability and the question's own parameters — usually difficulty, discrimination and, for multiple choice, guessing. Because ability and item difficulty are placed on the same scale, scores from different sets of questions remain comparable.

  • Norm group

    A norm group is the reference sample a raw score is compared against in order to produce a percentile or standard score. The same performance can be an unremarkable result against one norm group and an exceptional one against another, which makes the identity, size and date of the norm group part of the score itself.

  • Percentile

    A percentile is a position, not a proportion. A percentile of 70 means the person scored at or above 70 per cent of the norm group; it does not mean they answered 70 per cent of the questions correctly. Percentiles are also compressed in the middle of a distribution and stretched at the extremes.

  • Predictive validity

    Predictive validity is the correlation between a score measured before a hiring decision and an outcome measured afterwards, such as supervisor-rated performance. It is the form of validity evidence that matters most in selection, because it is the only one whose study design matches the way the test is actually used.

  • Psychometrics

    Psychometrics is the science of measuring psychological attributes — ability, personality, knowledge, attitude — that cannot be observed directly. It covers how questions are written and calibrated, how responses become scores, how much error surrounds a score, and what evidence is required before a score may support a decision about a person.

  • Reliability

    Reliability is the consistency of a measurement: how much of the variation in scores is signal rather than noise. It is reported as a coefficient between 0 and 1, and it comes in several distinct kinds — internal consistency, test–retest, parallel forms and inter-rater — which answer different questions and are not interchangeable.

  • Social desirability bias

    Social desirability bias is the tendency of respondents to answer questionnaires in the way they believe will be viewed favourably rather than accurately. It inflates scores on desirable traits, compresses the differences between people, and is strongest exactly when something depends on the answer — such as a job application.

  • Standard error of measurement

    The standard error of measurement is the expected spread of a person's observed scores around their true score if they could be tested repeatedly. It converts a reliability coefficient into the units of the actual scale, which is what turns an abstract number into a usable statement about one individual candidate.

  • Standard setting

    Standard setting is the formal process of deciding where a cut score belongs. It uses a documented method — Angoff, bookmark, contrasting groups — with a panel, a written definition of the borderline candidate, and a record of the judgements, so that the resulting threshold can be defended rather than merely asserted.

  • Stanine

    A stanine is a standardised score on a nine-point scale with a mean of 5 and a standard deviation of 2. The name contracts 'standard nine'. Each stanine covers a fixed slice of the normal distribution, so it reports a person's position deliberately coarsely — in nine bands rather than a hundred percentile points.

  • Test–retest reliability

    Test–retest reliability is the correlation between the scores of the same people measured twice with an interval between sittings. It is the most intuitive form of reliability and the least forgiving, because it exposes both measurement noise and genuine change in the person, and the two are difficult to separate.

  • Validity

    Validity is not a property of a test. It is the degree to which evidence and theory support the interpretations made of test scores for a specific proposed use. The same instrument can be valid for one decision and invalid for another, so validity is always stated in relation to a purpose.

  • Z-score

    A z-score expresses a raw score as the number of standard deviations it sits above or below the mean of its norm group. A z of 0 is exactly average and +1 is one standard deviation above. It is the common currency into which most other standard scores are converted.

Put the vocabulary to work

The four-fifths rule and the adverse impact ratio have a calculator that shows the working and tells you how many additional selections would clear the threshold.

Open the adverse impact calculator