What is Item analysis?
Also called Classical item statistics
Item analysis is the statistical review of how each question in a test performed. Two numbers carry most of it: difficulty, the proportion of candidates who answered correctly, and discrimination, the correlation between getting that question right and scoring well on the test overall. Weak items are revised or retired on this evidence.
Difficulty, and why the extremes are wasted
Difficulty is conventionally written as p, the proportion correct — so a high p means an easy item, which reads backwards until you are used to it. An item everybody answers correctly and an item nobody answers correctly both carry no information about differences between candidates, however good the question looks.
For a test intended to spread people out, items around the middle of the range do most of the work. For a test intended to certify against a fixed standard, the useful items are those near the cut score, which is a different target and implies a different bank.
Discrimination is the number that finds broken items
Discrimination is usually the point-biserial correlation between the item and the total score. A value near zero means the item is unrelated to whatever the rest of the test measures. A negative value means strong candidates are getting it wrong more often than weak ones — almost always a miskeyed answer, an ambiguous stem, or two defensible options.
A common working rule is to review anything below about 0.20 and to treat any negative value as a defect until proven otherwise. Both figures depend on the sample, so they are unstable on small numbers of responses and should not be acted on before a bank has accumulated enough sittings.
Reading three items
- Item 12 — p = 0.52, point-biserial 0.41 → healthy, discriminating, keep
- Item 13 — p = 0.97, point-biserial 0.04 → everyone gets it; carries no information
- Item 14 — p = 0.44, point-biserial −0.18 → strong candidates failing it; check the key
Not the same as item response theory
Classical item analysis produces statistics that depend on the sample who happened to sit the test. Item response theory estimates item parameters that are, in principle, independent of the sample — which is what makes item banking and adaptive delivery possible.
Item response theoryRelated terms
More terms beginning with I
Check a selection process against the four-fifths rule
Free, no signup, computed in your browser — with the remedy, not just the verdict.
Last reviewed