Item analysis is the statistical review of individual test questions after candidates have answered them: how hard each item proved, how well it separated stronger candidates from weaker ones, whether every answer option attracted anyone, and whether it behaved differently across groups. It is the difference between a test you believe in and a test you can defend.
Almost nobody does it. Questions get written, reviewed by a subject-matter expert, approved, and then run for years without anyone looking at what the response data says. The review meeting feels like quality control. The evidence says it mostly isn't.
The distractor nobody picks
The clearest evidence comes from a study of the Swiss Federal medical licensing examinations. Researchers analysed 737 items containing 2,948 distractors administered between 2005 and 2007. Using a threshold of 1% selection — a distractor picked by fewer than one candidate in a hundred is not doing any work — they found:
- 31.5% of all distractors (929 of 2,948) were selected by under 1% of candidates
- 70% of five-option items contained at least one non-functional distractor
- Only 30.3% of items had all four distractors functioning
Raise the threshold to a still-generous 5% and the picture collapses further: 68.3% of distractors fell below it, and just 2.8% of items had a full set of working options. (Rogausch et al., *BMC Medical Education*, 2010)
This is a high-stakes national licensing exam with professional item-writing infrastructure. A four-option question where two options are dead is a two-option question — a coin flip dressed as a test. Candidates find these quickly. Your score distribution absorbs the noise silently.
Expert review is a weaker control than it feels
The same study did something more uncomfortable. It asked 37 practising clinicians to predict, for a set of internal medicine items, which distractors candidates would rarely select.
- Individual experts averaged 64% accuracy at identifying the under-1% distractors, with per-item hit rates ranging from 10% to 100%
- A consensus panel did better — 74% at the 1% threshold, 93% at the looser 5% threshold
Read that generously and the finding is still damning for the way most organisations work. A single reviewer — the standard model in corporate assessment, where one manager signs off on a question bank — misses roughly a third of the dead options. Panels help, but panels are expensive and most teams do not convene them. The response data, by contrast, is free and already sitting in your platform.
The point is not that experts are unskilled. Predicting how a thousand strangers will misread a question is simply not a task intuition is built for. That is what the statistics are for.
The four numbers worth pulling
You do not need item response theory to start. Four classical statistics catch most of the damage.
Difficulty index (p-value) — the proportion of candidates answering correctly.
- Useful range for a selection test is roughly 0.30 to 0.80
- Items above 0.95 or below 0.10 carry almost no information; they separate nobody
- In a 2024 distractor-efficiency study of 59 items, mean item difficulty was 37.5 (SD 19.1), with 72.9% of items in the acceptable band (*BMC Medical Education*, 2024)
Discrimination index — how much better top scorers do on this item than bottom scorers.
- Above 0.30 is good; 0.20 to 0.29 is marginal; below 0.20 needs review
- Negative discrimination is an emergency: weaker candidates outperforming stronger ones almost always means a miskeyed answer or an ambiguous stem
- That same study reported a mean discrimination index of 0.46 (SD 0.22), with 69.5% of items discriminating well
Distractor selection rate — what share of candidates chose each wrong option.
- Any option under 5% is a candidate for rewriting or deletion
- In that study, distractor efficiency correlated moderately and negatively with the difficulty index (r = −0.548, p < 0.001) — items with more working distractors were genuinely harder
Differential item functioning (DIF) — whether candidates of equal underlying ability from different groups have different odds of answering correctly.
- DIF is not the same as a group score gap; it isolates the item as the source
- The *Standards for Educational and Psychological Testing* (AERA/APA/NCME, 2014) treat item-level fairness review as part of test development, not an optional audit
- The EEOC's guidance on employment tests and the Uniform Guidelines' technical standards expect documented evidence for the procedure you actually used — and under Annex III of the EU AI Act, AI systems used in recruitment and worker evaluation sit in the high-risk category, where that documentation stops being optional
The AI item-generation problem is an item-analysis problem
Generating assessment items with a large language model is now trivial, which is why this matters more in 2026 than it did in 2016. The volume arrives; the quality control does not scale with it.
The AI-GENIE study, published in Behavior Research Methods in 2026, is instructive because its authors were sympathetic to the method. Their pipeline generates items with an LLM and then screens them statistically — embedding the items, running exploratory graph analysis, and removing redundant and unstable ones. In their GPT-4o worked example, 181 of 320 generated personality items survived the screen, meaning roughly 43% were discarded. Retention varied sharply by construct: 77% for openness, 42% for agreeableness and extraversion (*Behavior Research Methods*, 2026).
That is a well-designed system throwing away two items in five. An organisation that generates a hundred questions, reads them over, and ships them is skipping the step that made the research defensible. At AssessAll we treat AI-graded scenario items the same way we treat any other item — the model drafts, the response data decides which drafts survive contact with real candidates.
Fewer options, better options
There is a simple structural fix hiding in this evidence. If most distractors don't function, stop writing so many.
Rodriguez's meta-analysis of 80 years of research — 27 studies, 56 independent trials, 2,406 items, 12,591 participants — found that moving from four options to three produced a slightly easier item (difficulty +.044) with marginally better discrimination (+.031) and better reliability (+.019). Moving from five to three left discrimination essentially unchanged (−.004). Reliability dropped when distractors were deleted at random (−.059) but held when the ineffective ones were removed deliberately (+.006). One included study found the three-option form took 17% less time to administer (Rodriguez, *Educational Measurement: Issues and Practice*, 2005).
The practical reading: one correct answer and two genuinely plausible distractors beat one correct answer and four, three of which are filler. You get more items per hour of candidate time, and every option earns its place.
An operational routine
- Set a review trigger, not a review calendar. Pull item statistics once an item has 100+ responses, not once a quarter.
- Flag on four rules. Negative discrimination, discrimination below 0.20, difficulty above 0.95 or below 0.10, any distractor under 5%.
- Triage negatives first. Check the answer key before you touch the wording. Miskeying is the most common cause and the fastest fix.
- Rewrite, don't delete, by default. A dead distractor usually means an implausible option, not a bad question. Replace it with a real candidate misconception.
- Run DIF where your volumes allow it. You need reasonable subgroup sample sizes; below a few hundred per group, the statistic is unstable and should be read as a prompt for human review rather than a verdict.
- Document what you changed and why. This is the artefact regulators and litigators ask for, and the one nobody has.
Where judgment still rules
Statistics flag items; they do not diagnose them. An item can discriminate beautifully and still measure test-wiseness rather than the competency you care about — statistics cannot see construct irrelevance. And for genuinely low-volume assessments — a senior leadership panel exercise, a specialist role with six candidates a year — classical item statistics will never stabilise, and expert review plus a structured rubric remains the right control. Knowing which regime you are in is the kind of question our custom assessment work usually starts with.
The takeaway
The cheapest quality improvement available to most assessment programmes is not a new platform or a new competency model — it is reading the response data you already have, four statistics at a time. Expert review catches roughly two-thirds of broken items; the data catches the rest for free.