Score comparability means a candidate's result carries the same meaning no matter which version of the test they sat. Two forms are comparable only when they are built to the same content blueprint and statistically equated, so that a 62 on Form A represents the same ability as a 62 on Form B. Shuffling questions from a pool does not achieve this.
Most hiring teams running volume assessment have quietly crossed into multi-form territory without deciding to. You started with one test. Then candidates began sharing questions, so someone switched on randomisation. Now every candidate sees a slightly different set of items — and the single cut score you validated last year is being applied to scores that are no longer strictly the same thing.
This is the FAQ we get asked most often by teams at that point.
Forms, pools and what "the same test" means
Why do I need more than one version of the same test?
Because content that is used repeatedly leaks. The International Test Commission's Guidelines on the Security of Tests, Examinations, and Other Assessments are explicit that "exposure of test or assessment content should be actively monitored and controlled" and that designs should prevent "unplanned and unmonitored over-exposure of items." Multiple equivalent forms are the ITC's first-named remedy. In campus and BPO drives where thousands of candidates sit the same assessment inside a two-week window, a single fixed form is a content-publishing exercise with extra steps.
Is randomly drawing 30 items from a 200-item pool the same as having multiple forms?
No — and this is the single most common mistake. Random draw controls exposure. It does nothing about difficulty. Unless the pool is unusually homogeneous, some 30-item draws will be measurably harder than others, and a fixed cut score will then reject candidates who happened to draw the hard set. You have not built multiple forms; you have built a lottery on top of a test.
Two things fix this. Either assemble forms deliberately against a blueprint that constrains both content coverage and statistical difficulty, or select items adaptively using calibrated item parameters, which puts every candidate on a common scale by construction.
What does equating actually do?
Equating is the statistical adjustment that makes raw scores from different forms interchangeable. It is not cosmetic rescaling. Chapter 5 of the *Standards for Educational and Psychological Testing* (AERA, APA and NCME) treats the comparability of scores across forms as a basic obligation of any programme that reports scores as if they mean one thing — and requires that the procedures and the data supporting them be documented.
Practically, equating requires a design that links forms: a set of common anchor items appearing on both forms, or a common group of candidates who sit both, or item parameters calibrated onto a shared scale. If you have none of these, you cannot equate, and you should not claim your forms are equivalent.
How big does the item bank need to be?
Bigger than intuition suggests. The standard review of exposure control strategies by Georgiadou, Triantafillou and Economides (2007) flags the ratio of pool size to test length as the binding constraint — problems concentrate where that ratio is small, because the selection algorithm keeps returning to the same small set of maximally informative items regardless of what the pool nominally contains.
A useful working rule for a fixed-form programme: you need enough calibrated items to build the number of parallel forms your volume demands, plus anchor items, plus a retirement buffer of roughly a third. A 40-item test run across four forms is not a 160-item problem; with anchors and reserve it is closer to 250.
Exposure, leakage and retakes
What exposure controls actually work?
A 2026 simulation study in *Frontiers in Education* compared six exposure-control methods against an uncontrolled baseline on a 160-item bank. The results are worth knowing:
Progressive-Restricted (a hybrid of randomisation and item information)
- Used 100% of the 160-item bank
- Mean item exposure rate of 18%
- 31% of items exceeded the 0.20 exposure target
The other five methods (Randomesque, Sympson-Hetter, unconditional and conditional multinomial, Fade-Away)
- Used 73.1% to 76.9% of the bank
- Mean exposure rates of 24% to 26%
- 41% to 44% of items exceeded the 0.20 target
Critically, measurement precision was comparable across all conditions — RMSE between 0.29 and 0.30. That is the finding that matters operationally: controlling exposure did not cost accuracy. Teams often resist exposure control on the assumption that it degrades the measure. On this evidence, it does not.
How much does a candidate gain just from having taken the test before?
Enough to matter at a cut score. The largest meta-analysis of retest effects on cognitive ability tests — Scharfen, Jansen and Holling (2018), covering 122 studies, 174 samples and 786 test outcomes across 153,185 participants — found a mean gain of 0.33 standard deviations on the first retest, rising to about 0.50 by the third administration before plateauing.
The form matters too. Identical forms produced retest gains roughly 0.15 SD larger than alternate forms at the first retest, and 0.20 SD larger by the second. Reusing the same form for retakes measurably inflates scores, and the inflation is not ability.
What should my retake policy be?
The ITC guidance is direct: "re-testing policies should be developed to reduce the opportunities for item harvesting," and a candidate "should not be allowed to retake a test that he or she 'passed' or retake a test until a set amount of time has passed."
Three decisions to make explicitly and write down:
- A waiting period. Long enough that a retake is not a second attempt at recalling items. Ninety days is a common operational floor.
- A different form. Given the 0.15 SD penalty for identical forms, serving the same content to a retaker is knowingly accepting a contaminated score.
- A cap. Unlimited retakes against a fixed cut score guarantee that persistence eventually substitutes for ability.
How would I know an item had been compromised?
By watching it, not by hoping. Two signals do most of the work. First, item drift — a question whose difficulty estimate falls over time without any change to the item is almost certainly circulating. Second, response-time anomalies — a candidate who answers a hard item correctly in a fraction of the expected time is showing a pattern that ability does not produce. Recent work in the Journal of Educational Measurement (Zopluoglu, 2026) models exactly this, detecting compromised items and candidates with preknowledge simultaneously from response-time data.
Both signals need per-item statistics you actually look at. On AssessAll, response-time and behavioural signals feed the integrity band attached to each attempt, which is the same data stream that makes item-level drift visible — but the review has to be somebody's job.
Governance
Who is accountable if scores are not comparable?
The employer. Vendors supply instruments; the EEOC's guidance on employment tests and selection procedures places the burden of job-relatedness and consistent administration on the organisation making the decision. "The platform randomised it" is not a defence if two candidates were rejected against the same number computed from materially different tests.
What should we review every quarter?
A short list is enough to catch most of it:
- Exposure rate per item, with anything above your target flagged
- Proportion of the bank actually being used, not merely loaded
- Difficulty estimates compared against the previous calibration, to catch drift
- Retake volume, and the score gap between first and subsequent attempts
- Whether the current forms are still linked by live anchor items
- Adverse impact ratios recomputed on the current forms, not the original validation sample
If your assessment is assembled or scored by a vendor, ask for these as a standing report. If nobody can produce them, that is itself the finding. Teams building role-specific content — through AssessAll's custom assessment work or in-house — should budget for item replacement from the start rather than treating a bank as a one-time build.
The takeaway
Multiple forms are not a security feature you switch on; they are a measurement commitment you take on. If you are going to serve different questions to different candidates, you owe them a blueprint, a linking design, and a number that means the same thing either way — otherwise the cut score you defend so carefully is being applied to scores that are not comparable.