An assessment AI-use policy is a documented decision about which tools a candidate may use at each stage, enforced through administration conditions and recorded alongside the score. It is a measurement decision before it is an integrity rule: permitting a tool changes what the score ranks people on, and how far apart those ranks sit.
Most of the argument about candidates and AI is an argument about cheating rates. That is the wrong quantity to start with.
The incidence numbers are lower than the discourse
The best current evidence comes from Robie, Wingate, Baytalskaya and Butera (2026) in the International Journal of Selection and Assessment, who asked real applicants in real pre-hire assessments what they had used.
- In Study 1, run in Q3 2024 with 5,675 applicants, fewer than 3% reported using generative AI — though up to 19% reported using it in combination with algorithmic resources such as search engines.
- Study 2, with 3,356 participants in Q3 2025, found self-reported adoption rising year over year.
- All three warning statements tested — consequences-based, educational, and reasoning-based — reduced reported use relative to a control, with limited evidence favouring the consequences-based version. Warnings moved behaviour without changing candidates' underlying motivations.
- Most usefully: generative AI use on its own showed no systematic relationship with job-fit scores.
Self-report under-counts, and nobody should read 3% as a true base rate. But the finding that use did not systematically shift scores should reframe the problem: the risk is not that a few applicants gain an undeserved point. It is what happens to the distribution once use is normal.
Variance is the thing actually at risk
A score ranks candidates only to the extent that candidates differ on it. That is not a philosophical point: reliability, standard error and every selection decision are functions of true-score variance. In a study in Scientific Reports, Mittelstädt, Maier, Goerke, Zinn and Hermes put five chatbots through the Situational Judgment Test for Teamwork — 12 scenarios, four options each, effectiveness rated by 109 experts, scored from −24 to +24 — and compared them with 276 pilot applicants.
Machine scores, with their standard deviations
- Claude 3.5 Sonnet: 19.4 (SD 0.66)
- Microsoft Copilot: 17.5 (SD 1.36)
- You.com: 16.8 (SD 1.40)
- ChatGPT: 14.5 (SD 0.81)
- Google Gemini: 13.9 (SD 1.14)
Human scores
- 276 applicants: 14.2 (SD 3.27)
Three of the five beat the human sample significantly; two matched it. The means are the headline, but the standard deviations are the finding: human spread was two to five times the machine spread. A tool that answers a teamwork SJT around or above the human mean with almost no variability does not merely inflate scores — it collapses the top of the distribution into a band too narrow to rank. You cannot select on a variable that has stopped varying.
AI assistance does not move everyone in the same direction
The second reason to treat this as measurement rather than policing is that the evidence on assisted performance points both ways.
- Compression. In a randomised experiment published in Science, Noy and Zhang gave occupation-specific writing tasks to 453 college-educated professionals: time fell 40%, graded quality rose 18%, and "inequality between workers decreased."
- Compression, concentrated at the bottom. Brynjolfsson, Li and Raymond, in the Quarterly Journal of Economics, studied more than 5,000 customer-support agents: issues resolved per hour rose 14% on average, 34% for novice and low-skilled workers, with minimal impact on experienced ones.
- Compression inside the frontier, damage outside it. Dell'Acqua and colleagues randomised 758 BCG consultants. On tasks AI handled well, they completed 12.2% more, 25.1% faster, at more than 40% higher quality — with below-average performers gaining 43% against 17% for those above average. On a task placed deliberately beyond the model's competence, AI users were 19 percentage points less likely to reach the correct solution.
- Divergence. Otis, Clarke, Delecourt, Holtz and Koning gave Kenyan entrepreneurs a GPT-4 business assistant over WhatsApp. There was no significant average effect, but high baseline performers gained just over 15% while low performers did about 8% worse — a treatment-effect gap of 0.27 standard deviations.
Three settings where assistance narrowed the gap, one where it widened it, and one where it misled people on the tasks that mattered. No general law tells you which case your assessment is. That is an empirical question about your instrument and your candidates, so it has to be measured, not assumed.
Six decisions behind an AI-use policy
The walkthrough below is illustrative — a composite of the decisions such a policy requires, not an account of any specific organisation.
1. Name the construct, then ask whether the tool is part of it
Job-relatedness is the standard, as 29 CFR 1607 has required since 1978 and the *Standards for Educational and Psychological Testing* restate. If the role is performed with AI tools daily, an AI-free score measures something the job does not require. If it turns on unaided judgement under time pressure, it does not.
2. Split the funnel into a secured lane and an open lane
The University of Sydney's two-lane model is the cleanest version: lane one is secured and identity-assured, carrying the decisions; lane two is open, with tools permitted and scaffolded, carrying development. It transfers directly to hiring. A 2025 critique in *Higher Education Research & Development* argues that applying it as an all-or-none rule is insupportable — worth reading before you adopt it wholesale.
3. Decide per instrument, not per funnel
Verbal, numerical and situational-judgment content is most exposed, because the construct lives entirely in the response text. Structured interviews, physical work samples and interactive simulations are far less so. One blanket rule across a mixed battery is almost always wrong somewhere.
4. Put the rule where the candidate reads it, and warn
The UK Civil Service Fast Stream publishes its line explicitly: candidates may use generative AI for research, practice and feedback on their own drafts, but "must not use these tools during the actual completion of the tests and assessments." The Robie findings suggest a stated consequence does measurable work.
5. Enforce with conditions, not inference
Detection claims age badly; administration conditions do not. The ITC and ATP Guidelines for Technology-Based Assessment treat administration mode as part of a score's meaning. Record the mode, and report integrity as a graded band with the evidence behind it rather than a binary verdict — which is why AssessAll returns proctoring signals as integrity bands rather than a pass/fail flag a recruiter cannot interrogate.
6. Re-norm before you re-decide
A cut score set on unaided administrations does not survive a change of rules, and percentile norms built on an earlier candidate population quietly expire. If you permit tools, re-run item statistics and check the spread before you keep using the old threshold.
Permit or prohibit: the honest comparison
Permitting AI fits better when
- The job is done with the same tools, and tool fluency is part of the competency.
- You can score the process — what the candidate asked for, what they rejected, what they verified — not just the artefact.
- The exercise sits outside the model's reliable range, where the BCG result suggests discrimination may actually improve.
- The assessment is developmental, where Dunlop and Lievens (2026) place the "reimagine" end of the employer response spectrum.
Prohibiting AI fits better when
- The construct is unaided reasoning, recall under load, or live interaction.
- You are using published norms or a standard-set cut score that assumed unaided conditions.
- The stakes justify secured, identity-assured administration and you can actually deliver it.
- Volumes are high and consistency of conditions is what keeps the decision defensible.
Neither answer is universally right, and a mixed battery usually needs both.
What to write down
- Which tools were permitted, at which stage, in the candidate-facing instructions.
- The administration mode and integrity evidence for every score used in a decision.
- The date the norms or cut score were established, and under which rules.
- The job-analysis basis for treating tool use as part of the construct, or not.
The question worth asking is not how many candidates use AI, but whether your scores still spread candidates far enough apart to support the decision you make with them — and that is a number you can check this quarter, on data you already hold.