All guides

How AI grading works in assessments

AI can grade written and scenario answers reliably when it's paired with rubrics and rule-based scoring. Here's how a modern grading pipeline is built, how to check it agrees with human markers, why agreement is not the same as validity, and the EU AI Act date for high-risk employment AI that moved to 2 December 2027.

Last updated

Not everything needs AI

The most reliable, lowest-cost grading is rule-based: multiple choice, true/false, numerical, and psychometric items can be scored exactly, instantly, with no model involved.

AI grading is reserved for what rules can't handle — short answers, essays, and scenario responses where meaning matters.

Rubrics make AI grading consistent

AI grading is only as good as its rubric. A clear rubric tells the model what a strong answer contains, so scoring is consistent across candidates rather than impressionistic.

AssessAll grades free-text answers against the assessment's rubric and returns not just a score but the reasoning behind it.

Escalation protects against low-confidence calls

A robust pipeline checks the grader's confidence. When an answer is borderline or the model's reasoning looks inconsistent, it's re-graded by a stronger model before the score is finalised.

This keeps cost low on the easy cases while protecting accuracy on the hard ones.

How to audit an AI grader for adverse impact

Rubrics and escalation make a grader consistent. They do not, on their own, make it fair. Consistency means the same answer gets the same score twice; fairness is a question about outcomes across groups, and the only way to know is to compute them.

The standard check is the adverse impact ratio: take each group's selection rate at the gate the grader feeds, divide by the rate of the group selected most often, and compare the result to 0.80. That threshold is the four-fifths rule from section 4(D) of the US Uniform Guidelines on Employee Selection Procedures. You can run the arithmetic on your own numbers with the adverse impact ratio calculator.

Two refinements matter when the gate is an AI grader specifically. Run the calculation per rubric criterion as well as on the overall pass, because a gap usually concentrates in one criterion rather than spreading evenly — written-fluency criteria are the usual suspects when candidates are answering in a second language. And re-run it after every model or rubric change: an AI grading pipeline is not a fixed instrument, so last quarter's audit describes a system that no longer exists.

The regulatory direction of travel now assumes this is being done. New York City's Local Law 144 has required an independent annual bias audit of automated employment decision tools, reporting selection rates and impact ratios by sex and by race or ethnicity including intersectional categories, since 5 July 2023. California's FEHA regulations on automated-decision systems took effect on 1 October 2025 and require four years of record retention for selection criteria and ADS data — and they make evidence of anti-bias testing, and of what the employer did about the results, relevant to a defence.

How to tell whether an AI grader agrees with human markers

Consistency claims are cheap. The statistic the automated-scoring field actually uses is quadratic weighted kappa, a version of Cohen's kappa that weights a disagreement by how far apart the two scores are, so a grader that is one band out is penalised far less than one that is three bands out. The conventional acceptance threshold is 0.70, following the Williamson framework for evaluating automated scoring engines against human raters.

Ask for it computed the right way: on a held-out set of responses that were double-scored by trained humans, drawn from your own candidate population rather than from the vendor's benchmark corpus, and reported alongside exact and adjacent agreement rates. A single headline accuracy figure with no kappa behind it is not evidence.

Know the statistic's limits too, because they are documented. Doewes, Kurdhi and Saxena (Educational Data Mining, 2023) show that quadratic weighted kappa moves depending on how two human scores are resolved into one, that high raw percent agreement can sit beside a low kappa, and that the statistic handles only two raters at a time. It is the best available single number and it is not the whole answer — which is why reliability figures should always be read next to the standard error of measurement in the units of the scale a candidate actually sees.

Agreement is not validity, and the difference is where AI grading fails

A grader that reproduces human scores perfectly has demonstrated agreement, not construct validity. If the human markers were themselves rewarding length, vocabulary or surface fluency, a model trained to match them will reward the same things faster and more consistently. High agreement with a flawed key is not an achievement.

The check for this is to look at what the model's scores correlate with that they should not: response length, reading-ease scores, and whether the candidate wrote in their first language. If score tracks length more strongly than it tracks any rubric criterion, the grader is measuring effort, not judgement.

Then treat every model or prompt change as a new instrument. An AI grading pipeline is not fixed hardware: a version upgrade can shift the score distribution without shifting any rubric wording, which moves everyone relative to a fixed cut score that nobody re-examined. Re-run agreement and the impact audit on the same held-out set after every change, and keep the previous version's numbers so drift is visible rather than inferred.

The kappa benchmark everybody quotes, and the sentence its own authors wrote next

Ask a vendor for an agreement figure and you will be told the kappa is 0.68, which is "substantial". That word comes from a single table in Landis and Koch's 1977 paper in Biometrics (33(1), 159–174): below 0 poor, 0.00–0.20 slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, 0.81–1.00 almost perfect. It is quoted everywhere in assessment procurement as though it were a standard.

The authors did not think it was one. In the same paper, immediately after presenting the table, they wrote that the divisions "are clearly arbitrary" and that they "do provide useful benchmarks for the discussion of the specific example" — that is, for the worked example in front of them, not as a general threshold for anyone else's instrument. A label that its own authors disclaimed in the sentence after they published it should not be the thing a buyer relies on.

What to ask for instead. First, the confusion matrix, not the summary statistic: kappa collapses every kind of disagreement into one number, so a grader that is one band out on many responses and a grader that is catastrophically wrong on a few can report the same figure. Second, agreement computed per rubric criterion, because pipelines usually fail on one criterion rather than uniformly. Third, the base rates, since kappa falls when a category is rare even where raw agreement is high — the prevalence paradox that makes a high-stakes, low-frequency criterion look worse than a trivial one.

For free-text grading with more than two categories, Krippendorff's alpha is the better-behaved statistic — it handles ordinal data, more than two raters and missing judgements, none of which kappa does gracefully. Krippendorff's own recommendation, in Content Analysis: An Introduction to Its Methodology (2nd ed., 2004, pp. 241–243), is to rely on data at α ≥ 0.800, to draw only tentative conclusions between 0.667 and 0.800, and to discard data below 0.667. Those bands are stricter than the ones the market quotes, and they come with a stated rationale rather than a disclaimer.

The EU AI Act date moved — and the one everyone is quoting is now wrong (verified 6 September 2026)

Anyone who built an AI-grading compliance plan in 2024 or 2025 put 2 August 2026 in it. That was the correct reading of Article 113 of Regulation (EU) 2024/1689: the general application date, which is when the Chapter III obligations for Annex III high-risk systems were to bite. Annex III point 4(a) covers AI systems intended for "the recruitment or selection of natural persons, in particular to place targeted job advertisements, to analyse and filter job applications, and to evaluate candidates" — which is an AI grader used for hiring, without ambiguity.

That date no longer applies. Regulation (EU) 2026/1744 of 8 July 2026, the Digital Omnibus on AI, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026. It moves the Chapter III obligations for Annex III systems to 2 December 2027, and those for Annex I embedded-product systems to 2 August 2028. The new dates are fixed calendar dates in the adopted text; they are not conditional on harmonised standards being ready, even though standards delays are what the recitals give as the reason.

Two practical warnings. The Commission's own AI Act Service Desk was, at the time of checking, still displaying the unamended Article 113 with a banner saying the page had not yet been updated — so a plan checked against that page will read the superseded date. And Annex III itself was not amended: the scope of what counts as high-risk employment AI is unchanged, only the date it applies from. A vendor that describes the delay as a narrowing of scope has misread it.

What the delay is not is a reason to defer the work. Every obligation the Act imposes on a high-risk employment system — logging, human oversight, a documented risk assessment, technical documentation, accuracy and robustness testing, and a bias examination of the training data — is something a buyer should already be asking a grading vendor for, because each one is also the evidence that the grader works. Sixteen extra months is time to have the documentation rather than a reason not to.

Frequently asked questions

How do you check an AI grader for bias?

Compute the adverse impact ratio at the gate the grader feeds: each group's selection rate divided by the highest group's rate, compared against 0.80. Do it per rubric criterion as well as overall, because a gap usually concentrates in one criterion, and re-run it after every model or rubric change — an AI grading pipeline is not a fixed instrument, so an old audit describes a system that no longer exists.

What agreement statistic should I ask an AI grading vendor for?

Quadratic weighted kappa against trained human double-scoring, with 0.70 as the conventional acceptance threshold, reported alongside exact and adjacent agreement rates. Ask for it on a held-out sample from your own candidate population rather than the vendor's benchmark corpus, because agreement is population-specific and a corpus chosen by the vendor is not a test.

Does high agreement with human raters mean the AI grader is valid?

No. Agreement means the model reproduces what the human markers did. If those markers were rewarding length or surface fluency, a model that matches them reproduces that too. The diagnostic is to check what the scores correlate with that they should not — response length, reading ease, whether the candidate wrote in a second language. If score tracks length more strongly than it tracks any rubric criterion, the grader is measuring effort.

How often should an AI grading pipeline be re-validated?

After every model version change, every prompt change and every rubric change, and on a fixed calendar in between. A version upgrade can shift the score distribution without any wording changing, which quietly moves every candidate relative to a cut score nobody re-examined. Keep the previous version's agreement and impact figures so that drift is visible rather than inferred.

Try it yourself
Take a free assessment and start your Skill Passport.
Browse catalogue