All articles
L&D & Capability21 September 2026·6 min read

The Rater Is the Instrument: Why Calibrating Assessors Beats Redesigning the Rating Form

Only 21-25% of variance in performance ratings comes from the person being rated. Frame-of-reference training moves that number - trained assessors tracked actual performance in 54.6% of rating variance vs 33.5% untrained. How to run a calibration session that works.

By AssessAll Editorial

Rater calibration — known in the research literature as frame-of-reference training — is the practice of giving every assessor the same mental model of what each performance level looks like, by having them score shared reference performances and then reconciling their scores against an agreed standard. It targets the judge, not the rating form.

Most organisations that are unhappy with their capability ratings respond by redesigning the form. They add behavioural anchors, rewrite the level descriptors, move from five points to four, commission a new competency dictionary. Then the ratings come back looking much the same, and the conclusion is that "ratings just don't work."

The evidence points somewhere less comfortable. The largest single source of noise in a rating is not the instrument. It is the standard sitting in the head of the person holding it.

How little of a rating is actually about the person being rated

Scullen, Mount and Goff's variance decomposition of job performance ratings found that only 21–25% of the variance in ratings was attributable to ratee main effects — that is, to the person being assessed (as summarised in *Industrial and Organizational Psychology*, 2024). The remainder is carried by the rater: their idiosyncratic view of the dimension, their internal severity setting, their sense of what "good" means for this role.

That is not a claim that raters are careless. It is a claim that two conscientious assessors, given the same form and the same candidate, are working from different pictures of what a 4 out of 5 looks like — and nothing in the form tells them otherwise.

What frame-of-reference training actually does

Frame-of-reference (FOR) training has four components, in this order:

  1. Define the dimension in behavioural terms — not "communication" but the specific observable actions that constitute it in this role.
  2. Specify what each performance level looks like, with behavioural examples at the low, borderline and high ends.
  3. Have assessors independently score shared reference performances — recorded, written or live — before seeing anyone else's score.
  4. Reveal the expert standard and reconcile the gaps, discussing why a behaviour maps to one level rather than the adjacent one.

The mechanism is a shared schema. Raters stop importing their private theory of the dimension and start applying a common one.

A 2024 study in Industrial and Organizational Psychology put numbers on the effect. Comparing FOR-trained assessors with an untrained control group, the researchers found that the candidate's actual performance level accounted for 54.6% of rating variance among trained assessors, versus 33.5% in the control group — and that assessor-specific effects shrank correspondingly, with the assessor-by-performance-level interaction falling from 22.8% to 13.0% (Gorman, Jackson, Meriac, Himmler & Contreras, 2024). Training also helped most where it is needed most: trained assessors outperformed untrained ones specifically at identifying poor performance, on five of six dimensions.

Two caveats worth stating plainly: that study used only two stimulus videos, and the authors themselves note it is unclear which components of FOR training carry the effect. The broader meta-analytic picture is more established — Roch, Woehr, Mishra and Kieszczynska's 2012 review in the Journal of Occupational and Organizational Psychology found FOR training effective across settings — but the exact size of the gain in any given programme is an empirical question for that programme, not a number to borrow.

The sequencing detail most calibration sessions get wrong

Running a calibration session badly is easy: show the reference performance, announce the expert score, ask if everyone agrees, move on. Everyone agrees. Nothing changes.

A 2025 realist evaluation of video-based benchmarking with 16 clinical examiners found the opposite sequence matters (Edwards, Yeates, Lefroy & McKinley, *Advances in Health Sciences Education*). Their findings map cleanly onto corporate calibration:

  • Assessors must score before they see the benchmark. Examiners who scored first and then saw the expert standard engaged far more than those given the standard up front. The discrepancy is the learning event; removing it removes the learning.
  • Two to three reference performances at different levels is the working range — enough to bracket the standard, not so many that cognitive load swamps the exercise.
  • Timing matters. The researchers found calibration was most useful shortly before live assessment — within about a day — rather than weeks earlier or in the minutes immediately preceding it.
  • Examiners remained genuinely uncertain about the standard until they saw concrete examples, despite having read the materials conscientiously. Reading the rubric is not calibration.

Fix the scale while you are in there

Calibration is the high-leverage intervention, but scale design is not free. Preston and Colman's study of response category counts in Acta Psychologica found reliability, validity and discriminating power all lowest for two-, three- and four-point scales, rising steeply and then plateauing at roughly seven categories.

If your capability rubric uses a four-point scale because "it forces a decision," you are paying for that decisiveness in measurement quality. Five to seven anchored levels, each with behavioural descriptors your assessors have practised applying, is the better default.

Your AI scorer has a frame of reference too — and it may not be yours

The same logic applies to machine raters, and the 2026 evidence is instructive. A study of ChatGPT scoring 192 EFL essays against an analytic rubric found the model highly consistent with itself across two rounds three weeks apart — quadratic weighted kappa of 0.862 and an intraclass correlation of 0.892 on total score — while agreement with human teachers was only moderate, at QWK 0.555, and the model scored lower than humans on 67.2% of essays (AlAmir, *Frontiers in Education*, 2026).

That is a calibration problem, not a reliability problem. An AI scorer with a stable but different standard will reproduce that difference at scale, every time, silently. The remedy is the same as for humans: score a set of reference performances, compare machine scores against your expert standard, and adjust the rubric and scoring prompt until they converge — then re-check periodically. When we build AI-graded scenario assessments at AssessAll, the reference set and the agreement check against client subject-matter experts are part of the build, not an optional extra.

A calibration session you can run next week

  1. Pick one dimension. Do not try to calibrate a twelve-competency framework in one sitting.
  2. Write behavioural descriptors for the low, borderline and high levels — the borderline one matters most, because that is where decisions are contested.
  3. Assemble three reference performances: one clearly weak, one borderline, one strong. Recorded responses, work samples or written submissions all work.
  4. Have a small expert panel score them independently and agree a defensible consensus score with reasons.
  5. In the session, assessors score all three independently, in silence, before any discussion.
  6. Reveal the panel scores. Spend the time on the gaps, especially the borderline case.
  7. Record the session: who attended, which references were used, what the agreed standard was. Both the SIOP Principles and the *Standards for Educational and Psychological Testing* treat documentation of scoring procedures as part of the validity evidence for a selection or promotion decision — and US EEOC guidance treats rating-based judgments used in employment decisions as selection procedures like any other.
  8. Re-calibrate before each assessment cycle. Standards drift, and new assessors arrive uncalibrated.

When this is not the right investment

Calibration earns its keep where human judgment is unavoidable and consequential: promotion panels, development centres, structured interviews, manager ratings feeding a HiPo decision. It is wasted effort where the judgment can be removed entirely. If a capability can be measured with a scored work sample or a machine-marked scenario, do that instead — a calibrated panel is still slower and more expensive than an instrument that does not need one.

The takeaway: if your capability ratings are noisy, the cheapest experiment available is not a new rubric — it is three reference performances, an expert consensus score, and one hour in a room where your assessors discover how far apart they actually are.

#rater-calibration#frame-of-reference-training#performance-ratings#assessor-training#rubrics

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free