Marking and Rating Consistency Assessment for Assessors and ReviewersAgreement is not accuracy, and severity is not rigour. This tells you which one you are.
Sixteen work samples against a four-level rubric printed as the four answers, with the credit table published on your report.
Everybody on a marking panel knows this happens. Almost nobody knows which of the two they are.
Two assessors read the same piece and give it two different levels. It happens on marking panels, in moderation meetings, at interview scoring, in performance calibration and in peer review. The statistics that describe it are computed on a panel after the fact, and there is nothing an individual assessor can sit that tells them where they personally stand before the panel meets.
Sixteen short work samples are put in front of you — a handover note, a closing argument, a risk register, a test plan, a discharge summary — with the rubric printed as the four answers, so nothing here tests whether you remembered the standard. Four samples warrant each level, and several sit deliberately on a boundary, which is where panels actually lose agreement.
Credit is graded by distance from the warranted level, using a table that is printed on your report. It is built so that the four columns have identical means: calling everything Meets scores exactly what calling everything Not yet scores, and every fixed answer scores behind reading the work. That property is the difference between a scoring model and an opinion with a number attached, and it is demonstrated on the page rather than asserted.
Three numbers come out and they answer different questions. How closely you agree with the standard. Whether you sit above it or below it. And whether you use the whole scale or crowd the middle two levels, which looks careful and removes the information the rubric exists to produce.
The table deliberately does not pay more for marking down than for marking up near the middle of the scale, because a table that did would train harshness and call it rigour. Direction is reported instead as its own signed number, beside the agreement figure and never folded into it.
What you walk away with
What the credit table awarded, corrected against what a typical respondent would score, with both spans drawn and written out.
How far above or below the standard you sit, on average, reported apart from agreement because they are different questions.
Whether you use the range the work needs, or crowd the middle, or spend the top and bottom levels on work that does not warrant them.
All sixteen cells with their column means, so you can see why no fixed answer pays and what value judgement is inside the scoring.
The level you called against the level warranted, sorted by gap, so the ones worth arguing about are at the top.
Set at the boundary your own calls actually moved at, naming a specific check before you award a level.
Inside your report
Illustrative sample — your report is generated from your own responses.
Sorted by gap, largest first, so the samples worth arguing about are the ones at the top.
| Warranted | Not yet | Approaching | Meets | Exceeds |
|---|---|---|---|---|
| Not yet | 3 | 2 | 0 | 0 |
| Approaching | 2 | 3 | 2 | 1 |
| Meets | 1 | 1 | 3 | 2 |
| Exceeds | 0 | 0 | 1 | 3 |
| Column mean | 1.50 | 1.50 | 1.50 | 1.50 |
The four column means are identical by construction: calling everything Meets scores exactly what any other fixed answer scores, and all of them score behind reading the work.
Built for
- Assessors, examiners, moderators and internal verifiers in education and vocational training
- Interview panels and assessment-centre assessors who score against a rubric
- Performance calibration groups and grading committees
- Peer reviewers, quality reviewers and content moderation QA leads
Find out whether you are the severe one
40 exercises across six formats · about 40 minutes · agreement, severity and spread reported as three separate numbers.
₹1,199 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
No. The rubric is printed as the four answers on every call, and the samples are short and drawn from several fields deliberately, so what is measured is how you apply a stated standard rather than what you know about nursing, law, engineering or teaching.
Because they are different questions with different remedies. An assessor can be reliably one level below the standard on every piece and agree perfectly with themselves. Folding direction into the headline would also mean paying more for harshness, which trains harshness and calls it rigour.
Because it is a value judgement rather than a measurement: it says which confusion between two levels costs least. A judgement made on your behalf should be visible to you, and the table also shows why calling everything Meets scores no better than any other fixed answer.
No. It carries no pass mark, and a low result is a reason for calibration and anchor papers rather than for removing somebody from a panel. Agreement with a standard is not the same thing as accuracy, and the report says so.
About forty minutes for forty exercises. ₹1,199 in India, inclusive of GST, or US$11.99 elsewhere, one time, for the sitting and the full report.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Forty original exercises across six formats: sixteen anchored level calls against a printed four-level rubric, six most-and-least forced choices, five matching exercises, six keyed claims, four select-every-that-applies exercises and three panel-expectation estimates. Construct statement: it measures how closely somebody's level calls match the level the work warrants, where they sit relative to that standard, and whether they can name the evidence and the rating errors behind a set of marks. It does not measure subject expertise in any of the domains the samples are drawn from, does not certify anybody as an assessor, and is not a selection instrument for a marking panel. Declared response instruction: knowledge and rule application, one instruction for the whole instrument. Scoring is a warranted-level match (C16). Credit comes from a four-by-four table, one asymmetric row per warranted level, published on the report. The rows are chosen so that the four columns have identical means across an evenly warranted bank, which is the property that makes every fixed answer score the same and all of them score behind reading the work. The bank is evenly warranted by construction, four samples per level, and the balance is audited empirically rather than asserted. The direction of the error is deliberately NOT priced into the headline near the middle of the scale, because a table that pays more for harshness trains harshness; direction is reported instead as a signed severity index beside the headline, and the spread of levels used is reported as a third number, so a consistent, severe, narrow marker and a consistent, generous, wide one are visibly different people. Every strand is chance-corrected against each item's own authored answer priors rather than against a uniform guess. Competencies carry three-way placements rather than numbers wherever fewer than eight exercises support them, McDonald's omega is reported rather than alpha, and the omega is an assumption stated in the open until live data exists. Constructs and sources: the accuracy components of interpersonal and performance judgement, elevation, differential elevation, stereotype accuracy and differential accuracy (Cronbach 1955); rating errors and their distributional fingerprints, halo, leniency, severity, central tendency, contrast and recency (Thorndike on halo; Saal, Downey and Lahey on rating quality; Murphy and Cleveland on performance appraisal); chance-corrected agreement and its interpretation (Cohen's kappa; Fleiss; Landis and Koch on benchmarks, with their arbitrariness noted); the many-facet measurement of rater severity and its separation from candidate ability (Linacre); criterion-referenced standard setting and the modified Angoff method (Angoff; Cizek and Bunch); anchor papers and exemplar-based calibration in marking (Sadler on criteria and standards); the effect of scale length on agreement; and item-writing guidance from Haladyna, Downing and Rodriguez. All items are original works written for this instrument, the rubric wording is our own, and no commercial instrument, exam board scheme or trademarked scale is reproduced or implied.