Applied Judgment Assessment
Applied skill assessment · teachers, trainers, assessors and moderators · browse the full catalogue

Marking Consistency and Assessment Design Assessment for Teachers, Trainers and AssessorsTwo markers, one script, two grades. One of them is yours.

Twelve pieces of student work a moderation panel has already marked. You mark them too — and two of the twelve are the same script twice.

45 minutes36 scored exercisesEvidence-keyed scoringGlobal · INR & USD

The market sells courses about marking consistency and software for statisticians. There is nothing in between.

A chartered assessment qualification runs to about fourteen hundred pounds across three stages. A one-day assessor and moderator course is published at nearly two thousand dollars. At the other end, the standard software for modelling judge severity costs a hundred and forty-nine dollars for a single user — and it is analysis software that needs a rater-by-script matrix you do not have and expertise most markers do not want.

Between those two there is no product that hands an individual marker a number describing their own harshness and their own consistency. Exam boards run standardisation and publish nothing about it.

Three figures come out of your twelve marks and they do not move together. How harsh or generous you are is one number, and it is the easiest of the three to fix because a moderator can correct a steady offset in one step. How widely you spread the marks is a second, and it decides how much any single misjudgement costs a student. Whether you agree with the panel about which script is better than which is a third — and it is measured AFTER your harshness and your spread have been removed, so a severe marker with perfect ordering scores high on it.

That third figure is corrected against what markers actually give each script, script by script, computed by simulation. Not against a uniform draw across the scale. This matters more than it sounds: real markers cluster in the middle, so marking thirteen to everything lands close to the key by accident, and a scorer that does not correct for it produces the result where everybody is broadly in line with the standard.

Two of the twelve are the same work put in front of you twice under different labels, several exercises apart. One is the fluent script that says nothing and one is the badly written script that says everything — the two traps in the set. The distance between the two marks you gave is the only figure on the page about you against yourself.

Four parts, the first measured by twelve marked scripts and the rest by at least six exercises each:
Agreeing with the standardWriting a question that measuresWhere marking goes wrongFeedback that changes the work

What you walk away with

Agreement, with harshness removed

Whether you order the scripts as the panel does, independent of how hard you mark.

Your severity, as one number

Your average against the panel's, which is correctable in a single step and is not a fault.

Your spread against theirs

Marking toward the middle is invisible on any one script and costs students the top and the bottom.

How close you are to yourself

Two seeded repeat scripts, and the distance between the two marks you gave them.

Every script, biggest gap first

Your mark and the panel's on one line, sorted so the script worth re-reading is at the top.

Question design and feedback

Two keyed parts on writing a question that measures and feedback a learner acts on.

Inside your report

Illustrative sample — your report is generated from your own responses.

Biggest gap first
C — fluent, says nothing
D — badly written, says it all
A — all three parts present

Square is the moderation panel, circle is you. Two shapes, not two colours, so it survives a laser printer.

Three figures that do not move together
Agreement, harshness removed
62
band 44 to 80
Severity
-2.4
marks out of 20
Spread
0.6×
narrower than the panel

Two marks harsh with a perfect ordering, and dead-on average with none, produce the same single accuracy percentage everywhere else. They are two entirely different problems.

Built for

  • Teachers and lecturers marking coursework against a rubric
  • Trainers and vocational assessors signing off competence
  • Moderators and internal verifiers running standardisation
  • Heads of department deciding where a marking disagreement actually comes from

Find out where your marking sits against a moderated standard

36 exercises across five formats · about 45 minutes · agreement, severity, spread and self-consistency, each printed apart.

₹899 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and Razorpay sends your receipt to that address.

Secure Razorpay payment · ₹899 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

Do I need to teach a particular subject?

No. The twelve scripts answer one short, general workplace task and the rubric is printed with every one of them. Nothing in the set requires knowledge of any discipline, syllabus or exam specification.

Is being a harsh marker a bad result?

No, and the report says so twice. A steady offset is the easiest rater effect to handle because it is a single number a moderator can correct. Erratic marking with no offset is far more damaging and much harder to see, which is why the three figures are kept apart.

What if I give every script the same mark?

Then the agreement figure is withheld and the reason is printed. A set of marks with no spread has no ordering in it, and inventing one would be manufacturing a finding out of the arithmetic rather than out of your marking.

Is this a qualification, or does it certify anything?

No. It awards nothing and accredits nothing. It is a developmental report with no pass mark, and it should never be used on its own to decide whether somebody marks, assesses or teaches.

How long is it and what does it cost?

About forty-five minutes for thirty-six exercises. ₹899 in India, inclusive of GST, or US$8.99 elsewhere, one time, for the sitting and the full report.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue

Methodology: Thirty-six original exercises across five formats: twelve marking exercises in which a piece of student work is marked out of twenty on a printed three-part rubric, ten keyed single-choice exercises, five select-every-that-applies exercises, five binary claims and four matching exercises. Two response instructions are declared. The marking strand is a performance task scored against a moderated panel mark. The other twenty-four exercises carry a knowledge instruction - what is most likely to be true, or most effective. The two are reported apart and never averaged, because agreeing with a panel and knowing why markers disagree are different things and a marker can be good at one and poor at the other. Construct statement: it measures how closely somebody's marking matches a moderated standard once their own severity and spread have been removed, how consistent they are with themselves on a script they mark twice, and what they know about writing questions that measure, about the rater effects that distort a mark, and about feedback a learner acts on. It does not measure subject knowledge in any discipline, teaching quality, classroom practice, or fitness to hold any teaching or assessing role, and it awards no qualification. Scoring is chance-corrected proximity, built in the order the method requires and not in the order that is convenient. First the twelve marks are ipsatised within the person - each mark expressed as a distance from that person's own mean in units of their own spread - and the panel key is ipsatised identically. Then the distance between the two ipsatised vectors is taken. Then that distance is corrected against an EMPIRICAL null: the mean distance produced by a respondent whose mark on each script is drawn from that script's own declared marginal distribution of marks, computed by simulation, and not by a uniform draw across the scale. A uniform null is wrong here and wrong in the direction that flatters: real markers cluster in the middle of a scale, so a respondent answering thirteen to everything lands close to the key by accident, and a scorer that does not correct for it produces the flat baseline where everybody is moderate. Ipsatisation is what makes the agreement figure independent of severity and of spread, and those two are then reported as their own numbers rather than being lost - a marker three marks below the panel throughout has a severity of minus three and an agreement figure that is unaffected by it, which is the whole point. Self-consistency is computed from the two seeded repeat scripts and is reported as a distance in marks, never as a pass or a fail. Every strand carries its standard error and reliability is reported as McDonald's omega estimated from item count and a stated assumed inter-item correlation; a strand whose omega will not carry a number gets a three-way placement and the refusal is printed. No percentile appears anywhere. Constructs and sources: rater severity, central tendency and the many-facet Rasch treatment of judge effects (Linacre 1989; Eckes 2011); halo and fluency effects on marking (Thorndike 1920; Klein and Hart 1968); contrast effects between consecutively marked scripts (Daly and Dickson-Markman 1982; Attali 2011); inter-rater reliability and the limits of percentage agreement (Shrout and Fleiss 1979); the effect of exemplars and standardisation on marker agreement (Bloxham, den-Outer, Hudson and Price 2016); reliability of essay marking in public examinations (Murphy 1979); item difficulty, discrimination and the uninformative ceiling item (Crocker and Algina 1986); assessment literacy as a distinct professional competence (Popham 2009; Stiggins 1991); feedback at task, process and self-regulation level and the harm of self-level feedback (Hattie and Timperley 2007); the grade crowding out the comment (Butler 1988); and ipsative within-person standardisation as a control for elevation and scatter in profile matching (Cronbach and Gleser 1953; Fisher 2004). All items and all student scripts are original works written for this instrument. No examination board's mark schemes, rubrics, examiner materials, scale names or report layouts are reproduced or implied, and no accreditation or affiliation is claimed.