Data Annotation and Labelling Quality Assessment for Artificial Intelligence Data Operations TeamsTwenty labelling trials in ten matched pairs. Ten the guideline covers, ten it does not. What does the difference cost you?
Inside each pair the printed guideline is the same, the four options are the same words in the same order, and one thing differs: whether a rule names the case in front of you. Your score is the difference between the two sets, taken on the raw scale, with what a respondent answering at the declared priors loses across the same two sets printed beside it. Nineteen more exercises on routing an uncovered case, reading agreement between annotators and keeping a batch auditable, plus two written escalation notes reported entirely apart. No annotation platform, dataset, model, vendor or client is named anywhere.
A data labelling test that measures what happens when the written guideline stops covering the case
The Data Annotation and Labelling Quality Assessment is a forty-minute check of what happens to a person's labelling when the written guideline stops covering the case. Twenty labelling trials in ten matched pairs, differing in one thing only, plus nineteen exercises on routing, agreement and batch hygiene. The score is the difference between the two sets.
Every other labelling test measures accuracy against a key, which measures how much the guideline was carrying as much as it measures the person. This one holds the guideline still. Inside a pair you get the same purpose line, the same three named-case rules, the same default route and the same four options in the same order. On one member a rule reaches the case. On the other no rule does, and the answer comes from the guideline's stated purpose, or the case goes to the guideline owner. Ten pairs across ten different labelling jobs: support topics, bounding boxes, span marking, transcription, preference ranking, review intent, entity typing, event timing, field extraction and search relevance.
The difference is taken on the RAW scale, and that is the design rather than a shortcut. The two sets have different nulls by construction, and that difference is itself the ordinary cost, so correcting each set against its own prior and then subtracting would make the figure move with how accurate a person is rather than with how much the silence costs them. The correction goes on the comparison: a respondent answering at the declared priors loses 18.6 points across the same two sets, and that number is printed beside yours and never divided into it. Zero means no cost at all, so somebody accurate everywhere and somebody inaccurate everywhere both land near it, and the report says so in words so the figure is read as a cost and not as a level.
The silence is not one situation, and the report never pretends it is. On six of the ten uncovered trials the guideline's stated purpose still reaches the case, and the answer is a label with the reading noted. On four nothing in the document reaches it, and the answer is the owner. The two errors are counted apart and never netted: reaching for a label the guideline could not give, and raising a question the purpose already settled. They have different costs and different fixes, and the escalation option sits in every response set and is keyed in both sets, so answering every silent case with the escalation route scores five of twenty, which is what guessing scores on a four-option item.
The report is a strength and shadow pair. Every reading on the page is printed in two ruled columns: what it gets you, and what it costs when it is overused. Eight pairs, four on the contrast and four on the supporting areas, and not one is left unpaired. That is the strongest anti-flattery device a report has, because a universally flattering sentence is accepted by everybody who reads it and a named cost cannot be. The ten matched pairs are also drawn as a ledger, one row each, so the mechanism of the figure can be counted rather than trusted.
Two written exercises ask you to draft the escalation note itself: the question you would send mid-batch, and the note asking for a clause after the same unnamed case has arrived ten times. They are graded against a printed rubric, reported entirely apart from every keyed figure, and because rubric grading is a service you do not control, an ungraded rewrite leaves BOTH sides of the ratio and is printed as DEFERRED rather than scored zero.
Every guideline extract is written for this instrument. No annotation platform, dataset, model, vendor, client, standards body or certification scheme is named anywhere, no exercise needs any country's content-moderation or data-protection law, and no trial uses a real person, brand or publication. Agreement statistics are named and used as the public methods they are. The report says in plain words what it did not measure: labelling speed, tool skill, output volume, eyesight or care. It is not a certification.
What you walk away with
Accuracy on the ten trials a rule named minus accuracy on the ten matched trials no rule named, on the raw scale, with the 68 and 95 per cent spans drawn on the axis and the ordinary cost marked beside your own figure as a hollow diamond. Where the 95 per cent span includes zero the page says the sitting shows a direction and not a measured size.
Ten rows, one per pair, with the named cell and the silent cell side by side. A pair held, a pair lost when the rule ran out, a pair gained because the written rule was read past, or a pair missed on both. The mechanism of the headline, countable rather than asserted.
Cases that needed the guideline owner and were labelled anyway, and cases the stated purpose already reached that were sent on instead. Each with its own count, its own placement, what it costs and how it is fixed. Never netted against each other, because a single figure would tell you to do neither.
Ten covered trials and fourteen exercises on the route a guideline gives an unnamed case: label it from the purpose with the reading noted, follow the precedence line, raise one question for the batch, or ask for a clause when the same case keeps arriving.
Eight exercises on per-cent agreement against a chance-corrected coefficient, what a gold set answers that agreement cannot, which disagreements a rewrite can remove, and why perfect agreement between two people is not evidence that a guideline is unambiguous.
Six exercises on rows labelled under a rule that has since changed, open questions shipped inside a batch, a sample drawn so it overstates the batch, and the uses a gold set supports against the uses that spend it. Then the two notes in your own words: the guideline question you would send mid-batch, and the note asking for a clause after ten similar cases. Both graded against a printed rubric, reported apart from every keyed figure, and printed as DEFERRED rather than zero when the grader does not come back.
Inside your report
Illustrative sample — your report is generated from your own responses.
How to read it: the difference is taken on the RAW scale, because the two sets have different nulls by construction and correcting each side before subtracting would make the figure move with how accurate you are rather than with how much the silence costs you. The correction goes on the comparison: the hollow diamond is what a respondent answering at the declared priors loses across the same two sets. Zero is no cost at all, so somebody accurate everywhere and somebody inaccurate everywhere both land near it.
| The difference is built from | Raw | At the priors | Corrected |
|---|---|---|---|
| ● The ten trials a rule named | 80% | 54.8% | 56 |
| ○ The ten trials no rule named | 49% | 36.2% | 20 |
| The difference, raw scale | 31 | 18.6 | +12.4 |
The two corrected figures are printed as supporting readings and are never subtracted from each other.
You apply what is written, including the rules that cut against what the case looks like.
Fluency with written rules makes the document feel complete, and the reader who applies rules best is the one most likely to find a rule for a case that has none.
This sitting. 8 of 10 answered, 80 per cent, against 54.8 per cent at the declared priors.
You work from the guideline's stated purpose when its rules run out, which keeps a batch consistent between versions.
Reading a purpose confidently is how an unwritten rule gets made. Your reading may be right and it is still yours, so it belongs in a note.
This sitting. 5 of 10 answered, 49 per cent, against 36.2 per cent at the declared priors.
You tell a case the purpose still reaches from a case nothing in the document reaches, which is the difference between a label and a question.
Getting that call right most of the time makes it feel easy, and the case that does need a clause looks most like the ones that did not.
This sitting. 3 of 4 cases that needed the owner were labelled anyway; 1 of 6 the purpose reached was sent on instead.
| Reading | Placement | Carries a shadow |
|---|---|---|
| Accuracy where a rule named the case | ▲ Above the ordinary respondent | yes |
| Accuracy where no rule named the case | ▬ Near the ordinary respondent | yes |
| Knowing which kind of silence you are looking at | ▼ Below the ordinary respondent | yes |
| Eight pairs on the full page: four on the contrast and four on the supporting areas. None is left unpaired. | ||
A flattering sentence is accepted by everybody who reads it. A named cost cannot be read as generic praise, and no sentence on the page would be accepted as accurate by somebody who scored the opposite.
Built for
- Data annotators and reviewers who want to know whether their labelling holds up on the cases the guideline never named
- Annotation QA leads choosing between more training and a guideline rewrite, who need the two told apart on evidence
- Data-operations and delivery managers staffing a labelling programme, who need a provider-neutral check of judgement rather than of tool vocabulary
- BPO and data-services training teams building an induction, who want a before-and-after figure with a printed threshold for what counts as real change
Find out what the guideline's silence costs you
39 exercises across five formats · about 40 minutes · one difference between twenty matched trials, two error directions counted apart, four areas placed rather than scored, and every reading printed with its shadow.
₹599 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
One thing: what happens to a person's labelling when the written guideline stops covering the case. Twenty labelling trials are presented in ten matched pairs. Inside a pair the printed guideline is the same, the four response options are the same words in the same order, the keyed decision is one, and the share of trials where the guideline's answer is also the intuitive answer is the same on both sides, four of ten. One thing differs: whether one of the three named-case rules reaches the case. Your score is the difference between your accuracy on the two sets. It does not measure labelling speed, tool skill, output volume, eyesight or care, and it is not a certification.
Because correcting each set against its own prior and subtracting the two corrected figures looks careful and is wrong. The two sets have different nulls by construction, and that difference is itself the ordinary cost, so the same raw accuracy maps to two different corrected scores and the corrected difference then moves with how accurate the respondent is rather than with how much the silence costs them. The correction belongs on the comparison instead. A respondent answering at the declared priors is right on about 54.8 per cent of the covered trials and 36.2 per cent of the uncovered ones, so the ordinary cost is about 18.6 points, and that number is printed beside your difference rather than divided into it.
No, and the form is built to make that habit worthless. The escalation option sits in every response set on both sides of every pair, and it is keyed in both sets: once across the ten covered trials, where a rule explicitly routes a named case to the guideline owner, and four times across the ten uncovered ones. Escalating everything therefore scores five of twenty, which is exactly what guessing scores on a four-option item. On six of the ten uncovered trials the guideline's stated purpose still reaches the case and the answer is a label, so the real discrimination is telling the two kinds of silence apart. The report counts those two errors separately and never nets them.
No. Every guideline extract, response set and case is written for this instrument. No annotation platform, labelling tool, dataset, model, vendor, client, standards body or certification scheme is named anywhere in the exercises, the report or this page, and the instrument is not affiliated with, endorsed by or derived from any of them. No exercise requires any country's content-moderation or data-protection law, and no trial uses a real person, brand or publication. Agreement statistics — per-cent agreement, chance-corrected agreement coefficients and overlap measures — are public methods and are named and used as generics. It works whatever your team labels and whatever it labels in.
About forty minutes for thirty-nine exercises across five formats. The sitting is free to take; the report is the product, priced at ₹599 in India, inclusive of GST, or US$5.99 elsewhere, one time. An unanswered exercise leaves the denominator rather than scoring zero, and an entirely empty sitting scores exactly zero and claims nothing. A sitting with fewer than twenty-four of the thirty-nine answered, or with fewer than six answered trials on either side of the contrast, is not reported at all: the refusal is printed where the figure would have been, because a difference of two accuracies read on a handful of trials would reward the short sitting. The two written exercises are printed as deferred rather than scored if the grader does not come back.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Thirty-nine original exercises across five formats. Twenty are labelling trials in TEN MATCHED PAIRS: inside a pair the printed guideline is the same, the four response options are the same words in the same order, the time envelope is the same, and the share of trials on which the guideline's answer is also the answer common sense reaches is the same, four of ten in each set. The pairs differ in one thing only. In Set A one of the three named-case rules covers the case; in Set B no rule names it, and the reader has to work from the guideline's stated purpose or raise a guideline question. Both sets carry the escalation option in every response set and both sets key it sometimes, once in Set A and four times in Set B, so answering every silent case with the escalation route scores five of twenty, which is what guessing scores on a four-option item. Seven select-all exercises, six true-or-false claims, four match-the-following exercises and two written exercises complete the form. DECLARED RESPONSE INSTRUCTION, one for the whole instrument and it is a KNOWLEDGE instruction: which label the supplied guideline supports, and what to do when it supports none. Never what the respondent would prefer, would feel, or would do under pressure. CONSTRUCT STATEMENT: this instrument measures whether a person applying a written labelling guideline keeps applying it when the guideline stops covering the case, measured as the difference between twenty matched trials in one sitting, together with four supporting areas: applying a stated rule, routing an uncovered case, reading agreement between annotators, and keeping a labelled batch auditable. It does not measure labelling speed, tool skill, output volume, eyesight, language ability or diligence, it says nothing about any annotator's honesty or employability, it is not a certification and confers none, and it is not affiliated with, endorsed by or derived from any annotation platform, dataset, model, vendor, client, standards body or certification scheme. SCORING DESIGN: a matched-trial within-person contrast. Accuracy is computed on each set over the trials the respondent answered, and the headline is the DIFFERENCE between the two accuracies TAKEN ON THE RAW SCALE. The two sets have different nulls by construction, that difference is itself the ordinary cost, and correcting each set against its own declared prior before subtracting would make the figure move with how accurate the respondent is rather than with how much the silence costs them. The chance correction therefore goes on the COMPARISON: the mean keyed prior across the ten covered trials is 54.8 per cent and across the ten uncovered trials 36.2 per cent, so a respondent answering at the declared priors loses about 18.6 points across the same two sets, and that figure is printed beside the respondent's own difference rather than divided into it. The two per-set figures are also corrected against their own priors and printed as supporting readings, and they are never subtracted from each other. Zero on the headline means the same accuracy on both sets, which is no cost at all; somebody accurate everywhere and somebody inaccurate everywhere both land near zero, and the page says so in words so the figure is read as a cost and not as a level. On the silent set the two error directions are counted apart and never netted: reaching for a label the guideline could not give, and raising a question where the stated purpose already settled the case. Priors are authored estimates carried per option as option_base_rate, per left-hand item as pair_base_rate, and duplicated as option_select_rate on the select-alls; a binary claim is never given an even split, because an even prior on a two-option item is the coin the correction exists to remove. Observed shares replace every prior once enough people have sat this. REFUSAL RULES: an unanswered exercise leaves the denominator rather than scoring zero. An entirely empty sitting scores exactly zero and claims nothing. A sitting with fewer than twenty-four of the thirty-nine answered, or with fewer than six answered trials in either set, is NOT REPORTABLE: the refusal is printed where the figure would have been rather than beside it. Every one of the four supporting areas carries a three-way classification and never a number and never a percentile, because none of them reaches the eight exercises a figure needs. Omega is reported, not alpha, on the unrounded value, and a part below .70 carries no number. Every reported figure carries its standard error and both its 68 and 95 per cent bands. The two written exercises are graded by the platform's own rubric pipeline, are reported entirely apart from every keyed figure, and because grading is a service the reader does not control, an ungraded rewrite LEAVES BOTH SIDES OF THE RATIO and is printed as DEFERRED, never scored zero. NAMED SOURCES for the public methods this instrument uses and reports on: Cohen's kappa for chance-corrected agreement between two raters; Scott's pi; Fleiss's kappa for more than two raters; Krippendorff's alpha for any number of raters and any measurement level; Gwet's AC1 as the paradox-resistant alternative; Feinstein and Cicchetti on the two kappa paradoxes that arise on an unbalanced label set; Landis and Koch for the widely quoted agreement benchmarks and the caution that they are arbitrary; Artstein and Poesio on inter-coder agreement for annotation work; the Jaccard index and intersection over union as overlap measures for spans and drawn regions; Gebru and colleagues on datasheets for datasets and Bender and Friedman on data statements, for the documentation practice the batch-hygiene exercises draw on; Haladyna, Downing and Rodriguez for the item-writing rules the options follow; Crocker and Algina for difficulty and discrimination targets; Ebel and Frisbie for the discrimination thresholds; McDonald for omega; and Baumgartner and Steenkamp on response styles. Every one of these is a published method or a published guideline, named factually, with no implication of endorsement by any author or publisher. ORIGINALITY: every exercise, guideline extract, response set, rationale and report sentence in this instrument is an original work written for it. No commercial instrument's items, scale names, report section names or product name are used, reproduced or implied, no annotation platform, dataset, model, vendor or client is named anywhere, and no competitor is named anywhere.