Trust & Safety and Content Moderation JudgmentTrust & Safety and Content Moderation Judgment. Knowing the policy is not the job.
The job is choosing how big an action to take, in minutes, on incomplete evidence, where leaving harm up and taking down speech that should have stayed both cost something. A scored diagnostic of that judgment.
Not a policy-recall quiz — a proportionality test
Certification modules test whether you can recite a rulebook. This puts you inside 19 situations where the rulebook runs out: satire that quotes the slur it attacks, a graphic clip that reporters are also citing, two hundred reports in an hour against one ordinary user, a first offence beside a documented pattern, an appeal where the first call was wrong, a distress signal that needs a support path rather than a takedown, and a policy that plainly does not fit the case in front of you.
The Severity Ramp is a new scoring design. Every option sits on a five-rung ladder of restrictiveness — leave, label, limit, remove, escalate — and carries a separate evidence weight for whether the policy text, the surrounding context or the account history was checked first. Over-enforcement and under-enforcement are measured separately and then combined into one Enforcement Balance, with a mean error printed beside it, so a reviewer whose opposite mistakes cancel out is never reported as calibrated.
Keying follows published evidence rather than one platform's house style: signal-detection theory and its two error types, the precision-recall trade-off, base-rate neglect, over-blocking and chilling-effect research, the least restrictive effective measure, procedural-justice research on why people accept decisions, inter-rater reliability work on why reviewers disagree, and decision-fatigue findings from high-volume queues.
What you walk away with
Five rungs from leave to escalate, showing how often the key called for each rung next to how often you chose it.
Whether you run heavy-handed or permissive, and whether a balance near zero is calibration or two opposite errors cancelling.
How often you landed on the exact rung the situation called for, not merely in the right direction.
How often you checked the policy text, the context or the account history before acting, and recorded the reason.
Posts, accounts, marketplace, messaging and appeals, plus policy application, harm calibration, proportionate action and process integrity.
Inside your report
Illustrative sample — your report is generated from your own responses.
Built for
- Content moderators, trust and safety analysts and community managers at any level
- Platform-policy, marketplace and social-platform operations teams building review capacity
- Technology operations recruiters and outsourcing partners hiring reviewers at volume
Find out whether your enforcement fits the situation
30 scored exercises · about 32 minutes · full bespoke report with your enforcement ladder, Enforcement Balance, Proportionality Hit Rate and Evidence Discipline Index.
₹999 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
Four competencies: applying policy to ambiguity, calibrating harm and severity, choosing the least restrictive effective action, and escalation, appeals and records. It measures what a reviewer would actually do across 19 realistic queue situations, rather than whether they can recall a rulebook they were shown in training.
Every option carries a graded quality score keyed to published research, plus two hidden weights: a rung on a five-step ladder of restrictiveness from leave to escalate, and an evidence weight for whether the policy text, the context or the account history was checked first. Your answers are compared with the keyed rung to produce separate over-enforcement and under-enforcement figures, one Enforcement Balance, a mean error that exposes cancelling mistakes, a Proportionality Hit Rate and an Evidence Discipline Index.
Content moderators, trust and safety analysts, community and platform-policy teams, marketplace and social-platform operations staff, and the technology operations teams and outsourcing partners hiring them. It is platform-neutral: no named platform, no country's law and no company's community guidelines are assumed or tested.
About 32 minutes for 30 scored exercises. You get a full report drawn as an enforcement ladder, with your Enforcement Balance and mean error, Proportionality Hit Rate, Evidence Discipline Index, craft across five product surfaces, four banded competencies and a short list of things to take into your next shift — downloadable as a colour PDF.
Rs 999 in India (including GST) or US$9.99 elsewhere, one-time, for one full sitting and report. Employer-only trust and safety assessment suites are sold as annual subscriptions, and moderator training courses are priced per seat; this is a single purchase, and the person assessed gets the full report as well. Organisations can assess at volume through AssessAll credits at 26 credits per person.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Measures applied judgment in content moderation and trust and safety operations through original situational items keyed to published evidence. The two-error frame comes from signal-detection theory (Green & Swets, 1966; Macmillan & Creelman) and its false-positive and false-negative costs, and from the precision-recall trade-off in classification, which is why over-removal and under-removal are scored separately rather than netted. Report-volume items are keyed against base-rate neglect and representativeness (Kahneman & Tversky, 1973; Tversky & Kahneman, 1974), and against the published research on over-blocking and the chilling effects of removal on lawful expression. Action selection is keyed to the principle of the least restrictive effective measure and to proportionality in enforcement, as set out in published human-rights and platform-accountability commentary and in transparency-reporting and audit practice. Appeal, explanation and consistency items follow procedural-justice research on why people accept decisions they dislike (Thibaut & Walker, 1975; Lind & Tyler, 1988; Tyler, 1990), built on voice, explanation, consistent treatment and a real route of appeal. Reviewer-disagreement items draw on inter-rater reliability and rubric research (Cohen, 1960; Krippendorff; Hallgren, 2012). Rapid-decision items reflect anchoring and framing effects (Tversky & Kahneman, 1974; 1981) and the decision-fatigue literature on sequential decisions in high-volume queues (Danziger, Levav & Avnaim-Pesso, 2011; Baumeister on depletion). Escalation and hand-off keying draws on high-reliability organisation research (Weick & Sutcliffe, 2007) and on published clinical hand-off studies. Wellbeing framing reflects published research on moderators exposed to distressing material (Roberts, 2019; Steiger et al., 2021; Spence et al., 2023). All items are original works: no platform's policy text, community guidelines, internal rulebook or trademarked instrument is reproduced, no country's law is tested, and no affiliation with any platform, regulator or publisher is claimed or implied. This assessment gives skills feedback for hiring and development. It is not legal advice, not a clinical or wellbeing screen, and not a certification of compliance with any law or standard.