AI Draft Supervision: Reviewing What the Machine WroteAI Draft Supervision. Everyone can use the tool. Almost nobody supervises it.
Your name goes on the draft. This is a scored profile of what you catch, what you leave, and what you change that did not need changing - built from drafts with defects planted in known classes, sitting next to passages that are genuinely good.
Not an AI-literacy quiz. A supervision test.
Every catalogue on the market now measures whether you can use an AI tool. None of them measures whether you supervise one, and that is the axis every knowledge job quietly added in the last two years. This assessment puts you in front of thirty-two exercises where the machine has already produced something: a customer reply, a set of minutes, a job advert, a board summary, a note to a client board after an incident. Six of them are Ghost Drafts - a new item format - where the defects were planted deliberately: a specific detail nobody established, a statistic with no source, a cause invented to explain a delay, an absolute assurance given before the review that could support it had finished, an instruction so hedged that nobody can act on it. Sound passages sit right next to them, and one of the sound passages is always the one it would cost most to delete.
The Supervision Yield is a new scoring design. Every option carries a declared cost in one published unit - how much harm it does if the draft goes out that way - and the cost matrix is printed on your report rather than buried in the scoring. Your result is expressed against what a typical reviewer would have let through, not against a coin toss, because every option also carries the share of reviewers expected to choose it. A score of zero means no better than the room. That is a much harder and much more useful zero than the one most instruments print.
Two errors are reported separately and neither is treated as the only one that matters. Deference is a defect left standing. Interference is sound work rewritten, or a caveat deleted because somebody found it unhelpful. Somebody who defers on half the material and over-edits the rest cancels out in a single figure and would be told they were balanced. Here the absolute figure decides the verdict, and the pattern is named Erratic.
What you walk away with
One number against the typical reviewer, drawn with the range a retest would land in, and the assumed reliability stated on the page rather than implied.
What you left as the machine wrote it, and what you rewrote that was already right, with the combined figure given the final say over the verdict.
Every class of planted defect, weakest first, in two columns: what happened, and what a stronger review did instead, with the difference written out in a sentence.
How often you cut or rewrote a passage that was carrying something the reader needed - a caveat, a commitment, an instruction. The quiet cost of over-review.
A single specific sentence you can actually run tomorrow, chosen from the error pattern your answers produced rather than from a template.
Inside your report
Illustrative sample - your report is generated from your own responses.
Built for
- Knowledge workers whose drafts are now produced or part-produced by an assistant
- Managers approving assisted work who are answerable for what it says
- Employers screening for the judgment that AI adoption has made scarce rather than abundant
Find out whether you supervise the machine or relay it
32 scored exercises - about 40 minutes - a full bespoke report with your Supervision Yield, both errors counted apart, the counterfactual panel and one change to make this week.
₹999 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
Four capabilities: catching what is wrong in an assisted draft, leaving sound work alone, owning what you send, and matching the depth of review to what a mistake would cost. It works through thirty-two exercises, six of which give you a draft with defects planted in known classes - an invented specific, a cause nobody established, an absolute assurance made too early, an empty hedge, wording that describes a person rather than a job - alongside passages that are genuinely good and expensive to delete.
By a decision-cost matrix. Every option carries a declared cost in one published unit: how much harm the choice does if the draft goes out that way. A fabricated delivery date reaching a customer costs three, a pointless rewrite of sound prose costs one, and the whole matrix is printed on your report because it is a value judgement rather than a measurement. Your Supervision Yield is the share of that harm your review removed, expressed against what a typical reviewer would have let through rather than against random guessing.
About you. No item names a tool, a vendor or a model, and nothing turns on knowing how a particular product works. What is measured is editorial judgment: whether a well-written passage gets read as an accurate one, whether an added inference is visible, and whether the review is targeted or turns into a rewrite. It does not measure your writing, your prompting, your knowledge of any tool, or your intelligence, and the report says so on its own last page.
About 40 minutes for 32 scored exercises. You get a bespoke report built as a counterfactual panel: your Supervision Yield with its band, a cost ledger against a perfect review and a typical one, both errors counted apart with the combined figure deciding the verdict, every class of planted defect in two columns with the difference annotated, four capability classifications, the base rates the scoring runs on, and one if-then change. Downloadable as a colour PDF that also reads in black and white.
Rs 999 in India including GST, or US$9.99 elsewhere, one time, for one full sitting and report. Commercial assessment suites sell this kind of thing only as an annual recruiter subscription, AI-literacy courses charge per seat and teach rather than measure, and free online quizzes test whether you can define a hallucination. Organisations can use AssessAll credits at 26 credits per person.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Measures applied supervision of machine-written work through original items keyed to published constructs. Construct statement: it measures whether a reviewer catches planted defects in an assisted draft while leaving sound content alone, and whether the depth of review tracks what an error would cost. It does not measure writing quality, prompt-writing skill, knowledge of any particular tool, or general ability. Declared response instruction: behavioural tendency throughout - every stem asks what the respondent would actually do, never what should be done. Sources drawn on across the construct area: the automation-bias literature on omission and commission errors in decision support, and the finding that both increase when the aid is usually right; the complacency and vigilance-decrement work on monitoring a reliable automated partner; research on the calibration of trust in automation, in which appropriate reliance depends on the operator's model of where the system fails rather than on how much they like it; the documented failure classes of generated text, including fabricated citations and quotations, invented specifics, and confident causal statements over correlational material; work on fluency and processing ease, in which a well-written passage is judged more accurate than the same content written plainly; signal-detection theory and its separation of discrimination from response threshold, which is why over-flagging and under-flagging are reported apart here; the editing and revision literature distinguishing substantive defects from preference edits, and its finding that untrained reviewers spend most effort on surface features; decision-cost framing, in which the depth of a check is set by the cost of the error it prevents rather than by the visibility of the artefact; the reversibility principle in decision quality, that a decision cheap to undo warrants less scrutiny than one that binds; disclosure and provenance work on assisted authorship, in which naming a tool without stating what was checked transfers the reader's risk rather than reducing it; confidentiality and data-stewardship guidance on inputs to third-party systems, where absence of a policy is not consent; and situational-judgment-test validity together with the behavioural-tendency response instruction. All items are original works. No trademarked instrument is reproduced, no tool, vendor or model is named, and no affiliation with any of the sources above is claimed. Scores are provisional: the base rates used for chance correction are authored priors, published in the report, to be replaced by observed marginals once live response data exists.