Signing It Off: AI Governance JudgmentSigning It Off: AI Governance Judgment. Plenty of tests ask whether you can use AI. This one asks whether you should have said yes to it.
You are not the user in these thirty-four exercises. You are the person who approves the tool, buys it, switches it on for a team, writes the rule for how it will be used, answers when it goes wrong, and decides whether to switch it off.
The approval decision, scored — and a reference curve that refuses to rank you
AI governance now sits alongside bias, privacy, intellectual property and ethical decision-making on the list of compliance programmes organisations say they most need. What that market sells is policy: an annual e-learning module, a knowledge check, a certificate that says somebody read the rules. Nothing in it scores the judgment of the person who actually signs a deployment off. That is what this measures, and it measures nothing adjacent to it — not technical knowledge of how these systems are built, not familiarity with any product or supplier, not any country's law or any published standard, and not your own skill at using such a tool.
Twenty-two situations put the decision in front of you. A supplier whose demonstration ran on their own cases. A tool that ranks job applications, with nobody able to say what it learned from. A month in which no manager changed a single draft the tool wrote for them. A reviewer with forty cases an hour to sign, and a tool that produces forty an hour. A discount offered in exchange for your customer records. Forty customers sent the wrong price by something still running. Every situation offers the same four moves — switch it on, attach a condition, hand the decision up, hold it back — and each of them is worth full credit somewhere and worth nothing somewhere else. Attaching a control is the strongest move in some of these and an expensive piece of paper in others, which is the whole point. Four further exercises ask which checks you would insist on before a tool goes live, scored so that insisting on everything costs exactly what insisting on too little costs.
The scoring is construct-weighted: every option carries a sparse credit vector naming at most two of the four competencies, and most name exactly one. That sparseness is the difference between four readings and one reading printed four times, so the report publishes the measured sparsity and the six correlations between the four readings from a simulated mixed-ability sitting rather than asserting they are separate. Your result is chance-corrected against declared per-option base rates, because a graded partial-credit percentage floors near half by construction; the raw figure is printed too, so the correction hides nothing. Your position is drawn on a reference density curve, and the curve is labelled a model with its simulated size on the same panel. No percentile appears anywhere in the report, for the composite or for any competency, because the only distribution available is a modelled one and a rank against a model is a rank against an assumption.
What you walk away with
Drawn as a density, not a rank, so you can see how crowded the middle is before you read anything into a few points. The curve's simulated size sits on the same panel, and no percentile is claimed for you anywhere.
One per competency, each with your mark on it. A number appears beside a competency only where the reliability table permits one, and where it does not you get a standing and a plain statement of why.
The measured sparsity of the credit vectors, and the six inter-construct correlations from a simulated sitting, against a printed line at .80. If they were too close together the report would say so.
Two times in three a retest would land between here and here; nineteen times in twenty between here and here. The narrower span is what decides whether your reading is on a boundary.
Always switch it on, always attach a condition, always hand it up, always hold it back, always take the first option. Each pushed through the real scorer on your own exercises, with the figures printed.
Inside your report
Illustrative sample - your report is generated from your own responses.
92 of 116 options count towards exactly one. None counts towards more than two.
Measured on a simulated mixed-ability sitting and printed rather than asserted. If these ran near .80 the report would be telling you the same thing four times.
Built for
- Managers, product owners, team leads and functional heads who approve, procure or switch on an AI tool that other people will then have to live with
- Compliance, risk, audit and data-protection teams building an AI governance programme who need a measure of judgment rather than a record of policy attendance
- Boards, executive committees and L&D teams who want a defensible before-and-after on deployment judgment across the layer of managers who actually sign these decisions off
Find out what you would actually approve, and what you would let through
34 scored exercises - about 40 minutes - a full bespoke report with your position on a modelled reference curve, four competency curves, both confidence spans, the sparsity and correlation evidence, and the reliability table that decides what may carry a number.
₹1,199 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
How a person decides whether an AI tool should be switched on for other people to use. Four competencies are reported: what you check before you switch it on; whether the oversight you put in place is actually doing something; how you treat personal data, other people's work and who gets the credit; and what you do when it gets something wrong. It does not measure technical knowledge of how these systems are built, familiarity with any product or supplier, knowledge of any country's law or any published standard, your own skill at using such a tool, or your general attitude towards the technology.
A policy course tests recall of a rule. This tests the decision. Every situation offers four courses of action a capable manager could defend, graded by degree rather than one right answer among fillers, and framed as what you are most likely to do rather than what a person should do, because the second is far easier to fake. It is also deliberately jurisdiction-neutral: no country's statute, no named regulation, no named standard, no named supplier and no named product appears anywhere, so the same form is fair to read in any market.
No, and the instrument is built so that habit cannot pass. Each of the four moves - switch it on, attach a condition, hand the decision up, hold it back - is the full-credit move somewhere and worth nothing somewhere else, and roughly a quarter of the situations key going ahead. Attaching a condition is the strongest move in some situations and an expensive piece of paper in others. Every one of those habits is pushed through the real scorer on your own exercises and printed on your report, so you can see for yourself that none of them pays.
Rs 1,199 or $11.99 for one sitting and the full report, inclusive of GST for buyers in India. AI governance training is normally sold as an institutional compliance subscription or a per-seat annual e-learning licence that examines recall of policy. This sits above the standard applied-skill price point because the instrument runs 34 scored exercises across four competencies, publishes its reference distribution, its sparsity and correlation evidence, its reliability assumptions and its strategy probe on the page rather than in a technical annexe nobody reads. Team and organisation licensing is available on credits at 40 per sitting.
You will not get a percentile, and that is deliberate. There is no norm group yet, so the curve on your report is modelled: simulated respondents drawn from the declared per-option base rates and pushed through the same scorer that produced your own figure. It is labelled a model with its simulated size on the same panel. A curve like that honestly answers the question a percentile pretends to - how crowded is the middle, and is this difference worth acting on - without claiming a rank the data cannot support. When enough real sittings exist, the observed distribution replaces it, and the shape of the curve is the first thing that will change.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Measures the approval decision behind an artificial-intelligence deployment - the judgment of the person who signs it off, buys it, switches it on for a team, writes the rule for how it will be used, and answers when it goes wrong - through original situational, selection and self-report items keyed to published construct areas. Construct statement: It measures how a person decides whether an artificial-intelligence tool should be switched on for other people to use: what they check before approving it, whether the human oversight they put in place can actually change an outcome, how they treat personal data, other people's work and credit, and what they do when the tool gets something wrong. It does not measure technical knowledge of how these systems are built, familiarity with any particular product or supplier, knowledge of any country's law or any published standard, a person's own skill at using such a tool, or their general attitude towards the technology. Declared response instruction: behavioural tendency throughout - every situation asks what the respondent is most likely to do rather than what a person should do, because instructed-knowledge framing is the more fakeable of the two, and situational judgment is a method rather than a construct, so the instruction is declared once and never mixed within the instrument. Item format: twenty-two graded situational judgment exercises, each offering four courses of action a capable manager could defend, carrying a non-negative effectiveness gradient with at least three distinct values, a per-option authored base rate and a full per-option rationale; four selection exercises asking which checks the respondent would insist on before a tool goes live, graded by overlap with the keyed set so that insisting on everything and insisting on too little cost the same; and eight balanced-keyed self-check statements, half of them worded so that agreeing is not the flattering answer. Every option additionally carries a sparse construct-credit vector naming at most two of the four competencies, and the report publishes the measured sparsity and the correlations between the four readings, because a credit vector that loads every option on every competency produces four numbers that say one thing. Scoring design: construct-weighted graded situational judgment. Each competency earns the sum of the credit-weighted effectiveness of the options taken, divided by the most that competency could have earned on the same exercises, and both the composite and each competency are then chance-corrected against the declared per-option base rates, because a graded partial-credit percentage has a floor near half by construction and would otherwise place almost every respondent in the middle. Only answered exercises enter either side of any ratio, so an unfinished sitting reduces coverage rather than scoring zero. The composite carries its standard error at both the sixty-eight and the ninety-five per cent span, printed in words as well as drawn; the reliability behind those spans is an assumption stated in advance rather than a measurement, and the report prints the table that decides, competency by competency, whether a number may be shown at all. The reference distribution on the report is modelled from the declared base rates and is labelled as a model with its simulated size, and no percentile is claimed anywhere. Construct areas and source families drawn on: risk tiering of automated decisions by their consequence for the person affected; intended-purpose and out-of-scope declaration for a deployed capability; pre-deployment evaluation on the deploying organisation's own work rather than on a supplier's benchmark; named accountability for automated output; the conditions that make human oversight meaningful, namely time, reasons and the standing to overturn; automation bias and complacency in human-machine decision making; selective adherence, in which a reviewer accepts a machine's call more readily than a colleague's; notice and transparency towards the people a system is used on; purpose limitation and data minimisation when records held for one reason are put to another; attribution and credit for work produced with assistance, and third-party rights in generated output; incident response, logging and the ability to switch a capability off; procurement due diligence and the evaluation of supplier claims; unapproved tools spreading inside a team; just-culture incident review, which separates a system failure from an individual's conduct; and redress and notification for people affected by an automated error. All items are original works. No trademarked instrument, branded methodology, named product, named supplier, named standard or country-specific law is reproduced or relied on, and no affiliation with any source is claimed.