Applied Judgment Assessment
Applied skill assessment · every employee, every role, every level, including people who have used an assistant twice · browse the full catalogue

Artificial Intelligence Literacy Assessment for Employees in Every RoleTwenty-two machine outputs, half of them fine. Find out which of the two mistakes you make, and what it costs you.

Not a quiz about how the technology works. Each of the twenty-two printed outputs sits beside the source or the request it came from, and asks one thing: would you send this on as it stands, or stop it? Half are fine. Half carry a named failure written to read fluently: an invented citation, a figure that contradicts the source above it, a summary that dropped the one caveat, a percentage over the wrong group, a reply whose tone is wrong for a regulator, an answer that is correct today and stale, a translation that reversed a negation. Reported as a two-error panel: what you sent that should have been stopped, what you stopped that was fine, counted apart and never averaged.

25 minutes39 scored exercisesEvidence-keyed scoringGlobal · INR & USD

An AI literacy assessment that measures the thing that costs money: the send-or-stop call

The Artificial Intelligence Literacy Assessment for Employees in Every Role is a twenty-five-minute check of whether a person can tell a machine output they should send from one they should stop, on twenty-two printed outputs, with the two mistakes counted separately, plus what they know about what may be shared with such a tool and the checks before sending.

Corporate training on this subject teaches what the technology is and what it can do. That is worth knowing and it is not what costs an organisation money. What costs money is one person, on one afternoon, forwarding a fluent paragraph with an invented source under their own name, or holding a correct reply for two days because it read unfamiliar. Neither shows up on a knowledge quiz. Both show up here, because every spine exercise is a real artefact and one call.

Nothing on the form requires knowing a product, a model, a menu or a price. The tool is described as the assistant your employer provides, and every exercise is answerable by somebody who has used such a tool twice and unanswerable by somebody who has not read the output in front of them. The difficulty lives in the judgement: the wrong outputs are written to read fluent and the fine ones to read plain, because fluency is the one thing the tool always supplies.

The scoring is asymmetric-lure detection with a declared take-up prior. Every stimulus carries the share of ordinary employees expected to call it correctly, a subtle one is worth more than an obvious one, and each of the two figures is corrected against the ordinary employee on the same stimuli. Zero means indistinguishable from most people; a hundred means every call right. Somebody who stops everything scores a hundred on lures caught and far below zero on fine outputs kept, and the report prints both, because half the stimuli are genuine and a rule of thumb in either direction is priced rather than rewarded.

There is no pass mark, and the report says why where a pass mark would have gone. The two errors carry different costs and are never netted into one number, so there is no honest single figure to hang a certificate on; a person who refuses every output and a person who forwards every output are both failing, in opposite directions, and one figure would call them the same. The declared cost ratio, three to one, is printed above the two columns because it is a value judgement and not a measurement, and the expensive column is drawn heavier so the page cannot be read as two equal problems.

Beneath the two columns, one line says which of the two you do more and what that habit costs at the printed ratio, and each column carries a worked example taken from your own answers and its own fix as an if-then rule. Three shorter strands are scored apart and never averaged in: what may and may not be pasted into a tool an outside company runs, which check catches which error and in what order, and how much of a passage a reader can actually verify. A strand too short to carry a figure carries a placement word and the refusal is printed where the figure would have gone.

The five areas follow, in our own words, the five competence areas of DigComp 2.2, the European Digital Competence Framework for Citizens, published by the European Commission's Joint Research Centre and free to download and reuse; the framework belongs to the European Commission, and this assessment is not affiliated with it, endorsed by it, or a conformity assessment against it. Organisations operating in the European Union carry a literacy obligation for staff who use these systems on their behalf; nothing here claims to discharge that obligation for anyone, and no vendor, product, model or competitor is named anywhere.

Two error directions counted apart, three knowledge strands scored separately, five areas of the work read as contributors:
Claims and sourcesFigures and denominatorsMeaning and caveatsAudience and toneSharing and the checks before sending

What you walk away with

Sent something that should have been stopped

How many of the fluent lures went out on your say-so, out of the lures you answered, with a worked example from your own answers: what you chose, what the output warranted, and what sending it would have cost in a working week. Drawn as the heavier column, because at the declared ratio it is the expensive error.

Stopped something that was fine

How many correct outputs you held, out of the fine ones you answered, with its own example, its own cost sentence and its own fix. Cheaper is not free: every fine output held is paid for by the person waiting for it, and a habit of holding trains colleagues to route around you.

Which of the two you do more, at the printed ratio

One line beneath the columns: the rate in each direction, the cost of each at three to one, and which habit is costing you more. The ratio is printed above the columns as a declared judgement, so a reader who disagrees with it can see exactly what would change.

The weighted pair, with both bands

Lures caught and fine outputs kept, each on its own axis from minus one hundred to one hundred with the ordinary employee at zero, the 68 and 95 per cent bands drawn and written out, and the two fixed habits, stop everything and send everything, ticked on the same axis so the shape of your pair can be read against them.

What may be pasted, which check, in what order

Three strands scored apart and never averaged into the panel: six select-every-line exercises on what an outside-run tool may be given, four matching and four ordering exercises on which check catches which error and the sequence before sending, and three counts of what a reader can actually verify in a passage.

One change and a named re-measurement

A single if-then rule drawn from the column that costs you more, naming the moment and the replacement behaviour, and a three-month re-sit with the threshold printed that a change must exceed before it counts as real rather than as measurement error.

Inside your report

Illustrative sample — your report is generated from your own responses.

The two-error panel: the cost ratio printed above, two columns counted apart, the expensive one heavier
The cost ratio, printed because it is a judgement and not a measurement
You decideOutput was fineOutput was wrong
Send it03 sent, should have stopped
Stop it1 stopped, was fine0

In words: three to one.A sent lure costs more than a held fine output, but not infinitely more: a ratio that treats false confidence as free trains the refuser.

The expensive error
Sent something that should have been stopped
3
of 11 answered27 per cent

Priced at 3 against 1: three to one. Drawn heavier because it is.

A worked example from your own answers
A summary that lost the exception. You chose: Send it on just as it stands. The output warranted: Stop it and redo it yourself. Sent, every contractor who read the noticeboard was told something the policy does not say, and the correction went out under your name.
58 in 100 ordinary employees also call this one wrong.
The fix for this column

If an output reads fluent and complete, as "a summary that lost the exception" did, then I will put the summary beside the source and look for the exception before I read anything else.

The cheaper error
Stopped something that was fine
2
of 11 answered18 per cent

Priced at 1 against 3. Cheaper is not free.

A worked example from your own answers
An output that admits it found nothing. You chose: Stop it and redo it yourself. The output warranted: Send it on just as it stands. Held, the manager waited a day to learn that the source was missing, which is the one thing they needed to know that morning.
54 in 100 ordinary employees also call this one wrong.
The fix for this column

If an output reads plain or unfamiliar, as "an output that admits it found nothing" did, then I will forward an output that names its own gap, because the gap is the useful part.

Which of the two you do more, and what it costs at the printed ratio

You do the expensive one more: 3 of 11 lures sent, against 2 of 11 fine outputs stopped.

At three to one, the 3 sent lures cost 9 units and the 2 held fine outputs cost 2 units, where one unit is one fine output held for a second look. The sending habit is the one costing you more.

filled header and thick rule = the expensive column · thin rule = the cheaper column. The weight is the ratio, not decoration.

The weighted pair: two figures, each with both bands, the two fixed habits ticked on the same axis, never one number
Caught: lures you stopped
41
-1000 = ordinary employee100 = every call rightstop allsend all

Weighted by how subtle each one was, 74 per cent against the ordinary employee’s 56. About two chances in three that a re-sitting lands between 29 and 53; nineteen in twenty between 17 and 65.

Kept: fine outputs you sent
58
-1000 = ordinary employee100 = every call rightstop allsend all

Weighted by how subtle each one was, 82 per cent against the ordinary employee’s 57. About two chances in three that a re-sitting lands between 46 and 70; nineteen in twenty between 34 and 82.

The contributors under the figures, by area
  • A precise figure credited to a report you never gave it · Claims and sources · right call · weight 0.62
  • A percentage over the wrong group · Figures and denominators · sent, should have stopped · weight 0.44
  • A summary that kept the exception · Meaning and caveats · right call · weight 0.34
  • A translation in different words that means the same · Audience and tone · stopped, was fine · weight 0.54
Tabular fallback, and the two fixed habits priced on the same stimuli.
WhoCaughtKept
You41 (17 to 65)58 (34 to 82)
Stop everything, or check everything100-100
Send everything-93100
The ordinary employee00

Built for

  • Employees in any role who use an assistant their employer provides, or are about to, and want to know which way their own judgement leans before it costs them
  • Learning and compliance teams who need one instrument for the whole workforce that measures the send-or-stop call rather than knowledge of the technology
  • Managers who want to know whether a team's risk is the fluent output that goes out or the correct output that never leaves, because the two need opposite fixes
  • Organisations with staff in the European Union who need a provider-neutral, framework-anchored read on literacy, with its reliability and its refusals printed on the page

Find out which of the two mistakes you make

39 exercises across six formats · about 25 minutes · ₹399 in India inclusive of GST, or US$3.99 elsewhere · a two-error panel with the cost ratio printed above it, both figures with their bands, three strands scored apart, and every refusal printed where the figure would have been.

₹399 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and your receipt is sent to that address.

Secure Razorpay payment · ₹399 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

Do I need to know how the technology works, or which product my company uses?

No. Nothing on the form requires knowing a product, a model, a menu or a price, and no vendor or product is named anywhere. The tool is described as the assistant your employer provides. Every exercise prints a short machine output beside the source or the request it came from, and asks whether you would send it on as it stands or stop it. Somebody who has used such a tool twice can answer every exercise; somebody who has not read the output in front of them cannot.

Why is this an assessment and not a certification?

Because the construct has two error directions whose costs are not symmetric and which are never netted into one number: sending an output that should have been stopped, and stopping an output that was fine. A person who refuses every output and a person who forwards every output are both failing, in opposite directions, and one pass mark would score them the same. The report prints this refusal where a pass mark would have gone, and prints the two figures each with its own band instead.

How is it scored, and what does zero mean?

Every one of the twenty-two printed outputs is either a lure or genuine and carries a declared prior: the share of ordinary employees expected to call it correctly. A subtle output is weighted more than an obvious one. Lures caught and fine outputs kept are each computed as a weighted share and corrected against the ordinary employee on the same outputs with the same weights, so zero means indistinguishable from most people, a hundred means every call right, and below zero means more of that error than most people make. The priors are declared assumptions printed on the report, and observed data will replace them.

What happens if I just stop everything to be safe?

You score a hundred on lures caught and far below zero on fine outputs kept, and the report prints both. Half the outputs are fine by design, so stopping everything, and checking everything, are priced on the page as habits with a cost rather than rewarded as caution. The declared cost ratio, three to one, says a sent lure costs more than a held fine output, but not that a held fine output costs nothing: every one of them is paid for by the person waiting for it.

How long does it take, what does it cost, and can I sit it again?

Thirty-nine exercises across six formats, about twenty-five minutes. One payment covers the sitting and the report together: ₹399 in India inclusive of GST, or US$3.99 elsewhere, one time, with nothing further to unlock. Re-sitting in three months is free, and the report prints the threshold a change on each figure must exceed before it counts as real rather than as measurement error. An unanswered exercise leaves the denominator rather than scoring zero, and a sitting with fewer than twelve of the twenty-two printed outputs answered is not reported at all.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue →

Methodology: Thirty-nine original exercises across six formats: sixteen single-choice calls and six true-or-false calls on a printed machine output, which together form the spine; six select-every-line exercises on what may be pasted into a tool an outside company runs; four matching exercises with more entries on the right than on the left; four ordering exercises on the checks before a machine-written message leaves; and three numeric estimates on a slider. CONSTRUCT STATEMENT: this measures whether a person can tell a machine output they should send from one they should stop, on twenty-two printed outputs in one sitting, and separately whether they know what may be shared with such a tool, which check catches which error, and the order of checks. It does not measure the person's job, their honesty, their intelligence, their skill with any product, or any tool's accuracy; it is not a qualification or a certification and carries no pass mark. DECLARED RESPONSE INSTRUCTION: the spine carries a behavioural-tendency instruction with a keyed call (would you send this on as it stands?), the same four routes on every stimulus; the other strands carry a knowledge instruction. SCORING DESIGN: asymmetric-lure detection with a declared take-up prior. Every spine stimulus is a lure or a genuine output and carries a declared per-option answer prior, the share of ordinary employees expected to choose each route. The weight of a stimulus is one minus the share expected to call it correctly, so a subtle stimulus is worth more. Caught is the weighted share of lures stopped by any route; kept is the weighted share of genuine outputs sent. Each is corrected against the ordinary employee on the same stimuli with the same weights, so zero means indistinguishable from somebody answering the way people typically answer and one hundred means every call right. The two are never netted: a person who stops everything scores one hundred on caught and far below zero on kept, and the report prints both. Eleven of the twenty-two stimuli are genuine, so refusing everything cannot win. THE NULL is each stimulus's declared marginal, never a uniform draw across the four routes, and no two-option prior is set at an even split. The standard deviation of the spine priors is computed and printed, with a floor of .06, because a prior set with no dispersion collapses the design into plain accuracy. THE COST RATIO between the two errors is three to one and is a value judgement rather than a measurement: it is printed above the two columns on the report and defended there. It is not larger because a report that treats false confidence as free trains the refuser. RELIABILITY ASSUMPTION: omega is the Spearman-Brown form from the answered item count and an assumed inter-item correlation of .24, stated on the report as an assumption; a strand carries a figure only with eight or more answered items and omega at or above .70 unrounded, and otherwise carries a three-way placement with the refusal printed where the figure would have gone. Every figure carries its standard error and its 68 and 95 per cent bands. An unanswered exercise leaves the denominator and the expected term together; an empty sitting scores exactly zero on both strands and is marked not reportable. No percentile appears anywhere, because no norm group exists yet. FRAMEWORK ANCHOR: the five areas follow, in our own words, the five competence areas of DigComp 2.2, the European Digital Competence Framework for Citizens, published by the European Commission's Joint Research Centre and free to download and reuse; its 2022 edition's examples on interacting with artificial intelligence systems informed which failures the stimuli carry. The framework belongs to the European Commission; this instrument is not affiliated with, endorsed by, or a conformity assessment against it, and none of its wording is reproduced as a scale. The regulatory context is Article 4 of Regulation (EU) 2024/1689, which asks providers and deployers to ensure a sufficient level of literacy among staff dealing with artificial intelligence systems; nothing here claims to satisfy that obligation for any organisation. SOURCES: Vuorikari, Kluzer and Punie, DigComp 2.2: The Digital Competence Framework for Citizens (Joint Research Centre, 2022); Regulation (EU) 2024/1689, Article 4; Green and Swets, Signal Detection Theory and Psychophysics (1966); Macmillan and Creelman, Detection Theory: A User's Guide (2005), for the two-error structure and why a single accuracy figure destroys it; Swets, Dawes and Monahan, Psychological science can improve diagnostic decisions (2000), for the decision-cost framing; Parasuraman and Riley, Humans and automation: use, misuse, disuse, abuse (1997); Lee and See, Trust in automation: designing for appropriate reliance (2004); Bucinca, Malaya and Gajos, To trust or to think: cognitive forcing functions can reduce overreliance on AI (2021); Bansal and colleagues, Does the whole exceed its parts? The effect of AI explanations on complementary team performance (2021); Ji and colleagues, Survey of hallucination in natural language generation (ACM Computing Surveys, 2023), for the failure modes the lures carry; Haladyna, Downing and Rodriguez, A review of multiple-choice item-writing guidelines (2002), for cue control; McDonald, Test Theory: A Unified Treatment (1999), for omega; Jacobson and Truax, Clinical significance: a statistical approach to defining meaningful change (1991), for the reliable change threshold; Gollwitzer and Sheeran, Implementation intentions and goal achievement (2006), for the if-then action. Every exercise is an original work written for this instrument. No source named here is affiliated with it, and none endorses it. No vendor, product, model or competitor is named anywhere.