AI Agent Oversight and Delegation Assessment for Knowledge WorkThe question is never whether the tool is good. It is what it may do without asking.
Sixteen authority calls on one printed ladder, four of them matched pairs differing only in whether the action can be taken back.
The decision arrives forty times a week and almost nobody has written it down
Everybody now has tools that act rather than just draft, and the grant of authority gets made by feel: this one can just get on with it, that one I had better read. Forty exercises make the decision explicit and then check whether it holds up.
Sixteen of them print the same four-rung ladder — draft only, show me first, tell me at once, summary later — and ask which rung a real task warrants. Filing your own photographs and deciding which job applications go forward are not the same decision, and a person who answers them the same way is applying a habit.
Eight of those sixteen are four matched pairs: the same task twice, differing only in whether the action can be taken back. Two of the pairs should move your answer and two should not, because something else already fixes the rung. That balance is what stops always granting less when something is irreversible from scoring as judgement.
Credit is graded by distance from the warranted rung on a table whose four columns have identical means. Always asking first and always letting it run score exactly the same, and both score behind reading the case. The table is printed on your report with those means beside it.
Whether you grant more or less than the tasks warrant is reported as its own signed number and is deliberately kept out of the headline. A table that paid more for caution would train caution and call it judgement, and granting too little is a real cost that is paid invisibly, as work nobody automates.
What you walk away with
With its error band, and against a typical respondent rather than against a coin.
Granting more or less than the tasks warrant, drawn on a beam with both directions labelled in words.
Which of them should move with reversibility and which should not, and what yours did on each.
Four by four, with the column means beside it, so you can see why no fixed answer pays.
Including what carrying every outcome personally costs, which is authority never granted.
Drawn from whichever pattern your own sitting showed most clearly.
Inside your report
Illustrative sample — your report is generated from your own responses.
| If the task warrants… | Draft only | Show me first | Tell me at once | Summary later |
|---|---|---|---|---|
| Draft only | 3 | 2 | 0 | 0 |
| Show me first | 2 | 3 | 2 | 1 |
| Tell me at once | 1 | 1 | 3 | 2 |
| Summary later | 0 | 0 | 1 | 3 |
| Mean for always answering this rung | 1.50 | 1.50 | 1.50 | 1.50 |
The four column means are identical by construction. Always asking first and always letting it run score exactly the same, and both score behind reading the case.
Both directions are labelled in words. Granting too little is a real cost paid invisibly, and a report that only warns about granting too much is training caution and calling it judgement.
Built for
- Anybody whose tools have started acting rather than only drafting
- Managers deciding what an assistant, agent or automation may do unsupervised
- Operations, finance, HR and support teams writing their first grants of authority
- Risk, audit and compliance functions who need a measure rather than a policy acknowledgement
Write down what it may do, before it does it
40 exercises across six formats · about 40 minutes · the credit table and the matched pairs printed.
₹899 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
No. It tests no model, product, vendor or framework, and it does not test prompting. Every exercise is about the task and what the task changes, because that is what sets the level of authority rather than how good the tool is.
Because reversibility is not always the thing that decides. A reply a customer has already read cannot be recalled in any useful sense whether or not the software has an undo. Half the pairs move and half do not, so moving on all of them and moving on none both score at chance.
Yes, and it is the one nobody measures. It is paid as work that never gets automated and as approvals performed at a rate nobody can sustain. The report gives severity as a signed number with both directions labelled, and neither direction is treated as the good one.
No. It is a developmental measure and it carries no pass mark. It does not certify anybody, and it says nothing about whether any particular tool is safe to use.
About forty minutes for forty exercises. ₹899 in India, inclusive of GST, or US$8.99 elsewhere, one time, for the sitting and the full report.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Forty original exercises across six formats: sixteen authority-level calls printed against the same four-rung ladder, six select-every-that-applies exercises, six keyed claims, four matching exercises, four ordering exercises and four most-and-least forced choices. There is deliberately no situational judgement item: the declared response instruction is knowledge and applied judgement throughout - what level this task warrants - so that one instruction covers the whole form. Construct statement: it measures how somebody decides how much authority an automated tool should have over a piece of work, whether reversibility actually moves that decision, how they would detect an error after the fact, and where they place accountability. It does not measure technical knowledge of any model, product or vendor, it does not test prompting, it is not a governance or compliance certification, and it says nothing about whether any particular tool is safe. Scoring is a warranted-level match. Every call carries the rung the situation warrants and credit is graded by distance from it on a four-by-four table which is published on the report with its column means printed beside it. The columns have identical means by construction, which is the property that stops any fixed answer paying: always asking first and always letting it run score the same, and both score behind reading the case. Severity - whether a respondent grants systematically more or less than the situations warrant - is reported as its own signed number rather than priced into the headline, because a table that pays more for caution trains caution and calls it judgement. The eight matched reversibility items are balanced two pairs to two: on two of them the warranted level changes with reversibility and on two it does not, so moving on every pair and moving on none both land at chance as arithmetic rather than as a threshold. Constructs and sources: levels of automation and the allocation of function between person and machine (Sheridan and Verplank's levels; Parasuraman, Sheridan and Wickens on types and levels of automation); automation bias and complacency, including the finding that both omission and commission errors rise with automated aids (Parasuraman and Manzey; Skitka, Mosier and Burdick); the irony that automating the easy part leaves the person monitoring, which people do badly (Bainbridge); vigilance decrement under high-rate approval work; the reversibility of a decision as the basis for how much scrutiny it deserves (the one-way and two-way door distinction); accountability that does not transfer to a tool or its supplier, and disclosure as a duty to the reader rather than a transfer of responsibility; detection as a separate control from prevention, and sampling as a stronger real-world check than exhaustive review performed at speed (Deming and Juran on inspection); and item-writing guidance from Haladyna, Downing and Rodriguez. All items are original works written for this instrument. No vendor, product, framework or standard is reproduced, named or implied, and no endorsement or affiliation exists or is suggested.