Applied Judgment Assessment
Leadership battery · people managers, team leads and the HR partners who advise them · browse the full catalogue

Team Read Accuracy Assessment for People ManagersYou rate them every year. Nobody ever tells you whether you were right.

Sixteen readings of four people, matched against what the record in front of you supports. About forty minutes.

40 minutes40 scored exercisesEvidence-keyed scoringGlobal · INR & USD

Every management instrument points at the manager. This one points at the people they manage.

Ask what a management assessment measures and the answer is always the manager: their style, their traits, their judgement in a hypothetical. That is a real thing to measure and it is not the thing a manager is paid to be accurate about. A manager is paid to be accurate about the people in front of them, and almost nobody in this market measures accuracy against evidence at all.

So this sitting gives you four people and a record of what each of them did over a year. Dates. Counts. What was delivered and when. What was raised and how late. Sixteen readings then ask you to place each person on four things that matter: finishing on the date agreed, naming a problem early, taking on work nobody owns, and staying accurate when the week is busy. The record supports a level on every one of those, and the four people have genuinely different shapes on purpose.

Your sixteen readings are matched against the sixteen the record supports as a CORRELATION rather than a distance, and that choice is the design. A correlation is unchanged if every one of your numbers moves up or down together, so reading the whole team warmly costs you nothing on the main figure — it is reported separately, as its own finding, with its own fix. What the correlation measures is whether you see the shape: who is strong where, and who is not.

Three readings come out of it and they are independent of each other. The shape, which is the headline. The level, which is whether you run warm or cold against the record. And the spread, which is whether you separate people at all — because sixteen readings all between four and five carry almost no information whatever their shape, and a report that scored that as average would be hiding the most useful thing on the page.

A respondent who gives every reading the same number gets no figure at all, and the report says why in plain words. A flat answer has no profile to match, and placing it in the middle would score it above somebody who actually read the record. Twenty-four further exercises cover which evidence is strong, what a record supports and where a reading goes past it, whether a change in somebody is about them or about what changed around them, and three written exercises asking for the one sentence you would actually say.

Four parts of one skill, each measured by at least six independent exercises:
Telling the record from the impressionReading the shape, not just the levelSituation or person: what the behaviour is aboutPutting the read into words somebody can use

What you walk away with

A shape figure with both spans

How closely your readings follow the record's pattern, drawn as a band rather than a point, with the 68 and 95 per cent spans written out in plain words.

Your level, reported separately

Whether you read the team warmer or cooler than the record. The headline ignores it on purpose, because it is a different problem with a different fix.

Your spread, reported separately

How far apart your numbers sit against how far apart the record's sit. Rating everyone between four and five is a finding, not an average.

A dumbbell of every reading

One row per person and attribute, your reading and the record joined by a line, sorted so the widest gap is the top row.

Four parts as strength and cost

Each part with what it gets you and what it costs when it runs unchecked, so no sentence in the report reads as generic praise.

One if-then change

Chosen from your own lean — warm, cool, or flat — naming a specific situation and a specific behaviour rather than a competency.

Inside your report

Illustrative sample — your report is generated from your own responses.

Your read against the record, widest gap first
Nadia · Taking on unowned work-3Tom · Naming a problem early+2Meera · Accuracy under pressure+2Ravi · Finishing on the date0

Shape as well as hue: a filled circle is your reading, an open square is the record. The row to look at is the top one.

Three readings, kept apart on purpose
Shape
41

Whether your readings follow the record's pattern. This is the headline.

Level
+0.9

How warm or cool you run overall. A correlation ignores this, so it is reported on its own.

Spread
0.6

How far apart your numbers sit against how far apart the record's sit.

Reading a whole team warmly and reading the wrong person as the strong one are different problems with opposite fixes.

Every reading, as a table as well as a drawing
PersonAttributeYouRecord
RaviNaming a problem early42
MeeraTaking on unowned work45
TomTaking on unowned work42

Nothing on the page is load-bearing in a hover, and every chart has the same information underneath it as a table.

Built for

  • People managers who write ratings and never find out whether they were accurate
  • Team leads about to run their first performance cycle
  • HR partners calibrating managers before a review round
  • Anybody who has been surprised by somebody they thought they knew well

Four people, sixteen readings, one record to match them against

40 exercises across six formats · about 40 minutes · a shape figure with both spans, your level and spread reported separately, and every reading drawn against what the record supports.

₹1,699 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and Razorpay sends your receipt to that address.

Secure Razorpay payment · ₹1,699 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

Are the four people real?

No. They are invented, and their records are authored so that the four of them have genuinely different shapes across the four attributes. That is what makes a correlation meaningful: if everybody in the set looked the same, there would be nothing for a reading to match.

What happens if I read the whole team warmly?

Nothing on the headline, and that is deliberate. The headline is a correlation, which is unchanged when all your numbers move together. Your level is then reported on its own, with its own reading and its own fix, because reading a team warm and reading the wrong person as the strong one are different problems.

What if I give every reading the same number?

You get no figure, and the report says so in plain words. A response with no variation in it expresses no profile, so there is nothing to match it against, and placing you in the middle would score a flat answer above a varied one. The sitting is flagged for review rather than scored.

Does this tell me whether I am a good manager?

No, and the report says so directly. It measures the accuracy of your read against a written record, at what level you use the scale, and whether you can say a read in words somebody could act on. It says nothing about how you coach, how your team feels, or how you perform.

How long is it and what does it cost?

About forty minutes for forty exercises. ₹1,699 in India, inclusive of GST, or US$16.99 elsewhere, one-time, for the sitting and the full report.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue

Methodology: Forty original exercises across six formats: sixteen evidence-anchored ratings of four people on four attributes, eight keyed single-choice items on the strength of evidence and the attribution of behaviour, six select-every-that-applies items on what a record supports, four match-the-following exercises pairing a pattern with the reading it supports, three true or false claims about person perception, and three written exercises graded against a published rubric and reported as their own strand rather than folded into the score. Construct statement: it measures the accuracy of a manager's read of individual people against the evidence in front of them - the shape of the read, the level at which the scale is used, and the spread across people - and the ability to say the read in words the person could act on. It does not measure how good a manager somebody is, how much their team likes them, their own performance, their personality, or their fitness for any particular role. Declared response instruction: knowledge and judgment throughout. Every exercise has an answer the record supports, and none of them asks for a preference. Scoring: the sixteen ratings are scored as a PROFILE-MATCH CORRELATION between the respondent's sixteen readings and the sixteen levels the record supports, chance-corrected against a Monte-Carlo null drawn from an authored prior on every scale point of every rating. A correlation is invariant to linear transformation, so elevation and scatter drop out of the headline for free and are reported separately as leniency and as differentiation. A respondent who gives every rating the same value has no shape to match and receives no figure at all, with the reason printed, rather than being placed at the centre of the scale where a flat answer would outscore a varied one. Reliability is reported as McDonald's omega rather than alpha, stated as an assumption from the item count until observed data replaces it, and any part too short to carry a number is given a three-way placement with the arithmetic printed. Construct grounding, drawn across sources rather than from one framework: the accuracy of interpersonal judgment and the separation of elevation, differential elevation and stereotype accuracy from pattern accuracy (Cronbach, 1955; Kenny, 1994); the lens model and the decomposition of judgment accuracy into achievement, consistency and matching (Brunswik, 1952; Hammond, 1955; Karelaia and Hogarth, 2008); halo error and the collapse of distinct attributes into a general impression (Thorndike, 1920; Cooper, 1981); leniency, severity and central tendency as rater effects distinct from accuracy (Landy and Farr, 1980; Murphy and Cleveland, 1995); the correspondence bias and the actor-observer asymmetry in explaining behaviour (Ross, 1977; Jones and Nisbett, 1971; Malle, 2006 on the size of the asymmetry); frame-of-reference training as the rater intervention with the strongest evidence behind it (Woehr and Huffcutt, 1994; Roch et al., 2012); the availability heuristic and why a vivid event is read as a frequent one (Tversky and Kahneman, 1973); base rates and the diagnosticity of evidence that could have come out the other way (Meehl and Rosen, 1955); feedback that describes behaviour at the task and process level rather than the person (Kluger and DeNisi, 1996; Hattie and Timperley, 2007); implementation intentions as the form of a change instruction that actually gets carried out (Gollwitzer and Sheeran, 2006); McDonald's omega as the reliability estimate reported in place of alpha (McDonald, 1999); and the standard error of measurement as the reason a band rather than a point is reported (AERA, APA and NCME Standards, 2014). All items are original works. No item, scale name or report section is taken from any commercial instrument, and no affiliation with or endorsement by any vendor is claimed or implied.