Applied Judgment Assessment
Self-understanding · anybody who has to work out what somebody meant · browse the full catalogue

Reading People: Are You Right, and Do You Know When You Are Not?Most emotional intelligence tests ask you to rate yourself. Then they tell you what you just said.

This one has a key. Sixteen short exchanges, one specific claim about each, and a second question after every claim: how sure are you? Being right and knowing when you are not are different things, and almost nothing measures the second.

20 minutes32 scored exercisesEvidence-keyed scoringGlobal · INR & USD

A key, and a confidence rating after every single answer

Self-report is the cheapest way to build an emotional intelligence test and the least informative way to take one. If the instrument asks how well you read people and then reports how well you read people, it has measured your self-image. The instruments that score emotion judgment against an actual key do exist, and they are released through trained practitioners at a few hundred dollars a head, so hardly anybody ever sits one. This is a keyed one, priced so that sitting it is a decision you can make on your own.

Sixteen short exchanges, each written out in three or four sentences. A deadline moved without warning and somebody snapped. A colleague who describes a problem in detail and then goes quiet. A manager who praises three things and names no problems. A person who agreed calmly in the room and raised the same point in writing two days later. After each one comes a single specific claim, and you answer true or false. Eight of the claims are true and eight are false, and several of them are the ones most people get wrong — which is why your accuracy is scored against what people typically answer rather than against a coin toss.

Then the part that makes this different. After every claim you say how sure you are, on a six-point scale, and half of those probes ask about doubt rather than certainty so that agreeing with everything cannot produce a confident-looking profile. From those sixteen pairs the report computes a proper calibration reading: whether your confidence runs ahead of your accuracy or behind it, and — separately, and more usefully — whether your confidence rises when you are right and falls when you are not. Two people with identical accuracy, one of whom can feel the difference between a read and a guess, are not the same person to work with, and one number would show them as identical.

The calibration score is broken into its parts and all of them are printed: how far your stated confidence sat from what actually happened, how much your confidence separated your right answers from your wrong ones, and how much of the total was a property of these sixteen claims rather than of you. Every figure carries its error band drawn on the chart. Each of the four kinds of reading carries a standing and no number, and the reliability table that explains that refusal is printed on the page.

Four kinds of everyday reading, each carried by four keyed claims and four confidence ratings:
What the person is feelingWhat they are actually asking forWhat has changed between two momentsWhat is being left out

What you walk away with

An accuracy figure with its band

How often your read matched the key, scored against what a typical reader would get on these particular claims rather than against fifty per cent, with both spans drawn and written out.

A calibration reading, kept separate

Over-sure or under-sure, with the gap in points and both directions named. It is never averaged with your accuracy, because they are different facts about you.

Whether your doubt means anything

Your average confidence where you were right against where you were wrong. This is the half of calibration that is actually useful, and almost nothing on the market reports it.

The Brier score, taken apart

Reliability, resolution and uncertainty printed separately, so a person who is well calibrated and undiscriminating is not confused with one who is discriminating and over-sure.

Four kinds of reading, classified

What somebody is feeling, what they are actually asking for, what has changed between two moments, and what they are leaving out — each with a standing and a written interpretation, and no number.

One if-then sentence

Chosen from your own pattern: the over-sure and the under-sure get different actions, because they have different problems.

Inside your report

Illustrative sample — your report is generated from your own responses.

Accuracy, against the typical reader
+31
Confidence, against your own accuracy
+14
0 = the typical reader

Being right and knowing when you are not are different facts. They are printed side by side and never averaged.

How sure you were, against how right you were
050100Not sureFairly sureVery surefilled = how sure you saidoutlined = how often right

The Brier score is split into reliability, resolution and uncertainty, and all three are printed, because a person can be well calibrated and undiscriminating.

Does your doubt mean anything?
Average confidence where you were right84%
Average confidence where you were wrong71%
The difference+13

This is the useful half of calibration and almost nobody measures it: whether the feeling of not being sure is doing real work for you, or telling you nothing.

Built for

  • Anyone whose job depends on working out what somebody actually meant from a short message
  • People who suspect they are good at reading others and have never had it checked against anything
  • Managers, teachers, clinicians, salespeople and support staff who want a keyed reading rather than a self-rating
  • Anyone curious whether the feeling of being sure is telling them anything at all

Find out whether you are right, and whether you know

16 keyed claims and 16 confidence ratings · about 20 minutes · no per-item timer · full bespoke report with your accuracy, its band, and your calibration.

Take free · full report ₹249 (incl. GST)

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

How is this different from the emotional intelligence tests I have taken before?

Those ask you to rate yourself and report what you said. This one has a key: sixteen specific claims about sixteen short exchanges, each keyed to published work on emotion, inference and everyday reading, with the reasoning printed for each one. Your score is how often you matched the key, not how highly you rate yourself. A self-report and a keyed instrument answer different questions, and the report says which one this is.

Why does it ask how sure I am after every answer?

Because accuracy on its own hides the thing most worth knowing. Two people can read the same sixteen exchanges equally well and differ completely in whether they can feel which of their reads is shaky. Confidence is collected as a probability, scored with a Brier score, and reported as three separate quantities — how far your confidence sat from reality, how well it separated your right answers from your wrong ones, and how much of the total was a property of the claims rather than of you.

Is my accuracy scored against fifty per cent?

No. Every claim carries an authored prior for how often people are expected to answer it each way, and your accuracy is measured against the average of those. Some of these claims are ones most people get wrong and some are ones most people get right, and they are not worth the same. The report prints the range of those priors and states that observed answer rates will replace them.

Does it measure empathy?

No, and the report says so in plain words. It measures how often your read of a short written exchange matches what the published evidence supports, and how well your confidence tracks that. It is not a measure of empathy, kindness, how much you care about people, your own emotions, your personality or your mental health, and it says nothing about how well you read the specific people you actually work with, about whom you know things no exercise here can give you.

How much is it, and how long does it take?

₹249 in India, inclusive of GST, or US$2.99 elsewhere, one-time. Because it sits below the ₹300 threshold the sitting itself starts free and the report is what you buy. About twenty minutes for sixteen exchanges, with no per-item timer and nothing in the score that uses how fast you answered. Keyed instruments of this kind elsewhere are certification-gated and cost a few hundred dollars a head.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue

Methodology: Sixteen original keyed claims about short written exchanges, each paired with a confidence rating, scored for accuracy and for calibration. Construct statement: It measures two things about reading other people from short written exchanges: how often your read matches what the published evidence on emotion and everyday inference supports, and how well the confidence you attach to each read tracks whether that read was right. It does not measure empathy, kindness, how much you care about people, your own emotions, your personality, your mental health, or how well you read the specific people you actually work with, who you know things about that no exercise here can give you. Declared response instruction: knowledge throughout - every claim asks what is more likely to be true of the exchange, and no item asks what the respondent would do. Construct grounding, drawn across sources rather than from one framework: appraisal theory on which appraisals produce which emotions, and in particular other-blame and perceived preventability behind anger (Lazarus, 1991; Scherer, 2001; Roseman, 1996); attribution theory on controllability and the emotions it selects (Weiner, 1985); display rules and the difference between what is felt and what is shown (Ekman and Friesen, 1969; Matsumoto, 1990); emotion regulation and the behavioural signature of expressive suppression (Gross and John, 2003); the ability model of emotion understanding, used as a public construct with no branded instrument, vocabulary or scale reproduced (Mayer, Salovey and Caruso, 2004); social sharing of emotion and the finding that narration is usually for being heard rather than for being solved (Rime, 2009); politeness theory on softeners as deference rather than optionality (Brown and Levinson, 1987); conversational implicature and the informativeness of an omission (Grice, 1975); the weakness of single-channel behavioural cues and the value of convergent signals in person perception (Funder, 1995, and the realistic accuracy model); and the calibration literature on confidence and accuracy as separable quantities (Lichtenstein, Fischhoff and Phillips, 1982; Fischhoff, Slovic and Lichtenstein, 1977). Scoring design C9: accuracy is chance-corrected against an authored per-claim base rate rather than against fifty per cent, and calibration is a Brier score decomposed into its reliability, resolution and uncertainty parts, so that a person who is well calibrated but undiscriminating is not confused with one who is discriminating and over-sure. Both error directions are named in the instrument's own words - over-sure and under-sure - and both are reported. Reliability is stated as an assumed omega until live data exists; no domain of fewer than eight items carries a number or a percentile. All items are original works. No trademarked instrument, scale name, report section or item is reproduced, and no affiliation with any assessment publisher is claimed or implied.