Applied Judgment Assessment
Speak-up judgment · all staff and their managers · browse the full catalogue

Speak-Up: Raising a Concern and Hearing OneSpeak-Up: Raising a Concern and Hearing One. Buy it before the programme. Run it again after. It tells you, on day one, how much movement would count as real.

Almost every assessment sells a snapshot. This one is built as a matched pair of half-forms, so a single sitting can publish its own re-test threshold: the number of points a later score has to move before the movement is bigger than the measurement error. Then it books the date.

42 minutes32 scored exercisesEvidence-keyed scoringGlobal · INR & USD

An instrument designed to be taken twice, and honest about what a second reading can prove

Speak-up judgment is what a person does with a concern at work: whether they raise it while it is still small, whether they say what actually happened rather than how they feel about it, what they do when somebody brings one to them, and whether the person who raised it is ever told what happened next. Four behaviours, twenty-four situations, eight self-checks. Not courage, not honesty, not how safe the team happens to be — those are different things and this does not claim to measure them.

The twenty-four situations are twelve matched twins. Each twin is the same underlying behaviour and the same keyed action through a different cover story, one in half A and one in half B, at comparable difficulty. That structure is why the report can do the thing almost nothing else on the market does: state its own change threshold from one sitting. It publishes SEdiff = SD × √(2 × (1 − ω)) and a threshold of 1.96 × SEdiff in score points, declares the standard deviation and the reliability it used as assumptions rather than measurements, and commits in advance to the rule — a later movement smaller than the threshold is drawn flat and greyed, captioned within measurement error, never as an arrow. Publishing that number before the second sitting exists is the only moment at which it costs nothing to publish, and it is why hardly anybody does it: a change threshold is an uncomfortable figure, and a vendor who states one in advance cannot quietly relax it later to make a programme look effective.

Two more things it does that a speak-up instrument usually will not. Six of the twenty-four situations key the more restrained action — wait until you have the fact, go to the person before you go over them, let a first instance be a first instance — because raising everything immediately to everybody is a real failure and it is the one that gets called courage. And the agreement between the two halves is reported as a property of the FORM. If your half A and your half B disagree by more than the band, the report says the halves did not behave as matched versions of each other for this sitting, that the question that raises is about the instrument, and that nothing in your result is marked down because of it. Charging our own measurement error to the person sitting the test is the one thing we will not do.

Four working behaviours, each carried by six matched situations and two self-checks, and each reported as a classification rather than a number:
Saying It While It Is SmallSaying What Actually HappenedHearing It Without DefendingTelling Them What Happened Next

What you walk away with

Your own re-test threshold, in points

How far a later score has to move before the movement is real rather than noise — with the arithmetic, the assumed reliability and the assumed spread all printed, not buried.

A now / 30 / 90 rail that ends in a date

One action and one observable at each station, chosen from your weakest behaviour and the single situation inside it where the gap was widest, terminating in a booked re-measurement.

The half-to-half check, framed as ours not yours

Twelve twins, one mark each, showing where the two halves agreed. If they did not, the report says that is a question about the instrument and adjusts nothing about your result.

Both directions of error, named and opposed

Took it up before you had the fact, against let it go until it was somebody else's problem. The signed count is printed and the absolute count decides the reading.

Four behaviours classified, never scored

Six situations each cannot support a number or a percentile, so each carries a three-way classification with its interpretation, plus each half separately so you can see where they diverged.

Inside your report

Illustrative sample - your report is generated from your own responses.

Now, day 30, day 90 — then measure again
NowDay 30Day 90Re-measure1 Dec 2026needs a move over 35.2 ptsday 0
Now

Go back to the oldest concern you never answered.

One open loop is closed, and they heard it from you.

Day 30

When you cannot act yet, give a date rather than a reassurance.

Every concern you took this month has a date on it.

Day 90

Go back on one where the answer was no, with the reasoning.

Somebody has heard a no and could argue with it.

One action and one observable per stage, chosen from your weakest behaviour and the situation inside it where the gap was widest — ending in a booked date with the threshold already printed on it.

What would count as real movement

A later sitting has to land more than 35.2 points away from today’s score before the movement is bigger than the measurement error.

today 41drawn flat — within measurement errora real falla real riseworth having (15)edge of the flat zone (35.2)

The hatched band is the rule, published before there is a second sitting to apply it to: a movement inside it is drawn flat and greyed, never as an arrow. The dashed marks are a separate and less stringent number, and the report says plainly that it is a value judgement rather than a measurement.

Judgement-pattern analysis
21
Strongest chosen
8
Over-acting
3
Under-acting
2
Passed on
The Doer Reflex — emerging pattern · counter-habit included

Built for

  • L&D teams about to run a speak-up, raising-concerns or psychological-safety programme who want a baseline they can defend and a matched re-measurement rather than a happy sheet
  • All staff and their managers — half the situations put you in the position of the person raising a concern and half in the position of the person hearing one
  • Anyone who has been told to speak up more, or to listen better, and would like to know which of the four behaviours the advice is actually about

Get the baseline, and the number that will tell you whether the programme worked

32 scored exercises - about 42 minutes - a bespoke report with your re-test threshold, a now/30/90 rail ending in a booked date, and both directions of error named.

₹1,199 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and Razorpay sends your receipt to that address.

Secure Razorpay payment · ₹1,199 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

What does this measure, and does it tell me how safe my team is?

It measures one person's judgment about concerns at work: raising them while they are still small, describing what actually happened rather than how it felt, hearing one without defending, and telling the person who raised it what came of it. It is NOT a measure of how safe your team actually is. Psychological safety is a property of a group and of the people running it, and it cannot be read off one person's answers to twenty-four situations in one sitting. If you want to know whether people in your organisation feel able to speak, run a climate survey; this tells you what an individual does with a concern once they have decided to raise or receive one, which is a different and complementary thing.

What is the re-test threshold, and is the second sitting included?

The second sitting is a SEPARATE purchase at the same price; nothing about a retake is bundled into this one. The threshold is what the first sitting buys you: from an assumed population spread and an assumed reliability, both declared on the report as assumptions rather than measurements, it publishes SEdiff = SD x root(2 x (1 - omega)) and a 95% threshold of 1.96 x SEdiff in score points. A later score has to land more than that many points away before the movement is bigger than the measurement error. A smaller movement is drawn flat and greyed and captioned within measurement error, never as an arrow. A second, less stringent number - the minimally important difference - is printed beside it and labelled clearly as a value judgement about what would be worth having, not a measurement of anything.

How much does it cost, and how does that compare?

Rs 1,199 in India including GST, or US$11.99 elsewhere, one time, for one full sitting and report. Organisations can use AssessAll credits at 40 credits per person. For comparison: per-seat speak-up and raising-concerns e-learning typically runs a few hundred rupees to a few dollars a seat and teaches rather than measures, so it can tell you completion but not change. Organisation-priced psychological-safety survey platforms are usually sold as an annual contract in the tens of lakhs or tens of thousands of dollars and measure the climate of a team, which is a different construct from an individual's judgment, and they do not give a person a report they own. This sits between the two on purpose and is priced per person with nothing recurring.

Why is the assessment built as two matched halves?

Because a change statistic needs a defensible estimate of measurement error, and a matched design is how you get one from a single sitting. The twenty-four situations are twelve twins: each twin is the same behaviour and the same keyed action through a different cover story, at comparable difficulty, with one member in each half. The report then compares the two halves and tells you whether they behaved as matched versions of each other for your sitting. It reports that as a property of the instrument, not of you, and if they diverged nothing in your result is marked down.

Does it depend on any country's law, or on my reporting an incident?

No. No item depends on any country's employment law, whistleblowing statute, regulator or named reporting product, and nothing in it involves harassment, abuse, self-harm or criminal conduct. The concerns are ordinary and work-shaped: a checking step skipped to save time, an expense line that looks wrong before you have seen the detail, a colleague spoken over once, a deadline nobody believes, a risk somebody flagged a month ago that was never answered. That is deliberate: it keeps the instrument usable in every market we sell in, and it keeps the difficulty in the judgement rather than in the vocabulary.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue

Methodology: Measures what a person does with a concern at work, through original situational and self-report items keyed to published constructs. Construct statement: It measures what a person does with a concern at work - whether they raise it while it is still small, whether they say what actually happened rather than how they feel about it, what they do when somebody brings a concern to them, and whether the person who raised it is ever told what happened next. It does not measure courage, honesty, how safe a person's team actually is, their standing at work, or whether they were right about the concern. Declared response instruction: behavioural tendency throughout - every stem asks what the respondent is most likely to do, never what a person should do, because instructed-knowledge framing measures knowledge of the espoused answer and is the more fakeable of the two. Item format: twenty-four four-option situational items authored as TWELVE MATCHED PARALLEL TWINS, twelve in half A and twelve in half B, each twin carrying the same underlying behaviour, the same competency, the same partial-credit gradient and the same keyed action type through a different cover story; alignment across the two halves is carried by the twin tag and by aligned per-option arrays, never by option index, because the key-balancer permutes each item's options independently. Twelve items are written from the position of the person raising a concern and twelve from the position of the person receiving one, and both halves contain both positions. Six of the twenty-four key the more restrained action - wait for the fact, go to the person first, let a small thing be small - so that raising everything immediately to everybody cannot be a winning strategy. Eight balanced-keyed self-check statements carry the trait, four of them reverse-scored and worded so that agreeing is not the flattering answer. Scoring: graded partial credit across the answered situations, chance-corrected against authored per-option base rates, because a graded percentage has a floor near half by construction and an uncorrected figure would put nearly everybody in the middle band; the raw percentage is printed beside the corrected one. The instrument's product is a RELIABLE CHANGE INDEX: from an assumed population standard deviation and an assumed reliability, both declared on the report as assumptions pending live data, it publishes SEdiff = SD * sqrt(2 * (1 - omega)) and a threshold of 1.96 * SEdiff, and states that a later sitting must land more than that many points away before the movement is larger than the measurement error. A separate and less stringent minimally-important difference is printed alongside it and labelled a value judgement rather than a measurement. The agreement between the two halves is reported as a property of the INSTRUMENT and never as a finding about the reader, and the twin-level agreement count is labelled descriptive rather than a reliability coefficient, because reliability is a population statistic and cannot be computed from one person. Each of the four competencies rests on eight items and therefore carries a three-way classification rather than a number or a percentile. Both directions of error are named in the instrument's own words and the absolute count decides the reading. Construct areas drawn on: employee voice and silence research, including the distinction between prosocial voice and defensive silence; psychological safety as the condition under which a concern reaches a listener; upward communication and the mum effect, the tendency to withhold unwelcome news from those above; the specificity principle in behavioural description, that a concern lands when it names an observable event rather than an attribution; feedback-seeking and non-defensive listening research on what a first response teaches an observer; the bystander and diffusion-of-responsibility literature on why a visible problem goes unraised; escalation and voice-target research on the cost of going upward before going across; procedural and interactional justice, where being told the outcome and the reasoning predicts acceptance more than the outcome itself; loop closure and response-to-voice research on why an unanswered concern predicts silence next time; graded partial-credit situational judgement measurement; the reliable change index and its distinction from a minimally important difference; and parallel-forms reliability as the basis for a matched half-form design. No item depends on any country's employment law, whistleblowing statute, regulator or named reporting product, and no item involves harassment, abuse, self-harm or criminal conduct. All items are original works, no trademarked instrument or branded methodology is named or reproduced, and no affiliation with any source is claimed. AssessAll original design.