What is a situational judgement test (SJT)?
A situational judgement test presents realistic work scenarios and asks what you would do. Here's how SJTs work, what the meta-analytic evidence says they predict, and how the response instructions change what you are measuring.
Last updated
A situational judgement test in one sentence
A situational judgement test (SJT) is an assessment that puts a person inside a realistic work scenario — an unhappy customer, a conflicting deadline, a struggling teammate — and measures the quality of their judgement by how they choose to respond.
Unlike a knowledge test, an SJT has no formula to recall. It measures the practical, applied judgement that separates people who know the right policy from people who make the right call under pressure.
How an SJT works
Each item describes a concrete situation and offers several plausible responses. Depending on the format, the candidate picks the best response, ranks all of them, or rates each option's effectiveness. The best SJT options are all defensible — the test discriminates between good and better judgement, not right and obviously wrong.
Scoring is anchored to a key built by subject-matter experts: each response option carries a value based on how effective experienced practitioners judge it to be. Some platforms, AssessAll included, also use rubric-based AI grading for open-ended scenario responses, so a written 'what would you do and why' answer can be scored consistently at scale.
What SJTs predict — and why employers use them
Decades of selection research place SJTs among the stronger predictors of job performance, particularly for roles where interpersonal judgement matters: customer service, sales, team leadership, healthcare, and frontline operations.
Employers use them for three reasons. They are job-relevant, so candidates experience a fair preview of the actual work. They are harder to fake than a self-report questionnaire, because socially desirable answers are built in as tempting-but-suboptimal options — though how much harder depends on the response instructions, which is the design choice covered below. And they work at volume: an SJT can screen thousands of applicants consistently before a single interview is scheduled.
What the evidence actually says, with the numbers
The largest meta-analysis on the question — McDaniel, Hartman, Whetzel and Grubb, Personnel Psychology, 2007 — put the corrected correlation between situational judgement scores and job performance at 0.26, and found the same 0.26 whether the test asked what a candidate should do or what they would do. That is a real, useful predictor and it is not a large one; treat any vendor quoting a dramatically higher figure as owing you the study.
Read it next to the wider revision of selection validity published in 2022, when Sackett, Zhang, Berry and Lievens re-analysed the field's meta-analytic estimates and found that corrections for range restriction had been applied too aggressively for decades. Their revised figures put structured interviews at 0.42, job knowledge tests at 0.40, work samples at 0.33 and cognitive ability at 0.31 — where cognitive ability had been quoted at 0.51 since 1998. An SJT at 0.26 sits inside that band rather than below a much higher one, and the practical conclusion is unchanged: combine instruments that measure different things rather than stacking two that measure the same thing.
What the coefficients do not capture is why organisations keep buying SJTs: they are the cheapest way to put a realistic preview of the job in front of a large applicant pool, and face validity — how relevant the test looks to the person taking it — drives completion and acceptance even though it is not itself validity evidence.
"Should do" or "would do": the design choice that changes what you measure
An SJT is a method, not a construct. What it measures depends heavily on how the question is asked, and the same meta-analysis quantified the difference. Knowledge instructions — what should be done — correlate with cognitive ability at 0.35. Behavioural tendency instructions — what you would do — correlate with cognitive ability at only 0.19, and instead correlate with personality: agreeableness 0.37, emotional stability 0.35, conscientiousness 0.34, against 0.19, 0.12 and 0.24 for the knowledge form.
So the two forms predict performance equally well and measure noticeably different things. If a battery already contains a reasoning test, a knowledge-instruction SJT partly duplicates it; a behavioural-tendency SJT adds something the reasoning test does not have. If the battery already contains a personality inventory, the reverse holds.
The trade-off runs the other way on faking. Behavioural tendency instructions ask a candidate to describe themselves, which invites social desirability bias in exactly the setting where the incentive to distort is highest. Knowledge instructions ask what a competent person should do, which is much harder to inflate and, for that reason, easier to coach for.
SJT vs aptitude vs personality tests
An aptitude test measures whether someone can reason through problems; a personality or behavioural profile describes how they tend to operate; an SJT measures what they would actually do in a specific situation. The three answer different questions, which is why strong screening batteries combine them.
A contact-centre battery, for example, might pair a customer-interaction SJT with a listening test and a typing test; a leadership pipeline might pair a management SJT with a reasoning test and a DISC profile.
SJTs on AssessAll
AssessAll supports situational judgement as a first-class item type, scored by expert keys for structured formats and by rubric-based AI grading for written scenario responses.
SJTs appear across the catalogue — customer interaction judgement for contact-centre hiring, workplace reliability screening, and leadership scenarios — and every result feeds the candidate's competency scores and Skill Passport, so a good judgement score becomes verifiable evidence, not a one-off number.
Whose judgement is the key? The question that decides what the score means
The common claim that an SJT has no right or wrong answers is wrong in a way that matters, because you are scored against a key. What is true is that the key is not an objective fact. The US Office of Personnel Management describes SJT scoring plainly: it relies on subject-matter experts' judgements of the best and worst alternatives. So an SJT measures agreement with expert judgement about a situation — and which experts, from which organisation, is a question every buyer should ask and almost none do.
That has a consequence for anyone buying an off-the-shelf SJT. A key built by one organisation's experts encodes that organisation's norms about escalation, disclosure and when contradicting a peer is a correction rather than an attack. On items where the top two options are both honest responses separated only by timing, a borrowed key can be defensibly wrong for you. Ask whether the key can be re-weighted by your own subject-matter experts; on AssessAll forms it can.
Three ways to build a key are recognised in the literature. A consensus key takes the judgement of experts or of a reference group. An empirical key weights options by how they correlate with an actual performance outcome. A rational or theoretical key derives weights from job analysis or from stated principles. AssessAll's workplace judgement bank uses the third: option weights follow professional-conduct principles — transparency, escalation order, scope discipline — as recorded in the assessment's own methodology note. That is a real basis and it is not an empirical one, and the difference is worth stating rather than blurring.
How a selection becomes a score, with the arithmetic
Each option carries a weight on a 0–3 gradient: 3 for the best-supported response, 2 for a reasonable one, 1 for weak-but-not-harmful and 0 for the least effective. The item's points are then multiplied by the selected option's weight divided by the highest weight on that item, so choosing a 2 where the maximum is 3 earns two-thirds of the item's points. The rule lives in `lib/ai/sjt-scoring.ts`.
Two things follow that are worth knowing before you buy. First, single-select SJT scoring on AssessAll is rule-based rather than model-graded — the same answer always earns the same score, and no part of that result depends on a language model. (Written scenario responses, where a candidate types an answer rather than choosing one, are a different item type and are graded against a rubric.) Second, partial credit is the whole argument for a graded key: a right/wrong key can only tell you that someone avoided the worst option, while a graded key separates the candidate who discloses immediately from the candidate who discloses eventually, which is usually the distinction the hiring decision turns on.
What the evidence adds up to, in one paragraph
Two large meta-analyses bracket the method. McDaniel, Morgeson, Finnegan, Campion and Braverman (Journal of Applied Psychology 86(4), 2001, 730–740) reported an operational validity of .34 against job performance; McDaniel, Hartman, Whetzel and Grubb (Personnel Psychology 60(1), 2007, 63–91) reported .26, and — the figure that should govern the purchase — found that SJTs add roughly one to two percent of criterion variance once general mental ability and personality are already in the model. An SJT is a real predictor and a modest incremental one, which is an argument for using it where it measures something your battery does not, not for using it as the deciding score. The literature has since been mapped in an integrative review of 524 documents (Kepes, Keener, Lievens & McDaniel, Journal of Management 51(6), 2025, 2278–2319).
For how a numerical or reasoning score behaves differently — and why a timed form and an untimed form of the same construct are not interchangeable — see what a numerical reasoning score actually means.
Frequently asked questions
Do situational judgement tests have right answers?
Not in the way a maths test does. Every option is usually plausible; scoring reflects how effective experienced practitioners judge each response to be. Some options score full marks, others partial credit, and a few — typically the avoidant or aggressive choices — score poorly.
How do I prepare for an SJT?
Learn the role and the organisation's values, then practise with realistic scenarios. When answering, favour responses that take ownership, address the problem directly, and consider both the person and the task. Avoid options that pass the problem to someone else or delay action without reason.
Are SJTs reliable enough to hire with?
Used well, yes, with the numbers in front of you. The 2007 McDaniel meta-analysis puts the corrected correlation with job performance at 0.26 — a genuine predictor, and a modest one. That is an argument for combining an SJT with an instrument that measures something different, such as a structured interview or a work sample, rather than for deciding on any single score.
Should our SJT ask what a candidate would do, or what they should do?
Both forms predict job performance about equally well (0.26 in the McDaniel meta-analysis). They measure different things: 'should do' correlates with cognitive ability at 0.35, while 'would do' correlates with personality traits in the 0.34 to 0.37 range. Choose the one that adds to what the rest of your battery already measures, and note that 'would do' is the more fakable of the two.
How much does an SJT add to a battery that already has a cognitive test?
Less than most vendors imply, and it is measurable. The 2007 McDaniel meta-analysis found SJTs explain roughly one to two percent more criterion variance once general mental ability and personality are already in the model. That is a genuine gain and a small one, so the case for an SJT is strongest where it measures something the rest of your battery does not touch — judgement about disclosure, escalation and competing obligations — and weakest where it is added as a second general-ability measure under a different name.
Can we change the scoring key?
You should expect to. A key encodes a judgement about how a particular organisation expects people to behave, and on items where the two strongest options are both honest responses, a key built elsewhere can be defensibly wrong for you. On AssessAll forms the option weights can be re-set by your own subject-matter experts, and the scenarios can be rewritten to your sector and seniority. Ask any vendor the same question; an instrument whose key cannot be examined cannot be defended.
Can an SJT be taken online without supervision?
Yes. Most SJTs run unsupervised online, and because the options are all plausible, they are harder to game than knowledge tests. For higher-stakes screening, AssessAll can add AI proctoring — webcam, screen, and behaviour monitoring — to an SJT sitting.