Test library
Test library · how it is scored

What is an abstract reasoning test, and is it really culture-fair?

An abstract reasoning test asks a candidate to infer an unstated rule from symbols, shapes or series rather than from words or numbers. Because the items carry little language, the format is widely sold as culture-fair. That claim describes the item medium, not the score, and the published evidence for it is considerably weaker than the market implies.

SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed

The form we publish, at a glance

Instrument
Abstract & Pattern Reasoning, Inductive Reasoning, and Cognitive Agility & Learning Potential
Length
22 items (Abstract & Pattern) · 20 (Inductive) · 24 (Cognitive Agility) — 55, 66 and 55 seconds per item respectively
Time
20 minutes (Abstract & Pattern Reasoning) · 22 minutes (Inductive Reasoning) · 22 minutes (Cognitive Agility & Learning Potential)
Level
Campus and early careers (Abstract & Pattern, Cognitive Agility) · analysts, researchers and experienced hires (Inductive Reasoning)
Item format
Multiple choice. Text-renderable symbol, letter and number patterns — no figural image bank. Inductive items use short business observations as well as series
Scoring
Right/wrong against a derivable key; no negative marking
What you get back
A banded score filed under Abstract & Pattern Reasoning, plus a second and sometimes a third competency that differs by form

What it measures

Abstract & Pattern Reasoning

Extracting a rule that was never stated, from structure rather than from vocabulary or schooling. All three forms report against this competency, which is why the label on the report does not tell you which form produced it.

Related competency definition

Cognitive Speed & Agility (Abstract & Pattern and Cognitive Agility forms)

How much correct work is produced inside a fixed window. On the Cognitive Agility form this is deliberate and stated in the product's own note: there are more questions than most candidates finish, and accuracy under time pressure is the measure.

Related competency definition

Logical Reasoning (Inductive Reasoning and Cognitive Agility forms)

Whether a generalisation is actually supported by the evidence given — including refusing the causal leap that a correlation invites, which several inductive items are keyed to reward.

Related competency definition

"Language-light" is a fact about the items. "Culture-fair" is a claim about the scores, and it is not the same claim

Start with our own copy, because it is the clearest example of the error. The product note for AssessAll's Abstract & Pattern Reasoning form says performance "depends on extracting rules, not on schooling or vocabulary, which also makes this the fairest test across backgrounds". The first half of that sentence is a description of the item format and it is accurate. The second half is a claim about group differences in scores, and nothing in the first half establishes it. That sentence is being raised with the product owner rather than quietly edited, because a public page correcting the vendor's own catalogue is worth more than a catalogue that has been silently tidied.

The lane says the same thing in softer words, checked at source on 24 September 2026. Testlify's abstract reasoning test page publishes both numbers a buyer needs — 12 questions, 15 minutes, which is more than most of this category manages — and then says the test will "provide a fair and unbiased measure of a candidate's cognitive abilities, reducing selection bias and ensuring a more equitable hiring practice". It publishes no reliability coefficient, no standard error of measurement and no norm group. TestInvite's employer-facing page on pre-employment abstract reasoning says the format reveals ability "without depending on prior knowledge or expertise" and "without relying on verbal or numerical cues", and publishes no duration, no item count and no psychometric properties at all. Neither vendor is doing anything improper. Both are describing the item medium and inviting the reader to hear a conclusion about fairness.

The meta-analysis the claim is usually hung on does not contain it. Roth, BeVier, Bobko, Switzer and Tyler (2001), in Personnel Psychology, is the standard reference for subgroup differences on cognitive measures in employment. It reports a Black-White standardised difference of .99 for measures of g in industrial applicant samples (N = 6,169 across 8 studies) and .41 in incumbent samples, the gap between the two reflecting range restriction from selection that has already happened. It analyses verbal and mathematical ability separately, at .76 for both, and concludes at p. 321 that "tests assessing g were generally associated with larger differences than verbal or mathematical abilities". What it does not do is analyse figural, abstract, spatial or non-verbal measures as a distinct category at all. A paper that never separated out the format cannot be evidence that the format is fairer.

And the evidence that bears most directly on the claim points the other way. Flynn's 1987 study of IQ gains in fourteen nations reports at p. 171 that "some of the largest gains occur on culturally reduced tests and tests of fluid intelligence". At p. 185 he gives the comparison as a rate: a median gain of 0.588 IQ points per year on culturally reduced tests against 0.374 on verbal tests. On Raven's Progressive Matrices — the archetype of the culture-reduced format, and the design the abstract items in this category descend from — the Netherlands gained 0.667 points per year across thirty years and France 1.005 points per year (Table 15, p. 175), and in the three nations where both test types were measured the culture-reduced gains ran at roughly twice the verbal ones (Table 16, p. 186). The argument is short. A score that rose by around twenty points in one country inside a single generation is measuring something that the environment reaches, and reaching it faster than the verbal tests did. Removing the words from an item removes the words. It does not remove the environment.

None of this makes the format a bad instrument, and it should not be read as one. Pattern-extraction items travel across languages without translation, which is a genuine and useful property for a multi-country pipeline, and the Roth paper's own complexity finding — d of .86 at low job complexity, .72 at moderate and .63 at high, the last from only two studies — says that the differences narrow in exactly the analytical roles these tests are usually bought for. What has to go is the inference in the middle: that because a candidate does not need English to attempt the item, the resulting score is equally fair to everyone who sits it.

Fairness in measurement is not a property a format confers. It is a property you test for, item by item, and the test has a name and a published threshold. Differential item functioning asks whether two candidates of the same underlying ability, drawn from different groups, have different probabilities of getting a particular item right. The classification used in the NAEP technical documentation, which follows the scheme ETS made standard, is explicit about where the lines fall: an item is category A when the Mantel-Haenszel common odds ratio on the delta scale does not differ significantly from 0 at alpha = .05 or is less than 1.0 in absolute value; category C when it is "significantly greater than 1 and larger than 1.5 in absolute magnitude"; and everything in between is category B. Items flagged C are "reviewed by a committee of trained test developers and subject-matter specialists to determine whether the differential functioning of a particular item is due to bias or not" — a human judgement, made on a flagged list, not an automatic deletion.

So the question to put to any vendor selling an abstract reasoning test, ours included, is not whether the items contain words. It is: have you run DIF on this bank, against which groups, at what sample size, and how many items came back category C. AssessAll cannot currently answer that question for any of these three forms. No DIF analysis has been run on them, no reliability coefficient is published for them, and no norm group exists — the research-readiness gate in the codebase is what holds the last of those, and it has not been cleared. That is the row on this page that rules us out, and it is on the page because a page arguing that buyers should demand these numbers is worthless if it exempts the vendor publishing it.

Five worked scenarios, with the whole key

Every option below carries its key, the reason it earns that key, and the arithmetic in full. This is the part that is normally invisible: prep sites publish a question and name an answer, and vendors publish neither the working nor what a wrong answer means.

These scenarios were written for this page. They are not taken from the live item bank and are not a practice test. Publishing live items would degrade the instrument for every organisation using it — which is also the answer to give any vendor who offers to show you theirs.

Scenario 1 · Alternating rule — the trap is assuming one operation

A sequence runs: 2, 6, 5, 15, 14, 42, 41, …

What comes next?

OptionKeyWhy it earns that weight
123CorrectKeyed answerCorrect. Two operations alternate: multiply by 3, then subtract 1. The last operation applied was a subtraction, so the next is a multiplication.
40IncorrectDiagnostic distractorChosen by continuing the most recent operation. This is the commonest error on alternating series and it is diagnostic: the candidate found a rule and stopped looking for the second one.
82IncorrectDiagnostic distractorDoubling. The candidate noticed that the numbers roughly double across each pair and fitted the nearest simple operation rather than deriving the exact one.
44IncorrectDiagnostic distractorAdding 3. The candidate read the 3 in the pattern as an addend rather than a multiplier — an arithmetic misread, not a reasoning failure, and worth separating from the other two.

The working. 2 × 3 = 6 · 6 − 1 = 5 · 5 × 3 = 15 · 15 − 1 = 14 · 14 × 3 = 42 · 42 − 1 = 41 · 41 × 3 = 123.

What the item separates. The item separates finding a rule from finding all of a rule. Both of the top two answers require the candidate to have spotted a pattern; only one requires them to have kept testing after the first one fitted.

Scenario 2 · Matrix completion — two independent routes to the same answer

A three-by-three grid reads, by row: 3, 6, 12 / 5, 10, 20 / 7, 14, ?

Which value belongs in the empty cell?

OptionKeyWhy it earns that weight
28CorrectKeyed answerCorrect by either route. Each row doubles from left to right (7 → 14 → 28), and each column advances by a constant that itself doubles: +2 down column one, +4 down column two, +8 down column three, giving 12 → 20 → 28.
21IncorrectDiagnostic distractorTripling the 7. The candidate applied a row rule read from the gap between 3 and 12 rather than from the step between adjacent cells — a scale error inside a correct instinct.
24IncorrectDiagnostic distractorAdding 4, the column-two constant, to 20. The candidate found the column rule but did not notice that the constant itself changes across columns.
26IncorrectDiagnostic distractorAdding 6 to 20, extending the row-one-to-row-two difference. A plausible arithmetic path that never tested itself against the row rule.

The working. Row rule: 7 × 2 = 14, 14 × 2 = 28. Column rule: column three runs 12, 20, 28 with a constant step of 8, which is twice column two's 4 and four times column one's 2. Both routes give 28.

What the item separates. A well-built matrix item is solvable two ways and the two ways agree. That is what makes it worth marking: a candidate who checks their answer against the second route is doing something a candidate who guesses cannot fake.

Scenario 3 · Rule-breaking detection — which group does not belong

Four groups of three numbers: (4, 7, 10) · (6, 11, 16) · (3, 9, 15) · (2, 5, 9).

Which group does not follow the same rule as the others?

OptionKeyWhy it earns that weight
(2, 5, 9)CorrectKeyed answerCorrect. Every other group has a constant step within it — 3, 5 and 6 respectively. This one steps 3 then 4, so no single constant describes it.
(3, 9, 15)IncorrectDiagnostic distractorChosen by candidates looking for a constant step that is the same across groups rather than within each group. The rule is local to the group, and this is the distractor that catches a reader who generalised one level too far.
(6, 11, 16)IncorrectDiagnostic distractorChosen on size rather than structure — the numbers are the largest in the set. This is a surface-feature answer and it is the cheapest one to give.
(4, 7, 10)IncorrectDiagnostic distractorChosen because it is first and looks like the template the others depart from. Position effects are real on this item type, which is why option order is shuffled at delivery.

The working. Within-group steps: 4 → 7 → 10 is +3, +3. 6 → 11 → 16 is +5, +5. 3 → 9 → 15 is +6, +6. 2 → 5 → 9 is +3, +4. Only the last is inconsistent with itself.

What the item separates. The diagnostic distinction here is between a candidate who tests the rule inside each group and one who tests it across groups. Both are reasoning; only one has read the question.

Scenario 4 · Induction — what the evidence supports, and the three ways to overshoot it

A support team records six months of weekly data. In every week in which more than 40 tickets arrived, the average resolution time also exceeded two days.

Which conclusion does that observation support?

OptionKeyWhy it earns that weight
So far, every week with more than 40 tickets has also had an average resolution time above two days.CorrectKeyed answerCorrect, and it is deliberately unexciting. It restates the observation and its scope — six months, this team — without adding a mechanism, a direction or a remedy.
High ticket volume causes slow resolution.IncorrectDiagnostic distractorThe causal leap. Nothing in the data rules out a third factor — a seasonal outage, a staffing gap, a product release — producing both. This is the single most common wrong answer on inductive items and the one most worth knowing about in an analyst.
Every week with resolution above two days had more than 40 tickets.IncorrectDiagnostic distractorThe converse. The observation runs one way only; slow weeks at low volume are entirely consistent with it and were never measured.
Adding staff would bring resolution times below two days.IncorrectDiagnostic distractorAn imported remedy. The candidate has supplied a cause from outside the data and then acted on it — a step further than the causal leap, and the one that costs money.

The working. The observation is a conditional with a sample attached: for the weeks observed, volume above 40 implies resolution above two days. Reversing the conditional, asserting a mechanism, or prescribing an intervention each add a premise the data does not supply.

What the item separates. On a well-built inductive form the keyed answer is often the most cautious one on the list. What the item measures is the discipline of stopping where the evidence stops, which is why a wrong answer here tells you which direction a candidate overshoots in.

Scenario 5 · Rule revision — the pattern changes, and the score is about noticing

A sequence runs: 1, 2, 4, 8, 16, 21, 26, 31, …

What comes next?

OptionKeyWhy it earns that weight
36CorrectKeyed answerCorrect. The sequence doubles up to 16 and then advances by 5, confirmed three times over. The current rule, not the original one, governs the next term.
62IncorrectDiagnostic distractorDoubling 31. The candidate found the first rule, committed to it, and did not revise when three consecutive terms stopped obeying it. This distractor exists to measure exactly that.
41IncorrectDiagnostic distractorAdding 10, on the assumption that the step size is itself growing. A reasonable hypothesis that the data already contradicts: the step has been flat at 5 for three terms.
32IncorrectDiagnostic distractorReverting to the power-of-two series and picking the next member of it, ignoring where the sequence actually is. A near-miss that looks like a small slip and is not.

The working. 1 × 2 = 2 · 2 × 2 = 4 · 4 × 2 = 8 · 8 × 2 = 16 · 16 + 5 = 21 · 21 + 5 = 26 · 26 + 5 = 31 · 31 + 5 = 36.

What the item separates. This item separates rule discovery from rule revision, and they are not the same ability. Candidates who score well on the first four items and choose 62 here are telling you something specific about how they behave when a system they have understood changes underneath them.

What can be changed

  • Choose the form by what you want in the number: Abstract & Pattern Reasoning if pace is genuinely part of the job, Inductive Reasoning if you want the evidence-handling construct with logical reasoning as the second competency, Cognitive Agility if adaptation to a changing rule is the thing you are hiring for.
  • Extend or shorten the window. The three forms sit at 55, 66 and 55 seconds per item; lengthening the window moves the score away from speed and towards accuracy, and it changes what the band means, so say which you did.
  • Mix item media. The forms are text-renderable by design, so a bank can be assembled from symbol series, letter rules, number matrices and short business observations without an image library.
  • Set the reporting competency explicitly if a downstream decision depends on it, rather than accepting the default — all three forms file under Abstract & Pattern Reasoning and the second competency is what differs.
  • Pair with a deductive form if the role needs rule-application as well as rule-extraction; the two dissociate, and one band will not show you both.

What an SJT will not do, including ours

A buyer would find most of this in ten minutes of reading, so it belongs on the page that sells the instrument rather than in a footnote somewhere else.

  • No criterion-validity study has been run on any of these three forms. Any validity figure you read elsewhere on this site or in the literature is about the method, not about this instrument.
  • No reliability coefficient and no standard error of measurement is published for these forms, so the confidence band around a score on them cannot currently be computed by a buyer or by us.
  • No norm group exists for them. A band is a band on this instrument, not a percentile against a national or industry population.
  • No differential item functioning analysis has been run on these banks. The page argues that buyers should ask for one; we cannot yet supply it.
  • AssessAll does not claim these forms are culture-fair, culture-free or free of subgroup differences, and the product copy that came close to claiming it is being corrected.
  • The five items above were written for this page. They are never taken from seeds, from the live bank, or from any form a candidate might sit — publishing live items would degrade the instrument for every client using it.

Questions buyers ask

Is an abstract reasoning test fairer than a verbal or numerical one?

It is less dependent on language, which is a real advantage across a multi-country pipeline. Whether the scores are fairer is a separate and much harder question, and the meta-analysis usually cited for it — Roth et al. (2001) — analysed verbal and mathematical measures separately and did not analyse figural or abstract measures as a category at all. Flynn's 1987 fourteen-nation study points the other way: the largest secular score gains occurred on culturally reduced and fluid-reasoning tests, at roughly twice the rate of verbal tests. Ask the vendor for differential item functioning results instead of accepting the format as an answer.

What is the difference between abstract, inductive and logical reasoning tests?

Abstract reasoning infers an unstated rule from symbols or series. Inductive reasoning does the same from evidence, including verbal and numeric observations, and is keyed to reward refusing a conclusion the evidence does not support. Logical reasoning in the narrow sense is deduction: applying a rule you were given to a case. They are correlated because all three sit under fluid reasoning, and they are not the same score.

How long should an abstract reasoning test be?

Long enough that the number means what you want it to mean. The three AssessAll forms run at 55, 66 and 55 seconds per item; a tighter window moves the score towards speed and a longer one towards accuracy. The published durations in this category range widely — Testlify publishes 12 questions in 15 minutes, which is 75 seconds per item — and the important thing is that you know which end of that range your score came from.

Can candidates practise their way to a higher abstract reasoning score?

Practice effects on this format are well documented, which is part of why the secular gains Flynn measured are so large. This is an argument for supervising the sitting when the score will carry a decision alone, for using a form with a bank large enough that repeat sitters do not see the same items, and for treating a retest score as a different observation rather than a correction of the first one.

What does differential item functioning actually test?

Whether two candidates of the same underlying ability, drawn from different groups, have different probabilities of answering a particular item correctly. The classification used in the NAEP technical documentation flags an item as category C when the Mantel-Haenszel value on the delta scale is significantly greater than 1 and larger than 1.5 in absolute magnitude; those items go to a review committee rather than being deleted automatically. It is the only direct evidence about fairness that a test publisher can produce at item level.

Which form should I use for campus hiring?

Abstract & Pattern Reasoning or Cognitive Agility, on the basis of what the role does. Both are speed-weighted by design, which is defensible for high-volume early-careers screening and should be disclosed to candidates. If the role is analytical and pace is not the point, Inductive Reasoning is the cleaner instrument, and the Roth complexity finding — smaller subgroup differences at higher job complexity — is relevant to that choice.

Sources

Checked at source 24 September 2026.

Related reading

The form described on this page

Abstract & Pattern Reasoning, Inductive Reasoning, and Cognitive Agility & Learning Potential — 22 items (Abstract & Pattern) · 20 (Inductive) · 24 (Cognitive Agility) — 55, 66 and 55 seconds per item respectively, 20 minutes (Abstract & Pattern Reasoning) · 22 minutes (Inductive Reasoning) · 22 minutes (Cognitive Agility & Learning Potential). Right/wrong against a derivable key; no negative marking.

See the listing