Test library
Test library · how it is scored

What is a logical reasoning test, and is it measuring deduction or induction?

A logical reasoning test measures whether someone reaches conclusions the information actually supports. The name covers at least three different abilities: deduction, which applies rules you are given; induction, which infers an unstated rule from examples; and abstract pattern reasoning, which does the same with symbols instead of words. A candidate can be strong at one and weak at another.

SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed

The form we publish, at a glance

Instrument
Logical Reasoning — Graduate, Deductive Logic & Syllogisms, Inductive Reasoning, and Abstract & Pattern Reasoning
Length
25 items (Graduate) · 22 (Deductive) · 20 (Inductive) · 22 (Abstract)
Time
22 minutes (Graduate) · 20 minutes (Deductive) · 22 minutes (Inductive) · 20 minutes (Abstract)
Level
Final-year students and early careers (Graduate, Abstract) · experienced hires and analysts (Deductive, Inductive)
Item format
Multiple choice; language-light on the abstract form; several deductive items keyed to "may or may not follow"
Scoring
Right/wrong against a derivable key; no negative marking
What you get back
Banded score filed under a named competency — Logical Reasoning, Abstract & Pattern Reasoning, or both

What it measures

Logical Reasoning

Whether a conclusion follows from what the candidate was actually given — including the discipline of refusing a conclusion that does not follow, which several items on the deductive form are keyed to reward.

Related competency definition

Abstract & Pattern Reasoning

Extracting an unstated rule from symbols, series and matrices, where the answer depends on structure rather than on vocabulary or schooling. This is the competency the inductive and abstract forms report against.

Related competency definition

Cognitive Speed & Agility (Graduate and Abstract forms only)

How much correct work a person produces inside a fixed window. It is a separate thing from reasoning accuracy, and on those two forms it is deliberately part of the score — which is exactly why the band they report is not a clean reasoning band.

Related competency definition

"Logical reasoning" is at least three different abilities, and the report tells you which one you bought

The taxonomy the measurement field uses has separated these for decades. In the Cattell-Horn-Carroll model, the broad ability Fluid Reasoning contains, among others, two distinct narrow abilities: Induction, "the ability to infer general implicit principles or rules that govern the observed behavior of a phenomenon or the solution to a problem", and General Sequential Reasoning, "the ability to reach logical conclusions from given premises and principles, often in a series of two or more sequential steps". One goes from examples to a rule; the other goes from a rule to a case. They are correlated, because both sit under fluid reasoning, and they are not the same score.

Vendor catalogues have not followed. Checked on 8 September 2026: TestGorilla's cognitive-ability library lists seventeen tests, each with a duration and none with an item count, and it has no separate inductive or deductive test — the two are collapsed into one twelve-minute "Critical Thinking" test whose own description says it measures inductive and deductive reasoning, alongside a separate ten-minute abstract reasoning test. Mercer Mettl's logical reasoning ability test publishes both numbers — 20 questions, 30 minutes, medium difficulty by default — and does not distinguish the two anywhere on the page. SHL's public product catalogue names only SHL Verify as its cognitive product, with no components, item counts or durations published at all. Thomas's employer-facing guide does list diagrams, inductive, deductive, abstract and critical thinking as things a logical reasoning test "can include", and then does not tell a buyer that these are separable scores or which one their role needs.

The consequence is the part worth acting on: a candidate can be strong on one and weak on another, and a single number labelled "logical reasoning" hides which. Someone excellent at applying a written policy exactly as specified may be mediocre at spotting the rule behind a set of examples, and the reverse is at least as common in analyst hiring. If you screen for the first and score the second, the test is working correctly and you are reading the wrong construct.

AssessAll ships the three separately, and the honest way to tell them apart is not the marketing copy — it is the competency tag each form reports against, which is set in the item-bank definition and is what the banded report is filed under. Deductive Logic & Syllogisms (22 items, 20 minutes) carries exactly one competency: Logical Reasoning. Inductive Reasoning (20 items, 22 minutes) carries two, and the first of them is Abstract & Pattern Reasoning — so its score is primarily a pattern-extraction score, whatever the word "inductive" suggests. Abstract & Pattern Reasoning (22 items, 20 minutes) carries Abstract & Pattern Reasoning and Cognitive Speed & Agility. The Graduate form (25 items, 22 minutes) carries Logical Reasoning and Cognitive Speed & Agility, which means its logical band has pace mixed into it by design.

So: if you want a clean logical-reasoning score, use the form whose competency list has one entry on it. If you want a graduate screen where working quickly is genuinely part of the job, the mixed form is the right instrument and you should say so to candidates. And the question to ask any vendor, ours included, is not "is this a logical reasoning test" but "which of the three is it, and what else is inside the number".

Five worked scenarios, with the whole key

Every option below carries its key, the reason it earns that key, and the arithmetic in full. This is the part that is normally invisible: prep sites publish a question and name an answer, and vendors publish neither the working nor what a wrong answer means.

These scenarios were written for this page. They are not taken from the live item bank and are not a practice test. Publishing live items would degrade the instrument for every organisation using it — which is also the answer to give any vendor who offers to show you theirs.

Scenario 1 · Deduction — the invalid form almost everyone accepts

A support policy states: if a ticket is tagged Priority 1, it is escalated to the duty manager. Ticket 4471 was escalated to the duty manager.

What follows about ticket 4471?

OptionKeyWhy it earns that weight
It may or may not have been tagged Priority 1.CorrectKeyed answerCorrect, and the keyed answer is a refusal. The policy says every P1 ticket gets escalated; it does not say escalation happens only for P1 tickets, so an escalation is consistent with the tag and with several other routes.
It was tagged Priority 1.IncorrectDiagnostic distractorAffirming the consequent — the single most commonly accepted invalid inference, and the one the AssessAll graduate bank explicitly builds traps for. It reads a one-way rule as if it ran both ways.
It was not tagged Priority 1.IncorrectDiagnostic distractorThe mirror error. Diagnoses a candidate who has noticed the first option is a trap and over-corrected into the opposite unwarranted claim.
The policy was breached.IncorrectDiagnostic distractorReads a rule that requires an action as one that forbids it in other cases. Diagnoses a reader importing an assumption the text does not carry — the same failure mode a Cannot Say item catches on a verbal form.

The working. Write the rule as P1 → escalated. We are told the consequent, escalated. From "P → Q" and "Q", nothing whatever follows about P. The two valid moves are modus ponens (P → Q with P, giving Q) and modus tollens (P → Q with not-Q, giving not-P). This item offers neither, which is the point.

What the item separates. A form that never keys "may or may not" cannot separate a careful reasoner from a confident one. Ask any vendor whether their deductive items include unwarranted-conclusion keys, because a bank without them rewards decisiveness and calls it logic.

Scenario 2 · Deduction — contrapositive, converse and inverse

An access rule states: every contractor with system access has completed the security module.

Which of these must be true?

OptionKeyWhy it earns that weight
Anyone who has not completed the security module does not have contractor system access.CorrectKeyed answerCorrect. This is the contrapositive, and it is the only one of the three rewrites that is guaranteed to be true whenever the original is.
Anyone who has completed the security module has contractor system access.IncorrectDiagnostic distractorThe converse. True only if the rule also runs the other way, which it does not say. This is the same error as item 1 in a different grammatical dress — which is why it is worth having both on a form.
Anyone without contractor system access has not completed the security module.IncorrectDiagnostic distractorThe inverse. Equivalent to the converse, and picked by candidates who negate both halves and assume that preserves the rule.
Contractors complete the security module before access is granted.IncorrectDiagnostic distractorA claim about sequence that the rule does not make. Diagnoses a candidate reasoning from how the process probably works rather than from what was stated — the most job-relevant error on the item, because that is exactly how policies get misapplied at work.

The working. The rule is: contractor-with-access → completed-module. Contrapositive: not-completed → not-contractor-with-access. Valid, always. Converse: completed → contractor-with-access. Invalid. Inverse: not-contractor-with-access → not-completed. Invalid, and logically equivalent to the converse.

What the item separates. Three of the four options are the same rule read in a direction it does not license. That is what a well-built deductive distractor set costs to write, and it is why the option a candidate picked is more informative than the fact that they were wrong.

Scenario 3 · Induction — where "nothing follows" becomes the wrong answer

A team observes that every support ticket resolved in under an hour last month came from a customer on the Enterprise plan.

Which is the best-supported conclusion?

OptionKeyWhy it earns that weight
Enterprise tickets appear to be resolved faster, and the claim is worth testing against the tickets that took longer.CorrectKeyed answerCorrect. Induction produces the best-supported generalisation plus the test that would break it — not a certainty. Naming the test is what separates a hypothesis from a hunch.
Being on the Enterprise plan causes faster resolution.IncorrectDiagnostic distractorThe causal leap. Enterprise customers may raise simpler tickets, be routed to senior agents, or sit in a friendlier timezone; the observation cannot choose between those.
All Enterprise tickets are resolved in under an hour.IncorrectDiagnostic distractorThe observation inverted. It said fast implies Enterprise, not Enterprise implies fast — the item-1 error reappearing inside an inductive question, which is why candidates who have only practised deduction still miss it.
Nothing can be concluded from a single month of data.IncorrectDiagnostic distractorThe most interesting wrong answer on this page. It applies the deductive standard to an inductive question. Induction never yields certainty, so a rule that refuses every uncertain conclusion refuses all of them — including the ones a business is supposed to act on.

The working. The observation has the form: resolved-under-1h → Enterprise. Option C reverses it. Option B adds a mechanism the data does not contain. Option D demands entailment from evidence that is, by construction, not entailing. Option A states the direction actually observed, marks it as provisional, and names the sample that would falsify it — the tickets that took longer, which were never examined.

What the item separates. Deduction and induction disagree about what a good answer looks like: on item 1 the key is a refusal to conclude, and here the same instinct is the distractor. A candidate applying the wrong standard fails an item they fully understand — which is the clearest practical proof that these are two scores, not one.

Scenario 4 · Abstract pattern — the rule is arithmetic, not perception

A sequence of symbol groups runs: ■□ , ■□□ , ■■□□ , ■■□□□ , ?

Which group comes next?

OptionKeyWhy it earns that weight
■■■□□□CorrectKeyed answerCorrect. Written as (filled, open) the sequence is (1,1), (1,2), (2,2), (2,3) — each step increases exactly one count, alternating, and the last step raised the open count, so this one raises the filled count to give (3,3).
■■□□□□IncorrectDiagnostic distractorRepeats the step just taken instead of alternating. Diagnoses a candidate who inferred the rule from the last transition only — a local rule that fits one step and none of the others.
■■■□□□□IncorrectDiagnostic distractorApplies both increments at once, giving (3,4). Total length would jump by two where every previous step added one. Diagnoses averaging the two rules rather than alternating them.
■□■□■□IncorrectDiagnostic distractorReads the groups as an alternating pattern of symbols rather than as counts of them. Diagnoses a candidate working from the visual surface, which is the failure mode a language-light item is specifically built to expose.

The working. Convert to counts: (1,1), (1,2), (2,2), (2,3). Differences alternate +open, +filled, +open — so the next is +filled, giving (3,3), written ■■■□□□. Total length runs 2, 3, 4, 5, so the next group has six symbols; any option of length seven is eliminated before the rule is even applied.

What the item separates. The check that settles this item is arithmetic on the counts, and length alone eliminates one distractor. Language-light is not the same as culture-free or preparation-proof, and no honest abstract form should be sold as either.

Scenario 5 · Process logic — the format most catalogues call "diagrammatic"

A refund flow. Step 1: if the order is less than 30 days old, go to step 2; otherwise decline. Step 2: if the item is unopened, approve; otherwise go to step 3. Step 3: if the customer has had fewer than two refunds this year, approve as a goodwill credit; otherwise escalate to a supervisor. The case: an order placed 12 days ago, item opened, and this is the customer's third refund this year.

What does the flow return?

OptionKeyWhy it earns that weight
Escalate to a supervisor.CorrectKeyed answerCorrect. 12 is less than 30, so step 2. The item is opened, so step 3. Three refunds is not fewer than two, so escalate.
Approve as a goodwill credit.IncorrectDiagnostic distractorThe boundary error, and the most valuable wrong answer here: it reads "fewer than two" as "two or fewer", or simply stops reading at the friendliest branch. In a live process this is the error that costs money quietly.
Decline.IncorrectDiagnostic distractorExits at step 1 despite the order being inside the window. Diagnoses a candidate who recognised that the case looks unfavourable and jumped to the outcome rather than tracing the branch.
Approve.IncorrectDiagnostic distractorApproves at step 2 despite the item being opened. Diagnoses a skipped condition rather than a misread one — a different remediation from the boundary error above.

The working. Step 1: 12 < 30, so continue. Step 2: opened, so continue. Step 3: is 3 fewer than 2? No, so escalate. The strict inequality is the whole item: a customer on their second refund also escalates, because two is not fewer than two.

What the item separates. This format is sold across the market as diagrammatic or flowchart reasoning, and it is deduction — the rules are handed to you and nothing has to be inferred. Every error it catches is a branch or boundary error, which is precisely the competency behind SOP compliance and no-code automation, and not the same thing as spotting a pattern.

What can be changed

  • Choose the construct before the form. One competency on the report means one construct in the score; two means the second one is in there too, whether or not you wanted it.
  • Form length and window are properties of the form you pick, not a slider — 20 items in 22 minutes reads very differently from 20 items in 10 minutes, and the Logical Sprint form states in its own item-bank note that it deliberately contains more questions than most candidates finish.
  • Option order can be shuffled per sitting, so two candidates in the same room do not see the same answer in the same place.
  • Combine a deductive form with an inductive or abstract one when the role needs both, and read the two bands separately rather than averaging them — the whole argument of this page is that the average of two constructs is not a construct.
  • Set the pass mark by a standard-setting method rather than by a round number, and check it against the standard error before you use it to reject anyone.

What an SJT will not do, including ours

A buyer would find most of this in ten minutes of reading, so it belongs on the page that sells the instrument rather than in a footnote somewhere else.

  • No criterion-validity study has been run on any AssessAll logical, inductive or abstract form. Every validity figure quoted anywhere on this page is about the METHOD as studied in the published literature, never about a coefficient for one of our forms — and no AssessAll instrument has yet met our internal gate for publishing a norm or a benchmark, so there is no AssessAll percentile that means anything outside the pipeline it was computed in.
  • A reasoning test is not the best single predictor of job performance, and the number the market still quotes for it is out of date. Schmidt and Hunter's widely cited operational validity of about .51 for general mental ability was revised in 2022 by Sackett, Zhang, Berry and Lievens in the Journal of Applied Psychology, who showed that applying range-restriction corrections from predictive studies to concurrent ones systematically inflated the estimates. Their corrected figure for general mental ability is about .31 — below structured interviews at about .42, job knowledge tests at about .40, and empirically keyed biodata at about .38. A reasoning test is a cheap, fast, defensible component of a selection process. It is not the centre of one.
  • Cognitive tests carry the largest subgroup differences of any common selection method, and a language-light format does not remove them. Roth and colleagues' 2001 meta-analysis in Personnel Psychology reports a Black-White standardised difference of about 0.99 across industrial samples on measures of general cognitive ability, larger than for verbal or mathematical ability specifically (both about 0.76), and smaller for higher-complexity jobs (about 0.63). That analysis did not separately examine figural or abstract measures, which means the claim commonly made for abstract tests — including in our own item-bank note for the abstract form, which calls it the fairest test across backgrounds — is not established by the evidence usually cited for it. Monitor your own selection rates rather than trusting a format.
  • Logical reasoning is among the most coachable formats in the market: series, syllogism and code items have fixed structures and a large free practice industry teaching them. A coached candidate and an uncoached one are not on the same starting line, and no vendor can honestly claim otherwise.
  • An unproctored, multiple-choice reasoning test taken at home in 2026 is not a secure measure of anything, and the specific limit here is worth stating. AssessAll's AI-answer detector reads free-text responses only, with a minimum length before it will look at all — so on an all-multiple-choice reasoning form it sees nothing. What remains are behavioural and device signals: tab switching, blocked paste, blocked print-screen, an additional face in frame, a suspected-automation device signal, and a too-fast completion flag, each carrying a published weight into an integrity band whose own code comments call those weights investigative rather than verdicts. Treat the result as a lead to follow, not as proof, and use a supervised sitting where the decision warrants one.

Questions buyers ask

Is a logical reasoning test the same as an IQ test?

No, though they overlap. An IQ test aims at a broad estimate of general cognitive ability across several domains and reports it against a standardisation sample. A logical reasoning test measures one slice — usually deduction, induction or pattern extraction — and reports against whatever comparison group the vendor has. In the Cattell-Horn-Carroll framework the abilities these items tap sit under Fluid Reasoning, which is one broad ability among several, so a reasoning score is a narrower and more job-anchored claim than an IQ score, and it should be described as one.

Which of the three should we use for which kind of role?

Match the construct to the work. Deduction is the fit where the job is applying rules exactly as written — compliance, claims, underwriting, operations, anything governed by an SOP or a policy document, and it is the same competency the flowchart and process items measure. Induction and abstract pattern reasoning fit work where the rule is not given and has to be found: analytics, research, quality investigation, early-career technical roles where trainability matters more than current knowledge. A graduate screen that must cover both usually runs two short forms rather than one long mixed one, because two bands can be read and an average cannot.

Can candidates prepare for a logical reasoning test?

Yes, and more effectively than for most instruments. The item structures are finite and publicly documented, and an entire free practice industry teaches them. Practice reliably lifts early scores, mostly by removing format unfamiliarity, which means a coached and an uncoached candidate are not comparable on a first sitting. Two things reduce the effect: an item bank large enough that a form is not memorable, and reporting attempted alongside correct so a fast guesser is visible. Neither eliminates it, and any vendor claiming a practice-proof reasoning test is claiming something the format cannot deliver.

Can a candidate just use an AI assistant to answer these at home?

On an unproctored sitting, assume yes for text-based deductive and series items — current general-purpose models handle them reliably. Be precise about what a platform can and cannot see. AssessAll's AI-generated-answer review reads free-text responses only and ignores anything shorter than a threshold, so on an all-multiple-choice reasoning form it contributes nothing at all. Detection rests on behavioural and device signals — tab switches, blocked paste and print-screen, an extra face in frame, suspected automation, and a too-fast completion flag — which roll into a banded integrity score that is designed as an investigative lead rather than a verdict, and never auto-rejects. If a reasoning score is going to carry a hiring decision on its own, supervise the sitting.

What pass mark should we set on a reasoning test?

Not one borrowed from another organisation, and not a round number. A defensible cut score comes from a standard-setting study in which people who know the job judge, item by item, how a borderline-competent candidate would perform. Then check the cut against the standard error of measurement: if the error band around a score is wider than the distance from your cut, the test cannot reliably separate the people either side of it, and the cut is a decision rule you cannot defend. Our Angoff cut-score helper runs both numbers.

How many sittings before a pass rate on this test means anything?

More than most teams assume, and the answer is different for a group than for a person. A pass rate observed on a few dozen sittings carries a confidence interval wide enough to cover most of the decisions you would want to make with it. The error band on one candidate's score does not shrink at all as the cohort grows — it is a property of the instrument, not of the sample. The sample-size and reliability calculator works both numbers for a cohort you specify.

Sources

Checked at source 8 September 2026.

Related reading

What is a numerical reasoning test?

The same graduate screen's other half — and why a speeded form and a power form of one construct tell you different things.

What is a verbal reasoning test?

Another name covering two instruments, and a Cannot Say key that is the verbal cousin of item 1 on this page.

What is a situational judgement test?

A method rather than a construct, with a graded key instead of a right answer.

Angoff cut-score calculator

Where a defensible pass mark on a reasoning test comes from, and the two errors that decide whether it can be used.

Assessment sample-size & reliability calculator

How many sittings a pass rate needs, and why one candidate's error band never shrinks.

Adverse impact ratio calculator

The four-fifths check that matters more on cognitive measures than on any other method.

Standard error of measurement

The band around a single reasoning score, and what it does to a cut.

Norm group

What a percentile on a reasoning test is a percentile of.

What is a psychometric assessment?

Where ability tests sit among the other instrument families, and what each one is for.

How to read an assessment report

The three lines to check before a banded reasoning score means anything.

Numerical reasoning test

Speeded form or power form, five worked items with the working, and what each wrong answer diagnoses.

Verbal reasoning test

Reasoning from text or English proficiency — two tests under one label, five worked True/False/Cannot Say items, and why a lower cut score is the wrong fix.

The form described on this page

Logical Reasoning — Graduate, Deductive Logic & Syllogisms, Inductive Reasoning, and Abstract & Pattern Reasoning25 items (Graduate) · 22 (Deductive) · 20 (Inductive) · 22 (Abstract), 22 minutes (Graduate) · 20 minutes (Deductive) · 22 minutes (Inductive) · 20 minutes (Abstract). Right/wrong against a derivable key; no negative marking.

See the listing