Test library
Test library · how it is scored

What is a verbal reasoning test, and is it measuring reasoning or English?

A verbal reasoning test measures how accurately someone draws conclusions that a written passage actually supports — usually by judging statements True, False or Cannot Say. It is not an English test: a language-proficiency screen measures vocabulary, grammar and error spotting, and the two are sold under the same name while measuring different things.

SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed

The form we publish, at a glance

Instrument
Verbal Reasoning — Professional (True/False/Cannot Say) and Verbal Ability — Foundation
Length
20 items (Professional) · 25 items (Foundation)
Time
22 minutes (Professional, reasoning) · 18 minutes (Foundation, language)
Level
Experienced hires and graduates · entry level and high-volume screening
Item format
True / False / Cannot Say and argument evaluation · multiple choice on vocabulary, grammar, error spotting and sentence ordering
Scoring
Right/wrong against a keyed answer; no negative marking
What you get back
Banded score with the competency profile, and the sentence that licenses each key

What it measures

Verbal Reasoning & Comprehension

Whether a person's conclusion is licensed by the text in front of them rather than by what they already believe about the subject. The entire discipline of the format is holding to what the passage says, including when the passage is silent.

Related competency definition

Critical Reasoning (Professional form)

Whether a person can separate what strengthens or weakens a claim from what merely sounds right — a causal reading of a self-selected comparison, a biased sample, survivorship, an unstated assumption, a base rate left out.

Related competency definition

Cognitive Speed & Agility (Foundation form)

How much correct language work a person produces inside a fixed window. On the 25-item Foundation form this is deliberately part of the score; on the reasoning form it is not the point.

Two different tests share this name — and the usual fix for the language problem is the one thing US law names

Two quite different instruments are marketed as verbal reasoning. One is a language-proficiency screen: vocabulary in context, subject-verb agreement, prepositions and phrasal verbs, error spotting, sentence ordering. The other is a reasoning-from-text measure: you are given a business passage and asked whether a statement is True, False, or Cannot Say on the strength of that passage alone. Both produce a number on a 0–100 scale labelled "verbal", and the two numbers are not measuring the same thing.

AssessAll ships both, and the difference is visible in the configuration rather than in the marketing. Verbal Ability — Foundation is 25 items in 18 minutes, speed-weighted, and its methodology note describes it as calibrated to workplace English rather than literary English; its named second competency is Cognitive Speed. Verbal Reasoning — Professional is 20 items in 22 minutes of True/False/Cannot Say and argument-evaluation items, and its named second competency is Critical Reasoning. Neither form is better. They answer different questions, and a score from one is not a score from the other.

The consequence lands on fairness, and this is where the lane goes wrong. A language screen loads heavily on English proficiency by design. A reasoning-from-text form loads on it too, but less, because the passage supplies the vocabulary and the work is inference rather than idiom. At least one employer-facing guide currently ranking for this term acknowledges that non-native speakers are disadvantaged and recommends the obvious-looking remedy: lower and more flexible cut-off scores for those candidates. In the United States that specific remedy is prohibited by statute.

Section 106 of the Civil Rights Act of 1991 (Public Law 102–166, 21 November 1991) added §703(l) to Title VII, now 42 U.S.C. §2000e-2(l): "It shall be an unlawful employment practice for a respondent, in connection with the selection or referral of applicants or candidates for employment or promotion, to adjust the scores of, use different cutoff scores for, or otherwise alter the results of, employment related tests on the basis of race, color, religion, sex, or national origin." The concern behind the advice is real and the advice is well meant. The mechanism it reaches for is the one the statute names.

The remedy that is available is a construct decision made before anyone sits the test. Ask what level of English the job genuinely requires and at what level the work is performed, then choose the form that measures that — and where the role needs reasoning rather than idiom, use the reasoning form, on one cut score applied to everyone. After that, monitor selection rates by group with the four-fifths rule as an early indicator, and treat a flag as a reason to re-examine the instrument, not as a reason to move the bar for some candidates. Changing the test is available to you; changing the score is what the statute forbids. That is the US rule, other jurisdictions handle score adjustment differently, and none of this is legal advice.

The limit on our side of that argument: AssessAll publishes no differential-item-functioning study, no subgroup analysis and no local validation study on either verbal form, so we cannot tell you how large the language effect is on ours. What we can tell you is which of the two things each form is measuring, which is the part that decides the question you started with.

Five worked scenarios, with the whole key

Every option below carries its key, the reason it earns that key, and the arithmetic in full. This is the part that is normally invisible: prep sites publish a question and name an answer, and vendors publish neither the working nor what a wrong answer means.

These scenarios were written for this page. They are not taken from the live item bank and are not a practice test. Publishing live items would degrade the instrument for every organisation using it — which is also the answer to give any vendor who offers to show you theirs.

Scenario 1 · A conditional is not an equivalence

Passage: "Following the redesign, the support desk resolved 72% of tickets at first contact, up from 64%. Every team that adopted the new triage script saw its first-contact rate rise. The regional team did not adopt the script."

Statement: "The regional team's first-contact resolution rate did not rise." True, False, or Cannot Say?

OptionKeyWhy it earns that weight
Cannot SayCorrectKeyed answerCorrect. The passage tells you what happened to every team that adopted the script. It says nothing at all about teams that did not, and a rate can rise for other reasons.
TrueIncorrectDiagnostic distractorThe classic inversion: "every team that adopted it improved" read as "only teams that adopted it improved". Diagnoses a candidate who converts a conditional into an equivalence — the habit that turns an association in a report into a causal recommendation in a meeting.
FalseIncorrectDiagnostic distractorTreats "we do not know" as "the statement is wrong". False is available only when the passage supports the negation, and here the passage supports nothing either way.

The working. The passage licenses adopted → rose. It licenses nothing about did not adopt. Both True and False require a fact the passage does not contain, which is precisely the situation Cannot Say exists for.

What the item separates. Separates candidates who read a conditional as a conditional from candidates who read it as an if-and-only-if. The second group is the one who will report that the script caused the improvement.

Scenario 2 · The plausible inference the reader brought with them

Passage: "The Chennai office opened in March and reached full headcount in July. Attrition across the company was 14% for the year, its lowest since 2021."

Statement: "Attrition at the Chennai office was below 14%." True, False, or Cannot Say?

OptionKeyWhy it earns that weight
Cannot SayCorrectKeyed answerCorrect. A company-wide figure constrains an average, not any one site. Chennai could sit well above 14% and be outweighed by larger, more stable offices.
TrueIncorrectDiagnostic distractorThe most common wrong answer here, and the one worth noticing, because it is not a stupid answer. It comes from a sensible real-world inference — a new office, everyone recently hired, so few leavers — which is knowledge the reader supplied rather than something the passage said.
FalseIncorrectDiagnostic distractorThe same error mirrored: reasoning that a new office must churn. Both wrong answers are the reader's own experience arriving in place of the text.

The working. 14% is an aggregate and nothing in the passage decomposes it by site, so no statement about a single site is decidable from it. Note that the wrong answer is a good guess; that is what makes the item diagnostic rather than merely hard.

What the item separates. This item separates knowing a thing from the passage having said it. The candidates who fail it are often the well-informed ones, which is why the error survives comfortably into senior roles.

Scenario 3 · Under-reading — refusing an entailment the passage does license

Passage: "The assessment was sat by 240 candidates. 90 scored above the cut. The report notes that no candidate who scored above the cut withdrew before interview."

Statement: "At least 150 candidates did not score above the cut." True, False, or Cannot Say?

OptionKeyWhy it earns that weight
TrueCorrectKeyed answerCorrect, and it takes one subtraction: 240 − 90 = 150 did not score above the cut, so "at least 150" holds exactly. An arithmetic step the passage licenses is still something the passage says.
Cannot SayIncorrectDiagnostic distractorThe under-reading error, and the one this format's own coaching produces. Candidates taught to "use only what is written" begin refusing valid entailments because the number does not appear as a numeral. It costs the same mark as over-reading and means the opposite thing about the person.
FalseIncorrectDiagnostic distractorUsually a misreading of "at least" as "exactly more than". The statement is satisfied by exactly 150, and 150 is what the passage gives you.

The working. 240 sat, 90 above the cut, so 240 − 90 = 150 at or below it, and "at least 150" is satisfied by exactly 150. The third sentence, about withdrawals, is irrelevant to the statement and is there on purpose: a passage containing only relevant sentences is not testing selection.

What the item separates. The most useful item on any True/False/Cannot Say form, because it catches the error the format's own reputation creates. Over-reading and under-reading both lose a mark and they are opposite problems — one candidate will act on a conclusion the data does not support, the other will refuse to draw one it does.

Scenario 4 · Checking the quantifier

Passage: "The programme ran in four regions. It improved average handling time in three of them. In the fourth, average handling time was unchanged."

Statement: "The programme improved average handling time in every region where it ran." True, False, or Cannot Say?

OptionKeyWhy it earns that weight
FalseCorrectKeyed answerCorrect. The passage names a region where handling time did not improve, so the universal claim is contradicted. False is available precisely because the passage supports the negation, explicitly.
Cannot SayIncorrectDiagnostic distractorThe over-cautious answer, and the most common wrong answer on items whose key is False. Candidates who have learned that Cannot Say is often right start choosing it as a hedge, which is a strategy rather than a reading.
TrueIncorrectDiagnostic distractorReads "three of four" as "the programme worked". Diagnoses summarising rather than checking the quantifier — the same habit that turns a mixed result into a green status in a steering pack.

The working. "Every region" is a universal claim and one counterexample in the passage falsifies it. Note that "unchanged" is not "worse": the item is False not because the programme harmed anyone but because "every" is a stronger claim than "three of four".

What the item separates. Tests whether a candidate checks quantifiers. Most reporting errors inside organisations are not arithmetic; they are a "most" promoted to an "all" somewhere between the data and the slide.

Scenario 5 · A comparison between groups that chose themselves

Passage: "Teams using the new triage script have a first-contact resolution rate 12 points higher than teams that do not. Adoption of the script was voluntary."

Statement: "Adopting the triage script raises first-contact resolution." True, False, or Cannot Say?

OptionKeyWhy it earns that weight
Cannot SayCorrectKeyed answerCorrect, and the word that decides it is "voluntary". Teams chose to adopt, and the teams that chose are plausibly the ones already better at triage. The passage reports an association and then hands you the reason it cannot be read as an effect.
TrueIncorrectDiagnostic distractorThe costliest wrong answer in this set, because outside a test it is a purchase decision. Diagnoses a reader who converts a comparison between self-selected groups into a causal claim — which is the shape most vendor case studies take, including ones written about us.
FalseIncorrectDiagnostic distractorOver-corrects: "this does not show the script works" read as "the script does not work". It may well work. The passage does not establish it, and absence of evidence is not evidence of absence.

The working. A 12-point gap between adopters and non-adopters is consistent with the script causing the gap, with better teams adopting the script, and with both at once. The single word "voluntary" is what makes the second reading unexcludable, and it is exactly the kind of word a fast reader skims.

What the item separates. The item worth keeping when you shorten a form. The person who reads a self-selected comparison as an effect is the person who will recommend a tool on the strength of its own case study.

What can be changed

  • Pick the construct before the length. If the bottleneck in the role is reading and writing accurate workplace English, the Foundation form measures that in 18 minutes. If the bottleneck is drawing conclusions someone will act on, the Professional form is the one, and the two are not interchangeable.
  • A third form runs the same construct in interactive item types — reordering jumbled paragraphs, reconstructing event chains, and judging the most and least effective phrasing for a workplace message, 20 items in about 19 minutes. It exists for teams who want a format candidates cannot have drilled in a multiple-choice prep book.
  • Passage domain can be matched to the work — operations, finance, policy, client correspondence — without changing the construct, because the reasoning demand lives in the question stem rather than in the subject matter.
  • Timing on the reasoning form can be relaxed. Doing so changes what the score means, so record that you did it: comparing a relaxed sitting against a standard one is not a valid comparison.
  • Report items attempted alongside items correct. On any timed verbal form that is the difference between a slow accurate reader and a fast careless one, and a single percentage hides both of them.
  • Run it as one component of a battery rather than as the screen. Verbal reasoning measures one thing well and it is not the measure that generalises furthest across roles.

What an SJT will not do, including ours

A buyer would find most of this in ten minutes of reading, so it belongs on the page that sells the instrument rather than in a footnote somewhere else.

  • There is no AssessAll norm group, benchmark or published percentile for either verbal form. The internal gate for publishing one — 400 completed sittings on a single form, from at least ten organisations, none contributing more than a quarter, with every reported cut backed by at least 50 — has not been passed on any form on the platform. A percentile in a report is a cohort percentile: the share of scored candidates in that same pipeline at or below this one.
  • No differential-item-functioning study, no subgroup analysis and no local criterion-validation study has been run on either verbal form. We therefore cannot quantify the language effect this page tells you to think about, and we would rather say so than estimate it.
  • Published validity figures for verbal ability belong to the method and to the studies that produced them. They are not evidence about an AssessAll form and this page does not present them as such.
  • A verbal reasoning score is not a measure of intelligence, of writing quality, or of whether someone will be good at the job. It measures whether conclusions drawn from a passage are conclusions the passage supports.
  • Practice improves scores on this format, particularly across the first few sittings, mostly by removing unfamiliarity with the Cannot Say convention. That is a real limitation rather than a marketing point, and it is one more reason to report attempted alongside correct.
  • Nothing here is legal advice. The statute quoted above is United States federal law; other jurisdictions treat score adjustment differently, and AssessAll holds no accreditation or certification in any jurisdiction.
  • The five items below were written for this page. They are not taken from the live item bank, they are not a practice test, and no live item is published anywhere on this site.

Questions buyers ask

Is a verbal reasoning test an English test?

That depends entirely on which of the two instruments you bought, which is the problem. A language-proficiency screen tests vocabulary, grammar, prepositions and error spotting, and it is an English test. A True/False/Cannot Say reasoning form gives you the vocabulary in the passage and tests whether your conclusion is licensed by it. Both are labelled "verbal". Ask the vendor which one their score is from before you compare it to anyone else's.

What does a "Cannot Say" answer actually mean?

It means the passage is genuinely silent — the statement is neither supported nor contradicted by what is written, so a competent reader would decline to judge it. It does not mean "probably not" and it is not a safe hedge. On our Professional form the keys are audited so that Cannot Say is used only where the passage really does not decide the question, which is what makes a wrong Cannot Say diagnostic rather than unlucky.

Our candidates are not native English speakers. Should we lower the cut score for them?

No, and in the United States it is unlawful: 42 U.S.C. §2000e-2(l), added by section 106 of the Civil Rights Act of 1991, makes it an unlawful employment practice to use different cutoff scores for employment-related tests on the basis of national origin, among other bases. The concern is legitimate and the fix is upstream. Decide what level of English the role genuinely requires, choose the form that measures the construct you actually need, apply one cut score to everyone, and monitor selection rates by group so that a problem shows up as a flag on the instrument rather than as a private adjustment. This is not legal advice.

Which form should we use — the language one or the reasoning one?

Ask what a failure at work would look like. If the failures you are trying to prevent are ambiguous emails, missed conditions in an SOP and grammar that costs the company credibility with clients, that is the language form. If the failures are conclusions drawn from a report that the report does not support, that is the reasoning form. If the honest answer is both, run both as short components rather than picking one and reading it as though it covered the other.

What score should we set as the pass mark?

Not one borrowed from another organisation and not a round number that looks reasonable. A defensible cut comes from a standard-setting method — the Angoff method is the most common — in which people who know the job judge, item by item, how a borderline-competent candidate would perform. Then check the cut against the standard error of measurement: if the SEM is 4 points and your cut is 60, candidates between 56 and 64 are not reliably distinguishable from one another, and a shortlist drawn on that boundary implies a precision the instrument does not have.

Is a wrong Cannot Say the same as a wrong True?

It costs the same mark and it means the opposite thing, which is why a total score alone wastes the most useful information a verbal form produces. A candidate who answers True where the key is Cannot Say is over-reading: importing world knowledge, converting a conditional, treating an association as an effect. A candidate who answers Cannot Say where the key is True is under-reading: refusing an entailment the passage supports. The first will act on conclusions the evidence does not carry; the second will not act on conclusions it does. Ask for the pattern, not only the percentage.

Sources

Checked at source 5 September 2026.

  • 42 U.S.C. §2000e-2(l) — prohibition of discriminatory use of test scores

    Added by section 106 of the Civil Rights Act of 1991, Public Law 102–166, 21 November 1991 (105 Stat. 1074–1076). Makes it unlawful to adjust scores, use different cutoff scores, or otherwise alter the results of employment-related tests on the basis of race, colour, religion, sex or national origin. Quoted verbatim above.

  • AssessAll item-bank methodology notes (seeds/aptitude-catalogue/)

    The Professional form's note states that keys are audited so that Cannot Say is used only where the passage is genuinely silent, at 20 items in 22 minutes; the Foundation form's note describes 25 items in 18 minutes, speed-weighted and calibrated to workplace English. Both state no affiliation with any third-party instrument.

  • AssessCandidates — verbal reasoning tests for recruitment

    Employer-facing guide ranking for this term; read 5 September 2026. It acknowledges that non-native speakers may struggle and recommends lower and more flexible cut-off scores, which is the recommendation the section above addresses. Recorded because the reader should be able to check that the advice exists rather than take our word for it.

  • Mercer | Mettl — verbal reasoning test

    Checked 5 September 2026: publishes duration (45 minutes) and item count (24 questions) for its verbal reasoning test, and does not publish a reliability coefficient, a standard error, a norm-group description, cut-score guidance or sample items. Recorded as a fact about what is published, which is the part of any comparison a reader can re-verify.

  • Standards for Educational and Psychological Testing (AERA, APA, NCME)

    The standards governing what a test score may be claimed to mean, including validity evidence, fairness in testing, and the justification required for a cut score.

  • US Office of Personnel Management — assessment and selection

    Public-sector reference on choosing assessment methods, the evidence each one requires, and the trade-offs between them.

Related reading

The form described on this page

Verbal Reasoning — Professional (True/False/Cannot Say) and Verbal Ability — Foundation20 items (Professional) · 25 items (Foundation), 22 minutes (Professional, reasoning) · 18 minutes (Foundation, language). Right/wrong against a keyed answer; no negative marking.

See the listing