Test library
Test library · how it is scored

What is a numerical reasoning test, and what does a score actually mean?

A numerical reasoning test measures how well someone draws correct conclusions from figures — percentages, ratios, rates and data tables — rather than how much mathematics they remember. Items are multiple-choice against a computed key, and the distractors are usually the results of specific miscalculations, so a wrong answer identifies the method that produced it.

SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed

The form we publish, at a glance

Instrument
Numerical Reasoning — Graduate and Numerical Reasoning — Professional
Length
25 items (Graduate) · 20 items (Professional)
Time
22 minutes (Graduate) · 25 minutes (Professional)
Level
Final-year students and early careers · experienced hires
Item format
Multiple choice, four options, every distractor a named wrong method
Scoring
Right/wrong against a computed key; no negative marking
What you get back
Banded score with the competency profile, and the worked method per item

What it measures

Numerical Reasoning

Whether a person selects the right operation for the question asked — which base a percentage is taken on, whether growth compounds, whether an average needs weighting — and executes it without arithmetic slips.

Related competency definition

Cognitive Speed & Agility (Graduate form only)

How much correct work a person produces inside a fixed window. This is a separate thing from accuracy, and on the Graduate form it is deliberately part of the score.

Related competency definition

Speed and accuracy are two different scores, and most tests report one number for both

A test is speeded when the time limit is tight enough that a meaningful share of candidates do not reach the end, and a power test when nearly everyone finishes and difficulty does the separating. Almost every published numerical reasoning test is speeded to some degree, and almost none of them say so on the results page — so an employer reads a low score as weak numeracy when it may be an accurate, deliberate candidate who ran out of clock.

The distinction is a configuration choice, not a quality difference, and it has a consequence you can act on. AssessAll ships both: the Graduate form is 25 items in 22 minutes and its own methodology note describes it as speed-calibrated, with Cognitive Speed as a named second competency. The Professional form is 20 items in 25 minutes and is described in the same notes as a power format that rewards method over speed. Same construct, deliberately different questions.

So the question to ask of any numerical reasoning score — ours or anyone's — is how many items the candidate attempted, not only how many they got right. A candidate who answered 14 of 25 and got 13 right is telling you something quite different from one who answered all 25 and got 13 right, and a single percentage hides both of them. If a vendor cannot give you attempted alongside correct, they cannot tell you which of those two people you are looking at.

The second thing worth knowing is that a wrong answer on a well-built form is not noise. The distractors on both AssessAll forms are anchored to documented miscalculation patterns — simple-versus-compound growth, adding successive percentages, discount-base errors, the arithmetic mean of two speeds — so the option a candidate picked names the method they used. That is the part of a numerical test that is genuinely diagnostic, and it is thrown away by any report that shows only a total.

Five worked scenarios, with the whole key

Every option below carries its key, the reason it earns that key, and the arithmetic in full. This is the part that is normally invisible: prep sites publish a question and name an answer, and vendors publish neither the working nor what a wrong answer means.

These scenarios were written for this page. They are not taken from the live item bank and are not a practice test. Publishing live items would degrade the instrument for every organisation using it — which is also the answer to give any vendor who offers to show you theirs.

Scenario 1 · Percentage points versus relative change

A support team's first-contact resolution rate rose from 64% to 72% over a quarter.

By how much did first-contact resolution improve?

OptionKeyWhy it earns that weight
8 percentage points, which is a 12.5% relative increaseCorrectKeyed answerCorrect, and it names both quantities. These are different numbers describing the same movement, and a report that gives only one of them is ambiguous.
8%IncorrectDiagnostic distractorThe most common error in business reporting: a difference in percentage points reported as a percentage change. It is out by more than a third here, and the gap widens as the base falls.
11.1%IncorrectDiagnostic distractorRight method, wrong base — 8 ÷ 72 instead of 8 ÷ 64. Diagnoses someone dividing by the new value rather than the starting value.
112.5%IncorrectDiagnostic distractorThe ratio of new to old read as the increase. Diagnoses a candidate who computed 72 ÷ 64 correctly and then did not subtract the original whole.

The working. 72 − 64 = 8 percentage points. 8 ÷ 64 = 0.125, so a 12.5% relative increase. Note 8 ÷ 72 = 11.1%, which is the same movement measured against the wrong base.

What the item separates. The item separates candidates who can compute a percentage from candidates who know which base the question is asking about. Only the second group can be trusted with a dashboard.

Scenario 2 · Growth compounds

Revenue grew 10% in the first year and a further 20% in the second.

What was the total growth across the two years?

OptionKeyWhy it earns that weight
32%CorrectKeyed answerCorrect. Successive growth multiplies rather than adds.
30%IncorrectDiagnostic distractorThe additive error — the single most documented miscalculation in numerical reasoning banks. It understates every multi-period growth figure.
15%IncorrectDiagnostic distractorThe mean of the two rates. Diagnoses a candidate who reached for an average because two numbers were present.
33.1%IncorrectDiagnostic distractorCompounding the wrong rate — 1.10 applied three times. Diagnoses someone who has understood compounding and misapplied it, which is a different and more recoverable error than the additive one.

The working. 1.10 × 1.20 = 1.32, so 32%. The additive answer, 10 + 20 = 30%, omits the 10% earned on the second year's growth. 1.10³ = 1.331 is where 33.1% comes from.

What the item separates. Distractors 2 and 4 are both wrong and they diagnose opposite problems. A report that says only 'incorrect' loses that distinction.

Scenario 3 · Successive discounts

A list price is reduced by 20%, and the reduced price is then cut by a further 15% in a clearance.

What is the total discount against the original list price?

OptionKeyWhy it earns that weight
32%CorrectKeyed answerCorrect. The second discount applies to 80% of the list price, not to the list price.
35%IncorrectDiagnostic distractorThe two discounts added. The same additive error as item 2, in the direction that overstates rather than understates.
68%IncorrectDiagnostic distractorThe fraction of the price that remains, reported as the discount. Diagnoses a candidate whose arithmetic was right and who then answered a different question.
30%IncorrectDiagnostic distractorThe most interesting wrong answer here: someone who knows the additive figure overstates and has shaved it by eye. The instinct is right and the correction goes too far — the true answer is 32%, not below 32%.

The working. 0.80 × 0.85 = 0.68 of the list price remains, so the discount is 1 − 0.68 = 0.32, or 32%. Adding 20 + 15 gives 35% and double-counts the 15% on the fifth of the price already removed.

What the item separates. Every option here is produced by a real method. That is what a good distractor set costs to build, and it is why the wrong answers are worth reading.

Scenario 4 · Averages need weighting

A team of 12 has an average tenure of 3 years. Four people leave; their average tenure was 6 years.

What is the average tenure of the 8 who remain?

OptionKeyWhy it earns that weight
1.5 yearsCorrectKeyed answerCorrect. Work in totals, then divide by the headcount that is actually left.
1 yearIncorrectDiagnostic distractorThe right numerator over the wrong denominator — the remaining tenure divided by the original headcount of 12. Diagnoses a candidate who set the problem up correctly and lost the last step.
4.5 yearsIncorrectDiagnostic distractorThe unweighted mean of the two averages. Diagnoses someone averaging averages, which is only valid when the groups are the same size and they are not.
3 yearsIncorrectDiagnostic distractorThe assumption that an average is unchanged by who leaves. Diagnoses a candidate treating the mean as a property of the team rather than of the people in it.

The working. Total tenure = 12 × 3 = 36 years. Leavers = 4 × 6 = 24 years. Remaining = 36 − 24 = 12 years across 8 people = 1.5 years. Dividing that 12 by the original 12 people gives the 1-year distractor.

What the item separates. This is the item that most often separates people who are fluent with percentages but have never had to reconstruct a total from an average — which is most of what workforce reporting asks for.

Scenario 5 · Rates do not average

A courier drives a delivery route at an average 40 km/h and returns along the same route at an average 60 km/h.

What is the average speed for the round trip?

OptionKeyWhy it earns that weight
48 km/hCorrectKeyed answerCorrect. More time is spent on the slower leg, so the average sits below the midpoint.
50 km/hIncorrectDiagnostic distractorThe arithmetic mean of the two speeds — the single most documented wrong method in the rate family. It is the answer most candidates produce in under five seconds.
It cannot be determined without knowing the distanceIncorrectDiagnostic distractorA tempting and wrong caution. The distance cancels, so the answer is the same for a 10 km route and a 1,000 km one. Diagnoses a candidate who has correctly noticed that something is missing and not tested whether it matters.
100 km/hIncorrectDiagnostic distractorThe two speeds summed. Rare, and when it appears it usually indicates the candidate was out of time rather than out of method — which is exactly the confound the speeded/power distinction above is about.

The working. For a one-way distance d, time out = d/40 and time back = d/60, so total time = d(1/40 + 1/60) = d/24. Total distance = 2d. Average speed = 2d ÷ (d/24) = 48 km/h. The d cancels, which is why the distance is not needed.

What the item separates. The best-designed items are ones where the fast intuitive answer is on the page. An item whose wrong options are implausible measures reading, not reasoning.

What can be changed

  • Choice of form — speeded (Graduate, 25 items in 22 minutes) or power (Professional, 20 items in 25 minutes). These are not interchangeable; pick the one that matches the question you want answered.
  • Time limit adjusted independently of item count, including an untimed administration where the construct you want is accuracy alone.
  • Data presented in your own tables and charts — a finance operation and a logistics operation read different figures, and a generic table is the weakest part of any off-the-shelf numerical test.
  • Calculator on or off. This changes the construct materially: with a calculator the item measures method selection, without one it also measures arithmetic execution.
  • Item mix weighted toward the families that matter for the role — percentage and base handling, growth and interest, ratio and mixture, rate and work, or data-table extraction.
  • Accessibility accommodations, including extended time, which for a speeded form is a change to what the score means and should be recorded as such.

What an SJT will not do, including ours

A buyer would find most of this in ten minutes of reading, so it belongs on the page that sells the instrument rather than in a footnote somewhere else.

  • A raw score is not a standing. There is no AssessAll norm group, benchmark or national percentile for this or any other instrument: the platform has not passed its own evidence gate on any form (scripts/qa/research-readiness.ts requires 400 sittings on one form, from at least 10 organisations, none above 25%, with every reported cut at n ≥ 50). A band tells you where a score sits against the key, not against a population.
  • AssessAll has not run a local criterion-validation study on either numerical form. Nothing here should be read as evidence that a score predicts performance in your roles; that is a study you would have to run, on your own criterion.
  • A numerical reasoning test measures reasoning with figures, not job knowledge and not judgement. It will not tell you whether someone will notice that a number is wrong in a document nobody asked them to check, which is usually the thing an employer is actually worried about.
  • The items on this page were written for this page. They are not from the live bank, not from any form a candidate sits, and not a practice test.
  • Any cognitive measure needs its own adverse-impact monitoring wherever it is used to select. The four-fifths check is a starting point, not a clearance, and it is the employer's obligation rather than the vendor's.
  • Neither form is affiliated with, equivalent to, or benchmarked against any third-party publisher's instrument, and no score from it converts to one.

Questions buyers ask

What is the difference between a numerical reasoning test and a maths test?

A maths test asks you to perform an operation you have been taught. A numerical reasoning test gives you figures in a work context and asks which operation the question calls for — whether a percentage is taken on the old base or the new one, whether growth adds or compounds, whether an average needs weighting. The arithmetic involved is usually school level. Choosing the right arithmetic is the part being measured, and it is the part that transfers to reading a report.

Should we use a timed or an untimed numerical test?

It depends on whether the job is bounded by time. If the role involves reading figures under deadline pressure, a speeded form is measuring something the job genuinely requires. If the role involves getting a number right and the deadline is a week, a power form tells you more and a speeded form will penalise exactly the careful candidate you want. Whichever you pick, record which one you used: the same person will score differently on the two, and comparing scores across forms is not valid.

What score should we set as the pass mark?

Not one borrowed from another organisation, and not a round number chosen because it looks reasonable. A defensible cut score comes from a standard-setting method — the Angoff method is the most common — in which people who know the job judge, item by item, how a borderline-competent candidate would perform. Then check the cut against the standard error of measurement: if the SEM is 5 points and your cut is 60, candidates scoring 55 to 65 are not reliably distinguishable from each other.

Can candidates prepare for a numerical reasoning test?

Yes, more so than for most instruments, and this is a real limitation rather than a marketing point. Practice reliably improves scores on the first few sittings, mostly by removing format unfamiliarity and by teaching the common traps — which means a coached candidate and an uncoached one are not on the same starting line. Two things reduce the effect: item banks large enough that a form is not memorable, and reporting attempted alongside correct so that a coached fast-guesser is visible.

Is a numerical reasoning test fair?

Fairness is a property of how a test is used, not of the test. A numerical measure can be defensible for a role where reading figures is genuinely part of the work and hard to defend as a blanket filter across a whole recruitment drive. In practice: check that the construct is required by the role, monitor selection rates by group with the four-fifths rule as an early indicator, and keep the evidence you relied on. Our adverse impact calculator does the arithmetic; the judgement remains yours.

How many candidates do we need before a pass rate means anything?

More than most people assume. A pass rate observed on a few dozen sittings carries a confidence interval wide enough to cover most decisions you would want to make with it, and the standard error on an individual score does not shrink as the cohort grows. The sample-size calculator works both numbers for a cohort you specify.

Sources

Checked at source 4 September 2026.

Related reading

The form described on this page

Numerical Reasoning — Graduate and Numerical Reasoning — Professional25 items (Graduate) · 20 items (Professional), 22 minutes (Graduate) · 25 minutes (Professional). Right/wrong against a computed key; no negative marking.

See the listing