---
title: "How reliable is a 10-question aptitude test?"
description: "About 0.69 to 0.73 — and three published routes agree on that band. Here is the arithmetic, checked against a real 60-item instrument, plus what a reliability of 0.69 does to one candidate's score and to a comparison between two of them."
canonical: https://www.assessall.com/guides/how-reliable-is-a-10-question-aptitude-test
updated: 2026-09-30
source: AssessAll
---

# How reliable is a 10-question aptitude test?

About 0.69 to 0.73 — and three published routes agree on that band. Here is the arithmetic, checked against a real 60-item instrument, plus what a reliability of 0.69 does to one candidate's score and to a comparison between two of them.

_Last updated 2026-09-30._

<!-- #the-short-answer -->
## The short answer

A 10-question aptitude test has a reliability of about 0.69 to 0.73. Three independent routes from published figures agree on that band. At 0.69, one candidate's score carries a 95% band of nearly eleven points on a scale whose standard deviation is ten, and two candidates need a fifteen-point gap before the test has separated them at all.

That is not an opinion about short tests. It is a formula from 1910 applied to numbers somebody else published, and every step below is one you can re-run yourself in a spreadsheet.

The practical consequence is narrower than *short tests are bad* and more useful. A ten-item test can tell you roughly where a large group sits. It cannot tell you which of two shortlisted people is stronger, and the arithmetic says by how much it cannot.

<!-- #what-the-pages-ranking-for-this-question-actually-say -->
## What the pages ranking for this question actually say

Searched on 30 September 2026, the phrase *how reliable is a 10 question aptitude test* returns nine results: a journal PDF, Wikipedia's entry on a mechanical aptitude test, Criteria Corp, AssessmentDay, PracticeAptitudeTests, TestInvite, TestDome and two Psico-smart blog posts. The same set, allowing for one or two substitutions, has been recorded three times across eleven days.

Three of them were opened in full and read on 30 September 2026. **Not one prints a reliability coefficient, a standard error of measurement, a confidence interval, the Spearman-Brown formula, or any arithmetic connecting the number of questions to the reliability of the score.**

Criteria Corp's *Are Aptitude Tests Accurate?* states that its tests undergo rigorous validation and prints two numbers, both about business outcomes rather than measurement: a 50% reduction in time-to-hire and a 3× improvement in retention. TestInvite's article on ensuring reliable aptitude tests prints no numeric value at all apart from its own publication dates, and never mentions test length. AssessmentDay's *What Is A Good Score For An Aptitude Test?* contains one reference to a number of questions — a worked example reading "In this test, we scored 7 out of a possible 12 which is 58%" — with no band around it.

The rank-one result was an academic pilot study on the validity and reliability of an aptitude test, and it could not be retrieved: the host's robots file blocked the fetch and was not worked around. Nothing is claimed about its contents here, and the other five results were not opened.

So the honest description of this lane is that it is contested on traffic and empty on substance. Every page answers *are aptitude tests accurate in general*. None answers *what happens when there are only ten of them*.

<!-- #the-formula-and-the-check-that-it-actually-works -->
## The formula, and the check that it actually works

The relationship between test length and reliability has had a closed form for 116 years. Spearman and Brown published it independently in the same 1910 volume of the *British Journal of Psychology*, and it is still called the Spearman-Brown prophecy formula. It says that if you multiply the length of a test by n, the new reliability is n times the old reliability, divided by one plus n minus one times the old reliability.

Set n below 1 and it runs backwards, which is what answers this question: a ten-item version of a twenty-five-item test is n = 0.4.

**A formula is worth nothing here unless it predicts something somebody measured, so here is the check.** The International Cognitive Ability Resource is a public-domain cognitive ability measure published by Condon and Revelle in *Intelligence* in 2014. It exists in a 60-item form and a 16-item sample test, the sixteen being a nested subset chosen to be representative of the sixty on difficulty and factor loading — which is close to the parallel-subset condition the formula assumes. Their Study 1 ran 96,958 participants, of whom 4,574 sat the 16-item form. **Both coefficients are published in the same table: α = 0.81 for the 16 items, α = 0.93 for the 60.**

Feed the formula the 16-item figure and ask it to predict a form 3.75 times longer. It returns **0.941. The measured value is 0.930 — an error of 0.011.** Run it the other way, from 0.93 down to sixteen items, and it predicts 0.780 against a measured 0.810, an error of 0.030.

Both directions and the size of each error are worth having. The formula is accurate enough to use, it is more accurate lengthening than shortening on this instrument, and when it errs shortening it errs **low** — so a shortening estimate is a floor rather than a flattering guess. That is the opposite of the direction a vendor would choose.

<!-- #three-routes-to-the-ten-item-number-and-they-agree -->
## Three routes to the ten-item number, and they agree

One estimate from one baseline is a guess. Three from different baselines that land in the same place is an answer.

**From the 60-item ICAR at 0.93, cut to ten items: 0.689.** **From the 16-item ICAR at 0.81, cut to ten items: 0.727.** And from the illustrative case this site already publishes — a 25-item graduate form at 0.85 — **a ten-item version comes out at 0.694.**

So the band is **0.69 to 0.73**, and the reason the two ICAR routes differ slightly is the same asymmetry the check above exposed rather than a disagreement about short tests.

For comparison, the thresholds the measurement literature conventionally uses are 0.70 as a floor for group-level research, 0.80 for provisional individual decisions and 0.90 or above for consequential ones. **A ten-item aptitude test sits on or just below the lowest of the three.**

<!-- #what-a-reliability-of-0-69-does-to-one-candidate-s-score -->
## What a reliability of 0.69 does to one candidate's score

Reliability is not a grade out of one. It converts into an error band, and the conversion is one line: the [standard error of measurement](https://www.assessall.com/guides/glossary/s/standard-error-of-measurement) is the scale's standard deviation times the square root of one minus the reliability.

On the standard-deviation-of-ten scaling most score reports use, a reliability of 0.689 gives a standard error of **5.58**, so the 95% band around a candidate's score is **plus or minus 10.9 points**.

Read that against the scale. **The band is wider than one whole standard deviation either side of the score.** A candidate reported at the population average could sit anywhere from the 14th percentile to the 86th, and the report would look identical.

The 60-item form for contrast: reliability 0.93, standard error 2.65, 95% band plus or minus 5.2 points. **Cutting sixty items to ten multiplies the width of one person's error band by 2.1.**

<!-- #what-it-does-to-a-comparison-between-two-candidates-the-row- -->
## What it does to a comparison between two candidates — the row that decides a shortlist

Nobody buys an aptitude test to put a band around one person. They buy it to choose between people, and that is a different and larger statistic: the [standard error of the difference](https://www.assessall.com/guides/glossary/s/standard-error-of-the-difference), which is 1.41 times the standard error of measurement on a single instrument, because both scores carry error and the gap between them carries both.

Divide a gap by it and compare against 1.96, and you have the smallest gap at which the instrument has separated two people at all.

**Ten items, reliability 0.689: the standard error of the difference is 7.89, so the smallest real gap is 15.5 points — 1.55 standard deviations.** Sixteen items at 0.81: 12.1 points. Sixty items at 0.93: **7.3 points, or 0.73 of a standard deviation.**

**Going from sixty items to ten more than doubles the gap two candidates need before the test has said anything about which of them is stronger — 7.3 points becomes 15.5.** On a ten-item test almost every shortlist is a tie, and the tie is not a close call to be broken by feel. It is the instrument reporting that it cannot tell.

This site's own worked table for that statistic runs at reliabilities of 0.95, 0.90, 0.85 and 0.70. The row a ten-item test actually needs was missing from it, which is why it is computed here.

<!-- #in-raw-scores-out-of-ten-how-many-scores-can-be-told-apart-a -->
## In raw scores, out of ten: how many scores can be told apart at all

The standard-deviation arithmetic above assumes a scaled score. Most ten-item tests do not have one — they report a raw number out of ten, so here is the same question asked in that currency, using the Wilson score interval this site already prints on its eight-item practice test.

There are 55 possible pairs of scores on a ten-item test. **Twelve of them can be distinguished at 95% confidence. Forty-three cannot.**

The full list of separable pairs is short enough to print: 0 against 6 or better, 1 against 8 or better, 2 against 9 or 10, 3 against 10, and 4 against 10. **That is the whole list. Every score from 5 to 10 out of 10 is statistically indistinguishable from every other score from 5 to 10** — so a candidate who got half the questions right cannot be told apart, at 95% confidence, from one who got all ten.

The smallest gap that ever separates two scores on a ten-item test is **six items, which is 60% of the test.** On a 25-item test it is seven items, or 28% of the test, and 150 of 325 pairs separate. On a 40-item test it is eight items, or 20%, and 450 of 820 pairs separate. **The proportional resolution roughly triples between ten items and forty.**

<!-- #so-how-many-questions-would-you-actually-need -->
## So how many questions would you actually need?

Run the formula forwards from the ten-item estimate and the answer is specific. **To reach 0.80, about 18 items. To reach 0.90, about 41.** Starting instead from the ICAR's measured 0.81 at sixteen items, reaching 0.90 needs about 34.

Those are counts of items of the *same quality*, which is the assumption doing the work. Adding worse items buys less than the formula promises, and the formula's own documented limitation says so: if a reliable test is lengthened with poor items the achieved reliability will be lower than predicted.

The practical reading for a buyer is that the jump from ten items to around twenty is where a short test stops being a rough sort and starts being defensible for individual decisions, and the jump to forty is where it becomes comfortable. **Between ten and twenty items you are buying most of the available improvement.**

<!-- #what-a-ten-item-test-is-genuinely-good-for-and-the-one-thing -->
## What a ten-item test is genuinely good for — and the one thing it did better

A low reliability is a statement about precision, not about validity, and conflating them would be the mirror image of the mistake this page is correcting.

The clearest evidence for that is in the same ICAR table. The 16-item form's general-factor saturation was **0.66, higher than the 60-item form's 0.61**. The short form was, on that measure, a slightly *purer* indicator of general cognitive ability while being far less precise about any one person's level of it. **A short test is not a bad test made small; it is a different trade.**

So ten items are genuinely fit for: practice and familiarisation, where the point is the format rather than the score; self-assessment, where the reader owns the uncertainty; routing a large group into broad bands where the bands are wide enough to survive an eleven-point band; and a tripwire early in a funnel where the cost of a wrong call is a second look rather than a rejection.

They are not fit for choosing between candidates, setting a [cut score](https://www.assessall.com/guides/glossary/c/cut-score) with a consequence attached, or reporting a percentile as though it were a property of the person. A ten-item test used for any of those is being asked a question it has the arithmetic to refuse.

<!-- #why-you-should-still-ask-for-the-measured-figure-rather-than -->
## Why you should still ask for the measured figure rather than trusting this page

Every number above is an estimate produced by a formula, and the formula has one assumption: that the shorter form is a parallel subset of the longer one. Real short forms are often not.

The direction of the error depends on how the items were chosen, and it runs both ways. A ten-item form built from the ten most discriminating items of a longer test can beat 0.69, sometimes comfortably. A ten-item form built from whichever items fitted on the page will land near it or below. The ICAR's sixteen items were selected to be *representative* rather than *best*, which is exactly why they make a fair test of the formula and why a vendor's hand-picked short form is not the same case.

**Which makes the estimate a floor to negotiate from, not a verdict.** Compute it, then ask the vendor for the coefficient they measured on the form they are actually selling you. If they publish one, the arithmetic above turns their number into an error band in one line. If they publish nothing, the estimate is the only figure anybody in the transaction has, and that is worth saying out loud.

<!-- #where-assessall-stands-on-this-including-what-it-does-not-pu -->
## Where AssessAll stands on this, including what it does not publish

**No AssessAll aptitude form is as short as ten items**, and the counts are on the public test pages rather than on request: 25 items in 22 minutes on the graduate [numerical reasoning](https://www.assessall.com/guides/tests/numerical-reasoning-test) form and 20 in 25 minutes on the professional one; 20 or 25 on the [verbal reasoning](https://www.assessall.com/guides/tests/verbal-reasoning-test) forms; 20 to 25 across the [logical](https://www.assessall.com/guides/tests/logical-reasoning-test) and [abstract](https://www.assessall.com/guides/tests/abstract-reasoning-test) forms; and 40 items in 40 minutes on the Complete Graduate Aptitude Battery.

The one AssessAll surface that is shorter than twenty items is the free [eight-item practice test](https://www.assessall.com/guides/practice/numerical-reasoning-practice-test), which prints its own confidence interval next to the score and states on the page that no selection decision should rest on eight items. It is a teaser, and the interval is there because the alternative is a number that flatters itself.

**And the part a competitor page would leave out. AssessAll publishes no reliability coefficient, no standard error of measurement and no norm group for its aptitude or English forms.** So the arithmetic on this page cannot be run on an AssessAll measured coefficient either — it can only be run on the item count, which is the figure AssessAll does publish. A reader holding this page to the standard it sets should ask us for the coefficient exactly as they would ask anybody else, and the honest current answer is that we do not have one to give.

That is not a comfortable paragraph to publish on a page about measurement precision. It is more useful than the alternative, because a buyer who cannot check a vendor's claim against the vendor's own admissions has no way to tell a published number from a marketing one.

<!-- #four-questions-to-send-any-vendor-including-this-one -->
## Four questions to send any vendor, including this one

Each is answerable in a sentence. A vendor who cannot answer them about their own instrument has told you something.

How many scored items are on the specific form you are quoting me — not on the longer parent form it was cut from?

What reliability coefficient do you publish for **that** form, is it internal consistency, split-half or test-retest, and if test-retest, over what interval?

What is the standard deviation of the reported scale, so I can compute the standard error and the smallest real gap myself?

Below what score gap do you tell your own customers that two candidates are tied?

Then run the fifth question at yourself: **the short screen already in your funnel has an item count, so compute its band before the next shortlist meeting.** If it has ten items and you have been reading gaps of three or four as rankings, the arithmetic above says you have been ranking noise — and that is a check on a tool you already bought rather than an argument for buying another one.

<!-- #sources -->
## Sources

Spearman, C. (1910). Correlation calculated from faulty data. *British Journal of Psychology*, 3, 271–295. And Brown, W. (1910). Some experimental results in the correlation of mental abilities. *British Journal of Psychology*, 3, 296–322. The two independent 1910 papers the prophecy formula is named for.

Condon, D. M., & Revelle, W. (2014). The international cognitive ability resource: Development and initial validation of a public-domain measure. *Intelligence*, 43, 52–64. Open at the author's own publications page. All four ICAR figures used above — α = 0.81 and ω-hierarchical = 0.66 for the 16-item sample test, α = 0.93 and ω-hierarchical = 0.61 for the 60-item form, 96,958 participants in Study 1 and 4,574 on the sample test — are read from that paper's Table 3 and its Study 1 description, confirmed on 30 September 2026.

The standard error of measurement and standard error of the difference formulas, their primary sources in the National Council on Measurement in Education's instructional module and in Jacobson and Truax (1991), and the four-reliability worked table this page extends, are set out at [standard error of the difference](https://www.assessall.com/guides/glossary/s/standard-error-of-the-difference) and [standard error of measurement](https://www.assessall.com/guides/glossary/s/standard-error-of-measurement).

The Wilson score interval is the method already used for the eight-item interval on AssessAll's free practice test, and the figures out of ten above were computed with it at 95% confidence rather than with the exact binomial, which gives slightly wider bands.

The three competitor pages named in the second section were opened in full and read on 30 September 2026, and what is reported about them is what those pages showed on that date. Search results are re-ranked continuously and this set will change; the observation is dated deliberately. The rank-one result was not retrievable and nothing is claimed about it.

AssessAll item counts and time limits are read from the live published specification on each test page rather than estimated, and the statement that AssessAll publishes no reliability coefficient, standard error or norm group for its aptitude and English forms is current as of 30 September 2026.

<!-- #faq -->
## Frequently asked questions

### How reliable is a 10-question aptitude test?

About 0.69 to 0.73. Three routes agree: cutting the published 60-item ICAR at α = 0.93 down to ten items gives 0.689, cutting its 16-item form at α = 0.81 gives 0.727, and cutting a 25-item form at 0.85 gives 0.694. That is on or just below 0.70, the conventional floor for group-level research and well under the 0.80 usually wanted before an individual decision.

### Can a 10-question test rank two candidates?

Almost never, and the arithmetic is specific. Of the 55 possible pairs of scores out of ten, twelve can be distinguished at 95% confidence and forty-three cannot. Every score from 5 to 10 out of 10 is statistically indistinguishable from every other score in that range, so a candidate who scored 5 cannot be separated from one who scored 10. On a scaled score the same point appears as a smallest real gap of 15.5 points where the scale's standard deviation is 10.

### How many questions should an aptitude test have?

Depends on what the score decides, and the formula gives counts rather than opinions. From a ten-item baseline at 0.69, reaching 0.80 takes about 18 items of the same quality and reaching 0.90 takes about 41. Most of the available improvement is bought between ten and twenty items. Beyond forty the gains are real but small, and the cost is candidate time.

### Is a short aptitude test better than nothing?

For some jobs, yes. Ten items are fit for practice, self-assessment, sorting a large group into wide bands, and an early tripwire whose cost of error is a second look. They are not fit for choosing between shortlisted candidates or for setting a cut score with a consequence attached. The distinction is precision, not validity — in the ICAR study the 16-item form was actually a slightly purer indicator of general ability than the 60-item form while being far less precise about any one person.

### Does a longer test always measure better?

Not automatically. The Spearman-Brown formula assumes the added items are as good as the ones already there, and its own documented limitation is that lengthening a reliable test with poor items yields less than predicted. Length improves precision; it does nothing for whether the test measures what the job needs. A 60-item test of the wrong construct is a precise wrong answer.

### Does AssessAll publish a reliability coefficient for its own tests?

No. AssessAll publishes no reliability coefficient, no standard error of measurement and no norm group for its aptitude or English forms, and says so on its test pages and its comparison pages. What it does publish, per form and without a sales call, is the item count, the time limit and the competencies each form reports — which is what the arithmetic on this page needs. No AssessAll aptitude form is shorter than 20 items; the only shorter surface is a free 8-item practice test that prints its own confidence interval.

---

Source: https://www.assessall.com/guides/how-reliable-is-a-10-question-aptitude-test
