---
title: "Spearman–Brown formula"
description: "The Spearman–Brown formula predicts what a test's reliability would become if it were lengthened or shortened by a given factor. Doubling a test of reliability 0.70 projects 0.82; quadrupling it projects 0.90. It assumes the questions being added are as good as the ones already there, which is the assumption that usually fails."
canonical: https://www.assessall.com/guides/glossary/s/spearman-brown-formula
updated: 2026-09-25
source: AssessAll
---

# Spearman–Brown formula

_Also known as: Spearman–Brown prophecy formula, Spearman–Brown prediction formula, Prophecy formula._

The Spearman–Brown formula predicts what a test's reliability would become if it were lengthened or shortened by a given factor. Doubling a test of reliability 0.70 projects 0.82; quadrupling it projects 0.90. It assumes the questions being added are as good as the ones already there, which is the assumption that usually fails.

<!-- #what-it-is-actually-for -->
## What it is actually for

A reliability coefficient on its own tells you nothing you can act on. It is not a grade; it is an input. The Spearman–Brown formula is what turns it into a decision — given where a test's reliability is now and where it needs to be, how many questions have to be added. The relation is n times r, divided by one plus n minus one times r, where r is the current reliability and n is the factor by which the test's length changes.

Rearranged, it answers the question a test designer actually has rather than the one the textbook asks. To move from a current reliability to a target one, n = target × (1 − current) ÷ (current × (1 − target)). Multiply that by the number of items you have and the answer comes back in questions, which is a thing you can commission, rather than in a coefficient, which is not.

The same relation applies to raters as well as to items, and that is where it does the most commercial work. At the interrater reliability the literature reports for ratings collected for administrative purposes, 0.45, one rater gives 0.45, two give about 0.62, three about 0.71 and five about 0.80. That curve — not a better rating scale — is the reason a 360 review asks for more raters.

<!-- #spearman-wrote-the-assumption-down-and-the-market-quotes-the -->
## Spearman wrote the assumption down, and the market quotes the formula without it

The formula comes from Charles Spearman's 1910 paper *Correlation Calculated from Faulty Data*, published in the *British Journal of Psychology* in the same volume as William Brown's paper, which is why both names are on it. A free scan of the original was read at source on 25 September 2026, and the condition the formula holds under is stated on page 281 in a single sentence: "Here, and also in the following section, all the measurements of x (or of y) are supposed to have been of equal general accuracy."

**Equal general accuracy is a strong requirement, and a test being lengthened in practice almost never meets it.** The items already in the test are the ones that survived [item analysis](https://www.assessall.com/guides/glossary/i/item-analysis). The items being added are the next best available, which by construction are worse than the ones that survived. So the projection is an upper bound, not an estimate — and a vendor who presents it as a plan is quoting Spearman without the sentence Spearman attached to it.

This is worth asking about directly, because the answer is cheap to give and revealing either way: were the additional items drawn from the same calibrated pool as the existing ones, and what did their discrimination indices look like? A projection made before the items exist is arithmetic. A reliability figure measured after they exist is evidence.

<!-- #the-case-the-formula-cannot-express-lengthening-that-makes-r -->
## The case the formula cannot express: lengthening that makes reliability worse

The bound can be breached in the other direction, and not by a rounding error. Yigal Attali's ETS research report on the reliability of speeded number-right multiple-choice tests reaches a conclusion the prophecy formula has no way to represent: adding an item can **lower** a test's reliability once more than half the cohort has to guess at it, because a guessed response contributes noise rather than information.

That inverts the standard remedy. A test that fails to separate people is usually diagnosed as too short, and this formula is what licenses the fix. But if the reason it fails to separate people is that the items are too hard for the cohort, or that the time limit stops most people reaching the end, then more items make it worse and the formula will still promise an improvement. Check the completion rate and the item difficulties before buying the projection.

---

Source: https://www.assessall.com/guides/glossary/s/spearman-brown-formula
