All articles
Assessment Science1 September 2026·7 min read

70% Is Not a Cut Score: How to Set a Passing Mark You Can Defend

A cut score is a policy judgment, not a property of the test. A step-by-step guide to setting a defensible passing mark: modified Angoff panels, selection-ratio reality checks, measurement error at the boundary, and an adverse-impact check before launch.

By AssessAll Editorial

A cut score is the point on an assessment scale that separates candidates who pass from those who do not. It is not a property of the test. It is a policy decision, made by human judges, about how much proficiency is enough for a specific job in a specific context. Two competent panels can defensibly set two different cut scores on the same test.

That last sentence is the one most hiring teams have never been told, and it explains why so many passing marks are indefensible. Somebody picked 70%. Nobody can say why.

The round number is the tell

There is no true cut score waiting to be discovered inside a test. There is only a judgment, made well or badly, about where acceptable performance begins.

The evidence is unambiguous. A 2025 case study in Language Testing in Asia ran two standard-setting methods on the same 150-item English proficiency test, using 15 subject-matter expert raters and data from 1,558 test takers. Angoff recommended a cut score of 72; Bookmark recommended 77. The university's own long-standing cut score was 84 — well above the number either evidence-based method supported, and candidates had been failing against it for years.

A critical review of Angoff methods in the Alberta Journal of Educational Research notes that Woehr and colleagues (1991) compared seven cut-score procedures and found that although results varied, all fell within measurement error of each other. Method choice matters less than having a method at all.

The US Uniform Guidelines on Employee Selection Procedures put it plainly at 29 CFR 1607.5(H): cutoff scores "should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force." Reasonable and consistent with expectations. Not round.

Step 1: Decide what the score is for before you decide what it is

A cut score and a ranking are different instruments with different evidential burdens. The Uniform Guidelines are explicit at 1607.5(G): evidence sufficient to support a selection procedure on a pass/fail basis "may be insufficient to support the use of the same procedure on a ranking basis." The practical translation: if you screen out everyone below 65 and interview the survivors in any order, you need to justify 65. If you rank survivors and offer to the top 40, you need to justify that a candidate scoring 82 is genuinely better than one scoring 78 — a much harder claim, because at that resolution you are usually reading noise. Decide first. Most volume-hiring funnels need a screen, not a ranking.

Step 2: Run a modified Angoff panel

The workhorse method: judgment-based, replicable and documentable, which is exactly what a defence requires.

| Stage | What happens | |---|---| | Define the MCC | The panel agrees, in writing, what a minimally competent candidate for this role knows and can do | | Assemble the panel | Practitioner guidance suggests 6–20 SMEs, with 8–10 preferred; include line managers, not just trainers | | Round 1 | Each SME independently estimates, item by item, the percentage of MCCs who would answer correctly | | Discussion | The facilitator surfaces items with the widest disagreement and makes raters argue their reasoning | | Round 2 | SMEs re-rate after discussion; dispersion should narrow | | Aggregate | The mean of the item estimates is the recommended cut score |

The narrowing is measurable. In the 2025 study above, Angoff rating variability fell from a standard deviation of 13.47 in round one to 8.36 in round two — the discussion is not ceremony, it is where the estimate becomes reliable.

Know the method's weakness before you rely on it. Berk (1996) called estimating minimally competent performance a "nearly impossible cognitive task"; Shepard (1995) argued it "exceeds human cognitive processing capacities." The record is mixed rather than damning — Goodwin (1999) found judges' estimates differed from empirical item difficulties by a mean of just 0.03, while Plake and Impara (2001) found judges were not particularly good at the borderline case specifically. Expert SMEs also overestimate what an entry-level candidate can do, because they last were one a long time ago. Expect round-one numbers to run high.

Step 3: Reality-check against your selection ratio

A cut score you cannot fill is a failed cut score. The Taylor-Russell framework makes the trade-off concrete: with a base rate of .50 — half of all applicants able to perform adequately — and a selection ratio of .10, or ten applicants per opening, a test with a validity of .25 yields roughly 67% successful hires among those selected, and one with a validity of .30 yields about 71%. Four points of accuracy across 20,000 hires is roughly 800 additional capable people.

Two things follow. A valid test buys you the most when your base rate sits near .50; as it approaches 0 or 1, no test improves much on chance. And your cut score is your selection ratio — set it high and you buy accuracy with volume you may not have. Run it both ways before committing: what share of last quarter's actual applicant pool clears this line, and does that leave enough people to fill the roles?

Step 4: Respect the error band at the boundary

Every score carries measurement error, and it bites exactly where it hurts — at the cut. A candidate one point below the line and one point above are statistically indistinguishable.

Some organisations adjust the recommended cut score down by one standard error of judgment. Others report a band rather than a point and route borderline candidates to a second hurdle — a structured interview, a work sample, a scored scenario — instead of rejecting them on a difference the instrument cannot detect. On AssessAll, AI-graded scenario responses and integrity bands are useful precisely here: a second, differently-sourced signal for the people sitting inside the noise, rather than a coin-flip on the raw number.

What you should not do is pretend the boundary is sharp because a spreadsheet renders it that way.

Step 5: Test for adverse impact before you go live, not after

Run your proposed cut score against historical applicant data and compute selection rates by group. The four-fifths rule at 29 CFR 1607.4(D) treats a selection rate below 80% of the highest group's rate as evidence of adverse impact — while cutting both ways, noting that "smaller differences in selection rate may nevertheless constitute adverse impact, where they are significant in both statistical and practical terms," and that larger ones may not, where they rest on small numbers and are not statistically significant.

Nor is this only a US concern. New York City's Local Law 144 requires that an automated employment decision tool have been subject to a bias audit within one year of use, with results publicly available. If your cut score is the mechanism producing a disparity, an audit will find it — and you will be explaining a number nobody documented.

The checklist

  • Written definition of the minimally competent candidate, agreed before any rating
  • Panel of 8–10 SMEs including people who supervise the role
  • Two rating rounds with structured discussion between them; record the dispersion in both
  • Recommended cut score expressed as a band, not a point
  • Yield modelled against a real applicant pool, not a hypothetical one
  • Group selection rates computed at the proposed cut before launch
  • A named owner, a dated rationale document, and a review date
  • Re-set when the job changes, the pool changes, or the test changes — not annually out of habit

When a cut score is the wrong tool

Sometimes it is. Where the role is scarce and every marginal point of ability translates to output, ranking with strong validity evidence may serve better than a threshold. Where the assessment measures several distinct things, a single compensatory total can hide a fatal gap in one of them, and a multi-hurdle design with a modest cut on each is the honest structure. And if you have fewer applicants than openings, you do not have a selection problem; you have a sourcing problem, and no cut score fixes it.

Takeaway

A defensible passing score is not a better guess; it is a documented process — a defined standard, an expert panel, two rounds, an error band, and an impact check run before launch rather than after a complaint. If nobody in your organisation can name who set your current cut score or why, that number is not a standard. It is a habit.

#cut-scores#standard-setting#angoff#adverse-impact#psychometrics

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free