All articles
Hiring Practice15 September 2026·6 min read

Multiple-Hurdle vs Compensatory Scoring: How You Combine Assessment Results Decides Who You Reject

Sequential hurdles or one weighted composite? The evidence on accuracy, adverse impact, cost and legal defensibility — and how to choose between them deliberately.

By AssessAll Editorial

Multiple-hurdle selection administers assessments in sequence, eliminating candidates who fail each stage before they reach the next. Compensatory selection administers everything to everyone and combines results into a single weighted score, so strength on one measure can offset weakness on another. The choice is not administrative housekeeping. It changes who gets rejected, how accurate your decisions are, and what you can defend.

Most hiring teams never make this choice deliberately. The funnel grows one stage at a time — a screening test here, a video interview there — and the multiple-hurdle model arrives by accident. That matters, because the two models fail in opposite directions.

The two models, precisely

Multiple-hurdle (sequential)

  • Each stage has its own cut score.
  • Failing any stage ends the application, regardless of performance elsewhere.
  • Only survivors incur the cost of the next stage.
  • A candidate who would have scored well overall can be removed by a single weak measure.

Compensatory (composite)

  • Every candidate completes the full battery.
  • Scores are standardised, weighted and summed into one composite.
  • One cut score, or a rank order, is applied at the end.
  • A strong structured-interview score can offset a middling test score, and vice versa.

Hybrids are common and legitimate: a minimum-qualification gate (a licence, a language floor) followed by a compensatory composite for everyone who clears it.

On accuracy, compensatory wins

The reliability argument is arithmetic. A composite of several measures is more reliable than any single measure within it, and reliability sets the ceiling on validity. A hurdle asks one imperfect instrument to make an irreversible decision.

Jisoo Ock and Frederick Oswald tested this directly with Monte Carlo simulation in the Journal of Personnel Psychology (2018). Their conclusion: compensatory selection "produced a higher level of expected criterion performance in the selected applicant subgroup, and a higher overall selection utility in most conditions." They were explicit that practitioners choose hurdles anyway, because "administering an entire predictor battery to every applicant can be time-consuming, labor-intensive, and costly." That is the real trade-off — reliability against cost, not accuracy against convenience.

There is a second, quieter cost to hurdles: they corrupt your own validity evidence. When later-stage data exists only for people who cleared earlier stages, the range on every predictor is restricted, and observed correlations with job performance understate the true relationship. If you are trying to prove your interview works using only the candidates your test let through, you are measuring a truncated sample. The range-restriction problem is severe enough that Sackett, Zhang, Berry and Lievens's re-analysis in the Journal of Applied Psychology found decades of meta-analytic validity estimates had been overcorrected for it, reducing mean validities by .10 to .20 points across most procedures.

On adverse impact, hurdles can win — if you order them well

This is where the comparison gets genuinely interesting, because the evidence does not run one way.

Finch, Edwards and Wallace simulated 43 multistage selection strategies against three single-stage compensatory baselines in the Journal of Applied Psychology (2009, vol. 94, pp. 318–340). Their headline: multistage strategies produced substantially less adverse impact than compensatory models, particularly when the first stage was not a cognitive ability test.

Some of their reported figures:

  • A single-stage composite of integrity, cognitive ability and a structured interview produced an adverse impact ratio of 0.255 at a 10% net selection ratio.
  • A two-stage version — integrity first, then cognitive ability plus the interview — reported a ratio of 0.545.
  • Their best-balanced strategy (integrity first, then biodata, interview and conscientiousness) reached a ratio of 0.651 at a 25% selection ratio, with expected performance of 0.611.
  • The best-performing multistage strategy reached expected performance of 0.742, against 0.76 for the single-stage composite.

Read those last two together. You give up a little expected performance and you buy a large reduction in impact — but only because of what you put first. Order is the lever. A cognitive test in stage one applies the largest subgroup difference to the largest applicant pool, and no downstream cleverness recovers the people it removed. These are simulations, not field results; treat them as a map of the trade-off surface rather than as numbers to quote at your legal counsel.

The legal asymmetry nobody prices in

A compensatory composite is evaluated as one procedure. A multiple-hurdle funnel is a series of procedures, and in the United States each one can be challenged on its own.

The Uniform Guidelines on Employee Selection Procedures contain a "bottom line" provision at §1607.4(C): where the total selection process shows no adverse impact, enforcement agencies "in usual circumstances, will not expect a user to evaluate the individual components." Employers read that as shelter for their hurdles. It mostly is not.

In Connecticut v. Teal, 457 U.S. 440 (1982), a written examination passed 54.17% of 48 Black candidates and 79.54% of 259 white candidates — roughly 68% of the white rate, below the four-fifths benchmark. The state then promoted 22.9% of Black candidates and 13.5% of white candidates, and argued the favourable bottom line was a complete defence. The Supreme Court rejected it: a "nondiscriminatory 'bottom line'" neither prevents a prima facie case nor provides a defence to one. Title VII protects the individual's opportunity to compete, not the group's final tally.

The practical consequence: every hurdle you add is a separate object of proof. The Guidelines also warn at §1607.5(G) that evidence sufficient to justify pass/fail use may not justify ranking, and §1607.5(H) that cut scores "should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force." The SIOP Principles and the EEOC's guidance on employment tests set the same expectation: justify the use you actually make of the score.

When hurdles are the right answer

Honestly, often.

  • True minimum qualifications. A licence, a legal work requirement, a hard language floor for a voice role. There is nothing to compensate; the requirement is binary.
  • Extreme selection ratios. At 40,000 applicants for 300 seats — routine in Indian campus and BPO drives — running a full battery on everyone is not a reliability decision, it is a budget one.
  • Expensive human stages. Assessment centres and panel interviews cannot be compensatory in practice.
  • Sequencing that protects impact. Where the evidence above applies, a well-ordered sequence can beat a composite on adverse impact.

The economics have shifted, though. The cost argument for hurdles was built when every assessment seat carried a licence fee. Consumption pricing changes the arithmetic: on AssessAll, assessment credits run ₹30 / US$0.50 per assessment, which makes a two- or three-instrument composite for an entire shortlist a different order of decision than it was under seat licences.

How to decide, in five steps

  1. Separate true minimums from preferences. Anything that is a preference belongs in the composite, not in a gate.
  2. Order remaining stages by subgroup difference, smallest first. Put the measure with the largest known difference last, where the pool is smallest.
  3. Set each cut score with a defensible method — Angoff or bookmark, tied to job content — not at a round number.
  4. Track pass rates by subgroup at every stage, not just at the offer. Teal is the reason.
  5. Convert flags into weights where you can. Integrity signals are a good example: AssessAll reports proctoring results as integrity bands rather than a pass/fail verdict, which lets a reviewer weigh a borderline session alongside everything else instead of ending the application on one signal.

The takeaway

Compensatory scoring is usually the more accurate model and the easier one to defend as a single procedure; multiple-hurdle scoring is usually the cheaper one and, when sequenced deliberately, can reduce adverse impact. Choose between them on purpose — and if you have hurdles you never chose, the first thing to check is what your stage-one instrument is doing to the largest pool you will ever see.

#multiple-hurdle#compensatory-scoring#selection-design#adverse-impact#hiring-funnel#psychometrics

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free