All articles
Hiring Practice26 September 2026·6 min read

18 of 391 Employers Published the Bias Audit. Here's How to Run One on Your Own Funnel.

A step-by-step guide to adverse impact analysis in hiring: what the four-fifths rule actually says, why passing it is not a clean bill of health, and the records US state rules now expect employers to keep.

By AssessAll Editorial

An adverse impact analysis compares each demographic group's selection rate, at every stage of your hiring process, against the rate of the group with the highest rate. If a group passes at less than four-fifths of that rate, enforcement agencies treat it as evidence of adverse impact, and the burden shifts to you to show the step is job-related.

That is the whole mechanic — arithmetic a hiring manager can do in a spreadsheet. Most organisations still do not. When researchers audited compliance with New York City's Local Law 144, which requires an annual independent bias audit of automated employment decision tools plus a public summary of results, they checked 391 employers and found 18 that had posted an audit report and 13 that had posted the required candidate transparency notice (Wright et al., FAccT 2024). The law had been in force for months. The analysis was not hard; it simply was not being run.

What the four-fifths rule actually says

The exact language, from the 1978 Uniform Guidelines on Employee Selection Procedures, is narrower than the shorthand suggests:

"A selection rate for any race, sex, or ethnic group which is less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact" — 29 CFR 1607.4(D).

Three words do a lot of work there. Generally — a rule of thumb, not a statute. Evidence — not proof, and not a verdict. And the group with the highest rate — the reference group is whichever group did best in your data, not a fixed majority group.

The same section sets the bottom-line convention: where the total process shows no adverse impact, agencies generally will not expect component-by-component analysis of the steps inside it — with exceptions where there is evidence of a discriminatory practice or a component already shown to be job-irrelevant.

Why passing the rule is not a clean bill of health

The four-fifths rule is a ratio of two proportions. It says nothing about whether the gap it measures is real or an artefact of small numbers, and that is its central weakness.

  • It is unstable at small sample sizes. Roth, Bobko and Switzer modelled the rule's behaviour and concluded practitioners should treat its output with caution rather than as a determination (*Journal of Applied Psychology*, 2006); Collins and Morris devoted a full paper to the small-sample case two years later (*JAP*, 2008). With 30 applicants per group, one extra hire can move the ratio across the 0.80 line either way.
  • Significance tests alone are not the answer either. They are sensitive to sample size in the opposite direction — trivial gaps turn significant in very large pools — and they say nothing about the magnitude of the difference. Practitioner guidance converges on using several measures together rather than any single one (Statistical Significance Standards for Adverse Impact Measurement).
  • A clean bottom line can hide a bad component. If your assessment stage runs at 0.62 and your interview stage over-corrects, the funnel passes overall while one step does damage you cannot see or defend in isolation.

Running the analysis: six steps

1. Model the funnel as stages, not one event

Compute a selection rate at every transition — applied to screened, screened to assessed, assessed to interviewed, interviewed to offered, offered to hired — then cumulatively. The end-to-end number tells you whether you have a problem; stage-level numbers tell you where it is.

2. Fix the applicant definition first

Decide in writing who counts as an applicant, how you treat candidates who abandoned an assessment midway, and how you handle duplicates and internal transfers — before you compute anything. A definition chosen after seeing the result is not an analysis; it is a defence.

3. Collect the demographic data on a lawful basis

The Uniform Guidelines expect US records by sex and by the EEO-1 race and ethnicity categories. Outside the US there is often no equivalent scheme — India has no adverse-impact framework comparable to the Uniform Guidelines — so use voluntary self-identification with a stated purpose, store it separately from the hiring decision record, and keep it invisible to assessors and interviewers.

4. Report three numbers, not one

For every group and every stage, record:

  • The impact ratio — the group's selection rate divided by the highest group's rate. Flag anything under 0.80.
  • A test of statistical significance — a two-proportion z-test for reasonable pools, an exact test for small ones, which tends to be conservative and will under-flag.
  • The shortfall — how many additional selections from the affected group would reach parity. This is what tells you whether a flagged ratio matters in practice or reflects two people out of eleven.

5. Analyse every automated component separately

Any step where software ranks, scores, screens or filters gets its own ratio, significance test and date stamp — resume parsers, knockout questions and scheduling logic included, not just the obvious "AI" tools.

6. Keep the evidence trail

The analysis is only as good as the record beneath it. If your assessment stage cannot tell you, six months later, which items each candidate saw and how each response was scored, you cannot re-run the numbers when someone asks. AssessAll retains per-attempt scoring against a fixed rubric alongside an integrity band for each session — the minimum needed to audit a stage after the fact instead of reconstructing it from memory.

What the current rulebook expects you to have on file

The rules have moved quickly and unevenly, but the direction is toward documentation, not prohibition.

  • New York City — Local Law 144 has required an annual independent bias audit, a published results summary and advance candidate notice since July 2023.
  • California — Civil Rights Council regulations on automated-decision systems took effect 1 October 2025. They define an ADS broadly enough to cover resume screeners, targeted job advertising and systems analysing candidate video or audio; they require ADS-related records to be kept for four years; and they make the quality, scope, recency, results and employer response to anti-bias testing relevant evidence — so an absence of testing can itself count against an employer (Civil Rights Council, summary).
  • Illinois — the HB 3773 amendment to the Human Rights Act took effect 1 January 2026. It prohibits AI use with a discriminatory effect on protected classes, requires notice when AI is used in employment decisions, and addresses ZIP codes used as proxies for protected classes; the Department of Human Rights and the Human Rights Commission enforce it (overview).
  • Colorado — SB 24-205 was repealed and re-enacted as SB 26-189 in May 2026, reframed around automated decision-making technology in covered domains including employment, effective 1 January 2027.
  • Federal US — the EEOC's dedicated AI guidance documents came down from its website in early 2025. Title VII and the 1978 Uniform Guidelines were not amended, and the underlying disparate-impact exposure did not leave with the web pages (analysis).

When the ratio is genuinely below 0.80

The instinct is to remove the offending step. Resist it until you have asked the prior question: is the step valid?

A well-validated, job-related assessment that shows adverse impact sits in a completely different position — legally and practically — from an unvalidated one with the same gap. The Uniform Guidelines contemplate exactly this: impact triggers a validity requirement, not an automatic prohibition. So the sequence is validate, then mitigate:

  1. Confirm with documented evidence that the step measures something the job requires.
  2. Strip construct-irrelevant difficulty — heavy verbal load in a numerical test, unfamiliar interface conventions, time limits tighter than the work demands.
  3. Consider reordering: in a multiple-hurdle design, an early high-impact screen removes candidates a later, lower-impact step might have passed.
  4. Re-run the analysis on new data, and keep both versions.

Steps 2 and 3 are where most of the recoverable impact lives, and neither requires lowering the standard.

The takeaway: the four-fifths rule is a monitoring instrument, not a compliance certificate — run it stage by stage, pair it with a significance test and a shortfall count, and fix your method before you see the result. The organisations that get caught out are rarely those whose numbers looked bad, but those who never generated the numbers.

#adverse-impact#four-fifths-rule#bias-audit#selection-process#compliance

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free