Free tools

How do you set a defensible pass mark?

The Angoff method sets a pass mark by expert judgement rather than by ranking candidates. A panel estimates, for every item, the probability that a just-barely-qualified candidate answers it correctly; each judge's estimates are summed into that judge's cut score, and the panel mean of those totals becomes the cut score. Because it is fixed to the content, the bar does not move when a stronger or weaker cohort arrives.

1. What cut score did the panel set?

One row per judge; one number per item — the probability, 0–100, that a just-barely-qualified candidate answers that item correctly. Paste straight from a spreadsheet. Nothing you type is sent anywhere: the whole calculation runs in your browser.

Panel cut score
6.6 / 1065.8% — 6 judges, 10 items
Standard error of the cut
± 0.25SD across judges 0.62
95% interval on the cut
6.1–7.11.0 score points wide

6 judges is below the researched target range.Hurtz and Hertz’s generalizability study puts 10 to 15 raters as the optimal target for Angoff cut scores. The interval above is real, but it will narrow — and may move — with a fuller panel.

1 item the panel does not agree on (a spread of 40 points or more between the highest and lowest rating). These are the round-2 agenda, in order: item 4 (40 pts), item 9 (25 pts), item 10 (25 pts). Disagreement this wide is usually about what the item is asking, not about how hard it is — which makes it an item-review finding as much as a rating one.

Per-judge and per-item detail
JudgeCutvs panel
A6.3-0.3
B6.3-0.2
C6.4-0.2
D6.0-0.6
E6.8+0.2
F7.7+1.1
ItemMeanRange
175.0%6585
264.2%5575
385.8%8095
446.7%3070
559.2%5070
674.2%6585
749.2%4060
882.5%7590
965.8%5580
1055.0%4570

2. What does that cut score do to real candidates?

A question about people, not about the panel. Enter what your scored cohort actually looks like. The cut score carries over from panel 1.

Cut score
65.8%6.6 of 10 items
Standard error of measurement
± 5.6 pts≈ 0.6 items on this form
Expected pass rate
29.0%normal approximation

26.9% of this cohort scores within one standard error of the bar (60.1%–71.3%). For that quarter and more of your candidates the test is not deciding the outcome — measurement error is. Something else has to make the call in that band: a second instrument, a structured interview, a work sample.

The two conventional adjustments, priced. Moving the bar down one SEM to 60.1% passes 43.9% — it buys back roughly 14.9 points of pass rate and trades false rejections for false acceptances. Moving it up one SEM to 71.3% passes 17.0%, and trades the other way. Neither is more scientific than the other: the SEM tells you how big the adjustment is, and the consequences of a wrong hire versus a wrong rejection tell you which direction to take it.

Two different errors, and they are not interchangeable. The ± 0.25 in panel 1 is how far the bar would move with a different panel. The ± 5.6points here is how far one candidate’s score would move on a retest. Both have to be small next to the decision before a cut score is a decision rule rather than a number.

Pass rates assume scores are approximately normal with the mean and SD you entered. This is an estimate for planning, not a substitute for counting your own sittings — and a pass rate is not fairness evidence: run it by group before you deploy the cut.

The number this returns is half an answer

The Angoff arithmetic itself is not scarce. MetricGate publishes a free calculator that returns the judges’ individual cuts, the panel mean, a standard error and 95% interval, and an ICC(2,k) rater-reliability coefficient; Assessment Systems Corp distributes an Angoff analysis tool behind a download form. Both stop at the same place, which is the standard-setting workshop’s output — the half of the problem that has been solved for fifty years. (Checked 7 September 2026.)

The half nobody prints is what happens next. A cut score does not stay in the workshop; it goes into a hiring system and starts rejecting people, and at that point a completely different error takes over. The panel’s standard error tells you how far the bar would move with a different panel. It tells you nothing at all about how far a candidate’s score would move on a retest, and it is the second number that decides whether the person just below the line belongs there.

So the second panel above asks the question the first one cannot: given this bar, this cohort and this reliability, how many of your candidates are being sorted by measurement error rather than by ability? On a defensible-looking cut with a good instrument, the answer is routinely a fifth to a quarter of the pipeline. Nothing about that is visible in the panel’s mean.

The method everyone uses is a footnote

William Angoff described the method in a chapter called Scales, norms, and equivalent scores in the second edition of Educational Measurement (1971). The procedure in his main text asked judges a yes/no question: would a minimally competent person answer this item correctly? The probability version — estimate what share of just-qualified candidates get it right — appeared as a variation in a footnote.

The footnote is what the world adopted. Almost every “Angoff study” run today, and everything called the “modified Angoff”, descends from it rather than from the method in the body of the chapter. A critical review in the Alberta Journal of Educational Research(52(1), 2006, 53–64) puts it plainly: the variation “which many people take to be the original method, has been modified in various ways”.

This is not trivia. The method arrived without the procedural apparatus a defensible study needs — how many judges, how they are briefed, whether they see item statistics, how many rounds — and every one of those decisions has since been shown to move the answer. There is no canonical Angoff. There is a family of procedures, and what makes one defensible is that you wrote down which one you ran.

What the panel is actually estimating — and where it goes wrong

The borderline candidate has to be defined before anyone rates anything. “Just barely qualified” is not self-explanatory, and a panel that has not agreed on it in concrete terms is not rating the same person. Write two or three sentences describing what that candidate can and cannot do on the job, agree them aloud, and keep them visible during the rating.

Judges confuse their own unfamiliarity with item difficulty. This is the best-measured failure mode in the literature and the least-known. Clauser, Hambleton and Baldwin compared Angoff ratings on items judges knew against items they did not, across two datasets, and found unfamiliar items were rated 0.107 to 0.138 lower on the probability scale(p < .001 in both). Excluding the unfamiliar items raised one panel’s passing score by more than ten raw points (Educational and Psychological Measurement77(6), 2016, 901–916). The practical remedy is cheap: check content coverage against your panel’s actual experience before the workshop, not after.

A rating below the chance level is a briefing failure, not a hard item. On a four-option multiple-choice item a candidate who knows nothing scores 25% by guessing. A judge who writes 15% has answered “is this item difficult?” rather than “what share of just-qualified candidates get it right?”. The calculator flags these because they are silently deflating the cut score and they are fixed by one sentence of re-briefing.

Wide disagreement is usually about the item, not the difficulty. When judges are 40 points apart on one item and close on the rest, the common cause is that they read the item differently. That makes it an item-review finding as much as a rating one — and it is the strongest argument for running a round 2 with discussion rather than averaging one round and calling it a study.

How many judges, and what a small panel is allowed to claim

Hurtz and Hertz answered this with a generalizability study on occupational licensing examinations and concluded that approximately 10 to 15 raters is an optimal target range, though fewer are sometimes sufficient (Educational and Psychological Measurement 59(6), 1999, 885–897).

Most panels in employment settings are smaller than that, and a small panel can still set a usable bar. What it cannot do is quote a confidence interval as though it were measured: a standard error computed from four judges is itself so imprecisely estimated that printing it overstates what you know. This calculator therefore reports the mean at any panel size, withholds the interval below five judges, and says why — a refusal is more useful than a reassuring number, because the reassuring number is what ends up in the defensibility file.

One judge sitting far from the panel is a facilitation signal, never a deletion rule. Dropping a dissenting rater tightens the interval and manufactures the agreement it appears to measure, and it is exactly the kind of decision that reads badly if the study is ever examined. Bring the difference into round 2 and let them explain it.

What the law actually says about cut scores

For selection decisions in the United States, the operative sentence is short. The Uniform Guidelines on Employee Selection Procedures say at 29 CFR 1607.5(H): “Where cutoff scores are used, they should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force.”

Read that as a question about the work, not about the score distribution. It is the reason a round number inherited from last year is weak and a documented panel of people who know the job is strong — and the reason the borderline definition, which feels like workshop admin, is the part that connects the cut score to the role.

One clause further up is quietly more consequential and gets missed. Section 1607.5(G) warns that “evidence which may be sufficient to support the use of a selection procedure on a pass/fail (screening) basis may be insufficient to support the use of the same procedure on a ranking basis.” An Angoff study justifies a bar. It does not, on its own, justify ordering everyone above the bar and hiring from the top.

And a cut score that is defensible in method can still be indefensible in effect. The pass rate this calculator returns is a whole-cohort number; the one that matters legally is the pass rate by group. Run it through the four-fifths rule calculator at the cut you actually intend to deploy, before you deploy it — an impact ratio is cheap to check in advance and expensive to discover afterwards.

Where cut scores sit inside AssessAll — including what is missing

The honest starting point is a refusal. AssessAll has no standard-setting workflow. There is no panel tool in the product, no rater grid, no round-2 facilitation, no cut-score study. This page is the only place on the platform where a cut score is derived from evidence rather than typed into a settings field, and it is a calculator on a marketing site, not a feature. If you need a documented Angoff study for a regulated credential, you need a standard-setting consultancy or a credentialing platform built for it.

What the platform does have is worth naming precisely, because the gap between the two is the argument for using this tool at all.

  • The default pass mark is 70, and 70 is a default, not a standard. In the submission route the pass mark is read as settings.passing_score ?? 70: an assessment that never sets the key inherits 70%. Nobody set that by a panel. It is a sensible-looking number that exists so scoring does not crash, and replacing it with a cut score somebody can defend is the entire point of the exercise above.
  • An explicit 0 means no pass mark at all, and is honoured as such. Trait profiles — workstyle, communication style, the premium aptitude profiles — set 0 deliberately because they have no pass concept, and no credential is issued for them. A credential certifies passing, which is meaningless on a profile.
  • Where band boundaries exist, they are design priors and are labelled that way. The English suite maps CEFR bands onto a 10–90 scale and ships role-recommended cuts (53 for a voice process, 45 for chat, 67 for a client-facing team lead) that are editable per product. They are derived from the CEFR mapping, not from an Angoff panel, and a candidate within four scale points below the cut is returned as borderline rather than as a rejection.
  • Borderline results are escalated, not resolved by arithmetic. On AssessAll Certified, a composite within 3 points of the nearest level cut is flagged and sent to human review before anything is issued. That is the contested band from panel 2, implemented as a rule.

What that means in practice, plainly: if you run assessments on AssessAll with the default pass mark, your cut score has no standard-setting evidence behind it. That is true of most cut scores in most hiring systems, and it is fixable in an afternoon with six people who know the job.

Running the study: an eight-step version that survives scrutiny

  1. Write down what the test is for.A cut score is defensible against a stated purpose. “Screening for a voice process at volume” and “certifying readiness to work unsupervised” are different bars on the same instrument.
  2. Define the borderline candidate in two or three concrete sentences drawn from the job, and get the panel to agree them aloud before any rating begins.
  3. Pick 10–15 judges who represent the range of practice. Not only your strongest performers, and not only the item writers. Record who they were and why.
  4. Check content familiarity before the workshop, not after. Judges rate unfamiliar content as harder, by a margin large enough to move a passing score by ten raw points.
  5. Run round 1 independently, with no discussion and no item statistics. Independence is what makes the spread meaningful.
  6. Discuss the items the panel disagrees on, then run round 2. Show the spread; showing empirical item difficulty is a legitimate variant but it changes what the study is, so record which you did.
  7. Compute the cut, its standard error, and the SEM. Then look at how much of your cohort sits inside the contested band, and decide what happens to those people before you know who they are.
  8. Check adverse impact at the cut you intend to deploy, and re-run the whole thing when the instrument changes. A new item set, a new rubric or a new delivery mode moves the distribution against a bar nobody re-examined.

Frequently asked questions

What is the Angoff method?

A panel of subject-matter experts estimates, for every item on a test, the probability that a just-barely-qualified candidate answers it correctly. Each judge's estimates are summed to give that judge's cut score, and the panel's mean of those totals is the cut score for the test. It is criterion-referenced: the bar is fixed to the content, so it does not move when a stronger or weaker cohort turns up.

How many judges does an Angoff panel need?

Hurtz and Hertz ran a generalizability study on this exact question and concluded that approximately 10 to 15 raters is an optimal target range, though fewer are sometimes sufficient (Educational and Psychological Measurement, 59(6), 1999, 885–897). Below about five, the standard error of the cut is itself estimated too imprecisely to be worth quoting, which is why this calculator withholds the interval there rather than printing a reassuring number.

Who should be on the panel?

People who know the work the test is for and who between them represent the range of practice — not only the strongest performers, and not only the people who wrote the items. Familiarity matters more than most panels assume: Clauser, Hambleton and Baldwin found judges rated unfamiliar content as systematically harder, by around 0.11 to 0.14 on the probability scale, which moved one panel's passing score by more than ten raw points.

Is 70% a defensible pass mark?

Not by itself. A round number carries no evidence about what the test measures or what the job requires, and 70% on an easy form is a different standard from 70% on a hard one. The US Uniform Guidelines say at 29 CFR 1607.5(H) that where cutoff scores are used they should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force — which is a question about the work, and the reason standard-setting methods exist.

Should you adjust the cut score by one standard error of measurement?

Both directions are defensible and neither is more scientific than the other. Lowering the bar by one SEM reduces false rejections and raises false acceptances; raising it does the reverse. The SEM tells you how large the adjustment is; the relative cost of a wrong hire versus a wrong rejection tells you which way to move. What is not defensible is adjusting until the pass rate looks right and then describing the result as a standard-setting outcome.

What is the difference between the standard error of the cut score and the standard error of measurement?

They answer different questions and are routinely confused. The standard error of the cut score is the standard deviation of the judges' totals divided by the square root of the number of judges: it says how far the bar would move if a comparable panel had done the exercise instead. The standard error of measurement is the score standard deviation times the square root of one minus reliability: it says how far one candidate's score would move on a retest. A cut score is only a decision rule when both are small next to the decision being made.

Does an Angoff cut score make a selection procedure legally defensible?

It is evidence, not immunity, and it is only one part of the record. The Uniform Guidelines also warn at 29 CFR 1607.5(G) that evidence sufficient to support a selection procedure on a pass/fail basis may be insufficient to support the same procedure used for ranking. A documented panel, a job-analysis basis for the borderline definition, and adverse-impact monitoring at the cut you actually deploy are what make the record. Take advice from an employment lawyer in your jurisdiction.

Does this calculator store the ratings I enter?

No. The whole calculation runs in your browser. Nothing you type is transmitted to AssessAll or to anyone else, there is no signup, and no result is saved.

Sources

  • Angoff, W. H. (1971). Scales, norms, and equivalent scores. In R. L. Thorndike (Ed.), Educational Measurement (2nd ed.). American Council on Education. The probability variation appears in a footnote.
  • Hurtz, G. M., & Hertz, N. R. (1999). How many raters should be used for establishing cutoff scores with the Angoff method? A generalizability theory study. Educational and Psychological Measurement 59(6), 885–897.
  • Clauser, J. C., Hambleton, R. K., & Baldwin, P. (2016). The effect of rating unfamiliar items on Angoff passing scores. Educational and Psychological Measurement 77(6), 901–916.
  • Setting cut-scores: a critical review of the Angoff and modified Angoff methods. Alberta Journal of Educational Research 52(1), 2006, 53–64 — on the footnote provenance.
  • EEOC et al. (1978). Uniform Guidelines on Employee Selection Procedures, 29 CFR Part 1607 — §1607.5(G) on ranking and §1607.5(H) on cutoff scores.
  • AERA, APA & NCME (2014). Standards for Educational and Psychological Testing, chapter 5 (scores, scales, norms, score linking and cut scores).
  • Nunnally, J. C. (1978). Psychometric Theory (2nd ed.), pp. 245–246 — on reliability standards for decisions about individuals.
SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed

This page describes measurement methods and quotes the Uniform Guidelines as published. It is not legal advice, and a cut score is a decision with legal consequences — take advice from a qualified employment lawyer in your jurisdiction before deploying one.

Related reading