Free tools

Differential item functioning calculator

Differential item functioning is present when candidates from two groups who have the same level of the ability being measured still answer an item correctly at different rates. It is not the same as a gap in overall scores, which is impact, and it is not the same as bias, which is a judgement about why. DIF is the difference that survives after you match people on ability.

1. What does the Mantel–Haenszel statistic say?

One row per ability band, and the bands must come from candidates’ total score on the rest of the test— that is what makes this a DIF analysis rather than a pass-rate comparison. Four counts per row: reference group correct and incorrect, then focal group correct and incorrect. Paste straight from a spreadsheet. Nothing you type is sent anywhere; the whole calculation runs in your browser.

MH odds ratio
1.098α̂(MH)
MH D-DIF
−0.22±0.31 (SE)
MH chi-square
0.44p = 0.506
ETS category
Anegligible

Category A — negligible. On the ETS rule this item would not be flagged: the chi-square is not significant at the 5% level. That is the verdict on the headline number alone. Read panel 2 before you act on it.

2. What the single number cannot show you

Mantel–Haenszel estimates one odds ratio pooled across every band. If an item advantages one group at the bottom of the range and the other group at the top, the two halves cancel in the sum and the headline comes back near 1.00. Here is the same data, band by band, which is the picture the pooled statistic cannot draw.

Ability bandSingapore poolMalaysia poolln(OR)Favours
0-3920.0% (n=100)40.0% (n=90)−0.98Malaysia pool
40-5435.0% (n=120)44.4% (n=108)−0.40Malaysia pool
55-6950.0% (n=130)50.0% (n=110)0.00neither
70-8470.0% (n=110)44.4% (n=90)+1.07Singapore pool
85+85.0% (n=80)61.1% (n=54)+1.28Singapore pool

The direction changes across the range. This item favours Malaysia pool in some bands and Singapore pool in others, a spread of 2.26 log-odds units, moving steadily with ability. That is non-uniform DIF, and Mantel–Haenszel is structurally blind to it: the pooled estimate above is an average of two opposite effects. Panel 1 classified this item A — negligible. Panel 1 is wrong about this item, and it is wrong in the way the method is known to be wrong. The remedy is a logistic-regression DIF model with an ability × group interaction term (Swaminathan & Rogers, 1990), which tests uniform and non-uniform DIF separately. Do not conclude the item is clean.

And the number your hiring team is actually looking at

Singapore pool pass rate
50.4%n = 540
Malaysia pool pass rate
46.9%n = 452
Impact (raw gap)
+3.5 ptsunconditioned

Impact is not DIF and DIF is not bias. The gap above compares everyone in one group with everyone in the other, including any real difference in preparation or ability; a group can score lower on an item for reasons that have nothing to do with the item. DIF compares people who scored the same overall, which is why it is the question worth asking about an item. And a DIF flag is still only a hypothesis: it says the item behaves differently, not why. Deciding it is bias requires someone who knows the construct and the population to look at the item and say what, other than the thing you meant to measure, it is picking up.

Why a regional hiring team runs into this first

A shared-services or GBS hub in Singapore or Kuala Lumpur is the clearest case in this region, because it is the structure that guarantees the problem. One role, one assessment, one cut score — and applicant pools from Singapore, Malaysia, the Philippines, Indonesia, Vietnam and India, sitting a test in English that is a first language for some of them and a third for others. Sooner or later someone puts the pass rates by country on a slide, and the meeting has two failure modes: ignore the gap because the test is “standardised”, or drop the test because the gap looks bad. Both skip the only question worth asking, which is whether equally able candidates from the two pools had different odds on a given item.

The standards body that governs cross-population testing is blunt about the order of operations. Guideline SSI-2 of the ITC Guidelines for Translating and Adapting Tests (second edition) says plainly: “Only compare scores across populations when the level of invariance has been established on the scale on which scores are reported.” Guideline C-2 asks for statistical evidence of construct, method and item equivalence for all intended populations. Most regional hiring programmes compare the scores first and never do the second part at all.

The honest caveat: this applies most strongly when the versions differ — different languages, different adaptations. A single English-language instrument sat by everyone is a weaker case for formal invariance work than a translated one, and the guideline is written for adaptation. It is still the right question, because a common language is not the same as a common construct: an item can carry an idiom, a workplace convention, a date format or a unit that is transparent in one market and opaque in the next while the underlying skill is identical.

The ETS rule, and the two conditions that get dropped

The A/B/C classification is the most widely used DIF flagging rule in the world, and the version that circulates most widely is a simplification of it. The short form says: negligible below 1.0 delta units, moderate from 1.0 to 1.5, large at 1.5 and above. The actual rule, quoted from Rebecca Zwick’s ETS review of the procedures, carries two significance conditions as well:

  • “An A item is one in which either the Mantel-Haenszel (MH) chi-square statistic is not significant at the 5% level orMH D-DIF is smaller than 1 in absolute value.”
  • “In order to qualify as a C item, the MH D-DIF statistic must be significantly greater than 1 in absolute value at the 5% leveland must have an absolute value of 1.5 or more” — the test being (|MH D-DIF| − 1) / SE > 1.645.
  • Items meeting neither definition are B items.

Dropping those conditions is not a rounding error, and it fails in the direction that costs you most. An effect-size-only rule will hand out a C on a noisy sample where ETS would return an A, which means an item gets pulled from a bank on evidence that does not support pulling it. This calculator applies the significance conditions, and shows the chi-square and the standard error next to the delta so you can see how much evidence sits behind the letter.

The delta scale itself is worth a sentence, because it is the part people mis-sign. MH D-DIF = −2.35 × ln(α̂MH), so a negativevalue means the item is harder for the focal group among candidates of equal ability. One delta unit corresponds to an odds ratio of 1.530, which is the arithmetic ETS publishes and the check this tool’s implementation is verified against.

Why the tool refuses to give you a letter on a small sample

Because the numbers required are much larger than a single hiring round produces, and a letter printed on 60 people is a decision dressed up as a statistic. ETS’s operational minimum for flagging at the test assembly phase is at least 200 members in the smaller group and at least 500 in total, rising to 300 and 700 at the preliminary item analysis phase. The ITC guidelines arrive at a comparable figure from a different direction: studies to identify potentially biased test items require a minimum of 200 persons per version.

Below those thresholds the calculator still computes and still shows everything — the odds ratio, the delta, the chi-square, the per-band table — and withholds only the A/B/C label. The remedy it points you at is the one that actually works: pool sittings across intakes until the smaller group reaches 200, rather than lowering the bar and flagging anyway. If you run 300 candidates a quarter through a regional screen, that is a three-quarter job, and knowing it is a three-quarter job is more useful than a letter you cannot defend.

One further caveat the tool cannot apply for you. The matching variable — the total score — contains the item being studied, which pulls the estimate towards finding nothing. The standard remedy is purification: run the analysis once, remove the flagged items from the matching score, and run it again. If several items flag on the first pass, do that before you act on any of them.

Where this sits in Singapore and Malaysia law, as at 15 September 2026

Singapore. The Workplace Fairness Act was passed by Parliament on 8 January 2025 and, per TAFEP, is “slated to take effect in end-2027”. It is not in force today. What makes it relevant to this page is which characteristics it covers: alongside age, sex and marital status, race, religion and disability, the protected list includes nationality and language ability — the two attributes a regional English screen is most likely to be sorting on without meaning to. Until then, the Tripartite Guidelines on Fair Employment Practices and the Fair Consideration Framework are the governing instruments, and they are guidelines rather than statute.

Malaysia. Section 69F of the Employment Act 1955, inserted by the Employment (Amendment) Act 2022 (Act A1651), which came into operation on 1 January 2023 after its original September 2022 date was deferred, provides that the Director General “may inquire into and decide any dispute between an employee and his employer in respect of any matter relating to discrimination in employment” and may make an order. Failure to comply with that order is an offence carrying a fine of up to RM50,000, and up to RM1,000 a day for a continuing offence. There is no codified statistical threshold — no four-fifths rule, no DIF standard — so item-level evidence functions as internal governance and as the documentation you would want to have, not as a compliance test.

This is not legal advice, it is a dated summary of provisions read at source on the verification date below, and neither country’s law requires a DIF analysis today. Take advice in the jurisdiction you hire in before you rely on any of it.

When to use something else instead

Checked on 15 September 2026. This tool is a screen for one item from counts you already have in a spreadsheet; it is not a psychometrics suite, and for a real DIF study it is the wrong instrument.

  • difR, in R.The complete answer, and the one to use for an actual study. Thirty-plus DIF methods including Mantel–Haenszel, Lord’s chi-square, logistic regression and generalised logistic regression, SIBTEST, standardisation and lasso-based detection; polytomous items; and built-in item purification through its anchor-item argument. Choose it instead whenever you have R, a full response matrix, and more than one item to examine — which is to say, whenever you are doing this properly.
  • MetricGate’s Mantel–Haenszel DIF calculator. A browser tool that takes a raw dataset — item response column, group indicator, total score — does the stratification for you, and returns α̂MH, ΔMH, a chi-square, an ETS classification, optional per-stratum 2×2 tables and a direction indicator. Choose it instead if what you have is a response-level export rather than a cross-tab and you would rather not decide your own ability bands. Its published statement of the ETS rule is the effect-size short form; ours applies the significance conditions quoted above, which is the difference worth knowing about when a result sits near a boundary.
  • Winsteps, and Rasch software generally. If your programme is already Rasch-calibrated, the DIF contrast in your existing software is a better fit than any free tool, because it uses the calibrated ability estimate as the matching variable rather than a raw total score. Choose it instead if you have a calibrated bank.

What this page adds to that list is narrow and deliberate: the sample-size gate applied as a refusal rather than a footnote, the crossing check that the pooled statistic cannot perform on itself, and impact, DIF and bias set out as three separate questions on one screen, for a reader who is hiring rather than publishing.

What AssessAll does and does not do here

AssessAll does not currently ship a DIF analysis, by country, by first language or by anything else. Saying so is more useful than implying otherwise, because the gap is specific and you can ask every other vendor the same question.

What does exist, read out of the codebase rather than out of a brochure: an opt-in adverse-impact monitorthat computes each group’s selection rate and applies the four-fifths rule, with groups smaller than 5 shown but excluded from flagging and individual answers never surfaced — and its dimensions are gender and age band only. There is no nationality dimension and no first-language dimension, which is precisely the pair Singapore’s Workplace Fairness Act will protect. Separately, an IRT calibration pipeline (2PL and Samejima’s graded response model) exists behind a gate of 200 responses per item, with adaptive delivery stopping at a standard error of 0.30 — but calibration is not DIF.

What AssessAll does do for a regional cohort today is the part upstream of the statistics: one assessment delivered across markets by share link or QR code with no participant accounts, so the same form reaches every pool on the same terms; English-for-work and skills-first screening reported as ranges rather than single points; and credit-based pricing in USD, so a hub running one screen across six countries buys one pool of credits rather than six contracts.

What to do when an item flags

  1. Check the bands before you check the item. Thin strata and a matching score that includes the studied item account for a large share of results that do not replicate. Purify and re-run.
  2. Look at the per-band direction. If the sign changes across the range, the pooled statistic is an average of two opposite effects and the letter it produced is not informative either way.
  3. Read the item with someone from the focal group. This is the step the statistics exist to trigger and the one most often skipped. What are you asking that is transparent in one market and not the next — an idiom, a date format, a unit, a workplace convention, an assumed piece of local context?
  4. Decide whether the difference is part of the construct. If the item is harder for a group because of something you genuinely meant to measure, it is doing its job. If it is harder because of something you did not mean to measure, it is not.
  5. Write down the decision and why.The record of a review that concluded “keep” is worth as much as the one that concluded “drop”, and it is the only thing that shows the check was run before anyone asked.

Questions people ask

What is differential item functioning?
Differential item functioning, or DIF, is present when candidates from two different groups who have the same level of the ability being measured still have different probabilities of answering a particular item correctly. The words that matter are 'the same level of the ability'. A group scoring lower overall is not DIF; DIF is a difference that survives after you match people on how they did on the rest of the test.
What is the difference between DIF, impact and bias?
Impact is the raw difference between two groups' scores, with nothing controlled for. DIF is the difference that remains once candidates are matched on the ability being measured. Bias is a judgement about why. Impact can exist with no DIF at all, because groups can genuinely differ on the thing you meant to measure. DIF can exist with no bias, because an item can be harder for a group for a reason that is part of the construct. Only a review of the item's content by someone who knows the construct and the population can move a DIF flag to a finding of bias.
How do you calculate the Mantel-Haenszel DIF statistic?
Sort candidates into ability bands using their total score on the rest of the test. In each band, build a 2x2 table of group against right or wrong. The common odds ratio is the sum across bands of (reference correct x focal incorrect / band total), divided by the sum of (reference incorrect x focal correct / band total). ETS then reports it on the delta scale as MH D-DIF = -2.35 x ln(alpha), where a negative value means the item is harder for the focal group among candidates of equal ability.
What are the ETS A, B and C categories for DIF?
Quoting the ETS review by Zwick (2012): an A item is one in which either the Mantel-Haenszel chi-square statistic is not significant at the 5% level or MH D-DIF is smaller than 1 in absolute value. To qualify as a C item, MH D-DIF must be significantly greater than 1 in absolute value at the 5% level — the test is (|MH D-DIF| - 1) / SE > 1.645 — and must have an absolute value of 1.5 or more. Items that meet neither definition are B items. The two significance conditions are part of the rule and are frequently dropped when it is restated.
How many candidates do you need for a DIF analysis?
More than most hiring teams have on one intake. ETS's operational minimum for flagging at the test assembly phase is at least 200 members in the smaller group and at least 500 in total, rising to 300 and 700 at the preliminary item analysis phase. The International Test Commission's test adaptation guidelines give a comparable figure: studies to identify potentially biased test items require a minimum of 200 persons per version. Below those numbers a DIF statistic is a screening signal, not a flag.
Can Mantel-Haenszel detect non-uniform DIF?
No, and this is its main limitation. Mantel-Haenszel estimates a single common odds ratio pooled across all ability bands. When an item favours one group at low ability and the other group at high ability, the two effects cancel in the sum and the pooled statistic can return a value close to 1.00 — classified as negligible — for an item that is behaving very differently at the two ends of the range. Inspect the per-band odds ratios, and use a logistic regression model with an ability-by-group interaction term when the direction changes.
Does this calculator store the numbers I enter?
No. The entire calculation runs in your browser. Nothing you type is transmitted to AssessAll or to anyone else, there is no signup, and no result is saved.
Does AssessAll run DIF analysis on its own assessments?
Not today, and the page says so rather than implying otherwise. AssessAll ships an opt-in adverse-impact monitor that computes selection rates and four-fifths ratios over self-declared gender and age band only — there is no nationality or first-language dimension — and an IRT calibration pipeline gated at 200 responses per item. Neither of those is a Mantel-Haenszel DIF analysis by country or language group. This calculator is published because the arithmetic is useful whoever runs the assessment, not as evidence of a feature.
SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed

Statutory and standards claims on this page were read at source on 2026-09-15. This is general information about a measurement method, not legal advice — neither Singapore nor Malaysia requires a DIF analysis today, and you should take advice on your own obligations from a qualified employment lawyer in the jurisdiction you hire in.