Test adaptation is the work of making an assessment usable and comparable in a new language or market. It is not translation. The construct, the administration method, the individual items and the score scale each need separate evidence of equivalence. Until that evidence exists, a score from your Manila version and a score from your Warsaw version cannot be placed on the same ranking.
Most multi-country assessment programmes skip straight from translation to a shared leaderboard. That single move is where comparability quietly dies — and because the resulting scores still look like numbers, nobody notices.
Translation is one of eighteen steps
The ITC Guidelines for Translating and Adapting Tests, Second Edition (International Test Commission, 2017) set out 18 guidelines across six categories:
- Pre-Condition — 3 guidelines
- Test Development — 5 guidelines
- Confirmation — 4 guidelines
- Administration — 2 guidelines
- Score Scales and Interpretation — 2 guidelines
- Documentation — 2 guidelines
Translation sits inside Test Development. Fifteen of the eighteen guidelines are about something other than wording — whether the construct exists in the target culture, whether the format is familiar, whether the items behave the same way, and what you are permitted to say about the resulting scores.
The guidelines are also pointed about the most popular shortcut. Back-translation is widely treated as proof of a good translation, but the ITC warns that a poor, overly literal translation can back-translate almost perfectly — precisely because it did not attempt meaning. The recommended design is double-translation and reconciliation by a panel that knows both cultures, followed by judgment-based review against rating criteria.
Three kinds of equivalence, not one
Guideline 10 asks for statistical evidence of three distinct things, and conflating them is the most common error in cross-border programmes.
Construct equivalence — does the thing you are measuring exist, and mean the same thing, in the target market? "Taking personal initiative" or "challenging your manager's decision" are not culturally neutral behaviours. If the construct does not overlap sufficiently, no amount of translation quality will rescue the comparison.
Method equivalence — are candidates equally familiar with the format, the device, the time pressure and the response scale? Response-style differences across cultures (extreme responding, midpoint avoidance) are method effects, not ability.
Item equivalence — do individual items function the same way for equally able candidates in each group? This is differential item functioning (DIF), and it is measurable with IRT, Mantel–Haenszel, logistic regression or restricted factor analysis.
What DIF looks like at international scale
The best-documented evidence comes from large international surveys. Gökçe, Berberoğlu, Wells and Sireci (2021), analysing TIMSS 2015 — 57 countries and 43 languages — found that DIF increased as countries became more distant in culture and language family. Their key result is more specific than "translation causes DIF": the magnitude of DIF was greatest when both language and country differed, and smallest when the language was the same but the country differed.
That has a practical read-out for employers hiring across India, the Philippines, the UAE and Singapore on one English-language instrument. Same-language, different-country is the lowest-risk configuration — not risk-free, but the comparison most likely to survive scrutiny. The configuration that demands real psychometric work is the one where you have commissioned local-language versions.
The invariance ladder decides what you are allowed to report
Measurement invariance is tested in steps, and each step unlocks a different claim:
- Configural — the same items load on the same factors. You may claim the test measures the same structure.
- Metric (weak) — loadings are equal across groups. You may compare relationships and correlations.
- Scalar (strong) — intercepts are also equal. Only now may you compare observed means or put candidates on one ranked scale.
- Strict — residual variances are equal too. Rarely achieved, rarely required.
ITC Guideline 16 states the rule plainly: compare scores across populations only when invariance has been established on the scale on which scores are reported. A shared percentile table is a scalar-level claim, whether or not anyone tested for it.
Full scalar invariance is uncommon in real data, and the honest standard is partial invariance with a documented proportion. A four-country study of job insecurity across the United States (n = 486), China (n = 629), Italy (n = 482) and South Africa (n = 345) found that full metric invariance was not supported; non-invariant parameters ranged from 7.6% to 19% across samples, below the roughly one-third threshold conventionally treated as tolerable, and the authors judged cross-country interpretation permissible on that basis. Note what makes that defensible: a number, a threshold stated in advance, and a published decision — not an assumption.
Adaptation does work, when the content is construct-driven
The cross-cultural literature is not uniformly discouraging. A German situational judgment test of personal initiative was translated and administered in Cuba, with 192 Cuban and 213 German participants across 11 analysed items; the authors reported similar measurement invariance, construct-related validity and reliability in both groups (McDonald's ω .61–.74), with no items requiring cultural adaptation beyond translation. Cuban participants rated the test slightly more favourably than German participants did.
The authors attribute this to the SJT being construct-driven — built around one defined behavioural construct rather than a broad sweep of job-specific scenarios — and caution that broad-skill SJTs may not transfer as cleanly. That is the design lever worth taking seriously: scenario content anchored to an explicit behavioural construct, with the local scenery kept thin, travels across borders far better than richly contextualised cases that encode one country's office norms. AssessAll's AI-graded scenarios are built this way for the same reason, and market-specific versions go through local adaptation rather than straight translation.
English is not a constant either
Running one English instrument everywhere avoids translation DIF but introduces a different confound. The EF English Proficiency Index 2025, drawn from 2.2 million EF SET test takers across 123 countries and regions, reports that speaking remains the weakest skill in more than half the countries measured. Where English is a second language, reading load becomes part of what a numerical-reasoning or judgment item measures. If the construct is not English, the language demand is construct-irrelevant variance and should be minimised deliberately — shorter stems, plainer syntax, no idiom.
Six checkpoints before you compare scores across markets
- Define the construct in writing, then have bilingual, bicultural subject-matter experts confirm the overlap is sufficient for your intended use. Do this before commissioning any translation.
- Double-translate and reconcile. Treat back-translation as one weak check among several, not as sign-off.
- Pilot in-market and inspect item difficulty, timing and completion drift against the source version.
- Run DIF item by item, and act on it — revise or drop flagged items rather than noting them.
- Test the invariance ladder and publish the level reached plus the percentage of non-invariant parameters. State your tolerance threshold in advance.
- Choose the score scale to match the evidence. One common scale only at scalar invariance. Otherwise: local norms, within-market ranking, or locally set criterion cut scores.
Then document all of it. Guideline 17 and 18 exist because an adaptation nobody can audit is indistinguishable from a guess — and under the EU AI Act, employment and worker-management AI systems are classified high-risk in Annex III, while the US EEOC's guidance on employment tests and selection procedures expects job-relatedness to be demonstrable. The Standards for Educational and Psychological Testing make the same demand on fairness grounds regardless of jurisdiction.
When one scale is the wrong goal
Sometimes the right answer is to stop trying. If you are filling 40 seats in Kuala Lumpur and 40 in Kraków, you do not need a global ranking — you need two defensible local decisions. Criterion-referenced cut scores set separately in each market, against the same documented construct, are usually easier to justify than a pooled percentile that quietly assumes scalar invariance nobody tested. A portable record such as a Skill Passport carries the score plus the evidence behind it, which is what lets a receiving employer judge how far the number travels.
Takeaway
A translated test is a new test until the evidence says otherwise. Decide early whether you need one comparable scale or several defensible local ones — that choice determines how much psychometric work you owe, and pretending it does not is how measurement programmes lose their credibility in the markets where they matter most.