Start with our own copy, because it is the clearest example of the error. The product note for AssessAll's Abstract & Pattern Reasoning form says performance "depends on extracting rules, not on schooling or vocabulary, which also makes this the fairest test across backgrounds". The first half of that sentence is a description of the item format and it is accurate. The second half is a claim about group differences in scores, and nothing in the first half establishes it. That sentence is being raised with the product owner rather than quietly edited, because a public page correcting the vendor's own catalogue is worth more than a catalogue that has been silently tidied.
The lane says the same thing in softer words, checked at source on 24 September 2026. Testlify's abstract reasoning test page publishes both numbers a buyer needs — 12 questions, 15 minutes, which is more than most of this category manages — and then says the test will "provide a fair and unbiased measure of a candidate's cognitive abilities, reducing selection bias and ensuring a more equitable hiring practice". It publishes no reliability coefficient, no standard error of measurement and no norm group. TestInvite's employer-facing page on pre-employment abstract reasoning says the format reveals ability "without depending on prior knowledge or expertise" and "without relying on verbal or numerical cues", and publishes no duration, no item count and no psychometric properties at all. Neither vendor is doing anything improper. Both are describing the item medium and inviting the reader to hear a conclusion about fairness.
The meta-analysis the claim is usually hung on does not contain it. Roth, BeVier, Bobko, Switzer and Tyler (2001), in Personnel Psychology, is the standard reference for subgroup differences on cognitive measures in employment. It reports a Black-White standardised difference of .99 for measures of g in industrial applicant samples (N = 6,169 across 8 studies) and .41 in incumbent samples, the gap between the two reflecting range restriction from selection that has already happened. It analyses verbal and mathematical ability separately, at .76 for both, and concludes at p. 321 that "tests assessing g were generally associated with larger differences than verbal or mathematical abilities". What it does not do is analyse figural, abstract, spatial or non-verbal measures as a distinct category at all. A paper that never separated out the format cannot be evidence that the format is fairer.
And the evidence that bears most directly on the claim points the other way. Flynn's 1987 study of IQ gains in fourteen nations reports at p. 171 that "some of the largest gains occur on culturally reduced tests and tests of fluid intelligence". At p. 185 he gives the comparison as a rate: a median gain of 0.588 IQ points per year on culturally reduced tests against 0.374 on verbal tests. On Raven's Progressive Matrices — the archetype of the culture-reduced format, and the design the abstract items in this category descend from — the Netherlands gained 0.667 points per year across thirty years and France 1.005 points per year (Table 15, p. 175), and in the three nations where both test types were measured the culture-reduced gains ran at roughly twice the verbal ones (Table 16, p. 186). The argument is short. A score that rose by around twenty points in one country inside a single generation is measuring something that the environment reaches, and reaching it faster than the verbal tests did. Removing the words from an item removes the words. It does not remove the environment.
None of this makes the format a bad instrument, and it should not be read as one. Pattern-extraction items travel across languages without translation, which is a genuine and useful property for a multi-country pipeline, and the Roth paper's own complexity finding — d of .86 at low job complexity, .72 at moderate and .63 at high, the last from only two studies — says that the differences narrow in exactly the analytical roles these tests are usually bought for. What has to go is the inference in the middle: that because a candidate does not need English to attempt the item, the resulting score is equally fair to everyone who sits it.
Fairness in measurement is not a property a format confers. It is a property you test for, item by item, and the test has a name and a published threshold. Differential item functioning asks whether two candidates of the same underlying ability, drawn from different groups, have different probabilities of getting a particular item right. The classification used in the NAEP technical documentation, which follows the scheme ETS made standard, is explicit about where the lines fall: an item is category A when the Mantel-Haenszel common odds ratio on the delta scale does not differ significantly from 0 at alpha = .05 or is less than 1.0 in absolute value; category C when it is "significantly greater than 1 and larger than 1.5 in absolute magnitude"; and everything in between is category B. Items flagged C are "reviewed by a committee of trained test developers and subject-matter specialists to determine whether the differential functioning of a particular item is due to bias or not" — a human judgement, made on a flagged list, not an automatic deletion.
So the question to put to any vendor selling an abstract reasoning test, ours included, is not whether the items contain words. It is: have you run DIF on this bank, against which groups, at what sample size, and how many items came back category C. AssessAll cannot currently answer that question for any of these three forms. No DIF analysis has been run on them, no reliability coefficient is published for them, and no norm group exists — the research-readiness gate in the codebase is what holds the last of those, and it has not been cleared. That is the row on this page that rules us out, and it is on the page because a page arguing that buyers should demand these numbers is worthless if it exempts the vendor publishing it.