All articles
L&D & Capability14 September 2026·6 min read

Testing Is Not Just Measurement: What 272 Effect Sizes Say About Quizzing After Training

Practice testing beats restudying by g = 0.51 across 272 effect sizes, and spaced retrieval beats massed by g = 0.74 after bias correction. Why the end-of-course quiz is mistimed, what retrieval practice does not do, and a seven-step design for L&D.

By AssessAll Editorial

Retrieval practice is the act of pulling information out of memory — answering a question with the material closed — rather than putting it back in by re-reading. Across 272 independent effect sizes, it beat restudying by g = 0.51. That makes the quiz after training not just a measurement of learning but one of the cheapest interventions that produces it.

Most L&D functions treat the post-course test as a gate: a score, a pass mark, a completion record. The evidence says the test is doing a second job entirely, and that scheduling it once — at the end, at the moment recall is highest and least informative — throws most of that job away.

The evidence, in numbers

The largest synthesis is Adesope, Trevisan and Sundararajan's meta-analysis of practice testing in Review of Educational Research (2017), covering 272 independent effects from 188 experiments. Practice tests outperformed restudying the same material at g = 0.51, and outperformed filler or no activity at g = 0.93.

The mechanism shows up most clearly in Roediger and Karpicke's 2006 experiments in Psychological Science, which is also why the effect is so easy to miss in practice:

  • Five minutes after learning: the restudy group recalled 81%, the tested group 75%. Restudying looks better.
  • One week later: the restudy group recalled 42%, the tested group 56%. The ranking reverses.
  • In their second experiment, study-only fell to 40% at one week while study-then-test-three-times held at 61%.

Read that sequence again from an L&D perspective. If you measure at the end of the session, re-reading wins. If you measure when the capability is actually needed, testing wins by a wide margin. A programme evaluated on a same-day score is systematically biased toward the weaker method.

This is not a fringe finding. In their review of ten common study techniques for Psychological Science in the Public Interest, Dunlosky and colleagues rated only two as high utility: practice testing and distributed practice. Highlighting, summarising and re-reading — the backbone of most corporate courseware — did not make the cut.

Why the money is going the other way

Training spend is rising while contact time is falling. The 2025 Training Industry Report puts total U.S. training expenditure at $102.8 billion, up 4.9%, with average spend per learner climbing to $874 from $774 — while average training hours per learner dropped from 47 to 40.

Fewer hours, more money per hour. That is precisely the condition under which the method you use inside the hour starts to matter more than the number of hours. Spending a fifth of a shrinking window on retrieval instead of exposition is not a sacrifice of coverage; on the one-week numbers above, it is what determines whether coverage survives at all.

What retrieval practice does not do

An article about measurement credibility should be honest about the limits, and they are specific.

Pan and Rickard's meta-analytic review of transfer in Psychological Bulletin (2018) pooled 192 effect sizes from 122 experiments across 67 studies and found transfer of d = 0.40 — real, but conditional. Transfer was strongest when the final test changed format, when questions demanded inference or application, and when retrieval practice was elaborated with explanatory feedback rather than bare right/wrong marking. It was weakest to near zero for material that was studied but never quizzed. The authors also reported evidence of publication bias affecting the overall estimate, though not the moderators.

The practical reading is blunt: you get durability on what you actually retrieve, not on everything in the deck. Quizzing module three does not protect module four. If a skill matters in ninety days, it needs an item.

Format matters too, and not in the direction most practitioners assume. Adesope and colleagues found multiple-choice practice tests carried a larger weighted effect (+0.70) than short-answer ones (+0.48) — but the authors caution against reading that as superiority. Multiple choice is efficient for fact retention; constructed response demands more of the higher-order reasoning most workplace capability actually consists of. The format that produces the bigger number in a memory study is not automatically the format that matches your job task.

Spacing is the multiplier

If there is one design decision worth more than the rest, it is not writing better questions. It is not asking them all at once.

Latimier, Peyre and Ramus's meta-analytic review in Educational Psychology Review (2021) found spaced retrieval practice beat massed retrieval practice at g = 1.01, falling to g = 0.74 after correction for publication bias — a larger effect than retrieval practice itself contributes over restudying. They also compared expanding schedules (gaps that lengthen) against uniform ones and found essentially nothing between them: g = 0.034, not significant.

That second result is the liberating one. The schedule does not have to be clever. It has to exist.

The workplace evidence points the same way. In a randomised trial of urology residents, participants who received the same content as daily interactive multiple-choice emails rather than in a single block still scored higher two years later — 70.2% versus 66.8%, effect size 0.35, p = 0.03. A three-point gap two years after a delivery-format change costing nothing is an unusually good return for L&D.

How to build it

A workable design, in seven steps:

  1. Name what must survive ninety days. Not the syllabus — the ten to fifteen decisions, rules or judgments a learner will be worse at their job without. This is the item blueprint.
  2. Write items to the decision, not the slide. "Which of these is the escalation threshold?" is a fact. "The customer has done X and says Y — what do you do first?" is the job.
  3. Quiz during, not only after. Retrieval inside the session is the intervention. The end-of-course test is the record.
  4. Space three to five passes over six to eight weeks. Uniform gaps are fine. Any spaced schedule beats one sitting.
  5. Give elaborated feedback. The correct answer plus why the attractive wrong answer is wrong. This is the moderator that made transfer work in Pan and Rickard's data.
  6. Mix formats deliberately. Multiple choice for coverage and volume; open-response scenarios where judgment is the construct. AI-graded scenario items make the second affordable at scale rather than only at assessment-centre prices.
  7. Separate low-stakes learning from the high-stakes record. Practice quizzes should be consequence-free — anxiety undermines the effect you are trying to buy. Reserve proctoring and integrity bands for the certification event at the end, where the score has to defend a decision.

Step four is where seat-licence economics quietly kill good design, because four spaced passes cost four times one. Pay-as-you-go pricing — AssessAll runs at ₹30 or US$0.50 a credit — makes a spaced sequence a rounding error against the $874 per learner already being spent.

When this is the wrong tool

Retrieval practice addresses knowledge durability. It does not fix a training problem that was never a knowledge problem. If the failure is psychomotor skill, quizzing will not substitute for supervised practice. If it is attitude or culture, no quiz schedule will move it. If nothing needs to be remembered because a job aid sits on every desk, skip the whole exercise and invest in the job aid. And if learners can recall the content perfectly but still do not apply it, the gap is transfer, not retention — a different diagnosis with a different remedy.

The takeaway

The test at the end of your programme is currently doing one job badly timed. Moved earlier, repeated a few times, and given real feedback, the same items do a second job the rest of the programme is struggling to do — and the evidence base for that claim is four decades deep and unusually consistent.

#retrieval-practice#testing-effect#spaced-learning#training-retention#l-and-d#learning-science

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free