Level 3 evaluation measures whether trained people behave differently on the job, usually 60 to 120 days after a programme ends — not whether they enjoyed the course or learned the content. Measuring it credibly needs three things: a defined target behaviour, a rating source other than the trainee, and a baseline taken before training.
Most L&D functions have none of the three. This walkthrough follows the decision one capability team faced when their CFO stopped accepting completion rates as evidence. The organisation described below is an illustrative composite built to show the reasoning — not a client, and not real data.
The number that taught L&D to give up
Ask why behaviour change goes unmeasured and someone will eventually say that only 10% of training transfers to the job anyway, so why bother. That figure has circulated for four decades. People who have chased the citation trail find it terminates in a 1982 Training and Development Journal piece by D. L. Georgenson, where the number appears inside a rhetorical question rather than as a result. No study sits behind it.
The actual evidence is considerably more encouraging. Arthur, Bennett, Edens and Bell's meta-analysis of 397 effect sizes in the Journal of Applied Psychology found a mean effect of d = 0.62 on behavioural criteria across 122 effect sizes and 15,627 participants — statistically indistinguishable from the effect on learning criteria (d = 0.63). Training moves behaviour. What collapses is not transfer; it is the measurement of transfer.
The same paper documents the mismatch precisely. Reaction measures accounted for just 4% of the effect sizes in the published literature, while 78% of the organisations in the authors' benchmarking sample said they used them.
The situation
A 900-person services firm promotes roughly 90 individual contributors into first-line manager roles each year and runs them through a three-day programme. Cost per cohort is real but not enormous. At budget review the CFO asks a fair question: what changed?
The team put three measurement designs on the table.
Design A: completion and reactions
The default. A post-course survey, an attendance record, a satisfaction score.
It costs almost nothing and answers almost nothing. It establishes that the programme ran and that participants did not resent it. It cannot distinguish a well-liked programme from an effective one — and the correlation between the two is weak enough that the industry has a nickname for the instrument.
Practice still leans on it. The TalentLMS 2026 L&D report, based on September 2025 surveys of 101 US HR managers and 1,000 employees, found completion rates to be the most common success metric, with 37% citing business impact, 31% career growth and 28% training satisfaction. Notably, only 20% named difficulty measuring learning ROI as a challenge — a level of confidence that sits awkwardly beside what is actually being measured.
Design B: self-reported behaviour change at 90 days
Send participants a survey three months later asking how often they now hold one-to-ones, give corrective feedback, or delegate. Cheap, fast, and it feels like Level 3.
It isn't. Blume, Ford, Baldwin and Huang's meta-analytic review of transfer in the Journal of Management — 89 studies, 93 independent samples, N = 12,496 — reports that self-ratings and supervisor ratings of the same person's transfer, taken at the same time, correlate only .28. The two sources are describing substantially different things.
Worse, when the trainee rates both the predictor and the outcome, relationships inflate sharply. The correlation between work environment and transfer was .54 under same-source measurement and .23 without it. A self-report design does not just add noise; it manufactures a finding.
Design C: two-source behavioural ratings plus scenario reassessment
Five or six observable behaviours, defined concretely, rated on behaviourally anchored scales by the participant's own manager and by two direct reports, at baseline and again at 90 days — plus a scenario-based reassessment scored against the same rubric used before the programme.
More expensive per head. Also the only one of the three that produces a defensible answer.
Comparing the three
| | Design A | Design B | Design C | |---|---|---|---| | What it evidences | The programme ran | Perceived change | Observed change vs baseline | | Rating source | Trainee | Trainee | Supervisor + reports + scenario | | Baseline required | No | No | Yes | | Known failure mode | Reaction ≠ effect | Same-source inflation | Cost; rater burden | | When it is the right choice | Pilot logistics; vendor triage | Fast pulse on perceived relevance | Any programme whose budget is being defended |
When Design A is genuinely right: a first-run pilot where the question is whether the content and logistics hold up at all. When Design B is right: as a low-cost supplement alongside a harder measure, never as the headline.
What the team chose
They ran Design C on a stratified sample of about 30 of the 90 managers, and Design A on everyone.
That is the move most teams miss. Level 3 evidence does not require a census. A sample large enough to detect a moderate effect, measured properly, beats a full population measured by self-report — and it costs less. Sampling also lets you spend the rating burden where it buys the most: on the supervisors whose judgements will be scrutinised.
For the scenario component they used a fixed rubric applied identically before and after, so the two scores were comparable. AI-graded scenarios of the kind AssessAll runs are useful here specifically because the grader does not drift between January and April the way a human panel does; pay-as-you-go credits at ₹30 / US$0.50 make a two-point design cost roughly what a single seat licence would.
Then design for transfer, not just measurement
Measuring behaviour change without engineering it is a good way to document failure. Taylor, Russ-Eft and Chan's meta-analysis of 117 behaviour modelling studies found that transfer improved when mixed positive-and-negative models were shown, when practice used trainee-generated scenarios, when trainees were instructed to set goals, when their superiors were also trained, and when rewards and sanctions existed in the work environment. Effects on declarative knowledge decayed over time; effects on skills and job behaviour held steady or grew.
The environment carries most of the load. In Blume et al., supervisor support correlated .31 with transfer and peer support only .14. The moderator that matters most for management training: environmental effects were far stronger for open skills like leadership (.26) than for closed skills like software use (.04). A CRM rollout can survive an indifferent manager. A delegation programme cannot.
Salas, Tannenbaum, Kraiger and Smith-Jentsch make the point plainly in *Psychological Science in the Public Interest*: what happens before and after training matters as much as what happens during it.
A working checklist
- Define the behaviours before the programme is designed. Five to seven, observable, at the level a supervisor could describe an instance.
- Take a baseline. Without it, a 90-day rating is an opinion.
- Rate from at least two non-trainee sources. Self-report may accompany them; it may not replace them.
- Sample, don't census. Spend the budget on measurement quality, not coverage.
- Fix the rubric across both time points. A drifting standard invents change.
- Train the supervisors too, and require goal-setting at course close. Both are evidence-backed transfer levers, not nice-to-haves.
- Wait 60 to 120 days. Earlier is recall; much later confounds with everything else that happened.
- Report the design alongside the number. A finance audience trusts a modest effect with a stated method over a large one with none.
The takeaway
The reason most training looks ineffective is not that it fails to transfer — the meta-analytic evidence says it transfers about as well as it teaches. It is that the measurement most organisations run cannot detect transfer even when it happens. Fix the design before you conclude anything about the programme.