Training Evaluation and Programme Effectiveness Assessment for L&D and HR TeamsA programme defended on how much people enjoyed it has not been defended. Do you know what would defend it?
Thirty-eight original exercises on the judgement behind evaluating a training programme — which level of evidence the decision needs, what a claim of effect rests on, what the workplace does to transfer, how to read the figures, and what to tell the sponsor — with a report that pairs every strength with its shadow and refuses the one figure it cannot yet honestly print.
Every strength has a shadow, and the report prints both
The Training Evaluation and Programme Effectiveness Assessment is a thirty-eight exercise judgement check for people who must say whether a training programme worked. It measures which level of evidence a decision needs, what a claim of effect rests on, how the workplace shapes transfer, how to read an evaluation figure honestly, and what to tell the sponsor.
The four levels of training evaluation — reaction, learning, behaviour on the job, organisational result — are a public structure and the exercises use them by description. Nothing here names a model, a product or a person, and the instrument claims no affiliation with any of them. The two errors it is keyed against are the two that practice makes most often: measuring only reaction, and claiming a business result from a design that cannot support the claim. Both are errors of doing; the wrong options also carry errors of not doing, such as deferring every finding or skipping the learning measure, so no blanket habit scores.
Two programme briefs ask you to split a fixed budget of evaluation effort across the four levels: a mandatory regulatory module for three thousand staff, and a coaching workshop in a contact centre where the behaviour and result figures are already being collected. The briefs were written to need different shapes, and the report draws your shape beside an authored consensus and beside ordinary practice, level by level, so it can be inspected rather than trusted.
That consensus is where the report is most honest. An expert allocation should come from a panel of five or more raters with a reported inter-rater agreement. No panel has sat for this instrument, so the consensus is one author's reading of the published literature, and the report says so in a ledger at the top of the strand: one rater, no agreement computed, match figure refused. The strand carries a three-way placement and both allocations in full. The number a measured agreement would have licensed is not printed, and the page says what would fill the blank.
The number on the page comes from the thirty keyed exercises instead. Every option, binary claim, matching pair, ordering position and slider carries a declared answer prior, and the score is corrected against it, so zero means a respondent who answers the way ordinary evaluation practice answers rather than a coin. The score carries its 68 and 95 per cent bands drawn on the chart, its reliability stated as an assumption until live data exist, and a three-way placement rather than a figure for each of the five parts, because six exercises cannot carry a number.
The report design is the strength/shadow pair. Every part of the subject is printed in two columns: what the strength gets you, and what it costs when it is overused. Reading the numbers carefully is shadowed by the report that is all caveats and never read; insisting on a comparison group is shadowed by the programme that ends up evaluated by nobody. A page that names a cost beside every strength cannot be read as generic praise, and every strength on this page has one.
What you walk away with
Which level of evidence the decision needs, what each measure is actually evidence of, and why an end-of-course test overstates what is kept.
Selection, testing, history and regression to the mean, the designs that remove each, and how to build a wait-list comparison from a rollout that was happening anyway.
Supervisor support, opportunity to perform and peer use, the order of the follow-up supports, and why more course is usually the wrong answer to low use.
A response rate before an average, a comparison before an attribution, an effect size before a percentage, and the completer-versus-dropout gap that is not an effect.
The sentence the evidence supports, the figure it does not, which outcome the design could not test, and how to decline a return figure without being heard as a refusal.
Your two allocations beside the authored consensus and beside ordinary practice, and the ledger that prints why no match figure is claimed and what would fill the blank.
Inside your report
Illustrative sample — your report is generated from your own responses.
Solid bar: you. Outlined bar: the authored consensus. Hatched bar: ordinary practice, read from the declared answer distribution. The shape is what is compared, and both allocations are printed in full so it can be inspected rather than trusted.
| Raters behind the consensus allocation | 1 | an authored reading of the literature; a measured standard needs five or more |
| Inter-rater agreement (ICC) | □ not computed | no panel has sat, so there is no agreement to compute |
| Match figure | □ refused | the figure a measured ICC would have licensed is not printed |
| Printed instead | a three-way placement | and both allocations in full |
| What replaces this | a rated panel | the placement becomes a figure when the panel's ICC is on the page |
An expert allocation needs five or more raters with a reported agreement, or it is one person's opinion. No panel has sat, so the report prints the blank in its own row and refuses the figure rather than rounding one into existence.
Built for
- Learning and development managers who have to defend a programme to a sponsor
- Training providers and learning consultants who write evaluation reports for clients
- HR business partners asked whether a programme should be run again or extended
- Anybody who has been handed a satisfaction score and asked what it proves
Find out what your evaluation judgement gets you, and what it costs
38 exercises across six formats · about 35 minutes · every strength paired with its shadow, both allocations printed in full, and one figure honestly refused.
₹1,199 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
Applied judgement about how to evaluate a training programme: which level of evidence a decision needs, what comparison a claim of effect rests on, what the work environment does to transfer, how to read an evaluation figure without over-claiming, and what an honest report to a sponsor says. It does not measure the ability to run an evaluation, statistical skill, or whether your own programmes work, and it is not a certification.
No. The four levels of training evaluation are a public structure and the exercises use them by description: reaction, learning, behaviour on the job and organisational result. No model, product or author is named anywhere in the instrument or the report, and no affiliation with any of them is claimed.
Because an expert allocation should come from a panel of five or more raters with a reported inter-rater agreement, and no panel has sat. The reference is one author's reading of the published literature. The report prints that in a ledger, gives the allocation strand a three-way placement and both allocations in full, and declines the figure that a measured agreement would have licensed.
Zero is what somebody scores who answers the way ordinary evaluation practice answers: measures reaction, reads a rise after the course as the course's effect, and gives the sponsor the figure. Every exercise carries a declared answer prior and the score is corrected against it, so zero is a typical practitioner rather than a coin toss. The score carries its 68 and 95 per cent bands on the chart.
About thirty-five minutes for thirty-eight exercises across six formats. ₹1,199 in India, inclusive of GST, or US$11.99 elsewhere, one time, for the sitting and the full report.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Thirty-eight original exercises across six formats: ten single-choice judgement items, eight allocation sliders (two programme briefs, each split across the four levels of evaluation), six select-every-that-applies exercises, six binary claims, four matching exercises and four ordering exercises. One response instruction is declared for the whole instrument and it is a KNOWLEDGE instruction: what the published evidence supports for the programme described, never what the respondent would personally do or what is usual in their organisation. Construct statement: this measures applied judgement about how to evaluate a training programme, meaning which level of evidence a decision needs, what comparison a claim of effect rests on, what the work environment does to transfer, how to read an evaluation figure without over-claiming, and what an honest report to a sponsor says; it does not measure the respondent's ability to run an evaluation, their statistical skill, their programme-design ability, whether their own programmes work, or their standing with any sponsor, and it is not a certification in evaluation. The four levels of training evaluation are treated throughout as a public structure and are named by their descriptions, reaction, learning, behaviour on the job and organisational result; no evaluation product, model name or author's name is used anywhere in the instrument, and no affiliation with any such product is claimed or implied. Scoring has two strands that are never combined. The KEYED strand is the thirty keyed exercises: each is scored on the platform's own rule for its type, expressed as a quality from zero to one, and chance-corrected against a declared answer prior authored on every option, claim, matching pair and ordering position, so that zero means a respondent who answers the way ordinary practice answers rather than a uniform guess; the strand carries the headline number with its standard error band, its reliability stated as McDonald's omega estimated from item count and a declared inter-item correlation of .20 until live data replace it, and a three-way placement rather than a figure for each of the five competencies because six items cannot carry a number. The ALLOCATION strand is an expert-vector allocation match: the respondent's eight shares are correlated with a reference allocation, so that the level of effort drops out and only the shape is scored, and the correlation is corrected against a Monte-Carlo null drawn from the declared answer distribution on each slider, in which the reaction-heavy allocation of ordinary practice carries the mass. The reference allocation is an AUTHORED CONSENSUS read off the published evaluation literature. No rater panel has sat, no inter-rater agreement has been computed, and an expert vector without five or more raters and a reported intraclass correlation is one reading of the literature rather than a measured standard; for that reason the allocation strand is refused the match figure that a measured agreement would have licensed, and the report prints a three-way placement, both vectors in full, and the reason, never a percentile and never a precise match score. The two errors the instrument is keyed against are the two that evaluation practice most often makes: measuring only reaction, and claiming an organisational result from a design that cannot support the claim. Both are errors of commission; the distractors also carry errors of omission, such as deferring every finding to a longer follow-up, skipping the learning measure, or leaving the sponsor's question unanswered. Keying sources: Alliger and Janak (1989) on the assumptions of the four-level structure; Alliger, Tannenbaum, Bennett, Traver and Shotland (1997), the meta-analysis of relations among training criteria; Sitzmann, Brown, Casper, Ely and Zimmerman (2008) on reactions and learning; Holton (1996) on the flawed four-level model; Bates (2004) on the limits of evaluation practice; Kraiger, Ford and Salas (1993) on cognitive, skill-based and affective learning outcomes; Baldwin and Ford (1988) and Blume, Ford, Baldwin and Huang (2010) on transfer of training; Burke and Hutchins (2007) on transfer; Ford, Quinones, Sego and Sorra (1992) on opportunity to perform; Gollwitzer and Sheeran (2006) on implementation intentions; Arthur, Bennett, Edens and Bell (2003) on training effectiveness and effect sizes; Salas, Tannenbaum, Kraiger and Smith-Jentsch (2012) on the science of training; Shadish, Cook and Campbell (2002) on experimental and quasi-experimental designs; Sackett and Mullen (1993) on evaluation beyond formal experimental design; Cohen (1988) on effect-size conventions; Rossi, Lipsey and Freeman (2004) on evaluation practice; and the recurring practitioner surveys of evaluation practice, which report reaction measured on nearly every programme and organisational results on a small minority, and which are the source of the reaction-heavy answer prior. Every exercise is an original work written for this instrument; none reproduces an item from any commercial instrument, and no affiliation with any evaluation product, framework owner or certification body is claimed.