The decision: proving a programme worked

The levy paid for the training. What proves it changed anything?

Proving a training programme worked is the decision of what evidence would change your mind about running it again. It is answered by measuring the same capabilities on the same people before and after the programme, reporting movement only for participants who completed both ends, and publishing the drop-out alongside the gain rather than behind it.

At a glance

Proving a Training Programme Worked: the facts, with their units

The decisionDecide what evidence would make you cancel, redesign or repeat this programme — then collect that evidence, not the evidence that is easiest to collect.
What attendance provesThat the course happened. An HRD Corp claim is evidenced by an attendance report, an invoice, a T3 form and a JD/14 form — none of which is evidence of effect (hrdcorp.gov.my claim process flow, checked 3 September 2026).
The unit of evidenceA matched pair: one participant with both a baseline and an outcome attempt. Gain is computed on pairs only; unmatched attempts are counted and shown, never averaged in.
When we refuse to reportBelow 3 matched pairs the platform withholds effect sizes and says so on the report — a standard deviation on two people is noise, not evidence.
Cost of the measurement layerCredits are US$0.50 each and a baseline plus an outcome is two sittings. Packs of 250 / 1,000 / 5,000 cost US$125 / 465 / 2,175, discounting volume up to 13%.
Cost to pilot250 free credits on signup — enough to baseline a small cohort, run the programme, re-measure and read the impact report before paying anything.
SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed

Completion is the metric because completion is the metric that exists

Attendance is captured automatically, so it becomes the report. It answers a funding question — did the training take place, and can we claim for it — and it is the correct evidence for that question. It is simply not evidence about capability, and a board that asked whether people can now do the thing has not been answered.

The happy sheet measures the trainer, not the training

A reaction form at the end of the last session records how the room felt about the experience. Reaction is the first of Donald Kirkpatrick's four levels — reaction, learning, behaviour, results — and the model's own point is that the first level does not predict the ones underneath it. An enjoyable course and an effective course are different findings and most programmes only ever collect the first.

The post-only test cannot tell learning from selection

Testing after the programme tells you where the cohort ended up. It cannot tell you what the programme contributed, because you never measured where they started — and the people who show up on the last day are not a random sample of the people who showed up on the first.

Nobody wants to be the one who measures it properly

A paired design can return a null result, and a null result is a difficult meeting. That is precisely why it is worth something: a measurement that could only ever have produced good news was never a measurement. The compensating fact is that a null result on one competency is usually accompanied by a real gain on another, which is a redesign brief rather than a verdict.

What you get

Built for proving a training programme worked

Learning Journeys — baseline, during, outcome

A journey sequences the assessments around a programme in stages: a baseline before it starts, optional checkpoints during, and an outcome wave after. The same competencies are carried through, so the comparison is between two readings of one thing rather than between two different tests.

Gain computed on matched pairs, by design

A participant contributes to a gain figure only when they have both a baseline and an outcome attempt. Unmatched attempts are counted and reported separately — outcome-only participants who cannot show gain, and baseline-only participants who started and never finished — so drop-out is visible in the report rather than hidden inside an average.

Effect sizes, and the refusal to report one

The impact report carries a paired Cohen's d alongside the raw gain, so movement is expressed against how much participants varied rather than in bare percentage points. Below three matched pairs the effect size is withheld and the report states why, rather than printing a number that a reader would over-trust.

Per-competency movement, not one headline number

Each competency the instrument measures gets its own reading: baseline mean, outcome mean, gain, effect size, how many participants below the pass threshold at baseline crossed it, and how many declined. A programme that moved one competency and not another is a design finding, and a single average would erase it.

Non-improvers named at participant level

The report flags every participant whose score did not improve, with their baseline, outcome and checkpoint completion beside it. That list is the follow-up action; a cohort average is not something you can do anything about on Monday.

Cost figures marked as your inputs, not our findings

Supply a programme cost and headcount and the report returns cost per participant and cost per point of gain — with a note stating in plain words that those figures are your inputs and are not measured by the platform. Cost per point of gain is withheld entirely when the gain is zero or negative, because a programme that did not move the needle has no meaningful cost per point.

360° cycles that re-measure against the earlier one

For behaviour rather than knowledge, a follow-up 360° cycle is linked to its predecessor and the reports show change against it — the same paired logic applied to observed behaviour, for programmes whose claim is about how someone leads rather than what they know.

How the attribution was decided, on the report

Every attempt is tied to a stage by an explicit, ranked rule — its own share link where one exists, otherwise the assessment sitting in exactly one stage, otherwise the participant's own chronology with the earliest attempt as baseline. Which rule fired, and for how many attempts, is printed in the method note so a reader can judge the evidence rather than trust it.

How the decision is made

The sequence, in order

  1. 1

    Write down the result that would make you stop

    Before anyone is enrolled, write the number that would mean this programme is not worth repeating — no movement on the two competencies it targets, or fewer than half the people below the bar crossing it. A programme with no failing condition cannot be evaluated, only described, and a year later the only honest report available is a completion percentage.

  2. 2

    Baseline the same competencies you intend to claim

    Measure before the first session, on the competencies the programme says it will move — not adjacent ones, and not a self-rating survey, which measures confidence and moves in the opposite direction to competence in exactly the people who need the training most. The baseline is also your enrolment filter: people already above the bar do not need the course.

  3. 3

    Re-measure the same people on the same instrument

    After the programme, run the same competencies again. Same items or a matched form, same conditions, same scoring. The comparison only means something because nothing except the programme changed between the two readings — which is also why a redesigned post-test quietly destroys the evidence it was meant to produce.

  4. 4

    Report the gain on matched pairs only

    A participant counts toward a gain figure only if they completed both ends. Comparing the average of everyone who sat the pre-test to the average of everyone who sat the post-test is the standard way an ineffective programme produces an impressive number, because the people who struggled are the people who did not come back.

  5. 5

    Publish the drop-out next to the gain

    Say how many started the baseline and never returned. If that number is large next to the number of matched pairs, the headline gain is flattered by survivorship and the report has to say so. This is the step no vendor enjoys writing and it is the one that makes the other four believable.

  6. 6

    Name the people it did not work for, and act on them

    Every cohort contains non-improvers — participants whose score did not move or fell. They are not a rounding error; they are the finding. Either the programme did not reach them, or they needed something else, and both answers change what you run next quarter.

  7. 7

    Re-run the reading one quarter later

    A gain measured on the last day of a course measures learning, not retention. Re-measuring the same competencies a quarter on is the cheapest way to find out whether anything survived contact with the job, and it is the reading a board actually asked for when it asked whether the training worked.

Compare the approaches

Four kinds of evidence that a programme worked, compared

Approach-level rather than vendor-level: the choice in front of an L&D lead is usually between these four, and three of them are the right answer some of the time. Every row carries the case for choosing it instead.

External facts verified at source on .

ApproachWhat it measuresWhat it costsWhen to choose it instead
Attendance and completion recordsThat the training happened, to whom, and for how long. In Malaysia this is the evidence an HRD Corp claim actually runs on: an attendance report generated by the training platform showing each trainee's log-in and log-out, an invoice addressed to HRD Corp, a T3 form signed by trainees and a JD/14 form completed by the training provider, submitted within six months of completion.Effectively nothing — it is a by-product of delivery and of the claim you were going to file anyway.Choose it when the question is a funding or audit question, because for that question it is not merely adequate, it is the correct and required evidence. The mistake is only ever in what it is then asked to prove: "HRD Corp claimable" is a statement about funding eligibility, not a finding about whether the course changed anything.
End-of-course reaction surveyHow participants felt about the session, the trainer and the materials — Kirkpatrick's first level.A few minutes per participant, and it is usually already in the delivery workflow.Choose it to manage the delivery: a trainer whose reaction scores collapse across three cohorts is a real signal and worth acting on quickly. Keep collecting it. Just do not report it upward as evidence of capability, because it is a measure of the experience and was never designed to be anything else.
Post-programme knowledge testWhere the cohort stands at the end, on the material as taught.One sitting per participant.Choose it when the requirement is a standard rather than a change — a mandatory pass mark, a certification gate, a regulator or client who needs to see that everyone who finished is above a line. That is a legitimate and common requirement, and a baseline adds nothing to it. It simply cannot tell you what the programme contributed.
Paired baseline and outcome measurement (AssessAll)Movement on named competencies for named participants: gain in percentage points, paired Cohen's d, how many below the pass threshold at baseline crossed it, how many declined, and the coverage split that shows who dropped out.Two sittings per participant. Credits are US$0.50 each; packs of 250 / 1,000 / 5,000 at US$125 / 465 / 2,175 discount up to 13%. New organisations get 250 free credits, enough to run a pilot cohort end to end.Not the answer for a cohort of five. Below three matched pairs this platform withholds effect sizes on purpose, and even at ten pairs you are reading individual movement rather than a programme-level result — so for a small leadership group, buy the individual reports and read them as individuals. It is also the wrong spend when the programme's claim is about behaviour six months out and nobody will be available to re-measure; an evaluation design nobody will finish produces worse evidence than the attendance record it replaced.

Facts about the HRD Corp claim process are taken from HRD Corp's own published claim-application process flow and employer FAQ on the date shown and are re-checked quarterly. AssessAll is not registered with HRD Corp, SkillsFuture Singapore or the incoming Skills and Workforce Development Agency, holds no accreditation or certification in either market, and nothing here determines whether a programme or a provider is claimable or funded — that is decided by the scheme, not by an assessment platform. Kirkpatrick's four-level model is described as published by its author and is not a proprietary AssessAll framework.

Sources: HRD Corp — Process Flow for Claim Application (supporting documents and the six-month window) · HRD Corp — employer FAQ · MOM Singapore — merger of WSG and SSG, 12 February 2026 · MOM Singapore — Second Reading, Skills and Workforce Development Agency Bill, 5 May 2026

What is different here

Singapore and Malaysia: two funding systems, one measurement gap

Both markets fund workforce training heavily and both hold providers to account for delivery. Neither scheme's paperwork is designed to tell an employer whether capability moved — which is the employer's own question to answer.

Malaysia — the levy is claimed on delivery evidence

HRD Corp's claim application runs on an attendance report showing each trainee's log-in and log-out, an invoice addressed to HRD Corp, a T3 form signed by trainees and a JD/14 form completed by the training provider, submitted within six months of training completion (HRD Corp's published claim process flow, checked 3 September 2026). That is a complete and appropriate evidence set for a funding decision. It is not an evidence set about capability, and an employer who files it and reports nothing else has answered the scheme's question rather than the board's.

Malaysia — "HRD Corp claimable" is a funding status

It is one of the most repeated phrases in Malaysian corporate training and it is regularly read as a quality mark. It says the programme sits inside the scheme and the levy can be applied to it. Whether the programme moved anything is a separate question that the phrase does not address, and the only way to answer it is to measure the same people twice.

Singapore — the agency changes, the question does not

MOM announced on 12 February 2026 that Workforce Singapore and SkillsFuture Singapore would merge into a single statutory board, jointly overseen by MOM and MOE. The Skills and Workforce Development Agency Bill had its Second Reading on 5 May 2026, and SWDA was stated to launch in the third quarter of 2026. Employers should expect continuity of services through the transition; what does not change is that a funded programme still has to be able to say what it changed.

Singapore — skills-first reporting raises the bar on evidence

The Second Reading speech describes SWDA supporting skills-first workforce planning, hiring and internal mobility, with a Careers and Skills Passport as a common verified information source. A workforce plan expressed in skills needs readings of those skills over time, which is a measurement requirement before it is a reporting one. AssessAll has no connection to the Careers and Skills Passport and no role in any government scheme; the point is only that skills-first planning is hard to do on completion data.

Both markets — the cohorts are often too small for a headline

A great deal of training here runs in cohorts of eight to twenty. That is enough to read individual movement and per-competency direction honestly, and not enough to publish a programme-level effect size with confidence. The right report for a cohort of twelve names who moved and who did not; the wrong one is a single average with a decimal point on it.

Both markets — English is often the hidden variable

In multilingual workplaces a knowledge gain measured only in English can be a language reading in disguise: the participant understood the concept and could not demonstrate it in the test's language. If the programme is delivered in one language and measured in another, say so in the report, and consider measuring workplace English separately so the two findings do not contaminate each other.

Common use cases

  • Levy-funded corporate training in Malaysia where the board wants a capability reading and not a claim summary
  • SkillsFuture-supported programmes in Singapore reporting to a sponsor who asked what changed
  • Training providers who want an outcome number in the proposal rather than a testimonial
  • Leadership programmes measured by a follow-up 360° cycle against the first one
  • Compliance and conduct training where completion is mandatory and judgement is the actual objective
  • Shared-services and GBS hubs proving an upskilling programme before scaling it across sites
  • Pilot cohorts that have to earn the budget for the full rollout

Pricing for proving a programme worked

Measurement is priced per sitting, in US dollars, with no platform fee, seat licence or annual contract: 1 credit = US$0.50, and a paired design is two sittings per participant — the baseline and the outcome — plus any checkpoints you choose to run in between. Credit packs (250 / 1,000 / 5,000 at US$125 / 465 / 2,175) discount volume by up to 13%. New organisations get 250 free credits, which is enough to baseline a pilot cohort, run the programme, re-measure and read the impact report before spending anything. The arithmetic worth running against us: divide the cost of two sittings by your cost per participant for the programme itself. If the measurement layer is a low single-digit percentage of the programme, the question is only whether you want evidence; if it is a large fraction of it, your programme is cheap enough that a pilot with a proper baseline is a better use of the money than measuring every cohort.

Frequently asked questions

How do you prove a training programme actually worked?+

Measure the same capabilities on the same people before and after the programme, and report the change only for participants who completed both measurements. Three things make that evidence rather than a claim: the competencies measured are the ones the programme said it would move; the gain is computed on matched pairs, so people who dropped out cannot flatter the average; and the drop-out is published beside the gain. A completion rate, a reaction survey and a post-only test each answer a different and narrower question.

Does an HRD Corp claim prove the training was effective?+

No, and it is not designed to. HRD Corp's published claim process runs on an attendance report showing each trainee's log-in and log-out, an invoice addressed to HRD Corp, a T3 form signed by trainees and a JD/14 form completed by the training provider, submitted within six months of training completion. That evidences that the training took place and that the levy may be applied to it. "HRD Corp claimable" is a statement about funding eligibility, not a finding about capability. Facts checked against HRD Corp's own claim-application process flow on 3 September 2026; AssessAll is not registered with HRD Corp and cannot determine what is claimable.

Why is a post-training test not enough on its own?+

Because it tells you where the cohort ended up, not what the programme contributed. Without a baseline you cannot separate what people learned from what they already knew, and the people sitting the final test are not a random sample of the people who started — the ones who struggled are disproportionately the ones who left. A post-only test is the right instrument when the requirement is a standard to be met rather than a change to be demonstrated: a mandatory pass mark, a certification gate, a client who needs everyone above a line.

How many participants do you need before the result means anything?+

For individual movement, one — a person's baseline and outcome are informative about that person. For a programme-level effect size, more than most cohorts contain. AssessAll withholds the effect size below three matched pairs and says so on the report, because a standard deviation computed on two people is noise. Between roughly three and thirty pairs, read direction and per-competency movement and name the individuals; above that, an effect size starts to be worth quoting. If you want the arithmetic behind that judgement, the sample-size and reliability calculator is the tool for it.

What is a matched pair, and why does it matter so much?+

A matched pair is one participant who completed both the baseline and the outcome assessment. It matters because the alternative — comparing the average of everyone who sat the pre-test to the average of everyone who sat the post-test — manufactures improvement out of drop-out. If the people who found it hardest do not return, the post-test average rises without anybody having learned anything. AssessAll computes every gain figure on pairs only and counts the unmatched attempts separately, so the coverage figures show exactly how much of the cohort the headline number rests on.

What if the result shows the programme did not work?+

Then you have learned the thing you were paying to learn, one cohort in rather than three years in. In practice a null overall result is rarely null everywhere: the per-competency table usually shows one capability that moved and another that did not, which is a redesign brief rather than a verdict on the programme. The participant-level view adds the second half of it by naming the people whose scores did not move, and those two readings together are what turns a disappointing number into next quarter's plan.

Can you measure behaviour change, or only knowledge?+

Both, with different instruments. Knowledge and applied judgement are measured directly by assessment, before and after. Behaviour is measured by observation, which on AssessAll means a 360° cycle: a follow-up cycle is linked to the earlier one and its reports show change against it, so the same paired logic applies to how someone is seen to lead rather than to what they know. Behaviour readings need more time between waves than knowledge readings — a quarter at minimum — because the thing being measured is habit.

Is AssessAll accredited or registered with HRD Corp or SkillsFuture?+

No. AssessAll is an assessment platform and holds no accreditation, certification or scheme registration in Malaysia or Singapore. Nothing on this page determines whether a programme, a provider or a course is claimable or funded — that is decided by the scheme. What the platform does is measure capability before and after a programme so that whoever is accountable for it can report what changed, whatever the funding route.

The levy paid for the training. What proves it changed anything?