A single hands-on practical exam is an unreliable measure of technical skill. Fifty years of generalizability research show that most of the measurement error in performance assessment comes from task sampling — a candidate who does well on one task often does poorly on another of equal difficulty. Reliable practical assessment needs many short tasks, not one long one.
That finding matters more in India this year than it has in a decade. In May 2025 the Union Cabinet approved a national scheme to upgrade 1,000 government Industrial Training Institutes with a total outlay of ₹60,000 crore — ₹30,000 crore central, ₹20,000 crore state, ₹10,000 crore industry — plus five National Centres of Excellence for skilling at Bhubaneswar, Chennai, Hyderabad, Kanpur and Ludhiana. The scheme targets 20 lakh trained youth and 50,000 trainers over five years.
Money is going into curriculum, equipment and trainers. Comparatively little attention is going to the thing that converts all of it into a hiring decision: the assessment at the end.
The exam at the end of the course is the weak link
Vocational assessment in India is a genuinely serious system on paper. The National Council for Vocational Education and Training regulates awarding bodies and assessment agencies, and the revised National Skills Qualifications Framework organises qualifications into eight levels described along five dimensions — professional theoretical knowledge, professional and technical skills, employability skills and aptitude, learning outcomes, and responsibility.
What the framework does not do is guarantee that any particular practical exam produces a stable score. And the standard design — one candidate, one assessor, one extended hands-on task, one day — is close to the worst case the measurement literature describes.
What the variance components actually show
The most useful evidence here comes from a body of generalizability studies that decomposed performance-assessment scores into their sources of variance. The CRESST technical report on sampling variability of performance assessments (Shavelson, Baxter and Gao) analysed science, mathematics and military job-performance datasets and found the same pattern in every one.
How much variance the person-by-task interaction absorbed
- Science performance assessment: 82% of total variability
- Mathematics performance assessment: 49%
- California Assessment Program tasks: 48%
- US military hands-on job-performance tests: 55–60%
Read that plainly: the largest single thing a practical exam measures is not "how skilled is this person" but "how does this person happen to do on this task." Fitter A is excellent at bench-fitting a dovetail joint and mediocre at marking out; Fitter B is the reverse. One task cannot tell them apart from a genuinely weaker candidate.
The reliability numbers follow directly. With one task and one rater, the relative generalizability coefficients were 0.15 for the science assessment, 0.21 for mathematics and 0.32 for the CAP tasks. For comparison, a selection instrument is normally expected to clear 0.70 before anyone makes a consequential decision with it.
To reach a generalizability coefficient of roughly 0.80, the same studies estimated the number of tasks needed: about 8 for the CAP assessment, 15 for mathematics, 17 for the Navy hands-on tests, 23 for science and 35 for the Marine Corps battery. The authors themselves called the implication "disquieting."
The assessor is not the main problem
This is the counter-intuitive part, and it is where most assessment-quality effort in vocational training is misdirected. Across those datasets, variance attributable to raters and to person-by-rater interaction was consistently near zero. Trained assessors working from a checklist agreed with each other.
So the usual fixes — more assessor training, stricter assessor accreditation, a second assessor in the room — buy far less reliability than they appear to. They address a source of error that was already small. Doubling the number of tasks does more for score stability than doubling the number of assessors.
That does not make assessor calibration worthless. It makes it insufficient on its own, and it is a poor place to spend the marginal rupee if task coverage is thin.
What this means for an employer hiring from the pipeline
A certificate tells you a candidate passed an assessment. It does not tell you what that assessment sampled. For a manufacturing, maintenance, logistics or field-service employer hiring at volume, three practical consequences follow.
First, treat the credential as an eligibility filter, not a ranking. An NSQF-aligned certificate is meaningful evidence that a person completed structured training in a trade. It is weak evidence about where they rank against 300 other certificate-holders, because the score behind it likely came from too few tasks.
Second, prefer breadth over depth in your own screening. If you are going to run a practical stage, ten ten-minute tasks across the competency map beat one hundred-minute showpiece task. Same clock time, substantially better score.
Third, remember what the predictor rankings say. In the 2022 re-analysis of personnel selection validity by Sackett, Zhang, Berry and Lievens, later extended in Industrial and Organizational Psychology, structured interviews came in at ρ = 0.42 and job knowledge tests at ρ = 0.40 — both above work sample tests at ρ = 0.33. Job knowledge tests are cheap, scalable and scored the same way every time. For trades, a well-built job-knowledge stage is not a consolation prize for employers who cannot afford a workshop; it is a strong predictor in its own right.
A design that survives scrutiny
For a technician-hiring funnel drawing on ITI and apprenticeship output:
- Map the competencies first. Use the qualification pack for the trade as a starting list, then cut it to the 8–12 competencies your actual job needs. Assess those.
- Run many short tasks. Target at least 8–10 independent tasks or scenarios per candidate across the map. Fewer than five, and the score is closer to noise than signal.
- Score each task on observable behaviour. Checklists of specific actions, not a global 1-to-5 impression. The evidence says checklist-based raters agree; impressions do not generalise.
- Separate knowledge from execution. Screen job knowledge and safety reasoning online and at scale; reserve bench time for the candidates who clear it.
- Set the cut score on evidence, not roundness. A standard-setting exercise with subject-matter experts, not 70% because 70% feels like a pass.
- Keep the documentation. Wherever you hire, a selection procedure needs a job-relatedness rationale — the EEOC's guidance on employment tests and selection procedures is the clearest statement of the principle, and the discipline it imposes is good practice regardless of jurisdiction.
The online layer is where volume becomes affordable. AssessAll's AI-graded scenarios let a plant or staffing team put 300 certificate-holders through a multi-task knowledge-and-judgement screen at ₹30 (US$0.50) per assessment, so the expensive bench slots go only to the shortlist — and new accounts get 100 free credits individually, 250 for corporates, which is enough to pilot the design before committing to it.
When one long practical is the right call
There are real cases for the extended single task. When the job is one long integrated performance — commissioning a line, completing a weld to a code-qualified procedure, executing a full service on a specific machine — then fidelity matters more than task sampling, and the exam should mirror the job. Certification to a technical standard is exactly this situation, and the standard, not a psychometrician, defines the task.
The error is importing that design into a hiring decision that has to discriminate among many candidates across a broad competency map. Different purpose, different instrument.
The takeaway
India's ITI upgrade will produce a far larger and better-trained pool of technical candidates over the next five years; the constraint on employers will shift from supply to selection. Whoever screens them should spend their reliability budget on more tasks rather than more assessors — because that is where the error actually lives.