In assessment centre research, the exercise effect is the finding that ratings cluster by the exercise a participant completed rather than by the competency being rated. Two different competency scores from the same role-play resemble each other more than two scores for the same competency from different exercises. Meta-analysis puts exercises at 34% of rating variance and competencies at 22%.
That single result has survived four decades of attempts to design it away, and it has a direct consequence for L&D: the five-bar competency profile a development centre hands back to a participant is, for the most part, not a measurement of five competencies.
The finding that will not go away
The pattern was first documented in the early 1980s, when researchers noticed that post-exercise dimension ratings correlated more strongly within an exercise than across exercises — the opposite of what a competency-measurement instrument should produce.
The quantified version comes from Bowler and Woehr's 2006 meta-analysis of exercise-by-dimension matrices (Journal of Applied Psychology, 91, 1114–1124): dimensions accounted for 22% of variance in post-exercise dimension ratings, exercises for 34%. Those figures are summarised in an open-access review of assessment centre construct validity for anyone without journal access. Later modelling work has not rescued the dimensions. A 2020 study in Frontiers in Psychology on indicator-to-dimension ratios reports that across convergent and admissible models, mean dimension variance was under 22% of total variance, and exercise-based factor models still outperformed dimension-based ones.
The strongest statement of the case arrived in 2024. Reviewing four large-scale generalizability-theory studies, Chris Dewberry concluded in Industrial and Organizational Psychology that assessment centres "do not reliably measure stable competencies, but instead measure general, and exercise-related, performance" — and argued that using them to measure competencies "can no longer be justified" (17(2), 154–175).
Generalizability theory matters here because it does what confirmatory factor analysis of a correlation matrix cannot: it partitions the variance in a rating into named sources — the person, the exercise, the assessor, the competency, and the interactions — and reports how much each contributes. When the competency term is small and the exercise and general-performance terms are large, the competency labels on your report are decoration.
A 2026 study shows the problem in miniature
A PLOS ONE study published this year put 64 mid-level managers through 16 virtual-reality development centre sessions — an immersive plant-crisis scenario — and compared five competency scores against parallel 270-degree evaluations and self-assessments.
The centre was well run by conventional standards. Inter-rater agreement was high: rwg of .82 to .95 and ICC(2) of .76 to .92. But convergence with the 270-degree ratings held for only three of the five competencies:
- Converged with others' ratings: managing people and tasks (r = .40), change management (r = .37), goal orientation (r = .29)
- Did not converge: decision-making, cooperation
- Self-assessment: correlated with the centre on cooperation only (r = .33), and negatively on managing people and tasks (r = −.29)
Read that carefully. Assessors agreed with each other almost perfectly and still produced two competency scores with no relationship to any external view of the same competency. High inter-rater reliability tells you assessors watched the same behaviour and applied the rubric consistently. It says nothing about whether the construct label on the score is real. The authors' own conclusion is the measured one — that virtual development centres are behaviourally informative and best treated as complementary to existing systems, not substitutive.
Why this bites L&D harder than hiring
For selection, the exercise effect is survivable, because selection uses the overall score. In Sackett and colleagues' revised validity estimates, assessment centres predict job performance at .29, revised upward to about .33 once criterion unreliability in managerial jobs is accounted for — behind structured interviews (.42) and job knowledge tests (.40), ahead of unstructured interviews (.19). A general performance factor is still a useful predictor. Rank people on the overall rating and you are on defensible ground.
Development consumes the profile, not the overall score. The competency bars are what get pasted into a development plan, what decide which leadership programme someone is nominated for, what feed the nine-box, and what a participant remembers for years ("I'm weak on strategic thinking"). If those bars are largely exercise variance and general impression, then the development budget is being allocated by noise, and a participant has been told something about themselves that the instrument cannot support.
What is defensible to report
You do not have to scrap the centre. You have to stop over-claiming from it.
- Report the overall rating as the evaluative score. It is the part with validity evidence behind it.
- Report exercise-level performance, not competency-level. "In the crisis simulation, here is what happened" is a claim your data supports. "Your decision-making is a 3.2" mostly is not.
- Give behavioural evidence, not scale points. Specific observed behaviours in a named situation are defensible, actionable in coaching, and immune to the construct-validity problem.
- Sample more situations if you need a general claim. One long exercise is one observation. Several short, varied exercises give a more stable overall estimate — the same logic that governs any work-sample design.
- Separate the development conversation from the promotion decision. Mixing them gives participants a reason to manage impressions, which inflates the general performance factor further.
- Check your own data before defending the design. Correlate the same competency across two exercises, then correlate two competencies within one exercise. If the second number is higher, you have the exercise effect in your own numbers.
- Document what you did. The international taskforce guidelines on assessment centre operations set out the job analysis, behavioural observation and assessor-training requirements that separate an assessment centre from a well-attended group exercise.
Where the constraint is assessor time rather than design quality, AI-graded scenario responses can widen the situation sample cheaply — more short, varied exercises scored consistently, rather than one long simulation asked to carry five construct claims. AssessAll's custom assessment work is built around that trade: scenario-level scores you can defend, not a competency radar chart you cannot.
When dimension-based design is still the right call
Not everyone accepts Dewberry's conclusion, and the dissent is legitimate. Melchers and König argued in the same journal that dimensions should not be dismissed, pointing to evidence that assessors can identify criteria and that frame-of-reference training improves construct validity. Hoffman and colleagues (2011), as summarised in the Frontiers review above, found that broad dimension factors carried incremental criterion-related validity over exercise factors — broad, not the six-to-eight fine-grained competencies most frameworks use.
So dimension-based design remains reasonable when the dimensions are few and broad rather than many and narrow, when assessors get frame-of-reference training rather than a briefing, when each dimension is observed in at least three genuinely different situations, and when your organisation needs a common vocabulary across cohorts more than it needs per-person precision. The mistake is not using dimensions. It is reporting fine-grained dimension scores from a two-exercise centre as if they were measurements.
Your assessment centre is probably a good general-performance instrument and a poor competency instrument. Report it as the first, and your development spending starts tracking something real.