All guides

How to assess leadership skills (a practical guide)

Leadership is measurable when you break it into competencies and use the right item types. Here's which competencies to assess, how many raters the evidence actually needs, where to set the bar, and how to turn the result into development rather than a rank.

Last updated

Start with competencies, not personality

Strong leadership assessment begins by defining the competencies that matter for the role — for example decision-making under uncertainty, stakeholder influence, and developing others — rather than a single 'leadership score'.

The AssessAll Capability Framework groups these into domains so an assessment can target exactly the competencies a role needs.

Use item types that reveal judgement

Multiple-choice questions test knowledge, but leadership shows up in judgement. Situational judgement items (realistic scenarios with trade-offs) and short written responses surface how a person actually reasons.

These richer item types need careful grading — rule-based scoring for the objective parts, and rubric-based AI grading for the written and scenario parts.

Which competencies to assess, and how each one shows up

Four to six competencies is the working range. More than that and the battery gets long, the panel loses discipline and every candidate scores middling on everything. Which four depends on the decision, and the competencies split cleanly by the kind of evidence they need.

Judgement competencies are best measured with scenarios, because the interesting part is the trade-off a person makes when both options cost something: Decision Quality, Prioritisation & Focus, Adaptability & Ambiguity. Competencies that only other people can see need multi-rater or structured behavioural evidence rather than self-report, because self-assessment is weakest exactly where the blind spot is: Self-Awareness, Coaching & Development, Empowerment & Delegation.

For each one, decide in advance what would count as evidence and what would count as its absence. Every competency page in the Competency Dictionary publishes both: the four scored band anchors, the observable signs the competency is missing, and interview questions that make a person describe a specific situation, a specific action and a cost — the format that separates a practised answer from a real one.

How much evidence a leadership rating actually needs

Leadership competencies are usually rated by people, and a single human observer is a noisy instrument. That is not a criticism of the observers; it is arithmetic. If one rater's judgement of a competency agrees with another's at around 0.30 — a realistic figure for a single unstructured observation — then a rating based on one person is mostly not about the person being rated.

The Spearman-Brown formula turns that into a planning number. Combined reliability for k raters is k × r ÷ (1 + (k − 1) × r). At a single-rater agreement of 0.30: three raters give 0.56, five give 0.68, eight give 0.77. Nothing about the raters changed between those lines — only how many of them there were. This is why multi-rater feedback is designed around a minimum number of respondents per relationship group, and why a 360 returned by two people should be read as a conversation starter rather than as a measurement.

The same arithmetic explains why structure is the cheapest quality lever available. Raising single-rater agreement from 0.30 to 0.45 — which is what a behaviourally anchored scale and a short calibration session buy you — gets three raters to 0.71, better than eight unstructured raters manage.

Set the bar before you see the scores

Decide what score means ready, and write down why, before anyone is assessed. A cut score chosen after the results are in is a rationalisation of the shortlist you already wanted, and it is the point at which a defensible process quietly stops being one.

Set it against the standard rather than against the cohort: the question is what a person moving into this role needs to be able to do, not who came top of the eleven people who happened to apply. That is standard setting, and even an informal version — a written description of the borderline candidate, agreed by the panel in advance — removes most of the argument later.

Then check the standard error of measurement against the distance from the bar. If a candidate sits closer to the cut than the instrument's own error, the assessment has not separated them from it, and the honest response is a review band with a second source of evidence rather than a decision dressed up as a measurement.

If you want to do this properly rather than informally, the method is the modified Angoff: a panel estimates, item by item, the probability that a just-barely-ready candidate answers correctly, and the sum of those estimates is the bar. Two findings from the research are worth knowing before you convene one. Hurtz and Hertz's generalizability study puts the optimal panel at approximately 10 to 15 raters (Educational and Psychological Measurement 59(6), 1999, 885–897) — most internal panels are half that, and a small panel can still set a usable bar but cannot honestly quote a confidence interval around it. And Clauser, Hambleton and Baldwin found that judges rate content they are personally unfamiliar with as systematically harder, by 0.107 to 0.138 on the probability scale, enough to move one panel's passing score by more than ten raw points (EPM 77(6), 2016, 901–916). For leadership panels that matters twice over, because the senior people you want on the panel are often furthest from the day-to-day work being assessed.

The Angoff cut score calculator runs both halves of this: the panel's cut with its standard error and the items your judges disagree on, and then what that cut does to a real cohort — how many of your candidates fall inside one standard error of the bar, and what the two conventional ±1 SEM adjustments cost in pass rate.

Turn results into development, not just a rank

The point of assessing leadership is to develop it. Good assessment returns specific feedback — what was strong, where to grow, and a concrete next step — and maps each person against the competencies a role requires.

On AssessAll, every attempt returns coaching feedback and a learning path, and results feed a Skill Passport and competency-gap view for teams.

When the question behind the assessment is a promotion or succession decision rather than development, the measurement changes: you assess judgement at the level the person is moving to, not skills at the level they hold. That approach — level-calibrated readiness measurement — is covered in how to measure leadership readiness, and productised as the Leadership Ladder.

Frequently asked questions

Which leadership competencies should we assess?

Four to six, chosen for the decision you are making. For promotion into a first management role, judgement under uncertainty, delegation, coaching and difficult conversations cover most of the risk. For a senior role, add strategic thinking and change leadership. More than six competencies produces a long battery and undifferentiated scores; fewer than four leaves the decision resting on too little.

Can leadership really be measured, or is it too subjective?

The abstract noun cannot be measured; the competencies underneath it can. Once a competency has a written definition, level anchors describing what it looks like at each level, and observable signs of its absence, two trained assessors can agree on evidence for it. What makes leadership assessment feel subjective is usually a competency that was named but never defined.

Is a personality test a leadership assessment?

No. A behavioural or personality profile describes working style and preference, which is useful context for development and for team composition. It does not measure judgement, and it is not designed to be scored against a job standard. Use scenario-based judgement measurement for the decision, and a behavioural profile alongside it for the conversation.

How many raters does a multi-rater leadership assessment need?

More than most programmes collect. If two observers agree at about 0.30 on a competency — realistic for unstructured observation — then by the Spearman-Brown formula three raters reach a combined reliability of 0.56, five reach 0.68 and eight reach 0.77. The cheaper lever is structure: behaviourally anchored scales and a short calibration session that lift single-rater agreement to 0.45 get three raters to 0.71, better than eight unstructured raters achieve.

Should we set the promotion bar before or after seeing the scores?

Before, always, and in writing. A threshold chosen after the results are in is a rationalisation of a shortlist rather than a standard. Set it against what the role requires rather than against the cohort that applied, and check the standard error of measurement against the distance from the bar — a candidate closer to the cut than the instrument's own error has not been separated from it by the assessment.

Should the person being assessed see their own results?

For development, yes — the report is the intervention, and a result withheld teaches only that the process is done to people rather than with them. For a live hiring decision, feedback is usually summarised rather than released item by item, so the items keep working for the next candidate.

Try it yourself
Take a free assessment and start your Skill Passport.
Browse catalogue