All guides

Succession planning beyond the 9-box grid (measure, don't vote)

The 9-box grid fails succession planning because its potential axis is opinion, not measurement. Here's how to replace calibration votes with level-calibrated readiness assessment — and run a succession slate on evidence.

Last updated

The short answer

Succession planning works when the 'potential' half of the decision is measured instead of voted. The 9-box grid plots performance (usually real data) against potential (usually a manager's opinion), which is why calibration meetings turn into advocacy contests and why so many 'ready now' successors struggle when the promotion lands. The fix is to keep the process — identify critical roles, build slates, develop gaps — but replace the potential vote with a readiness assessment: level-calibrated judgement scenarios, set at the level of the seat being planned for, producing a comparable verdict report per candidate.

This guide covers why the 9-box keeps producing bad slates, what a measured alternative for the potential axis looks like, how to run an assessment-backed succession round in practice, and what it costs — which is now low enough to measure the whole bench, not just the finalists.

Why the 9-box keeps failing

The 9-box's performance axis is usually defensible: it draws on ratings, targets, delivery. The potential axis has no such anchor. In most organisations it is a single manager's judgement, made without ever seeing the person operate at the level in question, shaped by visibility, recency, and how much the manager likes advocating. Two managers with identical evidence routinely place the same person in different boxes — which is why calibration sessions exist, and why they resolve disagreements by negotiation rather than by data.

The performance axis deserves a harder look than it usually gets, and the numbers are public. A 2019 reliability generalisation meta-analysis by Salgado and Moscoso in Frontiers in Psychology cumulated 219 coefficients across 43,203 people and found the observed interrater reliability of supervisory ratings of overall job performance to be 0.56 — close to the 0.52 Viswesvaran, Ones and Schmidt reported in 1996 across 40 studies and 14,650 people. Two competent supervisors rating the same person agree to about that degree and no better.

The finding inside that finding is the one a succession process should act on. Reliability split sharply by why the rating was collected: 0.45 observed for ratings made for administrative purposes, against 0.61 observed (0.69 corrected) for ratings collected for research — the corrected figure for research ratings is over half again as large as for administrative ones. Administrative ratings are exactly the kind a talent review runs on: they decide pay, promotion and standing, and everyone in the room knows it. So the axis the 9-box treats as its solid one is measured with the least reliable version of the rating in the literature.

That does not make performance data useless; it makes a single rater's number a weak input. The practical corrections are cheap. Ask how many independent raters stand behind each performance placement — one rater at r = 0.45 is a coin weighted only slightly. Ask whether the rating was produced for the review it is now being used in. And stop treating a one-box difference between two candidates as a difference at all, because it is comfortably inside the noise both axes carry. Ask one more question that the 2019 figures do not raise on their own: what kind of job is being rated. Zhou, Sackett, Shen and Beatty (Journal of Applied Psychology, 2024; 132 independent samples) report supervisory interrater reliability at 0.57 for managerial positions against 0.68 for non-managerial ones, and advise against using an overall grand mean at all. A nine-box grid is populated almost entirely with managers, so it is the lower of the two coefficients that describes the axis, and a single grand-mean figure — whichever one is quoted — is the wrong statistic for this table.

The deeper problem is what 'potential' is standing in for. What a succession decision actually needs to know is: can this person handle the dilemmas of the target seat? That is a question about judgement at a specific level — delegating at one altitude, holding standards through proxies at another, trading a quarter against the franchise at a third. A generic potential rating compresses all of that into one opinion. The grid isn't wrong to want a second axis; it's wrong about how the second axis gets filled in.

Replace the potential vote with a readiness measure

A readiness assessment answers the question the potential axis was guessing at. Instead of asking a manager to rate potential, you put each succession candidate into realistic judgement scenarios calibrated to the target level — the same instrument family as a situational judgement test, with the scenario bank written for the altitude of the seat being planned for. The output is a verdict per candidate per level — ready now, ready with development in named areas, or not yet — with the judgement evidence behind it. How to measure leadership readiness covers the instrument design in depth, including the anti-gaming properties a defensible readiness measure needs.

This changes the calibration meeting's job. With measured readiness on the table, the meeting stops adjudicating opinions about potential and starts doing what humans are actually good at: weighing verdicts alongside context the instrument can't see — track record in adjacent situations, retention risk, timing. The grid can even survive as a display: performance on one axis, measured readiness on the other. What goes is the vote.

One nuance the 9-box cannot express at all: readiness is level-specific and situation-specific. A candidate can be ready for the next level and unready for the chapter the seat is walking into — a turnaround, an integration, a founder transition. Where the situation is known, measure judgement in that situation specifically; AssessAll packages six of these as the Situational Suite alongside the Leadership Ladder's eight level-calibrated instruments (L0–L7).

Running an assessment-backed succession round

The process keeps the familiar succession scaffolding. First, name the critical roles and the level each sits at — on a ladder framing, most succession-relevant seats fall between senior manager and board (L3–L7). Second, build a slate of candidates per role — wider than feels natural, because the whole point of cheap measurement is that you no longer have to pre-filter on impressions. Third, assess every candidate at the relevant level or levels and compare verdict reports side by side. Fourth, treat 'ready with development' as the process working: a named gap plus a development plan, caught before the promotion instead of eighteen months into a struggling tenure.

On AssessAll this runs without any candidate onboarding: the organisation buys credits, picks a slate, and sends each candidate a share link — no candidate accounts. The Succession Slate runs a candidate through all of L3–L7 (375 credits) for five comparable verdict reports; the Full-Ladder Benchmark (450 credits) baselines a HiPo cohort across all eight rungs; the Transition Slate (405 credits) maps bench strength across the six situational moments. Reports return on submission, so a full succession round fits inside a normal talent-review cycle rather than stretching across one.

What evidence-based succession costs

The reason succession planning defaulted to opinion was never that anyone trusted opinion — it was that measurement cost too much to apply to more than a shortlist. Consultant-led assessment centres price per engagement, at thousands of dollars a head, so only final-round candidates for the biggest seats ever got measured. Metered assessment inverts that: on AssessAll, credits are ₹30 / US$0.50 each, a full five-level Succession Slate is 375 credits per candidate, and single readiness assessments run ₹599–3,499 (US$6.99–39.99) — pay-as-you-go pricing with no platform fee or seat licence.

At that price, measuring an entire slate costs less than one assessment-centre day for one finalist — which is what makes the wide-slate discipline affordable in the first place. A new organisation's 250 free credits are enough to run a real pilot slate before paying anything.

Frequently asked questions

What is a succession planning assessment?

An assessment used to determine whether succession candidates are ready for the roles they are slated for — typically judgement scenarios calibrated to the target seat's level, producing a verdict (ready now / ready with development / not yet) per candidate. It replaces or hardens the 'potential' judgement that grids like the 9-box leave to manager opinion.

How reliable is the performance axis of a 9-box?

Less than the process assumes. Meta-analytic work puts the observed interrater reliability of supervisory ratings of overall job performance at about 0.56 (Salgado & Moscoso, Frontiers in Psychology, 2019; 219 coefficients, 43,203 people), close to the 0.52 reported by Viswesvaran, Ones and Schmidt in 1996 and revised again to 0.65 by Zhou, Sackett, Shen and Beatty (Journal of Applied Psychology, 2024; 132 independent samples), whose own recommendation is to stop using a grand mean and use job-specific figures — 0.57 for managerial positions, 0.68 for non-managerial. More usefully, that reliability falls to about 0.45 for ratings collected for administrative purposes — pay, promotion, standing — which is precisely the kind of rating a talent review uses. Two consequences follow: ask how many independent raters sit behind each placement, and treat a one-box difference between two candidates as no difference at all.

Is the 9-box grid outdated?

The format isn't the problem; the inputs are. Performance-versus-readiness is a sensible display, but the classic 9-box fills the second axis with an unmeasured potential rating, which makes the whole grid only as reliable as the least calibrated manager in the room. Filling that axis with a measured readiness verdict keeps the grid's clarity and removes its central weakness.

What should replace the potential axis in succession planning?

A level-calibrated readiness measure: scenario-based judgement assessment set at the level of the target seat, scored against a research-keyed rubric, with self-ratings excluded from the score. Unlike a potential rating, it is the same instrument for every candidate, so slate comparisons are like-for-like.

How many candidates should be on a succession slate?

More than the traditional two or three. Pre-filtering to a short slate made sense when measuring each candidate cost thousands; at credit pricing (a five-level slate for 375 credits ≈ ₹11,250 / US$187.50 per candidate) the better discipline is to assess everyone plausible and let the verdicts narrow the slate — surfacing the strong candidates impressions would have missed.

How often should succession candidates be reassessed?

On decision moments rather than a fixed calendar: when a seat's timeline moves, when a candidate completes the development plan a previous 'ready with development' verdict named, or when the situation changes (a planned turnaround or integration makes the situational diagnostics relevant). An annual full-bench refresh alongside the talent review is a reasonable ceiling.

Do succession candidates need accounts on the platform?

On AssessAll, no. Organisations send each candidate a share link; the candidate takes the assessment in the browser and the verdict report returns to the organisation on submission. This matters in succession work, where discretion is often required and asking candidates to register on an assessment platform signals exactly what the process is trying not to announce.

Run your next succession round on evidence
Credit-priced succession slates on the Leadership Ladder (L3–L7) — share links, no candidate accounts, verdict reports on submission. 250 free credits to pilot.
Explore succession slates