A competency framework is a shared vocabulary for what good performance looks like — usually 8 to 14 named capabilities with behavioural definitions and proficiency levels. It is not, on its own, a measurement system. A framework tells you what to look for; only an evidence layer — tests, work samples, scored scenarios, calibrated ratings — tells you who actually has it.
Most organisations have built the first half and skipped the second. Mercer's 2025/2026 Skills Snapshot Survey found 38% of organisations now maintain a single enterprise-wide skills library, up from 30% in 2023, and 55% map skills directly to jobs, up from 47%. The vocabulary is spreading fast. What is not spreading at the same rate is any defensible answer to the question that follows: how do you know?
The framework is usually fine. The measurement layer is missing.
This is not an argument against competency frameworks. The research behind them is solid and the design guidance is well settled. Campion and colleagues' *Doing Competencies Well* (Personnel Psychology, 2011) remains the reference text, and its recommendations are unglamorous and correct: keep the model to roughly a dozen competencies, define each in observable behavioural terms, and specify proficiency levels with real anchors rather than adjectives. The paper cites Boeing capping each job family at 10–12 general plus 10–12 technical competencies, and Microsoft applying 8–14 competencies per role.
Campion et al. also name the thing that makes competency models different from job analysis: they are developed top down from business strategy, in the organisation's own language, as an organisational-development intervention. That is a genuine strength — it is why executives read them and job analyses gather dust. But it carries a cost that most implementations never pay off. A model built top down in house language has no inherent evidence behind its ratings. The rigour has to be added deliberately, afterwards, at the point where someone assigns a number.
Almost nobody adds it. The framework ships, it gets wired into the appraisal form and the internal mobility page, and from that day forward the organisation treats "Level 3 in Stakeholder Management" as data.
Three places the framework quietly stops being evidence
1. Self-rated proficiency
Ask people to rate themselves against the framework and you get a skills inventory in a week. You also get numbers with a well-documented ceiling. Mabe and West's meta-analysis of 55 studies on self-evaluation of ability found a mean validity coefficient of r = .29 against criterion performance measures, with high variability (SD = .25).
The useful part of that paper is what moved the number. Self-ratings got more accurate when raters expected their self-evaluation to be compared against a criterion measure, when they had prior experience self-evaluating, and when instructions emphasised comparison with others rather than with an abstract standard. Anonymous, uncompared, unanchored self-ratings — which is exactly how most skills inventories are collected — sit at the weak end of that range.
2. Manager ratings collected for a decision
The second failure is subtler, because manager ratings feel like real observation. Salgado and Moscoso's meta-analysis in *Frontiers in Psychology* (2019) pooled 219 independent coefficients across 14,730 individuals and reported an observed interrater reliability of .56 for overall job performance.
The split underneath that average is the finding that matters. Ratings collected for research purposes reached .61 observed and .69 corrected. Ratings collected for administrative purposes — promotion, pay, calibration, the real ones — came in at .45. The authors put the corrected difference at 53% in favour of research ratings.
Read plainly: the moment a competency rating carries consequences, two competent managers observing the same person agree with each other less than half the time. That is not a people problem. It is what happens when a scale has stakes attached and no anchoring behind it.
3. Proficiency levels that describe seniority instead of behaviour
Open most frameworks at Level 2 versus Level 3 and you find the difference expressed as scope ("across the team" versus "across the function") or as adverbs ("consistently", "proactively"). Neither is observable. A rater cannot verify an adverb, so they fall back on their overall impression of the person and pick a level that matches it — which is how a five-level scale collapses into a three-point popularity index.
What actually fixes it
Two things, and neither requires scrapping the framework.
Train the raters, or accept the .45. Frame-of-reference training — showing raters worked examples of behaviour at each level and calibrating them against known-correct ratings — has meta-analytic support as a method for improving rating accuracy (Roch et al., 2012, Journal of Occupational and Organizational Psychology), with the largest gains in differential and behavioural accuracy. A half-day of calibration before a promotion cycle does more for your data than another quarter of framework redesign.
Give each competency an evidence source that is not an opinion. Not every competency needs a test, but every competency needs a named answer to "what would count as proof?"
| Competency type | Defensible evidence | What to watch | |---|---|---| | Technical / functional | Work sample, scored task, structured knowledge test | Content validity — does the task look like the job? | | Judgement, prioritisation, service handling | Situational judgment or scenario exercise, consistently scored | Scoring key derived from expert consensus, not one author | | Communication, written and spoken | Rated production sample against a published rubric | Rubric anchors and rater calibration | | Behavioural / interpersonal at work | Structured behavioural interview, multi-rater with defined incidents | Multi-item scales outperform single-item ones (.70 vs .63 corrected) | | Learning gain from a programme | Pre/post measure on the same instrument | Same form, same conditions, controlled exposure |
This is where a platform earns its keep or doesn't. AssessAll exists for the middle rows: AI-graded scenario assessments with consistent scoring keys, proctoring that reports integrity bands rather than a pass/fail verdict, and results that travel with the person as a Skill Passport instead of dying in a spreadsheet. Pay-as-you-go credits at ₹30 / US$0.50 make it cheap enough to put a real measure behind a competency you currently rate by feel — but the discipline above matters more than the tool.
Retrofit checklist
- Count your competencies. Above roughly a dozen per role, adoption falls and rating quality follows.
- Mark every competency with its evidence source. Anything left blank is currently opinion. Say so out loud.
- Rewrite proficiency levels as behaviours a stranger could verify. Delete every adverb.
- Separate the inventory from the decision. Self-ratings are fine for finding candidates for a programme. They are not fine for selecting into one.
- Calibrate raters before any cycle with consequences, using worked examples at each level.
- Use multi-item scales, not one global rating per competency.
- Check adverse impact on any competency rating used for selection or promotion, as US federal Uniform Guidelines expect of any selection procedure.
- Document the linkage from business objective to competency to evidence — the SHRM–SIOP competency modeling documentation is the model to follow.
When the vocabulary alone is genuinely enough
There is an honest case for a framework with no measurement layer at all. If it is being used to write job adverts, structure development conversations, give a learning catalogue a spine, or get a leadership team to agree on what "commercial judgement" means here — a shared vocabulary is the entire deliverable, and bolting assessment onto it adds cost and resistance for no gain. Campion et al.'s point about organisational language is exactly right in that setting.
The line to hold is this: the moment a competency rating starts allocating something scarce — a promotion, a seat on the HiPo programme, a pay band, a job offer — it has become a selection instrument, and it inherits every obligation that comes with one.
Takeaway
A framework that names the right things and measures none of them is a communication asset, not evidence, and it should not be quietly promoted into a decision system. Mark each competency with the evidence behind it, calibrate the people assigning numbers, and be honest about which ratings are opinion.