All articles
L&D & Capability9 October 2026·7 min read

The Mean Is a Choice, Not a Summary: Turning Individual Scores into a Team-Readiness Decision

Mean, minimum and maximum aggregation of members' ability scores each predicted team performance at .21 to .29, so the meta-analysis cannot pick your rule. An illustrative walkthrough of the five decisions behind one defensible team-level capability claim.

By AssessAll Editorial

Aggregating assessment results to team level means choosing a composition rule — mean, minimum, maximum or spread — that turns a set of individual scores into one team-level claim. The rule is a measurement decision, not a clerical one. On an identical set of scores, those four rules can place the same team on different sides of a readiness threshold.

The scenario below is illustrative — a composite built to show the decisions, not an account of any client.

The question the assessment was not designed to answer

A shared-services organisation assesses 84 people across 12 delivery teams against a capability model, ahead of taking on a new product line. The individual reports are clean. Then the COO asks a different question: which teams are ready?

Nothing in the assessment design answers that. The instrument was built, validated and scored at the individual level. Rolling it up requires a rule nobody specified — and that rule, not the scores, will decide which of the 12 teams get the work.

The question arrives more often now because the capability gap is where executives are looking. The World Economic Forum's Future of Jobs Report 2025 found 63% of employers naming skill gaps as the biggest barrier to business transformation through 2030, with 39% of existing skill sets expected to be transformed or outdated by then. "Is my team ready" is the operational form of that worry.

Decision 1: Establish what the work does with individual contributions

Before picking a rule, describe how the task actually combines effort. Steiner's task taxonomy, still the cleanest frame for this, distinguishes several interdependence types:

  • Additive — individual contributions sum. Output volume per person aggregates upward.
  • Compensatory — contributions are averaged, so a strong member genuinely offsets a weak one.
  • Disjunctive — the group adopts one member's solution. One person's capability can carry the task.
  • Conjunctive — every member must contribute to the product, so the weakest contributor sets the ceiling.
  • Discretionary — the group decides for itself how to combine contributions.

In the illustrative case, two of the 12 teams run a hand-off chain where every step must clear a quality gate: conjunctive work. Three run parallel caseloads where volume aggregates: additive. One rule across all 12 would be wrong for at least five of them.

Decision 2: Derive the composition rule from the task, not from the dashboard

The measurement literature has a formal vocabulary for this. Chan's typology of composition models (Journal of Applied Psychology, 1998) sets out how a construct at one level relates to the same construct at another — additive, consensus, referent-shift, dispersion and process models — and insists the choice is theoretical, made before the data is touched. Mathieu and colleagues' review of team composition models (Journal of Management, 2014) maps the same ground for team attributes.

The empirical evidence will not choose for you. Devine and Philips' meta-analysis *Do Smarter Teams Do Better* (Small Group Research, 2001) compared three level-based indices of team cognitive ability — highest member score, lowest member score and mean score — and found all three yielded moderate positive sample-weighted estimates of the population relationship, in the range .21 to .29. Sampling error did not account for enough of the variation to rule out moderators. Read that correctly: aggregated ability predicts team performance across the literature, but which aggregate is right is a question about your task, answered locally.

A workable default mapping:

  • Compensatory or additive work → the mean, reported with the team's standard deviation.
  • Disjunctive work → the maximum, plus evidence that the team actually routes problems to its strongest member.
  • Conjunctive work → the minimum, treated as the team's operating capability.
  • Coordination-heavy work → the spread as a second number in its own right, not a footnote to the mean.

Decision 3: Check whether this is a skill question at all

Some team-level questions cannot be answered by any function of individual scores.

Woolley and colleagues' study of collective intelligence (Science, 2010), covering 699 people in groups of two to five, found a single factor whose initial eigenvalue accounted for more than 43% of the variance in group performance across tasks — and that factor was only modestly related to members' individual ability. Average member intelligence correlated with collective intelligence at r = 0.15, the highest-scoring member's intelligence at r = 0.19. What correlated more strongly was average social sensitivity (r = 0.26) and the evenness of speaking turns (r = –0.41 with the variance in turns taken).

The follow-up is more directly useful to practitioners. Riedl and colleagues' meta-analytic study (PNAS, 2021), pooling 22 studies covering 5,349 individuals in 1,356 groups, found the balance between individual skill and group process depends on the task: performance on tasks involving selection among alternatives was better predicted by members' individual skill, while idea-generation performance was better predicted by the group's interaction process. Their task-by-task breakdown is stark: more than 51% of the explained variation on a Sudoku task came from individual member skill, while 55% of the variation on a word-unscrambling task came from group processes.

So ask what kind of task the readiness question is about. For structured, judgement-under-rules work, individual scores aggregate meaningfully. For generative or coordination-dependent work, a team score built from individual scores answers a question nobody asked.

Decision 4: Price the error in the rule you picked

Every aggregation rule carries a different amount of measurement error, and this is where team-level reporting usually overreaches.

  • The mean is the forgiving rule. Independent errors partially cancel across members, so a team mean is more reliable than any single member's score. With six members, the standard error of the mean is roughly 40% of an individual's.
  • Minimum and maximum are the brittle rules. Each rests on one observation, inherits its full error, and selects on that error: the lowest observed score in a team of eight is more likely to belong to someone who had a bad day than to be the team's true floor.
  • The practical consequence: a minimum rule requires a confirmatory re-measure of the flagged member before any decision is attached to it. A second observation is cheap next to a wrong deployment call — on AssessAll, a re-test costs one pay-as-you-go credit at ₹30 / US$0.50.
  • The unit of analysis changes. Once the claim is about teams, the sample is 12, not 84. Any local validation of a team-level threshold is being estimated on 12 data points, which is a reason to report ranges rather than cut-offs.
  • Agreement indices are not a universal gate. The interrater agreement statistics reviewed by LeBreton and Senter in *Answers to 20 Questions About Interrater Reliability and Interrater Agreement* (Organizational Research Methods, 2008) license aggregation for shared-perception constructs, where members should converge. A pooled-ability composite is different: disagreement between members' scores is the signal, not noise to screen out. Chan's review of team-level constructs (Annual Review of Organizational Psychology and Organizational Behavior, 2019) is a good guide to telling the two apart.

Decision 5: Put the rule in the report header

The *Standards for Educational and Psychological Testing* require validity evidence to support the specific interpretation and use proposed for a score. A team-level use needs team-level reasoning on the record. In the illustrative case, the defensible report for each team states four things: the construct, the composition rule and why the task warrants it, the team value with an error band, and the decisions the number does not license.

That last line saves the exercise. "Team C's conjunctive operating capability is at level 2, lower bound level 1" supports a targeted intervention. It does not support a statement about Team C's potential, its leadership, or its members' prospects.

The check before you publish a team score

  1. Name the interdependence type of the work, in writing, before computing anything.
  2. Choose the composition rule from that type, not from what the dashboard offers.
  3. Confirm the question is about pooled individual capability rather than team process.
  4. Attach an error band, and re-measure any individual whose single score is carrying a team verdict.
  5. State in the report what the number does not license.

Where the mapping from task to rule is genuinely complex, that is work for a custom assessment design rather than a spreadsheet formula.

The takeaway: a team capability score is a claim about how the work combines people, not an average of the people. Decide the rule from the task before you see the numbers, and report it beside the result — because the rule, far more than the scores, determines who gets the work.

#team-capability#composition-models#aggregation#capability-measurement#l-and-d-measurement#psychometrics

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free