Coaching Conversation Assessment with a Spoken Round for Managers and Internal CoachesNot how you would coach. What you would say next, turn by turn, and where in the conversation it changes.
Three coaching conversations that are already running: a senior engineer whose release slipped, a first-year analyst whose report keeps coming back from two levels above her, and a peer team lead whose report goes over his head. At eighteen turns the other person has just spoken, in their words, and you choose what you are most likely to say next out of four things a real manager would say. The report is four heat strips, one cell per turn in conversation order, so the finding is a shape: most managers open asking and close telling, and no summary score can show it. One spoken turn per conversation is graded on its transcript and reported apart. Report ₹899 in India inclusive of GST, or US$8.99 elsewhere.
An assessment of the next utterance, not the coaching approach
The Coaching Conversation Assessment is a thirty-five-minute situational judgement instrument for managers and internal coaches. It measures what a person actually says next, turn by turn, inside three coaching conversations that are already running, across four parts: their agenda, opening rather than closing, the observable rather than the trait, and the next step staying theirs.
Most coaching assessments measure an approach: what a manager would do with a person over a relationship, or how much coaching theory they hold. Two such instruments already sit in this catalogue. This one measures something no approach-level instrument can see: the single next utterance, and where in a conversation it changes. A manager can hold an excellent approach and still close every conversation by telling, and that is the thing this report is built to show.
Eighteen turns across three conversations of six. Each turn prints what the other person has just said, in their words, and asks what you are most likely to say next out of four utterances a capable manager would defend. Some turns warrant a question; some warrant a reflection, an observation, or landing the step, and a reader who always asks a question scores below the ordinary respondent. Every utterance carries an effectiveness value and a sparse credit vector over the four parts, keyed to the published evidence on goal-focused coaching, autonomy support, feedback at the task and process level rather than the self level, reflective listening, and implementation intentions.
Eight forced-choice items ask which of four utterances you are most and least likely to say, matched so that no option is the transparently nice one. Six transcripts ask you to mark every line that moved the conversation onto your agenda, named a trait instead of a thing seen, or took the step away, and over- and under-marking both cost. Four matching and four ordering items ask what an utterance does to the other person's thinking. One instruction is declared for the whole instrument: what you are most likely to say, never what a coach should say.
The report is four heat strips, one per part, one cell per turn in the order the conversations ran, shaded by the credit the utterance you chose earned, with the value and a word in every cell. Because the x-axis is the conversation, the finding is a shape: the mean quality of your utterances in the opening third of each conversation against the closing third, printed with the difference and its measurement threshold, drawn as a slope only when it clears the threshold and drawn flat when it does not. Beside every strip, the within-strip spread as a word, because a manager who is even at a middling level and one who is spiky need different advice.
One spoken turn per conversation is recorded, transcribed and graded on the transcript against the same four content criteria, and reported entirely apart: it enters no strip and no figure. Nobody scored your tone, warmth, pace or accent, and no such score exists anywhere in the report. A coaching conversation is partly how it sounds, and the report says plainly that it does not measure that part.
Every figure is chance-corrected against a declared marginal, so zero is the ordinary respondent and not a coin; every figure carries its 68 and 95 per cent bands; a part too thin to carry a figure carries a placement word and its refusal printed where the figure would have gone; and the modelled correlation between the four parts is printed on the page, so you can see they are four things and not one thing drawn four times. No total, no percentile, no type label. One if-then utterance for your next one-to-one, and the movement a re-sitting in six weeks would have to clear to count as real change.
What you walk away with
One strip per part, one cell per turn in the order the three conversations ran, shaded in four greyscale steps by the credit the utterance you chose earned, with the value and a word in every cell. A turn you did not answer is drawn empty and labelled, never as a zero. The strips are the finding; nothing is averaged away.
The mean quality of your utterances in the opening third of each conversation against the closing third, per part and overall, printed with the difference and the measurement threshold it has to clear. A slope is drawn only when it clears; otherwise the line is drawn flat and the report says the shape is within measurement error.
Whether your utterances on a part were even, mixed or spiky across the turns, at declared cuts, with what each shape means. An even middling level is a habit; a spiky one is a set of turns to look at, and the strip shows which.
Their agenda, opening rather than closing, the observable, and the next step staying theirs, each chance-corrected against a declared marginal so zero is the ordinary respondent, each with its 68 and 95 per cent bands, sorted lowest first. The modelled correlation between the four is printed so you can see they are four things.
Three spoken turns, one per conversation, transcribed and graded on the transcript against four content criteria. Reported apart from every figure. A turn nobody has listened to yet is printed as deferred, in words, never as a score. Tone, warmth, pace and accent are not scored, and the report says so.
A single boxed if-then for your next one-to-one, drawn from the part furthest below the ordinary respondent: a specific moment and a specific replacement utterance. The strength your utterances carried best, paired with what it costs. The movement a re-sitting in about six weeks would have to clear to count as real change.
Inside your report
Illustrative sample — your report is generated from your own responses.
opening third 100, closing third 33: −67, clears the 66-point threshold; you start on their agenda and end on yours.
opening third 100, closing third 33: −67, clears the 66-point threshold; you open asking and close telling.
opening third 100, closing third 17: −83, clears the 66-point threshold; you start with the observable and end with praise.
in play on closing turns only, by design; no open-to-close shape is claimed.
| Turn | Their | Opening, | The | The |
|---|---|---|---|---|
| The report sent back, t1 | 3 of 3 (full) | not in play | not in play | not in play |
| The report sent back, t2 | not in play | not in play | 3 of 3 (full) | not in play |
| The report sent back, t3 | 2 of 3 (near) | not in play | not in play | 2 of 3 (near) |
| The report sent back, t4 | not in play | 2 of 3 (near) | not in play | not in play |
| The report sent back, t5 | 1 of 3 (weak) | not in play | 2 of 3 (near) | not in play |
| The report sent back, t6 | not in play | 1 of 3 (weak) | not in play | 3 of 3 (full) |
How to read it: the x-axis is the conversation, not the item list, so the finding is a shape. This reader earns full credit on the opening turns and closes by telling, taking the step or praising. A turn not answered is drawn empty and labelled, never as a zero.
| Part | Opening | Closing | Difference | Threshold | Reading |
|---|---|---|---|---|---|
| All turns | 94 | 39 | -55 | 39 | ▼ falls, clears: you open asking and close telling |
| Their agenda | 100 | 33 | -67 | 66 | ▼ falls, clears: start on theirs, end on yours |
| Opening, not closing | 100 | 33 | -67 | 66 | ▼ falls, clears: open asking, close telling |
| The observable | 100 | 50 | -50 | 82 | ▬ within error, flat |
How to read it:the threshold is computed from a declared per-turn spread and the number of turns on each side, and it is printed on the drawing. A slope is claimed only when the difference clears it; otherwise the line is dashed and grey and the report says the shape is within measurement error. “You open asking and close telling” is the phrase this reader earned; a reader whose thirds sit together earns no phrase at all.
Built for
- Managers who hold one-to-ones and want to know what they actually say when a person brings them a problem, rather than what they believe they would say
- Internal coaches and people who coach peers, who want a turn-by-turn read of where in a conversation their opening moves give way to telling
- Learning and development teams choosing which managers to send on coaching-skills training, who want a measure of the utterance rather than the theory
- Anybody who has taken a coaching-style questionnaire, been told their approach, and wondered what they would have said at turn five
Find out what you say next in a coaching conversation, and where in the conversation it changes
43 items across six formats · 18 turns in three conversations, 8 forced-choice, 6 transcripts, 4 matching, 4 ordering, 3 spoken turns · about 35 minutes · report ₹899 in India inclusive of GST, or US$8.99 elsewhere. Sit it again in six weeks; the report prints the movement that would count.
₹899 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
An approach assessment asks what a manager would do with a person over a relationship, or what they know about coaching. The Coaching Conversation Assessment asks what you are most likely to say next, at eighteen specific turns inside three conversations that are already running, and reports where in each conversation your utterances change. Its report is four heat strips ordered along the conversation, so the finding is a shape: many managers open asking and close telling, and no approach-level score can show that. Two approach-level coaching instruments sit in the same catalogue, and this one is built to be read beside them, not instead of them.
Every turn has four utterances a capable manager would defend, graded from best to worst against the published evidence, and one instruction is declared for the whole instrument: what you are most likely to say, never what a coach should say. Some turns warrant a question, some a reflection, an observation, or landing the step, and several questions on the form are the wrong move: advice with a question mark, or a what-question that moves the topic to yours. Always asking a question was pushed through the real scorer during authoring and lands below the ordinary respondent, as does always reflecting, always advising, always the longest option and every fixed position.
Their agenda: the next utterance stays on what the person came to work on. Opening, not closing: it widens what they are considering rather than narrowing it to your answer. The observable: it names what was seen or heard rather than a trait or a word of praise. The next step stays theirs: they leave owning the action with a date they chose. Every turn credits at most two of the four, so the parts stay distinct, and the report prints the modelled correlation between them. There is no total because a manager who opens every turn and never lands a step is not the same as one who does both halfway, and one number would say they were.
One spoken turn per conversation: the person has just said something, and you say what you would say next, out loud, in under forty-five seconds. It is recorded, transcribed, and graded on the transcript against four content criteria, the same four parts the written turns measure. It is reported entirely apart and enters no figure. Your tone of voice, warmth, pace, pauses, accent and fluency are not scored, the rubric does not ask about them, and the report says so in its refusal block, because a coaching conversation is partly how it sounds and this instrument does not measure that part. A turn nobody has listened to yet is printed as deferred, in words, never as a zero.
The report is ₹899 in India inclusive of GST, or US$8.99 elsewhere, one time. Forty-three items across six formats take about thirty-five minutes: eighteen turns, eight forced-choice items, six transcripts, four matching, four ordering and three spoken turns. The report names a re-sitting in about six weeks, after the one if-then utterance has been in use for ten one-to-ones, and prints per part the movement a re-sitting would have to exceed to count as real change rather than measurement error. No percentile appears anywhere, and you are ranked against nobody.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Forty-three original items across six formats: eighteen scenario turns in three conversations of six, eight forced-choice most-and-least items, six select-every-line transcripts, four matching items, four ordering items and three spoken turns. DECLARED RESPONSE INSTRUCTION: one instruction is declared for the whole instrument and it is BEHAVIOURAL TENDENCY: every turn asks what you are most likely to SAY next, never what a coach should say; the forced-choice items ask which of four utterances you are most and least likely to say; nothing tests knowledge of coaching. Behavioural-tendency framing is less fakeable and correlates better with real behaviour than knowledge framing (McDaniel, Hartman, Whetzel and Grubb, 2007). CONSTRUCT STATEMENT: this measures what a manager actually says next in a coaching moment, turn by turn, inside three conversations that are already running, across four parts: whether the next utterance stays on the other person's agenda, widens rather than narrows what they are considering, names what was seen or heard rather than a trait, and leaves the next step with them with a when they chose; and where in a conversation that changes. It does not measure whether the reader is an empathetic person, their coaching qualifications, how much coaching theory they know, the approach they would take across a whole relationship, how anything sounded, or how they would do in a conversation this instrument did not describe. HOW IT DIFFERS FROM THE COACHING INSTRUMENTS IN THIS CATALOGUE: the Coaching and Developing Others battery and the Sales Coaching Judgement instrument measure the coaching APPROACH a manager would take, across a relationship. This measures the single next utterance, turn by turn, inside one running conversation, and reports WHERE IN THE CONVERSATION the reader changes, which no approach-level instrument can see. SCORING DESIGN: a construct-weighted graded situational judgement test with a SPARSE credit vector. Every turn's four options carry an effectiveness value v from 0 to 3 and a construct-credit vector over the four parts, content.construct_credit, keyed to the published evidence: goal-focused coaching keeps the goal the coachee's (Grant, 2003; Theeboom, Beersma and van Vianen, 2014; Jones, Woods and Guillaume, 2016, for the meta-analytic effect sizes of workplace coaching and the finding that the coachee's own goal and reasoning carry them); autonomy support against control (Deci and Ryan, 2000); feedback at the task and process level against the self level, where praise and the Barnum effect live (Hattie and Timperley, 2007; Kluger and DeNisi, 1996); trait labels against behaviour (Dweck, 2006); reflective listening and open questions as the pairing that produces the other person's own change talk, and distorting or verdict-laden reflections as the pairing that produces defence (Miller and Rollnick, 2013; Magill, Apodaca, Borsari, Gaume, Hoadley, Gordon, Tonigan and Moyers, 2018, meta-analysis of the technical hypothesis); advice given before the person has reasoned as suppressing that reasoning and being discounted (Bonaccio and Dalal, 2006; Deelstra, Peeters, Schaufeli, Stroebe, Zijlstra and van Doornen, 2003, on unwelcome help at work); implementation intentions, a when and a where the person sets themselves, for the closing move (Gollwitzer and Sheeran, 2006); and the tell-versus-ask account of coaching practice (Whitmore, Coaching for Performance). Every turn puts AT MOST TWO parts in play, every option loads only on those, and the best option holds the maximum credit on each, so no option loads on more than two parts and picking the best utterance never depresses a part the turn does not credit. S_k is the credit earned on part k over the answered items crediting it, M_k the maximum attainable on those items, C_k the expected credit under each item's DECLARED marginal, and figure_k = 100 (S_k/M_k - C_k/M_k) / (1 - C_k/M_k), in -100 to 100. The forced-choice, transcript, matching and ordering items feed the parts on the same arithmetic, each worth up to three credit: forced choice (credit of MOST + 3 minus credit of LEAST) / 2 against two declared pole marginals; select-every-line the share of lines whose ticked state matches the key, so over- and under-selection both cost; matching and ordering the share of correct placements. The headline is the eighteen turns alone on v against option_base_rate. THE NULL AND ZERO: zero on every figure is the ORDINARY respondent, somebody choosing at each item's declared marginal (option_base_rate on turns and forced-choice MOST, option_base_rate_least on forced-choice LEAST, select_prior per transcript line, match_prior per matching left, position_prior per ordering position). The marginals are authored assumptions stated in the open, non-uniform within every item, and authored so that the ordinary respondent earns about the mean credit of the four options on every part a turn credits: a fixed position, which the platform's shuffle turns into a coin, must not beat the ordinary respondent, and the probe checks that every fixed habit (each position, longest, shortest, always ask a question, always a what-or-how question, always reflect, always advise, always the utterance that begins with I will, select all, select none, identity) lands within seven points of zero on the headline and on every part. Observed shares replace the declared marginals once live data exists. SPARSITY, PROVEN: the inter-competency correlation matrix from a reference simulation of three thousand sittings through this instrument's own builder, with four latent tendencies sharing half their variance and each option's attractiveness tilted by its own credits, is printed on the report (maximum off-diagonal .66, under the .70 above which a report tells the reader the same thing four times); a second run with independent tendencies gives the correlation the credit vectors alone manufacture (maximum .34). Both are declared assumptions and observed data replaces them. RELIABILITY: McDonald's omega in its Spearman-Brown form from the answered item count and an assumed average inter-item correlation of .24, printed as an assumption; every figure carries its standard error of measurement from a declared, modelled standard deviation, its 68 and 95 per cent bands, and a three-way placement against the ordinary respondent (McDonald, 1999). REPORT DESIGN: the heat strip item map ordered along the CONVERSATION: one strip per part, one cell per turn in conversation order, four greyscale-distinguishable steps each with a value label and a word, a turn not answered drawn empty and labelled no answer, a turn not crediting the part drawn not in play. Because the x-axis is the conversation the finding is a shape: the mean quality on the opening third (turns 1 and 2) against the closing third (turns 5 and 6), pooled across the three conversations, per part and overall, printed with the difference and the measurement threshold 1.96 times a declared within-person per-turn standard deviation times the square root of the sum of the reciprocals of the two counts; a difference inside the threshold is within measurement error and is drawn flat. The within-strip spread is printed beside each strip with a word, even, mixed or spiky, at declared cuts of .20 and .35. THE SPOKEN ROUND: the pipeline that commissioned this product asked for an AI-scored spoken role play; the platform delivers a spoken turn per conversation, recorded, transcribed and rubric-graded on the TRANSCRIPT against four content criteria, the same four things the written turns measure. No pronunciation, tone-of-voice, warmth, pace or fluency score reaches the builder and none is claimed; a coaching conversation is partly how it sounds and this instrument does not score that. The spoken round is reported entirely apart, enters no figure, and a turn whose transcription or grading did not complete is printed DEFERRED in words rather than as a score. LOCALE: the three conversations are set in an Indian workplace with mixed-seniority pairs, and where deference to seniority changes the better next utterance the item says so in its scenario rather than hiding it in the key; that item is the flagship, tagged for expert review, and its keying follows the finding that autonomy support in a high power-distance setting means scaffolding the step rather than withdrawing from it (Hofstede, 2001; Deci and Ryan, 2000). REFUSAL RULES: an unanswered item leaves numerator, denominator and chance term; an empty sitting scores exactly zero everywhere and is not reportable; fewer than twenty-two of the forty keyed items answered is withheld with every refusal printed where its figure would have gone; a part with fewer than eight answered items or omega under .70 on the unrounded value carries a placement word and no figure; a shape with fewer than two answered turns in either third is not computed; no total across the four parts, no percentile anywhere. RE-MEASUREMENT: the report names a re-sitting in about six weeks and prints, per part, the movement it would have to exceed to count as real change, 1.96 times the declared standard deviation times the square root of 2 (1 - omega), from the Reliable Change Index (Jacobson and Truax, 1991), and distinguishes it from the smaller difference that would matter to the people the reader coaches. FURTHER SOURCES: de Haan, Duckworth, Birch and Jones, 2013, for what carries coaching outcomes; Ely, Boyce, Nelson, Zaccaro, Hernez-Broome and Whyman, 2010, for the evaluation of leadership coaching; Baumgartner and Steenkamp, 2001, for response styles. No proprietary coaching model, questioning framework or credential is named or used; the constructs are public and are cited; no item from any published instrument is reproduced or adapted. All forty-three items are original works written for this instrument. No competitor is named anywhere, and the instrument is not affiliated with or endorsed by any author or organisation named above.