Performance Appraisal Conversation Assessment for Managers, with a Spoken RoundThe rating is decided. This is about the conversation it has to survive.
Forty-four items for managers on the appraisal conversation: the meeting in which a rating already decided is explained, supported, disputed and followed by a development point. Twenty-four items ask how you actually run that conversation across four legitimate approaches, evidence-led, forward-led, two-way and direct, none better than the others and each scored against your own average. Eighteen keyed exercises test, apart from that profile, what an appraisal conversation gets right or wrong. Two spoken exercises are recorded, transcribed and graded on what you said. The report is a trade-off beam: your own average is the fulcrum, each approach is an arm, and the longest arm carries the situation in which that approach is the wrong one.
The first appraisal assessment that scores the shape of the conversation and refuses to call it a score
The Performance Appraisal Conversation Assessment is a forty-five-minute assessment for managers who run performance appraisals. It measures the shape of how a manager runs the appraisal conversation across four legitimate approaches, each scored against the manager's own average so the four sum to zero; and, kept apart, whether the manager knows what an appraisal conversation gets right or wrong.
Appraisal training teaches a model and a form. What it cannot tell a manager is which way they actually lean in the room: towards the file of examples, towards next year's plan, towards the person's own account, or towards stating the rating and holding it. Each of those is a defensible way to run the conversation, each is the right one in some meetings, and each costs in others. The evidence-led manager puts a person who already agrees on trial. The forward-led manager moves to next year before a lower-than-expected rating has been explained. The two-way manager asks for a self-rating and then has to overrule it. The direct manager states a rating before hearing the fact that would have changed it. No assessment we know of measures that shape without turning it into a type, and this one does, and refuses the type.
The measurement is a within-person centred profile, done the way the response-style literature says it has to be done. Each approach has six items, half of them worded at the opposite pole so that agreeing with everything scores nothing. Each is read against the way managers typically answer that item, declared per item and printed on the report, never against the scale midpoint, because rating scales cluster above the midpoint and a midpoint reference flatters everybody. Then your own average across the four is removed. The four arms sum to zero by construction, high on everything is arithmetically impossible, and a manager who presses the same button throughout is refused a beam in plain words rather than drawn a shape that would be entirely our wording.
Apart from the profile, eighteen keyed exercises ask what should be done, keyed to the published evidence on appraisal: whether a justification rests on observed, dated behaviour or on a trait word; which rating error a pile of identical marks shows; whether a development point is one behaviour a person can start on Monday or a personality they cannot; whether pay news is kept out of the development conversation, as the split-roles evidence says it must be; and what to do in the sixty seconds after a person says the rating is unfair. These are chance-corrected against declared marginals and reported beside the profile, never averaged into it. Tendency and knowledge are different questions, and the report keeps them in different places.
Two spoken exercises close the form. In the first you open an appraisal where the rating is lower than the person expects; in the second you answer the sixty seconds after they call it unfair. Both are recorded through the microphone, transcribed, and graded on the transcript against four content criteria each: did you name a dated example, did you state the rating in plain words, did you say the objection back before answering it, did you name a date. Nothing is filmed. Camera presence, appearance, body language, tone and accent are not assessed, and the report says so. Where nobody has graded a recording yet, the report prints deferred in words and never a zero.
The report is the Trade-off Beam. Your own average is the fulcrum, drawn as a rule down the middle of the page; each approach is an arm extending left or right by its centred value, with its 68 per cent band as a solid block and its 95 per cent band as an outlined box, both written in words. The fact that organises the page is printed at the top in the largest type on it: the four arms sum to zero, this is the shape of how you run the conversation and not how good you are at it, and a long arm is a choice and not a score. The longest arm carries, immediately beside it, the situation in which that approach is the wrong one, because a beam without that reads as a ranking. Every arm carries what leaning on it gets you and what it costs, with the direction your answers showed printed heavy and the other printed light and labelled as not shown. A flat beam is printed as a finding at full size. The page never says you are an evidence-led manager, or any other kind.
Beneath the beam: the keyed strand on its own scale with its bands and its three parts as three-way words, the spoken round in its own panel, the items behind every arm, what zero means with the declared marginals, the reliability table with every refusal printed, the honest downside paired with the strength it belongs to, and one boxed if-then change drawn from your longest arm, with the number of points a re-sitting would have to move to count as real change. At Rs 1,199 inclusive of GST in India, or US$11.99 elsewhere, that report is the product: appraisal-skills courses and certifications for managers in India run at several multiples of it, and none we know of records the conversation, profiles the approach or prints its own refusals.
What you walk away with
Four approaches to the appraisal conversation, each an arm from your own average: evidence-led, forward-led, two-way and direct. Every arm is drawn with its 68 per cent band as a solid block and its 95 per cent band as an outlined box, and both are written in words. The arms sum to zero, and the page says so first, in its largest type.
Whichever approach you lean on most carries, printed beside it, the meeting in which that approach costs: the person who already agrees, the rating that has not yet been explained, the decision that is not open, the fact you have not checked. A beam without that reads as a ranking, so this one never appears without it.
Each approach carries what leaning on it gets you and what it costs when overused, and what leaning away from it gets and costs. The direction your answers showed is printed heavy; the other is printed light and labelled as not shown by your answers, so the page cannot be read as generic praise or generic warning.
Eighteen exercises on the evidence behind a rating, the development point and the pay news, and the disputed rating, chance-corrected against the way managers typically answer them. A figure with its bands, three parts as three-way words, and the repair for each part in one sentence. Knowledge and tendency are reported in different places and never averaged.
Opening a lower-than-expected rating, and the sixty seconds after the person calls it unfair, recorded and transcribed and graded against four content criteria each. Nothing is filmed; camera presence, appearance and body language are not assessed. A recording nobody has graded yet is printed as deferred, never as a zero.
A single boxed if-then sentence drawn from your longest arm: the situation in which that approach is wrong, and the replacement behaviour. Beside it, the number of points an arm would have to move on a re-sitting after four to six weeks to count as real change rather than measurement error.
Inside your report
Illustrative sample — your report is generated from your own responses.
The four arms sum to zero. This is the shape of how you run the conversation, not how good you are at it, and a long arm is a choice and not a score.
Two times in three a re-sitting would put this arm between 4 and 18; nineteen times in twenty between -3 and 25.
When the person already accepts the rating and has come to talk about next year, and when the written record is thin, so that proof means recency.
Two times in three a re-sitting would put this arm between -10 and 4; nineteen times in twenty between -17 and 11.
Two times in three a re-sitting would put this arm between -6 and 8; nineteen times in twenty between -13 and 15.
Two times in three a re-sitting would put this arm between -16 and -2; nineteen times in twenty between -23 and 5.
| Approach | Arm | 68% band | 95% band | Placement |
|---|---|---|---|---|
| Evidence-led | +11 | 4 to 18 | -3 to 25 | longer arm |
| Forward-led | -3 | -10 to 4 | -17 to 11 | near average |
| Two-way and invitational | +1 | -6 to 8 | -13 to 15 | near average |
| Direct and decisive | -9 | -16 to -2 | -23 to 5 | shorter arm |
| Sum of the arms | 0 | zero by construction · placement cut 7 points · no type label is printed | ||
Two exercises recorded through the microphone, transcribed, and graded on the transcript against four content criteria each. Nothing is filmed: camera presence, appearance, body language, tone and accent are not assessed. A recording nobody has graded yet is printed as deferred, never as a zero.
Graded on the transcript: three of four content criteria met. This is about what was said. It enters no keyed figure and no arm.
- · Says what the meeting is for and that the rating will be given in it
- · Names at least one specific, dated example before or alongside the rating
- · States the rating in plain words, without a softening or hedging phrase
- · Asks for her view of the year or of the rating, as a question
DEFERRED. A recording was made and the grading has not completed, so nobody has listened to it yet. This is not a zero and it is not a poor answer: it is a service that did not run.
- · Restates her objection in your own words before answering it
- · Asks her for the detail or the evidence of the migration work
- · Neither changes the rating in the room nor refuses to look at it again
- · Names a specific next step with a date
| Exercise | Status | Criteria met | Enters a figure |
|---|---|---|---|
| Opening the rating | Graded | 3 of 4 | No |
| After the word unfair | Deferred | — | No |
Built for
- Line managers and team leads who run annual or half-yearly appraisals and want to know which way they lean in the room before the next cycle, with the cost of that lean printed beside it
- New managers running their first appraisal round, who want the evidence on what the conversation gets right or wrong and a spoken rehearsal of the two hardest minutes
- HR business partners and L&D teams preparing managers for an appraisal cycle, who need one specific change per manager rather than a type label
- Experienced managers who have had a rating disputed and want to see how they handle the sixty seconds after the words unfair, graded on what they actually said
Find out which way you lean in the appraisal room before your team tells you
44 items across seven formats · about 45 minutes · four approaches on one beam with both bands, the longest arm with the situation in which it is wrong, a keyed strand beside it, two spoken exercises graded on the transcript, and one boxed change. Rs 1,199 in India inclusive of GST, or US$11.99 elsewhere. Sit it again after four to six weeks of the boxed change.
₹1,199 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
It measures the shape of how a manager runs the appraisal conversation across four legitimate approaches, evidence-led, forward-led, two-way and direct, each scored against the manager's own average so that the four always sum to zero and none is better than the others. Kept apart from that profile, eighteen keyed exercises measure what the manager knows about what an appraisal conversation gets right or wrong: whether a rating rests on observed behaviour or a trait word, which rating error a set of marks shows, whether a development point is one behaviour, whether pay news is kept out of the development conversation, and what to do when the rating is disputed. Two spoken exercises are graded on the transcript and reported apart.
The Marking and Rating Consistency Assessment measures rating accuracy: calling the warranted level on a set of work samples and spotting rating errors in a set of marks. This measures the conversation in which a rating already decided is explained, supported and disputed, and calls no level. The Coaching Conversation Assessment with a Spoken Round measures a conversation in which the manager has made no judgement and the next step is the other person's to name; an appraisal is a judgement already made, the evidence for it, and news the other person may not accept, and every item here lives on that side of the line. The Written Feedback Quality Assessment grades the written line; this is the spoken conversation.
Each approach is measured against your own average across the four, so the four arms sum to zero by construction, for you and for everybody who sits it. That is deliberate: it removes how high you rate yourself in general and the habit of agreeing with statements, both of which would otherwise flatter some people and not others. It also means high on everything is arithmetically impossible, so the report prints no total across the approaches and never says you are one kind of manager. The keyed strand does carry a number, against the way managers typically answer those exercises, and it is printed beside the beam on its own scale and never averaged into it.
No. The platform has no video capture. The spoken round is two exercises recorded through your microphone: opening an appraisal where the rating is lower than the person expects, and the sixty seconds after they say the rating is unfair. Each recording is transcribed and graded on the transcript against four content criteria, such as naming a dated example, stating the rating in plain words, saying the objection back before answering it, and naming a date. Camera presence, appearance, body language, tone of voice and accent are not assessed and no score for them exists in the report. Where a recording has not yet been graded, the report prints deferred in words, never a zero.
Rs 1,199 in India inclusive of GST, or US$11.99 elsewhere, one time; 40 credits on a business plan. Forty-four items across seven formats take about forty-five minutes, including the two spoken exercises of up to sixty seconds each. The report names a re-sitting after four to six weeks of the boxed change and prints the number of points an arm would have to move to count as real change rather than measurement error. No total is printed across the four approaches, no percentile appears anywhere, and you are ranked against nobody. Appraisal-skills courses and certifications for managers in India run at several multiples of this price, and none we know of records the conversation or profiles the approach.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Forty-four original items across seven formats: twenty-four profile items (twelve six-point frequency ratings and twelve five-point agreement statements, four approaches by six, half reverse-drawn), six scenario exercises with graded options, six single-choice exercises, four select-every-line exercises, two ordering exercises and two spoken exercises recorded on the platform's voice response, transcribed, graded on the transcript and reported entirely apart. DECLARED RESPONSE INSTRUCTIONS, one per strand and never mixed inside a scale: the profile strand is BEHAVIOURAL TENDENCY (how often you did this; whether you agree you do it), and the keyed strand is KNOWLEDGE (what should be done; what the evidence supports). CONSTRUCT STATEMENT: This measures the shape of how a manager runs a performance appraisal conversation, across four legitimate approaches, evidence-led, forward-led, two-way and direct, each scored against the manager's own average so that the four always sum to zero and none is better than the others; and, kept entirely apart from that profile, whether the manager knows what an appraisal conversation gets right or wrong: a rating supported by observed behaviour rather than a trait word, the rating errors and their fingerprints, a development point that is one behaviour, pay news kept out of the development conversation, and a disputed rating heard rather than closed. It does not measure rating accuracy against a set of marks, which is the Marking and Rating Consistency Assessment in this catalogue; it does not measure the coaching conversation, where the next step is the other person's to name, which is the Coaching Conversation Assessment with a Spoken Round; and it does not measure written feedback, which is the Written Feedback Quality Assessment. The spoken round is graded on what was said in two transcripts and nothing else: it does not assess camera presence, appearance, body language, tone or accent, and it enters no figure. It does not measure how good a manager you are, your personality, or how a particular appraisal you are worried about will go. FRAMEWORK: the due-process model of performance appraisal (notice, a fair hearing, judgement based on evidence) and the split-roles finding that salary and development belong in separate conversations are public constructs cited by author below; no trademarked instrument, appraisal system or vendor product is named or used. SCORING DESIGN, C17, the within-person centred style profile: each approach is scored on its own six balanced-keyed items, mapped to the item's own scale, read against the item's DECLARED marginal rather than the scale midpoint, and the reader's own mean across the four approaches is then removed, so the four always sum to zero and high on everything is arithmetically impossible. Flatness is tested on the RAW pressed values before any mapping, reversing or centring; a sitting with no spread is refused a beam, with the careless flag, rather than drawn a shape that would be entirely our authoring. Six items per approach is a placement at a printed cut of seven points from the reader's own average, never a number and never a percentile. The keyed strand is chance-corrected against declared per-option marginals, carries a number only with at least ten answered and omega at or above .70 unrounded, and is reported beside the profile and never averaged into it; its three parts of six items carry three-way words only. REPORT DESIGN, D27, the Trade-off Beam, new this run: a horizontal beam whose fulcrum is the reader's own mean, each approach an arm extending left or right by its centred value with its 68 per cent band as a solid inner block and its 95 per cent band as an outlined box, both written in words; the sum-to-zero fact printed at the top in the largest type on the panel; the longest arm carrying, immediately beside it, the situation in which that approach is the WRONG one; every arm carrying a paired fits-here and costs-here sentence with the evidenced direction heavy and the other light and labelled as not shown; no type label anywhere; a flat beam printed as a finding at full size in the beam's own place; and a sitting with no raw spread refused a beam entirely in plain words. Beneath it the keyed strand on its own scale, the spoken round apart, the contributors, what zero means with the declared marginals, the reliability table with refusals printed, the honest downside paired with its strength, and one boxed if-then change with the named re-measurement and the movement that would count as real change. SOURCES: Meyer, Kay and French, Split roles in performance appraisal (Harvard Business Review, 1965), for separating salary from development and for the effect of criticism on goal achievement, keyed on the pay-news and meeting-content exercises; Kluger and DeNisi, The effects of feedback interventions on performance: a historical review, a meta-analysis, and a preliminary feedback intervention theory (Psychological Bulletin, 1996), for task-level over self-level feedback, keyed on the development-point exercises; Ilgen, Fisher and Taylor, Consequences of individual feedback on behavior in organizations (Journal of Applied Psychology, 1979), for acceptance of feedback and the source's credibility; Landy and Farr, Performance rating (Psychological Bulletin, 1980), for the rating errors and their fingerprints, keyed on the evidence exercises; Thorndike, A constant error in psychological ratings (Journal of Applied Psychology, 1920), for the halo effect; Murphy and Cleveland, Understanding Performance Appraisal: social, organizational and goal-based perspectives (Sage, 1995), for rating as a social and goal-directed act and for leniency and central tendency; Smith and Kendall, Retranslation of expectations (Journal of Applied Psychology, 1963), and Latham and Wexley, Increasing Productivity Through Performance Appraisal (Addison-Wesley, 1981), for behavioural anchoring and behavioural observation, keyed on the trait-versus-behaviour exercises; Greenberg, Determinants of perceived fairness of performance evaluations (Journal of Applied Psychology, 1986), for soliciting input and two-way communication as fairness determinants; Folger, Konovsky and Cropanzano, A due process metaphor for performance appraisal (Research in Organizational Behavior, 1992), for notice, a fair hearing and judgement based on evidence, keyed on the dispute exercises; Cawley, Keeping and Levy, Participation in the performance appraisal process and employee reactions: a meta-analytic review (Journal of Applied Psychology, 1998), for the effect of voice on acceptance; Locke and Latham, Building a practically useful theory of goal setting and task motivation (American Psychologist, 2002), for specific over vague goals, keyed on the development-point exercises; DeNisi and Murphy, Performance appraisal and performance management: 100 years of progress? (Journal of Applied Psychology, 2017), for the state of the evidence; Hattie and Timperley, The power of feedback (Review of Educational Research, 2007), for the feed-up, feed-back and feed-forward structure of the report; Gollwitzer and Sheeran, Implementation intentions and goal achievement: a meta-analysis (Advances in Experimental Social Psychology, 2006), for the if-then form of the one change; McDaniel, Hartman, Whetzel and Grubb, Situational judgment tests, response instructions, and validity: a meta-analysis (Personnel Psychology, 2007), for the declared response instructions; Baumgartner and Steenkamp, Response styles in marketing research (Journal of Marketing Research, 2001), for acquiescence and the balanced keying the centring removes; Cronbach and Gleser, Assessing similarity between profiles (Psychological Bulletin, 1953), for elevation and shape; Haladyna, Downing and Rodriguez, A review of multiple-choice item-writing guidelines (Applied Measurement in Education, 2002), for cue control. All forty-four items are original works written for this instrument; no item from any published questionnaire, appraisal system, training course or commercial instrument is reproduced or adapted, no competitor is named anywhere, and the instrument is not affiliated with or endorsed by any author, publisher or professional body named above.