Written Feedback Quality Assessment for Managers, Teachers and ReviewersThey read it. Could they act on it?
Thirty-five exercises on the feedback you write — and the first fourteen ask one question about a real line of feedback: could the person act on that without asking what it means?
Every product in this category measures attendance or reputation. None of them scores the artefact.
Courses on giving feedback run from about two hundred dollars a year for a library subscription to nearly three thousand dollars a seat for two days in a room, and a day with the researcher behind the evidence base costs about four hundred pounds. Every one of them measures whether you turned up.
The 360 tools measure something else again: they ask your colleagues to rate you on a scale for “gives helpful feedback”. That is a measure of reputation, and reputation and output are different things. Nothing in this market takes a piece of written feedback and scores it.
This does. Fourteen exercises print a real line — some blunt, some warm, some long and specific-sounding with nothing to do in them — and ask whether a receiver could act on it. Seven of the fourteen can be acted on and seven cannot, which is what makes accepting everything and rejecting everything score exactly the same as a guess.
That balance turns the sitting into a signal-detection measure rather than a percentage. You get two numbers instead of one: how well you separate usable feedback from the rest, and where you set your line. A lenient reader sends vague feedback on; a strict reader sends back feedback that would have worked. They are opposite problems and the fix for one makes the other worse, which is why no report should fuse them.
Ten more exercises give you a line aimed at the person and four rewrites. The wrong answers are the three things people write instead: the same verdict with more words, a process comment that stops before the next step, and a compliment sandwich that buries the ask. Two written exercises ask you to write the feedback yourself; they are graded on their own rubric, reported entirely apart, and if the grader cannot be reached they leave the arithmetic rather than scoring you a zero for something nobody read.
What you walk away with
Fourteen balanced calls on real feedback lines, scored as discrimination and threshold rather than as a percentage.
Lenient or strict, named and drawn, because the two failures need opposite corrections.
Ten rewrite exercises keyed to the four-level model: task, process and self-regulation help, the person level does not.
Nine exercises on the three conditions a gap has to meet before it can be closed.
Your own feedback, rubric-graded, reported apart and never averaged into a keyed figure.
A single boxed sentence naming a situation and a behaviour, chosen from where your line actually sits.
Inside your report
Illustrative sample — your report is generated from your own responses.
The band is drawn on the chart you read, not printed in a manual you never see. Zero is somebody answering the way people answer these, not half marks.
Two readers with the same accuracy and opposite lines are two different problems, and the fix for one makes the other worse. A single percentage destroys that distinction; this report prints both halves.
Built for
- Managers who write more feedback than they say out loud
- Teachers and tutors whose marking has to be worth the hours it takes
- Reviewers of code, drafts, plans and applications
- Anybody who has been told their feedback is fair and still sees nothing change
Find out whether your feedback can be acted on
35 exercises across five formats · about 35 minutes · the verdict, the band, the line you set, and one change.
₹599 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
Two written exercises do, on their own rubric, and they are reported apart from everything else. The other thirty-three exercises measure how you read somebody else's feedback, which is a different and more reliably measurable thing. The report says which is which and never averages them.
Because two readers with the same accuracy and opposite habits are different problems. One lets vague feedback through; the other sends back feedback that would have worked. A single percentage fuses them, so this report prints discrimination and threshold separately and names both in plain words.
No. Several of the keyed lines are blunt and several of the unkeyed ones are warm. The question is only whether a receiver could act on it without asking what it meant, and warmth is neither evidence for nor against that.
The four-level model of feedback, the meta-analytic finding that about a third of feedback interventions reduce performance when they direct attention to the person, the three conditions for formative feedback, and the implementation-intention literature behind the if-then form. The methodology note names all of them.
About thirty-five minutes for thirty-five exercises. ₹599 in India, inclusive of GST, or US$5.99 elsewhere, one time, for the sitting and the full report.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Thirty-five original exercises across five formats: fourteen binary calls on real feedback lines, ten single-choice rewrites, five select-every-that-applies exercises, four matching exercises and two written exercises. One response instruction is declared for the whole keyed instrument and it is a KNOWLEDGE instruction - what is most likely to be true, or most useful to the person receiving it - never what you would do. The two written exercises carry a production instruction, are rubric-graded, and are reported apart and never averaged into any keyed figure. Construct statement: it measures whether somebody can tell feedback a receiver could act on from feedback that only tells them how they did, and whether they can move a comment off the person and onto the work. It does not measure writing ability, warmth, subject expertise, how well liked somebody is, whether their feedback is kind, or whether the people they manage are performing. Scoring is signal detection. The fourteen binary calls are a balanced set - seven that can be acted on and seven that cannot - so discrimination (d-prime) and threshold (c) are computed separately and reported separately, with the log-linear correction applied so that an extreme rate cannot make d-prime infinite. Two readers with the same discrimination and opposite thresholds are a different problem from each other and a single accuracy percentage destroys that distinction. Accepting every line and rejecting every line both land at zero discrimination as arithmetic rather than as a threshold. The keyed strands beside it are corrected against an authored answer prior per option rather than against a uniform draw, because a uniform draw is not what anybody answers. Reliability is McDonald's omega estimated from item count and a stated assumed inter-item correlation; it is printed with what it permits, and where it will not carry a number the refusal is printed rather than performed silently. Every reported figure carries its standard error band. Constructs and sources: the four-level model of feedback - task, process, self-regulation and self - and the finding that the self level is the one that does not help (Hattie and Timperley 2007); the meta-analytic finding that about a third of feedback interventions REDUCED performance, and that directing attention to the self is the mechanism (Kluger and DeNisi 1996); the three conditions for formative feedback - know the standard, see the gap, act to close it (Sadler 1989); specificity, timing and elaboration in formative feedback (Shute 2008); the finding that information-rich feedback substantially outperforms reinforcement and punishment feedback (Wisniewski, Zierer and Hattie 2020); feedback literacy and uptake on the receiver's side (Carless and Boud 2018); implementation intentions and the if-then form, meta-analytic d = 0.65 across 94 studies (Gollwitzer and Sheeran 2006); goal specificity (Locke and Latham 2002); the speaker's overestimation of their own clarity (Keysar and Henly 2002); signal-detection theory and the separation of discrimination from criterion (Macmillan and Creelman 2005); the log-linear correction for extreme rates (Hautus 1995); and the Barnum effect as the reason a report that names only attractive things reads as a horoscope (Forer 1949). All items are original works. No commercial feedback instrument, 360 questionnaire, competency vocabulary or report layout is reproduced or implied, and no affiliation with any assessment publisher or training provider exists or is claimed.