Technical Explanation Assessment for Engineers and AnalystsFluent is not the same as correct. This can tell them apart.
Four faults described in plain words. You write the explanation a non-technical reader could use — then judge sixteen claims about those same faults, eight of which are true and eight of which are the specific ways an explanation goes wrong.
A writing test scores how it reads. This scores whether it is true.
The moment that decides whether a technical explanation helps is not when it is written. It is a week later, when somebody repeats a sentence from it in a meeting. If that sentence went past what the evidence carried, it is now a fact in the organisation and nobody can tell where it came from. Every assessment of technical communication we could find scores structure, tone and readability, and none of them scores that.
Four faults are described here in ordinary words. A retry loop that kept a failing service failing. A cached price that stayed wrong for four hours because the refresh never ran. An index added in working hours that made orders queue. A daily report that counted some orders on the wrong day. For each one you write the explanation a shop manager, a service lead, an operations manager or a finance lead could actually use, and it is graded against four criteria: does it name the cause rather than the symptom, does it assert anything you were not told, does it use a word the reader would have to be taught, does it say what they now have to do.
Then the sitting does the thing that makes it a measurement rather than an opinion. Sixteen claims about those same four faults arrive one at a time and you say whether the description supports each. Eight are true. Eight are the named failure modes: the amplifier called the origin, the duration read off the wrong start point, the component swapped for a different one, a direction nobody stated, a promise that it will not recur. Because the two halves are balanced, accepting everything and rejecting everything both score exactly zero separation — as arithmetic, not as a threshold somebody chose.
That balance is what lets the report say something a percentage cannot. Two readers with the same score can be opposite people: one waves unsupported claims through, the other rejects things that were actually stated. Both write bad explanations and they need opposite advice. The report prints how well you separate the two kinds of claim and, separately, where you set your line, with the band around the first computed from your own two rates rather than assumed.
Your written answers come back beside a stronger version of the same explanation, with the difference annotated rather than a model answer printed alone. Not this is better — yours carried the cause, and a stronger one also named what is still unknown. That is feed-forward at the level generic reports skip.
What you walk away with
A detection figure with a measured standard error rather than an assumed reliability, because the balanced trial structure supplies one.
Two people at the same accuracy can be a reader who lets claims through and a reader who queries things that were stated. The report names both directions in the words of the work.
Four counterfactual panels, with the difference annotated. What you had, what a stronger answer also had, and which specific line closes the gap.
All sixteen, whether each was supported, what you said, and what happened. Receipts rather than an assertion, and none of it in a hover.
A single sentence naming the situation and the behaviour, chosen from the weakest of your four parts, because four changes cannot be told apart afterwards.
What this did not measure, in the reader's language and in body text, so the report cannot be read as a verdict on somebody's engineering.
Inside your report
Illustrative sample — your report is generated from your own responses.
Eight claims are false and eight are true, so accepting everything and rejecting everything both land at zero separation as arithmetic rather than as a threshold.
- why it lasted: our own retries kept it failing
- what is still unknown: how many customers
- what the reader does now: nothing, and here is the signal
The difference is annotated rather than a model answer printed alone, because the useful feedback is which specific thing the stronger version also carried.
Receipts as endnotes rather than a hover, so the two directions of the error can be checked against the claims that produced them.
Built for
- Engineers and analysts who write incident notes, status updates and post-mortems
- Technical leads who review what their team sends to people outside it
- Support and service teams who have to translate an engineering answer for a customer
- Anybody whose written explanation has been quoted back with a sentence they did not mean
Write four explanations. Then find out what you asserted.
4 written explanations, 16 balanced claims and 17 keyed exercises · about 45 minutes · a counterfactual panel per answer and a separation figure with a measured band.
₹799 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
Partly. Four exercises ask you to write an explanation and they are graded against four stated criteria. The other thirty-three are keyed, and the largest block is sixteen balanced claims about the same four faults, which is what turns an opinion about writing into a measurement. Style, tone and grammar are not scored anywhere.
Because it makes both lazy strategies worthless. Eight claims are true and eight are false, so accepting everything and rejecting everything land at the same place: no separation at all. That is arithmetic rather than a threshold somebody picked, and it means the score cannot be produced by a habit.
Separation is how far apart the supported and unsupported claims sit for you. Where you set your line is how readily you reject when unsure. Two people can have identical separation and opposite lines, and they are a reader who waves things through and a reader who queries things that were stated. They need opposite advice and a single percentage hides which one you are.
They are rubric-graded against four criteria: names the cause rather than the symptom, asserts nothing unsupported, uses no term needing a glossary, and says what the reader must do. They are reported on their own and are never averaged into the detection figure, because a rubric score and a detection statistic are not the same measurement.
About forty-five minutes for thirty-seven exercises. ₹799 in India, inclusive of GST, or US$7.99 elsewhere, one-time. Nothing is timed exercise by exercise.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Thirty-seven original exercises across four formats: four written explanations graded against a four-criterion rubric, sixteen claims about a described fault to accept or reject, nine keyed single-choice items and eight select-every-that-applies items. Four short technical faults are described in plain words, and every exercise works on one of them. Construct statement: it measures whether somebody can explain a technical fault to a reader who cannot read the code - naming the cause rather than the symptom, asserting nothing the evidence does not carry, removing the vocabulary without removing the mechanism, and saying what the reader now has to decide. It does not measure engineering skill, debugging ability, writing style, vocabulary, spelling, or how well somebody presents out loud. Declared response instruction: knowledge and demonstrated skill throughout - each exercise has a defensible answer that does not depend on preference. Scoring: the sixteen claims are scored by signal detection rather than by a percentage correct. Eight claims are false and eight are true of the description, so rejecting everything and accepting everything both score zero discrimination. The report separates how well somebody tells a supported claim from an unsupported one, which is d prime, from where they set their threshold, which is the criterion - a reader who waves claims through and a reader who rejects things that were actually stated are two different people at the same accuracy, and they need different advice. The log-linear correction is applied so an extreme rate cannot make the statistic infinite. The four written explanations are graded separately by rubric and are never folded into the detection figures. Construct grounding, drawn across sources rather than from one framework: signal detection theory and the separation of sensitivity from response bias (Green and Swets, 1966; Macmillan and Creelman, 2005), with the log-linear correction for extreme rates (Hautus, 1995); the curse of knowledge, under which an expert cannot recover what a reader does not know (Camerer, Loewenstein and Weber, 1989); the illusion of explanatory depth, under which people believe they understand a mechanism until asked to explain it (Rozenblit and Keil, 2002); the finding that explaining a mechanism to somebody else exposes gaps that self-assessment does not (Chi and colleagues on self-explanation, 1994); plain-language and readability work on how technical prose fails a lay reader (Kimble, Writing for Dollars, Writing to Please; Cutts, Oxford Guide to Plain English); the distinction between a cause and a symptom in incident analysis, and the finding that a single root cause is usually a narrative rather than a finding (Dekker, The Field Guide to Understanding Human Error, 2014); the slip-versus-mistake distinction (Reason, Human Error, 1990); overclaiming and the tendency to assert familiarity with things that were never presented (Paulhus and colleagues, 2003); rubric design and the reliability of analytic over holistic scoring (Jonsson and Svingby, 2007); McDonald's omega as the reliability estimate reported in place of alpha (McDonald, 1999); and the standard error of measurement as the reason a band rather than a point is reported (AERA, APA and NCME Standards, 2014). All items are original works. No item, scale name or report section is taken from any commercial instrument, and no affiliation with or endorsement by any vendor is claimed or implied.