Applied Judgment Assessment
Applied skill assessment · voice-process, support and client-facing roles · browse the full catalogue

Speech Intelligibility and Repair Assessment for Voice-Process, Support and Client-Facing RolesThe listener never lost your accent. They lost the word. This assessment finds out whether you can hear which word, name why, and fix it.

Forty-four items for anybody whose voice carries the job. Eighteen lines from real support calls, each printed as it was said, and one question: what will a listener on a phone line in the UK, the US or Australia lose? Twelve carry a real loss among harmless accent features; six carry only accent, and flagging those counts against you. Then the repairs: which change actually removes the risk, and which only makes a speaker sound like somebody else. A spoken round is graded on content only and reported apart. The report is a breakdown register: five named mechanisms, ordered by what each costs a listener, with your own lines reproduced as the evidence.

40 minutes44 scored exercisesEvidence-keyed scoringGlobal · INR & USD

An intelligibility assessment that scores the ear and the repair, and never the accent

The Speech Intelligibility and Repair Assessment is a thirty-five-minute test for voice-process agents, support desks and client-facing specialists. It measures whether a speaker can tell what a listener will lose in connected speech, name the mechanism that costs the listener the word, and choose the change that removes it, and it does not measure accent, nationality or how the speaker sounds.

Most of what a voice floor buys under the name of accent training scores the wrong thing. It rates how a speaker sounds against a native model, or it hands back a fluency number nobody can act on. The published framework this instrument is built on, the phonological control scale of the CEFR Companion Volume, makes intelligibility the criterion and says in terms that nativeness is not the measure. So every item here asks one question: what does a listener actually lose, and why.

Five mechanisms are named and reported: thought grouping, speaking rate and where the pauses fall, word stress, consonant endings and clusters, and vowel contrast. Each is a cause of lost intelligibility the framework recognises, with its sources. A retroflex t, a syllable-timed rhythm, a sounded r and a regional vowel are not on the list, because they cost a listener nothing, and six of the eighteen detection trials carry only features like those. A habit of flagging everything cannot score here, because a flag counts only when the mechanism is named.

The scoring is signal detection with a named hit condition: d-prime, which is how well you tell a line that loses a word from one that does not, and c, which is where you set the bar. Two people can share a d-prime and be a flagger and a filterer, and a single accuracy figure would hide that. Because naming is hard, the headline is read against the ordinary respondent, a person answering every trial at its declared marginal, not against a coin, and the report prints what saying nothing, naming at random and naming one mechanism every time would each have scored beside you.

The repair strand is reported apart. Six select-every-change items, four ordering items and a matching exercise ask which change removes the risk in a specific line, and both over-selecting and under-selecting cost. No keyed change anywhere amounts to sounding less Indian; the changes that do, copying the caller's vowel or flattening an r, are the distractors, so a reader who reaches for them is told so.

The spoken round is three recorded exercises: say a line in groups and name where you paused; say a line as you would repair it and name the two changes; say your own opening with a reference number and name the detail a caller would lose. Each is transcribed and graded on content only. Nobody scores how you sound, the report says so in plain words, and an exercise nobody could listen to is printed as deferred rather than as a zero.

The report is a breakdown register. At the top, d-prime and c as a pair with their bands drawn and written, the shape in one phrase, and the counts underneath. Then the cost table, printed before the register because a cost ordering is a value judgement, and then one row per mechanism in that order: how often you caught it, how often you over-called it, your own line reproduced word for word, and the repair in one sentence. A mechanism none of your answers touched is printed as not evidenced today, at full size, in its own row, because the shape of the register is the advice.

Five named causes of a lost word, a detection pair with its bands, and the repair for each, in your own lines:
Speaking rate and where the pauses fallWord stressThought groupingConsonant endings and clustersVowel contrast

What you walk away with

d-prime and c, as a pair

How well you tell a line that loses a word from one that does not, and where you set the bar, each with its 68 and 95 per cent bands drawn and written. Read against the ordinary respondent, never against a coin, with what every fixed habit would have scored printed beside you.

The shape in one phrase

A flagger who can name it, a filterer who can name it, an even ear: one sentence from the shape of eighteen answers, not from the level, guarded by the report's own statement that it describes the sitting and not you.

Five mechanisms in cost order, with your counts inside

Thought grouping, rate and pauses, word stress, consonant endings, vowel contrast: how often you caught each, how often you over-called it, and a figure with its band where the evidence carries one. Ordered by what each costs a listener, never by your score.

Your own line, reproduced as the evidence

Every row carries one of the lines you judged, word for word, with how it was said, what you named and what the listener actually loses. A row none of your answers touched is printed as not evidenced today, at full size, because the gap is a finding.

The repair, and the one change

Given the mechanism, whether you picked the change that removes it, reported apart from detection. Then one boxed if-then sentence for a specific situation on a call, drawn from the highest-cost row your evidence puts below the ordinary respondent.

A spoken round graded on content, and every refusal printed

Three recorded exercises transcribed and read for what was said, never for how it sounded, with a deferred exercise printed as deferred. A reliability table with omega unrounded says what each part is allowed to claim, and an empty sitting scores exactly zero.

Inside your report

Illustrative sample — your report is generated from your own responses.

The detection header: d-prime and c as a pair, the shape in one phrase, the counts beneath
solid inner block = 68% band · outlined box = 95% band · dashed tick = zero · grey tick = the ordinary respondent · ▲ above · ▬ not distinguishable · ▼ below
Against the ordinary respondent

+1.62 Above the ordinary respondent

0 = the ordinary respondent+1.62-2.80+3.70

About two chances in three that a re-sitting would land between +0.90 and +2.34; nineteen in twenty between +0.21 and +3.03.

Raw d-prime +1.16
0 = hit rate equals false-alarm rateordinary -0.46+1.16-3.50+3.50
c, where you set the bar −0.41
0 = a level bar-0.41-2.50+2.50
The shape, in one phrase

A flagger who can name it

You told a line that loses a word from one that does not, better than the ordinary respondent, and your bar sits low: when in doubt you flagged. You caught 9 of 12 losses and over-called 3 of 6 clean lines.

What the pair was built from (0.5 added to every cell before a rate is taken)
● Caught and named9of 12 lines that lost a word
◆ Sensed but misnamed2a miss for d-prime, counted apart
○ Missed1said nothing is lost
○ Over-called3of 6 clean lines
● Left alone, rightly3accent, not loss

What a habit scores here. Saying nothing is lost on every line scores d-prime −0.30; naming a mechanism at random −1.66; naming one mechanism on every line −2.08. A flag counts only when the mechanism is named.

How to read it: two readers can share a d-prime and be a flagger and a filterer, so both figures are printed with their bands, and the headline is read against the ordinary respondent rather than a coin. The phrase is the shape of eighteen answers on one day, not a type.

The breakdown register: five mechanisms in cost order, your counts inside, your own line as the evidence
The cost table, printed before the register it orders (a value judgement, with its sources on the report)
RankMechanismWhat the listener losesCost
1Thought groupinga whole clause5 of 5
2Speaking rate and where the pauses fallthe detail4 of 5
3Word stressthe word3 of 5
4Consonant endings and clustersthe grammar2 of 5
5Vowel contrastone word, when both fit1 of 5
● caught · ◆ sensed but misnamed · ○ missed or over-called · ▲ ▬ ▼ placement · ○ Not evidenced today, full size in its own place
1. Thought grouping cost 5 of 5
Above the ordinary respondent
Caught
3 of 3
Over-called
0 of 6
0 = the ordinary respondent74-100100
Your own line, reproduced

“Bed four is nil by mouth from midnight, the bloods are back, and the consultant wants them chased before the round.”

All of it in one breath, with no pause at either comma.

● Caught and named. You named thought groups, and that is what the listener loses here.

The repair, in one sentence. Break the line where each clause ends, and give any number or condition a group of its own.

2. Speaking rate and where the pauses fall cost 4 of 5
Below the ordinary respondent
Caught
1 of 2
Over-called
2 of 6
0 = the ordinary respondent-14-100100
Your own line, reproduced

“The engineer will call you between nine and twelve on Thursday and the job number is four seven one nine.”

Fast and even from start to finish, the digits at the same speed as the words.

○ Missed. You said nothing is lost; the listener loses this line to rate and pauses.

The repair, in one sentence. Slow only at the number, say digits in pairs or threes, and pause after the group.

3. Word stress cost 3 of 5
Above the ordinary respondent
Caught
3 of 3
Over-called
1 of 6
0 = the ordinary respondent51-100100
Your own line, reproduced

“That comes to fifteen dollars a month, and the first payment leaves your account on the first.”

'Fifteen' is weighted on its first syllable, FIF-teen.

● Caught and named. You named word stress: weighted on the first syllable, fifteen is heard as fifty.

The repair, in one sentence. Move the weight to the syllable the listener stores the word under, or say the digits.

4. Consonant endings and clusters cost 2 of 5
Above the ordinary respondent
Caught
2 of 2
Over-called
0 of 6
0 = the ordinary respondent38-100100
Your own line, reproduced

“We can't restore the old number today, but the new one is live and calls will forward.”

The t at the end of 'can't' is not sounded.

● Caught and named. You named final consonants: without its t, can't arrives as can.

The repair, in one sentence. Sound the final consonant where it carries the negative, or say cannot.

5. Vowel contrast cost 1 of 5
Not evidenced today
Not evidenced today

None of your answers touched this mechanism, so nothing is said about it: no figure, no placement, no verdict. The gap is printed at full size because a register with a gap in it is a different finding from one without.

The shape of the register, which is the advice

Clustered: speaking rate and where the pauses fall

3 of your 4 detection errors sit on one row. A cluster is one thing to work on: the repair in that row, and nothing else, until a re-sitting moves it.

Tabular fallback for the register, in cost order.
MechanismCaughtOver-calledFigure95%Placement
1. Thought grouping3 of 30 of 67443 to 100Above the ordinary respondent
2. Speaking rate and where the pauses fall1 of 22 of 6-14-47 to 19Below the ordinary respondent
3. Word stress3 of 31 of 65120 to 82Above the ordinary respondent
4. Consonant endings and clusters2 of 20 of 6387 to 69Above the ordinary respondent
5. Vowel contrast————Not evidenced today

How to read it: the rows are ordered by what each mechanism costs a listener, never by your score, so you work on whichever row costs a caller most rather than whichever you happened to top. Every row carries one of your own lines, word for word, and a row none of your answers touched is printed as not evidenced today rather than dropped.

Built for

  • Voice-process and support agents on an Indian floor talking to callers in the UK, the US or Australia, who want to know which of the five mechanisms actually costs their callers a word
  • Team leads and voice coaches who need a diagnosis per agent that names a mechanism and a repair, rather than a fluency score or an accent rating
  • Client-facing specialists, account managers and consultants whose calls carry numbers, dates and names that have to arrive on the first saying
  • Anybody who has been sent to accent training and would rather find out what a listener loses, and what to change, without being told to sound like somebody else

Find out which word your caller loses, name why, and fix that one thing

44 items across seven formats · about 35 minutes · 18 detection trials, 6 select-every-change, 5 match-the-following, 5 true-or-false, 4 ordering, 3 listening and 3 spoken · d-prime and c with their bands, five mechanisms in cost order with your own lines as evidence, one boxed change, and a spoken round graded on content only. Report ₹599 in India inclusive of GST, or US$5.99 elsewhere.

₹599 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and your receipt is sent to that address.

Secure Razorpay payment · ₹599 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

Does this test my accent?

No. It measures whether you can hear what a listener will lose in a line, name the mechanism that costs them the word, and pick the change that removes it. Six of the eighteen detection trials carry only accent features, a retroflex t, a syllable-timed rhythm, a sounded r, a regional vowel, and flagging those counts against you, not for you. No keyed repair anywhere amounts to sounding less Indian. The published framework the instrument is built on, the CEFR Companion Volume's phonological control scale, makes intelligibility the criterion and not nativeness, and the report says so on the page.

What are the five mechanisms, and why those five?

Thought grouping, speaking rate and where the pauses fall, word stress, consonant endings and clusters, and vowel contrast. Each is a cause of lost intelligibility the framework recognises, under its prosodic-features or sound-articulation sub-scale, and each has published evidence behind it: grouping and stress from the work on how listeners hold a line and locate a word, rate from the work on speaking rate and comprehensibility, endings and vowels from the work on which sounds carry meaning. The report prints a cost table ordering them by what a listener loses, from a whole clause down to one recoverable word, with the sources beside it.

How is it scored?

The eighteen detection trials are scored by signal detection with a named hit condition: a flag counts only when you name the mechanism that actually costs the listener the word. That produces d-prime, how well you separate a line that loses a word from one that does not, and c, where you set the bar. Both are printed with their 68 and 95 per cent bands. Because naming is hard, the headline is read against the ordinary respondent, a person answering every trial at its declared marginal, and the report prints what saying nothing, naming at random and naming one mechanism every time would each have scored. Repair is a separate chance-corrected figure with its own reliability, and the spoken round enters no figure at all.

Is the spoken round scored for pronunciation or fluency?

No. The three recorded exercises are transcribed and graded on content only: whether you said the line in full, named where you paused, named the mechanism and the change, and never described the aim as sounding native. Nobody scores how you sound, and the report says so where the spoken results are printed. If the transcription or the grading did not run, the exercise is printed as deferred in plain words, never as a zero. This instrument measures your ear and your repair, not your own output, and it says that too.

How long does it take, what does it cost, and can I sit it again?

Forty-four items across seven formats take about thirty-five minutes. The report is ₹599 in India inclusive of GST, or US$5.99 elsewhere, one time. The report names the next sitting at about six weeks and prints, for each figure, the movement a re-sitting would have to exceed to count as real change rather than measurement error. An empty sitting scores exactly zero and describes nothing, a mechanism none of your answers touched is printed as not evidenced today, no percentile appears anywhere, and you are ranked against nobody.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue →

Methodology: Forty-four original items across seven formats: eighteen detection trials on a fixed closed option set, six select-every-change repair items, five match-the-following, five true-or-false statements of current guidance, four ordering items, three listening items with a written equivalent, and three spoken items graded on the transcript and reported apart. CONSTRUCT STATEMENT: this measures whether a speaker can tell what a listener will lose in connected speech, name the mechanism that costs the listener the word, and change it so the listener does not lose it; it does not measure accent, nationality, vocabulary range, grammar, or how the reader personally sounds. DECLARED RESPONSE INSTRUCTION: knowledge, one instruction for the whole instrument. Every keyed item asks what a listener will lose, which mechanism causes it, or which change removes it, judged against the line as described; no item asks the reader to rate themselves and none is a behavioural-tendency item. FRAMEWORK ANCHOR: the Common European Framework of Reference for Languages, Companion Volume (Council of Europe, 2020), scale for phonological control, whose three sub-scales are overall phonological control, sound articulation and prosodic features, and whose organising criterion is intelligibility rather than nativeness. The five reported mechanisms are placed on it: speaking rate and pausing, word stress and thought grouping under prosodic features; consonant endings and clusters and vowel contrast under sound articulation. CEFR is cited as a published framework, with no implication of endorsement by the Council of Europe. No commercial English test is named, reproduced or referred to anywhere in this instrument. SCORING DESIGN: C6 signal detection with a NAMED hit condition, on the eighteen detection trials. Each trial prints a line as it was said, with three features of its delivery, and the same six categories as options on every trial: the five mechanisms and 'nothing is lost'. Twelve trials carry one real mechanism among two harmless accent features (signal); six carry only harmless features (noise). A HIT is a signal trial on which the reader flagged the line AND named the mechanism that costs the listener the word; a signal trial on which the reader named a different mechanism is counted as SENSED BUT MISNAMED and is a miss for d-prime; a noise trial on which the reader named any mechanism is a FALSE ALARM; naming nothing on a noise trial is a CORRECT REJECTION. d' = z(hit rate) - z(false alarm rate); c = -0.5 (z(hit rate) + z(false alarm rate)); the log-linear correction adds 0.5 to every cell before a rate is taken, so an extreme rate cannot make z infinite. Both are reported, because a flagger and a filterer can share a d-prime. THE NAMING REQUIREMENT AND THE ORDINARY RESPONDENT (playbook §6.32, §6.79): requiring the mechanism to be named stops indiscriminate flagging from earning a high hit rate, and it also puts the prior-drawing respondent below zero on raw d-prime. The declared marginals put that respondent at d' of about -0.46 on a full sitting, and the say-nothing habit at about -0.30, so the two are within two tenths of each other; naming a fixed mechanism on every trial lands between -2.0 and -2.4 and naming at random lands near -1.7. The headline is therefore d-prime AGAINST THE ORDINARY RESPONDENT: the reader's d-prime minus the d-prime of a respondent drawing every answer from the declared per-option marginals on the same answered trials, with the same correction. Zero on that scale is the ordinary respondent, not the coin. Raw d-prime and c are printed beside it. REPAIR STRAND, reported apart: six select-every-change items, four ordering items and one match-the-following, each reduced to a quality in [0,1] over the answered item (the share of lines whose ticked state matches the key; the share of positions correct; the share of pairs correct) against the expected quality of a respondent drawing from the item's declared per-line, per-position or per-pair marginal; chance-corrected as 100 (mean quality - mean chance) / (1 - mean chance), with its own omega and never averaged into d-prime. THE REGISTER: each mechanism pools every keyed unit tagged to it (its signal trials, the noise trials whose lure it is, its repair items, its guidance statement, its listening item, and the left-hand items of a matching exercise that name it) into a chance-corrected figure with a declared modelled standard deviation, an assumed omega and a three-way placement; a mechanism with fewer than eight answered units, or omega under .70 on the unrounded value, carries a placement word and no figure. A mechanism with no answered unit at all is printed as not evidenced today, full size, in its own row. Rows are ordered by the mechanism's DECLARED cost to a listener, printed above the register with its sources, never by the reader's score. THE SPOKEN ROUND: three voice_response items, transcribed and rubric-graded on the transcript against four content criteria each. The platform places no pronunciation, accent, pace or fluency score on a daily-catalogue voice response, and the rubric asks only about what the transcript can carry: whether the line was said in full, where the pauses were named, which mechanism was named, which change was named, and that the aim was not described as sounding native. The spoken round is reported entirely apart, enters no keyed figure, and prints DEFERRED in plain words where the transcription or the grading did not run; a zero with no transcript is nobody listening, never a poor answer. THE NULL: every chance correction runs against the item's declared marginal: option_base_rate on the detection trials, the guidance statements and the listening items, select_prior per line, pair_base_rate per left-hand item, position_base_rate per position. These are authored assumptions stated in the open; observed shares replace them once live data exists. RELIABILITY: McDonald's omega in its Spearman-Brown form from the answered unit count and an assumed average inter-item correlation of .24, printed as an assumption; the standard error of d-prime is the larger of its binomial standard error (Gourevitch and Galanter 1967) and the omega-based one on a declared modelled standard deviation. Every figure carries a 68 and a 95 per cent band, drawn and written. REFUSALS: an unanswered item leaves numerator, denominator and chance term; an empty sitting scores exactly zero everywhere and is not reportable; fewer than twenty-two answered keyed items is withheld; the detection figures need at least ten answered trials with at least six signal and four noise, else a placement word; no percentile appears anywhere because there is no norm group; no repair keyed anywhere amounts to sounding less Indian, because the construct is intelligibility and CEFR's whole point is that intelligibility is not nativeness. LOCALE: every line is one an agent on an Indian support floor would say to a listener in the UK, the US or Australia; the harmless features on the noise trials (retroflex stops, unaspirated stops, a syllable-timed rhythm at a moderate pace, a sounded r, th as t in 'thank you', a short a in 'can't', the British stress on 'advertisement') are accent features attested in the descriptive literature that a listener in those countries follows without effort. SOURCES: Council of Europe, Common European Framework of Reference for Languages: Learning, teaching, assessment, Companion Volume (2020), phonological control scale and its three sub-scales; Jenkins, The Phonology of English as an International Language (Oxford, 2000), the Lingua Franca Core, for consonant clusters, vowel length and nuclear stress as the segmental and prosodic features that carry intelligibility; Derwing and Munro, Pronunciation Fundamentals: Evidence-based Perspectives for L2 Teaching and Research (John Benjamins, 2015), and Munro and Derwing, Foreign accent, comprehensibility and intelligibility in the speech of second language learners (Language Learning, 1995), for the separation of accentedness from intelligibility; Munro and Derwing, Modeling perceptions of the accentedness and comprehensibility of L2 speech: the role of speaking rate (Studies in Second Language Acquisition, 2001), for rate; Derwing, Munro and Wiebe, Evidence in favor of a broad framework for pronunciation instruction (Language Learning, 1998), for prosody-focused instruction outperforming segmental instruction on comprehensibility; Hahn, Primary stress and intelligibility: research to motivate the teaching of suprasegmentals (TESOL Quarterly, 2004); Field, Intelligibility and the listener: the role of lexical stress (TESOL Quarterly, 2005); Cutler, Native Listening: Language Experience and the Recognition of Spoken Words (MIT Press, 2012), for stress-based lexical access; Brown, Functional load and the teaching of pronunciation (TESOL Quarterly, 1988), and Munro and Derwing, The functional load principle in ESL pronunciation instruction (System, 2006), for which vowel contrasts cost; Wells, Accents of English (Cambridge, 1982), volume 3, for the descriptive account of Indian English features treated as harmless here, and Wells, Longman Pronunciation Dictionary (3rd edition, 2008), for the British and American stress variants; Gourevitch and Galanter, A significance test for one parameter isosensitivity functions (Psychometrika, 1967), for the standard error of d-prime; Hautus, Corrections for extreme proportions and their biasing effects on estimated values of d' (Behavior Research Methods, 1995), for the log-linear correction; Macmillan and Creelman, Detection Theory: A User's Guide (2nd edition, 2005). All items are original works written for this instrument. No trademarked instrument, test or scale is reproduced or named; the features, lines and repairs are the authors' own. Not affiliated with, endorsed by or derived from any commercial English test or any certification body.