All sample reports
Sample report

What does a BPO voice assessment report look like?

A voice assessment report is the recruiter-facing output of a spoken-English and service screen. It gives a composite score against a role benchmark, a suitability verdict, competency scores for pronunciation, fluency, listening and service judgment, and quoted evidence from the candidate's own recorded answers. This page walks through a real AssessAll report, panel by panel.

Voice & Service Suitability Report

The complete 3-page report, exactly as the platform renders it. No form, no email address, no signup. The person in it is fictitious.

Download the sample PDF
Report
Voice & Service Suitability — the default recruiter template for a multi-modal battery
Length
3 pages, seven numbered sections
Built from
A five-section screen: listen-and-repeat, listening comprehension, a spoken service response, a service-to-sales SJT and analytical reasoning
Scored on
8 competencies on a 0–100 scale, plus a composite against a configured role benchmark
Scoring locale
English (India) — the platform default, printed on the cover
Delivered
On screen and as PDF, generated once when the attempt is scored
SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed

The report, panel by panel

Every published sample report in this market assumes its reader already knows how to read one. This is the walkthrough for someone who does not: what is on each page, and — the part that matters when you are comparing instruments — how to read it.

Cover

The identity strip — Test ID, duration, scoring locale and the CEFR badge

What is on the page. A photo slot, the candidate's name, the assessment title, a Test ID that resolves at a public verification URL, the date and time of the sitting, the duration, the scoring locale, the composite score in a dial with the role benchmark marked on it, and a CEFR spoken-English level.

How to read it. Read the scoring locale and the timestamp before you read the score. On this sample they say English (India) and 13 Sept 2026, 7:34 am IST — and that sitting was taken at 10:04 in the morning in Manila. Every time on the report, including the proctoring photo captions, is rendered in Indian Standard Time, because the renderer formats in Asia/Kolkata. Nothing is wrong with the measurement, but a Philippine recruiter reconciling a report against a shift roster is reading a clock two and a half hours behind theirs, and a screenshot of this cover in a client pack will be read as the local time it is not.

Cover

The CEFR badge, and the five scores it is made of

What is on the page. A single CEFR letter — C1 on this sample — under the heading 'CEFR Spoken English Level'.

How to read it. This is the panel most likely to be over-read, so it is worth knowing exactly how it is produced. It is the plain average of five competency scores — pronunciation, fluency, spoken understanding, vocabulary and grammar — pushed through a fixed banding table. Active listening, articulation and service judgment are not in it. On this sample those five average 75.2, and the C1 band starts at 75: one point lower on vocabulary and the same candidate prints B2. The mapping is documented in the code and described there as tunable, which is the honest way to describe it. It is an internal mapping to a public scale, not a CEFR certification, and if a client contract specifies a CEFR level you should be buying a certificated instrument, not reading this badge.

Section 1

Spoken English & Service Competencies — eight cards, and the line under the heading

What is on the page. Eight competency cards, each with a score out of 100, a ten-cell signal meter, and a written interpretation: Pronunciation, Fluency, Active Listening, Spoken Understanding, Vocabulary, Grammar, Articulation & Voice Energy, and Service-to-Sales Judgment.

How to read it. The most important sentence on this page is the small grey line under the heading, and it is printed on every copy of this report: pronunciation and fluency are evaluated against the Indian English (en-IN) acoustic model, with the note that candidates are not penalised for a natural Indian accent. There is no Philippine-English acoustic model in the platform today. For a Manila or Cebu drive that is a fact to weigh, not a disqualifier — the features being scored are intelligibility features rather than accent-matching, and the line exists to prevent exactly the wrong inference — but if you are buying a voice screen for Philippine agents you should ask every vendor, including this one, which acoustic model scores the audio and what the reference population is. The interpretation text under each score comes from a fixed, owner-reviewed bank with no model call, so identical scores always print identical wording; that is a feature when you are comparing two candidates and a limitation when you want nuance.

Section 1

The delta chips and the percentile — and why this sample has neither

What is on the page. Where a norm exists, each competency card carries a chip reading '▲ +6 vs India norm', and a line under the section gives the candidate's percentile 'among all candidates who completed this battery on AssessAll (India norm group)'.

How to read it. Both are suppressed on this sample, and their absence is the annotation. Deltas appear only once eight prior scored attempts of the same product exist, and the percentile needs eight completed attempts too. So the first several candidates in any new drive get a report with no comparison at all, and the ninth suddenly gets one. Then read the label: the comparison group is the India norm group, because that is where the platform's volume is. Eight people is not a norm in any psychometric sense, and this report does not claim it is — it is a running average of the people who sat the same battery before. Treat a delta chip as a sorting aid inside one drive, never as a population statistic, and never quote it to a client as a benchmark.

Section 2

Job Suitability Profile — four rings, and the number that is not in the report

What is on the page. Four ring gauges with a fit percentage and a fit label: Direct Customer Interaction, Domestic Voice Profile, International Voice Profile and Backend Processing Profile — each with a caption and a 'Driven by' line naming its two strongest inputs.

How to read it. Three things to check here, all visible on the sample. First, 'Domestic' means the platform's domestic, not yours — the caption under that ring reads 'india-market customer support', which on a Philippine screen is the ring a recruiter would most naturally misread. Second, the rings are weighted blends, and the weights redistribute when an input is missing: if a battery has no vocabulary section, the International ring is computed from three of its four inputs and prints the same label with the same confidence. You cannot tell from the artefact which rings were computed on partial input, so check the ring's inputs against the sections actually in your battery. Third, the Backend Processing ring's heaviest input is not in the report. It is 40% an analytical-reasoning score which appears in the Section Performance table but is deliberately excluded from the competency cards and from the 'Driven by' line — on this sample the ring prints 65% and names Spoken Understanding 72 and Service-to-Sales Judgment 64, neither of which is the thing weighting it most.

Section 3

Section Performance — check that the weight column sums to 100

What is on the page. A table of every section in the battery with its weight, its score out of 100 and the same score rescaled to a five-point scale.

How to read it. This is the table that makes the composite auditable, and it has one trap. The weight column prints whatever weight the battery was configured with, followed by a per-cent sign. On this sample the five sections were configured 30 / 20 / 25 / 15 / 10 and the column reads correctly. A battery configured with relative weights — 3, 2, 3, 2, 1, which is a perfectly normal way to configure one — prints '3%', '2%', '3%' and so on, a column summing to 11%. The scores and the composite are unaffected; only the printed column is misleading. Add the column up before you send the report to a client, and if it does not reach 100 the weights are relative and the percentages are not.

Section 4

Evidence from the candidate's own responses

What is on the page. Up to three quoted transcripts, one per section, each with the item score, an assessor note, and a marker saying audio replay is available in the online report.

How to read it. This is the hardest panel in any vendor's report to fake, which is why it is the first one to look for when you compare samples. Interpretation text can be assembled from a lookup table; a verbatim transcript of what this person said cannot. Two practical notes. The panel appears only where the battery captured transcripts longer than thirty characters, so a screen built entirely from multiple-choice items produces a report with no Section 4 at all — and a voice screen whose report has no quoted speech is worth a question. And the audio itself lives in the online report, not in the PDF, which matters in the Philippines: under RA 4200 recording is all-party and criminal, so the disclosure the candidate acknowledged before the sitting is what makes the capture lawful, and forwarding the audio onward to a client is a separate act the Act also reaches.

Sections 5–6

Work Style Profile, and the interview probes

What is on the page. A table of behavioural traits banded out of five where the battery includes them, then up to three suggested interview questions.

How to read it. The probes are selected mechanically: any competency scoring below 70, the three lowest first. That is a useful default and a visible one — on this sample only Service-to-Sales Judgment qualifies, at 64, so a single probe prints. It also means a strong candidate's report carries no probes at all, which is not the same as there being nothing to ask. Judge the probes the way you would judge any coaching text: a usable one names a behaviour the interviewer can hear in the answer. The work-style table appears only when the battery carries likert trait items; a purely voice-and-listening screen prints no Section 5.

Section 7

Assessment Integrity — the band is a violation band wearing an integrity label

What is on the page. A coloured band pill headed 'Integrity Band', and nine counters: photos captured, tab switches, no-face events, multiple faces, paste blocked, print screen, window blur, resumes and looked away. Where photo capture is on, a strip of captured images with a baseline marked 'VERIFIED LIVE CAPTURE'.

How to read it. Read this pill carefully, because the word on it means the opposite of what it looks like. It reports violations, not integrity: this sample prints LOW in green, which means few violations and is the good outcome, while a report with many violations prints REVIEW in red. And there is a default worth knowing: if proctoring is enabled but no band was recorded, the pill prints 'HIGH' in green — the best-looking label on the report is also what an absence of data produces. So do not read the pill alone; read the counters beside it, and if every counter is zero and the photo count is zero, you are looking at a report where nothing was captured rather than a candidate who did nothing wrong. On this sample the photo slot reads 'NO IDENTITY PHOTO' because the illustrative attempt carries no images; with capture on, the baseline photo appears there and in the strip.

Cover strip

The verdict — one composite, one threshold, one fixed band

What is on the page. A verdict pill reading SUITABLE, BORDERLINE or NOT SUITABLE, with a sentence of standard text beneath it.

How to read it. The arithmetic is simple enough to state in full, and stating it is the point. The composite is compared to the role benchmark configured on the product: at or above it the verdict is SUITABLE; within fifteen points below it, BORDERLINE; anything lower, NOT SUITABLE. The fifteen-point band is a fixed constant in the code, not a standard error estimated from this battery, so it does not widen when the screen is short or the sections are few. This sample scored 71 against a benchmark of 60. What follows is the honest reading: the verdict is a sorting label computed from one number and one threshold you chose, and it is the right tool for ranking three hundred applicants on a Monday and the wrong tool to hand a candidate as a reason.

8 questions to ask of any sample report

Including this one. If you are choosing between behavioural instruments, download three or four samples — several publishers post them without a form — and put the same six questions to each.

  1. 1

    Does it name the norm group the scores are compared against?

    A percentile with no stated comparison group is not a measurement. Look for it on the cover, not in a footnote.

  2. 2

    Does any paragraph cite the reader's own answers?

    Interpretation text can be assembled from a lookup table. A line quoting which option the person actually chose cannot, and it is the clearest sign the report was generated rather than templated.

  3. 3

    Is there a response-quality or validity check, and what does it do?

    A check that lowers a confidence band is a measurement instrument. A check that auto-rejects the person is a filter, and should be evaluated as one.

  4. 4

    Does the critical column cost anything?

    If the blind-spots section reads as flattery in the shape of a critique, the report is a sales document. A usable one names something the reader would rather not read.

  5. 5

    Is the advice specific enough to act on this week?

    'Communicate more clearly' is unusable. 'Send the numbers before the meeting, not during it' is a behaviour someone can change on Monday.

  6. 6

    How many people is the comparison built from, and is that number printed?

    A percentile or a 'vs norm' delta computed from a handful of prior attempts is a running average of whoever sat first, not a norm group. If the sample size is not on the artefact, ask for it before you quote the number to anyone.

  7. 7

    What does the integrity or proctoring summary print when nothing was recorded?

    Some reports default to the most favourable label when no data exists, which makes an unproctored sitting look like a clean one. Read the counters beside the badge: all-zero counters with a zero photo count is an absence of evidence, not evidence of absence.

  8. 8

    Does the report say what it does not measure?

    Every behavioural instrument has hard limits — no ability signal, weak links to job performance, no clinical meaning. A report that never mentions them is more likely to be over-read by whoever receives it.

What this report does not do

The limits of an instrument are part of describing it honestly, not a disclaimer to bury at the bottom.

  • The CEFR level on the cover is an internal mapping from five competency averages to a public scale, documented in the code as tunable. It is not a CEFR certification, not equivalent to a certificated language test result, and should not be quoted in a client contract that specifies a CEFR level.
  • Pronunciation and fluency are evaluated against an Indian English (en-IN) acoustic model — the report says so on its own first section — and the platform has no Philippine-English acoustic model today. Norm deltas and percentiles are labelled as an India norm group for the same reason.
  • Norm deltas and the percentile appear only once eight prior attempts of the same battery exist. Eight is a running average, not a norm group in the psychometric sense, and this report does not present it as one.
  • The suitability rings are weighted blends whose weights redistribute silently when an input is missing, and the Backend Processing ring's heaviest input is not printed anywhere on the report. A ring is a shortlisting heuristic, not a validated job-fit prediction.
  • The integrity band reports proctoring violations, not identity assurance. A report with proctoring enabled but no band recorded prints the most favourable label. It is evidence to review, never a finding of cheating.
  • The verdict is a comparison of one composite against one threshold you configured, with a fixed fifteen-point borderline band. It carries no validity coefficient for your role and no adverse-impact analysis; those are separate pieces of work.
  • This report measures spoken English and service judgment. It is not an ability test, not a personality assessment, and not a measure of whether someone will stay in the job.
  • The candidate in the sample is fictitious and the responses are illustrative. It is a real report from the real renderer, generated on 13 September 2026, not a real person's results.

Frequently asked questions

What does a BPO voice assessment report look like?

A voice assessment report usually opens with an identity strip — candidate, test ID, duration and the composite score against the role benchmark — then gives a suitability verdict, competency scores for spoken English and service behaviours, a section-by-section score table, quoted evidence from the candidate's recorded answers, and a proctoring or integrity summary. The AssessAll Voice & Service Suitability Report runs to three pages in that shape, with eight competencies, four job-suitability rings and nine proctoring counters. The complete report is downloadable on this page with no signup.

How is spoken English scored in an AI voice assessment?

Speech engines return per-utterance features — pronunciation, fluency, rhythm and intonation — from the candidate's own audio, and rubric-graded open responses contribute vocabulary and grammar criteria. Those are averaged into competency scores out of 100, and a subset of five of them is averaged again and banded into a CEFR-style level. The scoring is acoustic-model dependent: on this report the model is Indian English (en-IN), which the report states on its own first section. Ask any vendor the same question before you buy — which acoustic model, and what reference population.

Can I download a sample voice assessment report without giving my email?

Yes. The sample on this page is a direct PDF link with no form, no signup and no email wall, and it is the complete three-page report rather than an extract. It was generated on 13 September 2026 from the same renderer that produces live reports, using an illustrative candidate.

Is a CEFR level on an assessment report the same as a certificated CEFR result?

No. A CEFR level printed by an assessment platform is that platform's own mapping from its scores onto a public scale, and different platforms map differently. On this report the level is the average of five competency scores pushed through a fixed banding table, and a single point of movement on one competency can change the letter. If a client contract or a visa process specifies a CEFR level, buy an instrument that certificates one. A platform-mapped level is for internal shortlisting.

What should a Philippine BPO recruiter check on a voice screening report?

Five things, all visible on the artefact. The scoring locale and the timestamps — this report prints English (India) and Indian Standard Time, so a 10:04 am Manila sitting shows as 7:34 am. Whether the weight column in the section table sums to 100, because it prints the configured weight with a per-cent sign whatever the configuration. Whether the norm deltas are present at all, since they need eight prior attempts. What the integrity band actually reports — violations, where LOW is the good outcome and a missing band prints the best-looking label. And whether Section 4 quotes the candidate's own words, which is the panel that cannot be templated.

Is it lawful to record a candidate's voice for a screening assessment in the Philippines?

Recording is lawful when every party knows before it starts. RA 4200, the Anti-Wiretapping Act, is all-party and criminal, and the Supreme Court held in Ramirez v. Court of Appeals (1995) that even a participant who records without the other party's knowledge violates it — the governing word is secretly. An assessment that discloses on screen exactly what will be captured and requires acknowledgement before the attempt begins is compliant on that point, and replaying or forwarding the audio to a client is a separate act the Act also reaches. Note separately that the National Privacy Commission has said consent is a weak legal basis inside an employment relationship; the instrument-by-instrument treatment is on the Philippines compliance page.

Do other voice assessment vendors publish sample reports, and how do they compare?

Some do, and it is worth downloading theirs alongside this one. Checked at source on 13 September 2026: Pearson publishes an ungated two-page Versant English Test sample score report with four diagnostic subscores — sentence mastery, vocabulary, fluency and pronunciation — on the Global Scale of English with a CEFR equivalent, plus capability statements and improvement tips for each subscale; it does not name a norm group or reference population. SmallTalk2Me links a shared live report example rather than a PDF and describes thirty measured speech parameters with a CEFR level in half-level increments. iMocha's BPO assessment test page gives a ten-question, ten-minute screen with AI-enabled proctoring and no published sample report. When to choose them instead is straightforward: if a client contract or a regulator requires a certificated, independently validated language benchmark, buy a dedicated language instrument such as Versant rather than reading a platform-mapped level off a composed battery. What none of these publishes is an annotated walkthrough — what each panel is, how to read it, and what it hides — which is what this page is, and the eight questions above are written to be applied to their artefacts as readily as to ours.

Should a voice assessment verdict decide who gets hired?

No. The verdict on this report is one composite compared against one threshold the employer configured, with a fixed borderline band — it is built to rank a large applicant pool quickly, which it does well. A hiring decision needs the evidence panel, the interview stage the probes are written for, and a check that the threshold you set is not screening out a protected group at a rate you cannot justify. A screening platform can give you the first; the last two are the employer's work.