---
title: "How to Compare Skills Assessment Platforms: TestGorilla, SHL, Mercer Mettl, iMocha and Pay-As-You-Go Alternatives"
description: "To compare skills assessment platforms, look past the test-library size and score each option on six dimensions: pricing model, how assessments are built, delivery friction, proctoring evidence, grading depth, and whether the platform measures training as well as hiring. Here is the framework, where the incumbents are strong, and when a pay-as-you-go alternative wins."
canonical: https://www.assessall.com/guides/how-to-compare-skills-assessment-platforms
updated: 2026-09-13
source: AssessAll
---

# How to Compare Skills Assessment Platforms: TestGorilla, SHL, Mercer Mettl, iMocha and Pay-As-You-Go Alternatives

To compare skills assessment platforms, look past the test-library size and score each option on six dimensions: pricing model, how assessments are built, delivery friction, proctoring evidence, grading depth, and whether the platform measures training as well as hiring. Here is the framework, where the incumbents are strong, and when a pay-as-you-go alternative wins.

_Last updated 2026-09-13._

<!-- #the-short-answer -->
## The short answer

Comparing assessment platforms by test-library size — the number most vendor comparison pages lead with — answers the wrong question. The platforms differ far more in how you pay, how an assessment gets built, what a candidate has to do before they can sit it, what evidence backs each result, and whether the platform can measure anything beyond a hiring funnel. Score your shortlist on those six dimensions and the right choice usually becomes obvious from your volume and use case, not from a feature grid.

The market splits into two pricing worlds. Most established platforms — TestGorilla, SHL, Mercer Mettl, iMocha and their peers — are generally sold as subscriptions or annual licences: you pay for access, sized by plan tier, seats or candidate bands. Pay-as-you-go platforms such as AssessAll instead meter per candidate assessed (1 credit = US$0.50, with a typical screen running a few to a few dozen credits) with no platform fee or annual contract. Neither model is universally better — the arithmetic depends entirely on your volume pattern, which is why it is dimension one below.

<!-- #dimension-1-pricing-model-subscription-licence-or-metered -->
## Dimension 1 — pricing model: subscription, licence, or metered

A subscription or annual licence is priced for continuous use: if you assess steadily all year at meaningful volume, access pricing amortises well and is easy to budget. It fits enterprises with permanent hiring programmes and dedicated TA teams.

Metered pricing is priced for bursty or uncertain use: seasonal hiring drives, project-based screening, a training provider whose client waves come and go, or a team running its first structured assessments and unwilling to commit before seeing results. On AssessAll the unit maths is public: credits cost US$0.50 each, a typical hiring screen runs roughly US$7.50–20 per candidate depending on depth, volume packs (250 / 1,000 / 5,000 credits at US$125 / 465 / 2,175) discount up to 13%, and new organisations get 250 free credits — so a pilot costs nothing. The comparison to run for your own case: estimate candidates assessed per year, multiply by per-candidate cost, and set that against the quotes you get for annual access. Whichever number is smaller at YOUR volume wins dimension one; there is a crossover point, and high-volume year-round programmes can sit on the other side of it.

<!-- #dimension-2-how-an-assessment-gets-built -->
## Dimension 2 — how an assessment gets built

Library-first platforms give you a catalogue of pre-built tests — TestGorilla and iMocha are known for breadth here — which is fast when a stock test matches your role and constraining when it does not: your ability to measure a niche or hybrid role is bounded by what the catalogue anticipated.

Consultancy-grade providers such as SHL and Mercer Mettl sit at the other end: deeply validated, norm-referenced instruments, often configured with expert involvement — the standard for large regulated enterprises running formal psychometric programmes, with procurement timelines to match.

AI-composed platforms are the newer third option: describe the role or training brief in plain language and the platform composes the instrument — sections, items, scenarios and scoring rubric — in hours, alongside a ready catalogue for common needs. This is AssessAll's model (composition and grading run on Claude, Anthropic's AI), and it changes what is measurable: a hybrid role, a client-specific training-needs analysis, or a scenario set written in your industry's vocabulary stops being a custom-development project.

<!-- #dimension-3-delivery-friction-what-stands-between-a-candidat -->
## Dimension 3 — delivery friction: what stands between a candidate and the first question

Every account a candidate must create, app they must install, or desktop they must locate costs you completions — and the cost lands hardest in high-volume and mobile-first markets. If you screen walk-in applicants, job-fair queues, or candidates across the Philippines, the Gulf, Africa or India, delivery friction can matter more than any scoring feature.

The questions to ask any vendor: can a candidate go from link to first question on a phone with no account and no payment step? Can you hand delivery to a QR code at a venue? Does a dropped connection pause the attempt or destroy it? On AssessAll, delivery is a share link or QR code, participants create no accounts, the player is mobile-ready and attempts are resumable — that combination is the difference between screening everyone who applies and screening everyone who survived your funnel's friction.

<!-- #dimension-4-proctoring-a-score-you-can-defend-or-just-a-scor -->
## Dimension 4 — proctoring: a score you can defend, or just a score

Since AI assistants became universal, an unproctored result mostly measures who had the better model open in the next tab. Most serious platforms now offer proctoring; the comparison point is what evidence you can actually inspect afterwards. A cheating 'flag' you cannot open is just another claim.

Look for identity capture at the start of the sitting, continuous monitoring for face absence and additional faces, tab-switch and screen-capture detection, and — critically — the raw evidence attached to each attempt so a human can adjudicate. AssessAll issues each attempt a High/Medium/Low integrity rating with the photo evidence and event log attached; whatever platform you choose, insist on that inspectability before you rely on remote results for real decisions.

<!-- #dimension-5-grading-depth-multiple-choice-is-the-floor-not-t -->
## Dimension 5 — grading depth: multiple choice is the floor, not the ceiling

Auto-scored multiple choice is table stakes on every platform. The differentiating layer is open response: written answers, spoken English scored from the candidate's own recorded voice, role-play and scenario responses. Platforms handle this with human review queues (slow, per-review fees), rubric-assisted marking, or AI grading at submission time.

AI grading changes the economics of measuring anything expressive: on AssessAll, written and spoken responses are graded by Claude on submit, so an AI-scored speaking task or a written judgement scenario costs the same workflow as a multiple-choice item — scores, per-section breakdowns and an AI-written strengths-and-risks summary arrive together in the report. If spoken English or written communication matters to your roles, weight this dimension heavily and ask every vendor exactly how — and how fast — open responses get scored.

<!-- #dimension-6-beyond-hiring-can-it-measure-training -->
## Dimension 6 — beyond hiring: can it measure training?

Most assessment platforms are built around a hiring funnel and stop at the offer. If you also run L&D — or you are a training provider — the comparison widens: can the platform diagnose training needs before you spend, and prove capability gain after? That is a different product surface from a test library, and few hiring-first platforms have it.

On AssessAll this is native: TNA Studio composes a diagnostic from a plain-language brief and returns individual plus group reports splitting trainable gaps from structural ones, and Learning Journeys chain baseline, during-programme and outcome measurements so the capability gain is evidenced per person rather than asserted in a completion certificate. If your assessments serve training decisions, put this dimension first — it filters the market fastest.

<!-- #dimension-7-data-handling-which-got-materially-harder-in-202 -->
## Dimension 7 — data handling, which got materially harder in 2025–26

This dimension did not exist in most 2024 comparison checklists and now belongs on every one, because the rules moved. Malaysia's Personal Data Protection Act amendments, phased in through the first half of 2025, made breach notification to the Commissioner mandatory, required both controllers and processors to appoint a data protection officer, added biometric data to the sensitive-personal-data definition, replaced the old cross-border transfer whitelist with a risk-based test, and raised the maximum fine from RM300,000 to RM1,000,000 with imprisonment up to three years. Singapore's PDPC issued advisory guidelines on 1 March 2024 for personal data in AI recommendation and decision systems, pressing minimisation, pseudonymisation, and publishing your policy on such use rather than waiting to be asked.

Four questions follow, and they are worth asking every vendor on your shortlist in writing. Where are attempt records physically stored, and if that is outside the country you hire in, what is the transfer basis now that a whitelist is no longer available to point at? What exactly does the proctoring capture — a stored photograph a human reviews, or a face template used for matching — because that distinction is what determines whether you are handling sensitive personal data? What is the default retention period for scores, and separately for proctoring imagery, and can you set it? And is there a documented breach-notification path, given that notification is now a legal deadline rather than a courtesy?

None of these are exotic. What makes them useful in a comparison is that they separate vendors quickly: a platform that has thought about them answers in specifics, and a platform that has not answers with the word 'enterprise-grade'. AssessAll's own answers, and the instrument-by-instrument detail behind the questions, are set out on its [Singapore and Malaysia compliance page](https://www.assessall.com/solutions/fair-hiring-compliance-singapore-malaysia) — including the parts where the honest answer is that the judgement is yours and your advisers', not the vendor's.

<!-- #where-the-incumbents-are-strong-and-when-the-alternative-win -->
## Where the incumbents are strong — and when the alternative wins

An honest comparison cuts both ways. If you need norm-referenced psychometrics with decades of validation behind them for a regulated, high-stakes selection programme, SHL and Mercer Mettl are strong for exactly that, and an enterprise licence fits an enterprise programme. If your roles are standard and continuous, and a large ready-made library plus an annual plan suits your volume, TestGorilla and iMocha serve that model well. Vervoe and similar AI-graded platforms share part of the grading-depth story.

The pay-as-you-go, AI-composed alternative tends to win when volume is bursty or uncertain (pilots, drives, client waves), when roles or briefs do not match stock libraries, when candidates are mobile-first and account-creation friction is fatal, when spoken English and open responses must be scored at scale without per-review fees, and when the same platform must cover training measurement — TNA, learning journeys, capability evidence — not just hiring. Run your own numbers on dimension one, then let dimensions two to six break the tie.

<!-- #ask-for-the-artefact-not-the-demo-an-evaluation-step-most-sh -->
## Ask for the artefact, not the demo — an evaluation step most shortlists skip

A platform demo shows you the dashboard the vendor controls. The thing your organisation will actually live with is the report a hiring manager opens six weeks later, and it is the cheapest part of the evaluation to do properly: download the sample, put the same questions to every vendor's PDF, and see which ones survive.

Eight questions, all answerable by looking at an artefact rather than by asking a sales engineer. Does it name the norm group the scores are compared against, on the cover rather than in a footnote? Does any paragraph quote the candidate's own answers — the one panel that cannot be assembled from a lookup table? Is there a validity or response-quality check, and what does a flag actually *do*: lower a confidence band, or reject the person? Does the critical section cost anything, or is it flattery shaped like a critique? Is the advice specific enough to act on this week? **How many people is the comparison built from, and is that number printed?** **What does the integrity or proctoring summary print when nothing was recorded?** And does the report state what it does not measure?

The last three are the ones that separate vendors, and the reason is arithmetic rather than marketing. A percentile or a *vs norm* delta computed from a handful of prior attempts is a running average of whoever happened to sit first — on AssessAll's own voice report those chips appear only once eight prior attempts of the same battery exist, which is why the published sample carries none and says so. And some reports default to their most favourable integrity label when no proctoring data was recorded at all, which makes an unproctored sitting look like a clean one; the fix is to read the counters beside the badge, not the badge.

AssessAll reports are published in full and annotated panel by panel, with no form, among them the [Voice & Service Suitability report](https://www.assessall.com/guides/reports/voice-service-suitability-report). Each page carries the eight questions, and each is written to be used on a competitor's sample as readily as on ours — several publishers in this market post ungated samples, so a serious shortlist can run this comparison in an afternoon. The [full gallery is here](https://www.assessall.com/guides/reports).

<!-- #faq -->
## Frequently asked questions

### What is the best way to compare skills assessment platforms?

Score each option on six dimensions: pricing model (subscription or licence vs metered per candidate), how assessments are built (library, consultancy-configured, or AI-composed), delivery friction (accounts, apps and payment steps between candidate and first question), proctoring evidence you can actually inspect, grading depth for open and spoken responses, and whether the platform measures training as well as hiring. Volume pattern usually decides the pricing dimension; use case decides the rest.

### What is a pay-as-you-go alternative to subscription assessment platforms like TestGorilla?

A metered platform where you pay per candidate assessed instead of for annual access. AssessAll is one: credits cost US$0.50 each, a typical screen runs a few to a few dozen credits (roughly US$7.50–20 per candidate), volume packs discount up to 13%, there are no seat licences or platform fees, and new organisations get 250 free credits to pilot with. Subscription platforms amortise better at continuous high volume; metered pricing wins when volume is bursty, seasonal or uncertain.

### Are subscription assessment platforms ever the better choice?

Yes. If you assess at meaningful volume continuously all year, annual access pricing can cost less per candidate than metering, and enterprise programmes needing decades-validated, norm-referenced psychometrics — SHL's and Mercer Mettl's home ground — generally run on enterprise licences. The honest method is arithmetic: candidates per year × per-candidate metered cost, compared against the access quotes you receive. The crossover point is yours, not the vendor's.

### What is an AI-composed assessment and how is it different from a test library?

A test library offers pre-built tests; you pick the closest match. An AI-composed assessment is generated from your plain-language brief — sections, items, scenarios and scoring rubric — so niche roles, hybrid roles and client-specific training diagnostics become measurable without custom development. On AssessAll, composition and the grading of written and spoken responses run on Claude, Anthropic's AI, alongside a ready catalogue for common roles.

### How important is proctoring when comparing assessment platforms?

Decisive for any remote or high-stakes use. Since AI assistants became universal, unproctored results are unreliable, and a cheating flag you cannot inspect is just another claim. Compare platforms on inspectable evidence: identity photo at the start, monitoring for face absence and extra faces, tab-switch and screen-capture detection, and the raw evidence attached to each attempt — AssessAll surfaces this as a High/Medium/Low integrity rating with the evidence included.

### What data-handling questions should I ask an assessment vendor in 2026?

Four, in writing. Where are attempt records stored, and what is the cross-border transfer basis — Malaysia replaced its transfer whitelist with a risk-based test in the 2025 amendments, so pointing at a list is no longer an answer. What exactly does the proctoring capture, a stored photograph reviewed by a human or a face template used for matching, since biometric data is now sensitive personal data in Malaysia. What is the default retention for scores, separately for proctoring imagery, and can you change it. And what is the documented breach-notification path, now that notifying the Commissioner is a legal obligation rather than a courtesy. Vendors that have thought about this answer in specifics; vendors that have not answer with 'enterprise-grade'.

### Which assessment platforms also cover training needs analysis and learning measurement?

Few hiring-first platforms do — most stop at the offer. If assessments must also serve L&D, look for a training-needs analysis capability with individual and group reporting, and baseline-to-outcome measurement across a programme. On AssessAll these are native (TNA Studio and Learning Journeys), metered on the same credits as hiring screens, which is why training providers and L&D teams often weight this dimension first.

---

Source: https://www.assessall.com/guides/how-to-compare-skills-assessment-platforms
