How to Compare Skills Assessment Platforms: TestGorilla, SHL, Mercer Mettl, iMocha and Pay-As-You-Go Alternatives
To compare skills assessment platforms, look past the test-library size and score each option on six dimensions: pricing model, how assessments are built, delivery friction, proctoring evidence, grading depth, and whether the platform measures training as well as hiring. Here is the framework, where the incumbents are strong, and when a pay-as-you-go alternative wins.
The short answer
Comparing assessment platforms by test-library size — the number most vendor comparison pages lead with — answers the wrong question. The platforms differ far more in how you pay, how an assessment gets built, what a candidate has to do before they can sit it, what evidence backs each result, and whether the platform can measure anything beyond a hiring funnel. Score your shortlist on those six dimensions and the right choice usually becomes obvious from your volume and use case, not from a feature grid.
The market splits into two pricing worlds. Most established platforms — TestGorilla, SHL, Mercer Mettl, iMocha and their peers — are generally sold as subscriptions or annual licences: you pay for access, sized by plan tier, seats or candidate bands. Pay-as-you-go platforms such as AssessAll instead meter per candidate assessed (1 credit = US$0.50, with a typical screen running a few to a few dozen credits) with no platform fee or annual contract. Neither model is universally better — the arithmetic depends entirely on your volume pattern, which is why it is dimension one below.
Dimension 1 — pricing model: subscription, licence, or metered
A subscription or annual licence is priced for continuous use: if you assess steadily all year at meaningful volume, access pricing amortises well and is easy to budget. It fits enterprises with permanent hiring programmes and dedicated TA teams.
Metered pricing is priced for bursty or uncertain use: seasonal hiring drives, project-based screening, a training provider whose client waves come and go, or a team running its first structured assessments and unwilling to commit before seeing results. On AssessAll the unit maths is public: credits cost US$0.50 each, a typical hiring screen runs roughly US$7.50–20 per candidate depending on depth, volume packs (250 / 1,000 / 5,000 credits at US$125 / 465 / 2,175) discount up to 13%, and new organisations get 250 free credits — so a pilot costs nothing. The comparison to run for your own case: estimate candidates assessed per year, multiply by per-candidate cost, and set that against the quotes you get for annual access. Whichever number is smaller at YOUR volume wins dimension one; there is a crossover point, and high-volume year-round programmes can sit on the other side of it.
Dimension 2 — how an assessment gets built
Library-first platforms give you a catalogue of pre-built tests — TestGorilla and iMocha are known for breadth here — which is fast when a stock test matches your role and constraining when it does not: your ability to measure a niche or hybrid role is bounded by what the catalogue anticipated.
Consultancy-grade providers such as SHL and Mercer Mettl sit at the other end: deeply validated, norm-referenced instruments, often configured with expert involvement — the standard for large regulated enterprises running formal psychometric programmes, with procurement timelines to match.
AI-composed platforms are the newer third option: describe the role or training brief in plain language and the platform composes the instrument — sections, items, scenarios and scoring rubric — in hours, alongside a ready catalogue for common needs. This is AssessAll's model (composition and grading run on Claude, Anthropic's AI), and it changes what is measurable: a hybrid role, a client-specific training-needs analysis, or a scenario set written in your industry's vocabulary stops being a custom-development project.
Dimension 3 — delivery friction: what stands between a candidate and the first question
Every account a candidate must create, app they must install, or desktop they must locate costs you completions — and the cost lands hardest in high-volume and mobile-first markets. If you screen walk-in applicants, job-fair queues, or candidates across the Philippines, the Gulf, Africa or India, delivery friction can matter more than any scoring feature.
The questions to ask any vendor: can a candidate go from link to first question on a phone with no account and no payment step? Can you hand delivery to a QR code at a venue? Does a dropped connection pause the attempt or destroy it? On AssessAll, delivery is a share link or QR code, participants create no accounts, the player is mobile-ready and attempts are resumable — that combination is the difference between screening everyone who applies and screening everyone who survived your funnel's friction.
Dimension 4 — proctoring: a score you can defend, or just a score
Since AI assistants became universal, an unproctored result mostly measures who had the better model open in the next tab. Most serious platforms now offer proctoring; the comparison point is what evidence you can actually inspect afterwards. A cheating 'flag' you cannot open is just another claim.
Look for identity capture at the start of the sitting, continuous monitoring for face absence and additional faces, tab-switch and screen-capture detection, and — critically — the raw evidence attached to each attempt so a human can adjudicate. AssessAll issues each attempt a High/Medium/Low integrity rating with the photo evidence and event log attached; whatever platform you choose, insist on that inspectability before you rely on remote results for real decisions.
Dimension 5 — grading depth: multiple choice is the floor, not the ceiling
Auto-scored multiple choice is table stakes on every platform. The differentiating layer is open response: written answers, spoken English scored from the candidate's own recorded voice, role-play and scenario responses. Platforms handle this with human review queues (slow, per-review fees), rubric-assisted marking, or AI grading at submission time.
AI grading changes the economics of measuring anything expressive: on AssessAll, written and spoken responses are graded by Claude on submit, so an AI-scored speaking task or a written judgement scenario costs the same workflow as a multiple-choice item — scores, per-section breakdowns and an AI-written strengths-and-risks summary arrive together in the report. If spoken English or written communication matters to your roles, weight this dimension heavily and ask every vendor exactly how — and how fast — open responses get scored.
Dimension 6 — beyond hiring: can it measure training?
Most assessment platforms are built around a hiring funnel and stop at the offer. If you also run L&D — or you are a training provider — the comparison widens: can the platform diagnose training needs before you spend, and prove capability gain after? That is a different product surface from a test library, and few hiring-first platforms have it.
On AssessAll this is native: TNA Studio composes a diagnostic from a plain-language brief and returns individual plus group reports splitting trainable gaps from structural ones, and Learning Journeys chain baseline, during-programme and outcome measurements so the capability gain is evidenced per person rather than asserted in a completion certificate. If your assessments serve training decisions, put this dimension first — it filters the market fastest.
Where the incumbents are strong — and when the alternative wins
An honest comparison cuts both ways. If you need norm-referenced psychometrics with decades of validation behind them for a regulated, high-stakes selection programme, SHL and Mercer Mettl are strong for exactly that, and an enterprise licence fits an enterprise programme. If your roles are standard and continuous, and a large ready-made library plus an annual plan suits your volume, TestGorilla and iMocha serve that model well. Vervoe and similar AI-graded platforms share part of the grading-depth story.
The pay-as-you-go, AI-composed alternative tends to win when volume is bursty or uncertain (pilots, drives, client waves), when roles or briefs do not match stock libraries, when candidates are mobile-first and account-creation friction is fatal, when spoken English and open responses must be scored at scale without per-review fees, and when the same platform must cover training measurement — TNA, learning journeys, capability evidence — not just hiring. Run your own numbers on dimension one, then let dimensions two to six break the tie.
Frequently asked questions
What is the best way to compare skills assessment platforms?
Score each option on six dimensions: pricing model (subscription or licence vs metered per candidate), how assessments are built (library, consultancy-configured, or AI-composed), delivery friction (accounts, apps and payment steps between candidate and first question), proctoring evidence you can actually inspect, grading depth for open and spoken responses, and whether the platform measures training as well as hiring. Volume pattern usually decides the pricing dimension; use case decides the rest.
What is a pay-as-you-go alternative to subscription assessment platforms like TestGorilla?
A metered platform where you pay per candidate assessed instead of for annual access. AssessAll is one: credits cost US$0.50 each, a typical screen runs a few to a few dozen credits (roughly US$7.50–20 per candidate), volume packs discount up to 13%, there are no seat licences or platform fees, and new organisations get 250 free credits to pilot with. Subscription platforms amortise better at continuous high volume; metered pricing wins when volume is bursty, seasonal or uncertain.
Are subscription assessment platforms ever the better choice?
Yes. If you assess at meaningful volume continuously all year, annual access pricing can cost less per candidate than metering, and enterprise programmes needing decades-validated, norm-referenced psychometrics — SHL's and Mercer Mettl's home ground — generally run on enterprise licences. The honest method is arithmetic: candidates per year × per-candidate metered cost, compared against the access quotes you receive. The crossover point is yours, not the vendor's.
What is an AI-composed assessment and how is it different from a test library?
A test library offers pre-built tests; you pick the closest match. An AI-composed assessment is generated from your plain-language brief — sections, items, scenarios and scoring rubric — so niche roles, hybrid roles and client-specific training diagnostics become measurable without custom development. On AssessAll, composition and the grading of written and spoken responses run on Claude, Anthropic's AI, alongside a ready catalogue for common roles.
How important is proctoring when comparing assessment platforms?
Decisive for any remote or high-stakes use. Since AI assistants became universal, unproctored results are unreliable, and a cheating flag you cannot inspect is just another claim. Compare platforms on inspectable evidence: identity photo at the start, monitoring for face absence and extra faces, tab-switch and screen-capture detection, and the raw evidence attached to each attempt — AssessAll surfaces this as a High/Medium/Low integrity rating with the evidence included.
Which assessment platforms also cover training needs analysis and learning measurement?
Few hiring-first platforms do — most stop at the offer. If assessments must also serve L&D, look for a training-needs analysis capability with individual and group reporting, and baseline-to-outcome measurement across a programme. On AssessAll these are native (TNA Studio and Learning Journeys), metered on the same credits as hiring screens, which is why training providers and L&D teams often weight this dimension first.