All articles
Hiring Practice24 September 2026·6 min read

Candidates Quit in the First 20 Minutes. Shortening Your Assessment Won't Help.

Across 69 selection systems and ~250,000 applicants, assessment length had no relationship to withdrawal - 21% quit, most inside the first 20 minutes. Why shortening a validated test widens your error band without fixing drop-off, and what to fix instead.

By AssessAll Editorial

Assessment drop-off is the share of candidates who begin a pre-hire assessment and never finish it. Across 69 selection systems and nearly 250,000 applicants, about 21% did not complete — and the length of the assessment had no relationship to whether they quit. Drop-off is a first-ten-minutes problem, not a duration problem.

That finding is inconvenient, because "cut the assessment down" is the first thing almost every hiring team does when completion rates look bad. It is a fix that costs measurement precision and buys very little.

The evidence against the length theory

The largest direct test comes from Hardy, Gibson, Sloan and Carr (2017) in the Journal of Applied Psychology. They analysed roughly a quarter of a million applicants across 69 selection systems in 29 organisations, with a mean assessment length of 67 minutes. Two results matter:

  • Assessment length was not related to withdrawal once candidates had started.
  • Assessment length was not related to whether candidates accepted the initial invitation to take the assessment in the first place.

What did concentrate was the timing of quitting. The majority of applicants who abandoned an assessment did so inside the first 20 minutes, and lengthening an assessment beyond that point did not add incremental attrition. A 30-minute assessment and a 60-minute assessment lose most of the same people, for the same reasons, in the same opening stretch.

Van Iddekinge, Lievens and Sackett (2023), reviewing the selection literature in Personnel Psychology, draw the direct conclusion: the evidence does not support sacrificing validity or reliability in order to lift completion rates through shorter assessments.

The people you lose are not the people you think

The unspoken fear behind shortening is that long assessments drive away the strongest candidates — the ones with other offers and no patience. The data points the other way.

Hardy and colleagues (2020), writing in the International Journal of Selection and Assessment, examined over two million job applicants across distribution centre, retail, bank manager and urgent-care nursing roles. Candidates who abandoned an assessment had lower scores on the components they had already completed than candidates who finished. Abandonment behaved as a weak negative signal about capability, not a filter that selectively removed top talent.

That does not make drop-off costless — it is still wasted funnel, wasted advertising spend, and a poor experience for people who deserved better. But it undercuts the premise that every abandoned assessment is a lost star candidate.

Applicant reactions research points the same way. In Speer, King and Grossenbacher (2016), Journal of Personnel Psychology, applicants who took the longer cognitive assessment reported more favourable reactions than those who took the shorter version. Length is not what candidates are reacting to.

What shortening actually costs

Reliability scales with test length in a way you can calculate before you cut anything. The Spearman–Brown prophecy formula gives the expected reliability of a test whose length changes by a factor of n:

Halving a well-built test

  • Starting reliability: α = .90
  • Reliability after halving: α = .82
  • Standard error of measurement, on a scale with SD = 10: rises from 3.16 to 4.27 points
  • 95% confidence band around a candidate's score: widens from roughly ±6.2 to ±8.4 points

A ±8.4-point band means that around a cut score of 70, candidates scoring anywhere from about 62 to 78 are statistically indistinguishable from the boundary. You did not make the assessment friendlier. You made the decision blurrier — and you widened the zone where measurement error, not capability, decides who advances.

That has consequences beyond accuracy. Under the Uniform Guidelines on Employee Selection Procedures and the EEOC's guidance on employment tests, a selection procedure that produces adverse impact must be justified by validity evidence. Shortening a test does not suspend that obligation; it erodes the reliability ceiling your validity evidence sits on. The *Standards for Educational and Psychological Testing* treat score precision as a property you are expected to report, not a detail you quietly trade away.

What actually drives drop-off

If length is not the lever, what is? The literature points at friction, redundancy and unexplained purpose — all of which live in the opening minutes.

Redundancy. Hartwell, Orr and Edwards (2020), IJSA 28, 200–208, tested what happens when an employer removes duplicated content from an online application — the classic "upload your CV, now retype your CV" pattern. Attrition fell, and applicant quality did not.

Perceived job-relatedness. The meta-analysis by Hausknecht, Day and Thomas (2004) in Personnel Psychology found applicants react more favourably to interviews, work samples and résumé screens than to personality inventories, integrity tests and biodata — and that favourable reactions are associated with job-offer acceptance intentions and willingness to recommend the employer. Candidates stay for tasks that visibly resemble the job.

Explanation and consistency. The SIOP white paper on applicant reactions (Bauer, McCarthy, Anderson, Truxillo and Salgado, 2012) identifies the consistently validated levers: job-relatedness, opportunity to demonstrate ability, consistent treatment, explanations for what is being measured, and basic respect. None of those require a shorter test.

Device and conditions. Aggregate vendor data published by HireVue in 2025 reports that desktop and tablet users complete at higher rates than mobile users by several percentage points — a difference of rendering and interruption, not stamina. In India and across BPO and frontline hiring more broadly, where the majority of candidates start on a phone, this is the single most expensive unforced error in assessment design.

A checklist for the first ten minutes

  1. Instrument the drop-off curve before changing anything. Log the timestamp of the last interaction for every non-completer. If your losses cluster before minute ten, length is not your problem.
  2. Delete every field the candidate has already given you. Parsed CV data, contact details re-requested at a second step, duplicate demographic forms.
  3. Open with the most job-like task you have. A scenario, a work sample, a short simulation — not a demographic form and not the least engaging item in the bank.
  4. State the contract on screen one. How long it takes, how many sections, what is measured, whether it can be paused, and what happens to the result.
  5. Make pause-and-resume real for anything over 20 minutes, and say so upfront.
  6. Test the whole flow on a mid-range Android phone on a weak connection — not on your laptop.
  7. Explain the integrity requirements before the camera turns on, not at the moment it does. Unannounced proctoring is a spike in early abandonment. Tools that report an integrity band rather than a pass/fail verdict — the approach AssessAll uses — are easier to explain honestly to candidates because the output is a flag for human review, not an automatic rejection.
  8. Re-check the curve after each change, one change at a time.

When shortening is the right call

Sometimes it genuinely is. Cut length when the assessment contains sections with no validity evidence behind them; when two instruments measure the same construct and you are paying for both; when item analysis shows a large block of items contributing nothing to score variance; or when the role is a high-volume, low-margin hire where a 60-minute screen cannot be justified against the cost of a mis-hire. Those are all decisions about content, made on evidence. They are different from trimming a validated assessment by a third because a dashboard showed 70% completion and somebody wanted a rounder number.

Per-candidate pricing changes this calculus too. When assessments were bought as annual seat licences, teams over-assessed to justify the licence. Usage-based models — AssessAll charges ₹30 / US$0.50 per assessment credit, with free credits on signup — make it cheaper to run a short, well-targeted screen first and a deeper assessment only on the candidates who clear it. That is multiple-hurdle design, and it reduces average candidate time without reducing the precision of any single decision.

The takeaway

Drop-off is a design problem located in the first ten minutes of the candidate's experience — redundancy, unexplained purpose, bad mobile rendering, a dull opening — and shortening a validated assessment treats none of them while measurably widening the error band around every decision you make. Fix the opening, not the duration.

#candidate-experience#assessment-drop-off#completion-rates#selection-design#reliability#applicant-reactions

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free