All articles
Assessment Science5 October 2026·6 min read

Applicants Score Half a Standard Deviation Higher on Conscientiousness. Here Is How to Audit Faking on Your Own Assessment.

Meta-analysis puts applicant score inflation at d = .52 on conscientiousness and .50 on emotional stability, and in one experiment Likert validity fell from .15 to .01 once the ideal-employee factor was removed. A seven-step audit you can run on your own applicant data.

By AssessAll Editorial

Faking an assessment means deliberately distorting self-report answers to match what the job appears to reward. Auditing it means comparing your applicant score distributions against low-stakes or incumbent data on the same instrument, scale by scale, and checking whether the gap is large enough to change who you shortlist. It is a data analysis, not a judgement about candidates.

Almost every organisation running a personality or behavioural questionnaire assumes some candidates inflate their answers. Very few have ever measured it on their own data. The measurement is not difficult, and the published evidence tells you roughly what to expect before you start.

What the gap actually looks like

The reference point for applicant inflation is Birkeland and colleagues' meta-analysis comparing job applicants with non-applicants on the Big Five. The gaps are uneven, and they are largest on exactly the two traits most selection systems weight most heavily:

  • Conscientiousness — Cohen's d = .52
  • Emotional stability — d = .50
  • Agreeableness — d = .19
  • Openness — d = .15
  • Extraversion — d = .13

Source: Birkeland et al., *International Journal of Selection and Assessment*, 2006.

That is the applicant-versus-incumbent gap, which mixes real population differences with distortion. The ceiling — what people can do when explicitly told to fake — is higher. Viswesvaran and Ones' within-subjects work put directed faking at .47 SD on agreeableness, .54 on extraversion, .76 on openness, .89 on conscientiousness and .93 on emotional stability.

A 2025 study of 448 Austrian teacher-education applicants closed the loop by measuring the same people twice, once in a real high-stakes admission exam and once in a low-stakes online study. Emotional stability moved d = 0.94; conscientiousness 0.30; agreeableness 0.19. Critically, intelligence positively predicted the size of the shift on conscientiousness and emotional stability (Weissenbacher, Jud & Krammer, *Frontiers in Psychology*, 2025). Faking is not random noise distributed evenly across your pool — the candidates best equipped to work out what you want are the ones who do it best.

Against this, the strongest counter-evidence is Hogan, Barrett and Hogan's study of over 5,000 rejected applicants who reapplied and retook the same inventory: virtually none of their scores changed beyond the standard error of measurement. Prevalence estimates sit around 30% of applicants on selection assessments, not everyone. So the honest summary is that inflation is real, concentrated on a few scales, unevenly distributed, and smaller in live hiring than in laboratory faking conditions.

Why "it doesn't hurt validity" is the wrong reassurance

The usual defence of self-report is that criterion validity survives faking. It does, on the surface — and the reason is the problem.

In a 652-participant experiment using modern forced-choice measures, Huber and colleagues found Likert scales retained a mean validity of .15 under fake-good instructions, higher than the forced-choice alternatives. But faking introduced a general "ideal employee" factor across all the scales. When that shared method variance was removed statistically, validity for the Likert scales fell from .15 to .01 (Huber, Kuncel, Huber & Boyce, *Personnel Assessment and Decisions*, 2021).

Read that carefully. The composite still predicts, because wanting to look like a good employee is itself mildly predictive. The individual dimensions stop meaning anything. If you report scale-level feedback, build development plans on profiles, or weight conscientiousness differently from agreeableness, you are acting on numbers that have lost their separate signal.

The audit: seven steps

1. Split your data by stakes. Pull applicant responses and any low-stakes responses on the same instrument — incumbent benchmarking, internal development use, practice administrations. Keep them separate from the start.

*2. Compute d per scale, not per instrument.* Standardised mean difference between the two conditions, one figure per scale. Compare against the Birkeland benchmarks above. A conscientiousness gap near .5 is unremarkable; a gap of 1.0 is a finding.

3. Look at the top of the distribution. Selection happens at the tail, not the mean. Compare the 90th percentile in each sample and count how many applicants cluster near the scale ceiling. A compressed top decile means your instrument has stopped discriminating exactly where you use it.

4. Check scale intercorrelations in both samples. If applicant scales correlate far more tightly with each other than low-stakes scales do, you are seeing the ideal-employee factor. That single check catches the problem Huber's study isolated.

5. Run measurement invariance if your samples are large enough. Configural, metric and scalar invariance across the two conditions tells you whether the scores are on the same metric at all. If scalar invariance fails, applicant and norm scores are not comparable and any norm-referenced percentile you report is wrong.

6. Do not "correct" with a social desirability scale. This is the most common intervention and the evidence against it is long-standing. Social desirability measures largely fail to capture applicant faking behaviour (Griffith & Peterson, *Industrial and Organizational Psychology*, 2008), and in the Huber experiment a social desirability correction "did not significantly improve convergence" with honest scores. A correction that does not work, applied to everyone, adds error rather than removing it.

7. Test the fix before you adopt it. Two fixes have real evidence behind them, and both are partial:

  • Warnings. Dwight and Donovan's meta-analysis put text-based warnings at d = .23 — roughly a 30% reduction in score inflation. Cheap, ethical, and not a solution on its own.
  • Multidimensional forced choice. In the Huber data, directed faking moved Likert scores by a mean of .81 SD and forced-choice scores by .27–.28. Under incentivised rather than instructed faking, the figures were .13 for Likert and .04–.08 for forced choice.

Unproctored remote testing is now a different problem

Generative AI has changed the risk profile of the unproctored self-report. Phillips and Robie tested 655 business students against large language models on high-stakes personality measures and found the models scored higher than the human participants, with forced-choice formats proving more resistant to manipulation than single-stimulus items (*Personality and Individual Differences*, 2024).

The practical implication is not to abandon self-report but to stop letting it carry a hurdle on its own in unproctored conditions. Two design responses hold up: move self-report to post-shortlist, informational use rather than a screening gate, and put the screening weight on evidence a model cannot produce on a candidate's behalf — work samples, AI-graded scenario responses tied to the actual job, and session-level integrity signals. Platforms including AssessAll report integrity bands alongside scores from AI proctoring precisely so the score and its conditions travel together rather than the score arriving alone.

When self-report is still the right instrument

Faking is a selection problem, not a measurement problem in general. For development work, team feedback, coaching and self-insight — where the respondent has no incentive to distort and often the opposite — a well-constructed Likert inventory remains the most efficient instrument available, and forced-choice formats trade away intuitive feedback for resistance the setting does not need. The same applies to low-stakes internal skills inventories. The cost of forced choice, longer questionnaires and warning language is only worth paying where a real decision hangs on the score.

Nothing here is a reason to drop personality assessment. The relevant professional standards are unambiguous that validity claims belong to a specific use in a specific population, which is exactly what a faking audit tests — see the *Standards for Educational and Psychological Testing* and the EEOC's guidance on employment tests and selection procedures.

The short checklist

  1. Separate applicant and low-stakes data on the same instrument.
  2. Compute d per scale and compare to published benchmarks.
  3. Examine the 90th percentile and the ceiling, not just the mean.
  4. Compare scale intercorrelations across conditions for an ideal-employee factor.
  5. Test measurement invariance before reporting norm-referenced percentiles.
  6. Delete any social desirability correction you are applying.
  7. Pilot warnings and forced-choice formats on your own data, and document the result.

Takeaway: the question is never "do candidates fake?" but "does the inflation on my instrument change who I shortlist?" That is a question your own data can answer in an afternoon, and until you have run it, every percentile you report to a hiring manager rests on an assumption you have not tested.

#faking#response-distortion#personality-assessment#forced-choice#measurement-invariance#psychometrics

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free