The AssessAll Blog
Assessment science, decoded weekly.
Measurement, hiring, proctoring, and the future of skills — written for people who make talent decisions.
Reliability vs Validity: Your Assessment Can Be Perfectly Consistent and Completely Wrong
Reliability tells you whether a score repeats; validity tells you whether it means what you claim. A practitioner comparison of the two checks, with the maths on standard error of measurement, attenuation ceilings and what validity evidence actually requires.
Read article70% Is Not a Cut Score: How to Set a Passing Mark You Can Defend
A cut score is a policy judgment, not a property of the test. A step-by-step guide to setting a defensible passing mark: modified Angoff panels, selection-ratio reality checks, measurement error at the boundary, and an adverse-impact check before launch.
No Single Right Answer, Real Predictive Power: The Science of Situational Judgment Tests
Five decades of research and a 2025 review of 524 studies show SJTs predict job performance with smaller group differences than cognitive tests. What they measure, why instruction wording matters, and how AI-scored open-response formats change the economics.
Can You Trust an AI-Graded Answer? What 2026's Scoring Studies Actually Show
Three 2026 studies settle the AI grading question: hybrid LLM scoring now matches trained human raters on open-response tests and beats them on consistency. What agreement benchmarks, bias checks, and architectures to demand before trusting machine-scored assessments.
Half the Questions, Twice the Signal: How Adaptive Testing Works
Computerized adaptive testing reaches the same measurement precision as fixed tests with roughly half the questions. How CAT engines work, what the evidence shows, and why test length is quietly killing your hiring funnel.
How Biased Is AI Assessment, Really? What 2026's Landmark Audits Reveal
A Stanford-led audit of 4 million applications and 150+ independent bias audits give the first real answer on AI assessment bias: fairer than humans on average, yet failing in one in ten jobs. How to audit your own stack.
AI Proctoring in 2026: What Integrity Bands Actually Measure
Remote assessment is only as good as its integrity signal. Here is how modern AI proctoring works — identity baselines, face detection, screen monitoring — and why a High/Medium/Low band beats a binary "cheated / didn't".
Want this thinking applied to your organisation?
Talk to the team behind the platform.
solutions@bodhih.com