Applied Judgment Assessment
Applied skill · software engineers, data engineers and the people who lead them · browse the full catalogue

Debugging Judgment Assessment with a Mid-Task Requirement ChangeThe spec moved at rung five. Half your work survived. Which half?

Two trails of eight decisions through a real debugging task. Partway through each one the requirement changes — and half of what you have done still stands.

50 minutes40 scored exercisesEvidence-keyed scoringGlobal · INR & USD

Nothing else changes the requirement halfway through, which is the only way to measure this

Two things end more engineering work than any bug does. One is defending a piece of work because of how much went into it. The other is throwing the whole thing away because something moved. They feel completely different from the inside and they cost about the same, and no assessment on the market can tell them apart — because measuring them requires a requirement that actually changes while somebody is working, and every instrument in this category is a static list of questions.

This one changes it. Two trails, eight decisions each. A test suite that passed yesterday and does not today. A nightly import that has been writing duplicate rows for a fortnight. At every rung you are told exactly where the work has got to, and you choose what you would do next. Then at rung five the product owner arrives: discounts now apply before tax rather than after; the feed will start sending updates as well as new records. Neither change is negotiable and neither invalidates everything.

What follows is balanced on purpose. Three of the six pieces of prior work survive each change and three do not, so carrying on regardless and clearing the desk both score at chance — as arithmetic, not as a threshold. The report names the two directions in the words of the work and counts them separately, because they need opposite advice and a single figure would call them the same.

The trails are fixed-state and the report says so in plain words. The position at every rung is written into the exercise and is identical for everybody, because a fixed item bank cannot branch. That matters: if it responded to your answers, a shared rung would be a consequence of your earlier choice and reading it otherwise would be reading a fact about our pipeline as a fact about you. And every rung is scored against the moves available at that rung, so the best move out of a bad position earns full marks. Recovery is measured rather than punished, and a running total would have counted one early error four more times.

Twenty-four further exercises come at the same ground from other directions: what you want in hand before changing a line, what belongs in a change and what belongs in its own, what makes a change hard to undo, and what an honest reason to keep work looks like next to the three arguments that are the sunk cost wearing different clothes. No code is written and none is read, which is why this works for an engineer in any language.

Four parts of one judgement, each measured by at least eight independent exercises:
Establishing the fault firstThe smallest change that holdsReading what actually changedLetting go of what no longer counts

What you walk away with

Your path, drawn against what was available

Both trails as a line, with the ceiling at the best move available at each rung — the only visual that makes recovery credit legible.

Before and after the change, separately

Your level on each trail either side of the requirement moving, so holding your ground across a change reads as the finding it is.

Carried on regardless, and cleared the desk

Counted separately, on a balanced set where both habits score at chance. Opposite errors, opposite fixes, and nobody else measures either.

A stepped table of every rung

What you chose, what the best available move was, and what it was worth. The path with its receipts underneath it.

One if-then change

Chosen from your own lean, naming a specific situation and a specific behaviour rather than a competency to work on.

The fixed-state disclosure

Printed in the report rather than assumed, because a reader who thinks the trail responded to them will read a shared rung as their own consequence.

Inside your report

Illustrative sample — your report is generated from your own responses.

Your path, against the best move available at each rung
best move available therethe requirement changes12345678round = before the change · square = after it

The ceiling is the best move available at that rung, which is what makes recovery from a bad position worth full marks rather than nothing.

What you did when the requirement moved
Carried on regardless
2
Kept work the change had already invalidated
Cleared the desk
1
Dropped work the change never touched
Best move after
5
Of six post-change rungs

Three of the six pieces of prior work survive each change and three do not, so keeping everything and starting over both score at chance.

Before and after, on the same trail
Before the change
83%
After it
89%
Every rung is scored against the moves available at that rung, so holding your level across the change is a finding rather than the change being easy.

A running total would have counted one early error four more times and called it your judgement.

Built for

  • Software and data engineers at any level, in any language
  • Tech leads and engineering managers assessing judgement rather than syntax
  • Teams whose requirements move mid-sprint and want to know how people handle it
  • Anybody who has finished a piece of work that stopped being needed halfway through

Two trails, sixteen decisions, one requirement change each

40 exercises across five formats · about 50 minutes · both trails drawn as a path, recovery credit made visible, and the two directions of the error counted separately.

₹999 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and Razorpay sends your receipt to that address.

Secure Razorpay payment · ₹999 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

Do I have to write code?

No. Every exercise is written in plain English about a described fault, and no code is written or read. That is what makes it work for an engineer in any language and for the people who lead them, and it is why it measures judgement rather than syntax.

Do the trails respond to my answers?

No, and the report says so plainly. A fixed item bank cannot branch, so the position at every rung is authored and identical for everybody. Each rung is scored against the moves available at that rung, which is what makes the best move out of a bad position worth full marks. If the trail did respond, a shared rung would be a consequence of your own earlier choice, and reading it that way would be reading a mechanical fact about our pipeline as a fact about you.

How can you measure sunk cost in an assessment?

By changing the requirement partway through and balancing what follows. Three of the six pieces of prior work survive each change and three do not, so keeping everything and dropping everything both score at chance. The two directions are counted separately and named in the words of the work rather than as over and under.

What is recovery credit?

Each rung is scored against the value of the moves available at that rung, not against the whole trail. So a strong move out of a weak position earns full marks there. Without that, one early error propagates and the instrument measures rung one eight times over and presents it as your failure.

How long is it and what does it cost?

About fifty minutes for forty exercises. ₹999 in India, inclusive of GST, or US$9.99 elsewhere, one-time. Nothing is timed rung by rung and no part of the score uses how fast you answered.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue

Methodology: Forty original exercises across five formats. Sixteen are the rungs of two fixed-state trails through a real debugging task, eight rungs each; the other twenty-four are eight keyed single-choice items, eight select-every-that-applies items, four true or false claims and four ordering exercises. Construct statement: it measures whether somebody establishes a fault before changing anything, makes the smallest change that restores the behaviour without breaking what already passed, reads what a changed requirement does and does not touch, and lets go of work the change has invalidated without discarding the work it has not. It does not measure programming language knowledge, algorithm design, system design, speed, or any particular technology. No code is written and none is read. Declared response instruction: behavioural tendency throughout the two trails - each rung asks what the respondent is most likely to do. The twenty-four supporting exercises are keyed knowledge items and are reported as their own strands rather than folded into the trail score. The trails are FIXED-STATE and the report says so in plain words. A static item bank cannot branch, so the position at every rung is authored and is identical for every respondent, and each rung is scored against the moves available at that rung rather than against the whole trail. That is what makes recovery real: the best move from a bad position earns full credit at that rung, so one early error is not counted four more times. Without that normalisation the instrument would be measuring rung one, four times over, and presenting it as the respondent's failure. The drift: at rung five of each trail the requirement changes, and it changes in a way that invalidates part of the work and leaves the rest standing. The keep-or-drop exercises are balanced - half the pieces of prior work survive the change and half do not - so carrying on regardless and starting over both score at chance. The two directions are reported separately and are named in the language of the work rather than as over and under. Construct grounding, drawn across sources rather than from one framework: the sunk-cost effect and escalation of commitment (Arkes and Blumer, 1985; Staw, 1976), and the finding that it survives expertise and is reduced by separating the decision from the person who made it; the psychology of debugging, in which the dominant failure is changing code before the fault has been located (Zeller, Why Programs Fail, 2009; Gould, 1975); delta debugging and the value of the smallest change that reproduces or removes a behaviour (Zeller and Hildebrandt, 2002); the distinction between a slip and a mistake (Reason, Human Error, 1990); confirmation bias in hypothesis testing (Wason, 1960) as the reason a fix that passes is not a fix that is understood; requirements volatility and the measured relationship between late change and defect density (Boehm, 1981; Curtis, Krasner and Iscoe, 1988); psychological safety as the condition under which a team can abandon work in progress (Edmondson, 1999); the path-dependence problem in scoring a sequence of decisions, and normalisation against the moves available at each node; the criterion validity of work-sample and situational-judgment methods (Roth, Bobko and McFarland, 2005; McDaniel and colleagues, 2007); McDonald's omega in place of alpha (McDonald, 1999); and the standard error of measurement as the reason a band rather than a point is reported (AERA, APA and NCME Standards, 2014). All items are original works. No item, scale name or report section is taken from any commercial instrument, and no affiliation with or endorsement by any vendor is claimed or implied.