Applied Judgment Assessment
Applied skill assessment · software engineers, tech leads and engineering managers · browse the full catalogue

Code Review Assessment for Software Engineering TeamsFlagging everything is not thoroughness. It is a review that gets answered in bulk.

Forty-two exercises priced by what each mistake actually costs, with the price list printed on your report.

40 minutes42 scored exercisesEvidence-keyed scoringGlobal · INR & USD

Every team reviews code. Almost none of them measures whether anybody is good at it.

The two failure modes look nothing alike and both are expensive. One reviewer approves in ninety seconds and lets a change through that drops a column another service still reads. The other leaves eleven comments about naming, so the twelfth, which was the real one, gets answered in bulk with the rest.

Forty-two exercises put changes in front of you in plain words: a cache key that quietly drops the tenant, a retry wrapped around a payment call, a delete that runs before its confirmation is wired up, a one-line fix in a file nobody has touched for two years. You decide what is worth a comment, what is worth blocking, and how the comment should be written.

No programming language is needed anywhere. Every change is described in ordinary English, because what is being measured is the judgement rather than the syntax, and a reviewer who is good at this is good at it in any language.

Scoring is a decision-cost matrix. Missing a defect that destroys data and spending a comment on a variable name are priced differently, and the price list is printed on your report with your own counts against every row. It encodes an opinion about what a defect is worth, so it is published rather than buried in the scoring code.

Every claim the report makes is footnoted to the specific thing you selected or left alone. That turns the report from an assertion into an argument you can disagree with line by line, and the receipts are on the page rather than in a hover.

Four parts, each measured across the whole form:
What is worth raisingWhat is worth leaving aloneHow the comment landsWhen to block and when to approve

What you walk away with

Cost avoided, not accuracy

How much of the available cost you kept off the table, against a reviewer commenting the way people typically do.

Which way you lean

Findings left unsaid on one side and comments spent where none was needed on the other, both counted in words.

The cost table, printed

Seven classes with what each costs to miss, and your own counts against every row.

Receipts

Every claim footnoted to the specific option you selected or left alone, as endnotes rather than tooltips.

Four parts with their shadows

Including what a reviewer who reliably finds the expensive thing costs a team in queue time.

One if-then change

Drawn from the pattern your own sitting showed most clearly, not from a list of good habits.

Inside your report

Illustrative sample — your report is generated from your own responses.

Which way your reviewing leans
3 unsaid5 spentsolid bar = cost left unsaid · dashed bar = comments not needed

Missing a finding and spending a comment on a preference are different problems with different fixes. One accuracy percentage would have hidden both.

The cost table, printed with your own counts
If nobody comments on itCostOn your formYou caught
Data destroyed or unrecoverable922
Somebody sees what they should not832
A failure path nothing would catch453
Harder to change safely later241
Read once by somebody on a Tuesday160

The table encodes an opinion about what a defect is worth, so it is published rather than buried in the scoring code. Somebody could argue for different numbers, and now they can.

Built for

  • Software engineers at every level who review other people's changes
  • Tech leads and engineering managers deciding how review time is spent
  • Teams where AI-generated changes have made review the bottleneck
  • Anybody who has watched a real finding get answered in bulk with eight naming comments

Find out where your review attention is actually going

42 exercises across six formats · about 40 minutes · no language required · the cost table printed.

₹1,199 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and Razorpay sends your receipt to that address.

Secure Razorpay payment · ₹1,199 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

Do I need to know a particular programming language?

No. Every change is described in ordinary English. What is measured is which findings you raise, which you leave, when you block, and how you write the comment — and that judgement is the same in any language.

Why is commenting on everything scored badly?

Because attention is the budget. A review with twelve comments is read as noise and answered in bulk, so the finding that mattered arrives with the same weight as the naming. The cost table prices both errors and the score is the cost avoided.

Where do the costs come from?

They are a value judgement, and that is said on the report. The table says losing data costs about nine times what a naming preference costs. Somebody could argue for different numbers and get a different score, which is why it is printed with your own counts against it.

Is this a hiring test?

It is not a hiring decision on its own, and it does not measure programming ability or the quality of your own code. It is a developmental measure of how a review's limited attention is spent.

How long is it and what does it cost?

About forty minutes for forty-two exercises. ₹1,199 in India, inclusive of GST, or US$11.99 elsewhere, one time, for the sitting and the full report.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue

Methodology: Forty-two original exercises across six formats and no situational judgement item at all: twelve select-every-that-applies review budgets, twelve keyed choices, six keyed claims, four matching exercises, four ordering exercises and four most-and-least forced choices. The declared response instruction is knowledge and applied judgement throughout - what a reviewer should do here - with the forced choices asking what the respondent would most and least likely do, reported beside rather than inside the composite so the two instructions are never averaged. Construct statement: it measures how somebody spends a review's limited attention, which findings they raise, which they leave, when they block, and how they write a comment so it is acted on. It does not measure programming ability, it does not test any language, framework or tool, it is not a hiring decision on its own, and it says nothing about how good the respondent's own code is. Scoring is a decision-cost matrix. Every option in every review budget carries a class - data loss, security, correctness, test gap, maintainability, style, or not a defect - and each class carries the cost of missing it and the cost of spending a comment on it. The score is the cost avoided as a share of the distance between a reviewer commenting at the item's own authored answer priors and a reviewer who spends attention perfectly, so commenting on everything and commenting on nothing both land near zero rather than one of them scoring well by volume. The cost table is a value judgement rather than a measurement, so it is published in full on the report with the respondent's own counts against each row. Constructs and sources: the empirical literature on code review effectiveness and the sharp fall in defect detection as change size grows (the Cisco and SmartBear review study; Rigby and Bird on convergent contemporary review practices); modern code review as communication rather than defect hunting (Bacchelli and Bird on expectations and outcomes; Sadowski and colleagues on review at scale); the effect of review comment tone and volume on author behaviour and on people leaving projects (Bosu and colleagues on useful review feedback; Egelman and colleagues on interpersonal conflict in code review); signal detection and the cost of a false alarm as distinct from a miss (Green and Swets); decision analysis with an explicit loss matrix rather than an accuracy percentage; the reversibility of a decision as the basis for blocking (the one-way and two-way door distinction); implementation intentions for the single change the report recommends (Gollwitzer and Sheeran); and item-writing guidance from Haladyna, Downing and Rodriguez. All items are original works written for this instrument. No company's review guidelines, checklist or training material is reproduced, named or implied.