Applied Judgment Assessment
Leadership Battery · tech leads and new engineering managers · browse the full catalogue

Engineering Manager Judgment Assessment for Tech Leads and New Engineering ManagersThe calls where the code is not the problem. Which way you lean, and what each lean costs.

Forty-four exercises for tech leads and new engineering managers, twenty of them forced choices: a senior engineer who has gone quiet, a date agreed without the team, a pull-request war that arrives from both sides, a rumour nobody official will confirm, a fix you could do yourself in twenty minutes. Four moves each time, none of them the obviously nice one; you mark the one you would most likely make and the one you would least likely make. The report opens with what it did not measure, then draws five leans as bands against the ordinary respondent, each paired with its honest cost. No total, no type, no percentile.

40 minutes44 scored exercisesEvidence-keyed scoringGlobal · INR & USD

The first engineering-manager assessment built around what it refuses to say

The Engineering Manager Judgment Assessment is a forty-minute forced-choice battery for tech leads and new engineering managers. It measures the lean a manager brings to five calls where the code is not the problem, reports each as a band against the ordinary respondent with its cost attached, and opens with what it did not measure.

Most of what a new engineering manager gets wrong is not technical, and most instruments for managers either test general first-line management or hand back a type. This does neither. Five leans are measured, and no end of any of them is scored as right: taking a problem on yourself or routing it; shielding the team from the organisation or carrying its asks in with their reasons; holding the date or holding the quality; raising a problem upward early or keeping it in the team; making the technical call yourself or fixing how the team decides. Each lean gets you something and costs you something, and the report says both, in sentences written so that a manager who leans the other way would not accept them as theirs.

Twenty of the forty-four exercises are forced-choice blocks: a situation, four moves, one on each of four leans, written in four registers and matched on how good they sound, so that there is no nice answer and no decisive-sounding answer to pick every time. You choose the move you would most likely make and the one you would least likely make. Because every block gives one most and one least across different leans, a profile high on everything is arithmetically impossible, and the report prints the sum to show it. Faking resistance is designed against the effect sizes the applied literature actually finds in real selection settings, not the inflated ones from laboratory instructions, and the page claims no more than that.

Each block also carries a key: the move the published evidence on new managers, incident practice, technical debt and bad news on troubled projects says works in that situation, and the one that most reliably backfires. Eight scenario questions, three select-every-one exercises and three match-the-following exercises add to that. Together they score fit, apart from every lean and never added to one: a lean says which way you go, fit says whether the way you went was the one the situation rewarded. Ten short ratings, two per lean and one drawn each way, ask which way you think you lean, and the report sets that beside what you actually pressed, in words only, because two ratings carry a direction and not a figure.

Every figure is chance-corrected against the way people ordinarily answer these blocks, declared on every item and replaced by observed data once it exists, so zero is the ordinary respondent and not a coin. Every band uses the wider of two standard errors, one from the assumed reliability and one from press noise alone, and a lean is only called when the whole ninety-five per cent band sits on one side of zero. A lean too thin to carry a figure carries a placement word instead, with the reason printed where the figure would have gone. The competence descriptors are anchored on the European e-Competence Framework, EN 16234-1, a published standard cited as such; nothing here is endorsed by its custodians.

The report is built around its refusal. It opens, in the largest type on the page, with what the instrument did not measure: not your technical ability, not your intelligence or integrity, not whether you are a good person or a good manager, not how you would do in a job it has not described. Then the five bands with the ordinary respondent marked and named on each, the cost of every lean shown, the contributors under every band, fit apart, the stated lean beside the pressed one, one boxed if-then change with the movement a re-sitting would have to clear to count as real, the reliability table with every refusal printed, and what zero means. At Rs 1,699 inclusive of GST in India, or US$16.99 elsewhere, that report is the product.

Five leans as bands, fit apart, and a refusal where a score would have gone:
Taking it on yourselfShielding the teamHolding the dateRaising it upwardMaking the technical call

What you walk away with

What it did not measure, first

The page opens with the refusal in the place a composite score would have taken: not technical ability, not intelligence or integrity, not whether you are a good manager, not performance in a job it never described. It is the largest type on the page and never a footer, because what the instrument does not say is the first thing you need before reading what it does.

Five leans drawn as bands, not points

Taking it on yourself, shielding the team, holding the date, raising it upward, making the technical call. Each drawn against the ordinary respondent with the 68 per cent band as a solid block and the 95 per cent band as an outlined box, both pole names printed at the ends, and the band written out in words. A lean is called only when the whole band sits on one side of zero.

The honest cost of every lean you showed

A strong pull toward taking it on yourself makes you the slowest path to every fix within a quarter; a strong pull toward routing means you learn what the work is like only when it fails. Every lean the report calls carries its cost in a sentence a manager leaning the other way would not accept as theirs.

Fit, scored apart

Whether the move you chose was the one the evidence supports in that situation, over thirty-four keyed exercises, chance-corrected so that zero is the ordinary respondent. It is never added to a lean: a strong lean can fit well or badly, and no lean at all can fit either way.

The lean you said beside the lean you pressed

Ten short ratings, two per lean and one drawn each way, ask which way you think you lean. The report sets each beside what your presses showed, in words and by direction only, and names the leans where the two point opposite ways, which is the most useful line on the page.

One change and the movement that would count

One boxed if-then sentence naming a specific situation and a specific replacement behaviour, drawn from the lean furthest from ordinary on your sitting, with the number of points a re-sitting would have to move to count as real change rather than measurement error.

Inside your report

Illustrative sample — your report is generated from your own responses.

The page opens with the refusal, in the place and the type size a score would have had
Read this first: what this report does not say

This did not measure whether you are a good manager.

  • It did not measure your technical ability, your coding skill or your system design knowledge. No item asked you to read, write or review code.
  • It did not measure your intelligence or your integrity. Nothing in a forced choice between four defensible moves can.
  • It did not measure whether you are a good person or a good manager. There is no total anywhere on this page, and the five bands below do not add up to one.
  • It did not measure how you will do in a job it has not described. It measured which way you leaned in twenty forced choices, in one sitting, on one day.
  • It gave you no type and no percentile. A lean is a band, not a box, and you are ranked against nobody.

What it did measure: five leans, each a band against the ordinary respondent, each paired with the cost of the pole you showed; apart from them, whether the moves you chose were the ones the evidence supports.

Why nobody can be high on everything here

Every forced choice gave one MOST and one LEAST across four different leans, so a press toward one lean is a press away from another. The five raw sums add to zero on any fully answered sitting.

Own hands+7Shield-2Date-6Raise+3The call-2left = toward the low pole · right = toward the high pole · total 0
Tabular fallback for the raw press sums
Own hands +7Shield -2Date -6Raise +3The call -2total 0

How to read it: the refusal sits where a composite would have gone, in the largest type on the page, because what the instrument does not say is the first thing to know. The arithmetic beneath it is why there is no total: the five leans cannot add up.

Three of the five leans, each a band against the ordinary respondent, furthest from ordinary first, no total
▲ leans toward the high pole · ▼ leans toward the low pole · ▬ no lean shown (the 95% band includes zero) · solid block = 68% band · outlined box = 95% band · dashed tick = ordinary respondent
1. Taking it on yourself 16 of 16 blocks answered
Leans toward: take it on yourself
0 = the ordinary respondentyou +48-100 = route it to its owner, on every block+100 = take it on yourself, on every block

About two chances in three that a re-sitting today would land between 32 and 64; nineteen in twenty between 17 and 79. Zero is the ordinary respondent.

What it costs, in the same manager

The cost of taking it on yourself: the people whose job it was learn that hard things go to you, and within a quarter you are the slowest path to every fix and the person who cannot be away.

2. Raising it upward 16 of 16 blocks answered
Leans toward: keep it in the team
0 = the ordinary respondentyou -36-100 = keep it in the team, on every block+100 = raise it early, on every block

About two chances in three that a re-sitting today would land between -52 and -20; nineteen in twenty between -67 and -5. Zero is the ordinary respondent.

What it costs, in the same manager

The cost of containing: the day a contained problem becomes a visible one, your director learns about it late and from someone else, and every quiet week before that is re-read as concealment.

3. Holding the date 16 of 16 blocks answered
No lean shown
0 = the ordinary respondentyou +9-100 = the quality holds, on every block+100 = the date holds, on every block

About two chances in three that a re-sitting today would land between -7 and 25; nineteen in twenty between -22 and 40. Zero is the ordinary respondent.

No lean shown, so no strength is claimed. The whole 95% band includes zero; the two costs are printed for when you do lean.

Tabular fallback for the three bands shown
LeanYou68%95%Placement
Taking it on yourself+4832 to 6417 to 79Leans toward: take it on yourself
Raising it upward-36-52 to -20-67 to -5Leans toward: keep it in the team
Holding the date+9-7 to 25-22 to 40No lean shown

How to read it:a lean is called only when the whole outlined box sits on one side of the ordinary respondent’s tick; a box that crosses it reads “no lean shown”, which is a finding. Every lean called carries its cost in the same box, and the five are never added.

Built for

  • Tech leads about to become engineering managers, who want to know which way they will lean before their team finds out
  • Engineering managers in their first year, whose hardest calls this month were not about code
  • Engineering directors and HR partners building a leadership programme for new managers, who need one specific change per person with its cost named
  • Senior engineers deciding whether the manager track is for them, who want the honest cost of each lean rather than a type

Find out which way you lean when the code is not the problem

44 exercises across five formats · about 40 minutes · five leans as bands, fit apart, one boxed change, and a refusal where a score would have gone. Rs 1,699 in India inclusive of GST, or US$16.99 elsewhere. Sit it again after six weeks of the boxed change.

₹1,699 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and your receipt is sent to that address.

Secure Razorpay payment · ₹1,699 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

Does this test my coding or system design ability?

No. Nothing on the form asks you to read, write or review code, and no item scores technical knowledge. It measures the lean you bring to five calls where the code is not the problem: taking it on yourself or routing it, shielding the team or carrying the organisation's asks in, holding the date or the quality, raising a problem upward or keeping it in the team, and making the technical call yourself or fixing how the team decides. The report opens by saying so, in the place a score would have gone.

How is this different from the First-Time Manager Readiness Assessment, the Code Review Assessment or the Mid-Level Backend Engineer Assessment?

The First-Time Manager Readiness Assessment measures readiness for general first-line management and reports a level. The Code Review Assessment measures judgment inside the code review itself. The Mid-Level Backend Engineer Assessment measures engineering ability. The Decision-Making Skills Assessment for Junior IT Managers is a client-specific diagnostic of decision quality. This one measures the tendency an engineering manager brings to the non-technical calls, as a forced-choice profile with no total and no level, and reports fit apart from it. The report names each difference on the page.

Is there a right answer, and can I just pick the decisive-sounding option every time?

Each forced-choice block does carry a key, the move the published evidence says works in that situation, and that key scores fit, apart from every lean. The leans themselves have no right end. The four moves in every block are matched on how good they sound and written in four registers spread evenly across the leans, so always picking the decisive-sounding move, or the people-first one, moves no lean and fits the evidence no better than the ordinary respondent. Because every block gives one most and one least, nobody can be high on everything, and the report prints the arithmetic.

Why is there no total score, no type and no percentile?

Because none of them would be true. Five leans that are forced against each other cannot add up to one number. A type cut at a midpoint would call a person at plus twelve and a person at plus nine two kinds of manager, so the report calls a lean only when the whole 95 per cent band sits on one side of zero, and otherwise says no lean was shown. And a percentile needs a norm group with a stated size, which does not yet exist; zero is the ordinary respondent on declared marginals that observed data will replace, and the page says so.

How much does it cost, how long does it take, and can I sit it again?

Rs 1,699 in India inclusive of GST, or US$16.99 elsewhere, one time; 45 credits on a business plan. Forty-four exercises across five formats take about forty minutes. The report names a re-sitting after six weeks of the boxed change and prints, for every lean that carries a figure, the number of points it would have to move to count as real change rather than measurement error.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue →

Methodology: Forty-four original items in five formats: twenty quasi-ipsative forced-choice blocks (most_least), each a situation where the code is not the problem and four moves, one per dimension on four of the five dimensions, each move at a declared pole (high or low) of its dimension, matched on social desirability within the block and written in four registers (decisive, people-first, cautious, process) crossed with dimension across the bank so that no register tracks a dimension; eight scenario questions graded zero to three; three select-every-one exercises; three match-the-following exercises; and ten six-point ratings of the lean the respondent says they have, two per dimension with one drawn in each direction so that every stated trait is acquiescence-balanced. DECLARED RESPONSE INSTRUCTION: behavioural tendency throughout. Every block asks which move the respondent is most and least likely to make, every scenario asks what they are most likely to do, and nothing asks what a manager should do. CONSTRUCT STATEMENT: This measures the tendency a tech lead or new engineering manager brings to the calls where the code is not the problem: how much of a problem they take on themselves rather than route, whether they shield the team from the organisation or carry its asks in, whether the date or the quality gives when both cannot be had, how early they raise a problem upward, and what they do with their own technical opinion now that they are the manager. It reports those five leans as bands against the ordinary respondent, with no total, no level and no type. It does not measure technical ability, coding skill or system design knowledge; it does not measure whether you are a good person or a good manager; it does not measure intelligence or integrity; and it says nothing about how you would perform in a job this instrument has not described. FRAMEWORK ANCHOR: competence descriptors are anchored on the European e-Competence Framework (e-CF), EN 16234-1, a published European standard whose competence content is made available by its custodians for reference use: D.9 Personnel Development (taking it on yourself, as the choice between doing and developing), E.4 Relationship Management (shielding the team, as the boundary with the organisation), E.2 Project and Portfolio Management with E.6 ICT Quality Management (holding the date against quality), E.3 Risk Management (raising it upward, as the escalation of risk), and E.9 IS Governance for its descriptor on decision-making structures (making the technical call against building how the team decides). The e-CF is cited factually as a published framework; no endorsement by CEN, by any national standards body or by the e-CF custodians is implied, and nothing here is affiliated with them. SCORING DESIGN, C5, a Thurstonian forced-choice tendency index with the ordinary respondent as zero: for each dimension k and each answered block in which k appears, the press on k's statement scores s = +1 when MOST lands on the high pole or LEAST lands on the low pole, s = -1 when MOST lands on the low pole or LEAST lands on the high pole, and 0 otherwise; q = (s + 1) / 2. Every block declares option_base_rate for MOST and option_base_rate_least for LEAST, both dispersed, and the expected q under them is E[q] = 1/2 + pole * (p_most - p_least) / 2 per pole pressed. The tendency index is T_k = 100 * (mean q - mean E[q]) / (1 - mean E[q]) over the answered blocks containing k, so T_k = 0 is the respondent pressing at the declared marginals, +100 is the high pole pressed on every block, and the negative reach -100 * mean E[q] / (1 - mean E[q]) is printed as the design's own floor. The five raw press sums add to zero by construction, because every block gives one MOST and one LEAST, so a profile high on everything is arithmetically impossible and the page says so. The index is quasi-ipsative, not ipsative: desirability is mixed across the form, both poles of every dimension appear eight times each, and the keyed moves lean no more than a few presses on any dimension (net keyed presses by dimension are printed in scoring.keyed_net), so a respondent who always chooses the keyed move reads near the ordinary respondent on four dimensions and mildly process-leaning on the fifth, which the page states. Faking resistance is designed against APPLIED effect sizes, the applicant-versus-incumbent meta-analytic figures of about a quarter to half a standard deviation for single-stimulus measures and the near-zero to small figures for forced choice in high-stakes settings, not against instructed-faking laboratory effects near one standard deviation; within-block desirability matching is what carries that resistance, and the methodology claims no more than the applied literature supports. A FIT strand is scored apart: each block carries a keyed most_index and least_index, the move the published evidence says works in that situation and the one that most reliably backfires, graded as the platform grades most_least, together with the eight scenarios, three select-every-one and three matching exercises, all chance-corrected against their declared marginals (option_base_rate, option_base_rate_least, select_prior, match_prior), and reported as F = 100 * (mean q - mean E[q]) / (1 - mean E[q]) with 0 the ordinary respondent and 100 every keyed move chosen. Fit is never added to any tendency. The ten ratings are the STATED lean, two per dimension, one drawn each way, reported in words beside the shown lean and never as a figure, because two ratings per dimension carry a direction and not a reliable figure. RELIABILITY is McDonald's omega in its Spearman-Brown form from the answered count and an assumed inter-item correlation of .30, printed as an assumption until live data replaces it; every figure carries its SEM and its 68 and 95 per cent bands from a declared, modelled standard deviation of 30 per tendency and 28 for fit; a tendency under eight answered blocks or omega under .70 unrounded carries a placement word and no number; under four, nothing; the sitting under twenty-six of forty-four answered is not reported. The re-measurement threshold is the reliable change index at 95 per cent. No percentile appears anywhere, no type is cut from a midpoint, and no total exists. REPORT DESIGN, D16, the Anti-Barnum Clause as the organising object: the page opens with what the instrument did not measure, in the place and type size a composite would have taken; then the five tendencies as bands with the ordinary respondent drawn and named, each paired with its honest cost for the pole shown; then the fit strand, the stated-versus-shown comparison, one boxed if-then action with its re-measurement threshold, the reliability table with every refusal printed, and what zero means with the declared marginals. SOURCES: Hill, Becoming a Manager: how new managers master the challenges of leadership (Harvard Business School Press, 2003), for the transition from doing to enabling and the new manager's pull toward hands-on work; Fournier, The Manager's Path (O'Reilly, 2017), for the tech lead's and new engineering manager's handling of their own technical opinion; Edmondson, Psychological safety and learning behavior in work teams (Administrative Science Quarterly, 1999), for the keying of the in-the-room and silent-senior situations; Keil, Smith, Pawlowski and Jin, Why didn't somebody tell me? Climate, information asymmetry, and bad news about troubled projects (Data Base for Advances in Information Systems, 2004), and Smith and Keil, The reluctance to report bad news on troubled software projects (Information Systems Journal, 2003), for the keying of the bad-news situations; Argyris, Teaching smart people how to learn (Harvard Business Review, 1991), for inquiry before advocacy; Brown and Maydeu-Olivares, Item response modeling of forced-choice questionnaires (Educational and Psychological Measurement, 2011), for the Thurstonian model of forced-choice blocks and the recovery of normative information from quasi-ipsative designs; Cao and Drasgow, Does forcing reduce faking? A meta-analytic review of forced-choice personality measures in high-stakes situations (Journal of Applied Psychology, 2019), and Birkeland, Manson, Kisamore, Brannick and Smith, A meta-analytic investigation of job applicant faking on personality measures (International Journal of Selection and Assessment, 2006), for the applied faking effect sizes the design is built against; McDaniel, Hartman, Whetzel and Grubb, Situational judgment tests, response instructions, and validity (Personnel Psychology, 2007), for the behavioural-tendency instruction; Luo, Hariri, Eloussi and Marinov, An empirical analysis of flaky tests (FSE, 2014), for the keying of the flaky-test situation; Fowler, Technical debt quadrant (martinfowler.com, 2009), for deliberate and prudent debt in the workaround and feature-flag situations; CEN, EN 16234-1, European e-Competence Framework, for the competence descriptors; Jacobson and Truax, Clinical significance: a statistical approach to defining meaningful change (Journal of Consulting and Clinical Psychology, 1991), for the reliable change index; Gollwitzer and Sheeran, Implementation intentions and goal achievement (Advances in Experimental Social Psychology, 2006), for the if-then form of the one change; Haladyna, Downing and Rodriguez, A review of multiple-choice item-writing guidelines (Applied Measurement in Education, 2002), for cue control; Baumgartner and Steenkamp, Response styles in marketing research (Journal of Marketing Research, 2001), for the balanced reverse drawing of the ratings. All forty-four items are original works written for this instrument; no item from any published questionnaire, leadership inventory, commercial manager assessment or client test file is reproduced or adapted; no branded instrument is named or used; no competitor is named anywhere; and the instrument is not affiliated with or endorsed by any author, standards body or organisation named above.