Applied Judgment Assessment
Applied skill · bid, proposal, pre-sales, consulting and grant-writing teams · browse the full catalogue

Bid and Proposal Quality Assessment for Sales, Consulting and Grant TeamsThe bid that lost was not the cheapest. It was the one that never answered the question.

Twelve extracts from bid and proposal responses, six lines at a time, and you mark the lines a reviewer should comment on. Half the lines are fine by design, so flagging everything scores exactly what flagging nothing scores. Two figures, never one: how well you tell the line that costs from the line that earns, and where you set the bar.

30 minutes36 scored exercisesEvidence-keyed scoringGlobal · INR & USD

A ledger with your calls in the margin, not a score for reading

The Bid and Proposal Quality Assessment for Sales, Consulting and Grant Teams is a thirty-minute knowledge assessment of whether a person can tell the line in a bid response that would cost marks from the line that is doing its job, reported as a line-by-line margin ledger with d-prime and criterion drawn beneath it.

Most proposal training teaches a method and most bid reviews are opinion: one reviewer covers a page in comments, another signs it off, and nobody can say which of them read it better. This instrument puts the extract on the page and asks for the reviewer's calls, line by line, against a stated buyer question. Every line carries a hidden class, a genuine defect that an evaluator would score down or a line that is earning its place, and thirty-seven of the seventy-two lines are fine by design. That is what makes the marking a measurement: a reviewer who comments on everything and a reviewer who comments on nothing both land at zero by arithmetic, and are told apart only by where they set their bar.

The scoring is signal detection, the same arithmetic used to measure radiologists and inspectors, with the log-linear correction so that a perfect sitting stays finite. Two figures are printed and never combined. D-prime is how far apart your two rates sit: the share of defects you caught against the share of fine lines you flagged. Criterion is where your bar sits: a low bar comments freely and pays in flagged lines that were fine; a high bar holds back and pays in defects that go out with the bid. Both are drawn on one axis each, with the 68 and 95 per cent bands drawn and written in plain words, and with a modelled typical reviewer marked on the same line so the reader can see where an ordinary reviewer lands.

The report is the ledger. Every extract is reproduced line by line, in the order it was written, and beside each line sits your own call: a mark, a word, and one sentence saying what that call would have cost the bid. Caught, missed, flagged but fine, left rightly. The extract that cost the most is printed first, the two error directions are counted apart with a worked example from your own ledger under each, and the report ends with one if-then sentence drawn from the direction you err in, and a re-measurement threshold that says how much change would be real.

The other twenty-four exercises ask what an evaluator actually scores: what happens to evidence offered on request, how a cross-referenced answer is marked, what a price with stated assumptions tells a buyer, what a missing signed form does to a bid, which fix belongs to which weakness, and the order a response is built in. The keying draws on published work on argument structure, on persuasion and social proof, on plain-language readability, and on procurement evaluation practice and the compliance-matrix tradition; the methodology note names every source. No competency carries a number, because none has enough exercises for one; each carries a placement in words with the reason printed beside it.

No real buyer, company, tender portal or country's procurement rules are named anywhere, so the same sitting is fair wherever it is taken. Every extract is an original, generic response to a generic question, and the instrument is not affiliated with any commercial instrument or certification body. It measures marking, not writing: a person who can see what costs a bid and a person who can write one are often the same person, but this measures the first.

Five parts of bid quality, each a placement in words, feeding two figures that are never combined:
Answering the question that was askedEvidence and proof behind claimsCommercials and pricingCompliance and completenessClarity and structure

What you walk away with

D-prime: telling the line that costs from the line that earns

How far apart your two rates sit, drawn on one axis with its 68 and 95 per cent bands, the modelled typical reviewer, and both fixed habits marked on the same line.

Criterion: where you set the bar

Whether you comment freely or hold back, as its own figure beside d-prime, because a flagger and a filterer at the same d-prime are two different reviewers.

The ledger

Every extract line by line, your call in the margin beside each with a mark, a word and a sentence on what the call would have cost. Costliest extract first.

The two error directions, counted apart

Defects left in and fine lines flagged, each with its own count, its own cost, a worked example from your own ledger and its own fix. Never added together.

Five parts, placed in words

Answering the question, evidence and proof, commercials and pricing, compliance and completeness, clarity and structure. Each a three-way placement against a modelled typical respondent, with the refusal of a number printed.

One change and a re-measurement

One if-then sentence drawn from the direction you err in, and the size of change on d-prime that would count as real rather than noise when you sit it again.

Inside your report

Illustrative sample — your report is generated from your own responses.

The ledger: the extract, with your call in the margin
Pricing section · the buyer asked: “Provide a fixed price that includes all travel and expenses”
1The fixed price is 84,000 in total.
Left, rightly8 of 100 flag it

A single fixed figure is what was asked for.

2Travel is billed at cost in addition.
Caught55 of 100 flag it

The buyer asked for travel inside the fixed price.

3This price is valid for ninety days.
Flagged, but fine20 of 100 flag it

A validity period does not qualify the price.

4Our day rates are highly competitive.
Missed40 of 100 flag it

Unsupported, and a day-rate framing undercuts the fixed price.

● 1 caught○ 1 missed▲ 1 flagged, fine▬ 1 left

Every extract is reproduced line by line, and your own call sits beside each with a mark, a word and one sentence on what the call would have cost the bid. The costliest extract is printed first.

D-prime on one axis, bands drawn and written
fine lines flagged more than defectsdefects told apartflag all / flag none 0.00◇ typical (modelled) 1.05● you 1.62shaded = 68% band 1.18 to 2.06 · outlined = 95% band 0.76 to 2.48
● You1.62
◇ Typical respondent (modelled)1.05
▲ Flags everything / flags nothing0.00

Criterion is drawn on a second axis of its own, because a reviewer who flags freely and one who holds back can share a d-prime and need opposite fixes. Both habits sit at zero by arithmetic; the typical respondent is a model drawn from the declared priors, and observed data will replace it.

Built for

  • Bid, proposal and pre-sales teams who review responses before they go out and want to know who reads them well
  • Consultants, agencies and professional-services firms that answer tenders and requests for proposals
  • Grant writers and fundraising teams whose applications are scored against published criteria
  • Sales managers and bid directors deciding who should hold the red pen on the next submission

Find out which lines you would have caught, and which you would have let go out

36 exercises across five formats · about 30 minutes · d-prime and criterion never combined, every extract reproduced with your calls in the margin, and the priors printed beside every line.

₹999 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and Razorpay sends your receipt to that address.

Secure Razorpay payment · ₹999 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

Is this a test of bid writing?

No. It is a test of bid reviewing: whether you can tell the line that would cost marks from the line that is earning its place, on twelve extracts against stated buyer questions. Writing and reviewing overlap, but this measures the reading. The other twenty-four exercises ask what an evaluator actually scores, how price and compliance are checked, and the order a response is built in.

Why are half the lines fine by design?

Because without them, commenting on everything would score full marks, and the instrument would be measuring suspicion rather than judgement. Thirty-seven of the seventy-two lines are earning their place. A reviewer who flags everything and a reviewer who flags nothing both land at d-prime zero by arithmetic, and the report prints both of those habits on the same axis as your own figure.

What are d-prime and criterion?

D-prime is how far apart your two rates sit: the share of defects you caught against the share of fine lines you flagged. Criterion is where your bar sits: a low bar comments freely and pays in flagged lines that were fine; a high bar holds back and pays in defects that go out. They are printed separately because a flagger and a filterer at the same d-prime are different reviewers with different fixes.

Does it name real buyers, tenders or procurement law?

No. Every extract is an original, generic response to a generic buyer question, and no country's procurement rules, tender portal, company or government department is named anywhere. The keying rests on published work on argument structure, persuasion, plain-language readability and evaluation practice, and the methodology note names every source. The instrument is not affiliated with any certification body.

How long is it and what does it cost?

About thirty minutes for thirty-six exercises across five formats: marking extracts, single-choice judgements, true-or-false claims, matching a weakness to its fix, and ordering the steps a response is built in. ₹999 in India, inclusive of GST, or US$9.99 elsewhere. For teams, 26 credits per sitting. The report is the product: a margin ledger of your own calls with two figures beneath it.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue

Methodology: Thirty-six original exercises across five formats. The spine is a marking task: twelve extracts, each six sentences from a bid or proposal response to a stated buyer question, and the respondent marks every line a reviewer should comment on. Behind each line is a hidden class, a genuine defect that would cost the bid or not a defect at all, and thirty-five of the seventy-two lines are defects against thirty-seven that are not, so marking everything and marking nothing are both worthless by arithmetic rather than by a threshold. The other twenty-four exercises are eight single-choice judgements about what a buyer's evaluator is actually scoring, six binary claims about how competitive bids are evaluated, five match-the-following exercises pairing a weakness with the change that fixes it, and five ordering exercises on the order a response is built in. CONSTRUCT STATEMENT: this measures whether a person can tell the line in a bid response that would cost marks from the line that is doing its job, and what they know about how evaluators score, price and check compliance; it does not measure writing ability, sales skill, persuasiveness in a meeting, domain expertise in any sector, whether a real bid would be won, or anything about the person's character. RESPONSE INSTRUCTION: knowledge, declared once for the whole instrument: what is worth a comment here, never what the respondent would do. SCORING DESIGN: C6 signal detection. Marking a defect is a hit, marking a fine line is a false alarm; d-prime is the distance between the two rates in standard units and criterion c is where the respondent sets the bar, and both are reported because a person who flags everything and a person who flags nothing both land at d-prime zero and are told apart only by c. The log-linear correction of one half per cell keeps both finite. The hit condition is the classic flag-or-leave call rather than a naming condition, because a respondent drawing from the declared priors on this bank scores a d-prime of about plus one half, above the scale's zero, so no fixed habit beats a real reviewer by doing nothing. Every marking line carries a declared prior, the share of ordinary reviewers expected to flag it, and every choice exercise carries a declared prior per option; the matching and ordering exercises carry a declared per-pair and per-position prior. A modelled typical respondent is computed from those priors and marked on the same axis as the reader, and it is a model rather than a norm group: observed answer shares replace it once live sittings exist, and no percentile is claimed anywhere. The two error directions, defects left in and fine lines flagged, are counted apart and never netted. REFUSAL RULES: an unanswered extract leaves the numerator, the denominator and the expected term together; a sitting with fewer than six extracts marked is refused a d-prime and the refusal is printed where the number would have been; every competency has fewer than eight exercises or an assumed omega under .70 and so carries a three-way placement, no number and no percentile, with the refusal printed; the assumed omega is estimated from item count and an assumed inter-item correlation of .22 and printed unrounded. A careless-responding flag count is computed for the operator and never shown to the respondent as a judgement. Sources drawn on: Green and Swets, Signal Detection Theory and Psychophysics, and Macmillan and Creelman, Detection Theory: A User's Guide, for d-prime, criterion and the log-linear correction; Toulmin, The Uses of Argument, for the claim-data-warrant structure behind the evidence items; Cialdini, Influence, on social proof and authority as persuasion levers and their misuse as unsupported claims; Oppenheimer, Consequences of Erudite Vernacular Utilized Irrespective of Necessity (2006), on needless complexity and judged competence; Kimble, Writing for Dollars, Writing to Please, and the plain-language readability findings collected there; Flesch and Kincaid on readability measurement; Newman, Proposal Guide for Business Development Professionals, and Sant, Persuasive Business Proposals, on proposal evaluation practice and the compliance matrix tradition; Lewis, Bids, Tenders and Proposals: Winning Business Through Best Practice, on evaluator behaviour; Thai, International Handbook of Public Procurement, on scored criteria, pass-or-fail gates and clarification practice; Kahneman, Thinking, Fast and Slow, on substitution, answering an easier question than the one asked; and Haladyna, Downing and Rodriguez (2002) for the item-writing rules. Every exercise is an original work written for this instrument, and no real company, buyer, government department, tender portal or country's procurement law is named anywhere. APMP is a trademark of the Association of Proposal Management Professionals; this instrument is not affiliated with, endorsed by or derived from that body, its certifications, or any other commercial instrument or certification body.