Applied Judgment Assessment
Applied skill · insight, research, HR listening, CX and evaluation teams · browse the full catalogue

Survey and Questionnaire Design Assessment for Research, HR and Customer Insight TeamsNine hundred people answered the question. It had already decided what they would say.

Forty-one exercises on the difference between a survey question that measures something and one that biases, wastes or loses the data. Nine draft questionnaires are printed question by question with their answer formats, and you mark what a reviewer should send back. Behind every question is a named defect class or nothing at all, and every class carries a printed price, so sending everything back costs more than a typical reviewer's calls. Three defective questions are yours to rewrite as neutral ones.

35 minutes41 scored exercisesEvidence-keyed scoringGlobal · INR & USD

A cost table you can read, printed above the ledger it prices

The Survey and Questionnaire Design Assessment for Research, HR and Customer Insight Teams is a thirty-five-minute check of whether a person can tell a draft survey question that will bias, waste or lose data from one that will do its job, scored against a published cost table and reported as a ledger of their own calls beside every draft question.

Most questionnaire-design quizzes score a count of defects spotted, which means marking everything scores full marks. This one prices every call. A leading question, an assumed direction or a scale with no bad end left in costs three units, because its numbers come back clean and wrong and get quoted. A double-barrelled question, an undefined term or a day nobody can remember left in costs two, because everyone answers it and nobody can read the answers. A box that overlaps its neighbour costs one. And a comment spent on a question that was doing its job costs two, because the writer rewrites a working question and the comparison with the last wave is lost. The table is printed on the report because it is a judgement, not a measurement, and you are meant to be able to disagree with it.

The prices were set against the two extreme strategies through the real scoring, not against intuition. On this bank a respondent who marks each line as an ordinary reviewer would sits at zero. Sending every line back scores minus forty-seven and sending nothing back scores minus forty-two. Both are drawn on the same axis as you, so the question of whether flagging everything pays is answered on the page rather than asserted in the copy.

The report is a ledger. Each draft questionnaire is reproduced as the questionnaire it was, question by question with its answer format set beneath, and beside each question sits your own call with a mark, a word, the class of the line in plain words and what the call cost. The two error directions, defects left in and comments spent on sound questions, are counted apart in lines and in units and never netted into one accuracy figure, because they are fixed by opposite habits. The heavier direction is drawn heavier.

Zero on every figure is a modelled typical respondent, not half marks. Every draft line carries a declared share of ordinary reviewers who send it back, every single-choice, binary, matching and ordering exercise carries a declared prior, and each score is corrected against those. They are authored assumptions stated in the open, and observed shares replace them once live sittings exist. Two of the five parts carry too few exercises to earn a figure on any sitting; those parts print a three-way placement and the refusal in the same place the number would have been.

Three of the exercises are written: a leading question, a double-barrelled question and a question with no permission to say no, each to be rewritten as a neutral one. They are graded against criteria printed on the report and reported in their own panel, never averaged into the headline, and a rewrite the grading pipeline has not yet returned is printed as deferred rather than as a zero. Nothing here is a qualification, and the report says what it did not measure: statistics, sampling, interviewing, and whether a real study would succeed.

Five parts of questionnaire design, three of them carrying drafts to mark and all five carrying knowledge exercises, plus the written strand reported apart:
Wording that leads, loads or asks two thingsResponse scales and answer optionsQuestions a respondent can actually answerQuestion order and context effectsPiloting, non-response and dropped items

What you walk away with

Wording that leads, loads or asks two things

The question that names the direction before it asks, the sentence that tells respondents what everyone else thinks, the assumed struggle, the double negative, the term only staff know, and the two things asked in one number.

Response scales and answer options

The scale with more points on one side, the numbers with no words, the boxes that overlap at every boundary, the boxes that leave someone out, and the Yes/No with no route for those it does not apply to.

Questions a respondent can actually answer

The day nobody remembers, the fact the respondent cannot know, the question nobody can say No to, the always that turns a frequency into a confession, and the attitude asked where a behaviour was needed.

Question order and context effects

Where the overall question sits relative to the specific ones, where sensitive questions go, what a screening question is for, and the order a questionnaire is built and tested in before fieldwork is paid for.

Piloting, non-response and dropped items

What a skip rate is telling you, what a think-aloud reading of one word shows, why a response rate is not a bias figure, and what a reworded item does to a trend with the last wave.

The cost of each call, printed

A cost table above the ledger, your own calls beside every draft question with the class and the price, the two error directions counted apart, and the two extreme habits drawn on the same axis as you.

Inside your report

Illustrative sample — your report is generated from your own responses.

The ledger: the draft, question by question, your call in the margin
Draft: an online course feedback form · Response scales and answer options
Q1How clear were the video lessons?
Answer: Very unclear to Very clear
Left alone, rightlySound
cost 0 units · 8 of 100 reviewers send it back
Q2How useful was the course?
Answer: Useful / Very useful / Extremely useful
Left inBiased: unbalanced scale
cost 3 units · 44 of 100 reviewers send it back
Q3How many hours a week did you study?
Answer: 0-2 / 2-5 / 5-10 / 10+
Sent back, rightlyCannot be coded: boxes overlap
cost 0 units · 36 of 100 reviewers send it back
Q4Would you recommend this course to a colleague?
Answer: Yes / No / Not sure
Comment spentSound
cost 2 units · 18 of 100 reviewers send it back

Every draft is reproduced as the questionnaire it was, and beside each question sits your own call with a mark, a word, the class of the line and what the call cost at the printed prices. The mark never carries the call alone, so it reads the same in greyscale.

The cost table, printed above the ledger, and the score it produces
Class of lineSent backLeft inWhy
Biased03Comes back clean-looking and wrong, and gets acted on.
Cannot be read02Answered by everyone, dropped after fieldwork is paid for.
Cannot be coded01Loses answers on the boundaries; partly recoverable.
Sound20A working question rewritten; comparability lost.
costlier than the prior-drawing respondentevery call right▲ send everything back -47▲ send nothing back -42◇ prior-drawing respondent at 0● you 38shaded = 68% band 23 to 53 · outlined = 95% band 9 to 67

The prices are a value judgement, so they are printed, not buried. Sending every line back and sending none back are both drawn on the same axis as you, well below the prior-drawing respondent at zero, across 5 parts.

Built for

  • Insight and market-research analysts who review questionnaires before they go to field and want to know which of their comments were worth the writer's time
  • HR and employee-listening teams building engagement, pulse and exit surveys whose results will be acted on
  • Product and customer-experience researchers writing in-app and post-purchase surveys under time pressure
  • Programme evaluators, NGO teams and academics running surveys where a biased item costs a finding rather than a metric

Find out which of your comments were worth a rewrite, and which questions you would have let through

41 exercises across six formats · about 35 minutes · a cost-priced score on nine drafts, five parts with figures where the reliability allows and placements where it does not, three rewrites reported apart, and every draft reproduced with your calls in the margin.

₹699 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and Razorpay sends your receipt to that address.

Secure Razorpay payment · ₹699 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

Does marking every draft question as defective score well?

No, and the report shows you why on its own axis. Every comment spent on a sound question costs two units at the printed prices, and more than half the draft questions are sound as written. On this bank, sending every line back scores minus forty-seven and sending nothing back scores minus forty-two, against a typical respondent at zero. Both habits are drawn on the same scale as your own score.

Where do the prices in the cost table come from?

They are a value judgement made by the instrument's authors and printed so that you can read and disagree with them. A biased question left in is priced highest because its numbers come back clean and wrong and get acted on; an unreadable one next, because it is dropped after the fieldwork is paid for; an uncodable one lowest, because its loss is partly recoverable. The price of a comment on a sound question was set by checking what the two extreme strategies score through the real scoring, not by intuition.

Why do two of the five parts carry a placement instead of a score?

A part needs at least eight answered exercises and enough assumed internal agreement, measured as omega and compared unrounded, before a figure is honest to print. Two parts carry seven exercises by design and so print a three-way placement, below, near or above a typical respondent, with the reason in the same place the figure would have been.

How are the three written rewrites graded?

By the platform's rubric pipeline against criteria printed beside each rewrite on the report: no direction named in the question, one thing asked, balanced answer options, plain words. They are reported in their own panel and never averaged into the headline or any part. If the pipeline has not returned a grade, the rewrite is printed as deferred and leaves both sides of its ratio; it is never scored as a zero.

Does it test knowledge of a specific survey tool or a particular field?

No. The drafts are original questionnaires for a canteen, a course, a commute, a library, an installation, a school, a workplace policy, a clinic and a park, written so that no domain expertise is needed to read them. It measures whether you can see what a question will do to its data, not familiarity with any software or sector, and it does not measure sampling, statistics or interviewing.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue

Methodology: Forty-one original exercises across six formats. The spine is a marking task: nine draft questionnaires, each six questions with their answer format, written for a stated decision, and the respondent marks every question a reviewer should send back. Behind each question is a named defect class (leading or loaded wording, a presupposition, an unbalanced scale, a socially desirable direction with no permission to say no, a double-barrelled question, a double negative, an undefined term, a recall window nobody can answer, a question the respondent cannot know the answer to, an absolute quantifier, overlapping options, non-exhaustive options, an unlabelled scale, a missing not-applicable route) or nothing at all; 26 of the fifty-four lines carry a defect and 28 are sound as written, and the number of defects per draft varies from two to four so that counting to a fixed number is not a strategy. The other thirty-two exercises are twelve single-choice judgements about what a wording, a scale, a window or an order does to the data, eight binary claims, five match-the-following exercises pairing a draft question with the name of its defect (one of which is always 'nothing wrong'), four ordering exercises on question and development order, and three short written rewrites of a defective question into a neutral one. CONSTRUCT STATEMENT: this measures whether a person can tell a survey question that will bias, waste or lose data from one that will do its job, and what they know about how wording, scales, recall, order and non-response act on survey data; it does not measure statistical analysis, sampling design, interviewing skill, writing fluency, subject expertise in any field the drafts touch, whether a real study would succeed, or anything about the person's character. RESPONSE INSTRUCTION: knowledge, declared once for the whole instrument: what is worth sending back, never what the respondent would do. SCORING DESIGN: C7, a decision-cost matrix. Every defect class belongs to a cost tier and the tier carries the cost of leaving that line in: 3 for a biased line, because it comes back looking clean and is acted on; 2 for a line that cannot be read, because the item is dropped after the fieldwork is paid for; 1 for a line that cannot be coded, because the loss is at the boundaries and partly recoverable. A comment spent on a sound line costs 2: a working question is rewritten, comparability with the last wave is lost, and rewrites of sound questions gain defects. The matrix is a value judgement, not a measurement, and it is printed on the report above the ledger. The score is the summed cost of the reader's calls, normalised so that zero is a modelled respondent drawing each call from the declared share of ordinary reviewers who send that line back, and one hundred is every call right. That null is authored, not observed: observed marking shares replace the declared priors once live sittings exist, and no percentile is claimed anywhere. The false-alarm price was calibrated against the two extreme strategies through the real arithmetic: on this bank, marking every line scores -47 and marking nothing scores -42, both well below the prior-drawing respondent at zero. The two error directions, defects left in and comments spent on sound lines, are counted and printed apart and never netted. The knowledge exercises are corrected against their declared per-option, per-pair and per-position priors. REFUSAL RULES: an unanswered exercise leaves the numerator, the denominator and the expected-cost term together; an empty sitting scores exactly zero; a sitting with fewer than twenty-four exercises answered, or fewer than seven drafts marked, is refused a headline and the refusal is printed where the number would have been; a part with fewer than eight exercises answered, or an assumed omega under .70 taken unrounded, carries a three-way placement and no number; omega is estimated from item count and an assumed inter-item correlation of .25, stated as an assumption that observed data will replace, and printed unrounded; every printed figure carries its 68 and 95 per cent band from an assumed score standard deviation. The three written rewrites are graded by the platform's rubric pipeline against criteria printed on the report, and they are reported apart from every keyed figure and never averaged into the headline; a rewrite for which no grade has come back is printed as deferred and leaves both sides of its ratio rather than scoring zero. A careless-responding flag count is computed for the operator and never shown to the respondent as a judgement. Sources drawn on: Sudman and Bradburn, Asking Questions (1982), and Bradburn, Sudman and Wansink, Asking Questions (2004), on wording, recall and threatening questions; Tourangeau, Rips and Rasinski, The Psychology of Survey Response (2000), on comprehension, retrieval, judgement and reporting; Schuman and Presser, Questions and Answers in Attitude Surveys (1981), on question order, balance and don't-know options; Krosnick, Response Strategies for Coping with the Cognitive Demands of Attitude Measures in Surveys (1991), on satisficing and acquiescence; Krosnick and Presser, Question and Questionnaire Design, in the Handbook of Survey Research (2010), on scale labelling and midpoints; Fowler, Improving Survey Questions (1995), on double-barrelled, loaded and undefined terms; Groves, Fowler, Couper, Lepkowski, Singer and Tourangeau, Survey Methodology (2009), on non-response bias and the response-rate fallacy; Groves and Peytcheva, The Impact of Nonresponse Rates on Nonresponse Bias (2008); Willis, Cognitive Interviewing (2005), on think-aloud and probing in pretesting; Presser, Rothgeb, Couper, Lessler, Martin, Martin and Singer, Methods for Testing and Evaluating Survey Questionnaires (2004); Dillman, Smyth and Christian, Internet, Phone, Mail and Mixed-Mode Surveys (2014), on visual design and ordering; Converse and Presser, Survey Questions: Handcrafting the Standardized Questionnaire (1986); Belson, The Design and Understanding of Survey Questions (1981), on misreading; Payne, The Art of Asking Questions (1951); and Haladyna, Downing and Rodriguez (2002) for the item-writing rules applied to the exercises themselves. Every exercise is an original work written for this instrument; no real employer, school, council, clinic, course or place is named anywhere, and no item, scale or section name is taken from any commercial questionnaire or instrument. The defect classes are the public vocabulary of survey methodology and are named as such; this instrument is not affiliated with, endorsed by or derived from any professional body, training provider or commercial instrument.