Survey and Questionnaire Design Assessment for Research, HR and Customer Insight TeamsNine hundred people answered the question. It had already decided what they would say.
Forty-one exercises on the difference between a survey question that measures something and one that biases, wastes or loses the data. Nine draft questionnaires are printed question by question with their answer formats, and you mark what a reviewer should send back. Behind every question is a named defect class or nothing at all, and every class carries a printed price, so sending everything back costs more than a typical reviewer's calls. Three defective questions are yours to rewrite as neutral ones.
A cost table you can read, printed above the ledger it prices
The Survey and Questionnaire Design Assessment for Research, HR and Customer Insight Teams is a thirty-five-minute check of whether a person can tell a draft survey question that will bias, waste or lose data from one that will do its job, scored against a published cost table and reported as a ledger of their own calls beside every draft question.
Most questionnaire-design quizzes score a count of defects spotted, which means marking everything scores full marks. This one prices every call. A leading question, an assumed direction or a scale with no bad end left in costs three units, because its numbers come back clean and wrong and get quoted. A double-barrelled question, an undefined term or a day nobody can remember left in costs two, because everyone answers it and nobody can read the answers. A box that overlaps its neighbour costs one. And a comment spent on a question that was doing its job costs two, because the writer rewrites a working question and the comparison with the last wave is lost. The table is printed on the report because it is a judgement, not a measurement, and you are meant to be able to disagree with it.
The prices were set against the two extreme strategies through the real scoring, not against intuition. On this bank a respondent who marks each line as an ordinary reviewer would sits at zero. Sending every line back scores minus forty-seven and sending nothing back scores minus forty-two. Both are drawn on the same axis as you, so the question of whether flagging everything pays is answered on the page rather than asserted in the copy.
The report is a ledger. Each draft questionnaire is reproduced as the questionnaire it was, question by question with its answer format set beneath, and beside each question sits your own call with a mark, a word, the class of the line in plain words and what the call cost. The two error directions, defects left in and comments spent on sound questions, are counted apart in lines and in units and never netted into one accuracy figure, because they are fixed by opposite habits. The heavier direction is drawn heavier.
Zero on every figure is a modelled typical respondent, not half marks. Every draft line carries a declared share of ordinary reviewers who send it back, every single-choice, binary, matching and ordering exercise carries a declared prior, and each score is corrected against those. They are authored assumptions stated in the open, and observed shares replace them once live sittings exist. Two of the five parts carry too few exercises to earn a figure on any sitting; those parts print a three-way placement and the refusal in the same place the number would have been.
Three of the exercises are written: a leading question, a double-barrelled question and a question with no permission to say no, each to be rewritten as a neutral one. They are graded against criteria printed on the report and reported in their own panel, never averaged into the headline, and a rewrite the grading pipeline has not yet returned is printed as deferred rather than as a zero. Nothing here is a qualification, and the report says what it did not measure: statistics, sampling, interviewing, and whether a real study would succeed.
What you walk away with
The question that names the direction before it asks, the sentence that tells respondents what everyone else thinks, the assumed struggle, the double negative, the term only staff know, and the two things asked in one number.
The scale with more points on one side, the numbers with no words, the boxes that overlap at every boundary, the boxes that leave someone out, and the Yes/No with no route for those it does not apply to.
The day nobody remembers, the fact the respondent cannot know, the question nobody can say No to, the always that turns a frequency into a confession, and the attitude asked where a behaviour was needed.
Where the overall question sits relative to the specific ones, where sensitive questions go, what a screening question is for, and the order a questionnaire is built and tested in before fieldwork is paid for.
What a skip rate is telling you, what a think-aloud reading of one word shows, why a response rate is not a bias figure, and what a reworded item does to a trend with the last wave.
A cost table above the ledger, your own calls beside every draft question with the class and the price, the two error directions counted apart, and the two extreme habits drawn on the same axis as you.
Inside your report
Illustrative sample — your report is generated from your own responses.
Every draft is reproduced as the questionnaire it was, and beside each question sits your own call with a mark, a word, the class of the line and what the call cost at the printed prices. The mark never carries the call alone, so it reads the same in greyscale.
| Class of line | Sent back | Left in | Why |
|---|---|---|---|
| Biased | 0 | 3 | Comes back clean-looking and wrong, and gets acted on. |
| Cannot be read | 0 | 2 | Answered by everyone, dropped after fieldwork is paid for. |
| Cannot be coded | 0 | 1 | Loses answers on the boundaries; partly recoverable. |
| Sound | 2 | 0 | A working question rewritten; comparability lost. |
The prices are a value judgement, so they are printed, not buried. Sending every line back and sending none back are both drawn on the same axis as you, well below the prior-drawing respondent at zero, across 5 parts.
Built for
- Insight and market-research analysts who review questionnaires before they go to field and want to know which of their comments were worth the writer's time
- HR and employee-listening teams building engagement, pulse and exit surveys whose results will be acted on
- Product and customer-experience researchers writing in-app and post-purchase surveys under time pressure
- Programme evaluators, NGO teams and academics running surveys where a biased item costs a finding rather than a metric
Find out which of your comments were worth a rewrite, and which questions you would have let through
41 exercises across six formats · about 35 minutes · a cost-priced score on nine drafts, five parts with figures where the reliability allows and placements where it does not, three rewrites reported apart, and every draft reproduced with your calls in the margin.
₹699 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
No, and the report shows you why on its own axis. Every comment spent on a sound question costs two units at the printed prices, and more than half the draft questions are sound as written. On this bank, sending every line back scores minus forty-seven and sending nothing back scores minus forty-two, against a typical respondent at zero. Both habits are drawn on the same scale as your own score.
They are a value judgement made by the instrument's authors and printed so that you can read and disagree with them. A biased question left in is priced highest because its numbers come back clean and wrong and get acted on; an unreadable one next, because it is dropped after the fieldwork is paid for; an uncodable one lowest, because its loss is partly recoverable. The price of a comment on a sound question was set by checking what the two extreme strategies score through the real scoring, not by intuition.
A part needs at least eight answered exercises and enough assumed internal agreement, measured as omega and compared unrounded, before a figure is honest to print. Two parts carry seven exercises by design and so print a three-way placement, below, near or above a typical respondent, with the reason in the same place the figure would have been.
By the platform's rubric pipeline against criteria printed beside each rewrite on the report: no direction named in the question, one thing asked, balanced answer options, plain words. They are reported in their own panel and never averaged into the headline or any part. If the pipeline has not returned a grade, the rewrite is printed as deferred and leaves both sides of its ratio; it is never scored as a zero.
No. The drafts are original questionnaires for a canteen, a course, a commute, a library, an installation, a school, a workplace policy, a clinic and a park, written so that no domain expertise is needed to read them. It measures whether you can see what a question will do to its data, not familiarity with any software or sector, and it does not measure sampling, statistics or interviewing.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Forty-one original exercises across six formats. The spine is a marking task: nine draft questionnaires, each six questions with their answer format, written for a stated decision, and the respondent marks every question a reviewer should send back. Behind each question is a named defect class (leading or loaded wording, a presupposition, an unbalanced scale, a socially desirable direction with no permission to say no, a double-barrelled question, a double negative, an undefined term, a recall window nobody can answer, a question the respondent cannot know the answer to, an absolute quantifier, overlapping options, non-exhaustive options, an unlabelled scale, a missing not-applicable route) or nothing at all; 26 of the fifty-four lines carry a defect and 28 are sound as written, and the number of defects per draft varies from two to four so that counting to a fixed number is not a strategy. The other thirty-two exercises are twelve single-choice judgements about what a wording, a scale, a window or an order does to the data, eight binary claims, five match-the-following exercises pairing a draft question with the name of its defect (one of which is always 'nothing wrong'), four ordering exercises on question and development order, and three short written rewrites of a defective question into a neutral one. CONSTRUCT STATEMENT: this measures whether a person can tell a survey question that will bias, waste or lose data from one that will do its job, and what they know about how wording, scales, recall, order and non-response act on survey data; it does not measure statistical analysis, sampling design, interviewing skill, writing fluency, subject expertise in any field the drafts touch, whether a real study would succeed, or anything about the person's character. RESPONSE INSTRUCTION: knowledge, declared once for the whole instrument: what is worth sending back, never what the respondent would do. SCORING DESIGN: C7, a decision-cost matrix. Every defect class belongs to a cost tier and the tier carries the cost of leaving that line in: 3 for a biased line, because it comes back looking clean and is acted on; 2 for a line that cannot be read, because the item is dropped after the fieldwork is paid for; 1 for a line that cannot be coded, because the loss is at the boundaries and partly recoverable. A comment spent on a sound line costs 2: a working question is rewritten, comparability with the last wave is lost, and rewrites of sound questions gain defects. The matrix is a value judgement, not a measurement, and it is printed on the report above the ledger. The score is the summed cost of the reader's calls, normalised so that zero is a modelled respondent drawing each call from the declared share of ordinary reviewers who send that line back, and one hundred is every call right. That null is authored, not observed: observed marking shares replace the declared priors once live sittings exist, and no percentile is claimed anywhere. The false-alarm price was calibrated against the two extreme strategies through the real arithmetic: on this bank, marking every line scores -47 and marking nothing scores -42, both well below the prior-drawing respondent at zero. The two error directions, defects left in and comments spent on sound lines, are counted and printed apart and never netted. The knowledge exercises are corrected against their declared per-option, per-pair and per-position priors. REFUSAL RULES: an unanswered exercise leaves the numerator, the denominator and the expected-cost term together; an empty sitting scores exactly zero; a sitting with fewer than twenty-four exercises answered, or fewer than seven drafts marked, is refused a headline and the refusal is printed where the number would have been; a part with fewer than eight exercises answered, or an assumed omega under .70 taken unrounded, carries a three-way placement and no number; omega is estimated from item count and an assumed inter-item correlation of .25, stated as an assumption that observed data will replace, and printed unrounded; every printed figure carries its 68 and 95 per cent band from an assumed score standard deviation. The three written rewrites are graded by the platform's rubric pipeline against criteria printed on the report, and they are reported apart from every keyed figure and never averaged into the headline; a rewrite for which no grade has come back is printed as deferred and leaves both sides of its ratio rather than scoring zero. A careless-responding flag count is computed for the operator and never shown to the respondent as a judgement. Sources drawn on: Sudman and Bradburn, Asking Questions (1982), and Bradburn, Sudman and Wansink, Asking Questions (2004), on wording, recall and threatening questions; Tourangeau, Rips and Rasinski, The Psychology of Survey Response (2000), on comprehension, retrieval, judgement and reporting; Schuman and Presser, Questions and Answers in Attitude Surveys (1981), on question order, balance and don't-know options; Krosnick, Response Strategies for Coping with the Cognitive Demands of Attitude Measures in Surveys (1991), on satisficing and acquiescence; Krosnick and Presser, Question and Questionnaire Design, in the Handbook of Survey Research (2010), on scale labelling and midpoints; Fowler, Improving Survey Questions (1995), on double-barrelled, loaded and undefined terms; Groves, Fowler, Couper, Lepkowski, Singer and Tourangeau, Survey Methodology (2009), on non-response bias and the response-rate fallacy; Groves and Peytcheva, The Impact of Nonresponse Rates on Nonresponse Bias (2008); Willis, Cognitive Interviewing (2005), on think-aloud and probing in pretesting; Presser, Rothgeb, Couper, Lessler, Martin, Martin and Singer, Methods for Testing and Evaluating Survey Questionnaires (2004); Dillman, Smyth and Christian, Internet, Phone, Mail and Mixed-Mode Surveys (2014), on visual design and ordering; Converse and Presser, Survey Questions: Handcrafting the Standardized Questionnaire (1986); Belson, The Design and Understanding of Survey Questions (1981), on misreading; Payne, The Art of Asking Questions (1951); and Haladyna, Downing and Rodriguez (2002) for the item-writing rules applied to the exercises themselves. Every exercise is an original work written for this instrument; no real employer, school, council, clinic, course or place is named anywhere, and no item, scale or section name is taken from any commercial questionnaire or instrument. The defect classes are the public vocabulary of survey methodology and are named as such; this instrument is not affiliated with, endorsed by or derived from any professional body, training provider or commercial instrument.