Flagger or Filterer: What You Notice, and What You Let ThroughMost tests give you one accuracy score. It cannot tell a person who misses things from a person who flags everything.
Those are opposite problems, and one number hides both. This one reads thirty short everyday records with you and reports two separate facts: how well you tell a real defect from a record that only looks wrong, and where you set the line at which you speak up.
Two facts about you, reported side by side, never added together
Noticing is not one skill, it is two, and almost every test of it collapses them into a percentage. Imagine two people who read the same thirty records and score the same. The first raises a query on every record they are not certain about; almost nothing gets past them, and their colleagues spend their afternoons proving that things were fine. The second lets the doubtful ones go; their interruptions are worth stopping for, and the borderline record — which is where real defects actually live — goes through. Same score. Opposite problems. A single accuracy percentage destroys the distinction entirely, and this instrument exists because of that.
So the sitting is built as a signal-and-noise stream. Thirty short records, each written out in full and read once: a repair bill, a train ticket over midnight, a message from payroll, a stock sheet, a voucher that looks expired, a driver's rota, a handover sheet. After each one you say what, if anything, is wrong. Thirteen of the thirty really do have a defect, and you have to name which thing it is — three specific defect claims are offered and only one is the real one, so sensing that something is off and pointing at the wrong part earns nothing. Seventeen of the thirty are completely sound, and on those the defect claims are the tempting ones: a number that looks odd and reconciles, a date that looks out of order and is not, a request that sounds irregular and is routine. The sound records deliberately outnumber the defective ones, so flagging everything scores below a reader who cannot tell them apart at all.
Scoring is signal-detection theory, the same arithmetic used to separate what a detector can see from how eager it is to say so. Your separation figure is the distance between your hit rate and your false-alarm rate. Your threshold figure is where the two of them sit together. They are printed beside each other and are never averaged, because they are different facts. The report carries the standard error of the separation figure taken from the spread of your own two rates, both the two-in-three and the nineteen-in-twenty span written out in words, the five counts the figure is made of, an honest statement of why nothing was corrected for guessing, and the reliability arithmetic that is the reason no section of it carries a percentage. Consumer one-time personality and ability reports sell at roughly $19–69 and none of them publish a standard error, a reliability estimate, or a list of what they refuse to claim. This one publishes all three.
It is also careful about what it is not. Sustained attention is a disability-relevant construct, so this instrument is built not to be a test of it: the form is short, no record is timed on its own, nothing in the scoring uses how fast you answered, retakes are allowed, and the report says in plain words that this is not a measure of attention, concentration, stamina or anything that rises and falls with health, medication, tiredness or the room you are sitting in. It says, on the page, that it must never be the only basis for a decision about anybody.
What you walk away with
Flagger, filterer, or between them — built from the shape of your answers rather than the level, and withheld outright when your answers cannot support it.
How far apart a defective record and a sound one are for you, drawn on a signed axis where zero means chance by construction, with both spans in plain words.
Caught and named, sensed but misnamed, passed over, false alarms, and left alone rightly. A composite without its parts is unusable, so the parts sit inside the card.
What a flagger gets and what a flagger costs; what a filterer gets and what a filterer costs. Your line marked between them, on its own axis.
Of the defects you claimed on records that really had one, the share you named correctly. Knowing something is wrong and naming the wrong thing is a separate, fixable failure.
Money, dates, messages and forms, each as a three-way classification with no number and no percentile — and the reliability table that explains why.
Inside your report
Illustrative sample — your report is generated from your own responses.
You are a flagger, not a filterer.
You separate a real problem from a clean record about as well as most people do. You just speak up sooner.
A composite without its parts is unusable, so the parts sit inside the same card as the number.
Built for
- Anyone who reads bills, bookings, forms and messages all week and wonders whether they are missing things or over-checking
- People who have been told they worry too much about detail, and people who have been told they do not worry enough
- Students, jobseekers and career changers who want a scored, self-owned profile rather than a free personality quiz
- Anyone curious whether their instinct to speak up is a strength or a habit — and what it costs either way
Find out which of the two you are
30 short records · about 10 minutes · no per-record timer · full bespoke report with your separation figure, its band, its five parts, and your flagger–filterer line.
Take free · full report ₹249 (incl. GST)
Frequently asked questions
Two separate things. The first is separation: how well you tell a record that has a real defect from a record that only looks irregular, expressed as d-prime, the distance between your hit rate and your false-alarm rate. The second is your criterion: where you set the line at which you speak up. They are reported side by side and never combined, because two people with the same separation and opposite lines are a flagger and a filterer, and one accuracy percentage would show them as identical.
By signal-detection theory. You read 30 short records; 13 carry a real defect and 17 are sound. A hit means naming the right defect on a defective record — naming the wrong one counts as a miss and is reported separately. A false alarm means claiming a defect on a sound record. A log-linear correction keeps the arithmetic finite even on a perfect or an empty sitting, and the standard error of the separation figure comes from the spread of your own two rates, printed as a two-in-three and a nineteen-in-twenty span. There is no percentage because a percentage is exactly the number this instrument exists to replace.
No, and the report says so in plain words. Sustained attention is a disability-relevant construct and this instrument is deliberately built not to measure it: the form is short, no record is timed on its own, nothing in the scoring uses how fast you answered, and you may sit it again. It is not a measure of attention, concentration, stamina, eyesight, memory, arithmetic ability or English fluency, and it must never be used as the only basis for a decision about anybody.
About ten minutes for 30 short records, with no per-record timer. You get a full bespoke report: the verdict card with your archetype sentence, your separation figure on a signed axis with its confidence band drawn and written out, the five counts the figure is made of, your flagger–filterer line on its own axis with both ends costed, a naming-precision reading, four content areas as three-way classifications, a modelled reference curve that claims no percentile, and the reliability table that explains every refusal on the page.
₹249 in India, inclusive of GST, or US$2.99 elsewhere, one-time, for one full sitting and the report. Because it sits below the ₹300 threshold the sitting itself starts free and the report is what you buy. Comparable one-time consumer personality and ability reports run roughly $19–69, and none of them publish a standard error, a reliability estimate or a statement of what they refuse to claim.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Measures how well a reader separates a defective everyday record from a sound one, and where that reader sets the threshold at which they call a defect, through thirty original short records. Construct statement: It measures two separable things about how a person reads a short everyday record: how well they tell a record with a real defect from a record that only looks irregular, and where they set the line at which they speak up. It does not measure attention, concentration, vigilance, or any capacity that rises and falls with health, medication, tiredness, pain, mood or the room a person is sitting in. It is not a test of speed, memory, arithmetic ability, eyesight, English fluency, honesty or care about other people, and it must never be used as the only basis for a decision about anybody. Declared response instruction: knowledge of the record in front of the reader, and one instruction only - each record is read and the reader states what, if anything, is wrong with it. No item asks what a person would do, would feel, or is most likely to do, so no behavioural-tendency framing is mixed in anywhere. Item format: B14 SIGNAL VS NOISE STREAM, carried on thirty mcq items and used here for the first time on this platform. Each item prints one short everyday record in full - a bill line, a booking, a request, a form entry, a delivery note, a receipt, a rota line - in a few short sentences, and asks what, if anything, is wrong with it. Four options follow: three specific and individually plausible defect claims and one option saying nothing is wrong. Thirteen records carry a real defect and seventeen are sound. On a defective record the two non-keyed defect claims are specific misreadings of parts that are sound, so sensing that something is off without naming it earns nothing. On a sound record all three defect claims point at something that looks irregular and reconciles once the whole record is read. Noise records outnumber defective ones seventeen to thirteen and every one of the four content areas carries both kinds, so flagging everything cannot pay on any part of the form. Scoring design: C6 SIGNAL-DETECTION SCORING. A hit is naming the defect correctly on a defective record; naming the wrong defect there is a miss and is also counted and reported on its own. A false alarm is claiming any defect on a sound record. Discrimination is d prime, the difference of the two probit-transformed rates, and threshold is c, minus one half of their sum; the two are reported side by side and are never combined into one figure. The log-linear correction adds half an observation to every cell so that a perfect or an empty sitting still returns a finite number, and the standard error of d prime is taken from the binomial variance of both rates and printed as a sixty-eight and a ninety-five per cent span in words as well as drawn. Only answered records enter either side, and below sixty per cent coverage no figure is printed. No chance correction is applied to d prime and the report says why: d prime is centred on zero for a reader who cannot tell the two kinds of record apart, so zero already means chance by construction. An authored option_base_rate is nevertheless declared on every option, published as a range and a mean, and used to draw a modelled reference curve; observed marginals will replace those priors. Naming precision and consistency across the four content areas are computed and reported separately from both headline figures. Named construct areas and source families: signal detection theory and the separation of sensitivity from response bias; criterion setting and the payoff structure that moves it; the probit sensitivity index d prime and its binomial standard error; the log-linear correction for extreme hit and false-alarm rates; low-prevalence effects, in which rare targets are missed more often than their difficulty predicts; vigilance and sustained-attention research, and the reasons it is a disability-relevant construct that this instrument deliberately does not measure; proofreading and error detection in written records; omission blindness, the finding that a missing element is far harder to see than a wrong one; arithmetic reconciliation and self-checking in everyday documents; anchoring on a plausible but wrong reading; misconception-anchored distractor writing and two-sided cue control; and reliability reporting through McDonald's omega, the standard error of measurement and subscore discipline. All items are original works. No trademarked instrument, branded methodology or published item is named or reproduced, and no affiliation with any source is claimed.