Estimation & Forecast CalibrationEstimation & Forecast Calibration. Every estimate is two numbers, and only one ever gets written down.
The figure, and how sure you are of it. This is a scored profile of the second one - whether your stated confidence matches how often you are right, and whether your ranges are narrow enough for anybody to use.
It scores both failures, and almost nothing else does
Ten exercises give you data and ask for the range you are eighty per cent sure holds the answer. Some answers land inside what the data has already shown and some land outside it, which is the point: a range built from the observed spread of five past cases misses roughly half the time, and this is the single most reproducible finding in the whole judgement literature. But the opposite failure is real too. A range wide enough never to be wrong is correct and useless, and nobody can plan against it.
So both are scored, and reported apart. You get a hit rate against the eighty per cent you were asked for - that is reliability - and an average width against the tightest range that would still have worked - that is resolution. A person can be perfectly calibrated and useless, with every range enormous, or highly informative and overconfident, and fusing the two into one accuracy figure destroys exactly the distinction that makes the result actionable. The underlying maths is a proper interval score: the width you chose, plus a penalty proportional to how far outside it the answer fell.
Eighteen further situations test the judgment around the number rather than the number itself. Whether you start from what similar work actually took or from what this team expects. What survives when a meeting wants one figure instead of a range. What you do on the day you know an estimate is wrong and nobody has asked you. What you give somebody whose preference for a small number is obvious. Every range exercise supplies its own data, so nothing turns on world knowledge, a country, a currency or a named product.
What you walk away with
How often your eighty per cent ranges actually contained the answer, drawn as a band, with the gap between claim and performance named.
How wide your ranges ran against the tightest one that would still have worked. Being right is not the same as being useful.
Overconfident, Vague or Calibrated, decided across ten exercises rather than from any single one, because interval width varies with risk appetite.
Contained and tight, contained but wide, or missed - each with your width, the tightest workable width and the answer, so nothing rests on a chart.
Where your estimates come from, what happens to the range in a meeting, and what you do the day you know it has moved.
Inside your report
Illustrative sample - your report is generated from your own responses.
Built for
- Delivery, product and engineering leads whose dates other people plan against
- Operations, finance and planning professionals who forecast for a living
- Employers and programmes who want calibrated judgment rather than confident judgment
Find out whether your eighty per cent is really eighty per cent
34 scored exercises - about 45 minutes - a full bespoke report with every score drawn as a band, your hit rate against your own claim, your width against the tightest workable range and the ten exercises laid out in full.
₹1,199 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
Four capabilities: hitting the confidence level you claim, giving ranges narrow enough to be useful, starting from what comparable work actually took rather than from the plan, and how an estimate is given and revised. Ten exercises ask for an eighty per cent range on data supplied in the item itself; eighteen situations test the judgment around the number. It does not measure arithmetic, numeracy, planning tools or delivery speed.
By a proper interval score: the width of the range you chose, plus a penalty proportional to how far outside it the answer fell. That rewards narrow ranges and punishes misses, which is why a range wide enough never to be wrong does not win. Your result is also expressed against what a typical estimator would score, because every range on offer carries a published base rate - the share of estimators expected to pick it - and a uniform null would flatter everybody. Hit rate and width are then reported apart, because one accuracy figure would hide which of the two failures you have.
No. Every range exercise supplies its own data inside the item - six job times, five delivery delays, three past projects, a week of customer counts - so nothing depends on world knowledge, a country, a currency, an industry or a named product. What is being tested is how you turn a handful of observations into a range and how much confidence you attach to it, which is why the same form works anywhere.
About 45 minutes for 34 scored exercises. You get a bespoke report where every score is drawn as a band rather than a point, with the estimate a tick inside it and one plain line about what a retest would do - which is the only honest frame for an instrument whose whole subject is the width of a claim. Plus your hit rate against your claim, your width ratio, the interval-score ledger, the ten exercises as a full table, four capability classifications and one if-then change.
Rs 1199 in India including GST, or US$11.99 elsewhere, one time, for one full sitting and report. Estimation and forecasting workshops are priced per seat in the hundreds and teach a technique rather than measuring calibration; superforecasting-style training programmes cost more again; and no commercial pre-hire suite scores calibration at all. Organisations can use AssessAll credits at 40 credits per person.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Measures estimation and calibration through original interval-selection and situational items keyed to published constructs. Construct statement: it measures whether a person's stated confidence matches how often they are right, whether their ranges are narrow enough to inform a decision, and whether they estimate from what comparable work actually took. It does not measure arithmetic ability, numeracy, planning tools, or delivery speed. Declared response instruction: behavioural tendency throughout - every situational stem asks what the respondent would actually do, and every interval item asks which range they would give. Sources drawn on across the construct area: the overconfidence and interval-estimation literature, in which stated ninety per cent ranges typically contain the answer between forty and sixty per cent of the time; calibration training research showing that the miscalibration is trainable and that feedback across many items, never one, is what moves it; proper scoring rules for interval forecasts, which reward narrow intervals and penalise misses in proportion to the distance outside; the decomposition of a calibration score into reliability and resolution, so that being well calibrated and being informative are reported apart; the outside-view and reference-class work on planning, in which what similar work actually took outperforms a detailed inside estimate; the planning-fallacy literature and its finding that decomposition improves estimates without removing the optimistic bias; base-rate neglect; anchoring effects on numbers supplied by other people, including the requester's evident preference; the small-sample fallacy, in which the observed range of a handful of cases is mistaken for the range of the process; regression and the tendency to extend a recent trend as though it were a distribution; verbal probability research showing that words such as likely and fairly sure are read as widely different numbers by different readers; the point-versus-interval reporting problem, in which a range converted to its midpoint for a meeting loses the uncertainty while appearing to keep it; and situational-judgment-test validity with behavioural-tendency response instructions. All items are original works, and every interval item supplies its own data so that no item depends on world knowledge, a country, a currency or a named product. No trademarked instrument is reproduced and no affiliation with any of the sources above is claimed. Scores are provisional: the base rates used for chance correction are authored priors, published in the report, to be replaced by observed marginals once live data exists.