Applied Judgment Assessment
Applied skill assessment · drug safety associates, pharmacovigilance officers, clinical operations and medical information teams, CRO and pharma training functions · browse the full catalogue

Pharmacovigilance and Adverse Event Case Handling Assessment for Drug Safety and Clinical Operations TeamsThe rash cleared when the tablet was stopped. Did you call the case right, and did you know you had?

Forty exercises on invented adverse event cases: eight case calls, each followed by a slider asking how sure you are, and thirty-two keyed exercises across seriousness and the reporting clock, causality reasoning, case completeness and follow-up. Scored as accuracy corrected against the ordinary case handler, and as calibration in two figures that are never added. Reported as a dumbbell: one row per case family, how often you were right beside how sure you said you were, sorted by the gap.

32 minutes40 scored exercisesEvidence-keyed scoringGlobal · INR & USD

A pharmacovigilance test that measures the call and the confidence in it, separately

The Pharmacovigilance and Adverse Event Case Handling Assessment is a thirty-minute check of whether a person handling an adverse event case reaches the right call under a stated procedure, across seriousness, causality, case validity and follow-up, and whether they know how sure to be of it, reported as one accuracy figure and two calibration figures that are never combined.

Most drug safety tests ask what the rule says. This one puts you inside the case. A woman admitted overnight after a new blood pressure tablet. A man whose kidney function fell after a new painkiller during a week of vomiting. A journal article naming its authors, the patient's age and the active substance. A serious liver injury with no dates and three days left on the clock. Each asks for one call, and after each of eight of them, a slider asks for a number: the chance, from 0 to 100, that the call you just made was right. Four probes ask the chance you were right and four the chance you were wrong, so a habit of stating one high number is tested on calls it would get wrong.

Accuracy is corrected against the ordinary case handler. Every option, pair, position and value carries a declared share of working case handlers expected to choose it, and your score is how far you sat above or below a respondent answering at those shares. Zero is that respondent, not half marks, and the shares are printed so they can be disagreed with. An unanswered exercise leaves the denominator; a sitting with fewer than twenty keyed answers is refused rather than scored.

Calibration is a Brier score decomposed into reliability and resolution, and the two are reported apart: reliability is whether, when you said 80, you were right about 80 per cent of the time; resolution is whether your confidence rose and fell with your accuracy at all. A person can be well calibrated and undiscriminating, or discriminating and over-confident, and a single number would hide which. Eight probes is the floor for a calibration figure, and the page says so beside the figure.

The report is a dumbbell of difference, drawn self against self. One row per case family: a filled square for how often you were right on that family's exercises and a hollow circle for how sure you said you were on that family's probes, joined by a line and sorted by the size of the gap, largest first. Where a row's two bands cross, it prints not distinguishable at this length instead of a gap, because two probes per family is a short length and the instrument declines to claim a difference it cannot support.

Every rule is stated inside the exercise that uses it, as the criteria your procedure applies and the clock your procedure sets. No country's timeline is assumed, no regulator's form is reproduced, no terminology dictionary and no safety database product is named. The report says in plain words what it did not measure: any jurisdiction's regulations, clinical judgement about a real patient, speed, or how you work under a backlog. It is not a regulatory qualification, it authorises nobody to submit a safety report, no real case is assessed, and nothing in it is medical advice.

One accuracy figure, two calibration figures that are never added, and four case families drawn as a dumbbell:
Seriousness and the reporting clockCausality reasoningCase completenessFollow-up and query

What you walk away with

Seriousness and the reporting clock

Which outcomes meet a seriousness criterion and which do not: an admission, a lasting loss of sight, a convulsion in the emergency room, against an overdose without harm and a mild rash that is merely unexpected. And the day the clock starts, including when seriousness arrives with a follow-up.

Causality reasoning

Time to onset, dechallenge and rechallenge, and the alternative explanation that lowers a category without clearing the drug. Two medicines stopped together, a reporter who says definitely, and nausea two weeks after a short-acting tablet was stopped.

Case completeness

The four minimum criteria and what identifies a patient: an elderly mother, a boy of about seven, a ward of patients, a customer whose tablet is not named. A journal article as a valid source, and a second reporter of the same patient as a duplicate rather than a new case.

Follow-up and query

What to ask for first, which gaps are worth the one request a reporter will answer, submitting on time and following up in parallel, closing follow-up as lost only after documented attempts, and which new facts need a follow-up report.

Calibration, in two figures

Reliability: when you said 80, were you right 80 per cent of the time. Resolution: did your confidence separate the calls you got right from the ones you got wrong. Each with its standard error and bands, and three reference respondents priced on your own pattern: fifty on everything, ninety on everything, and the declared ordinary chance on every call.

The dumbbell, and one change

The family where your confidence ran furthest ahead of your accuracy, the strength paired with what it costs, one if-then sentence written for the gap the page found, and a re-sit date with the number of points that would count as real change.

Inside your report

Illustrative sample — your report is generated from your own responses.

The dumbbell of difference: how often you were right beside how sure you said you were, one row per case family, largest gap first
Largest gap first050100The gap1. Causality4 of 8 rightsaid about 88ordinary 41▲ over-sure by 38 points2. Follow-up7 of 8 rightsaid about 66ordinary 58▬ not distinguishableat this length3. Seriousness6 of 8 rightsaid about 81ordinary 55▬ not distinguishableat this length4. Completeness6 of 8 rightsaid about 74ordinary 52▬ not distinguishableat this lengthfilled square: share righthollow circle: stated chance · bar 68%, dashed tick ordinary

How to read it: the rows are sorted by the size of the gap, so the family at the top is where you are most wrong about yourself. Here the causality row is the only one whose two bands do not cross: four of eight right, and a stated chance of about 88. The other three rows print not distinguishable at this length, because two probes per family is a short length and the page will not claim a difference it cannot support.

FamilyRightShare right (68%)OrdinaryStated (68%)The gap
1. Causality4 of 8■ 50 (36 to 64)41○ 88 (75 to 100)▲ over-sure by 38 points
2. Follow-up7 of 8■ 88 (77 to 99)58○ 66 (53 to 79)▬ not distinguishable at this length
3. Seriousness6 of 8■ 75 (60 to 90)55○ 81 (68 to 94)▬ not distinguishable at this length
4. Completeness6 of 8■ 75 (60 to 90)52○ 74 (61 to 87)▬ not distinguishable at this length
Calibration in two figures that are never added, with the direction, the bins, and three reference respondents priced on your own pattern
+12
▲ Over-suredirection: mean stated chance 75 minus share right 63 · 68%: +1 to +23 · 95%: −10 to +34
Reliability, never added to the other
0.046
Smaller is better; 0 is perfect

68%: 0.018 to 0.074 · 95%: 0.000 to 0.101

How far, within each confidence bin, your stated chance sat from your actual hit rate. Zero would mean that when you said 80 you were right 80 per cent of the time.

Resolution, never added to the other
0.109
Larger is better; 0 means no separation

68%: 0.061 to 0.157 · 95%: 0.015 to 0.203

How far your hit rate within each bin departed from your overall hit rate of 63 per cent. Zero would mean your confidence did not separate the calls you got right from the ones you got wrong.

Brier 0.171, of which uncertainty 0.234 belongs to the calls and not to you. Eight probes is the floor for a calibration figure, and eight is exactly what this instrument carries, so the figures are printed with their standard errors and read as a placement.

Stated chance binProbesMean statedActually right
0 to 491400%
50 to 6926250%
70 to 8427850%
85 to 100392100%
Reference respondent, on your own right-and-wrong patternBrierReliabilityResolution
● you0.1710.0460.109
◇ fifty on every probe0.2500.0160.000
◇ ninety on every probe0.3100.0760.000
◇ the declared ordinary chance on every call0.2430.0310.022

There is no zero on a calibration scale, so the three reference respondents are priced instead. Fifty on everything and ninety on everything both have zero resolution by construction, because a number that never moves cannot separate right calls from wrong ones.

Built for

  • Drug safety associates and case processors who want to know which family of call they are surest about and least right about
  • Pharmacovigilance officers, medical information and clinical operations staff who handle adverse event reports and their follow-up
  • CRO and pharma training functions who want a check that says over-sure or under-sure per case family, rather than a percentage
  • Managers hiring or onboarding case handlers, who need a report with its null, its bands and its refusals printed on the page

Find out where you are most wrong about yourself

40 exercises across six formats · about 30 minutes · accuracy corrected against the ordinary case handler, calibration in two figures never added, and a dumbbell per case family sorted by the gap.

₹999 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and your receipt is sent to that address.

Secure Razorpay payment · ₹999 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

What does the test actually ask?

Forty exercises about invented adverse event cases. Twelve are case calls with one keyed answer: is this serious, when did the clock start, which causality category fits, is this a valid case, what is asked for first. Eight of those are followed by a slider that names the case and asks the chance, from 0 to 100, that your call was right or, on four of them, wrong. The other twenty are true or false claims, select-all lists, match-the-following and ordering exercises on the same four families. No real case, no regulator's form and no code appears anywhere.

How is confidence scored, and why is it not just added to accuracy?

Confidence is scored as a Brier score over the eight probes and decomposed into reliability and resolution. Reliability is whether your stated chance matched how often you were actually right within each confidence bin. Resolution is whether your confidence rose and fell with your accuracy at all. They are printed apart because a person can be well calibrated and undiscriminating, or discriminating and over-confident, and a single figure would hide which. Eight probes is the floor for a calibration figure; with fewer, only the direction is printed in words.

What is the dumbbell and why are some rows marked not distinguishable?

One row per case family. A filled square marks how often you were right on that family's keyed exercises, and a hollow circle marks how sure you said you were on that family's two probes. The two are joined by a line and the rows are sorted by the size of the gap, largest first. Each dot carries a band. Where the two bands cross, the row prints not distinguishable at this length instead of a gap, because two probes per family is a short length and the instrument will not claim a difference it cannot support.

Is this a regulatory qualification, and does it depend on one country's rules?

No to both. Every rule is stated inside the exercise that uses it, as the seriousness criteria your procedure applies, the causality categories your procedure uses and the reporting clock your procedure sets; no jurisdiction's timeline is assumed. The public frameworks the rules follow are named once in the methodology note, with affiliation disclaimed. It is not a qualification, it does not authorise anybody to submit a safety report, no real case is assessed, and nothing in it is medical advice. The report says this in a refusal block, never in a footer.

How long is it, what does it cost, and what if I leave exercises unanswered?

About thirty minutes for forty exercises across six formats. The sitting is free; the report is the product, priced at ₹999 in India, inclusive of GST, or US$9.99 elsewhere, one time. An unanswered exercise leaves the denominator rather than scoring zero. A sitting with fewer than twenty of the thirty-two keyed exercises answered is not reported: the refusal is printed where the dumbbell would have been, and no family row, no accuracy figure and no calibration figure is printed beneath it.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue →

Methodology: Forty original exercises across six formats: twelve single-choice case calls, eight confidence probes on a slider, seven true-or-false claims, five select-every-one-that-applies lists, four match-the-following exercises and four ordering exercises. One response instruction is declared for the whole instrument and it is a KNOWLEDGE instruction: what the case supports under the procedure stated in the exercise, never what the respondent would do or feel. The eight confidence probes are part of the same knowledge instruction: each names a case the respondent has just called and asks for a percentage chance that their own earlier call was right (four probes) or wrong (four probes), which is a metacognitive judgement about their own answer and is never scored or described as a behavioural tendency, a trait or a personality style. Construct statement: this measures whether a person handling an adverse event case reaches the right call under a stated procedure, across seriousness and the reporting clock, causality reasoning, case completeness and follow-up, AND whether they know how sure to be of it; it does not measure knowledge of any one country's regulations, any regulator's forms or timelines, any terminology dictionary, any safety database product, clinical judgement about a real patient, medical writing, speed, or how a person works under a case backlog, and it is not a qualification. Scoring is C9, calibration scoring. ACCURACY is computed over the thirty-two keyed exercises: each answered exercise yields a quality in the unit interval and a chance term from its declared prior, which is the empirical marginal the instrument declares for that exercise: option_base_rate on single-choice and true-or-false items (the declared share of working case handlers expected to choose each option), per-option select rates on select-all items combined through the expected F1, pair_base_rate on matching items and position_base_rate on ordering items. The chance-corrected score is (raw minus expected chance) over (one hundred minus expected chance), so zero is the ordinary case handler answering at the declared marginals, never half marks; the marginals are printed on the page with their spread, and observed data will replace them. An unanswered exercise leaves the numerator, the denominator and the chance term together rather than scoring zero. An empty sitting scores exactly zero. Fewer than twenty keyed answers of thirty-two is refused, not reported: the refusal is printed where the dumbbell would have been, and no family, no calibration figure and no placement is printed beneath it. CALIBRATION is computed over the eight probes. The stated chance p is the slider value over one hundred on a probe phrased as the chance of being right, and one minus that on a probe phrased as the chance of being wrong; the outcome o is one if the named call was right. The Brier score B is the mean squared difference between p and o, decomposed after Murphy into reliability (the mean squared gap between stated and observed within four declared confidence bins, smaller is better), resolution (how far the hit rate within each bin departs from the overall hit rate, larger is better) and uncertainty (a property of the calls, not the respondent). Reliability and resolution are reported SEPARATELY and never summed, because a person can be well calibrated and undiscriminating, or discriminating and over-confident. The signed gap between mean stated chance and hit rate is the direction: over-sure or under-sure. Four probes attach to calls declared hard (keyed prior at or under .40) and four to calls declared easy (at or over .55), so a respondent who states the same high number on every probe is scored on calls they are likely to get wrong as well as right; the builder counts the split it received and prints it, and a respondent whose raw slider values do not move between right-phrased and wrong-phrased probes is flagged for review rather than scored as calibrated. Eight probes is the floor for a calibration figure, and the page says so: with all eight answered and their calls answered, the Brier score, reliability, resolution and direction are printed each with a standard error from a deterministic bootstrap and a 68 and 95 per cent band; with four to seven, only a placement in words is printed for the direction and the figures are refused; with fewer than four, the calibration strand is refused entirely. There is no zero on a calibration scale, so three reference respondents are priced on the reader's own right-and-wrong pattern instead: fifty on every probe, ninety on every probe, and the declared prior of each call stated as the confidence. The report is D4, the dumbbell of difference, adapted to self against self: one row per case family, a filled square for the share of that family's keyed exercises the respondent got right and a hollow circle for the mean stated chance on that family's two probes, each carrying a word as well as a shape, joined by a line and sorted by the size of the gap, largest first. The accuracy dot carries the declared expected-chance tick for that family and a binomial standard-error band; the confidence dot carries a band from a declared standard deviation of stated confidence of eighteen points, an assumption observed data will replace. A gap is printed in words and points only where the two 68 per cent bands do not cross; where they cross, the row prints not distinguishable at this length instead of a gap, and two probes per family is stated on the page as the reason that will often be so. Subscore discipline: each family carries eight keyed exercises; at the assumed inter-item correlation of .24 a family carries a chance-corrected figure only when all eight are answered and its omega, printed unrounded, is at or above .70; otherwise a three-way placement in words, never a number and never a percentile. The composite over thirty-two exercises has an assumed omega of about .91 and carries a figure with its 68 and 95 per cent bands from SEM equal to SD times the square root of one minus omega, with SD declared at twenty-four on the corrected scale from a modelled mixed-ability reference population run through this builder. Reliability weighting is used only when the omega spread across families exceeds .15; otherwise unit weights, stated. Every omega is McDonald's omega estimated from item count and the assumed inter-item correlation because no live data exist yet, and the page says so. Fixed habits are priced through the real builder before shipping: always serious, never serious, always the reporter's view, ninety on every probe and fifty on every probe each land at or below the ordinary case handler on the strand they touch. Two true-or-false claims tagged as a pair probe the same rule, that a sex and an age group identify a patient, through different cover stories, and disagreement between them is printed for review rather than scored. A careless responding composite is printed to whoever reviews the sitting as a flag count, never to the respondent as a judgement. Sources drawn on for the constructs and the scoring: Brier (1950) on the verification of probability forecasts; Murphy (1973) on the reliability-resolution-uncertainty decomposition; Lichtenstein, Fischhoff and Phillips (1982) on the calibration of probabilities; Yates (1982) on external correspondence in probability judgement; Koriat, Lichtenstein and Fischhoff (1980) on the reasons for confidence; Moore and Healy (2008) on the three forms of overconfidence; Hill (1965) on association and causation; Naranjo and colleagues (1981) on a structured probability scale for adverse drug reactions; Meyboom and colleagues (1997) on causality assessment in pharmacovigilance; Agbabiaka, Savovic and Ernst (2008) on the systematic review of causality assessment methods; Edwards and Aronson (2000) on the definition and diagnosis of adverse drug reactions; Waller and Evans (2003) on the conduct of pharmacovigilance; Haladyna, Downing and Rodriguez (2002) on item writing and cue control; and the three public frameworks named in framework_note, each cited once and with affiliation disclaimed. All forty exercises are original works written for this instrument. No case describes a real patient, reporter, product, company or regulator. No commercial instrument's items or name are used or implied, no licensed terminology dictionary is named, no safety database product is named, and no competitor is named. This is not a regulatory qualification, it does not authorise anybody to submit a safety report, no real case is assessed, and nothing in it is medical advice.