Second Look: Research Integrity in PracticeSecond Look: Research Integrity in Practice. Integrity training tests whether you can recall a policy. Nothing has ever gone wrong at the moment somebody recalled a policy.
It goes wrong when a senior colleague asks to be on the paper, when three runs make a figure look cleaner, when an editor asks a direct question, and when a tool writes a paragraph that reads better than the truth. This measures what you would actually do — and then, much later in the same sitting, how well you judge the choice you made.
It asks you a second time — and scores how well you judged yourself
Research integrity is not a belief. It is a set of ordinary decisions: who goes on the paper and in what order, what happens to the data and the records behind a figure, what an editor or a reviewer is told, and what a machine is allowed to write. This measures those decisions. Ten cases, four options each, and every option is something a working researcher does — a supervisor's instruction accepted, an exclusion made without a rule, a link to an author judged too old to mention, a fluent paragraph left standing. The difference between the options is degree, not decency, so the paper about ethics that everyone passes is not the paper you are sitting.
Then the instrument does something no other integrity assessment does. Sixteen positions later, each of those ten cases comes back in a sentence and asks you how strong the choice you made on it actually was. Your appraisal is compared against the measured quality of that same choice, case by case, and reported as three separate readings rather than one score: bias, whether you read your own work as stronger than it was or sold it short; resolution, whether your appraisal tracks the quality at all; and scatter, how much you move around that lean. They are kept apart on purpose, because a person can be well calibrated and useless, or discriminating and permanently over-confident, and a single number hides exactly that difference. The accuracy of your self-appraisal is scored. Its severity is not — no sentence anywhere praises or criticises you for being hard on yourself.
One limitation is printed near the top of the report rather than in a footer, because it is ours and not yours: the engine cannot replay your earlier answer to you, so the second look is appraised from memory. Below that, the spine of the report is a stack of case panels, and each panel has three bands — the situation, then what you chose against what a stronger response contained, and then the band that is the whole point: the difference, stated in one sentence. Not the model answer on its own. A model answer is something to copy on a case you have already seen. The sentence naming what the stronger response contained that yours did not is something you can use on a situation the report never describes.
What you walk away with
Three readings with three scales and three plain sentences. Whether you lean high or low, whether your appraisal tracks the quality at all, and how far it moves around that lean.
A ladder of the ten cases with the measured quality and the rating you gave it, and the same figures again as a table so nothing depends on seeing colour.
For each of your widest gaps: the situation, your choice, what a stronger response contained, and one sentence naming what it had that yours did not.
Expressed against a respondent choosing at the declared rates, with the 68% and 95% spans both printed and the assumed reliability stated as a target rather than a measurement.
Six exercises can place you in one of three bands. They cannot support a figure with a decimal in it, and the report says so rather than printing one.
Inside your report
Illustrative sample - your report is generated from your own responses.
Three runs in your dataset sit far from the rest. The figure is cleaner without them.
Drop the three, and say in the methods that outliers were removed.
Keep all runs, and state the rule you would use to exclude any.
A stronger answer let the reader see what the exclusion changed instead of naming it only in general terms.
The bottom band is the point. A model answer is something to copy on a case you have already seen; the difference is something you can use on one you have not.
Three readings, never fused into one. A person can be well calibrated and useless, or discriminating and permanently over-confident, and a single number hides the difference.
Built for
- Doctoral researchers, postdocs and academic staff who want a measure of research judgment rather than another certificate of completion
- R&D scientists and industry research groups where authorship, data records and AI-assisted drafting are live questions and no course covers them together
- Research offices, graduate schools and journal staff running integrity programmes who need a diagnostic that works in any country and names no regulation
Find out what you would actually do — and how well you judge your own call
32 scored exercises - about 40 minutes - a bespoke report with three separate calibration readings, the case ladder, and the difference annotated on every panel.
₹999 (incl. GST) · assessment and full report, nothing further to pay
Frequently asked questions
It measures judgment on the ordinary decisions that make research trustworthy - who goes on the paper and in what order, what happens to the data and the records behind a figure, what gets disclosed to an editor or a reviewer, and what a machine may write in a manuscript - and, separately, how accurately a person judges the quality of their own decision afterwards. It does not measure honesty, scientific ability, publication record, knowledge of any journal's or funder's policy, or English writing skill.
No. It is about the conduct of research itself, and it is written for the people doing that research: academic staff, doctoral researchers, R&D scientists, and the research office and journal staff around them. Nothing in it is about a student sitting an examination. If that is what you need, the catalogue has a separate assessment for AI use in student work.
No journal, publisher, funder, tool or country-specific rule is named anywhere in it. Every case is written around the principle - contribution decides credit, the record has to survive the person who made it, the editor weighs the link rather than you, and generated text is a draft until somebody verifies it - so it reads the same in any research system and does not go out of date when a policy is amended. It complements a local policy module rather than replacing it.
Because reviewing your own work is the control that runs most often and is measured least. The second-look items are matched to your earlier cases by a case tag rather than by option position, so the comparison survives the option shuffling, and the report gives you three separate readings instead of one: your lean, whether your appraisal tracks the actual quality, and how much it moves around. One honest limitation is printed near the top of the report: the engine cannot replay your earlier answer to you, so the appraisal is from memory. That is our constraint, not your failure, and the report says so in those words.
Rs 999 in India including GST, or US$9.99 elsewhere, one time, for one full sitting and report. The alternatives are not priced like this: research-integrity training is sold to institutions as an annual e-learning subscription per seat, and the funder and publisher modules that individuals can reach for free are recall tests on a policy document, which is a genuinely different thing from a scored measure of judgment - they are cheaper than this and they answer a different question. Nothing we can find scores research judgment for an individual, and nothing at all scores how well a researcher appraises their own decision. Organisations can use AssessAll credits at 26 credits per person.
Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.
Browse the catalogue →Methodology: Measures judgment on the everyday decisions that make research trustworthy, and separately the accuracy with which a person appraises their own decision, through original situational, sorting and self-report items. Construct statement: It measures judgment on the ordinary decisions that make research trustworthy - who goes on the paper and in what order, what happens to the data and the records behind a figure, what gets disclosed to an editor or a reviewer, and what a machine may write in a manuscript - and, separately, how accurately a person judges the quality of their own decision afterwards. It does not measure honesty, scientific ability, publication record, knowledge of any journal's or funder's policy, or English writing skill. Declared response instruction: behavioural tendency throughout - every situational stem asks what the respondent is most likely to do rather than what a person should do, because instructed-knowledge framing is the more fakeable of the two. Item format: ten four-option case items in which every option is an action a working researcher could defend and the spread between them is degree; ten retrospective re-scoring items placed late in the sitting, each restating one of those cases and asking the respondent to appraise the strength of the choice they made earlier, aligned to their case by a case tag rather than by option position; four balanced diagnosticity sorts, each listing six findings about a disputed situation, three of which could have come out the other way and three of which sound weighty while leaving the answer where it already was; and eight balanced-keyed self-check statements, four of them worded so that agreement is not the flattering answer. Platform limitation, stated plainly because it is ours and not the respondent's: The engine cannot show a respondent their earlier answer again, so the second-look exercises are appraised from memory. That is a limit of this platform, not a fault in the person taking it. Scoring is calibration scoring in the spirit of a Brier decomposition. The appraisal and the measured quality of the same choice are compared case by case and reported as three separate readings - a signed and absolute bias, a resolution correlation, and the mean absolute scatter around the bias - because fusing them into one number hides the difference between a person who is well calibrated and useless and a person who discriminates well but sits high. The accuracy of self-appraisal is scored and its severity is not; no reading praises or criticises a respondent for being hard on themselves. Calibration is reported only when at least eight cases have both halves answered, and resolution is withheld with a plain-language note when the appraisal vector has no variation in it, because a correlation on a flat vector is a zero that reads like a finding. The judgment composite is a graded partial-credit percentage corrected against authored per-option base rates, since an uncorrected graded percentage sits near half by construction, with the raw figure printed beside the corrected one and the assumed reliability stated in the open as a target rather than a measurement. Each of the four areas rests on eight exercises and therefore carries a three-way classification and never a number or a percentile. Construct areas drawn on: contribution-based authorship criteria and the difference between an author and an acknowledgement; gift, guest and ghost authorship and the seniority pressure behind them; author-order conventions and contribution statements; data availability, stewardship and the difference between personal files and the record behind a paper; selective reporting, undisclosed exclusion rules and analytic flexibility; version history as the evidence of when an analysis decision was made; conflict-of-interest disclosure in peer review and the editor rather than the reviewer as the person who weighs it; prior-publication and preprint disclosure; funder influence over publication; the diagnosticity principle, that a finding is worth its weight only if it could have come out the other way; automation bias and the supervision of generated text, including unverified citations and invented method steps; and metacognitive accuracy, calibration and resolution as separable properties of self-appraisal. All items are original works, no trademarked instrument or branded methodology is named or reproduced, no journal, publisher, funder, tool or country-specific rule is named, and no affiliation with any source is claimed. AssessAll original design.