Skip to content

Preregistration

Procedural commitment — instantiates Multiple-Testing Discipline

Timestamps the hypotheses, primary outcome, and analysis plan before the data exist, so what counts as confirmatory is fixed in advance rather than chosen after.

Version
v1 · 2026-08-24 · History
Mechanism #
6581
Type
Procedural Commitment
Form family
Rule, Policy & Commitment
Solution family
Evidence, Inference & Validation
Problem family
Uncertainty, Evidence & Inference Failure
Problem subfamily
Experimental Comparison & Hypothesis-Test Design
Origin domain
Statistics & Experimental Design
Also from
Medicine & Healthcare
Instantiates
Multiple-Testing Discipline

Preregistration is the ex-ante commitment that gives the confirmatory label its meaning. Before any data are collected, the researcher writes down — and timestamps in a place they cannot later edit — the hypotheses, the primary outcome, the sample and stopping rule, and the exact analysis. Anything that later matches the frozen plan is confirmatory; anything else is exploratory. The defining idea, and what separates it from every sibling, is time: preregistration is a document created before the data exist, and its whole power comes from being un-revisable after results are seen. It does not run a test, reserve a data partition, or correct a threshold; it fixes, in advance, which claims are eligible to be confirmatory at all.

Example

A psychology lab plans a study on whether a brief mindfulness exercise improves exam performance. Left to write up afterward, the team would face a hundred small temptations — report the outcome that moved, drop the participants who "clearly didn't comply," add the covariate that sharpens the effect, or announce a hypothesis that only occurred to them once they saw the pattern. So before collecting a single data point, they post a preregistration to a public repository such as the Open Science Framework: the hypothesis (mindfulness raises test scores), the single primary outcome (final exam percentage), the target sample size and stopping rule, the exclusion criteria, and the exact statistical model. The document is timestamped and locked. When the study is done, readers can compare what was planned against what was reported: the pre-specified test of exam scores is confirmatory, and if the team also notices an intriguing effect on self-reported stress that was not in the plan, it is honestly flagged as exploratory. The frozen record is what makes "we predicted this" checkable rather than a claim of memory.

How it works

  • Specify before data. Write the hypotheses, primary outcome, sample and stopping rule, exclusions, and analysis while no results yet exist.
  • Fix the confirmatory family. Name exactly which tests are pre-specified; the set is closed, so new comparisons invented later cannot slip into the confirmatory tier.
  • Timestamp immutably. Deposit the plan somewhere it is dated and cannot be silently altered, so priority of the prediction is verifiable.
  • Adjudicate against the plan later. After results, each reported claim is checked against the frozen document and labeled pre-specified/confirmatory or post-hoc/exploratory. Deviations are permitted but must be disclosed as such, which is what keeps hypothesizing-after-results honest.[n1]

Tuning parameters

  • Specificity — how tightly the plan pins each choice; a fully specified analysis leaves no wiggle room but is rigid if reality departs from expectations, while a looser plan is flexible but easier to reinterpret.
  • Bindingness — a private timestamp, a public registration, or a peer-reviewed Registered Report where the plan is accepted before data; stronger forms deter reinterpretation more but cost more up front.
  • Deviation policy — how departures from plan are handled and disclosed; transparent deviation preserves credibility, silent deviation destroys it.
  • Scope — a single primary hypothesis versus a fuller pre-specified family; broader plans cover more but demand foreseeing more in advance.

When it helps, and when it misleads

Its strength is that it draws the exploratory–confirmatory line before anyone can see which side flatters them, which is the only moment the line is incorruptible. It makes prediction verifiable, deters outcome switching and HARKing, and — crucially — does all this while leaving exploration fully legitimate, since off-plan findings are welcome as long as they are labeled.

Its failure mode is that a preregistration constrains only what it covers, and only if anyone checks. A vague plan preregisters nothing meaningful; a detailed plan can still be quietly deviated from if reviewers never compare registration to report; and preregistration says nothing about whether the study was well-powered, unconfounded, or important — a pre-specified result can be pre-specified and still wrong. The classic misuse is the decorative preregistration: a plan loose enough to permit any analysis, waved as a credibility badge. The guarding discipline is to make the plan specific enough to bind, deposit it where deviations are visible, and treat registration as a claim about process, not a warrant of truth.

How it implements the components

  • exploratory_confirmatory_boundary — the timestamped plan is the boundary: on-plan analyses are confirmatory, everything else exploratory, drawn before results could bias the line.
  • claim_family — pre-specifying which tests are planned closes the confirmatory family, so comparisons invented after seeing data cannot be smuggled into it.
  • result_status_label — every later claim inherits a status (pre-specified vs post-hoc) by comparison against the frozen document.

It commits the plan but does not itself execute the fresh confirmatory test a promising lead must pass — that confirmation_requirement is Confirmatory Follow-Up's stage; preregistration is the ex-ante document, not the follow-up study.

Editorial Notes

Form Classification

Form family: Rule, Policy & Commitment

Rationale: Preregistration operates as a standing rule, threshold, contractual commitment, or policy constraint governing future conduct because it timestamps the hypotheses, primary outcome, and analysis plan before the data exist, so what counts as confirmatory is fixed in advance rather than chosen after.

Independent corroboration: The frozen evidence defines Preregistration as 'Timestamps the hypotheses, primary outcome, and analysis plan before the data exist, so what counts as confirmatory is fixed in advance rather than chosen after', so its operative form is Rule, Policy & Commitment.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Statistics & Experimental Design

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Preregistration is most plausibly rooted in the statistics_experimental_design tradition because its characteristic form depends on probability, calibrated inference, experimental design, and uncertainty analysis. The assignment tracks that formative lineage, not the many settings in which the mechanism can now be applied.

Related originating lineages:

  • Medicine & Healthcare — The medicine_healthcare tradition materially shaped Preregistration through its own practice of clinical practice, patient safety, and health-system intervention.

Review resolution: Both blind reviewers agree that statistics experimental design is the primary origin. Explicit reconciliation resolves origin mode disagreement. Formative alternate lineages are retained as medicine_healthcare; later breadth of use is recorded separately as domain_reach=multi_domain, while origin_mode=cross_disciplinary_synthesis describes the relationship among origin lineages.

Review outcome: Reconciled after independent review; high confidence.

Notes

[n1] HARKing — Hypothesizing After the Results are Known — is presenting a hypothesis formulated after seeing the data as if it had been predicted in advance; preregistration defeats it by fixing and timestamping the hypotheses before data exist, so any post-hoc idea is visibly labeled exploratory.