Skip to content

Multiple Testing Discipline

Control false discoveries when many comparisons, claims, or tests are being tried.

The Diagnostic Story

Symptom: A result looks significant, or a pattern looks real, but the search that found it involved dozens of comparisons, subgroups, metrics, or time windows that nobody is counting. The one finding that got presented was selected because it was the most interesting — not because it was pre-specified. Failed comparisons disappear from the record, so the claimed discovery looks far more credible than the actual search process warrants. When someone tries to reproduce it, the pattern is gone.

Pivot: Define the family of related tests or claims being made, track the full search process rather than just the selected result, and choose a decision rule that accounts for how many chances existed for an accidental pattern. Separate exploratory evidence — leads worth investigating — from confirmatory evidence that is ready to justify action.

Resolution: The rate of false discoveries is bounded because the opportunity set that produced any selected finding remains visible and is factored into interpretation. High-stakes action requires confirmation stronger than the initial search result, so the distinction between a lead and a finding is preserved rather than collapsed. The discovery record becomes more credible because null results and failed comparisons are retained alongside the claimed success.

Reach for this when you hear…

[clinical research] “We ran thirty-two subgroup analyses and found one that was significant — that's not a finding, that's what chance looks like when you look often enough.”

[A/B testing] “The dashboard shows a green metric and everyone wants to ship, but we've been peeking at this experiment every day for two weeks and we never preregistered a stopping rule.”

[financial audit] “If we flag every anomaly that crosses a threshold after slicing by region, quarter, and account type, we'll bury the real findings in false positives — we need a family-level rule before we start slicing.”

When This Archetype Applies

Partial catalog groundingSome structural conditions are represented by existing abstractions, but no sufficient condition set is fully represented.

A system runs, inspects, or informally considers many tests, comparisons, metrics, groups, models, or stories, but reports the most appealing result as though it came from one pre-specified test. The more opportunities there are for a chance pattern, the easier it becomes to mistake noise for evidence.

What this problem means

When many tests are tried, chance has many opportunities to produce an impressive-looking pattern. If the failed, null, or unreported comparisons disappear from view, the remaining selected result looks more surprising than it really is. This is the structural source of p-hacking, metric shopping, subgroup fishing, data dredging, repeated interim looks, and overfitting to validation evidence.

The problem is not simply that people use a p-value, a threshold, or a dashboard. The problem is that the evidentiary context of the selected result is incomplete. A single result is being interpreted without the family of attempts that generated it. Multiple-Testing Discipline restores that missing context.

Show the applicability expression

Applicability expression5 distinct conditions

Many simultaneous testsandPost hoc result selectionandReused discovery datasetandIncentivized favorable comparisonsandResearcher degrees of freedom
Algebraic12345

groundedpartly groundedopen

5 conditions, all required.

5Required in every casenumbered 1–5

These hold no matter which pattern applies.

1

Many simultaneous tests · grounded

Many hypotheses, outcomes, subgroups, windows, thresholds, features, treatments, or metrics are tested.

2

Post hoc result selection · grounded

A reported result was selected after examining alternatives rather than pre-specified.

3

Reused discovery dataset · grounded

One dataset, dashboard, experiment, audit stream, or corpus supports repeated discovery attempts.

4

Incentivized favorable comparisons · open

Institutional incentives reward whichever comparison appears favorable or significant.

5

Researcher degrees of freedom · open

Analysts can keep slicing, filtering, modeling, or reframing until a preferred pattern appears.

Other requirements and context (2)

Why these sit outside the expression

Supporting contextit may accompany or help interpret the situation, but it is not a load-bearing condition in a sufficient diagnostic set.

Application gateit governs whether applying the archetype is appropriate or material, rather than defining the structural problem itself.

  • Supporting contextA finding will trigger expensive action, reputation change, policy change, clinical choice, enforcement, or strategic commitment.

  • Application gateExploratory search is valuable, but the organization needs a clear boundary between generating leads and confirming claims.

3 of 5 conditions grounded · 2 open.

Read the methodologyDownload the trigger-logic data

Mechanisms / Implementations

  • Bonferroni-Like Correction: Stiffens each test's significance bar in proportion to how many tests share the family, so that clearing it stays hard even after many simultaneous attempts.
  • False Discovery Rate Control: Ranks a whole family of results and draws the significance line to hold the expected share of false discoveries below a chosen rate, trading a little purity for far more power.
  • Preregistration: Timestamps the hypotheses, primary outcome, and analysis plan before the data exist, so what counts as confirmatory is fixed in advance rather than chosen after.
  • Holdout Validation: Seals a slice of the evidence away untouched during all discovery, then judges the selected finding once against that fresh partition.
  • Claim Registry: A living ledger of every attempted claim — its status, owner, and follow-up burden — so selective memory can't erase the failed tries that made a discovery look surprising.
  • Replication Study: Re-runs the finding from scratch in independent hands to see whether it survives outside the conditions and choices that first produced it.
  • Confirmatory Follow-Up: Turns one promising exploratory lead into a single pre-specified confirmatory test, so a pattern found by searching must earn its status on fresh ground.
  • Alpha-Spending Plan: Treats the total false-positive budget as a currency spent in pre-planned fractions across repeated interim looks, so peeking at accumulating data never inflates the error rate.
  • Metric Hierarchy: Ranks metrics into primary, secondary, and exploratory tiers before the data land, so a disappointing primary can't be quietly swapped for a flattering secondary.
  • Multiverse Analysis Report: Runs the analysis across every defensible analytic choice at once and shows the whole spread of results, exposing whether the headline depends on one lucky path.

Abstractions this archetype builds on — directly (a source ingredient) or as a related pattern. Links follow the typed catalog namespace.

Built directly on (3)

Also references 6 related abstractions

Variants

Narrower or domain-specific specializations that share this archetype's core structure. Recognized variants are established; candidate variants are provisional.

Familywise Error Control Variant · risk or failure variant · recognized

A stricter variant that tries to avoid even one false positive across a defined family of claims.

False Discovery Rate Variant · risk or failure variant · recognized

A screening-oriented variant that permits many discoveries while controlling the expected share that are false.

Exploratory–Confirmatory Partition Variant · temporal variant · promote to full archetype candidate

A staged variant that explicitly labels idea-generation evidence separately from confirmation evidence.

Claim Registry Variant · governance variant · recognized

A governance variant that controls multiplicity by making attempted claims, status, and follow-up obligations visible across a portfolio.

Editorial Notes

Problem Classification

Classification: Uncertainty, Evidence & Inference FailureExperimental Comparison & Hypothesis-Test Design

Problem kernel: multiple unplanned tests inflate chance findings

Rationale: Earliest causal condition: A system runs, inspects, or informally considers many tests, comparisons, metrics, groups, models, or stories, but reports the most appealing result as though it came from one pre-specified test. The more opportunities there are for a chance pattern, the easier it becomes to mistake noise for evidence.

Independent corroboration: The earliest necessary condition in the frozen evidence is: A system runs, inspects, or informally considers many tests, comparisons, metrics, groups, models, or stories, but reports the most appealing result as though it came from one pre-specified test. That is a experimental comparison and hypothesis test design problem because Treatment, control, assignment, blinding, power, and evidence thresholds are insufficiently designed to support the intended comparison.

Review outcome: Independent reviewer agreement; high confidence.