Multiple Testing Discipline¶
Control false discoveries when many comparisons, claims, or tests are being tried.
Essence¶
Multiple-Testing Discipline is the intervention pattern for situations where many comparisons, claims, metrics, screens, or analytic choices are tried and the most attractive result is at risk of being interpreted as though it came from a single planned test. The archetype does not prohibit exploration. It makes exploration honest: the selected result must be understood against the larger search space that made it possible.
The core move is to change the unit of credibility. Instead of asking only, "Does this one result look unlikely by chance?" the pattern asks, "How many chances did the system have to find something that looked unlikely, and what confirmation burden follows from that?" That shift turns a potentially misleading discovery ritual into a disciplined evidence process.
Compression statement¶
When many comparisons are searched, separate the family of attempted claims from any single attractive result, adjust the evidentiary standard or discovery procedure, record exploratory search, and require confirmation before treating selected findings as reliable.
Canonical formula: many attempted claims + ordinary single-claim interpretation -> inflated false positives; claim family + adjusted threshold/procedure + exploration log + confirmation -> credible discovery
When This Archetype Applies¶
Partial catalog groundingSome structural conditions are represented by existing abstractions, but no sufficient condition set is fully represented.
Diagnostic problem
A system runs, inspects, or informally considers many tests, comparisons, metrics, groups, models, or stories, but reports the most appealing result as though it came from one pre-specified test. The more opportunities there are for a chance pattern, the easier it becomes to mistake noise for evidence.
What this problem means
When many tests are tried, chance has many opportunities to produce an impressive-looking pattern. If the failed, null, or unreported comparisons disappear from view, the remaining selected result looks more surprising than it really is. This is the structural source of p-hacking, metric shopping, subgroup fishing, data dredging, repeated interim looks, and overfitting to validation evidence.
The problem is not simply that people use a p-value, a threshold, or a dashboard. The problem is that the evidentiary context of the selected result is incomplete. A single result is being interpreted without the family of attempts that generated it. Multiple-Testing Discipline restores that missing context.
Applicability expression5 distinct conditions
groundedpartly groundedopen
5 conditions, all required.
5Required in every casenumbered 1–5
These hold no matter which pattern applies.
Many simultaneous tests · grounded
Many hypotheses, outcomes, subgroups, windows, thresholds, features, treatments, or metrics are tested.
The source archetype describes the situation as follows: Many hypotheses, outcomes, subgroups, time windows, thresholds, features, treatments, or metrics are being tested or inspected. The normalized requirement above isolates the load-bearing portion used in this condition set.
primeMultiple Comparisons Correction— Adjust the thresholds or p-values of a defined family of simultaneous tests so a chosen family-level error criterion remains bounded despite multiplicity.
Post hoc result selection · grounded
A reported result was selected after examining alternatives rather than pre-specified.
The source archetype describes the situation as follows: A result was selected after looking across several alternatives rather than being specified before the evidence was seen. The normalized requirement above isolates the load-bearing portion used in this condition set.
domainHARKing (Hypothesizing After the Results are Known)— The research practice of building a hypothesis by inspecting already-collected data and then presenting it as if it had been specified in advance, silently inflating the reported false-positive rate because the test's independence assumption is violated.
How this was matched — 4 requirements, all needed
result is selected post hoc from examined alternatives without pre-specification
All of
- roleAn analyst has evidence, several alternative results, and one result that is reported.
- quantifierSeveral alternatives are examined before the reported result is chosen.
- timingSelection of the reported result occurs after the alternatives and evidence have been examined.
- polarityThe reported result was not specified before the evidence was seen.
Reused discovery dataset · grounded
One dataset, dashboard, experiment, audit stream, or corpus supports repeated discovery attempts.
The source archetype describes the situation as follows: The same dataset, dashboard, experiment, audit stream, or corpus is reused for many discovery attempts. The normalized requirement above isolates the load-bearing portion used in this condition set.
domainHARKing (Hypothesizing After the Results are Known)— The research practice of building a hypothesis by inspecting already-collected data and then presenting it as if it had been specified in advance, silently inflating the reported false-positive rate because the test's independence assumption is violated.
context guardThe post-hoc inspection examines many alternative patterns in the same dataset.
suppliesThe same evidence source is used repeatedly for discovery attempts. · The reuse supports many discovery attempts.
How this was matched — 3 requirements, all needed
one evidence source supports repeated discovery attempts
All of
- roleOne identifiable evidence source is available for exploratory discovery.
- relationThe same evidence source is used repeatedly for discovery attempts.
- quantifierThe reuse supports many discovery attempts.
Incentivized favorable comparisons · open
Institutional incentives reward whichever comparison appears favorable or significant.
The source archetype describes the situation as follows: Teams have incentives to publish, ship, escalate, or celebrate whichever comparison looks significant or favorable. The normalized requirement above isolates the load-bearing portion used in this condition set.
Researcher degrees of freedom · open
Analysts can keep slicing, filtering, modeling, or reframing until a preferred pattern appears.
The source archetype describes the situation as follows: Different analysts or stakeholders can keep slicing, filtering, modeling, or reframing until a preferred pattern appears. The normalized requirement above isolates the load-bearing portion used in this condition set.
Other requirements and context (2)
Why these sit outside the expression
Supporting context — it may accompany or help interpret the situation, but it is not a load-bearing condition in a sufficient diagnostic set.
Application gate — it governs whether applying the archetype is appropriate or material, rather than defining the structural problem itself.
Supporting contextA finding will trigger expensive action, reputation change, policy change, clinical choice, enforcement, or strategic commitment.
It is especially important when selected findings will guide decisions, reputations, resource allocation, safety action, publication, enforcement, or product launch. In this archetype, the relevant contextual consideration is: A finding will trigger expensive action, reputation change, policy change, clinical choice, enforcement, or strategic commitment. It helps interpret the situation or strengthens the practical case for examining the archetype.
Application gateExploratory search is valuable, but the organization needs a clear boundary between generating leads and confirming claims.
It becomes necessary when exploratory results are about to be presented as confirmed claims or used for consequential action. In this archetype, the relevant application gate is: Exploratory search is valuable, but the organization needs a clear boundary between generating leads and confirming claims. It narrows when choosing or applying the archetype is warranted or decision-relevant.
Coverage
3 of 5 conditions grounded · 2 open.
When to Use This Archetype¶
Use this archetype when the same evidence base can generate many candidate findings: many outcomes in a study, many metrics in an experiment, many subgroups in a report, many model variants in a benchmark, many anomaly screens in an audit, or many investigative leads in a case file. It is especially important when selected findings will guide decisions, reputations, resource allocation, safety action, publication, enforcement, or product launch.
The archetype is not needed for every open-ended exploration. It becomes necessary when exploratory results are about to be presented as confirmed claims or used for consequential action. In low-stakes exploration, the right move may be simple labeling: "this is a lead." In high-stakes contexts, the right move may require formal correction, holdout validation, independent replication, a registry of attempted claims, or a staged confirmation process.
Structural Problem¶
When many tests are tried, chance has many opportunities to produce an impressive-looking pattern. If the failed, null, or unreported comparisons disappear from view, the remaining selected result looks more surprising than it really is. This is the structural source of p-hacking, metric shopping, subgroup fishing, data dredging, repeated interim looks, and overfitting to validation evidence.
The problem is not simply that people use a p-value, a threshold, or a dashboard. The problem is that the evidentiary context of the selected result is incomplete. A single result is being interpreted without the family of attempts that generated it. Multiple-Testing Discipline restores that missing context.
Intervention Logic¶
The intervention begins by defining the claim family: the related tests, outcomes, metrics, models, filters, time windows, and segments that created opportunities for a selected discovery. It then records the multiplicity inventory, including failed or unselected attempts. Next, it chooses an error-risk policy: strict false-positive avoidance, discovery-rate control, exploratory labeling, or staged confirmation. Finally, it applies a rule or process that matches the search space, and it withholds decision-ready status until the finding has survived the appropriate confirmation burden.
A disciplined process can still be creative. Exploration generates leads; confirmation changes their status. The archetype succeeds when a stakeholder can see both the promising result and the path by which it was found, then judge whether the result is exploratory, adjusted, confirmed, replicated, or ready for action.
Key Components¶
Multiple-Testing Discipline restores the missing evidentiary context around a selected finding by making the broader search space visible and matching the confirmation burden to it. The Claim Family defines which comparisons belong together for interpretation — a single product experiment may contain one primary metric, several secondary metrics, dozens of segments, and multiple time windows, and the family makes those opportunities legible. The Multiplicity Inventory records the actual number and kind of attempted looks, including tested metrics, model variants, filters, thresholds, subgroups, interim checks, and abandoned analyses, so hidden flexibility cannot vanish from the evidence trail. The Error-Risk Profile then states what kind of mistake matters most for the context — a safety screen, a scientific discovery program, and a product experiment need different balances between false positives and missed discoveries — and chooses the policy by which the family will be governed.
The remaining components translate that policy into discipline at the moment of decision. The Multiplicity Adjustment Rule implements the chosen approach: a formal threshold correction, a false-discovery-rate procedure, a holdout requirement, an alpha-spending plan, or a staged confirmation gate. The Exploratory–Confirmatory Boundary prevents patterns found during search from being silently relabeled as planned tests, preserving exploration's value while protecting confirmatory credibility. The Discovery Record keeps failed, null, alternative, and selected analyses visible together so selective memory does not recreate the same inflation risk. Finally, the Confirmation Requirement specifies what must happen before a promising lead can guide consequential action — a new experiment, held-out data, independent replication, a second site, or a stricter review gate — so credibility scales with the stakes of the decision the finding is asked to support.
| Component | Description |
|---|---|
| Claim Family ↗ | Defines which comparisons belong together for interpretation. A single product experiment might include one primary metric, several secondary metrics, dozens of segments, and multiple time windows; the claim family makes those opportunities visible. |
| Multiplicity Inventory ↗ | Records the number and kind of attempted looks. This includes tested metrics, model variants, filters, thresholds, subgroups, interim checks, and abandoned analyses. |
| Error-Risk Profile ↗ | States what kind of mistake matters most. A safety screen, scientific discovery program, and product experiment may need different balances between false positives and missed discoveries. |
| Multiplicity Adjustment Rule ↗ | Implements the chosen discipline. It might be a formal correction, a false-discovery-rate procedure, a holdout rule, an alpha-spending plan, or a staged confirmation requirement. |
| Exploratory–Confirmatory Boundary ↗ | Prevents patterns found during search from being treated as planned tests. It preserves exploratory value while protecting confirmatory credibility. |
| Discovery Record ↗ | Keeps failed, null, alternative, and selected analyses visible. Without it, selective memory recreates the same false-discovery risk. |
| Confirmation Requirement ↗ | Defines what must happen before a selected result can guide high-stakes action. Confirmation may require a new experiment, held-out data, independent replication, a second site, or a stricter review gate. |
Common Mechanisms¶
Mechanisms implement the archetype; they are not the archetype itself.
- Bonferroni-Like Correction (
bonferroni_like_correction): A strict threshold-adjustment mechanism for contexts where even one false positive across the family is costly. - False Discovery Rate Control (
false_discovery_rate_control): A screening mechanism that permits many discoveries while controlling the expected share that are false. - Preregistration (
preregistration): A procedural commitment mechanism that records planned claims and analyses before results are observed. - Holdout Validation (
holdout_validation): An evidence-partitioning mechanism that reserves fresh data or cases for confirmation after discovery. - Claim Registry (
claim_registry): A governance mechanism that tracks attempted claims, statuses, owners, and follow-up requirements. - Replication Study (
replication_study): An independent-confirmation mechanism that tests whether a selected finding persists in new evidence. - Confirmatory Follow-Up (
confirmatory_follow_up): A staged-validation mechanism that turns a promising lead into a targeted confirmatory test. - Alpha-Spending Plan (
alpha_spending_plan): A sequential-testing mechanism for repeated interim looks. - Metric Hierarchy (
metric_hierarchy): A priority mechanism that prevents teams from promoting a secondary metric after the primary result disappoints. - Multiverse Analysis Report (
multiverse_analysis_report): A transparency mechanism that shows whether a result depends on one favorable analytic path.
10 documented mechanisms across 4 implementation forms.
The grouping reflects forms represented among the mechanisms currently documented for this archetype; an absent form is not necessarily an impossible implementation.
Analysis, Modeling & Optimization · 3 mechanisms
- Bonferroni-Like Correction — Stiffens each test's significance bar in proportion to how many tests share the family, so that clearing it stays hard even after many simultaneous attempts.
- False Discovery Rate Control — Ranks a whole family of results and draws the significance line to hold the expected share of false discoveries below a chosen rate, trading a little purity for far more power.
- Multiverse Analysis Report — Runs the analysis across every defensible analytic choice at once and shows the whole spread of results, exposing whether the headline depends on one lucky path.
Experiment, Test & Rehearsal · 3 mechanisms
- Confirmatory Follow-Up — Turns one promising exploratory lead into a single pre-specified confirmatory test, so a pattern found by searching must earn its status on fresh ground.
- Holdout Validation — Seals a slice of the evidence away untouched during all discovery, then judges the selected finding once against that fresh partition.
- Replication Study — Re-runs the finding from scratch in independent hands to see whether it survives outside the conditions and choices that first produced it.
Record, Log & Register · 1 mechanism
- Claim Registry — A living ledger of every attempted claim — its status, owner, and follow-up burden — so selective memory can't erase the failed tries that made a discovery look surprising.
Rule, Policy & Commitment · 3 mechanisms
- Alpha-Spending Plan — Treats the total false-positive budget as a currency spent in pre-planned fractions across repeated interim looks, so peeking at accumulating data never inflates the error rate.
- Metric Hierarchy — Ranks metrics into primary, secondary, and exploratory tiers before the data land, so a disappointing primary can't be quietly swapped for a flattering secondary.
- Preregistration — Timestamps the hypotheses, primary outcome, and analysis plan before the data exist, so what counts as confirmatory is fixed in advance rather than chosen after.
Parameter / Tuning Dimensions¶
Important tuning dimensions include the size of the claim family, the dependence among tests, the cost of false positives, the cost of false negatives, the strength of prior theory, the number of exploratory degrees of freedom, the availability of held-out evidence, the stakes of action, and the acceptable delay before confirmation.
A strict familywise approach is appropriate when any false claim is dangerous. A false-discovery-rate approach is often better when screening many candidates and expecting later follow-up. A governance-heavy approach is useful when the main risk is organizational cherry-picking rather than a single statistical formula. A staged exploration-confirmation approach is useful when broad search is necessary but action must wait.
Invariants to Preserve¶
The claim family must remain visible when interpreting a selected result. Exploration and confirmation must not be merged into a single status. Failed or unreported attempts must not disappear from the evidence context. The evidentiary burden should increase as the search space expands. High-stakes action should require stronger confirmation than initial discovery. Mechanisms should remain mechanisms: a p-value, Bonferroni correction, preregistration form, or holdout dataset is not the archetype by itself.
Target Outcomes¶
The archetype aims to reduce false discoveries, reduce p-hacking and metric shopping, improve trust in selected findings, preserve the value of exploratory search, and make reported claims more reproducible. A successful implementation does not eliminate uncertainty; it makes the credibility of discoveries better calibrated to how they were found.
Tradeoffs¶
Multiple-Testing Discipline trades some speed and sensitivity for credibility. Stricter correction can miss real signals. Broad exploration can generate useful leads but weak confirmation. Simple rules are easy to explain but may be too conservative or poorly matched to correlated tests. Claim registries improve memory but add administrative burden. Confirmation delays action, but premature action can be far more costly when the discovery is false.
The best version of the archetype is not always the strictest. It is the version that fits the error-risk profile and clearly labels what status each finding deserves.
Failure Modes¶
A common failure is undefined claim family, where only the reported tests are corrected and hidden analytic flexibility is ignored. Another is correction theater, where a formal adjustment is applied while metric shopping, data leakage, confounding, or selective reporting continues. Overcorrection paralysis happens when strict rules suppress useful discovery even when false leads could be cheaply followed up. Exploratory label laundering occurs when a claim is labeled exploratory in methods text but presented as confirmed in decisions or headlines. Confirmation contamination happens when holdout or replication evidence is influenced by the original search.
The most subtle failure is treating a result that survived multiplicity discipline as automatically important. A disciplined discovery can still be tiny, biased, confounded, practically irrelevant, or ethically unsafe to act on.
Neighbor Distinctions¶
Hypothesis Testing Frame structures a single claim against a default with evidence thresholds and error costs. Multiple-Testing Discipline adds the many-claim layer: what counts as evidence changes when many opportunities for a false alarm exist.
Reproducibility Protocol makes work rerunnable and auditable. A reproducible workflow can still overclaim a cherry-picked result if the discovery process had many unacknowledged attempts.
Uncertainty Explicitness communicates uncertainty. Multiple-Testing Discipline changes the discovery and confirmation process so reported uncertainty is not falsely narrow because of hidden search.
Confounder Control addresses third-variable distortion. Multiple-Testing Discipline addresses false discoveries created by repeated or flexible search.
Power-Aware Design and Effect Size Reporting remain merge-review neighbors in this batch. They address false negatives and practical magnitude, while this archetype addresses inflated false positives across many claims.
Cross-Domain Examples¶
In genomics, thousands of candidate genes may be screened; the archetype uses discovery-rate control and replication before treating candidates as credible. In product analytics, a feature may be examined across many metrics and segments; the archetype keeps primary metrics separate from exploratory subgroup leads. In machine learning, many hyperparameters may be tried; the archetype protects a locked test set for final confirmation. In public policy evaluation, many regions, outcomes, and demographic subgroups may be inspected; the archetype reports the family and labels subgroup findings appropriately. In safety monitoring, many anomaly screens may create false alarms; the archetype tracks alert families and requires corroboration before escalation.
Non-Examples¶
A single pre-specified pass/fail test is better handled by Hypothesis Testing Frame. A brainstorming session that produces speculative ideas without evidence claims does not require this archetype. A causal comparison distorted by a third variable calls for Confounder Control. A report that lacks uncertainty intervals calls for uncertainty representation. A statistically significant but practically tiny result calls for effect magnitude and decision-relevance checks, not primarily multiplicity discipline.
Related Abstractions¶
Abstractions this archetype builds on — directly (a source ingredient) or as a related pattern. Links follow the typed catalog namespace.
Built directly on (3)
- Multiple Comparisons Correction: Adjust the thresholds or p-values of a defined family of simultaneous tests so a chosen family-level error criterion remains bounded despite multiplicity.
- Reproducibility & Replicability: Repeatable results.
- Type I & Type II Errors: False positive/negative.
Also references 6 related abstractions
- Confirmation Bias: Favor confirming evidence.
- Hypothesis Testing (Null vs. Alternative): Null vs alternative evaluation.
- Probability: Quantifies uncertainty and likelihoods.
- Statistical Significance (p-Value): Likelihood results are random.
- Threshold: Safe vs harmful levels.
- Uncertainty: Incomplete knowledge.
Variants¶
Narrower or domain-specific specializations that share this archetype's core structure. Recognized variants are established; candidate variants are provisional.
Familywise Error Control Variant · risk or failure variant · recognized
A stricter variant that tries to avoid even one false positive across a defined family of claims.
- Distinct from parent: The parent includes several ways to discipline many-claim discovery; this variant emphasizes strict family-level false-positive prevention.
- Use when: Any false claim in the family could cause serious harm, liability, wasted resources, or irreversible action; The number of tests is moderate enough that strict control is still usable; The environment values high specificity over broad exploratory discovery.
- Typical domains: clinical safety, regulatory testing, quality acceptance, high stakes audit
- Common mechanisms: bonferroni like correction, closed testing procedure
False Discovery Rate Variant · risk or failure variant · recognized
A screening-oriented variant that permits many discoveries while controlling the expected share that are false.
- Distinct from parent: The parent is broader; this variant accepts a calibrated false-discovery burden to preserve useful discovery.
- Use when: Large numbers of candidate signals are screened and some false leads are tolerable; The goal is to prioritize follow-up candidates, not to make final irreversible claims from the first screen; Discovery value is high enough that overly strict familywise control would be counterproductive.
- Typical domains: genomics, screening programs, anomaly detection, lead generation
- Common mechanisms: false discovery rate control, ranked candidate follow up
Exploratory–Confirmatory Partition Variant · temporal variant · promote to full archetype candidate
A staged variant that explicitly labels idea-generation evidence separately from confirmation evidence.
- Distinct from parent: The parent controls many-claim false-discovery risk; this variant may become broader as a staged epistemic-status pattern for discovery work generally.
- Use when: Broad exploration is necessary but later claims must be credible; Discovery and confirmation can be separated by time, dataset, team, site, or protocol; Stakeholders frequently overstate exploratory patterns.
- Typical domains: scientific research, product analytics, machine learning, investigative analysis
- Common mechanisms: preregistration, holdout validation, replication study
Claim Registry Variant · governance variant · recognized
A governance variant that controls multiplicity by making attempted claims, status, and follow-up obligations visible across a portfolio.
- Distinct from parent: The parent includes statistical and procedural strategies; this variant centers on portfolio-level claim memory and accountability.
- Use when: Many teams, analysts, experiments, dashboards, alerts, or investigations generate claims over time; The main risk is not a single formula but organizational forgetting, selective reporting, or repeated re-analysis; Reviewers need to see what has been tried, what failed, and what remains exploratory.
- Typical domains: regulatory review, analytics governance, audit portfolios, research program management
- Common mechanisms: claim registry, review gate
Near names: Multiplicity Control, Multiple Comparisons Correction, False Discovery Control, Data-Dredging Guardrail, p-Hacking Guardrail, Look-Elsewhere Effect Control, Trials Factor Correction.
Editorial Notes¶
Problem Classification¶
Classification: Uncertainty, Evidence & Inference Failure → Experimental Comparison & Hypothesis-Test Design
Problem kernel: multiple unplanned tests inflate chance findings
Rationale: Earliest causal condition: A system runs, inspects, or informally considers many tests, comparisons, metrics, groups, models, or stories, but reports the most appealing result as though it came from one pre-specified test. The more opportunities there are for a chance pattern, the easier it becomes to mistake noise for evidence.
Independent corroboration: The earliest necessary condition in the frozen evidence is: A system runs, inspects, or informally considers many tests, comparisons, metrics, groups, models, or stories, but reports the most appealing result as though it came from one pre-specified test. That is a experimental comparison and hypothesis test design problem because Treatment, control, assignment, blinding, power, and evidence thresholds are insufficiently designed to support the intended comparison.
Review outcome: Independent reviewer agreement; high confidence.