Replication Study¶
Independent confirmation — instantiates Multiple-Testing Discipline
Re-runs the finding from scratch in independent hands to see whether it survives outside the conditions and choices that first produced it.
A Replication Study confirms a finding by having someone else run the whole thing again — new sample, new team, often a new site — and checking whether the effect reappears. Where in-house confirmation reserves data or runs a targeted follow-up within the original program, replication deliberately changes hands: the point is to break any dependence on the original lab's particular participants, apparatus, analytic quirks, or unspoken conditions. The defining idea, unique among its siblings, is independence of execution: the finding must survive being reproduced by people who did not discover it and have no stake in its survival. That is the strongest form of confirmation because it tests not just the claim but the claim's portability out of the circumstances that first produced it.
Example¶
A cancer-biology finding reports that a particular compound shrinks tumors in a mouse model — a striking result that could steer years of drug development. Before anyone builds on it, an independent lab, with no involvement in the original work, sets out to reproduce it: they obtain the same compound and cell lines, follow the published protocol (requesting materials and clarifications from the original authors where the methods are ambiguous), run the experiment on a fresh cohort of animals, and pre-specify what result would count as a successful replication. Nothing about the original team's hands touches the new run. If the tumor-shrinkage effect reappears at a comparable magnitude, the claim's status moves to replicated and gains real weight; if it vanishes or shrinks toward nothing, that too is informative — the original may have been a false positive, or the effect may depend on some uncontrolled condition specific to the first lab. Either outcome is a data point about how much the finding can be trusted outside the room it was born in — the kind of check that the broader replication crisis has shown is indispensable.[1]
How it works¶
- Take a completed claim. Replication starts from a finding already discovered and reported, not from a lead being explored.
- Reconstruct independently. A team with no stake obtains the materials and protocol and rebuilds the study, collecting genuinely new data rather than re-analyzing the original.
- Pre-specify the success criterion. Define in advance what result counts as a successful replication — same direction, comparable magnitude, or significance on the new sample — so the verdict is not negotiated afterward.
- Read the outcome both ways. A reappearing effect promotes the claim's status; a failure to replicate is itself evidence about the original's reliability or its hidden dependence on local conditions.
Tuning parameters¶
- Fidelity — direct replication (copy the protocol as closely as possible) versus conceptual replication (test the same claim with different methods). Direct probes reliability; conceptual probes generalizability, and each answers a different question.
- Independence degree — same field/different lab, different discipline, or a large multi-site consortium; more independence is more convincing but harder to organize.
- Success criterion — how strict a match counts as "replicated"; a demanding criterion resists false confirmation but can fail on real effects that are merely smaller.
- Power — how large the replication sample is; an underpowered replication that "fails" may simply have been too small to see a real but modest effect.
When it helps, and when it misleads¶
Its strength is that it is the hardest test a claim can pass: by re-running in independent hands it strips away dependence on the original team's sample, tacit skill, and analytic choices, so a survived replication says the effect is real and portable. It is the mechanism of last resort precisely when stakes are high enough that in-house confirmation is not enough.
Its failure mode has two faces. A failed replication is ambiguous — it may mean the original was spurious, or that the replication was underpowered, botched, or run under quietly different conditions the original depended on; declaring a claim dead on one weak replication is as much an error as trusting the original. Conversely, replications can be contaminated when the replicating team is not truly independent or leans on the original authors so heavily that they reproduce the same mistake. And even a clean, successful replication does not make a finding important — it can be robustly real and still tiny or irrelevant. The guarding discipline is to power replications adequately, pre-specify the success criterion, keep the replicating team genuinely at arm's length, and weigh replications in aggregate rather than staking everything on one.
How it implements the components¶
independent_replication_path— commissioning and running the study in independent hands is this mechanism; the independent path is its whole substance.confirmation_requirement— surviving an independent re-run is the highest confirmation bar in the archetype, the one reserved for the most consequential claims.result_status_label— the outcome stamps the claim replicated or failed to replicate, a status that travels with it thereafter.
It confirms by fresh, independent execution and does not lean on data reserved within the original dataset — that holdout_evidence partition is Holdout Validation's tool, and unlike a holdout, a replication can catch a bias baked into the whole original source.
Related¶
- Instantiates: Multiple-Testing Discipline — supplies the independent-confirmation form of multiplicity control.
- Consumes: Claim Registry flags which claims carry a standing replication obligation and are due to be re-run.
- Sibling mechanisms: Alpha-Spending Plan · Bonferroni-Like Correction · False Discovery Rate Control · Claim Registry · Confirmatory Follow-Up · Holdout Validation · Metric Hierarchy · Multiverse Analysis Report · Preregistration
Editorial Notes¶
Form Classification¶
Form family: Experiment, Test & Rehearsal
Rationale: Replication Study operates as an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation because it re-runs the finding from scratch in independent hands to see whether it survives outside the conditions and choices that first produced it.
Independent corroboration: The frozen evidence defines Replication Study as 'Re-runs the finding from scratch in independent hands to see whether it survives outside the conditions and choices that first produced it', so its operative form is Experiment, Test & Rehearsal.
Review outcome: Independent reviewer agreement; high confidence.
Origin Attribution¶
Primary origin: Statistics & Experimental Design
Origin pattern: Single lineage
Present-day reach: Universal
Rationale: Independent repetition of a finding is a foundational experimental-design and scientific-inference practice.
Review resolution: Both blind reviewers agree that statistics_experimental_design is the primary historical origin. Explicit reconciliation of alternate origin disagreement, domain reach disagreement adopts reviewer_a's evidence: Independent repetition of a finding is a foundational experimental-design and scientific-inference practice. The selected record uses alternates=none, origin_mode=single_lineage, and domain_reach=universal; the other review proposed alternates=data_science, mathematics, origin_mode=single_lineage, and domain_reach=multi_domain. The selected combination better preserves the mechanism-specific formative lineages and calibrated scope; broader present-day use is not treated as proof of additional historical origin.
Review outcome: Reconciled after independent review; high confidence.
Notes¶
Replication and Confirmatory Follow-Up are easily confused because both re-test a claim on fresh evidence. The line is who runs it and why: a confirmatory follow-up is the original program's own next stage, aimed at one lead it just found; a replication is an independent party re-running an already-reported finding to test whether it holds outside the hands that produced it. Follow-up asks "does this lead survive our dedicated test?"; replication asks "does it survive when strangers rebuild it from scratch?"
References¶
[1] Open Science Collaboration. "Estimating the Reproducibility of Psychological Science". Science 349(6251): aac4716 (2015). Treats both positive and negative replication outcomes as cumulative evidence about reproducibility, illustrating why replication is a defining scientific check. registry ↩