Skip to content

Independent Replication Protocol

Protocol — instantiates Independent Evidence Triangulation

A standing procedure for obtaining a separately executed repeat of a test by a different team, with controlled information sharing and explicit comparability conditions.

A single result — however careful — could be a fluke of one lab, one operator, one machine, one Tuesday. The Independent Replication Protocol is the standing procedure for finding out, by having the same test re-executed by a separately resourced team under stated comparability conditions and controlled information sharing. Its defining trait is that independence comes from separate execution: the replicating team runs its own materials, its own instruments, and its own hands, so that a result surviving the repeat is evidence about the phenomenon rather than an artifact of the original setup. It holds the method deliberately constant and varies the executor — the opposite emphasis from designs that vary the method itself.

Example

A behavioral-science lab reports that a brief "belonging" writing exercise raises first-year students' persistence, and a university wants to know whether to fund a campus-wide rollout — a costly, hard-to-reverse decision resting on one study. Rather than take the paper at face value, the office invokes a replication protocol. A second team at a different campus is handed the pre-registered materials and analysis plan — enough to run the identical procedure — but deliberately not the original outcome data or the first team's interim impressions, so their execution cannot be steered toward the expected result.

The protocol fixes the comparability conditions in advance: the same eligibility criteria, the same primary outcome measured the same way, a minimum sample that clears the quality floor, and a rule that a "success" must land within a stated effect envelope, not merely point the same direction. The second team runs it and finds a real but noticeably smaller effect, concentrated in students who arrive academically underprepared. That is not a failed replication and not a clean one — it is a narrowed claim, and because the protocol carries an update trigger, the finding reopens the rollout question with a sharper target population than the original study ever specified.

How it works

  • Transfer the recipe, not the answer. The replicating team receives the procedure and analysis plan; original results and interim judgments are withheld so execution stays uncontaminated.
  • Fix comparability before running. Population, primary outcome, measurement, and the effect envelope that counts as agreement are stated in advance, so "it replicated" is not decided after the fact.
  • Enforce a quality floor on the repeat. The replication must itself clear minimum sample, provenance, and integrity bars, or it cannot overturn or confirm the original.
  • Trip an update trigger on the outcome — confirmation, failure, or narrowing each reopens the downstream decision rather than closing it silently.

Tuning parameters

  • Information-sharing tightness — from full protocol transparency to withholding even the hypothesis direction. Tighter cuts contamination but risks the replicators diverging on an unstated procedural detail.
  • Comparability envelope — how close the repeat's result must sit to count as a replication. Too tight and honest sampling noise reads as failure; too loose and any directional echo passes.
  • Number of replications — one confirmatory repeat versus several across sites. More separates a robust effect from a site-specific one, at rising cost.
  • Directness — a close ("direct") repeat of the exact procedure versus a deliberately varied ("conceptual") one; direct tests reproducibility, varied tests robustness.[1]

When it helps, and when it misleads

Its strength is that it is the cleanest test of whether a result is real as opposed to reported: a separately executed repeat catches procedural artifacts, lucky samples, and lab-specific quirks that no amount of re-reading the original can. It is the concrete face of the reproducibility norm.

Its failure mode is that replications carry their own confounds. A "failed" replication may reflect a botched repeat, a subtly different population, or an unstated moderator rather than a false original — which is why the comparability conditions and quality floor are load-bearing, not paperwork. There is also a quieter trap: a replication that shares the original's hidden dependence — same flawed instrument design, same biased sampling frame reused — can reproduce the original's error and be mistaken for confirmation. The guarding discipline is to make the comparability envelope explicit up front, hold the repeat to the same quality bar as the original, and check that the replication does not inherit the very dependency the triangulation is trying to break.

How it implements the components

  • independence_criterion — it operationalizes independence as separate execution: a different team, its own materials and instruments, with the original's results withheld during the run.
  • evidence_quality_floor — the protocol sets minimum sample, provenance, and integrity bars the replication must clear before it can confirm or overturn anything.
  • review_and_update_trigger — the replication outcome is a defined event that reopens the downstream decision, whether it confirms, fails, or narrows the claim.

It varies the executor, not the method — designing complementary *different methods for one claim is method_diversity_plan, owned by its protocol twin Multi-Method Study Design. It also does not map shared lineage (dependency_and_common_cause_map, Source Dependency Graph).*

Editorial Notes

Form Classification

Form family: Experiment, Test & Rehearsal

Rationale: A different team deliberately repeats a test under controlled information-sharing and comparability conditions to generate independent evidence.

Nearest alternative: Protocol, Workflow & Routine — Execution is governed by a standing procedure, but the defining form is experimental replication.

Review outcome: Adjudicated after independent review; high confidence.

Origin Attribution

Primary origin: Statistics & Experimental Design

Origin pattern: Single lineage

Present-day reach: Multi-domain

Rationale: A controlled protocol for separately executed repeat studies belongs to the experimental-design lineage of replication and reproducibility.

Review outcome: Independent reviewer agreement; high confidence.

Notes

The nearest twin is Multi-Method Study Design, and the one-sentence separation is: replication holds the method fixed and changes the team to catch execution artifacts, whereas multi-method holds the claim fixed and changes the method to catch method-specific bias.

References

[1] Schmidt, S. "Shall We Really Do It Again? The Powerful Concept of Replication Is Neglected in the Social Sciences". Review of General Psychology 13(2), 90–100 (2009). Distinguishes direct repetition of an experimental procedure from conceptual replication with different methods, assigning the former to verification and the latter to broader theoretical corroboration. registry