Skip to content

Multiverse Analysis Report

Analytic transparency report — instantiates Multiple-Testing Discipline

Runs the analysis across every defensible analytic choice at once and shows the whole spread of results, exposing whether the headline depends on one lucky path.

A Multiverse Analysis Report attacks a multiplicity that hides inside a single analysis: the many small, defensible choices — which outliers to exclude, which covariates to include, how to transform a variable, where to cut a group — each of which could tip a result. Any one path can be defended, but the analyst quietly picked the path that produced the cleanest headline. The report refuses that privilege: it computes the result under all reasonable combinations of those choices and presents the entire distribution of outcomes side by side. The defining idea, distinct from every sibling, is that its unit of multiplicity is the set of analytic specifications of one claim, and its output is a transparency artifact — a spread showing how robust (or fragile) the finding is across the "garden of forking paths" — not a corrected threshold, a status ledger, or a confirmatory test.

Example

A social-science team studies whether a stretch of unusually hot weather raised local aggression, measured through incident reports. Along the way they face a thicket of defensible choices: exclude the top 1% of days as heat-wave outliers or keep them; control for weekday, for holidays, for both, or neither; define "hot" by absolute temperature or by deviation from seasonal norm; measure aggression by raw counts or per-capita rates. Each combination is a legitimate analysis, and reporting only the one that yields a crisp, significant effect would be researcher degrees of freedom at work.[n1] Instead the team builds a multiverse report: they enumerate every choice, generate the full grid of specifications — say a few hundred analyses — and plot the effect estimate and its significance across all of them. The picture tells the real story: if the effect is positive and significant in 90% of specifications, it is robust; if it is significant in only the handful of paths the authors happened to choose, the headline was an artifact of the fork they took. Either way, readers see the whole space rather than a single curated survivor.

How it works

  • Enumerate the choices. List every defensible analytic decision point and its reasonable options; the cross-product of these is the specification family.
  • Run the whole grid. Compute the result under each combination, producing an estimate (and p-value) for every specification in the space.
  • Show the distribution. Present the spread — a specification curve or histogram of effects — with the fraction of paths that reach significance and in which direction, so no single path is privileged.
  • Keep every path visible. The favorable and unfavorable specifications are reported together; suppressing the ones that "didn't work" would recreate the exact selective-reporting problem the report exists to expose.

Tuning parameters

  • Specification breadth — how many choice points and options enter the multiverse; a wider space is more honest but explodes combinatorially and can include indefensible paths that muddy the signal.
  • Inclusion criterion — how "defensible" a path must be to earn a place in the grid; loose inclusion is comprehensive but dilutes, strict inclusion is cleaner but invites the same curation it fights.
  • Summary statistic — whether robustness is reported as the share of significant paths, the median effect, or a formal specification-curve test; each frames fragility differently.
  • Display granularity — full curve versus summarized bands; more detail is more transparent but harder for a decision-maker to read.

When it helps, and when it misleads

Its strength is honesty about fragility: it converts an invisible degree of freedom into a visible distribution, so a reader can see at a glance whether a claim is a robust feature of the data or an artifact of one analytic fork. It is the natural remedy when the risk is not many tests but many ways to run one test.

Its failure mode is that a multiverse describes robustness but does not, by itself, adjudicate truth: a finding significant in 95% of specifications may still be confounded or trivially small, and breadth can be gamed — padding the grid with weak paths to dilute an inconvenient result, or omitting the one defensible choice that would overturn it. Presenting a reassuring curve as if it settled the question is the classic misuse. The guarding discipline is to fix the set of defensible specifications before seeing which ones are favorable, and to treat a robust multiverse as strong support that still owes independent confirmation, not as proof.

How it implements the components

  • multiplicity_inventory — the enumerated grid of analytic specifications is a literal inventory of the attempts hidden inside a single analysis.
  • discovery_record — by reporting every specification, favorable and unfavorable, it keeps the failed and null analytic paths visible instead of letting selective memory keep only the winner.
  • claim_family — it defines the family as the set of analytic paths for one claim, making the interpretive context of the reported effect explicit.

It maps how a claim behaves across analytic paths but does not assign the claim an evolving owned status or a follow-up bar over time — tracking result_status_label and confirmation_requirement across a program is Claim Registry's work; the multiverse report is a one-shot spread of a single claim.

Editorial Notes

Form Classification

Form family: Analysis, Modeling & Optimization

Rationale: Multiverse Analysis Report operates as a computation, comparison, model, or analytic representation used to infer, estimate, or choose because it runs the analysis across every defensible analytic choice at once and shows the whole spread of results, exposing whether the headline depends on one lucky path.

Independent corroboration: The frozen evidence defines Multiverse Analysis Report as 'Runs the analysis across every defensible analytic choice at once and shows the whole spread of results, exposing whether the headline depends on one lucky path', so its operative form is Analysis, Modeling & Optimization.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Statistics & Experimental Design

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Running all defensible analytic specifications and displaying the result distribution arose in statistics and metascience as multiverse or specification-curve analysis.

Related originating lineages:

Review resolution: Both independent reviews agree on primary origin statistics_experimental_design; reconciliation resolves secondary fields (alternate_origin_disagreement, origin_mode_disagreement). Alternate origins retained (data_science, tech_ethics_ai_governance) are the union of reviewer-supported formative lineages with explicit rationales, not a list of later application domains. Present-day breadth is represented separately as domain_reach=multi_domain; origin_mode=cross_disciplinary_synthesis records the historical relationship among lineages. Confidence is conservatively reconciled to high, and encyclopedia_synthesis=false preserves either reviewer's finding that the encyclopedia generalized the mechanism.

Review outcome: Reconciled after independent review; high confidence.

Notes

[n1] Researcher degrees of freedom — the many defensible data-processing and analysis choices available in a single study — create a "garden of forking paths" in which even honest analysts can arrive at a significant result without any deliberate p-hacking; multiverse and specification-curve analyses expose this by running all the paths at once.