Statistical Conclusion Validity¶
The warrantedness of inferences about whether and how strongly measured variables covary, given power, error control, model assumptions, measurement reliability, and analysis conduct.
Core Idea¶
Statistical conclusion validity is the degree to which evidence warrants an inference about whether two measured variables covary and how strongly they covary. In the Campbell validity framework, it addresses two linked questions: is there a relationship between the presumed treatment and outcome, and what is its magnitude? Shadish, Cook, and Campbell distinguish this from internal validity, which asks whether observed covariation reflects a causal relationship, and from construct and external validity.
The abstraction is not synonymous with “a significant \(p\)-value.” It audits the entire inferential chain: design sensitivity, sample size, measurement reliability, treatment implementation, range of observations, statistical assumptions, multiplicity, effect-size estimation, and heterogeneity.
Scope of Application¶
The framework was developed for experimental and quasi-experimental research but applies whenever a study makes statistical claims about association, difference, or covariation. Randomized trials, observational studies, interrupted time series, single-case designs, surveys, laboratory experiments, and model-based comparisons can all face conclusion-validity threats.
Shadish, Cook, and Campbell inventory nine recurring threats: low statistical power; violated assumptions of statistical tests; fishing and the error-rate problem; unreliability of measures; restriction of range; unreliability of treatment implementation; extraneous variance in the experimental setting; heterogeneity of units; and inaccurate effect-size estimation. These are a diagnostic catalog, not a claim that every threat is present in every study.
Clarity¶
Suppose a two-group experiment estimates a mean difference
with standard error \(\operatorname{SE}(\widehat\Delta)\). A test statistic may be
The validity of the conclusion does not follow from \(|t|>1.96\) alone. The standard error may be wrong if clustering is ignored; multiple outcomes may inflate the familywise false-positive rate; unreliable outcomes may attenuate effects and reduce power; attrition may change the analyzed units; and a flexible stopping rule may alter the reference distribution.
Manages Complexity¶
Statistical conclusions sit at the end of a long chain. The framework decomposes failure into interpretable threat mechanisms and directs remedies: increase precision or sample size for low power; use cluster-aware errors for dependence; correct or model multiplicity for many tests; improve measurement for unreliability; predefine analyses to constrain fishing; and report compatible effect sizes rather than significance alone.
Abstract Reasoning¶
Conclusion validity asks whether the sampling distribution used for an inference adequately represents uncertainty in the actual design and analysis. Formally, if \(\widehat\psi\) estimates a covariation parameter \(\psi\), then validity depends on bias, variance, coverage, calibration, and the selection mechanism:
Knowledge Transfer¶
The threat-audit pattern transfers from experiments to observational and computational studies by relabeling treatment and outcome as the variables whose covariation is inferred. Clustered logs, repeated measurements, cross-validation folds, and spatial fields all require dependence-aware uncertainty.
Transfer must preserve the boundary with causality. Better standard errors strengthen the association inference; they do not remove confounding. Likewise, increasing reliability strengthens measurement precision but does not prove that the measure captures the intended construct. The framework is valuable partly because it prevents these validities from substituting for one another.
Relationships to Other Abstractions¶
Current abstraction Statistical Conclusion Validity Domain-specific
Parents (1) — more general patterns this builds on
-
Statistical Conclusion Validity presupposes Statistical Inference Prime
Statistical Conclusion Validity compositionally presupposes Statistical Inference: it evaluates whether a sample-to-conclusion procedure correctly quantifies evidence and uncertainty about covariation.
Hierarchy paths (4) — routes to 4 parentless roots
- Statistical Conclusion Validity → Statistical Inference → Inductive Reasoning
- Statistical Conclusion Validity → Statistical Inference → Uncertainty
- Statistical Conclusion Validity → Statistical Inference → Probability → Measure → Set and Membership
- Statistical Conclusion Validity → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
Neighborhood in Abstraction Space¶
Statistical Conclusion Validity sits in a sparse region of the domain-specific corpus (84th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Unclustered & Miscellaneous (1565 abstractions)
Nearest neighbors
- External Validity — 0.81
- Propensity score matching — 0.81
- Differential effects — 0.81
- Suppressor variable — 0.80
- Standard error — 0.80
Computed from structural-signature embeddings · 2026-09-08