Misuse of p-values¶
Inferential errors that treat a p-value as evidence about hypothesis probability, causation, effect magnitude, practical importance, replicability, or categorical truth beyond its model-conditional tail-probability meaning.
Core Idea¶
A p-value is calculated in the direction from a specified model to possible data: assuming the model and analysis procedure, how extreme is the observed statistic or something more extreme? Common misuse reverses that condition and speaks as though the number were the probability the hypothesis is true given the data.
Other errors arise even when the definition is recited correctly. A conventional cutoff can be reified as a discovery boundary; selection and multiplicity can be hidden; a low value can be mistaken for a large, causal, or replicable effect; and failure to reject can be called proof of no effect. The defect lies in the relation between calculation and claim.
Structural Signature¶
Sig role-phrases:
- Specified model — Defines the null assumptions and stochastic reference under which the tail probability is computed. It is carrier. Counterfactual: A bare number without its model has no stable interpretation.
- Test statistic and extremeness rule — Determine which possible results count at least as incompatible as the observation. It is measurement. Counterfactual: Changing one- versus two-sided or selection rules changes the p-value.
- Observed data and design — Provide the realized statistic under sampling, assignment, measurement, and preprocessing choices. It is evidence. Counterfactual: A biased design is not repaired by a small p-value.
- Inferential claim — Is the conclusion the analyst draws from the conditional calculation. It is output. Counterfactual: The misuse occurs when this claim exceeds what the p-value licenses.
- Decision threshold — Optionally turns a continuum into a rule under predeclared error control. It is policy. Counterfactual: Treating 0.049 and 0.051 as opposite scientific truths creates a false cliff.
- Multiplicity and selection context — Accounts for the family of analyses, stopping, and reporting choices that shape the observed value. It is validity. Counterfactual: A nominal p-value after undisclosed selection is not calibrated as presented.
What It Is Not¶
- It is not the claim that p-values are never useful.
- It is not merely using the conventional 0.05 level.
- It is not evidence that a null hypothesis is probably true when p is large.
- It is not corrected by replacing one threshold with another while preserving the same overclaim.
- Closest near-miss. Dichotomous thinking treats results on opposite sides of a cutoff as qualitatively different; p-value misuse is broader and also includes inverse-probability, causal, magnitude, and replicability errors.
Scope of Application¶
- Scientific reporting. Audits whether prose claims match the design and model-conditional statistic.
- Multiple testing. Adjusts or contextualizes evidence across analysis families and selection.
- Policy and quality control. Separates a predeclared action threshold from a claim of scientific truth.
- Evidence synthesis. Combines effect estimates, intervals, prior knowledge, and replication rather than tallying significance.
- Statistical education. Teaches conditional direction and nearest misconceptions through explicit examples.
Clarity¶
Quote the exact claim, then write the probability conditioning represented by the p-value. State null model, statistic, sidedness, sample design, stopping and selection, multiplicity, effect estimate, interval, and substantive scale. A mismatch between these ingredients identifies the specific misuse instead of applying a generic accusation.
Manages Complexity¶
The abstraction turns a diffuse list of bad practices into auditable mappings from evidence to claim. Separating model, design, calculation, selection, decision, and substantive interpretation shows whether correction requires better computation, fuller reporting, a different estimand, or a narrower conclusion.
Abstract Reasoning¶
- Reconstruct the scientific question and effect quantity of interest.
- Write the exact null model and probability statement used to compute the p-value.
- Inspect design, stopping, preprocessing, multiplicity, and reporting selection.
- Compare the published conclusion with the limited implications of the calculation.
- Add effect size, uncertainty, sensitivity, and domain-relevant importance.
- Distinguish a decision rule from a graded evidential or causal statement.
Knowledge Transfer¶
The transferable cargo is a conditional-direction and claim-calibration audit: model, possible data, observed statistic, selection context, and conclusion. It transfers across experiments and observational studies when those roles are explicit; it stops at rote threshold replacement or at posterior and causal claims not supplied by the analysis.
Examples¶
Canonical¶
An article reports p=.03 and states there is only a three-percent chance the null is true and that the effect is important; both probability inversion and magnitude inference exceed the calculation.
Mapped back: p → .03; conditional direction → reversed; importance → unsupported.
Applied / In Practice¶
Two otherwise similar studies with p=.049 and p=.051 are described as demonstrating presence versus absence solely because one crosses 0.05.
Mapped back: continuity → nearly equal evidence; labels → opposite.
Applied / In Practice¶
A preregistered quality-control rule rejects a model at alpha=.05, reports the decision's error framework, and does not claim the null's posterior probability or effect importance.
Mapped back: rule → predeclared; claim → decision limited.
Structural Tensions¶
T1 — Simple Decision versus Graded Evidence. Thresholds support action but can erase the continuous and model-dependent character of evidence.
Diagnostic: Is the cutoff a decision policy or being presented as a natural truth boundary?
T2 — Single-Test Clarity versus Analysis Multiplicity. A reported p-value looks self-contained while its calibration can depend on all tests and selection choices.
Diagnostic: What analysis family and stopping rule produced this reported result?
T3 — Statistical Detectability versus Substantive Importance. Large samples can make tiny effects statistically detectable, while important effects may remain uncertain in small samples.
Diagnostic: What effect scale and interval answer the scientific question?
Structural–Framed Character¶
Misuse of P-Values is hybrid: structurally an evidence-to-claim mismatch and framed by statistical models, research designs, and decision conventions.
Structural Core vs. Domain Accent¶
The skeleton is a conditional measure interpreted beyond its support. Statistics supplies null models, sampling distributions, test statistics, error rates, multiple-comparison procedures, effect estimands, and the distinction among Fisherian evidence, Neyman–Pearson decisions, and other inferential frameworks.
Instantiates / Related Primes¶
This entry is a kind of Inferential Error.
-
Approved root. The frozen DAG keeps this family of misuse patterns distinct from the valid p-value and testing constructs it critiques.
-
Related — statistical significance, p-value, hypothesis testing, dichotomous thinking, multiple comparisons, effect size, and confidence interval. These supply the object, procedures, one special error, and corrective evidence.
Relationships to Other Abstractions¶
Current abstraction Misuse of p-values Domain-specific
Parents (1) — more general patterns this builds on
-
Misuse of p-values is a kind of Inferential Error Domain-specific
Misuse of p-values satisfies the defining boundary of Inferential Error: An inferential error is a conclusion, evidential interpretation, or uncertainty statement that is not warranted because the analysis misstates the target, unit, dependence structure, model, probability meaning, comparison, identification assumptions, multiplicity, or scope connecting observations to claims.Misuse of p-values satisfies the defining boundary of Inferential Error: An inferential error is a conclusion, evidential interpretation, or uncertainty statement that is not warranted because the analysis misstates the target, unit, dependence structure, model, probability meaning, comparison, identification assumptions, multiplicity, or scope connecting observations to claims.
Hierarchy path (1) — routes to 1 parentless root
- Misuse of p-values → Inferential Error
Neighborhood in Abstraction Space¶
Misuse of p-values sits in a crowded region of the domain-specific corpus (24th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.
Family — Statistical Hypothesis Tests & Diagnostics (9 abstractions)
Nearest neighbors
- Fitness-Proportionate Selection — 0.90
- Experiment (Probability Theory) — 0.90
- Chance-Constrained Programming — 0.89
- Approximate Bayesian Computation — 0.89
- Taguchi Loss Function — 0.89
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Statistical Significance / P-Value. Tell: That prime names the inferential quantity; this entry names systematic ways conclusions outrun it.
- Hypothesis Testing. Tell: A valid testing framework can use p-values without committing the listed interpretation errors.
- Dichotomous Thinking. Tell: Threshold cliff effects are one misuse, but inverse probability and magnitude or causality claims are others.
- Base-Rate Fallacy. Tell: Reversing conditional probability is related, while p-value misuse additionally includes design and selection errors.
- Absence of Evidence. Tell: Failure to reject can be misread as evidence of absence, but that is only one boundary case.
References¶
- Frozen Wikipedia discovery revision: https://en.wikipedia.org/wiki/Misuse_of_p-values (revision 1332777605).
- Preserved source candidate: https://www.stat.berkeley.edu/~aldous/Real_World/ASA_statement.pdf
- Preserved source candidate: https://xkcd.com/882/
- Preserved source candidate: https://books.google.com/books?id=FFCIBwAAQBAJ&pg=PA48
- Preserved source candidate: http://blog.minitab.com/blog/statistics-in-the-field/hypothesis-testing-and-p-values
- Preserved source candidate: https://www.cicm.org.au/CICM_Media/CICMSite/CICM-Website/Resources/Publications/CCR%20Journal/Previous%20Editions/June%202004/11_2004_Jun_Point-of-view.pdf
The frozen Wikipedia revision is discovery provenance. The retained source set was reviewed for identity, formal or operational relation, and scope. The encyclopedia's structural synthesis is bounded to those claims; a thin authority surface is recorded as a nonblocking source-strengthening repair rather than concealed.