Skip to content

Misuse of p-values

Inferential errors that treat a p-value as evidence about hypothesis probability, causation, effect magnitude, practical importance, replicability, or categorical truth beyond its model-conditional tail-probability meaning.

Version
v1 · 2026-09-28 · History
Domain-specific #
10743
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomains
Hypothesis Testing, Statistical Inference → Experimental Design & Statistics

Core Idea

A p-value is calculated in the direction from a specified model to possible data: assuming the model and analysis procedure, how extreme is the observed statistic or something more extreme? Common misuse reverses that condition and speaks as though the number were the probability the hypothesis is true given the data.

Other errors arise even when the definition is recited correctly. A conventional cutoff can be reified as a discovery boundary; selection and multiplicity can be hidden; a low value can be mistaken for a large, causal, or replicable effect; and failure to reject can be called proof of no effect. The defect lies in the relation between calculation and claim.

Scope of Application

  • Scientific reporting. Audits whether prose claims match the design and model-conditional statistic.
  • Multiple testing. Adjusts or contextualizes evidence across analysis families and selection.
  • Policy and quality control. Separates a predeclared action threshold from a claim of scientific truth.
  • Evidence synthesis. Combines effect estimates, intervals, prior knowledge, and replication rather than tallying significance.
  • Statistical education. Teaches conditional direction and nearest misconceptions through explicit examples.

Clarity

Quote the exact claim, then write the probability conditioning represented by the p-value. State null model, statistic, sidedness, sample design, stopping and selection, multiplicity, effect estimate, interval, and substantive scale. A mismatch between these ingredients identifies the specific misuse instead of applying a generic accusation. Inclusion test: Classify misuse only after reconstructing the null model, test statistic, design, analysis family, and exact claim made from the p-value. Exclusion test: Exclude legitimate model-conditional compatibility statements, predeclared decision rules with controlled errors, and criticism based solely on disliking frequentist methods. Nearest boundary: Dichotomous thinking treats results on opposite sides of a cutoff as qualitatively different; p-value misuse is broader and also includes inverse-probability, causal, magnitude, and replicability errors. Exit condition: The misuse disappears when the conclusion is limited to the specified model comparison and accompanied by effect estimates, uncertainty, design constraints, and selection context. Common misclassifications: It is not the claim that p-values are never useful. It is not merely using the conventional 0.05 level. It is not evidence that a null hypothesis is probably true when p is large. It is not corrected by replacing one threshold with another while preserving the same overclaim. Nearest named distinctions: Statistical Significance / P-Value: That prime names the inferential quantity; this entry names systematic ways conclusions outrun it. Hypothesis Testing: A valid testing framework can use p-values without committing the listed interpretation errors. Dichotomous Thinking: Threshold cliff effects are one misuse, but inverse probability and magnitude or causality claims are others. Base-Rate Fallacy: Reversing conditional probability is related, while p-value misuse additionally includes design and selection errors. Absence of Evidence: Failure to reject can be misread as evidence of absence, but that is only one boundary case.

Manages Complexity

The abstraction turns a diffuse list of bad practices into auditable mappings from evidence to claim. Separating model, design, calculation, selection, decision, and substantive interpretation shows whether correction requires better computation, fuller reporting, a different estimand, or a narrower conclusion.

Abstract Reasoning

  1. Reconstruct the scientific question and effect quantity of interest.
  2. Write the exact null model and probability statement used to compute the p-value.
  3. Inspect design, stopping, preprocessing, multiplicity, and reporting selection.
  4. Compare the published conclusion with the limited implications of the calculation.
  5. Add effect size, uncertainty, sensitivity, and domain-relevant importance.
  6. Distinguish a decision rule from a graded evidential or causal statement.

Knowledge Transfer

The transferable cargo is a conditional-direction and claim-calibration audit: model, possible data, observed statistic, selection context, and conclusion. It transfers across experiments and observational studies when those roles are explicit; it stops at rote threshold replacement or at posterior and causal claims not supplied by the analysis.

Relationships to Other Abstractions

Local relationship map for Misuse of p-valuesParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Misuse of p-valuesDOMAINDomain-specific abstraction: Inferential Error — is a kind ofInferentialErrorDOMAIN

Current abstraction Misuse of p-values Domain-specific

Parents (1) — more general patterns this builds on

  • Misuse of p-values is a kind of Inferential Error Domain-specific

    Misuse of p-values satisfies the defining boundary of Inferential Error: An inferential error is a conclusion, evidential interpretation, or uncertainty statement that is not warranted because the analysis misstates the target, unit, dependence structure, model, probability meaning, comparison, identification assumptions, multiplicity, or scope connecting observations to claims.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Misuse of p-values sits in a crowded region of the domain-specific corpus (24th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Statistical Hypothesis Tests & Diagnostics (9 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08