HARKing (Hypothesizing After the Results are Known)¶
The research practice of building a hypothesis by inspecting already-collected data and then presenting it as if it had been specified in advance, silently inflating the reported false-positive rate because the test's independence assumption is violated.
Core Idea¶
HARKing, coined by Norbert Kerr in 1998, is formulating a hypothesis from inspection of already-collected data, then presenting it in publication as if it had been specified before any data were seen. It typically bundles with selective reporting. The distortion is precise: the nominal type I error rate assumes the hypothesis was fixed before the data, so a hypothesis extracted from the data makes the reported p-value understate the true false-positive probability — inflation no analysis-side correction can reach.
Scope of Application¶
HARKing lives across the empirical-research-practice subfields of science — every field reporting quantitative hypothesis-confirmation tests against collected data — one substrate where the temporal order of hypothesis and result can be hidden.
- Experimental psychology — the birthplace and densest habitat, a leading contributor to the replication crisis.
- Biomedical and clinical research — outcome switching, with trial preregistration as the institutional cure.
- Neuroimaging — where voxel-by-condition cells make the garden of forking paths most severe.
- Economics — the identical move called specification searching.
- Machine-learning evaluation — test-set tuning or fitting the leaderboard.
Clarity¶
Naming HARKing makes legible a distinction the literature had blurred: the same analysis on the same data is confirmatory or merely exploratory depending entirely on the temporal order of hypothesis and result. It converts vague unease about a too-tidy paper into a specific, measurable practice with a specific remedy. The sharper question is "when was this hypothesis fixed relative to the data?" — clarifying that the cure is procedural, not a multiple-comparisons correction.
Manages Complexity¶
Before HARKing was named, judging a literature's trustworthiness meant weighing dozens of incommensurable signs study by study. HARKing collapses that sprawl onto one binary: was the hypothesis fixed before or after the data were seen? Once an analyst tracks that variable, the consequences follow without re-derivation — inflated false-positive rate, exploratory column, discount to the replication base rate. It also carves the hypothesis-side member cleanly out of the questionable-research-practices family.
Abstract Reasoning¶
HARKing licenses diagnostic reasoning — detecting reverse-engineered hypotheses from a fixed signature and inferring inferential damage from the ordering. It supports interventionist reasoning that restores the temporal order procedurally because no analysis-side fix reaches it, boundary-drawing that separates it from p-hacking and makes confirmatory-versus-exploratory a property of procedure, and predictive reasoning that a HARKed literature will under-replicate and accumulate contradictory theory.
Knowledge Transfer¶
Within empirical science HARKing transfers literally as mechanism across every field reporting quantitative hypothesis tests, even as the vocabulary changes (specification searching, test-set tuning) — but this is one substrate, which keeps it domain-specific. Beyond it, loose uses in legal or journalistic narrative are analogy, dropping the p-value machinery. The general lessons travel better through the parent patterns it instantiates: selection_bias, hindsight_bias/narrative_fallacy, and commitment-before-evidence.
Relationships to Other Abstractions¶
Current abstraction HARKing (Hypothesizing After the Results are Known) Domain-specific
Parents (3) — more general patterns this builds on
-
HARKing (Hypothesizing After the Results are Known) is part of Hypothesis Testing (Null vs. Alternative) Prime
The HARKing practice contains a hypothesis test whose apparent prespecification and nominal error rate are the objects being falsified.
-
HARKing (Hypothesizing After the Results are Known) is a decomposition of, typical Hindsight Bias Domain-specific
HARKing typically recruits hindsight by making an explanation fitted after the result appear to have been predictable beforehand.
-
HARKing (Hypothesizing After the Results are Known) is a decomposition of Selection Bias Prime
HARKing selects the claim by inspecting which outcomes favored it, so the reported hypothesis is conditioned on the evidence later used to test it.
Hierarchy paths (22) — routes to 12 parentless roots
- HARKing (Hypothesizing After the Results are Known) → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Inductive Reasoning
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Bias
- HARKing (Hypothesizing After the Results are Known) → Selection Bias → Bias
- HARKing (Hypothesizing After the Results are Known) → Selection Bias → Statistical Inference → Inductive Reasoning
- HARKing (Hypothesizing After the Results are Known) → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Uncertainty
- HARKing (Hypothesizing After the Results are Known) → Selection Bias → Statistical Inference → Uncertainty
- HARKing (Hypothesizing After the Results are Known) → Selection Bias → Vantage-Induced Omission → Viewpoint
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Reconstructive Memory → Schema → Abstraction
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Reconstructive Memory → Pattern Completion (Filling the Incomplete) → Inductive Reasoning
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Past-State Contamination → Temporal Dynamics → Time
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Reconstructive Memory → Pattern Completion (Filling the Incomplete) → Predictive Coding → Feedback
- HARKing (Hypothesizing After the Results are Known) → Hypothesis Testing (Null vs. Alternative) → Verification → Evaluation → Comparison → Self Checking
- HARKing (Hypothesizing After the Results are Known) → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Probability → Measure → Set and Membership
- HARKing (Hypothesizing After the Results are Known) → Selection Bias → Statistical Inference → Probability → Measure → Set and Membership
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Reconstructive Memory → Pattern Completion (Filling the Incomplete) → Interpretation → Representation → Abstraction
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Reconstructive Memory → Pattern Completion (Filling the Incomplete) → Predictive Coding → Compression → Abstraction
- HARKing (Hypothesizing After the Results are Known) → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
- HARKing (Hypothesizing After the Results are Known) → Selection Bias → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Reconstructive Memory → Pattern Completion (Filling the Incomplete) → Predictive Coding → Compression → Optimization
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Reconstructive Memory → Pattern Completion (Filling the Incomplete) → Predictive Coding → Encoding And Decoding → Transformation → Function (Mapping)
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Reconstructive Memory → Pattern Completion (Filling the Incomplete) → Predictive Coding → Compression → Aggregation → Micro Macro Linkage
- HARKing (Hypothesizing After the Results are Known) → Hindsight Bias → Reconstructive Memory → Pattern Completion (Filling the Incomplete) → Predictive Coding → Prediction Error → Baseline Deviation → Comparison → Self Checking
Neighborhood in Abstraction Space¶
HARKing (Hypothesizing After the Results are Known) sits in a sparse region of the domain-specific corpus (89th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Publication Bias & Research Artifacts (5 abstractions)
Nearest neighbors
- Funnel Plot Asymmetry — 0.82
- Jeffreys-Lindley Paradox — 0.82
- File Drawer Problem — 0.81
- Null Ritual — 0.81
- Congruence Bias — 0.81
Computed from structural-signature embeddings · 2026-07-12