Skip to content

HARKing (Hypothesizing After the Results are Known)

The research practice of building a hypothesis by inspecting already-collected data and then presenting it as if it had been specified in advance, silently inflating the reported false-positive rate because the test's independence assumption is violated.

Core Idea

HARKing — Hypothesizing After the Results are Known — is the research practice of formulating a hypothesis from inspection of data that have already been collected, then presenting that hypothesis in publication as if it had been specified before any data were examined. Coined and named by Norbert Kerr in 1998, HARKing is typically bundled with selective reporting: a researcher runs many analyses across many outcome measures, identifies the subset that reached statistical significance, and then constructs a theoretical rationale that is written in the paper's introduction as a prior prediction.

The practice generates a precise inferential distortion. The nominal type I error rate — the probability that a statistically significant result is a false positive — is computed on the assumption that the hypothesis was fixed before the data were seen. When the hypothesis is instead extracted from the data, the apparent test is not a test at all: the researcher has already seen which patterns emerged, so the "prediction" is guaranteed to match the result. The p-value is then the probability of the observed result under the null, but the relevant comparison is the probability that at least one of the many patterns examined would have reached nominal significance by chance — which is far larger. The gap between these two probabilities is the false-positive inflation HARKing produces.

Three structural consequences follow. First, positive findings in a HARKed literature will not replicate at rates the nominal alpha would predict, because each reported result bundles an unacknowledged multiple-comparison problem into a single p-value. Second, readers cannot distinguish genuine predictions that survived a test from patterns extracted from noise and retrolabelled as predictions — the literature is epistemically opaque about which studies were confirmatory and which were exploratory. Third, theory accumulates spuriously: because almost any post-hoc theoretical account can be made to fit almost any dataset, HARKing allows mutually contradictory theories to appear confirmed, preventing the literature from accumulating evidential weight. The institutional response that HARKing's identification made possible is preregistration — public time-stamped commitment to hypotheses and analysis plans before data collection — which restores the temporal order on which hypothesis-testing inference depends.

Structural Signature

Sig role-phrases:

  • the flexible study — a dataset with many measures, conditions, and possible analyses, so many patterns could be examined
  • the post-hoc inspection — an exploratory pass over the already-collected data, identifying which patterns happened to reach significance
  • the reverse-engineered hypothesis — a prediction generated from those observed patterns rather than fixed before the data were seen
  • the temporal misrepresentation — the post-hoc hypothesis presented in publication as if it had been specified a-priori, hiding the order of hypothesis and result
  • the inflated nominal alpha — the reported p-value no longer controls the false-positive rate, because the relevant comparison is the chance that any of the many patterns examined would reach significance, not this one under the null
  • the literature-level distortion — under-replication, contradictory theories all appearing confirmed, and inflated meta-analytic effect sizes accumulating across a HARKed body of work
  • the procedural remedy — preregistration / registered reports restoring the temporal order on which hypothesis-testing inference depends, where no analysis-side correction can reach

What It Is Not

  • Not exploratory analysis itself. Inspecting collected data for unanticipated patterns is legitimate, often valuable, science; what HARKing adds is the temporal misrepresentation — presenting the data-derived hypothesis as if it had been fixed a-priori. An exploratory finding honestly labelled as exploratory is not HARKing; the same finding dressed in the introduction as a prior prediction is.
  • Not p-hacking. Both are questionable research practices that inflate false positives, but they operate on different sides of the procedure: p-hacking manipulates the analysis (selectively running and reporting tests to reach significance), whereas HARKing manipulates the hypothesis (reverse-engineering the prediction from results already seen). They co-occur but are distinct, and their remedies differ — an FDR adjustment touches p-hacking but leaves HARKing's ordering violation intact.
  • Not a multiple-comparisons problem curable by correction. The inflation is not a defect in the statistical machinery that Bonferroni, Holm, or FDR could repair; it is a violation of the independence assumption on which the nominal alpha rests — the hypothesis was not fixed independently of the data. Analysis-side corrections address the testing side and leave the hypothesis-side reverse-engineering untouched, which is why the cure is procedural (preregistration), not analytic.
  • Not data fabrication or falsification. No datum is invented or altered; every reported number can be real and every analysis correctly computed. The deception is in the order — claiming foresight of a result that was in fact extracted from the data — not in the data themselves, which distinguishes HARKing from outright research fraud.
  • Not merely confirmation bias. Confirmation bias is a general cognitive tendency, present in all reasoning, to favor evidence consistent with a held belief. HARKing is a specific research-methodological practice with a formal inferential signature — a hidden temporal order that silently inflates a reported p-value — not a disposition of the individual reasoner.
  • Not a law that any post-hoc hypothesis is false. A hypothesis generated from the data may well be true; HARKing makes no claim about its truth-value, only that the test reported for it does not control the false-positive rate it advertises. The reverse-engineered prediction is unsupported as tested, awaiting a genuine confirmatory test on fresh data — not refuted.

Scope of Application

HARKing lives across the empirical-research-practice subfields of science — every field that reports quantitative hypothesis-confirmation tests against collected data and the metascientific apparatus built to police them; its reach is within that one substrate, where the temporal order of hypothesis and result can be hidden from the reader. (The loose narrative-construction analogues belong to narrative_fallacy / hindsight_bias, not here.)

  • Experimental psychology — the birthplace and densest habitat, where Kerr named it and the replication crisis identified it as a leading contributor to the low base rate of replication for published findings.
  • Biomedical and clinical research — where outcome switching between protocol and publication is the same hypothesis-side reverse-engineering, and trial preregistration (ClinicalTrials.gov) is the institutional cure.
  • Neuroimaging — where the sheer number of voxel-by-condition cells makes the garden of forking paths most severe, so post-hoc selection of "predicted" activations is acutely inflating.
  • Metascience / metaresearch — empirical measurement of HARKing's prevalence and of its effect-size inflation across literatures, treating "was the hypothesis fixed before the data?" as a measurable study attribute.
  • Editorial and journal policy — where the confirmatory-versus-exploratory distinction, registered reports, and pre-analysis plans are written into submission rules with HARKing as the named rationale.
  • Statistical-error analysis — separating analysis-side multiple-testing corrections (Bonferroni, Holm, FDR) from the hypothesis-side ordering violation they cannot reach, locating each remedy by where in the procedure it bites.
  • Economics — where the identical move is called specification searching: a regression specification chosen by inspecting which one yields significance, then presented as theory-driven.
  • Education research, sociology, and ecology — flexible-design observational fields where many measured cells and weak preregistration norms make data-derived hypotheses easy to relabel as predictions.
  • Machine-learning evaluation — where test-set tuning / fitting the leaderboard is the literal analogue: a model or configuration selected by repeatedly consulting held-out results, then reported as if committed in advance.

Clarity

Naming HARKing makes legible a distinction the published literature had blurred for decades: that the same analysis on the same data is epistemically confirmatory or merely exploratory depending entirely on the temporal order of hypothesis and result. Before the label, a reviewer's unease about an implausibly tidy paper — predictions that match the data too precisely, a theoretical rationale that arrives suspiciously well-fitted — had no name and invited only vague suspicion. HARKing converts that unease into a specific, prevalent, measurable practice with a specific structural remedy, and lets editors, reviewers, and meta-analysts speak about it precisely rather than gesturing at "fishing expeditions."

The sharper question it lets a methodologist ask is when was this hypothesis fixed relative to the data? — and that question reframes what a p-value even means here. The label clarifies that the inflation is not a statistical-machinery defect curable by a multiple-comparisons correction, but a violation of the independence assumption on which the nominal alpha rests; Bonferroni or FDR adjustments address the analysis-side multiple testing, yet leave the hypothesis-side reverse-engineering untouched. By separating confirmatory from exploratory as a property of procedure rather than of the analysis itself, HARKing isolates the failure to a stage no purely statistical fix can reach — and points to preregistration, a procedural rather than analytic remedy, as the corresponding cure.

Manages Complexity

Before HARKing was named, the trustworthiness of an empirical literature was a high-dimensional worry. A methodologist confronting a suspect field had to weigh, study by study, dozens of incommensurable signs — implausibly tidy predictions, theoretical rationales that fit too well, an excess of just-significant p-values, the analytic flexibility of the design, the publication incentives, the absence of any record of what was planned — with no principle telling which of these mattered for inference or how they combined. HARKing collapses that sprawl onto a single binary that does most of the work: was the hypothesis fixed before or after the data were seen? The whole question of whether a reported test controls its nominal false-positive rate reduces to the temporal order of hypothesis and result, because that order is exactly the assumption the alpha rests on. Once an analyst tracks that one variable, the qualitative consequences follow without re-derivation: a post-hoc hypothesis means the p-value understates the false-positive probability (it omits the unacknowledged multiple comparisons), means the finding sits in the exploratory rather than confirmatory column, and means it should be discounted to roughly the replication base rate rather than read at face value. The sign also separates fixes by where in the procedure they bite — analysis-side corrections (Bonferroni, FDR) cannot touch a hypothesis-side ordering violation, so only a procedural commitment that fixes the order (preregistration, registered reports) reaches it.

This compression also tames a second sprawl. The questionable-research-practices family is a multi-headed list — p-hacking, optional stopping, outcome switching, selective reporting — that otherwise has to be reasoned about case by case. HARKing carves off and names exactly the hypothesis-side member of that family, so a reviewer no longer asks the open-ended "is something wrong with this paper?" but the sharp, decidable "was this prediction reverse-engineered from the results?" — a question diagnosable from a small fixed set of signatures (no preregistration, post-hoc theoretical rationale in the introduction, a cluster of barely-significant results among many measures). The move is from a study-by-study integrity audit over many ill-defined cues to a single procedural parameter from which the inferential standing, the expected replication behavior, and the applicable remedy can all be read off.

Abstract Reasoning

Within research integrity and the study of questionable research practices, HARKing licenses reasoning moves that all turn on the temporal order of hypothesis and result and the independence assumption that order protects.

Diagnostic — detect reverse-engineered hypotheses from a fixed signature, and infer the inferential damage from the ordering. The signature move reads a study for the marks of post-hoc hypothesising: absence of preregistration, a theoretical rationale in the introduction that fits the results suspiciously well, and a cluster of barely-significant findings among many measures, conditions, and analyses. The analyst reasons FROM "predictions that match the data too precisely, with no a-priori record and many tested cells" TO "the hypothesis was probably extracted from the data and retrolabelled as a prediction." A second diagnostic move converts that ordering into the inferential consequence: reasoning FROM "the hypothesis was fixed after the data were seen" TO "the reported p-value understates the false-positive probability, because the relevant comparison is the chance that at least one of the many patterns examined would reach nominal significance, not the chance of this one pattern under the null." The reasoning is FROM the timing of the hypothesis TO whether the test controls its nominal alpha at all, and from the diagnostic signature TO the presence of HARKing.

Interventionist — restore the temporal order procedurally, because no analysis-side fix can reach a hypothesis-side violation. The decisive interventionist move targets the order itself: preregistration — a public, time-stamped commitment to hypotheses and analysis plans before data collection — is prescribed and predicted to restore the independence on which hypothesis-testing inference depends, with registered reports and results-blind review as stronger forms. The analyst reasons FROM "fix the hypothesis before the data are seen" TO "the apparent test becomes a genuine test." A second interventionist move locates fixes by where in the procedure they bite and predicts the failure of mismatched ones: reasoning FROM "Bonferroni and FDR correct analysis-side multiple testing" TO "they cannot touch hypothesis-side reverse-engineering," so a multiple-comparisons adjustment is predicted to leave HARKing's inflation intact. A third interventionist move, at the point of detection, prescribes a direct preregistered replication on a fresh sample, predicted to fail to reach significance if the original finding was a HARKed false positive — converting suspicion into a decisive test.

Boundary-drawing — separate the hypothesis-side practice from its analysis-side sibling, and confirmatory from exploratory as a property of procedure. A first boundary move distinguishes HARKing from p-hacking within the QRP family: both inflate false positives, but HARKing operates on the hypothesis side (constructing the prediction after seeing the result) while p-hacking operates on the analysis side (selectively running and reporting analyses to reach significance) — so the analyst reasons FROM "is the manipulation of the prediction or of the analysis?" TO "which pathway, and which remedy applies," noting preregistration covers both while FDR covers only p-hacking. A second boundary move makes confirmatory versus exploratory a property of procedure rather than of the analysis: the same analysis on the same data is confirmatory if the hypothesis was fixed first and merely exploratory if it was extracted from the results, so the analyst reasons FROM "when was the hypothesis fixed relative to the data?" TO "which column this result belongs in." A third boundary move classifies the defect: reasoning FROM "the inflation is a violation of the independence assumption the alpha rests on, not a flaw in the statistical machinery" TO "the cure is procedural (preregistration), not analytic (a correction factor)."

Predictive — a HARKed literature will under-replicate, accumulate contradictory theory, and overstate effects. A forward move predicts replication behaviour from the practice: because each reported result bundles an unacknowledged multiple-comparison problem into a single p-value, the analyst forecasts that positive findings in a HARKed literature will not replicate at the rate the nominal alpha implies, and predicts that a field with strong publication bias and weak preregistration will appear to confirm hypotheses at high rates while actually replicating at low ones. A second predictive move concerns theory: because almost any post-hoc account can be fitted to almost any dataset, the analyst predicts that mutually contradictory theories will all appear confirmed, so the literature fails to accumulate evidential weight. A third predictive move concerns aggregation: reasoning FROM "selection on the significant subset" TO "meta-analytic effect sizes are inflated and failed-replication base rates underestimated," the analyst forecasts the direction of distortion in any synthesis built on a HARKed body of work.

Knowledge Transfer

Within empirical science HARKing transfers as mechanism, literally and without translation, across every field that reports quantitative hypothesis-confirmation tests against collected data. The diagnostic signature (no preregistration, a too-well-fitting theoretical rationale in the introduction, a cluster of barely-significant results among many measured cells), the inferential consequence (the nominal alpha no longer controls the false-positive rate because the hypothesis was not fixed independently of the data), and the remedy (preregistration, registered reports, results-blind review) carry intact from experimental psychology — its birthplace — to biomedical and clinical research, to neuroimaging, where the sheer number of voxel-by-condition cells makes the garden of forking paths most severe, to education research, sociology, and ecology. The vocabulary travels too: economics calls the same move specification searching, machine-learning evaluation calls it test-set tuning or fitting the leaderboard, but in each the structure is identical — a hypothesis (or model, or specification) reverse-engineered from results that were already seen, then presented as if committed in advance. The currency of the test changes; the violated independence assumption, the false-positive inflation, and the procedural cure do not. This is the substrate over which HARKing is a genuine mechanism, and it is broad — but it is one substrate (empirical research practice in which a hypothesis is quantitatively tested against data and the temporal order of hypothesis and result is hidden from the reader), not several structurally distinct ones, which is exactly why it remains a domain-specific abstraction.

Beyond that substrate the term is used in two honestly different ways. First, by analogy. "HARKing" is borrowed loosely for legal narrative construction (a closing argument built backward from the verdict the advocate wants), for investigative journalism (writing the conclusion before assembling the evidence), and for policy analysis (reasoning back from a desired recommendation to the supporting findings). These rename the components — prediction becomes argument or narrative, the dataset becomes the case file or the brief — and borrow the shape of after-the-fact rationalization presented as foresight, but they drop the machinery that gives HARKing its force: there is no p-value whose nominal alpha is being silently inflated, no multiple-comparison structure across measured cells, no formal-inference penalty at all. The resemblance is real and sometimes illuminating, but it is pattern-by-resemblance, and the honest move is to mark it as analogy rather than as the mechanism traveling.

Second, and more usefully, the general lessons HARKing carries cross-domain are already available in more portable form from the patterns it instantiates, and those — not the named practice — are what should travel. The deep statistical pathology is selection on the outcome: choosing which hypothesis to "test" by looking at which results came out significant is selecting on the dependent variable, and that abstraction recurs wherever a claim is evaluated against the very evidence used to generate it, carried by selection_bias. The deep epistemic pathology is constructing a fit after the fact and mistaking it for prediction, which recurs in history, autobiography, and ordinary sense-making and is carried by hindsight_bias and narrative_fallacy. And the deep remedial pattern — fixing a claim publicly and verifiably before the evidence arrives, so that matching the evidence afterward actually means something — is one instance of commitment-before-evidence, the same structure as a sealed bid, a hash-commitment in cryptography, an escrowed forecast, or a pre-registered prediction in any tournament. When the cross-domain lesson is needed, it is these parent patterns that recur as co-instances; "HARKing," as named, carries research-methodological cargo — the alpha, the QRP family, the preregistration apparatus, the replication-crisis context — that stays home. The general pattern travels; the named concept does not, and that boundary is the point of Structural Core vs. Domain Accent below.

Examples

Canonical

Since Norbert Kerr coined the term in 1998, the defining illustration is the arithmetic of the hidden multiple comparison. Suppose a study measures 20 independent outcomes and the researcher tests each at the nominal significance level α = 0.05. Under the null hypothesis (no real effect anywhere), the probability that a given test does not reach significance is 0.95, so the probability that none of the 20 does is 0.95²⁰ ≈ 0.358. The probability that at least one reaches significance by chance is therefore 1 − 0.358 ≈ 0.64 — about 64%, not 5%. The HARKing researcher runs all 20, keeps the one significant outcome, and writes the paper's introduction as though that single result had been the a-priori prediction tested at p < 0.05. The reported test thus advertises a 5% false-positive rate while the true rate for finding some "confirmation" was roughly 64%.

Mapped back: The 20-outcome dataset is the flexible study; running all 20 and keeping the winner is the post-hoc inspection yielding the reverse-engineered hypothesis. Presenting it as the prior prediction is the temporal misrepresentation, and the gap between the advertised 5% and the real ~64% is exactly the inflated nominal alpha.

Applied / In Practice

Kaplan and Irvin's 2015 analysis of large NIH-funded cardiovascular trials shows the remedy working. The National Heart, Lung, and Blood Institute began requiring public preregistration of primary outcomes on ClinicalTrials.gov around 2000. Comparing 55 large trials before and after this rule, the authors found that the proportion reporting a statistically significant benefit on the primary outcome fell from about 57% (pre-2000) to about 8% (post-2000). Fixing the hypothesis and primary endpoint publicly before data collection removed the ability to reverse-engineer a "confirmed" outcome after the fact.

Mapped back: Before the rule, investigators could inspect results and elevate whichever endpoint reached significance — the post-hoc inspection feeding the temporal misrepresentation. Mandated preregistration is the procedural remedy restoring the temporal order; the collapse from 57% to 8% is the previously inflated nominal alpha deflating once the hypothesis-side ordering violation was blocked, a remedy no analysis-side correction could have supplied.

Structural Tensions

T1: Legitimate exploration versus temporal misrepresentation (the wrong is the relabeling, not the looking). HARKing is not the act of inspecting collected data for unanticipated patterns — that is legitimate, often valuable, science. What converts it into HARKing is a single added move: presenting the data-derived hypothesis as if it had been fixed a-priori. The tension is that the exploratory pass and the confirmatory claim run on the same observed pattern, so the very finding that is honest and useful when labelled "exploratory" becomes an inference-corrupting deception when dressed in the introduction as a prior prediction. A regime that polices HARKing too bluntly risks chilling the exploration itself; one that tolerates the relabeling destroys the confirmatory/exploratory distinction the whole inferential apparatus rests on. The defect lives entirely in the reporting of temporal order, not in the analysis, which is why the same result can be exemplary or fraudulent depending on one unobservable historical fact. Diagnostic: Is the objection here to the researcher finding the pattern in the data, or only to their claiming they predicted it in advance?

T2: Hypothesis-side violation versus analysis-side correction (the statistical instinct that cannot reach it). Confronted with inflated false positives, the trained reflex is to reach for a multiple-comparisons correction — Bonferroni, Holm, FDR. HARKing defeats that reflex: the inflation is not a defect in the statistical machinery but a violation of the independence assumption on which the nominal alpha rests, and analysis-side corrections address the testing side while leaving the hypothesis-side reverse-engineering untouched. The tension is that the problem looks like a multiple-testing problem — it even reduces, arithmetically, to the same "chance that at least one of many patterns reaches significance" — yet the fix that multiple-testing problems invite is structurally incapable of reaching it. Only a procedural remedy that restores the temporal order (preregistration) bites, and a researcher who applies FDR believes themselves cured while the ordering violation, and its inflation, remain fully intact. Diagnostic: Does the proposed remedy operate on how the analysis was run (correctable) or on when the hypothesis was fixed relative to the data (only preregistration reaches it)?

T3: Confirmatory versus exploratory as an invisible procedural fact (identical results, unverifiable status). HARKing makes confirmatory-versus-exploratory a property of procedure — the temporal order of hypothesis and result — rather than of the analysis or the data. This is the concept's sharpest clarification and also its deepest vulnerability: the epistemic standing of a result depends on a historical fact (when the hypothesis was fixed) that the finished paper, the dataset, and the p-value cannot reveal. The literature is therefore epistemically opaque by construction — a reader genuinely cannot distinguish a prediction that survived a test from a pattern extracted from noise and retrolabelled, because the two produce identical artifacts. The tension is that the very move that makes the distinction matter (procedure, not content) is the move that makes it invisible after the fact, so it can only be secured before the data exist, never audited from the result alone. Diagnostic: Can the confirmatory status of this result be established from any durable pre-data record, or does it rest entirely on the author's unverifiable account of when the hypothesis was set?

T4: Discredited test versus false hypothesis (HARKed is not the same as wrong). HARKing makes no claim about the truth of the reverse-engineered hypothesis; it claims only that the reported test does not control the false-positive rate it advertises. A data-derived prediction may well be true — it is unsupported as tested, awaiting a genuine confirmatory test on fresh data, not refuted. The tension is that this precision cuts against two opposite over-reactions: treating a HARKed finding as confirmed inflates the literature with false positives, but treating it as false discards what may be a real, replicable lead that exploration legitimately surfaced. The correct handling — discount the finding to the replication base rate and re-test it prospectively — is more demanding than either dismissal or acceptance, and it requires holding "the test is worthless" and "the hypothesis might be right" simultaneously. Diagnostic: Is the HARKed finding being treated as refuted, or as an untested candidate whose advertised confirmation is void but whose truth is still open to a fresh preregistered test?

T5: Diagnostic signature versus honest exploration (detection cues that legitimate work also trips). HARKing is diagnosed from a fixed signature — no preregistration, a theoretical rationale that fits the results suspiciously well, a cluster of barely-significant findings among many measured cells. But every one of those cues is also produced by honest, well-labelled exploratory work in a flexible-design field, and by genuine a-priori predictions that happened to land. The signature is circumstantial: it raises the posterior probability of HARKing without proving it, so acting on it risks false accusation of legitimate researchers, while demanding proof risks letting the practice pass undetected. The tension is that the practice is defined by an unobservable (temporal order) yet must be policed through observable proxies that do not cleanly separate it from its innocent look-alikes — the detector and the honest exploratory paper can be indistinguishable. Diagnostic: Do the HARKing signatures in this study reflect a hidden temporal misrepresentation, or an honestly-conducted exploratory analysis that simply shares the same surface marks?

T6: Autonomy versus reduction (a named research-integrity practice or the instance of its selection-and-commitment parents). HARKing is a specific, canonically-named methodological practice (Kerr 1998) carrying real research cargo — the nominal alpha, the QRP family, the preregistration apparatus, the replication-crisis context. It travels as mechanism across empirical science, even under other names (economics' specification searching, ML's test-set tuning) — but that breadth is one substrate: research practice in which a hypothesis is quantitatively tested against data with the temporal order hidden. Beyond it, "HARKing" for legal or journalistic narrative-building is analogy that drops the p-value and the inference penalty. What genuinely ports is not the named practice but the patterns it instantiates: selection_bias (choosing the hypothesis by looking at which outcomes came out significant is selecting on the dependent variable), hindsight_bias/narrative_fallacy (a fit constructed after the fact mistaken for foresight), and commitment-before-evidence (preregistration as a sealed bid or hash-commitment). The tension is between a construct specific enough to anchor journal policy and the recognition that its cross-domain lessons already live in those parents. Diagnostic: Resolve toward selection_bias + hindsight_bias/narrative_fallacy + commitment-before-evidence when carrying the lesson outside quantitative research; toward HARKing itself when a reported hypothesis-test's nominal alpha and temporal order are the live objects.

Structural–Framed Character

HARKing sits at the framed pole of the spectrum, closely parallel to how ad hominem is characterized: it is not a neutral mechanism but a normatively charged, practice-constituted verdict about a piece of research conduct, and its every distinctive feature is metascience furniture.

On evaluative_weight it scores high: to call a study "HARKed" is to convict it of a questionable research practice — a temporal misrepresentation that voids the test it advertises — so the label renders a research-integrity finding, not a value-free description the way "feedback" or "isostasy" names something neither good nor bad. (The entry's care that "HARKed is not the same as wrong" is a precision about the hypothesis's truth-value, not a softening of the verdict on the practice, which remains condemned.) Human_practice_bound is high in the strongest sense: HARKing is constituted by the practice of quantitative hypothesis-testing research and its publication norms, and it dissolves the instant that practice is removed — with no p-value whose nominal alpha can be inflated, no confirmatory/exploratory reporting convention, no introduction in which a prediction can be presented as prior, "hypothesizing after the results are known" is not a research offense but merely ordinary learning from data. Institutional_origin is equally pronounced: the entry is furniture of a specific metascientific tradition — Kerr's 1998 coinage, the QRP family, the nominal-alpha apparatus, preregistration and registered reports, ClinicalTrials.gov, the replication-crisis context — all distinctions drawn inside the institution of empirical publishing, not substrate-neutral form. On vocab_travels it is domain-bound: within empirical science the mechanism travels literally even under other names (specification searching, test-set tuning), but that breadth is one substrate, and off it the operative vocabulary loses its referents — "HARKing" for a lawyer's closing argument or a journalist's back-built narrative is import-by-analogy, borrowing the shape of after-the-fact rationalization-dressed-as-foresight while dropping the alpha, the multiple-comparison structure, and the formal-inference penalty entirely, so import_vs_recognize patterns as analogy off-substrate.

The one structural-looking feature is the formal inferential signature — selecting which hypothesis to "test" by inspecting which outcomes came out significant is selection on the dependent variable, a genuine, portable statistical pathology that recurs wherever a claim is judged against the very evidence used to generate it. That skeleton is real and substrate-portable, which is what tempts a more structural reading, but it does not pull HARKing off the framed pole, because it is precisely what HARKing instantiates from its umbrella (selection_bias), not what makes "HARKing" itself travel: the cross-domain reach belongs to the general selection-on-the-outcome pattern, while the alpha, the QRP taxonomy, and the preregistration machinery — the part that makes it HARKing in particular — stay home. Two further facets it instantiates deserve their own naming because they are demonstrably distinct parents, not padding: the epistemic pathology of mistaking an after-the-fact fit for foresight (hindsight_bias / narrative_fallacy) and the remedial structure of fixing a claim publicly before the evidence arrives (commitment-before-evidence, the sealed bid or hash-commitment). Its character: a normatively charged, research-practice-constituted integrity verdict whose every distinctive feature is metascience furniture, structural only in the selection-on-the-outcome skeleton it borrows from selection_bias and frames as a methodological offense.

Structural Core vs. Domain Accent

This section decides why HARKing is a domain-specific abstraction and not a prime, and it carries the case for its domain-specificity — there is no separate section for that. Its skeleton is genuinely multiple — three distinct portable pathologies are braided together — so each is named in turn rather than collapsed.

What is skeletal (could lift toward a cross-domain prime). Strip the metascience and three thin relational structures survive, each substrate-portable and each already carried by a parent prime. The load-bearing one is selection on the outcome: choosing which claim to "test" by inspecting which results came out favourable is selecting on the very evidence used to generate the claim, so the surviving claim is guaranteed to fit — this is selection_bias, and it recurs wherever a hypothesis is judged against the data that produced it. A second is mistaking an after-the-fact fit for foresight: constructing an account that matches known results and then experiencing it as a prior prediction — carried by hindsight_bias and narrative_fallacy. A third is the remedy's structure, commitment-before-evidence: fixing a claim publicly and verifiably before the evidence arrives so that a later match means something — the same shape as a sealed bid, a cryptographic hash-commitment, or an escrowed forecast. These three are genuinely doubled (tripled) skeletons, not padding: they are demonstrably distinct patterns, and together they are what HARKing instantiates. But they are the cores HARKing shares with its parents, not what makes HARKing distinctive.

What is domain-bound. Almost everything that makes it HARKing in particular is research-methodology furniture: the nominal alpha whose false-positive rate is silently inflated; the hidden multiple-comparison structure across measured cells; the confirmatory-versus-exploratory reporting convention; the temporal misrepresentation in a paper's introduction; the QRP family (p-hacking, optional stopping, outcome switching); and the preregistration / registered-reports apparatus (ClinicalTrials.gov, results-blind review) that constitutes its cure. The decisive test: remove the quantitative hypothesis-test — the p-value whose nominal alpha can be inflated — and "hypothesizing after the results are known" is no longer a research offense at all but ordinary, honest learning from data. The formal-inference penalty, the thing that makes it HARKing rather than mere hindsight, has no referent once the statistical test and its publication norms are gone.

Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose transfer is recognition of the same mechanism, not analogy. HARKing's transfer is bimodal. Within empirical research practice it travels intact and literally — even under other names (economics' specification searching, ML's test-set tuning / fitting the leaderboard), the reverse-engineered-and-relabelled hypothesis, the inflated alpha, and the procedural cure are recognized, not re-derived — but that breadth is a single substrate (quantitative hypothesis-testing with the temporal order hidden from the reader). Beyond it, "HARKing" for a lawyer's back-built closing argument or a journalist's conclusion-first narrative is analogy: it renames the components and drops the p-value, the multiple-comparison structure, and the inference penalty. And when the bare structural lessons are needed cross-domain, they are already supplied, in more general form, by the parents HARKing braids together — selection_bias for the outcome-selection pathology, hindsight_bias/narrative_fallacy for the after-the-fact-fit pathology, and commitment-before-evidence for the remedy. The cross-domain reach belongs to those parents; "HARKing," as named, carries research-methodological baggage — the alpha, the QRP taxonomy, the preregistration machinery, the replication-crisis context — that does not and should not travel.

Relationships to Other Abstractions

Local relationship map for HARKing (Hypothesizing After the Results are Known)Parents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.HARKing (Hypothesizi…DOMAINPrime abstraction: Hypothesis Testing (Null vs. Alternative) — is part ofHypothesis Test…PRIMEDomain-specific abstraction: Hindsight Bias — is a decomposition of, typicalHindsight BiasDOMAINPrime abstraction: Selection Bias — is a decomposition ofSelection BiasPRIME

Current abstraction HARKing (Hypothesizing After the Results are Known) Domain-specific

Parents (3) — more general patterns this builds on

  • HARKing (Hypothesizing After the Results are Known) is part of Hypothesis Testing (Null vs. Alternative) Prime

    The HARKing practice contains a hypothesis test whose apparent prespecification and nominal error rate are the objects being falsified.

  • HARKing (Hypothesizing After the Results are Known) is a decomposition of, typical Hindsight Bias Domain-specific

    HARKing typically recruits hindsight by making an explanation fitted after the result appear to have been predictable beforehand.

  • HARKing (Hypothesizing After the Results are Known) is a decomposition of Selection Bias Prime

    HARKing selects the claim by inspecting which outcomes favored it, so the reported hypothesis is conditioned on the evidence later used to test it.

Hierarchy paths (22) — routes to 12 parentless roots

Not to Be Confused With

  • p-hacking. The analysis-side member of the questionable-research-practices family: selectively running, transforming, or reporting analyses (dropping outliers, trying covariates, optional stopping) until a test crosses significance. Both inflate false positives, but p-hacking manipulates the analysis while HARKing manipulates the hypothesis — reverse-engineering the prediction from results already seen. Their cures differ: an FDR correction touches p-hacking but leaves HARKing's ordering violation intact. Tell: was the prediction fixed and the analysis flexed to reach it (p-hacking), or was the analysis fixed and the prediction retro-fitted to the result (HARKing)?

  • Outcome switching. The biomedical practice of changing which endpoint is reported as primary between a trial's protocol and its publication, elevating whichever measure reached significance. It is a close sibling — the hypothesis-side reverse-engineering applied to a pre-declared endpoint list — and is essentially HARKing operating against a registered protocol. HARKing is the general move (inventing the prediction post hoc); outcome switching is the specific case of swapping a pre-specified outcome. Tell: was there a pre-declared set of outcomes one of which got promoted after the fact (outcome switching), or was the hypothesis constructed from the data with no prior commitment at all (general HARKing)?

  • Publication bias. The literature-level distortion by which significant, positive studies are published and null studies are filed away, so the visible record overstates effects. This operates across studies at the journal/field level; HARKing operates within a study, at the hypothesis-formation stage. They compound each other but are distinct mechanisms. Tell: is the selection happening among which studies get published (publication bias), or among which patterns within one study get relabelled as predictions (HARKing)?

  • Data fabrication / falsification. Outright research fraud — inventing or altering data. HARKing invents no datum: every reported number can be real and every analysis correctly computed. The deception is purely in the temporal order — claiming foresight of a result extracted from the data — not in the data themselves. Tell: were any values manufactured or changed (fabrication/falsification), or is the data genuine and only the claim of prior prediction false (HARKing)?

  • Confirmation bias. The general cognitive tendency, present in all reasoning, to favour evidence consistent with a held belief. HARKing is not a disposition of the individual reasoner but a specific research-methodological practice with a formal inferential signature — a hidden temporal order that silently voids a reported test's nominal alpha. Tell: is this a broad psychological leaning toward confirming a prior belief (confirmation bias), or a discrete methodological act of presenting a data-derived hypothesis as pre-specified (HARKing)?

  • The parent cluster it instantiates (selection bias, hindsight bias / narrative fallacy, commitment-before-evidence). The substrate-neutral patterns HARKing braids together: selection_bias (choosing the hypothesis by which outcomes came out significant is selecting on the dependent variable), hindsight_bias/narrative_fallacy (an after-the-fact fit mistaken for foresight), and commitment-before-evidence (preregistration as a sealed bid or hash-commitment — the remedy's structure). These carry the cross-domain lessons; HARKing adds the alpha, the QRP taxonomy, and the preregistration apparatus that stay home. Tell: strip away the quantitative hypothesis-test and its publication norms and what remains is bare outcome-selection, hindsight, and pre-commitment — the parent patterns, not HARKing. (Treated fully in a later section.)

Neighborhood in Abstraction Space

HARKing (Hypothesizing After the Results are Known) sits in a sparse region of the domain-specific corpus (89th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Publication Bias & Research Artifacts (5 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12