Null Ritual¶
The institutionalised practice of mechanically executing a null hypothesis significance test — nil-null, p-value, p < .05 verdict — severed from the alternatives, priors, effect sizes, and decision context inference requires, yet retaining full editorial authority as if it had not been.
Core Idea¶
The null ritual, named and diagnosed by Gerd Gigerenzer (2004), is the institutionalised practice of mechanically executing a null hypothesis significance test — setting up a nil-null hypothesis, computing a p-value, applying a p < .05 threshold, reporting the binary reject/retain verdict — without engaging the theoretical commitments, alternative hypotheses, prior probabilities, effect sizes, or decision context that statistical inference is supposed to serve. The structural failure is that the procedure has been severed from the inferential reasoning it was designed to support yet retains full editorial and social authority as if it had not been: researchers perform the ritual because journals and reviewers require it, not because the question they are asking is one the ritual answers. The hybrid procedure the ritual performs is itself historically anomalous — it fuses Fisher's p-value significance test, which is a continuous measure of evidence against a null requiring no alternative and no decision, with the Neyman–Pearson framework of two competing hypotheses, error-rate control, and a decision rule, two frameworks whose architects publicly disputed their incompatibility. The resulting hybrid collapses power, prior probability, effect size, and decision context into a single bright-line verdict, producing a class of characteristic errors: treating p = 0.049 and p = 0.051 as qualitatively different states of evidence, reporting "significant" effects of magnitudes too small to matter, treating non-rejection as equivalence, and publishing without specifying the alternative the test was testing. Gigerenzer documents the pattern across psychology; Ioannidis (2005) in biomedical research and McCloskey and Ziliak (2008) in econometrics describe the same pathology in their fields.
Structural Signature¶
Sig role-phrases:
- the hybrid significance test — the spliced procedure: Fisher's evidential, alternative-free p-value welded to Neyman–Pearson's two-hypothesis error-controlled decision rule, frameworks whose architects denied compatibility
- the bright-line threshold — the p < .05 cliff that converts a continuous evidence measure into a binary reject/retain verdict
- the missing inferential furniture — no stated alternative, no prior, no effect size that would matter, no decision the verdict informs — all dropped at once by the splice
- the detachment — the procedure severed from the reasoning it was built to serve: the ritual proceeds but no inference is made
- the institutional enforcement — editorial and reviewer demand that keeps the empty form mandatory and fully authoritative, performed to satisfy gatekeepers rather than to answer a question
- the characteristic symptom cluster — p = .049 vs .051 treated as different worlds, "significant" effects too small to act on, non-rejection read as equivalence: the signature of form-without-function
- the reform-resistance branch (what it predicts) — comprehension-correcting cures (teach p-value semantics) leave the form intact and fail; only incentive-restructuring changes (registered reports, mandated estimation/effect sizes, bright-line bans) move the practice
What It Is Not¶
- Not a claim that significance testing is invalid. The target is not the apparatus but its detached, ritualized execution. A significance test engaged with a stated alternative, an effect size that matters, and a decision the verdict informs is legitimate inference; the null ritual is what remains when that furniture is stripped away and only the mechanical form is performed.
- Not p-hacking or the garden of forking paths. Those are active distortions — exploiting researcher degrees of freedom to manufacture significance. The null ritual is the opposite posture: mechanical compliance with a mandatory form, performed without manipulation and often without any inferential intent at all. The mechanisms differ, even when they co-occur.
- Not a mere cognitive error. It is tempting to read it as individual misunderstanding of what p-values mean, curable by education. But the practice is institutionally enforced — performed to satisfy journals and reviewers, not because anyone is confused — which is precisely why teaching p-value semantics barely dents it and only incentive-restructuring reforms bite.
- Not just p-value misinterpretation. Misreading a p-value is one symptom, not the thing itself. The null ritual is the underlying detached procedure that generates the whole symptom cluster at once — missing alternatives, ignored effect sizes, the p = .049/.051 cliff, non-rejection read as equivalence — rather than any single misreading.
- Not Fisher's test, nor Neyman–Pearson's. The procedure ritualized is neither pure framework but an incoherent hybrid welding Fisher's evidential, alternative-free p-value to Neyman–Pearson's two-hypothesis, error-controlled decision rule — frameworks whose own architects denied their compatibility. Attributing the ritual to either founder mistakes the splice for one of its sources.
Scope of Application¶
The null ritual lives wherever null hypothesis significance testing is the editorial-default inference apparatus of scientific publishing; its reach is bounded by that substrate — one specific hybrid test inside one specific publication culture — so it surfaces across every field that shares it. (The hollow-corporate-audit and agile-ceremony analogues are the parent cargo-cult / Goodhart pattern, not the ritual itself.)
- Psychology — Gigerenzer's original diagnostic turf, where the bulk of the documented cases live: mechanical p < .05 verdicts published without an alternative, an effect size that matters, or a decision the test informs.
- Biomedical research — Ioannidis's account of why most published findings are false rests on the same detached procedure inflating false positives across the clinical and life-science literature.
- Econometrics — McCloskey and Ziliak's "cult of statistical significance" documents the identical hollowed test, "significant" coefficients reported with no attention to the magnitude that would matter for policy.
- Ecology and environmental science — Anderson, Burnham, and Thompson's vocal critique targets ritualized testing in field studies, urging model-selection and effect-size reporting instead.
- Educational research and social-policy evaluation — NHST adopted as the default inference language, the binary verdict standing in for an engaged inference.
- Machine-learning benchmark evaluation — increasingly flagged: significance bars on metric deltas reported without effect-size or decision context, the ritual migrating into a new publication culture.
Clarity¶
Naming the null ritual lets a methodologist locate the disease precisely: the problem is not that significance testing is wrong, but that the procedure has been severed from the inferential reasoning it was built to serve while keeping all its editorial authority. Without the label, the field's recurring complaints — p-value misinterpretation, missing alternatives, absent priors, ignored effect sizes, the p = .049/.051 cliff, reading non-rejection as equivalence — look like a scattered list of bad habits to be corrected one by one. The label binds them into a single diagnosis: these are not independent mistakes but symptoms of one detached ritual. That reframing sharpens the question a critic can ask of any reported test from "is the p-value computed correctly?" to "was an inference actually made here — is there an alternative hypothesis, an effect size that matters, a decision this verdict informs?"
It also makes legible something the methods literature otherwise mystifies: why reform keeps failing. Treating the ritual as a cognitive error invites educational fixes (teach people what p-values mean), yet those fixes barely dent the practice. Recognizing it as an institutionally enforced ritual — performed because journals and reviewers demand it, not because it answers the researcher's question — explains the resistance and relocates the leverage: estimation statistics, Bayesian alternatives, pre-registration, and bright-line abandonment bite only insofar as they restructure the editorial incentives that sustain the form. A further clarity is historical: by exposing the tested procedure as a hybrid welding Fisher's evidential p-value to Neyman–Pearson's two-hypothesis decision framework — whose own architects denied their compatibility — the concept reveals that the ritual collapses power, prior, and effect size into one verdict because it was assembled from two frameworks that each supplied what the other dropped, and the splice quietly discarded both halves' safeguards.
Manages Complexity¶
The methods-reform literature, absent this concept, presents as a long, miscellaneous bill of complaints, each lodged and rebutted on its own: p-values are misinterpreted; the nil-null is the wrong null; alternatives go unstated; priors are ignored; effect sizes are absent; p = .049 and p = .051 are treated as different worlds; non-rejection is read as equivalence; "significant" is reported for magnitudes no one would act on. Treated as a scattered list, each item invites its own local fix and its own running dispute over whether it is really an error, and the field cycles through them without traction. The null ritual compresses the whole list into one structural diagnosis: these are not independent bad habits but symptoms of a single detached procedure — a hybrid significance test executed for its form while severed from the inferential reasoning it was built to serve. The analyst stops enumerating defects and instead asks one question of any reported test — was an inference actually made here? — and the presence or absence of an alternative hypothesis, an effect size that matters, and a decision the verdict informs settles the diagnosis at a stroke. A sprawling symptom-catalogue collapses to one mechanism with a small, checkable signature.
The deeper compression governs prediction, and specifically the prediction of why reform fails. The space of proposed cures is large — teach p-value semantics, switch to estimation statistics, adopt Bayesian methods, pre-register, ban bright-line thresholds, mandate effect sizes — and which will bite is otherwise an empirical muddle settled cure by cure. Locating the disease as institutional rather than cognitive reduces that muddle to a single tracked parameter: does the intervention restructure the editorial and reviewer incentives that make the ritual mandatory, or does it merely correct understanding? From that one coordinate the outcome reads off directly — educational fixes leave the practice intact because the ritual is performed to satisfy gatekeepers, not because anyone misunderstands it, whereas only changes that alter what journals require can move it. The concept thus partitions the entire reform space along one branch (incentive-restructuring versus comprehension-correcting) and predicts efficacy from which side a proposal falls on, rather than requiring each reform to be tried and its failure re-explained from scratch. The historical splice supplies the closing economy: recognising the tested procedure as a weld of Fisher's evidential p-value onto Neyman–Pearson's two-hypothesis decision rule explains in one stroke why power, prior, and effect size are all missing at once — the hybrid was assembled from two frameworks each of which supplied what the other dropped, and the join discarded both halves' safeguards — so the analyst need not catalogue the missing pieces severally but derives them from a single account of how the form was built.
Abstract Reasoning¶
The diagnostic move is a test of whether an inference was actually made — reasoning past the surface verdict to the inferential furniture that should accompany it. Confronting any reported significance test, the methodologist asks not "is the p-value computed correctly?" but "was an inference made here at all?", and reads the answer off a fixed signature: is there a stated alternative hypothesis, an effect size whose magnitude would matter, a prior the analyst held, a decision the verdict informs? The characteristic inference runs from the absence of that furniture to the procedure has been severed from the reasoning it was built to serve — the ritual proceeded but the inference did not. The move binds a scattered symptom-catalogue (p = .049 versus .051 treated as different worlds, "significant" effects too small to act on, non-rejection read as equivalence, missing alternatives) into one detached-procedure diagnosis, so the analyst stops cataloguing defects severally and settles the matter with a single check for whether the form is doing inferential work.
The load-bearing move is a cognitive-versus-institutional cause attribution that, in turn, predicts which reforms will bite. Reasoning from the durability of the practice under educational correction — teaching people what p-values mean barely dents it — the analyst infers that the ritual is sustained by editorial and reviewer incentives, not by misunderstanding: it is performed to satisfy gatekeepers, not because anyone is confused. This attribution partitions the entire reform space along one branch and lets the analyst forecast efficacy from which side a proposal falls on: comprehension-correcting fixes (explain the semantics) are predicted to fail because they leave the mandatory form untouched, while incentive-restructuring changes (editorial bans, registered reports, mandated effect sizes and estimation, pre-registration) are predicted to move the practice because they alter what publication requires. The inference runs from the disease is institutional to only interventions that change the gatekeeping incentives will work — replacing a cure-by-cure empirical muddle with a derivation from one tracked coordinate.
A distinctive historical-decomposition move explains the symptom cluster from the procedure's construction rather than enumerating its defects. The analyst reasons that the tested procedure is a hybrid — a weld of Fisher's evidential, alternative-free p-value onto Neyman–Pearson's two-hypothesis, error-controlled decision rule, frameworks whose own architects denied their compatibility — and infers from that splice why power, prior, and effect size are all missing at once: each framework supplied what the other dropped, and the join quietly discarded both halves' safeguards. The inference runs from how the form was assembled to which protections it must lack, so the missing pieces are derived from a single account of the weld rather than catalogued one by one.
Underwriting these is a boundary-drawing move that keeps the critique precise. The analyst reasons that the target is not significance testing as such — the apparatus is legitimate when an alternative, an effect size, and a decision are genuinely engaged — but the ritualized, detached execution of it, and separates this failure from adjacent ones it is easily fused with: it is not active gaming of researcher degrees of freedom (that is a different, distortion-driven mechanism) but mechanical compliance with a mandatory form. The inference is from the procedure became socially obligatory while detached from purpose to this is form-without-function under institutional enforcement, drawing the line that locates the disease in the practice's social role rather than in the arithmetic or in any individual's intent.
Knowledge Transfer¶
Within its home substrate — null hypothesis significance testing as the editorial-default inference apparatus of scientific publishing — the null ritual transfers as mechanism, and transfers across every field that shares that apparatus. Gigerenzer diagnosed it in psychology; Ioannidis describes the same detached procedure inflating false findings in biomedical research; McCloskey and Ziliak document it in econometrics; Anderson, Burnham, and Thompson attack the identical ritualized testing in ecology and environmental science; it is named in educational research and social-policy evaluation; and it is increasingly flagged in machine-learning benchmark evaluation, where significance bars on metric deltas are reported without effect-size context. Across all of these the transfer is literal because the substrate is literally the same — one specific hybrid test deployed within one specific publication culture — so the diagnostic signature carries untranslated (no stated alternative, no effect size that would matter, no prior, no decision the verdict informs; the p = .049/.051 cliff; non-rejection read as equivalence), and so does the causal account (an institutionally enforced form, not a comprehension failure) and the consequent prediction about reform: only interventions that restructure editorial incentives bite, which is why the field-internal cures — registered reports, mandated estimation statistics and effect sizes, pre-registration, bright-line abandonment, outright NHST bans of the kind Basic and Applied Social Psychology imposed in 2015 — are the ones that move the practice while teaching p-value semantics does not. The vocabulary and the remedies are NHST-and-publishing furniture, and they travel exactly as far as that furniture is in place, which is to say across all these fields at once.
Beyond that substrate the transfer is a shared abstract mechanism that belongs to the parent, not to the named ritual. What genuinely recurs across domains is form-without-function sustained by institutional enforcement — a procedure that keeps its full social and gatekeeping authority after being severed from the reasoning it was built to serve — which is the prime cargo cult (with Goodhart's law and Scott-style legibility as close kin: a once-meaningful form made mandatory and then hollowed). That general pattern co-instantiates widely — a compliance audit performed for the checklist rather than for safety, an agile ceremony run for ritual rather than coordination, an operations-research benchmark cited for ceremony rather than decision — and the cross-domain lesson (when a procedure becomes obligatory while detached from purpose, look to the incentives that make it mandatory, not to anyone's misunderstanding) is carried by cargo cult, not by "null ritual." The home-bound cargo is everything that makes the null ritual specifically itself: the Fisher / Neyman–Pearson hybrid and its incompatible splice, the nil-null, the p < .05 bright line, the missing power-prior-effect-size triad, and the estimation-and-pre-registration reform menu — none of which has a referent where there is no significance test. So calling a hollow corporate audit "a null ritual" is analogy: it borrows the form-without-function shape while dropping the statistical machinery that gives the original its diagnostic precision, and the honest move is to let the cross-domain weight ride on the cargo-cult/Goodhart pattern the ritual instantiates, reserving "null ritual" — and its NHST-specific diagnostics and cures — for the inference-practice substrate where they literally apply (see Structural Core vs. Domain Accent).
Examples¶
Canonical¶
Gigerenzer's defining illustration is the three-step liturgy he calls the null ritual: (1) set up a nil-null hypothesis of zero effect, (2) compute p, (3) if p < .05 reject the nil and report the result as "significant" — full stop. The archetypal published instance is a line like "the groups differed, t(38) = 2.1, p < .05," with no alternative hypothesis stated, no effect size reported, no prior considered, and no decision the verdict feeds. Gigerenzer pairs this with survey evidence that the ritual is performed without comprehension: Haller and Krauss (2002), replicating Oakes, found that not only students but a majority of the psychology methodology instructors they surveyed endorsed at least one demonstrably false interpretation of a significant p-value (e.g., that p < .05 gives the probability the null is true).
Mapped back: The t-test executed to its p < .05 verdict is the hybrid significance test run to its bright-line threshold; the absent alternative, effect size, prior, and decision are the missing inferential furniture; and the reported line is the detachment made visible — the form completed while no inference was made. The instructors' endorsement of false interpretations shows the ritual proceeding by rote, the characteristic symptom cluster of form-without-function.
Applied / In Practice¶
In 2015 the journal Basic and Applied Social Psychology, in an editorial by Trafimow and Marks, banned null hypothesis significance testing outright: authors could no longer report p-values, significance statements, or the reject/retain verdict, and were pushed toward descriptive statistics and effect sizes. Whatever its merits as statistics, the move is a clean natural experiment on the null ritual's causal account. Editors judged that decades of teaching p-value semantics had not dislodged the practice, and that only removing the gatekeeping requirement — changing what publication demanded — could. The ban targeted the institutional enforcement rather than researchers' understanding, exactly where the concept predicts leverage lies.
Mapped back: The ban is an intervention on the institutional enforcement — the editorial demand that keeps the empty form mandatory — rather than a comprehension fix. It sits on the incentive-restructuring side of the reform-resistance branch, which predicts that only changes to what journals require move the practice, while teaching the semantics of the hybrid significance test and its bright-line threshold leaves the detachment intact.
Structural Tensions¶
T1: Legitimate apparatus versus empty ritual (two identical-looking tests, one inference and one liturgy). The critique is careful not to condemn significance testing as such — a test engaged with a stated alternative, an effect size that matters, and a decision it informs is valid inference. The null ritual is only what remains when that furniture is stripped and the mechanical form is performed for its own sake. But this means the disease is invisible in the arithmetic: two papers can report the identical line "t(38) = 2.1, p < .05" while one made a genuine inference and the other performed a rote liturgy, and nothing in the computed statistic distinguishes them. The diagnosis rests entirely on absent context — was there an alternative, a magnitude that matters, a decision — which the published artifact may not record. The concept's precision (it targets execution, not apparatus) is also its evidentiary difficulty (the target is a use, not a mark on the page). Diagnostic: Setting the arithmetic aside, is there a stated alternative, an effect size whose magnitude would change a decision, and a decision the verdict feeds — or only the verdict?
T2: Institutional enforcement versus genuine misunderstanding (the attribution that locates leverage also removes culpability — and may overstate its case). The load-bearing move relocates the cause from cognition to incentives: the ritual persists not because researchers are confused but because gatekeepers require the form, which is why education fails and only incentive-restructuring bites. That attribution is what makes reform predictable. Yet the entry's own evidence — Haller and Krauss finding methodology instructors endorsing false interpretations of p-values — shows genuine comprehension failure coexisting with the institutional form, so "not a cognitive error" is an emphasis, not a clean separation. And locating the disease in the system risks absolving the individuals who nonetheless perform it and could resist. The tension is that a purely institutional diagnosis correctly identifies where leverage lives while understating both the real misunderstanding present and any agent's responsibility for enacting the form. Diagnostic: In this instance, would the practice survive if the editorial requirement vanished (institutional) or would the researcher still misread the p-value (cognitive) — and does treating it as purely institutional let the enactor off the hook?
T3: Mandatory standardization as coordination good versus its hollowing (the bright line is enforceable because it is useful). NHST became the editorial default not by accident but because a bright-line, mechanical verdict supplies real goods: a common inferential currency across a heterogeneous field, a low-cost gatekeeping filter reviewers can apply uniformly, and comparability across papers. The ritual is the dark side of a genuine standardization. This complicates the reform story: abandoning the p < .05 threshold (a predicted-to-bite intervention) also forgoes the coordination the threshold provided, and estimation or Bayesian alternatives that resist a single bright line are harder to enforce uniformly — which is part of why the mandatory form was stable. The very institutionalization that hollowed the test is also what let a sprawling literature share one decision rule. The tension is that the feature enabling the ritual (a cheap, uniform, mandatory verdict) is also a real coordination benefit its removal sacrifices. Diagnostic: Does abandoning the bright line here replace it with something reviewers can apply as uniformly, or does it trade a hollow standard for an unenforceable one?
T4: One detached mechanism versus a catalogue of separable defects (economy of diagnosis versus over-lumping). The concept's power is compression: the p = .049/.051 cliff, missing effect sizes, unstated alternatives, and non-rejection-read-as-equivalence are bound into symptoms of one detached procedure, so the analyst checks "was an inference made?" once instead of litigating each defect. But some of these errors have independent life — a researcher can misread a single p-value while otherwise reasoning inferentially, or omit an effect size for a bad reason unrelated to ritual — so binding the whole cluster to one mechanism risks explaining distinct failures with a single story and mistaking co-occurrence for common cause. The unifying economy that gives the concept traction is also a pressure to attribute every NHST pathology to the ritual. The tension is between the diagnostic parsimony of one detached procedure and the fidelity of recognizing that its symptoms can occur singly and for their own reasons. Diagnostic: Do the symptoms here co-occur as a full cluster pointing to a detached form, or is a single isolated defect present that has its own, non-ritual explanation?
T5: Restructuring incentives versus relocating the ritual (Goodhart on the cure). The reform prediction — only interventions that change what publication requires will move the practice — is vindicated by bans like BASP's. But the concept's own parent, cargo cult, warns that any newly mandated form is itself a candidate for hollowing: require effect sizes and they can be reported by rote; mandate estimation intervals or Bayesian priors and they can be produced to satisfy reviewers without inference, exactly as p-values were. Restructuring the gatekeeping incentive does not guarantee inference returns; it may install a new obligatory form that detaches from purpose in its turn. The reform that bites hardest — a mandatory replacement — is also the one most exposed to becoming the next null ritual, while the softer educational fixes that do not bite at least cannot be gamed into a fresh liturgy. Diagnostic: Does the proposed reform restore the inferential furniture, or merely swap one mandatory form for another that can be performed just as mechanically?
T6: Autonomy versus reduction (a named NHST pathology or the inference-practice instance of cargo cult). "Null ritual" is a named, historically specific diagnosis with irreducibly local cargo — the Fisher / Neyman–Pearson hybrid and its disputed splice, the nil-null, the p < .05 bright line, the missing power-prior-effect-size triad, the estimation-and-pre-registration reform menu — and within the NHST-and-publishing substrate it transfers as literal mechanism across psychology, biomedicine, econometrics, ecology, policy evaluation, and ML benchmarking, because that substrate is literally the same everywhere it appears. But beyond that substrate it does not travel as mechanism: what genuinely recurs — form-without-function sustained by institutional enforcement — is the prime cargo_cult (with Goodhart's law and legibility as kin), and calling a hollow corporate audit "a null ritual" is analogy that drops the statistical machinery. The tension is between a diagnosis that earns its own name through NHST-specific detail and the recognition that its cross-domain lesson belongs to cargo cult. Diagnostic: Resolve toward cargo_cult / Goodhart when the point is form-without-function outside significance testing; toward the named null ritual when diagnosing detached NHST practice within a publication culture that runs it.
Structural–Framed Character¶
The null ritual sits near the framed pole of the structural–framed spectrum — framed-leaning, close to the pole because it is a normatively charged critique of a hollowed institutional practice, held just short of it only by the genuine descriptive-mechanism content in its account of the Fisher/Neyman–Pearson splice. On evaluative_weight it scores high: "null ritual" is a diagnosis of malpractice — form-without-function, an empty liturgy performed for gatekeepers — a verdict that a practice has gone wrong, not a neutral description of a mechanism. On human_practice_bound it is maximal: the ritual is constituted by null hypothesis significance testing as the editorial-default inference apparatus of scientific publishing, and dissolves entirely without that practice — remove the journals, reviewers, and the mandatory p < .05 form, and there is no ritual, only whatever inference a researcher chooses to make. Institutional_origin is pronounced and definitional: the concept is Gigerenzer's diagnosis of an institutionally enforced form, and its own causal claim — that the practice persists through editorial incentives rather than misunderstanding — makes institutional constitution the load-bearing fact, with the tested procedure itself a historically contingent hybrid of two disputed frameworks. On vocab_travels it scores low: the nil-null, the bright line, the power-prior-effect-size triad, and the estimation/pre-registration reform menu have no referent where there is no significance test. And on import_vs_recognize the transfer is bimodal — within the NHST-publishing substrate it ports as literal mechanism across psychology, biomedicine, econometrics, ecology, policy, and ML benchmarking, but a hollow corporate audit is a co-instance of a shared parent, not an import of "null ritual."
The one portable structural skeleton is cargo_cult — form-without-function sustained by institutional enforcement, a once-meaningful procedure made mandatory and then severed from the reasoning it served — with Goodhart's law and legibility as close kin. That skeleton is genuinely substrate-independent and recurs as co-instance in compliance audits run for the checklist and agile ceremonies run for ritual. But it does not pull the null ritual off the framed pole, because the cargo-cult pattern is exactly what the ritual instantiates from its umbrella, not what makes "null ritual" itself travel: the cross-domain reach belongs to cargo_cult, while the NHST hybrid, the p < .05 bright line, the missing inferential furniture, and the field-specific reform menu stay home. Its character: a normatively charged, institutionally-constituted critique of a hollowed inference practice, structural only in the cargo-cult skeleton it borrows from its umbrella and specializes to detached significance testing.
Structural Core vs. Domain Accent¶
This section decides why the null ritual is a domain-specific abstraction and not a prime — why its cross-domain weight rides on a parent while its statistical machinery stays home.
What is skeletal (could lift toward a cross-domain prime). Strip the statistics and a thin relational structure survives: a once-meaningful procedure keeps its full social and gatekeeping authority after being severed from the reasoning it was built to serve, sustained not by anyone's misunderstanding but by institutional enforcement that makes the empty form mandatory. The portable pieces are abstract — a form detached from its function, retained authority, and an incentive structure that keeps the hollow form compulsory. This skeleton is genuinely substrate-portable, which is why the catalog carries it as the parent cargo_cult the null ritual instantiates (form-without-function made mandatory and then hollowed), with Goodhart's law and legibility as close kin. But it is the core the null ritual shares with a compliance audit run for the checklist or an agile ceremony run for ritual, not what makes it the distinctive thing it is.
What is domain-bound. Almost everything that makes the diagnosis the null ritual in particular is NHST-and-publishing furniture and none of it survives extraction. The hybrid significance test — Fisher's evidential, alternative-free p-value welded to Neyman–Pearson's two-hypothesis, error-controlled decision rule, frameworks whose architects denied compatibility; the nil-null hypothesis; the p < .05 bright line; the missing power-prior-effect-size triad that the splice drops all at once; the characteristic symptom cluster (p = .049 vs .051 treated as different worlds, "significant" effects too small to act on, non-rejection read as equivalence); and the estimation-and-pre-registration reform menu (registered reports, mandated effect sizes, bright-line bans) — these have no referent where there is no significance test. The decisive test: a hollow corporate audit has a form-without-function, but no p-value, no nil-null, no Fisher/Neyman–Pearson weld — remove the significance test and there is no null ritual in particular, only the bare cargo-cult parent.
Why this does not clear the prime bar. A prime's vocabulary travels and its transfer is recognition of the same mechanism, not analogy. The null ritual's transfer is bimodal. Within the NHST-publishing substrate it moves as literal mechanism across psychology, biomedicine, econometrics, ecology, policy evaluation, and ML benchmarking, because that substrate — one specific hybrid test inside one specific publication culture — is literally the same everywhere it appears, so the diagnostic signature, the institutional-cause account, and the reform prediction (only incentive-restructuring bites) carry untranslated. Beyond it the transfer is analogy: calling a hollow corporate audit or an empty agile ceremony "a null ritual" borrows the form-without-function shape while dropping the statistical machinery that gives the original its diagnostic precision. The genuinely portable structure is not the null ritual but its parent cargo_cult (with Goodhart's law and legibility), of which those hollow-procedure cases are fellow co-instances. So the cross-domain reach belongs to the parent; the disciplined move is to let the weight ride on the cargo-cult/Goodhart pattern outside significance testing, and reserve "null ritual" — and its NHST-specific diagnostics and cures — for the inference-practice substrate where they literally apply. It clears the domain-specific bar comfortably for statistical inference practice, but its only substrate-spanning content is already carried, in more general form, by the pattern it instantiates.
Relationships to Other Abstractions¶
Current abstraction Null Ritual Domain-specific
Parents (2) — more general patterns this builds on
-
Null Ritual is a kind of Ritual Prime
The Null Ritual is a prescribed, recurrent, socially enforced symbolic procedure specialized to mechanical significance testing.It inherits Ritual's fixed sequence, repetition, community enforcement, and standing-conferring performance, then fixes the sequence to nil-null, p-value, threshold, and verdict within scientific publishing. Its emptiness is a pathology of the ritual's relation to inferential purpose, not evidence that it is not a ritual.
-
Null Ritual is part of Hypothesis Testing (Null vs. Alternative) Prime
The Null Ritual contains a mechanically executed null-hypothesis test whose inferential furniture has been stripped away while its verdict remains authoritative.The nil null, p-value, bright-line threshold, and reject-or-retain decision are internal operations of the ritual. Removing that test leaves generic hollow procedure rather than the statistically specific diagnosis.
Hierarchy paths (8) — routes to 8 parentless roots
- Null Ritual → Ritual → Performativity
- Null Ritual → Ritual → Recurrence
- Null Ritual → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Inductive Reasoning
- Null Ritual → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Uncertainty
- Null Ritual → Ritual → Social Norms → Normativity → Constraint
- Null Ritual → Hypothesis Testing (Null vs. Alternative) → Verification → Evaluation → Comparison → Self Checking
- Null Ritual → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Probability → Measure → Set and Membership
- Null Ritual → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
Not to Be Confused With¶
-
Cargo cult (the parent prime). The substrate-general pattern of form-without-function sustained by institutional enforcement — a once-meaningful procedure made mandatory and then severed from the reasoning it served. The null ritual is the significance-testing instance of
cargo_cult. The parent travels to a hollow compliance audit or an empty agile ceremony; the NHST machinery does not. Tell: is there a p-value, a nil-null, and a Fisher/Neyman–Pearson weld (null ritual), or just a hollowed obligatory form in general (cargo cult)? (Treated more fully as the umbrella it instantiates in Structural Core vs. Domain Accent.) -
Goodhart's law. The kin principle that a measure ceases to be a good measure once it becomes a target. It illuminates why p < .05 hollows out (the threshold became the target of publication), but Goodhart is about metric-gaming in general, while the null ritual is the specific detached-NHST practice. Tell: is the point that optimizing to a metric corrupts it (Goodhart) or that a whole inferential procedure is performed by rote for gatekeepers (null ritual)?
-
p-hacking / garden of forking paths. Active distortions — exploiting researcher degrees of freedom to manufacture significance. The null ritual is the opposite posture: mechanical compliance with a mandatory form, performed without manipulation and often without inferential intent at all. Tell: is significance being actively engineered through flexible analysis (p-hacking) or a hollow form being executed by rote (null ritual)?
-
p-value misinterpretation. A single symptom — misreading what a p-value means (e.g. as the probability the null is true). The null ritual is the underlying detached procedure that generates the whole symptom cluster at once, not any one misreading. Tell: is the object one wrong interpretation of a statistic (misinterpretation) or the mechanical, purpose-severed execution that produces many such errors together (null ritual)?
-
Significance testing proper. The legitimate apparatus — a test engaged with a stated alternative, an effect size that matters, and a decision the verdict informs. The null ritual is only what remains when that furniture is stripped and the mechanical form is performed for its own sake; the arithmetic is identical, the inference is not. Tell: is an inference actually being made (legitimate NHST) or only the reject/retain verdict produced with no alternative, effect size, or decision (null ritual)?
-
Fisher's significance test / Neyman–Pearson decision theory. The two source frameworks the ritual welds — Fisher's evidential, alternative-free p-value and Neyman–Pearson's two-hypothesis, error-controlled decision rule, whose architects denied their compatibility. The null ritual is the incoherent splice, not either pure framework; attributing it to one founder mistakes the weld for a source. Tell: is the procedure a coherent framework used as its architect intended (Fisher or Neyman–Pearson) or the mandatory hybrid that discards both halves' safeguards (null ritual)?
-
The replication crisis. The broad phenomenon of published findings failing to replicate across science. The null ritual is one contributing mechanism (inflating false positives via detached testing), not the crisis itself, which also involves publication bias, low power, and fraud. Tell: is the subject the field-wide failure-to-replicate (the crisis) or specifically the hollowed-NHST practice feeding it (null ritual)?
Neighborhood in Abstraction Space¶
Null Ritual sits in a moderately populated region (49th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Fallacious Substitution in Argument (17 abstractions)
Nearest neighbors
- Cherry Picking — 0.86
- Suppressed Evidence — 0.84
- Jeffreys-Lindley Paradox — 0.84
- Benjamini–Hochberg Procedure — 0.83
- Congruence Bias — 0.83
Computed from structural-signature embeddings · 2026-07-12