Skip to content

Jeffreys-Lindley Paradox

The result that a frequentist significance test and a Bayesian posterior-odds comparison of the same data against the same point null can reach opposite verdicts, with the disagreement growing without bound as sample size increases — because the two answer different questions.

Core Idea

The Jeffreys-Lindley paradox (Harold Jeffreys, Theory of Probability, 1939; Dennis Lindley, "A Statistical Paradox," Biometrika 44, 1957) is the result that a frequentist significance test and a Bayesian posterior-odds comparison of the same data against the same point-null hypothesis can give diametrically opposed verdicts, and that this disagreement grows without bound as sample size increases. The canonical setup: a point-null hypothesis \(H_0\) specifies \(\theta = \theta_0\); the alternative \(H_1\) assigns \(\theta\) a diffuse prior (e.g., \(\theta \sim \text{Normal}(0, \sigma^2)\) with large \(\sigma^2\)); data \(D\) are observed with a test statistic just past the conventional significance threshold (say, \(z \approx 1.96\), \(p \approx 0.05\)). The p-value is small: \(H_0\) is rejected at the conventional level. The Bayes factor in favour of \(H_0\), however, is high — often substantially greater than 1 — because the diffuse prior on \(H_1\) spreads its probability mass across a wide range of \(\theta\) values, and the concentrated observation near \(\theta_0\) is much more probable under \(H_0\) than under the average parameter value the prior assigns to \(H_1\). As \(n \to \infty\) with the effect size held fixed, the p-value falls (making frequentist rejection easier) while the Bayes factor in favour of \(H_0\) grows (making Bayesian confirmation of \(H_0\) stronger) — the two verdicts diverge in opposite directions.

The disagreement is structural, not a calculation error. The p-value is a tail-area probability under \(H_0\) — it answers "how probable is a test statistic at least this extreme, if \(H_0\) is true?" and grows with \(n\) because estimation precision sharpens the test. The Bayes factor is an integrated likelihood ratio — it answers "how much more probable are these data under \(H_0\) than under \(H_1\), averaged over the prior?" and is sensitive to the prior's spread because a diffuse \(H_1\) pays a heavy Occam penalty when the data are concentrated near \(\theta_0\). The two procedures answer different questions, and large-\(n\) data with near-null effect sizes make those questions maximally divergent. The paradox is the standard proof that the choice between frequentist and Bayesian inference is not a matter of computational convenience but of the question being asked — and that the same dataset can simultaneously constitute frequentist grounds for rejection and Bayesian grounds for confirmation of the null.

Structural Signature

Sig role-phrases:

  • the point null — H₀ specifying θ = θ₀ exactly, a measure-zero hypothesis
  • the diffuse alternative — H₁ assigning θ a broad prior that spreads its mass across a wide parameter range
  • the large sample — data of size n big enough that the test statistic lands just past threshold (z ≈ 1.96, p ≈ 0.05) at a near-null effect
  • the two procedures — a frequentist tail-area p-value and a Bayesian prior-averaged Bayes factor, each answering a different question
  • the opposed verdicts — the engineered finding: the p-value rejects H₀ while the Bayes factor favors H₀, because the diffuse H₁ pays a heavy Occam penalty on concentrated data
  • the monotone-in-n divergence — as n grows with the effect fixed, the two verdicts move further apart rather than reconciling, the structural (not computational) heart of the result
  • the dissolving levers — its characteristic remedies: tighten the alternative prior, replace the point null with an interval null, or pre-register effect sizes — each shrinking the divergence the configuration manufactures

What It Is Not

  • Not a calculation error or a sign one procedure is wrong. The opposed verdicts are structural: the p-value (a tail-area probability under H₀) and the Bayes factor (a prior-averaged likelihood ratio) answer different questions, and both are computed correctly. Neither procedure has blundered; they diverge because they measure different things.
  • Not a disagreement that more data resolves. Adding data, with the effect size fixed, widens the gap rather than closing it — the p-value falls while the Bayes factor swings toward H₀. This inverts the naive expectation that more evidence brings agreement; here the divergence is monotone in n.
  • Not a verdict that Bayesian beats frequentist (or vice versa). The paradox does not crown a winner; it shows the choice of method is a choice of which question to ask — tail-area improbability under H₀ versus averaged relative likelihood of H₀ against H₁. It is an argument that the two are distinct, not that one is superior.
  • Not a logical contradiction. "Paradox" here names a clash of intuitions, not an inconsistency in the mathematics. There is nothing self-contradictory: a single dataset can be genuine frequentist grounds for rejection and Bayesian grounds for confirming the null, each coherent within its own framework.
  • Not merely an artifact of a poorly chosen prior. The prior's diffuseness is one lever, but the configuration also depends on testing a measure-zero point null; replacing it with an interval null dissolves much of the divergence. Blaming the prior alone misses that the strict point hypothesis is half the structure.
  • Not "significance equals weak evidence" as a general rule. The paradox is specific to large n with a near-null effect and a diffuse alternative; it does not say p ≈ 0.05 is always evidentially empty. At modest sample sizes with informative priors the two procedures broadly agree — the divergence is a property of the configuration, not of significance testing everywhere.

Scope of Application

The Jeffreys–Lindley paradox is a result about statistical inference, so its reach is wide but single-substrate: it surfaces in every field that does null-hypothesis testing with large samples, but each is importing the same inferential machinery (point-null testing, posterior odds, sample size), not re-instantiating an independent pattern. The broader lesson — two procedures answering different questions can diverge, and more data widens the gap — is the parent incommensurability (with the general measurement/inference framing), not the named paradox, and stays out of this map.

  • Bayesian statistics — the canonical motivating example for the p-value crisis literature, objective/non-local priors, and the proposal to redefine significance thresholds.
  • Frequentist statistics — cited as the standard warning that p ≈ 0.05 at large n carries negligible evidential weight, motivating effect-size reporting and confidence-interval-first practice.
  • Model selection — it surfaces whenever a small point-null model is compared to a parameter-rich alternative under diffuse priors, explaining why Bayes factors and AIC/BIC favor smaller models that p-tests reject.
  • Psychology and the replication crisis — many "failed" replications are Lindley-paradoxical originals (large n, near-null effect, p just past threshold), exposed rather than contradicted.
  • Particle physics — the choice of 5σ discovery thresholds rather than 3σ is partly a device to avoid Lindley-style false alarms in very large experiments.
  • Genomics and high-throughput science — multiple-testing corrections, empirical-Bayes shrinkage, and local false-discovery-rate methods are partly responses to the paradox at scale.

Clarity

Naming the paradox makes legible a separation that routine testing practice silently collapses: statistical significance and evidential weight are not the same quantity, and at large \(n\) they can point in opposite directions. Without the label, a result with \(p \approx 0.05\) on a huge sample reads as a clean rejection of the null, and a Bayes factor mildly favoring the null reads as a contradiction or a mistake. With it, the analyst recognizes a structural configuration — large sample, near-null effect, \(p\) just past threshold — and asks the sharp diagnostic question the surrounding literature now treats as standard: is this a genuine finding, or a Lindley-paradoxical artifact in which precision has been mistaken for importance? The same configuration explains a large share of the replication crisis's "failed" replications: many are not contradicting their originals so much as exposing them as small-effect-large-sample rejections carrying negligible evidential force.

The paradox also forces two distinctions that informal \(p\)-test practice tends to bury. First, it converts prior specification from an invisible default into an explicit, consequential decision: the disagreement is driven by how diffuse the alternative's prior is, so one can no longer pretend the choice of \(H_1\) is innocuous — a diffuse alternative pays an Occam penalty that a concentrated one does not. Second, it sharpens the difference between a strict point null (\(\theta = \theta_0\) exactly) and an interval null (\(|\theta - \theta_0| < \delta\)), a distinction usually elided, by showing that the paradox is largely an artifact of testing a measure-zero point hypothesis. The clarifying upshot is that the frequentist-versus-Bayesian choice is exposed as a substantive question about which question one is asking — tail-area improbability under \(H_0\) versus averaged relative likelihood of \(H_0\) against \(H_1\) — rather than a matter of computational taste, and that more data magnifies, rather than resolves, the divergence between those two questions.

Manages Complexity

Without the paradox as a named configuration, the cases it covers present as a scatter of separate puzzles, each demanding its own ad hoc resolution: a headline-significant drug trial whose effect is clinically negligible; a Bayes factor that mildly favors the null on data a \(p\)-test rejects; a "failed" replication that seems to contradict a published finding; a Bayes-factor-versus-AIC/BIC disagreement in model selection. Faced one at a time, each invites a wrong story — a calculation error, a contradiction between schools, a fluke, a methodological failing — and each gets re-litigated from scratch. The paradox compresses that whole family into a single recognizable structural signature with a small fixed parameter set. The analyst no longer asks open-endedly "why do these two analyses disagree, and which is wrong?" but checks three things: is the sample large, is the alternative's prior diffuse, and is the rejection just past threshold? When all three hold, the divergence is not an anomaly to be explained but the predictable behavior of the configuration, and the qualitative verdict — that significance here carries little evidential weight — is read off rather than re-derived.

The compression rests on a clean account of why the two procedures move as they do, which turns the disagreement into something monotone and trackable rather than mysterious. The \(p\)-value is a tail-area probability under \(H_0\) and shrinks with \(n\) as estimation sharpens; the Bayes factor is a prior-averaged likelihood ratio and, because a diffuse \(H_1\) pays a heavy Occam penalty when the data land near \(\theta_0\), swings toward \(H_0\) as \(n\) grows with the effect fixed. So the single dial of sample size predicts the direction of divergence — more data magnifies it rather than resolving it — and the analyst reads the qualitative outcome off that relationship instead of recomputing each case. The same recognition collapses several distinctions that informal practice keeps tangled into one decidable check: significance versus evidential weight, strict point null versus interval null, diffuse versus concentrated alternative prior. A high-dimensional interpretive problem — adjudicate, per result, between two inferential verdicts and decide whether a finding is real — reduces to locating the result by three parameters within a known structure whose behavior is already characterized, with the genuine-versus-artifact reading following directly.

Abstract Reasoning

The paradox licenses a sharp diagnostic move on any disagreeing test result: from the joint signature "\(p\) just past threshold, Bayes factor near or favoring the null, large sample, diffuse alternative," infer that the result is Lindley-paradoxical — that precision has been mistaken for importance — rather than that one procedure has erred or the two schools genuinely contradict. The reasoner goes FROM "a hugely-powered study rejects at \(p \approx 0.05\) yet the Bayes factor mildly supports \(H_0\)" TO "this is a small-effect-large-sample rejection carrying negligible evidential weight," and treats the configuration as a signature to be matched, not an anomaly to be explained anew. The same move re-reads the replication literature: confronted with a "failed" replication, the reasoner asks whether the original was Lindley-paradoxical, predicting that many such replications are exposing precision-driven rejections rather than contradicting real effects.

The predictive / order-of-events move is the monotone-in-\(n\) claim, which the paradox makes the reasoner reach for first. Holding the effect size fixed, the reasoner predicts that adding data widens the divergence rather than closing it: the \(p\)-value falls (rejection gets easier) while the Bayes factor swings toward \(H_0\) (because a diffuse \(H_1\) pays a heavier Occam penalty as the data concentrate near \(\theta_0\)). So from "we collected ten times more data" the reasoner predicts the two verdicts move apart, inverting the naive expectation that more evidence brings agreement — and this directional prediction is what flags large-\(n\) near-null settings as danger zones before any computation.

The interventionist moves follow from locating the paradox's two levers — the prior's diffuseness and the null's geometry. Because the divergence is driven by how spread the alternative's prior is, the reasoner predicts that tightening the alternative prior toward plausible effect sizes (or specifying it in advance) shrinks the Occam penalty and pulls the Bayesian verdict back toward the frequentist one, so prior specification becomes a consequential decision to be made deliberately rather than defaulted. Because the paradox is largely an artifact of testing a measure-zero point null, the reasoner predicts that replacing the point null with an interval null (\(|\theta - \theta_0| < \delta\)) dissolves much of the divergence — a concrete redesign whose effect the concept forecasts. And at the design stage the reasoner predicts that pre-registering effect sizes and priors forecloses the post-hoc diffuse-prior choice that manufactures the paradox.

The deepest move is boundary-drawing on what the verdicts mean: the paradox forces the reasoner to treat "frequentist or Bayesian?" not as computational taste but as a choice of which question is being asked — tail-area improbability under \(H_0\) versus prior-averaged relative likelihood of \(H_0\) against \(H_1\). So before adjudicating a disagreement the reasoner asks which question the decision actually needs answered, and refuses to read significance as evidential weight or evidential weight as significance, because the paradox proves they are distinct quantities that large \(n\) can drive in opposite directions.

Knowledge Transfer

The Jeffreys–Lindley paradox is a result about statistical inference, and within that home it transfers as the full result — the same diagnostic signature (large \(n\), diffuse alternative prior, rejection just past threshold), the same monotone-in-\(n\) divergence, the same levers (tighten the prior, replace the point null with an interval null, pre-register effect sizes), and the same boundary-drawing (significance is not evidential weight) all carry intact. Its reach across fields is wide but, crucially, single-substrate: biomedical trials, particle physics (5σ discovery thresholds are partly a Lindley-avoidance device), genomics and high-throughput science (multiple-testing corrections, empirical-Bayes shrinkage, local false-discovery-rate methods), psychology and the replication crisis (many "failed" replications are exposing precision-driven large-\(n\) rejections rather than contradicting real effects), economics, and ecology all encounter it — but every one of them is importing the same inferential machinery (probability, point-null testing, posterior odds, sample size). The data come from many fields; the substrate is one. So the paradox is literally the same result wherever null-hypothesis testing with large samples is done, and the diagnostics and interventions apply directly there.

The honest limit is that the paradox does not transfer outside that substrate at all: it is a result about statistical inference, not a phenomenon in any substantive domain, so there is no biological, social, or physical system in which "the Jeffreys–Lindley paradox" occurs except by way of the statistics applied to it. What genuinely generalizes is case (B) — two broader commitments the paradox instantiates, carried by parent abstractions rather than by the named result. The first is that two valid procedures answering different questions can disagree (tail-area improbability under \(H_0\) versus prior-averaged relative likelihood of \(H_0\) against \(H_1\)) — an instance of incommensurability between distinct measurement/inference questions. The second is that more data magnifies the divergence between answers to different questions rather than reconciling them — a structural lesson about comparing non-commensurable quantities. Those carry anywhere two differently-posed evaluations are confused for one; the paradox's specific cargo — the point null, the diffuse-prior Occam penalty, the \(p\)-versus-Bayes-factor monotonicity — stays inside formal statistical inference. So the honest cross-domain lesson should carry incommensurability (and the general measurement/inference framing), not "Jeffreys–Lindley" exported as a structural pattern, whose technical objects are home-bound to the inferential machinery. See Structural Core vs. Domain Accent.

Examples

Canonical

Lindley's 1957 demonstration is the defining instance. Take a point null H₀: θ = θ₀ against an alternative H₁ that places a diffuse prior on θ, and imagine collecting a sample so large that the observed mean lands exactly at the conventional boundary — a test statistic z ≈ 1.96, p ≈ 0.05. The frequentist rejects H₀ at the 5 percent level. But the Bayes factor tells the opposite story: because H₁ spread its prior mass over a wide range of θ while the data concentrated tightly near θ₀, the observation is far more probable under the sharp null than under the average alternative, so the Bayes factor favors H₀. The decisive feature is the scaling. Hold z fixed at 1.96 (so the actual effect shrinks like 1/√n as n grows) and the Bayes factor in favor of H₀ grows roughly in proportion to √n. So multiplying the sample a hundredfold, while keeping p pinned at 0.05, multiplies the evidence for the null by about ten — the two verdicts march apart without limit.

Mapped back: θ = θ₀ is the point null and the broad prior the diffuse alternative; the sample tuned to z ≈ 1.96 is the large sample. The tail-area p-value and the prior-averaged Bayes factor are the two procedures, and rejection alongside confirmation of H₀ is the opposed verdicts. That the √n scaling drives them apart as data accrue is the monotone-in-n divergence, the structural core of the result.

Applied / In Practice

Particle physics adopted its famously stringent "5σ" discovery threshold partly as a defense against Lindley-style false alarms. Experiments at facilities like the Large Hadron Collider analyze enormous datasets, and at such sample sizes a conventional p ≈ 0.05 (roughly 2σ) or even 3σ excess is exactly the regime where frequentist significance decouples from genuine evidential weight — a tiny, precisely-measured fluctuation can clear a modest significance bar while carrying little real support for a new particle. Requiring a five-standard-deviation excess (p ≈ 3×10⁻⁷) before claiming discovery, as the field did for the 2012 Higgs boson announcement, raises the bar so far that precision alone cannot manufacture a "discovery," aligning the frequentist verdict much more closely with what a Bayesian evidence assessment would license.

Mapped back: The LHC's massive datasets are the large sample where the paradox bites hardest, and a bare 2–3σ excess is the p-just-past-threshold half of the opposed verdicts. The 5σ convention is one of the dissolving levers in institutional form — raising the rejection threshold so that the monotone-in-n divergence cannot turn mere estimation precision into a spurious claim of a real effect.

Structural Tensions

T1: Precision mistaken for importance versus real small effects dismissed (the diagnostic that can over-fire). The paradox's chief service is a diagnostic: a large-n rejection at p ≈ 0.05 with a Bayes factor favoring the null is flagged as precision masquerading as importance, negligible evidential weight rather than a finding. This correctly deflates many headline-significant, clinically trivial results. But the same signature can be weaponized to dismiss effects that are small and genuinely real: a tiny mortality difference detected across millions of patients is exactly a large-n near-null rejection, yet it may matter enormously at population scale. The Lindley diagnosis reads "small effect at large n" as evidentially empty, but smallness is not unreality, and importance depends on the decision context the statistics do not carry. The tension is that the tool guarding against precision-inflated significance can, applied reflexively, license discarding true small effects by relabeling them paradoxical artifacts. Diagnostic: Is this large-n rejection an effect too small to be real, or an effect that is small but genuine and consequential at scale — and does the Lindley flag distinguish them?

T2: The diffuse-prior lever versus circularity of the fix (tightening the prior toward the answer you want). The divergence is driven by how diffuse the alternative's prior is, so the recommended remedy is to tighten that prior toward plausible effect sizes, pulling the Bayesian verdict back toward the frequentist one. This exposes prior specification as a consequential decision rather than an innocuous default — a genuine gain. But the lever cuts the other way: because the Bayesian verdict is a function of the prior, one can make the disagreement vanish by choosing a prior that agrees with the p-value, so "resolving" the paradox by prior choice risks engineering the conclusion. The paradox reveals the frequentist answer as prior-free but arguably question-wrong and the Bayesian answer as question-right but prior-dependent, and the fix trades the first's objectivity for the second's relevance. The tension is that dissolving the paradox through the prior means the Bayesian verdict becomes whatever the prior asserts. Diagnostic: Is the tightened alternative prior justified by independent knowledge of plausible effect sizes, or chosen to make the Bayesian and frequentist verdicts agree?

T3: Point null as artifact versus point null as the question of interest (the measure-zero hypothesis science actually asks). The entry notes the paradox is largely an artifact of testing a measure-zero point null, and that replacing it with an interval null (|θ − θ₀| < δ) dissolves much of the divergence. This correctly locates half the structure in the strict point hypothesis rather than the prior. But the point null is frequently the hypothesis science genuinely wants to test — is there exactly no effect, is this constant precisely zero — and the interval-null replacement requires choosing δ, the width of "negligible," which is itself arbitrary and domain-specific, smuggling a substantive judgment into the test. The tension is that "the point null is the problem" is true structurally yet the point null is often the object of real interest, and its replacement trades one idealization (measure-zero exactness) for a new free parameter (the negligibility threshold) the paradox does not fix. Diagnostic: Is a strict point null genuinely the scientific question here, or a convenient idealization replaceable by an interval null whose δ can be principled rather than arbitrary?

T4: "Different questions" resolution versus the need for one verdict (dissolved intellectually, unresolved practically). The paradox's deepest lesson is that there is no contradiction: the p-value and Bayes factor answer different questions — tail-area improbability under H₀ versus prior-averaged relative likelihood of H₀ against H₁ — both computed correctly. This is intellectually satisfying and defuses the appearance of a clash. But a decision-maker still needs one verdict — approve the drug or not, claim the discovery or not — and "they answer different questions" does not say which question the decision requires. The clean resolution explains why the two diverge without adjudicating between them, so the practical bind remains exactly where the philosophical one dissolves. The tension is that recognizing the disagreement as a difference of questions rather than a contradiction removes the paradox as a puzzle while leaving the actual choice — which answer to act on — untouched. Diagnostic: Does the decision at hand need tail-area improbability under H₀ or averaged relative likelihood of the hypotheses — and does "they ask different questions" actually resolve which verdict to act on?

T5: Autonomy versus reduction (the Jeffreys-Lindley paradox or the incommensurability it instantiates). The Jeffreys-Lindley paradox is a named result with specific inferential cargo — the point null, the diffuse-prior Occam penalty, the p-versus-Bayes-factor monotonicity in n, the dissolving levers. It transfers wherever null-hypothesis testing with large samples is done, but that reach is single-substrate: biomedical trials, particle physics, genomics, and psychology all import the same inferential machinery, not re-instantiate an independent pattern, and outside statistics there is no system in which the paradox occurs except via the statistics applied to it. What genuinely generalizes is incommensurability (with the measurement/inference framing): two valid procedures answering different questions can diverge, and more data magnifies the gap rather than closing it. The tension is between a technical statistical result and the flatter incommensurability lesson that is what actually carries anywhere two differently-posed evaluations are confused for one. Diagnostic: Resolve toward incommensurability when the confusion is between two differently-posed evaluations in general; toward the Jeffreys-Lindley paradox when a point-null test and a prior-averaged Bayes factor diverge on large-sample data in situ.

Structural–Framed Character

The Jeffreys-Lindley paradox is best placed mixed: it has a genuine mathematical kernel (a theorem-like necessity, structural) and carries no evaluative charge, yet it is a result about a human-constructed inferential practice whose every distinctive term is discipline-furniture (framed). The five criteria split cleanly. Evaluative_weight is nil and points structural: the paradox states that two correctly-computed procedures diverge; it crowns no winner, convicts nothing, and the word "paradox" names a clash of intuitions, not a verdict — this makes it more evaluatively neutral than a diagnostic like the is-ought problem. On the structural side too, the core fact — that the Bayes factor for H₀ grows like √n while the p-value stays pinned — is a mathematical necessity that holds with theorem-like force, not a contingent regularity. What pulls it framed is the rest. Human_practice_bound is high: the paradox exists only inside the practice of statistical inference and is vacuous outside it — the entry is explicit that "there is no biological, social, or physical system in which 'the Jeffreys-Lindley paradox' occurs except by way of the statistics applied to it." Its very objects — a p-value, a Bayes factor, a diffuse prior over an alternative — are constructs of a methodological practice, not features of a substrate that runs observer-free the way a lithosphere or a seed shadow does. Institutional_origin is likewise framed for the distinctive layer: the point null, the diffuse-prior Occam penalty, the dissolving levers (interval null, pre-registered effect sizes), and the 5σ-style thresholds are all furniture of statistical methodology and its traditions, even though the underlying inequality is math. Vocab_travels is low for that layer (point null, posterior odds, Occam penalty pin to statistics), and import_vs_recognize is single-substrate: every field encountering the paradox is importing the same machinery, not recognizing an independently-arising pattern, and beyond statistics only the parent carries.

The portable structural skeleton is incommensurability — two valid procedures answering genuinely different questions can diverge, and accumulating more data magnifies rather than reconciles the gap between answers to non-commensurable questions. That skeleton is substrate-portable and is why the paradox does not sit at the pure framed pole: it has a real, exportable logical core. But it is precisely what the entry instantiates from its umbrellaincommensurability (with the general measurement/inference framing) — not what makes "the Jeffreys-Lindley paradox" itself travel: the cross-domain reach belongs to the incommensurability parent, while the point-null, diffuse-prior, and p-versus-Bayes-factor machinery stays home in formal inference. Its character: an evaluatively neutral, theorem-grounded result about a human inferential practice, structural in its mathematical kernel and its incommensurability skeleton but framed in being constituted by, and stated in the vocabulary of, statistical methodology — mixed overall, with only the incommensurability core traveling.

Structural Core vs. Domain Accent

This section decides why the Jeffreys-Lindley paradox is a domain-specific abstraction and not a prime, and it carries the case for its domain-specificity in one place.

What is skeletal (could lift toward a cross-domain prime). Strip the statistics and a thin relational structure survives: two valid procedures that answer genuinely different questions can return opposed verdicts on the same evidence, and accumulating more evidence magnifies rather than reconciles the gap between them. The portable pieces are abstract — two evaluations posed against the same object, a difference in the question each actually asks, a divergence in their verdicts, and a monotone widening of that divergence as information grows. That skeleton is genuinely substrate-portable — it recurs wherever two differently-posed measurements of one thing are mistaken for one measurement — which is exactly why the entry instantiates the catalog's incommensurability (with the general measurement / inference framing). This is an unusually clean core: the underlying inequality (the Bayes factor for H₀ growing like √n while the p-value stays pinned) holds with theorem-like force. But it is the core the paradox shares, not what makes it distinctive.

What is domain-bound. Everything that makes it the Jeffreys-Lindley paradox in particular is statistical-inference furniture and does not survive extraction. The two procedures are not any evaluations but a frequentist tail-area p-value and a Bayesian prior-averaged Bayes factor; the object is a measure-zero point null \(\theta = \theta_0\); the divergence is driven by the Occam penalty a diffuse alternative prior pays on data concentrated near \(\theta_0\); the widening is the specific p-versus-Bayes-factor monotonicity in n; and the remedies are inference-internal levers (tighten the alternative prior, replace the point null with an interval null, pre-register effect sizes, raise the discovery threshold to 5σ). These are constructs of a methodological practice, not features of any substrate. The decisive test: there is no biological, social, or physical system in which "the Jeffreys-Lindley paradox" occurs except by way of the statistics applied to it — the paradox is a result about inference, not a phenomenon in a domain, so removing the inferential apparatus does not leave a looser version of the paradox, it leaves nothing for the paradox to be about.

Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose transfer is recognition of the same mechanism, not analogy. The paradox's transfer is unusually shaped: wide in reach but single-substrate. Within statistical inference it transfers as the full result — the diagnostic signature, the monotone-in-n divergence, the dissolving levers, and the significance-is-not-evidential-weight boundary all carry intact across biomedical trials, particle physics, genomics, psychology, economics, and ecology. But every one of those fields is importing the same inferential machinery (point-null testing, posterior odds, sample size), not re-instantiating an independently-arising pattern — so the breadth is one substrate seen many times, not cross-substrate travel. Beyond statistics the paradox does not occur at all except through the statistics laid over a domain. And when the bare structural lesson is wanted cross-domain — two differently-posed evaluations can diverge, and more data magnifies the gap — it is already carried, in fully general form, by the incommensurability prime (with the measurement / inference framing) the paradox instantiates. The cross-domain reach belongs to that incommensurability parent; "the Jeffreys-Lindley paradox," as named, packs the point null, the diffuse-prior Occam penalty, and the p-versus-Bayes-factor machinery that should stay home in formal inference.

Relationships to Other Abstractions

Local relationship map for Jeffreys-Lindley ParadoxParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Jeffreys-LindleyParadoxDOMAINPrime abstraction: Bayesian Updating — is part ofBayesianUpdatingPRIMEPrime abstraction: Hypothesis Testing (Null vs. Alternative) — is part ofHypothesis Test…PRIME

Current abstraction Jeffreys-Lindley Paradox Domain-specific

Parents (2) — more general patterns this builds on

  • Jeffreys-Lindley Paradox is part of Bayesian Updating Prime

    The Jeffreys-Lindley construction contains a Bayesian posterior-odds update whose prior-averaged likelihood supplies the opposed verdict.

  • Jeffreys-Lindley Paradox is part of Hypothesis Testing (Null vs. Alternative) Prime

    The Jeffreys-Lindley construction contains a frequentist point-null hypothesis test whose tail-area verdict supplies one side of the disagreement.

Not to Be Confused With

  • Statistical versus practical (clinical) significance. The single-framework observation that a hypothesis test can return a tiny p-value for an effect too small to matter, because power rises with n — so a "significant" result may be trivially small. This is a within-frequentist point about effect size versus significance; the Jeffreys-Lindley paradox is the cross-framework result that a p-value and a Bayes factor answer different questions and diverge in opposite directions with n. The two are cousins (both bite at large n and near-null effects) but the JL paradox specifically requires the Bayesian comparison to swing toward the null. Tell: is the point that a significant effect is too small to be practically important (statistical-vs-practical significance), or that the same data are frequentist grounds to reject while being Bayesian grounds to confirm the null (Jeffreys-Lindley)?

  • Overpowered-study / "large n makes everything significant." The purely frequentist warning that with enough data almost any point null is rejected because the test statistic scales with √n. This describes one arm of the paradox (the falling p-value) but not the paradox itself, which is the simultaneous opposite movement of the Bayes factor. Reading JL as merely "big samples over-reject" drops the Bayesian half that makes it a paradox. Tell: is there only a shrinking p-value under discussion (overpowered study), or a p-value and a Bayes factor marching apart on the same data (Jeffreys-Lindley)?

  • Simpson's paradox. A different named statistical paradox: an association present in aggregated data reverses (or vanishes) when the data are split by a confounding subgroup. It is about aggregation and confounding across strata; the JL paradox is about two inference procedures disagreeing on one undivided dataset. They share only the word "paradox" and a flavor of counterintuitive reversal. Tell: does the reversal come from pooling versus stratifying the data (Simpson), or from comparing a p-value against a Bayes factor on a single sample (Jeffreys-Lindley)?

  • The frequentist–Bayesian debate at large. The broad, open-ended dispute over the foundations, interpretation, and merits of the two schools. The Jeffreys-Lindley paradox is one specific, sharp result within that debate — the standard proof that the method choice is a choice of which question is asked, not of computational convenience. It does not crown a winner. Confusing the paradox with the whole debate inflates a precise theorem into a sprawling controversy. Tell: is the referent the general contest between schools (the debate), or the specific point-null/diffuse-prior configuration where p-value and Bayes factor provably diverge and widen with n (the paradox)?

  • Incommensurability (incommensurability, with measurement/inference). The substrate-neutral parent the paradox instantiates — two valid procedures answering genuinely different questions can diverge, and more information magnifies rather than reconciles the gap. This is what actually travels beyond statistics; the JL paradox is its formal-inference specialization. Tell: strip away the point null, the diffuse-prior Occam penalty, and the p-versus-Bayes-factor monotonicity and what remains — "two differently-posed evaluations of one thing can diverge, and more data widens the gap" — is the incommensurability parent, treated more fully elsewhere; carry it (not "Jeffreys-Lindley") wherever the confusion is between two differently-posed evaluations in general.

Neighborhood in Abstraction Space

Jeffreys-Lindley Paradox sits in a moderately populated region (44th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.

Family — Statistical Inference & Model Failure Modes (16 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12