Look-Elsewhere Effect¶
Discount an exciting best-of-many find by the size of the search that produced it — converting a local p-value at one scanned peak into a global p-value asking whether any peak this extreme would occur anywhere, via the trials factor.
Core Idea¶
The look-elsewhere effect is the inflation of apparent statistical significance that occurs when a search scans many locations, mass values, or parameter ranges for a signal and then reports the most extreme result found without correcting for the size of the search space. A bump or excess that would be genuinely striking if the search position had been pre-specified becomes much weaker evidence once we account for the fact that the same search procedure would have flagged any of many candidate locations had the excess appeared there instead. The local p-value — asking only whether this particular peak is improbable at this particular location — must be converted to a global p-value asking whether any peak this extreme would occur anywhere in the search window under the null, and the conversion shrinks significance sharply.
The name originated in high-energy physics, where analysts scanning mass spectra for new-particle resonances repeatedly observed that a 3-sigma local excess in one mass bin reduced to roughly 1.5–2 sigma after correcting for the number of independent mass hypotheses effectively tested across the search window. Gross and Vitells (2010) formalized the correction for continuous parameter scans using the expected number of upcrossings of the significance threshold, providing a closed-form approximation to the global p-value. The structural mechanism is the same in every application: the trials factor — the number of effectively independent tests the search performs — multiplies the false-positive probability at any single location, so that a significance threshold calibrated for one pre-specified test is far too permissive when applied to the best result from many simultaneous tests. In genomic association studies the same correction is implemented as Bonferroni or FDR adjustment across millions of SNPs; in astronomy and gravitational-wave detection it enters as a template-bank trials factor; in A/B testing platforms it arises from monitoring many metrics across many segments simultaneously. Each context applies the same underlying principle: discount the best-of-many find by the size of the effective search.
Structural Signature¶
Sig role-phrases:
- the search procedure — a scan over many positions, mass values, or parameter ranges for a signal, free to flag a peak anywhere in the window
- the reported extremum — the most extreme result found, presented as if its location had been pre-specified
- the local significance — the p-value asking only whether this excess is improbable at this location
- the trials factor — the number of effectively independent looks the procedure performed, the scalar that multiplies the single-location false-positive probability
- the global significance — the engineered correction: local significance discounted by the trials factor, asking whether any excess this extreme would appear anywhere under the null
- the local-versus-scanned discipline — the boundary refusing to read a scanned extremum as a pre-specified test, the cardinal over-read the effect guards against
- the discrete/continuous fork — the counting branch: a fixed set of pre-specified tests takes the multiplicative (Bonferroni-type) factor, a continuous scan takes the expected-upcrossings / Gumbel extreme-value approximation
- the design-attribution handle — the limitation-turned-lever that the trials factor is tied to design choices (window width, grid density, monitored metrics), so pre-specifying or blinding shrinks it before unblinding
What It Is Not¶
- Not a property of the data or the signal. The significance inflation lives in the freedom of the search procedure to have flagged any of many positions, not in the bump itself nor in the test at any single location. The same data with a pre-specified search position carries full significance; it is the scanning that must be paid for, which is why the effect is diagnosed on the procedure, not on the peak.
- Not a reason the local p-value is "wrong." The local p-value correctly answers its own question — is this excess improbable at this location. The look-elsewhere effect says that is the wrong question for a scanned extremum; the honest claim is the global p-value, asking whether any excess this extreme would appear anywhere under the null. Both are correct; only the global one is legitimate to report for a best-of-many find.
- Not p-hacking, optional stopping, or the garden of forking paths. Those are mechanisms by which the effective trials factor grows beyond what is reported — analytic flexibility, peeking, undisclosed choices. The look-elsewhere effect is the inflation from an honestly disclosed scan over many positions; it is what the trials-factor correction is for, not a form of misconduct that hides the count.
- Not applicable where there is no countable search. The trials-factor machinery requires genuine local p-values, a significance threshold, and an estimable number of effectively independent looks. In non-statistical settings — backtest dredging, dashboard alert fatigue — the intuition "you looked at too many things" travels by analogy, but there is no p-value being corrected, so invoking "look-elsewhere" there borrows the slogan, not the instrument.
- Not the general multiple-comparisons correction. The look-elsewhere effect is the continuous-search (or large-structured-set) special case, sibling to Bonferroni and Benjamini–Hochberg under the broader multiple-comparisons family. Its distinctive machinery — expected-upcrossings and extreme-value approximations over a parameter range — is what sets it apart from a simple correction across a fixed discrete list of pre-specified tests.
Scope of Application¶
Because the look-elsewhere effect is a significance-inflation phenomenon and its trials-factor correction — a statistical adjustment, not a mechanism — it is not bounded to a subject matter: it applies wherever its precondition holds, namely a search that scans many positions, parameter values, or hypotheses and reports the extremum, with a local significance and an estimable count of effectively independent looks. The fields below are real uses of the identical trials-factor logic, not analogies; the boundary to respect is instrument-reach versus over-reading (a scanned extremum must not be read as a pre-specified test, and the machinery applies only where genuine local p-values and a countable look-count exist — non-statistical "you looked at too many things" uses are analogy carried by selection_bias). The genuine scope is hypothesis-testing-under-search.
- Particle physics — the original coinage: peak-fitting in mass spectra, where a 3-sigma local excess reduces to ~1.5–2 sigma globally (Gross & Vitells 2010 formalised the continuous-scan correction).
- Astronomy and gravitational-wave detection — template-bank and sky-position trials factors discounting transient and signal candidates.
- Genome-wide association studies — the same problem under Bonferroni/FDR correction across millions of SNPs.
- A/B testing platforms — many metrics monitored across many segments, where apparent winners arise by chance without correction.
- Anomaly detection in monitoring — scanning many time-series for spikes, which inflates the apparent rate of "interesting" events.
Clarity¶
Naming the look-elsewhere effect makes visible the difference between two significances that an isolated bump conflates: the local p-value, which asks whether this excess is improbable at this location, and the global p-value, which asks whether any excess this extreme would appear anywhere across the search window under the null. Without the distinction, a 3-sigma peak reads as 3-sigma evidence; with it, the analyst sees that the figure is "striking conditional on this being where we happened to find it," and that the honest claim is about the chance of something somewhere clearing threshold, not about this peak in isolation. The sharper question a searcher can now ask is not "how improbable is this bump?" but "how large was the effective space I scanned, and what is the best-of-that-many result worth once the size of the search is paid for?"
Its load-bearing service is to make the trials factor — the number of effectively independent looks the search performed — an explicit, estimable quantity rather than a silent feature of the analysis pipeline. This localizes where significance leaked: not in the data and not in the test at any single location, but in the freedom of the procedure to have flagged any of many candidate positions. That framing turns a vague worry about "finding patterns by scanning" into a calculable correction (a multiplicative trials factor in the discrete case, an expected-upcrossings or extreme-value approximation for a continuous scan) and tells the experimenter exactly which design choices inflate it — a wider mass window, a denser grid of hypotheses, more monitored metrics or segments — so that pre-specifying the search window or blinding the analysis becomes a recognizable way to shrink the factor before the data are unblinded.
Manages Complexity¶
The worry the look-elsewhere effect addresses — that an exciting find was manufactured by the freedom of the search rather than present in the data — arrives, untamed, as an open-ended and substrate-specific anxiety. A particle physicist scanning a wide mass window, a geneticist testing millions of SNPs, an astronomer sweeping a template bank across sky positions, an experimentation platform monitoring dozens of metrics across dozens of segments: each could try to reason case by case about how much the best result should be discounted, weighing the breadth of the window against the density of the grid against the correlations among neighboring tests, a different bespoke argument for every pipeline. The effect collapses all of that onto a single scalar. Whatever the search is over and whatever the signal is called, the inflation of the best-of-many find is governed by one quantity — the trials factor, the number of effectively independent looks the procedure performed — which multiplies the false-positive probability at any single location. The sprawl of "how worried should I be about this particular scan?" reduces to "how large is the effective search, and what is the best result worth once that size is paid for?"
What the analyst tracks is therefore just two numbers: the local significance at the reported peak, and the trials factor of the search that found it. From those, the honest global significance reads off — local significance discounted by the effective number of looks — and the qualitative verdict follows: a 3-sigma local excess in a window with a trials factor near thirty is, globally, an unremarkable fluctuation, while the same local figure from a pre-specified single test stands. The structure also fixes the branch by which the trials factor itself is computed rather than leaving it to improvisation: a fixed discrete set of pre-specified tests takes the simple multiplicative (Bonferroni-type) correction, while a continuous scan over a parameter range takes an expected-upcrossings or extreme-value approximation, the two branches of one principle differing only in how the effective look-count is counted. And because the factor is an explicit, estimable quantity tied to design choices — window width, grid density, number of monitored metrics and segments — the analyst can read off in advance which choices inflate it and shrink it before unblinding by pre-specifying the window or blinding the analysis. A high-dimensional, per-pipeline "did my scanning fool me, and by how much" problem becomes a low-dimensional "estimate one trials factor, discount the local p-value by it" problem with a single correction principle and a clean discrete-versus-continuous fork.
Abstract Reasoning¶
The central move is deflationary discounting — reasoning from a striking local result and the size of the search that found it down to the honest global significance. The analyst takes the reported peak's local p-value, asks how many effectively independent looks the procedure performed across the window, and converts: a 3-sigma excess found anywhere in a window with a trials factor near thirty is, globally, an unremarkable fluctuation, while the identical local figure from a single pre-specified test stands at full strength. The characteristic inference runs from the freedom of the procedure to have flagged any of many positions to the best-of-many result being worth far less than it looks — the same procedure would have raised the same alarm wherever the excess happened to fall, so the find must be paid for at the scale of the whole scan, not the single bin.
This rests on a boundary-drawing move that distinguishes two significances an isolated bump conflates: the local p-value (is this excess improbable at this location?) versus the global p-value (would any excess this extreme appear anywhere in the window under the null?). The move is to recognize which question the data actually answer — "striking conditional on this being where we happened to look" — and to refuse the inflated reading that treats a scanned extremum as a pre-specified test. The inference is from how the peak was found (by searching, not by prediction) to which significance is the legitimate one to report.
A diagnostic-on-the-procedure move locates exactly where significance leaked: not in the data, not in the test at any single location, but in the trials factor — the number of effectively independent looks. The analyst reasons backward from an inflated apparent significance to the specific design freedoms that produced it — a wider mass window, a denser grid of hypotheses, more monitored metrics or segments — converting a vague worry about "finding patterns by scanning" into a calculable, design-attributable quantity. This makes the interventionist move available before the data are seen: because the trials factor is tied to design choices, the analyst predicts that pre-specifying the search window, blinding the analysis, or narrowing the grid will shrink the factor, and so reduces the eventual discount by construction rather than paying for it after unblinding. The inference runs from "which choices inflate the effective look-count?" forward to "constrain those before looking."
Finally there is a computation-routing move that fixes how the trials factor is counted rather than improvising it per pipeline. The analyst classifies the search by its structure and selects the matching correction: a fixed discrete set of pre-specified tests takes the simple multiplicative (Bonferroni-type) factor, while a continuous scan over a parameter range takes an expected-upcrossings or extreme-value (Gumbel-type) approximation that counts how often the significance threshold is crossed across the range. The inference is from the shape of the search space (discrete list versus continuum) to which counting method gives the effective number of looks — two branches of one principle, differing only in how the freedom of the scan is tallied.
Knowledge Transfer¶
The look-elsewhere effect is a significance-inflation phenomenon and its trials-factor correction — a statistical adjustment, not a causal mechanism — so "mechanism within / metaphor beyond" applies only loosely: there is nothing in the world to recognise, only a correction to apply to the best-of-many result of a search. Within statistical methodology the correction transfers literally wherever its precondition holds: a search procedure that scans many positions, parameter values, or hypotheses and reports the extremum, with a local significance and an estimable count of effectively independent looks. That precondition is substrate-blind, which is why the identical trials-factor logic serves particle physics (peak-fitting in mass spectra — the original coinage, formalised by Gross & Vitells 2010 for continuous scans), astronomy and gravitational-wave detection (template-bank and sky-position trials factors), genome-wide association studies (Bonferroni/FDR across millions of SNPs — the same problem under a different name), A/B testing platforms (many metrics across many segments), and anomaly detection (scanning many time-series for spikes). In every case the same two numbers — local significance and trials factor — yield the same global discount, with the discrete-versus-continuous fork (multiplicative Bonferroni-type factor versus expected-upcrossings / Gumbel extreme-value approximation) selected by the shape of the search space. The breadth is reuse of one correction across subject matter, not recurrence of a structural pattern under different vocabulary.
Because this is an instrument, the boundary to mark is instrument-reach versus over-reading, and the look-elsewhere effect is itself a guard against a specific over-read — the cardinal error of treating a scanned extremum as if it were a pre-specified test, reading a 3-sigma local excess found anywhere in a wide window as 3-sigma evidence when, paid for at the scale of the whole scan, it is an unremarkable fluctuation. The correction's discipline is to refuse that inflated reading and report the global significance the data actually warrant. The reach edge is equally important: the trials-factor machinery applies only where local p-values, a significance threshold, and a countable effective number of looks genuinely exist; invoking "look-elsewhere" where there is no such quantity is borrowing the slogan, not applying the instrument.
Where a genuinely cross-domain correction is wanted, it is not the look-elsewhere effect specifically but the broader prime it specialises that travels mechanically: multiple_comparisons_correction — as the number of opportunities for a false positive grows, the probability that at least one occurs approaches one, so the threshold must be tightened by the effective number of tests — of which the look-elsewhere effect is the continuous-search (or large-structured-set) special case, sibling to bonferroni_correction and benjamini_hochberg_procedure. Into non-statistical settings, by contrast, the reach is analogy, and should be marked so: machine-learning hyperparameter selection (test-set leakage by repeated selection), financial backtesting (data dredging over many strategies), and operational dashboards (false-alert inflation) all borrow the intuition "you looked at too many things" without using the trials-factor mathematics — there is no p-value being corrected. Where that analogy nonetheless bites mechanically — cross-validation overfitting, selecting a model on the same data used to evaluate it — the move is already housed by selection_bias and data_leakage, not by the look-elsewhere apparatus. So the portable correction belongs to the multiple-comparisons family; the portable intuition belongs to the selection-bias family; and "look-elsewhere effect," with its mass-window trials factors and expected-upcrossings approximations, stays within hypothesis-testing-under-search (see Structural Core vs. Domain Accent).
Examples¶
Canonical¶
The effect's home is the search for new-particle resonances in high-energy physics, and the Higgs boson search illustrates it exactly. When scanning a mass spectrum for a bump, an experiment considers many possible mass values across the search window — say the effective number of independent mass hypotheses is about 30. Suppose an excess appears at one mass with a local p-value of 0.0013, i.e. 3 sigma (a 0.13% chance of arising by fluctuation at that specific mass). But the search was free to find such a bump at any of the ~30 positions, so the relevant question is the chance that some bump this extreme appears anywhere. The global p-value is approximately 1 − (1 − 0.0013)³⁰ ≈ 0.038, about 2 sigma — a far less impressive result. This is precisely why particle physics demands a 5-sigma local discovery threshold: it buys back the significance a wide search space spends.
Mapped back: Scanning the mass spectrum is the search procedure and the single most extreme bin is the reported extremum, carrying its 3-sigma local significance. The ~30 effective mass hypotheses are the trials factor, and discounting to ~0.038 gives the global significance; counting looks by the multiplicative 1−(1−p)ⁿ form is the discrete branch of the discrete/continuous fork.
Applied / In Practice¶
Genome-wide association studies (GWAS) apply the identical logic under the name multiple-testing correction. A GWAS scans hundreds of thousands to millions of single-nucleotide polymorphisms (SNPs) across the genome, testing each for association with a disease or trait. Because roughly a million effectively independent tests are performed, a SNP reaching a nominally impressive p-value of, say, 10⁻⁵ is unremarkable — with a million looks, many such hits arise by chance. The field therefore adopted a stringent "genome-wide significance" threshold of 5×10⁻⁸, essentially a Bonferroni correction dividing a 0.05 target by roughly one million independent tests (0.05 / 10⁶ = 5×10⁻⁸). Only associations clearing that bar are treated as real. This is the look-elsewhere effect's discrete branch in a biomedical setting: the best-of-many find is discounted by the effective size of the genomic search before it is believed.
Mapped back: Testing every SNP is the search procedure and the top hit is the reported extremum with its local significance. The ~10⁶ independent tests are the trials factor, the 5×10⁻⁸ bar is the global significance threshold, and dividing 0.05 by the test count is the multiplicative Bonferroni side of the discrete/continuous fork — enforcing the local-versus-scanned discipline.
Structural Tensions¶
T1: Correct local p-value versus legitimate global p-value (the number that is right yet must not be reported). The look-elsewhere effect does not declare the local p-value wrong — it correctly answers its own question, whether this excess is improbable at this location — yet it forbids reporting it for a scanned extremum, where the honest claim is the global p-value asking whether any peak this extreme would appear anywhere. Both figures are mathematically correct; only one is legitimate for a best-of-many find. The tension is that the analyst holds a true number that is nonetheless misleading, and the error is not a miscalculation but a mismatch between the question the statistic answers and the question the search actually posed. A reader who trusts the arithmetic is misled precisely because the arithmetic is right. Diagnostic: Was the reported location pre-specified, making the local p-value the legitimate claim, or found by scanning, so only the global figure may be reported?
T2: Honestly disclosed scan versus uncountable analytic freedom (what the correction can and cannot buy back). The trials factor corrects an honestly disclosed search over a countable number of positions — that is exactly what it is for, and it is not misconduct. But the same "you looked too many times" intuition governs p-hacking, optional stopping, and the garden of forking paths, where the effective number of looks is inflated by undisclosed flexibility and has no countable value at all. The tension is that the instrument's discipline applies cleanly only where the search is transparent enough to tally, while the more dangerous inflation — the hidden one — leaves nothing to correct because the count itself is concealed. Discounting the disclosed scan can even lend false comfort, certifying the honest part while the undisclosed part goes unpaid. Diagnostic: Is the trials factor being computed over a fully disclosed, countable search, or is undisclosed analytic freedom inflating an effective look-count that no correction can reach?
T3: Effective independent looks versus raw position count (the scalar is itself an estimate). The whole correction rides on one number — the trials factor — presented as if it were a clean count, but "effectively independent looks" is not the raw number of scanned positions: neighbouring tests are correlated, so a naive multiplicative count over a fine grid is too conservative, while the continuous scan requires an expected-upcrossings or extreme-value approximation that is itself a model. The tension is that the quantity doing all the deflationary work carries modeling choices and estimation error under the guise of a single decisive scalar. Overstate independence and real signals are needlessly discounted; understate it and the search is under-corrected. The trials factor's authority as one clean number is exactly what hides that it was estimated. Diagnostic: Is the trials factor counting raw scanned positions, or the effectively independent looks after accounting for correlation among neighbouring tests — and how sensitive is the verdict to that estimate?
T4: Blinding-by-design versus exploratory freedom (paying the discount before versus after looking). Because the trials factor is tied to design choices, pre-specifying the window or blinding the analysis shrinks it by construction, so the discount is paid before unblinding rather than after. This is the disciplined path — but it costs the very latitude that makes a wide scan valuable: a pre-registered narrow window cannot chase the unexpected bump that a broad, free search surfaces. The tension is that the freedom to look everywhere is what finds surprises and what inflates the price of any find; constraining it up front buys significance at the cost of discovery reach. An analyst who blinds too aggressively forecloses the serendipitous excess; one who scans freely pays the full trials-factor toll. Diagnostic: Is the search window pre-specified to shrink the trials factor by design, or left wide to preserve discovery latitude at the cost of a larger discount at correction time?
T5: Buying back significance versus suppressing genuine signals (the type-II cost of a stringent threshold). Paying for the search — the 5-sigma local threshold in physics, the 5×10⁻⁸ bar in GWAS — buys back the significance a wide scan spends and guards against false positives. But the same stringency raises the bar for real effects: a genuine but modest resonance found in a wide window may never clear the global threshold and stays unclaimed. The tension is that the correction protecting against best-of-many false positives simultaneously inflates false negatives, and the discount scales with the search size, so the broadest, most exploratory searches — the ones most likely to encounter something new — are also the ones where a true but faint signal is hardest to establish. Diagnostic: Is the threshold set to control best-of-many false positives, and what genuine-but-modest signal would it discard as the price — is that type-II cost acceptable for this search?
T6: Autonomy versus reduction (a named physics effect or the continuous-search case of its parent correction). "Look-elsewhere effect" is a canonically named instrument with its own machinery — mass-window trials factors, the expected-upcrossings and Gumbel extreme-value approximations, Gross and Vitells' continuous-scan formalization, the 5-sigma convention. Yet within statistics it is the continuous-search special case of multiple_comparisons_correction, sibling to Bonferroni and Benjamini–Hochberg, and beyond hypothesis-testing-under-search it does not travel as an instrument at all — the transferable correction belongs to that parent, while the non-statistical "you looked at too many things" intuition is carried by selection_bias and data_leakage, with no p-value being corrected. The tension is between a distinctively named effect that earns its own formalism and the recognition that its portable substance already belongs to the multiple-comparisons family. Diagnostic: Resolve toward the parents (multiple_comparisons_correction; selection_bias) when asking what travels outside a countable statistical search, toward the named effect when correcting a scanned extremum in situ.
Structural–Framed Character¶
The look-elsewhere effect sits at the mixed position on the structural–framed spectrum — more structural than the rhetorical and code-smell entries near it, but held short of the mixed-structural band (where isostasy sits) by a distinctive fact: it is not a mechanism running in nature but a correction applied within an epistemic practice. The criteria split cleanly. Evaluative_weight is nil, its strongest structural mark: a significance-inflation phenomenon and its trials-factor discount are evaluatively neutral — the effect praises and blames nothing, it merely converts a local p-value into the global one the search actually warrants. On vocab_travels and import_vs_recognize it also reads structural within its substrate: the trials-factor correction transfers literally — not by analogy — from particle physics to genome-wide association studies to A/B-testing platforms to gravitational-wave template banks, the same two numbers (local significance, effective look-count) yielding the same global discount, because the underlying object is substrate-blind extreme-value mathematics. That is genuine recognition of one correction reused across subject matter, not metaphor.
What keeps it off the structural side, and out of the mixed-structural band, is the other pair. Human_practice_bound is high in a specific, telling way: as the entry itself stresses, "there is nothing in the world to recognise, only a correction to apply" — the effect is constituted by the practice of statistical inference (local p-values, a significance threshold, a null, a reported extremum) and dissolves the instant that practice is removed. Strip away the inference apparatus and what remains is only the bare mathematical fact that the maximum of many draws tends to be large; the "look-elsewhere effect" as such lives entirely inside hypothesis-testing-under-search, an epistemic activity, not in an observer-free process the way a rebounding lithosphere is. Unlike isostasy — which runs with every geophysicist removed — the look-elsewhere effect has no referent absent an analyst conducting a search and reporting a significance. Institutional_origin is correspondingly split: the phenomenon is mathematical, but its name, its formalization for continuous scans (Gross & Vitells 2010), and above all its decision thresholds (the 5-sigma discovery convention, the 5×10⁻⁸ genome-wide bar) are artifacts of specific disciplinary conventions, not facts of nature.
The one portable structural skeleton is multiple-comparisons correction: as the number of opportunities for a false positive grows, the probability that at least one occurs approaches one, so the acceptance threshold must be tightened by the effective number of tests. That skeleton is genuinely substrate-portable and lives in the catalog as multiple_comparisons_correction (with bonferroni_correction and benjamini_hochberg_procedure as siblings, and selection_bias/data_leakage carrying the non-statistical "you looked at too many things" intuition). But it does not lift the look-elsewhere effect to a prime, because that portable structure is precisely what the effect instantiates from its parent as the continuous-search special case, not what makes "look-elsewhere effect" itself travel: the cross-domain reach belongs to the multiple-comparisons family, while the entry's distinctive content — mass-window trials factors, expected-upcrossings and Gumbel approximations, the discrete/continuous fork, and the sigma conventions — stays home in hypothesis-testing-under-search. Its character: an evaluatively neutral, mathematically real correction whose portable core is substrate-blind but whose existence is constituted by the epistemic practice of significance testing, leaving it mixed — structural in its extreme-value skeleton, framed in being a within-inference correction rather than a mechanism in the world.
Structural Core vs. Domain Accent¶
This section decides why the look-elsewhere effect is a domain-specific abstraction and not a prime, and it carries the case for its domain-specificity — with the wrinkle that this entry is an instrument rather than a mechanism, so the boundary runs between reuse-across-subject-matter and the parent correction it specializes.
What is skeletal (could lift toward a cross-domain prime). Strip the physics and a thin statistical structure survives: as the number of opportunities for a false positive grows, the probability that at least one occurs approaches one, so the acceptance threshold must be tightened by the effective number of tests. The portable pieces are abstract: a search with many looks, a single-look false-positive probability, and a discount that pays for the best-of-many find at the scale of the whole search. That skeleton is genuinely substrate-portable — which is exactly why the entry names multiple_comparisons_correction as the parent it instantiates (with bonferroni_correction and benjamini_hochberg_procedure as siblings), and, for the non-statistical version of the "you looked at too many things" intuition, selection_bias and data_leakage. But this multiple-comparisons skeleton is the core the effect shares with the whole family, not what makes the look-elsewhere effect the distinctive continuous-search instrument it is.
What is domain-bound. Almost everything that makes the concept the look-elsewhere effect in particular is hypothesis-testing-under-search furniture, and none of it survives extraction intact. The search procedure scanning a mass window; the reported extremum presented as if pre-specified; the local significance and global significance the correction converts between; the trials factor estimated as effectively independent looks; the discrete/continuous fork (multiplicative Bonferroni-type factor versus the expected-upcrossings / Gumbel extreme-value approximation, the latter formalized by Gross & Vitells for continuous scans); the design-attribution handle; and the disciplinary sigma conventions (5-sigma discovery, the 5×10⁻⁸ genome-wide bar) — these are the worked apparatus. The decisive test: remove the inference practice — no local p-values, no significance threshold, no null, no estimable look-count — and the effect loses its referent entirely; what remains is only the bare mathematical fact that the maximum of many draws tends to be large, with none of the trials-factor machinery or the discrete/continuous counting that distinguishes this instrument from any other "you searched too much" worry. Invoking "look-elsewhere" where no p-value is being corrected borrows the slogan, not the instrument.
Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose cross-domain transfer is recognition of the same mechanism. The look-elsewhere effect's transfer is bimodal, but along an unusual seam because it is an instrument, not a mechanism. Within statistical methodology the correction transfers literally — not by analogy — across particle physics, astronomy and gravitational-wave detection, genome-wide association studies, A/B-testing platforms, and anomaly detection, because the precondition (a scan with local significance and a countable effective look-count) is substrate-blind and the same two numbers yield the same global discount; this is genuine recognition of one correction reused across subject matter. Beyond hypothesis-testing-under-search it does not travel as an instrument at all: machine-learning hyperparameter dredging, financial backtesting, and dashboard alert fatigue borrow only the intuition "you looked at too many things," with no p-value to correct — analogy, and where that analogy bites mechanically the move is already housed by selection_bias and data_leakage. And when the bare transferable correction is wanted cross-domain, it is already carried, in more general form, by multiple_comparisons_correction (of which the look-elsewhere effect is the continuous-search special case). The cross-domain reach thus splits between the multiple-comparisons family (the portable correction) and the selection-bias family (the portable intuition); "look-elsewhere effect," with its mass-window trials factors and expected-upcrossings approximations, carries hypothesis-testing-under-search baggage that should stay home.
Relationships to Other Abstractions¶
Current abstraction Look-Elsewhere Effect Domain-specific
Parents (1) — more general patterns this builds on
-
Look-Elsewhere Effect is a kind of Multiple Comparisons Correction Prime
The Look-Elsewhere Effect is multiple-comparisons correction specialized to a scanned or continuous search space, where a local extremum is converted to global significance through an effective trials factor.It inherits the defined-family multiplicity problem and the requirement to bound a family-level false-positive criterion. It specializes the family to locations or parameter values searched for an extremum and replaces a fixed discrete count, where necessary, with an effective trials factor estimated by upcrossing or extreme-value methods.
Hierarchy paths (6) — routes to 6 parentless roots
- Look-Elsewhere Effect → Multiple Comparisons Correction → Type I & Type II Errors → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Inductive Reasoning
- Look-Elsewhere Effect → Multiple Comparisons Correction → Type I & Type II Errors → Trade-offs → Constraint
- Look-Elsewhere Effect → Multiple Comparisons Correction → Type I & Type II Errors → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Uncertainty
- Look-Elsewhere Effect → Multiple Comparisons Correction → Type I & Type II Errors → Hypothesis Testing (Null vs. Alternative) → Verification → Evaluation → Comparison → Self Checking
- Look-Elsewhere Effect → Multiple Comparisons Correction → Type I & Type II Errors → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Probability → Measure → Set and Membership
- Look-Elsewhere Effect → Multiple Comparisons Correction → Type I & Type II Errors → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
Not to Be Confused With¶
-
Multiple-comparisons correction (the parent). The substrate-neutral principle the look-elsewhere effect instantiates — as the number of opportunities for a false positive grows, the probability of at least one approaches one, so the threshold must be tightened by the effective number of tests. This is the super-type, not a peer: the look-elsewhere effect is its continuous-search (or large-structured-set) special case. Tell: is the question the general "correct for how many tests you ran" (the parent, treated more fully in Structural Core vs. Domain Accent), or specifically the discount for a scanned extremum over a parameter range with an estimated trials factor (look-elsewhere effect)?
-
Bonferroni correction. The sibling correction for a fixed, discrete set of pre-specified tests — divide the target error by the number of tests (or equivalently multiply the p-value). It controls the family-wise error rate over a countable list. The look-elsewhere effect's distinctive machinery is the continuous-scan case, where the effective look-count must be estimated by expected-upcrossings / extreme-value (Gumbel) approximations rather than simply counted. Tell: is the search a fixed enumerated list of hypotheses (Bonferroni), or a continuous parameter scan whose independent-look count must be modeled (look-elsewhere)?
-
Benjamini–Hochberg / false discovery rate. The sibling procedure that controls the expected proportion of false positives among rejections (FDR) rather than the probability of any false positive (FWER). It answers a different error-control question and typically yields a less stringent threshold than a look-elsewhere / Bonferroni family-wise correction. Tell: is the goal to bound the chance of even one spurious peak (look-elsewhere / FWER), or to bound the fraction of claimed discoveries that are false (FDR)?
-
p-hacking / the garden of forking paths / optional stopping. Mechanisms by which the effective trials factor grows beyond what is disclosed — undisclosed analytic flexibility, peeking, choosing tests after seeing data. The look-elsewhere effect is the inflation from an honestly disclosed, countable scan; it is the correction such practices evade, not a form of them. Tell: is the number of looks fully disclosed and estimable so it can be corrected (look-elsewhere), or hidden and uncountable because of undisclosed choices (p-hacking)?
-
Selection bias / data leakage. The family that carries the non-statistical "you looked at too many things" intuition — repeated model selection on the same data, backtest dredging, evaluating on training data — where there is no p-value being corrected. When the look-elsewhere slogan is invoked off a genuine statistical search, this family, not the trials-factor instrument, is doing the work. Tell: is there a genuine local p-value, threshold, and countable look-count to discount (look-elsewhere), or just an informal over-search whose damage is best named as selection bias / leakage?
Neighborhood in Abstraction Space¶
Look-Elsewhere Effect sits in a sparse region of the domain-specific corpus (93rd percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Statistical Bias & Sampling Artifacts (6 abstractions)
Nearest neighbors
- Jeffreys-Lindley Paradox — 0.83
- Benjamini–Hochberg Procedure — 0.82
- Null Ritual — 0.81
- Type S Error — 0.81
- Congruence Bias — 0.81
Computed from structural-signature embeddings · 2026-07-12