Type M Error¶
Quantify how much a significant effect's reported magnitude is exaggerated by the significance filter under low power, via the exaggeration ratio — the expected significant estimate divided by the true effect — computable from the design before any data exist.
Core Idea¶
A Type M (magnitude) error is the systematic exaggeration of a statistically significant effect's size that results from filtering estimates through a significance threshold under low statistical power. When a study is underpowered, only estimates that land far out in the sampling distribution's tail — well above the true effect — are large enough to clear the significance threshold; the conditional distribution of effect-size estimates given significance is therefore shifted substantially upward from the true effect, and the expected reported magnitude among significant results is a multiple of the true magnitude. Gelman and Carlin (2014) formalized the exaggeration ratio as the key diagnostic: the ratio of the expected significant estimate to the true effect, computed prospectively from the design before data are collected. An underpowered study powered at 17% to detect a true effect of d = 0.10 will reject the null only when its sample estimate exceeds roughly d = 0.28, and the expected value of the estimate conditional on rejection is approximately d = 0.35 — an exaggeration ratio above three. The published headline number is thus not an overestimate by accident but by selection structure: the significance filter ensures that only the most extreme draws from the sampling distribution reach the literature, and under low power those extreme draws are far from the truth.
Type M error reorients the standard power analysis from "what is the probability that I detect the effect?" to "if I do detect the effect, what magnitude will I report, and how far from the truth will it be?" The former question addresses whether the study generates any publishable result; the latter addresses whether the publishable result is informative. The two questions give divergent answers when power is low, which is the regime that characterizes a substantial fraction of published empirical work in psychology, biomedicine, and the social sciences. The operationally important consequence is not merely that a single study's magnitude is inflated, but that inflated initial estimates — the winner's curse operating on effect-size estimates filtered by a significance gate — fail to replicate at full power, and the shrinkage from first study to replication is itself a diagnostic signature of systematic Type M error in a literature.
Structural Signature¶
Sig role-phrases:
- the noisy effect-size estimator — a sampling distribution of estimates centered on the true effect, the thing being reported
- the significance filter — the one-sided threshold gate that admits only estimates large enough to clear it
- the low-power regime — the condition that makes the bias bite: only far-tail draws clear the threshold, so the conditional distribution given significance is shifted upward from the truth
- the conditional-on-significance shift — the selection structure whereby a published magnitude is a draw from the upper tail, not an unbiased estimate (a 17%-power study against d=0.10 rejects only above ~d=0.28, expecting ~d=0.35)
- the exaggeration ratio — the prospective diagnostic scalar: the expected significant estimate divided by the true effect, computable from assumed effect and power before any data exist
- the unit-of-analysis reframe — the shift from the reject/fail-to-reject decision (Type I/II) to the reported magnitude conditional on rejection, with the low-power-versus-adequate fork deciding whether the headline size is trustworthy
- the replication-shrinkage signature — the retrospective tell: systematic shrinkage from an initial study to its higher-powered replication, diagnosing Type M in a literature without appeal to misconduct
- the remedy set — the interventions read off the mechanism: power up to push the ratio toward one, apply shrinkage estimators, pre-register, treat underpowered single-study magnitudes as upper bounds
What It Is Not¶
- Not a Type I or Type II error. Those frame the decision — reject or fail to reject — as the unit of analysis and leave the reported magnitude unexamined. Type M shifts the unit to the number itself, conditional on rejection, asking not "did I detect the effect?" but "if I did, how far from the truth is my estimate?" It is parasitic on the same significance scaffolding but answers a different question.
- Not fraud or bad luck. The inflation is a selection structure, not misconduct or an unlucky draw: under low power the significance filter admits only far-tail estimates, so the published magnitude is exaggerated by design. Reading a non-replicating effect as fabricated or merely unfortunate misses that the gate guaranteed an upward-biased number from an honest estimator.
- Not a flaw in the estimator. The sampling distribution is correctly centered on the truth; the exaggeration comes entirely from conditioning on having passed the threshold. An unbiased estimator still yields Type M error once filtered by significance — so the remedy is power (or shrinkage that undoes the conditional shift), not a "better" estimator.
- Not ordinary sampling noise. Type M is the systematic, directional shift in the conditional-on-significance distribution under low power, not the symmetric scatter of estimates around the truth. It biases the reported magnitude upward as a rule, where sampling noise alone would scatter estimates both ways; the exaggeration ratio quantifies the bias, not the variance.
- Not the winner's curse under its own name. Type M is the statistical-inference face of the winner's curse — select the extremum of noisy estimates through a one-sided gate and the selected value overstates the truth — but the portable cross-domain pattern is the winner's curse (and broader selection bias). An auctioneer overpaying commits no "Type M error": there is no p-value, no power, no significance filter, though both instance the same selection mechanism.
- Not a problem when power is adequate. When power is high the exaggeration ratio sits near one and a significant estimate's magnitude can be taken roughly at face value. The bias bites specifically in the low-power regime; it is not an unconditional indictment of all significant effect sizes, only of those filtered through an underpowered design.
Scope of Application¶
Type M error lives across the significance-testing sciences that share the substrate of estimation from a significance-filtered sample; its reach is within that one statistical domain, the mechanism recurring in identical form across them (always paired with its sign-sibling Type S). The broader pattern — select the extremum of noisy estimates through a one-sided gate and it overstates the truth — travels under winner_s_curse and the selection_bias family, not under this construct: an auctioneer overpaying commits no "Type M error," there being no p-value, power, or significance filter.
- Underpowered psychology, economics, and biomedicine — the original critique target: surprising effects that shrink dramatically on higher-powered replication.
- Early-phase clinical trials — effects clearing p < 0.05 that Phase III systematically fails to recover (a known regularity in oncology and cardiovascular medicine).
- Genome-wide association studies — winner's-curse shrinkage of top-hit effect sizes between discovery and replication.
- A/B testing — stopped-on-significance experiments exaggerating lift, met with sequential-testing corrections or shrinkage estimators.
- Meta-science and the publication-bias literature — PET-PEESE and related bias-adjustment methods explicitly correcting this magnitude inflation.
Clarity¶
Naming the Type M error breaks the equation that the Neyman-Pearson frame quietly encourages: that a significant result is a roughly-right estimate. Type I and Type II error treat the decision — reject or fail to reject — as the unit of analysis, leaving the reported magnitude unexamined; the Type M frame shifts the unit to the number itself, conditional on rejection, and asks not "did I detect the effect?" but "if I did, how large will my reported estimate be, and how far from the truth?" Those two questions are routinely conflated, and Type M shows they diverge sharply under low power — the regime that characterizes much of the published literature. The reader who has the exaggeration ratio stops treating "p < 0.05" as a license to believe the headline effect size, because the same significance gate that delivered the result also selected an upward-biased draw.
Its load-bearing service is to make the exaggeration ratio a concrete, prospective quantity rather than a vague worry about inflation. By defining it as the expected significant estimate divided by the true effect, computable from the design before any data are collected, the concept turns the question "is this study informative?" into a number an analyst can read off the power and the assumed effect size. This localizes the problem precisely: not in the estimator, not in the experimenter's honesty, but in the selection structure imposed by the significance filter under low power. And it supplies a distinct downstream signature — systematic shrinkage from an initial study to its higher-powered replication — that lets a reader diagnose Type M in a literature rather than blaming individual studies for failing to replicate, separating the genuine pathology (a magnitude inflated by selection) from the easy but wrong reading (the original result was fraudulent or merely unlucky).
Manages Complexity¶
The reasons a reported effect size might overstate the truth accumulate, across the empirical sciences, into a sprawling and seemingly disconnected catalog of worries: the winner's curse that shrinks top GWAS hits at replication, the systematic over-statement of early-phase trial effects relative to what Phase III recovers, the lift that stopped-on-significance A/B tests exaggerate, the publication-bias adjustments that PET-PEESE and related methods are built to undo, the general failure of "surprising" findings to hold up. Treated as separate phenomena, each invites its own correction folklore and its own post-hoc rationalization for why this particular literature inflated. Type M error compresses the whole catalog onto a single mechanism and a single prospective scalar. The common machinery is one selection structure — a significance threshold filtering a noisy estimator — and the magnitude inflation it produces is summarized by one number, the exaggeration ratio: the expected significant estimate divided by the true effect. A tangle of meta-research worries collapses to a quantity an analyst can compute from the design alone.
What the analyst tracks reduces to two design inputs — an assumed true effect and the study's power against it — from which the exaggeration ratio reads off, and the qualitative verdict on whether a significant result will be informative about magnitude follows by a fixed regime split rather than by case-by-case suspicion. Crucially the relevant question is reoriented at the same time: the standard power analysis asks "what is the probability I detect the effect?" while the Type M reframe asks "if I detect it, how large will my reported number be and how far from the truth?", and the two diverge precisely in the low-power regime. When power is high, the ratio sits near one and a significant estimate's magnitude can be taken roughly at face value; when power is low — as it is across a substantial fraction of published psychology, biomedicine, and social science — the threshold can be cleared only by draws far out in the tail, the ratio climbs to a multiple of three or more, and the headline magnitude must be read as selection-inflated rather than approximately true. The structure further supplies a downstream check that needs no appeal to misconduct: systematic shrinkage from an initial study to its higher-powered replication is the observable signature of Type M operating in a literature, so the same framework that predicts inflation prospectively also diagnoses it retrospectively, separating a magnitude inflated by the significance gate from the easy but wrong reading that the original was fraudulent or merely unlucky. A high-dimensional "why do effect sizes keep failing to hold up, and which of these many bias stories applies here" problem becomes a low-dimensional "assume an effect, read the power, compute one exaggeration ratio" problem with a single low-power-versus-adequate-power fork.
Abstract Reasoning¶
The signature move is conditioning on the filter — reasoning not about the distribution of effect-size estimates but about the distribution of estimates given that they cleared significance. The analyst recognizes that a published magnitude is not a draw from the sampling distribution centered on the truth but a draw from its upper tail, because only estimates far enough out to beat the threshold survive to be reported. The characteristic inference runs from low power + a significance gate to the conditional expectation of the reported effect being a multiple of the true effect: when power is 17% against a true d = 0.10, the null is rejected only when the sample estimate exceeds roughly d = 0.28, so the expected significant estimate is about d = 0.35 — an exaggeration above threefold. The move is to read the headline number as a selected order statistic, not an unbiased estimate.
This is made prospective and quantitative by the exaggeration-ratio move, which runs the reasoning forward from design to expected inflation before any data exist. The analyst supplies two inputs — an assumed true effect and the study's power against it — and computes the ratio of the expected significant estimate to the truth, converting a vague worry about inflation into a number read off the design. The inference is from power alone to how misleading a significant result will be, which licenses a design-stage verdict: a study can be adequately powered to detect an effect yet near-guaranteed to exaggerate it, so the analyst rejects or redesigns underpowered studies on the basis of their exaggeration ratio rather than their detection probability.
The reframing is itself a boundary-drawing move that re-chooses the unit of analysis. The Neyman-Pearson frame asks "what is the probability I detect the effect?" and treats the reject/fail-to-reject decision as the thing to control; the Type M move shifts the unit to the reported magnitude conditional on rejection and asks "if I detect it, how far from the truth will my number be?" The two questions are drawn apart precisely in the low-power regime — they converge when power is high (ratio near one, magnitude trustworthy) and diverge sharply when power is low (ratio large, magnitude selection-inflated). The inference is from where the study sits on the power axis to whether a significant estimate's size can be taken at face value, a fixed low-power-versus-adequate fork.
Finally there is a retrospective diagnostic move that detects the pathology in a literature without any appeal to misconduct: systematic shrinkage from an initial study to its higher-powered replication is the observable signature of Type M operating, the same exaggeration the ratio predicts prospectively now read backward from the replication record. The analyst infers from "surprising findings that consistently fail to hold up at full power" to "a significance gate selecting upward-biased draws," which separates the genuine mechanism (a magnitude inflated by selection) from the easy but wrong readings (the original was fraudulent, or merely unlucky). The interventions follow directly from the mechanism: power up to push the ratio toward one, apply shrinkage estimators to undo the conditional shift, pre-register to fix the analysis before the filter can select, and treat single-study magnitudes from underpowered designs as upper bounds rather than point estimates.
Knowledge Transfer¶
Type M error is a named statistical bias and its prospective diagnostic quantity — a selection effect on effect-size estimators summarised by the exaggeration ratio — not a causal mechanism in the world, so "mechanism within / metaphor beyond" applies only loosely. Within statistics and the significance-testing sciences it transfers as full mechanism, and what carries is the whole apparatus: the condition-on-the-filter move, the prospective exaggeration ratio computed from an assumed effect and the study's power, the reframe of the unit of analysis from the reject/fail-to-reject decision to the reported magnitude conditional on rejection, the low-power-versus-adequate fork, the replication-shrinkage retrospective signature, and the remedy set (power up, shrink, pre-register, treat single-study magnitudes as upper bounds). The mechanism recurs in identical form across fields that share the substrate of estimation from a significance-filtered sample: underpowered psychology, economics, and biomedicine (surprising effects that shrink dramatically on replication), early-phase clinical trials (effects clearing p < 0.05 that Phase III systematically fails to recover), genome-wide association studies (winner's-curse shrinkage of top hits), A/B testing (stopped-on-significance experiments exaggerating lift), and meta-science (PET-PEESE and related bias-adjustment methods explicitly correcting this inflation). In every case the exaggeration ratio, the conditional-on-significance reasoning, and the shrinkage diagnostic are the same objects; only the subject matter changes. The transfer is literal because the substrate — a noisy effect-size estimator filtered by a one-sided significance threshold under specified power — is held fixed.
Beyond the significance-testing scaffolding the situation is the shared-abstract-mechanism case, with an unusually exact parent: Type M error is the statistical-inference face of the winner_s_curse. The underlying machinery — select the extremum of a set of noisy estimates, and the selected value is biased away from the truth — genuinely recurs across domains: auction overpayment in its native bidding/order-statistic setting, the inflated apparent skill of the top performer in a noisy tournament, the disappointing realised value of whatever was chosen because it scored highest. But that recurrence is winner's curse travelling, not Type M: strip the Neyman-Pearson scaffolding and the significance filter, and the residual content is exactly selection-on-noisy-estimates, with selection_bias as the broader parent (Type M is a selection bias on effect-size estimators induced by a significance filter) and regression_to_the_mean as the cousin (the unconditional shrinkage of which Type M is the passed-a-significance-test special case). What does not travel under Type M's own steam is everything that makes it statistical: the exaggeration ratio defined against an NHST threshold, the design-stage retrospective power calculation, the p < 0.05 gate, and the companion Type S sign-error quantity. An auctioneer overpaying is not committing a "Type M error" — there is no p-value, no power, no significance filter — though both are instances of the same winner's-curse pattern.
So the cross-domain lesson — the extremum of noisy estimates, selected by a one-sided gate, systematically overstates the truth — should be carried by winner_s_curse (and the broader selection_bias family), not by "Type M error," whose contribution is to operationalise that pattern for the binary significance filter inside frequentist hypothesis testing. The portable structural reasoning belongs to the parent; the domain-specific operational content — the exaggeration ratio, the retrospective M/S calculation, the NHST-conditioned diagnosis — is what Type M uniquely contributes and what stays in statistics and experimental design (see Structural Core vs. Domain Accent).
Examples¶
Canonical¶
Gelman and Carlin's (2014) "retrodesign" makes Type M concrete as a design-stage computation. Take a study whose standard error equals the true effect it targets — say a true effect of d = 0.10 with a standard error of 0.10. A two-sided test at α = 0.05 rejects only when the estimate exceeds 1.96 standard errors, about d = 0.20, already twice the truth; the study's power is only Φ(1 − 1.96) ≈ 0.17. Conditional on clearing that gate, the reported estimate is a draw from the upper tail of a N(0.10, 0.10) distribution truncated at 0.196, whose expected value is 0.10 + 0.10·φ(0.96)/[1 − Φ(0.96)] ≈ 0.10 + 0.10·(0.252/0.169) ≈ 0.25. The exaggeration ratio — expected significant estimate over true effect — is therefore about 2.5, read off the design before any data exist.
Mapped back: The N(0.10, 0.10) sampling distribution is the noisy effect-size estimator; the 1.96-SE cutoff at d ≈ 0.20 is the significance filter; the 17% power is the low-power regime that makes only far-tail draws admissible. The truncated-upper-tail expectation of ≈ 0.25 is the conditional-on-significance shift, and its quotient with the truth, ≈ 2.5, is the exaggeration ratio computed prospectively — the whole point being that no data were needed to obtain it.
Applied / In Practice¶
Button, Ioannidis, Mokrysz, Nosek, Flint, Robinson, and Munafò's Power failure (Nature Reviews Neuroscience, 2013) surveyed meta-analyses across neuroscience and estimated the median statistical power of studies in the field at roughly 21%. Their argument runs beyond the familiar Type II worry that true effects go undetected: under such low power, the significant findings that do reach print are systematically inflated, because only estimates far out in the tail clear significance. A literature built from such studies therefore reports effect sizes that shrink when larger, higher-powered work revisits them — the observable tell they invoke to explain why so many headline neuroscience results fail to replicate at full strength.
Mapped back: The field-wide median power of ~21% is the low-power regime; the p < 0.05 publication gate is the significance filter admitting only tail draws, producing the conditional-on-significance shift in reported magnitudes. Their explanation for non-replication — inflated originals that shrink against higher-powered follow-ups — is precisely the replication-shrinkage signature, diagnosing Type M across a literature without invoking misconduct.
Structural Tensions¶
T1: Powered to detect versus powered to estimate (two questions the same design answers differently). A study can be adequately powered to detect an effect and near-guaranteed to exaggerate it, because detection probability and magnitude fidelity are distinct properties that diverge precisely in the low-power regime. The Neyman-Pearson frame optimizes the reject/fail-to-reject decision and treats a significant result as a job done; Type M shows that the very draw that cleared the gate is, under low power, a selected order statistic far from the truth. The tension is that "we found the effect" and "we measured the effect" feel like one achievement but come apart exactly where much of the literature lives — a design tuned to the first question can be silently disastrous for the second. Reporting a magnitude from a detection-powered study treats a selection-inflated number as an estimate. Diagnostic: Was this study powered to detect the effect, or powered so its exaggeration ratio sits near one — and which did the design actually optimize?
T2: Prospective computability versus dependence on the assumed truth (the ratio's virtue and its hostage). The exaggeration ratio's signal contribution is that it is computable from the design before any data exist — supply an assumed true effect and the power, read off the inflation. But that same prospectivity is bought by conditioning on the one quantity nobody knows: the true effect. Assume a smaller truth and the ratio balloons; assume a larger one and the alarm quiets, so the diagnostic can be talked up or down by the very effect-size guess it is meant to interrogate. The tension is that the ratio converts a vague worry into a concrete number only by importing an assumption whose uncertainty it cannot itself resolve — its greatest strength (no data needed) rests on its softest input (a true effect posited, not measured). Diagnostic: Is the exaggeration ratio being reported against a defensible range of plausible true effects, or against a single convenient assumption that flatters the design?
T3: Selection structure versus misconduct, luck, or heterogeneity (attributing the shrinkage). Type M's exculpatory power is real: an honest analyst with an unbiased estimator still reports an inflated magnitude, because the bias lives in the significance filter, not in fraud or an unlucky draw — so a non-replicating effect need not indict anyone. But that same framing can over-claim in the other direction, absorbing into "Type M" a shrinkage that was actually produced by p-hacking, selective reporting, or genuine effect heterogeneity across populations. The replication-shrinkage signature is consistent with Type M but not diagnostic of it alone. The tension is that the concept rightly rescues honest researchers from the fraud reading while making it tempting to file every disappointing replication under a mechanism that has innocent and non-innocent look-alikes. Diagnostic: Is the observed shrinkage explained by the significance filter under the study's actual power, or does its size exceed what Type M predicts, pointing to QRPs or true heterogeneity?
T4: Low-power bite versus adequate-power innocence (resisting the unconditional indictment). Type M is regime-specific: when power is high the exaggeration ratio sits near one and a significant magnitude can be taken roughly at face value, so the bias is an indictment of underpowered designs, not of significant effect sizes in general. The vivid low-power finding invites over-generalization — "all published effects are inflated" — which is false and corrosive, discarding trustworthy magnitudes from well-powered work alongside the selection-inflated ones. The tension is that the concept's rhetorical force comes from its dramatic low-power multiples, yet its correct application requires holding the line that adequate power neutralizes it. Applying the discount uniformly punishes the well-powered study for the sins of the underpowered one, while ignoring the regime split lets a genuinely inflated headline pass. Diagnostic: Does this study sit in the low-power regime where the ratio climbs, or is it powered enough that its significant magnitude needs no exaggeration discount?
T5: Autonomy versus reduction (a named statistical bias or the NHST face of the winner's curse). Type M is a genuine named construct with operational machinery of its own — the exaggeration ratio defined against a significance threshold, the retrospective power calculation, the companion Type S sign error — and within the significance-testing sciences it transfers as full mechanism. But its portable core is not proprietary: strip the Neyman-Pearson scaffolding and what remains is select the extremum of noisy estimates through a one-sided gate and the selected value overstates the truth — which is exactly winner_s_curse (with selection_bias as the broader parent and regression_to_the_mean as cousin). An auctioneer overpaying commits no "Type M error"; there is no p-value or power, yet it is the same selection pattern. The tension is between a standalone statistical construct worth its own operational toolkit and the recognition that its cross-domain lesson belongs to the winner's-curse family. Diagnostic: Resolve toward winner_s_curse and selection_bias when carrying the lesson outside hypothesis testing; toward named Type M error when the substrate is an effect-size estimate filtered by a significance threshold.
Structural–Framed Character¶
Type M error sits in the mixed band of the spectrum: a genuinely mathematical selection mechanism wrapped in — and inseparable from — the human methodological practice of null-hypothesis significance testing. The criteria split. Two point structural. Its evaluative_weight is essentially nil: the construct names a selection structure, not a verdict, and is emphatic that the inflation is "not fraud or bad luck" and "not a flaw in the estimator" — an honest analyst with an unbiased estimator still reports an exaggerated magnitude, so the word describes a mechanism rather than convicting anyone. And on import_vs_recognize, within the significance-testing sciences the mechanism transfers as recognition of the same object — underpowered psychology, early-phase trials, GWAS top-hit shrinkage, stopped-on-significance A/B tests, and meta-science bias corrections are recognized as the identical exaggeration-ratio-and-shrinkage structure, only the subject matter changing.
Three criteria point framed and hold it mid-spectrum. It is human_practice_bound in a specific sense: the construct presupposes the significance-testing scaffolding — a p-value, a power calculation, a one-sided significance gate — and dissolves without it; the entry states plainly that "an auctioneer overpaying is not committing a Type M error — there is no p-value, no power, no significance filter." Its institutional_origin is a methodological artifact: the exaggeration ratio defined against an NHST threshold, the retrospective design-stage power calculation, the p < 0.05 gate, and the companion Type S sign-error quantity are all constructs of the frequentist-hypothesis-testing framework (Gelman and Carlin's formalization), not facts of nature a survey reads off. And vocab_travels fails: exaggeration ratio, significance filter, power, conditional-on-significance, replication shrinkage are pinned to the significance-testing substrate.
The portable structural skeleton is single and mathematical: select the extremum of a set of noisy estimates through a one-sided gate, and the selected value is systematically biased away from the truth. That skeleton is exactly what Type M error instantiates from its parent prime winner_s_curse (the statistical-inference face of it), with selection_bias as the broader parent and regression_to_the_mean as the cousin of which Type M is the passed-a-significance-test special case: the cross-domain reach — auction overpayment, the inflated apparent skill of a noisy tournament's top performer, the disappointing realized value of whatever was chosen because it scored highest — belongs to those umbrellas, whose instances are co-instances of selection-on-noisy-estimates, not applications of Type M. The named construct's distinctive content — the exaggeration ratio against an NHST threshold, the retrospective power calculation, the p-value gate, the Type S companion — is precisely the home-bound cargo that does not lift. Its character: an evaluatively neutral, recognized-within-statistics operationalization of the winner's-curse selection mechanism for the binary significance filter, structural in the select-the-noisy-extremum skeleton it borrows from winner_s_curse but bound to its home domain by the NHST scaffolding that gives it its operational form, leaving it mixed rather than a free-floating prime.
Structural Core vs. Domain Accent¶
This section decides why Type M error is a domain-specific abstraction and not a prime — a case with an unusually exact parent, because Type M is precisely the statistical-inference face of a more general selection pattern.
What is skeletal (could lift toward a cross-domain prime). Strip the significance-testing scaffolding and a thin, mathematical relational structure survives: select the extremum of a set of noisy estimates through a one-sided gate, and the selected value is systematically biased away from the truth in the direction of the selection. Stated abstractly this is winner_s_curse — Type M is exactly its statistical-inference face — with selection_bias as the broader parent (Type M is a selection bias on effect-size estimators) and regression_to_the_mean as the cousin (the unconditional shrinkage of which Type M is the passed-a-significance-test special case). This skeleton is genuinely substrate-portable: it governs auction overpayment in its native bidding setting, the inflated apparent skill of a noisy tournament's top performer, and the disappointing realized value of whatever was chosen because it scored highest. It is the core Type M shares, not what makes it distinctive.
What is domain-bound. Everything that makes the construct Type M error in particular is null-hypothesis-significance-testing machinery that does not survive extraction. The one-sided gate is specifically a p < 0.05 significance filter; the noisy estimator is an effect-size sampling distribution; the diagnostic scalar is the exaggeration ratio defined against an NHST threshold (expected significant estimate over true effect); the design-stage computation is a retrospective power calculation run before any data exist; the unit-of-analysis reframe is stated against Type I/II decision error; the retrospective tell is replication shrinkage from an initial study to a higher-powered replication; and the construct travels paired with its Type S sign-error companion. The decisive test the entry supplies: an auctioneer overpaying is not committing a "Type M error" — there is no p-value, no power, no significance filter — even though it is the same winner's-curse pattern. Remove the significance scaffolding and what remains is bare selection-on-noisy-estimates, a looser thing that no longer wears this name.
Why this does not clear the prime bar. A prime's vocabulary travels and its cross-domain transfer is recognition of the same mechanism, not analogy. Type M's transfer is bimodal. Within statistics and the significance-testing sciences the mechanism travels in identical form by genuine recognition — the condition-on-the-filter move, the prospective exaggeration ratio, the decision-to-magnitude reframe, the low-power-versus-adequate fork, the replication-shrinkage signature, and the remedy set (power up, shrink, pre-register, treat single-study magnitudes as upper bounds) are the same objects across underpowered psychology/economics/biomedicine, early-phase clinical trials, GWAS top-hit shrinkage, stopped-on-significance A/B testing, and meta-science bias corrections, only the subject matter changing, because the substrate (a noisy effect-size estimator filtered by a one-sided significance threshold under specified power) is held fixed. Beyond the significance-testing scaffolding the named construct does not travel; what recurs is the winner's curse, not Type M. So when the bare structural lesson is needed elsewhere — the extremum of noisy estimates, selected by a one-sided gate, systematically overstates the truth — it is already carried, in general substrate-neutral form, by winner_s_curse (with selection_bias and regression_to_the_mean in the family), whose instances are co-instances of selection-on-noisy-estimates rather than exports of Type M. The cross-domain reach belongs to those parents; Type M's distinctive content — the exaggeration ratio against an NHST threshold, the retrospective power calculation, the p-value gate, the Type S companion — is exactly the home-bound cargo that should stay in statistics and experimental design. Type M error clears the domain-specific bar comfortably for the significance-testing sciences, but its only substrate-spanning content is the winner's-curse selection pattern the parent prime already carries; its own contribution is to operationalize that pattern for the binary significance filter.
Relationships to Other Abstractions¶
Current abstraction Type M Error Domain-specific
Parents (5) — more general patterns this builds on
-
Type M Error is a kind of Selection on Noisy Estimates Prime
Type M is the significance-threshold species in which selection inflates a reported effect magnitude.Selection on Noisy Estimates supplies the genus: When noisy estimates help determine which candidates, results, or options are selected, conditioning on selection shifts the selected estimates toward the favored tail and makes them systematically overstate their latent values. Type M Error preserves that general structure while adding its differentia: Quantify how much a significant effect's reported magnitude is exaggerated by the significance filter under low power, via the exaggeration ratio — the expected significant estimate divided by the true effect — computable from the design before any data exist. The parent can occur without those added commitments, whereas removing the parent structure leaves no basis for classifying the child as this subtype. That asymmetry establishes subsumption rather than mere association.
-
Type M Error is part of Effect Size Prime
Type M Error contains observed and assumed true Effect Sizes whose magnitudes form the selected estimate and comparison target.Effect Size supplies the direction-bearing magnitudes for the estimator and assumed truth. Type M adds their sampling distribution, selection through significance, expected admitted value, quotient diagnostic, and design-stage interpretation under low power.
-
Type M Error is part of Ratio Prime
Type M contains an exaggeration ratio whose numerator is expected significant effect magnitude and whose denominator is true effect magnitude.Ratio supplies an internal constituent: Compare one quantity with a nonzero reference quantity by division, so the quotient states how much numerator obtains per unit of denominator and stays interpretable only while both quantities, their units, and their scope are named. Type M Error requires that role within this mechanism: Quantify how much a significant effect's reported magnitude is exaggerated by the significance filter under low power, via the exaggeration ratio — the expected significant estimate divided by the true effect — computable from the design before any data exist. Remove the parent-role and the child loses a required internal operation, even though the parent can exist outside the child. The child is therefore built from the parent rather than being a taxonomic kind of it.
-
Type M Error is part of Statistical Power Prime
Type M Error contains the design's Statistical Power as the regime variable governing how far admitted estimates must lie from the truth.Power supplies the probability that the design rejects a false null for an assumed true effect. Type M reuses that design distribution conditionally: low power forces threshold-clearing draws into the far tail and raises the exaggeration ratio, while adequate power moves the ratio toward one.
-
Type M Error is part of Statistical Significance (p-Value) Prime
Type M Error contains a statistical-significance gate that selects the tail estimates whose conditional magnitude is evaluated.Statistical Significance supplies the null reference, tail probability, and rejection threshold. Type M adds a noisy effect-size estimator, low-power conditioning, the expected admitted magnitude, an exaggeration ratio, and a replication-shrinkage interpretation.
Hierarchy paths (24) — routes to 8 parentless roots
- Type M Error → Selection on Noisy Estimates → Selection Bias → Bias
- Type M Error → Effect Size → Scale
- Type M Error → Statistical Significance (p-Value) → Statistical Inference → Inductive Reasoning
- Type M Error → Effect Size → Comparison → Self Checking
- Type M Error → Ratio → Comparison → Self Checking
- Type M Error → Statistical Significance (p-Value) → Statistical Inference → Uncertainty
- Type M Error → Selection on Noisy Estimates → Selection Bias → Statistical Inference → Inductive Reasoning
- Type M Error → Statistical Significance (p-Value) → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Inductive Reasoning
- Type M Error → Statistical Power → Experimental Design → Comparison → Self Checking
- Type M Error → Statistical Power → Probability → Measure → Set and Membership
- Type M Error → Statistical Significance (p-Value) → Probability → Measure → Set and Membership
- Type M Error → Selection on Noisy Estimates → Selection Bias → Statistical Inference → Uncertainty
- Type M Error → Statistical Significance (p-Value) → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Uncertainty
- Type M Error → Selection on Noisy Estimates → Selection Bias → Vantage-Induced Omission → Viewpoint
- Type M Error → Statistical Power → Probability → Measure → Aggregation → Micro Macro Linkage
- Type M Error → Statistical Significance (p-Value) → Probability → Measure → Aggregation → Micro Macro Linkage
- Type M Error → Statistical Power → Experimental Design → Control Sample → Comparison → Self Checking
- Type M Error → Statistical Significance (p-Value) → Statistical Inference → Probability → Measure → Set and Membership
- Type M Error → Statistical Significance (p-Value) → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
- Type M Error → Statistical Significance (p-Value) → Hypothesis Testing (Null vs. Alternative) → Verification → Evaluation → Comparison → Self Checking
- Type M Error → Selection on Noisy Estimates → Selection Bias → Statistical Inference → Probability → Measure → Set and Membership
- Type M Error → Statistical Significance (p-Value) → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Probability → Measure → Set and Membership
- Type M Error → Selection on Noisy Estimates → Selection Bias → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
- Type M Error → Statistical Significance (p-Value) → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
Not to Be Confused With¶
- Type S (sign) error. The sign-sibling, computed jointly from the same design — right-direction-but-exaggerated (Type M) versus reversed-direction (Type S). They share the conditional-on-significance machinery, but Type M quantifies how much the magnitude is inflated (the exaggeration ratio), while Type S quantifies the probability the direction is wrong (Pr(wrong sign)). Type M is the common case; Type S bites only when the true effect is near zero relative to noise. Tell: is the worry that the significant estimate overstates a correctly-signed effect (Type M), or that it points the wrong way entirely (Type S)?
- Type I / Type II error. The classical Neyman–Pearson pair, framed around the reject/fail-to-reject decision — Type I falsely rejects a true null, Type II fails to detect a real effect. Type M leaves the decision aside and asks about the reported magnitude conditional on rejection: given a significant result, how far from the truth is its size? It is parasitic on the same significance scaffolding but changes the unit of analysis from the decision to the number. Tell: is the question whether the detect/miss decision was right (Type I/II), or how inflated the estimate is given that it was significant (Type M)?
- Regression to the mean. The cousin — extreme measurements tend to be followed by less-extreme ones, an unconditional shrinkage toward the average. Type M is the special case of this shrinkage induced specifically by passing a significance test: the significance filter selects upper-tail draws that then regress at replication. Regression to the mean is the broader statistical fact; Type M is its significance-gated, effect-size operationalization. Tell: is the shrinkage a general tendency of extreme values to moderate (regression to the mean), or specifically the inflation of estimates selected by a p < 0.05 gate under low power (Type M)?
- Publication bias. The field-level selection whereby significant/positive results are more likely to be published, distorting the literature. Type M operates within a single study's significance filter (the study's own gate under low power), though the two compound: publication bias is the file-drawer selection across studies, Type M the magnitude inflation each underpowered surviving study already carries. Tell: is the distortion that the literature over-represents significant studies (publication bias), or that each underpowered significant study's own reported size is inflated (Type M)?
winner_s_curse/selection_bias(parent primes), withregression_to_the_mean. The substrate-neutral skeleton Type M instantiates — select the extremum of noisy estimates through a one-sided gate and the selected value overstates the truth. This is what travels (auction overpayment, the top noisy-tournament performer's inflated apparent skill), while the exaggeration ratio and NHST gate stay home. It is the umbrella, not a peer confusable; an auctioneer overpaying commits no "Type M error." Tell: is the lesson the generic select-the-noisy-extremum-overstates pattern on any substrate (the parents), or the specific significance-filtered effect-size inflation with its exaggeration ratio (the named construct)? (Treated fully in a later section.)
Neighborhood in Abstraction Space¶
Type M Error sits in a moderately populated region (55th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Unclustered & Miscellaneous (309 abstractions)
Nearest neighbors
- Type S Error — 0.89
- Funnel Plot Asymmetry — 0.85
- File Drawer Problem — 0.85
- Small-Study Effects — 0.83
- Jeffreys-Lindley Paradox — 0.82
Computed from structural-signature embeddings · 2026-07-12