Skip to content

Type S Error

Quantify the risk that a statistically significant estimate points the wrong way by computing, before data collection, the probability that a two-sided significance filter is cleared from the opposite tail when the true effect is near zero relative to noise.

Core Idea

A Type S (sign) error is the failure mode in which a statistically significant estimate has the wrong sign — claiming a positive effect when the true effect is negative, or vice versa — because the study's power is so low relative to the true effect size that the significance threshold is passed almost exclusively by large-magnitude draws from the sampling distribution, and when the true effect is near zero relative to sampling noise, both tails of the distribution can generate significant results with similar probability. Gelman and Carlin (2014) introduced Type S alongside Type M to reframe inferential risk away from the classical null-versus-alternative binary and onto the question: given that I will only act on a significant result, how likely is my significant result to correctly identify the direction of the true effect? In extreme low-power settings — a study powered at 6% to detect a true effect of d = 0.05 — the probability of a sign-correct result given significance can fall to roughly 0.76, meaning roughly one in four significant results points the wrong way. At power levels near 50%, the conditional probability of a correct sign is still well below 1 if the true effect is small relative to sampling variation.

The mechanism is the same selection process that produces Type M exaggeration, but evaluated at the sign rather than the magnitude. A two-sided significance filter selects estimates whose absolute value exceeds a threshold; when the true effect is small, both positive and negative values above that threshold are reachable from the sampling distribution, and the relative frequencies with which each tail is reached depend on how far the true effect is from zero relative to the threshold. The closer the true effect is to zero, the more evenly the two tails contribute significant observations, and the closer Pr(sign correct | significant) is to 0.5. Sign flips between high-profile original studies and their high-powered replications — multiple reversal cases in behavioral science — are the downstream observable signature of Type S errors materializing across a literature. The diagnostic use of Type S is prospective: compute Pr(wrong sign | significant) at design time, before data collection, to determine whether a proposed study is informative about the direction of the effect even if it achieves significance.

Structural Signature

Sig role-phrases:

  • the near-zero noisy estimator — a sampling distribution whose true effect sits close to zero relative to sampling noise, the low-power regime that makes sign error possible
  • the two-sided significance filter — the gate admitting estimates whose absolute value clears a threshold, reachable from both tails when the truth is near zero
  • the two-tail reachability — the geometry whereby, the closer the true effect is to zero relative to the standard error, the more evenly both tails contribute significant draws
  • the conditional sign-flip — the failure: a significant estimate carrying the wrong direction because the opposite tail retains non-trivial mass (a 6%-power study against d=0.05 points wrong ~24% of the time)
  • the sign-error probability — the prospective diagnostic scalar Pr(sign correct | significant), computable at design time from the effect-to-noise ratio and power
  • the near-zero-versus-far fork — the boundary: far from zero only the correct tail is reachably significant and the sign is safe; near zero the probability slides toward 0.5 and the direction is barely better than a coin flip
  • the power-not-precision remedy — the diagnosis-fixed prescription that the cure is enough power to push opposite-tail mass toward negligible, not better measurement of magnitude
  • the sign-reversal signature — the retrospective tell: a direction flip between an original and its higher-powered replication, read as a materialized Type S risk rather than an unstable phenomenon
  • the joint pairing with Type M — the operational unit: exaggeration ratio and Pr(wrong sign) computed together from one design, jointly judging whether a significant result's size and direction are both worth believing

What It Is Not

  • Not a Type I or Type II error. Those frame error around the null-versus-alternative decision; Type S frames it around the sign of the estimate conditional on rejection — "given that I will only report a significant result, how often does it point the wrong way?" It presupposes the same hypothesis-testing scaffolding but asks a question about direction that the reject/fail-to-reject binary never poses.
  • Not the same as Type M error. Type M is the magnitude sibling — right direction, exaggerated size — and is the more common failure; Type S is the sharper, rarer alarm of a reversed direction. They share the conditional-on-significance machinery and are computed jointly, but a wrong sign and an inflated magnitude are distinct verdicts on the same significant result.
  • Not regression to the mean. Regression to the mean shrinks extreme follow-up measurements toward the average but does not, in general, flip the sign of a true effect that is far from zero. Type S flips signs precisely in the near-zero regime, where both tails of the filtered distribution are reachable — a different phenomenon with a different trigger.
  • Not evidence of fraud or an inherently unstable phenomenon. A sign reversal between a high-profile original and its high-powered replication is the predictable materialization of a Type S risk already present in the underpowered original design, not a scandal and not proof that the effect is capricious. The framework reads direction-reversals as a diagnostic of underpowered sign claims rather than as instability.
  • Not fixable by better measurement of magnitude. The danger is opposite-tail mass under low power, so the remedy is enough power to push that mass toward negligible, not more precise measurement of the effect's size. Improving magnitude precision without addressing power leaves the sign risk untouched.
  • Not a risk when the true effect is far from zero. When the true effect is large relative to sampling noise, only the correct tail is reachably significant and the sign is safe — Pr(sign correct | significant) sits near one. The sign-flip danger is specific to true effects near zero relative to the standard error; it is not a universal indictment of all significant direction claims.

Scope of Application

Type S error lives across the significance-testing sciences that share the substrate of estimation from a two-sided significance-filtered sample with true effects near zero relative to noise; its reach is within that one statistical domain, the mechanism recurring in identical form across them (always evaluated jointly with its magnitude-sibling Type M). The broader pattern — select noisy estimates near zero through a gate and the selected extremum can carry the wrong sign — travels under winner_s_curse and the selection_bias family, not under this construct, which needs a two-sided NHST filter, a p-value, and power to be the thing it is.

  • Behavioural-science replication crisis — priming and embodied-cognition reversals between high-profile originals and their high-powered replications, read as materialized Type S risk.
  • Health-policy A/B tests — underpowered studies reporting a harmful or beneficial direction when the true effect is opposite.
  • Genome-wide association studies — top discovery-stage SNP hits flipping sign at the replication stage when the true effect is small.
  • Election-poll subgroup analyses — small-cell subgroup swings reversing direction across consecutive surveys.

Clarity

Naming the Type S error punctures the rhetorical comfort of "at least the direction is right" — the implicit fallback a researcher reaches for when an effect size looks shaky but the sign of a significant result seems safe. Type S shows this fallback is exactly the assumption that fails under low power: when the true effect is small relative to sampling noise, a two-sided significance filter is reachable from both tails, so a significant result can carry almost no information about which way the effect actually points. The distinction it sharpens is between two questions the classical null-versus-alternative frame fuses — "is there an effect?" versus "given that I will only act on a significant result, is its direction trustworthy?" — and it makes vivid that significance can be achieved while the sign is barely better than a coin flip.

Its clarifying force is to convert "we found a positive effect" from an apparently safe claim into one with a computable failure probability, Pr(wrong sign | significant), available at design time before any data exist. This localizes the danger to a specific regime — true effects near zero relative to standard error — and tells the analyst that the remedy is not better measurement of magnitude but enough power to push the opposite-tail mass toward negligible. It also names the downstream signature precisely: a sign reversal between a high-profile original and its high-powered replication is not a scandal or a fluke but the predictable materialization of a Type S risk that was present in the original design, which lets a reader read direction-reversals in a literature as a diagnostic of underpowered sign claims rather than as evidence that the phenomenon is simply unstable.

Manages Complexity

Whether the direction of a reported effect can be trusted is, in practice, a worry that proliferates wherever estimates are read off small or noisy cells: an interaction term, a subgroup contrast, an exploratory comparison, a discovery-stage SNP, a subgroup swing in an election poll. Each such direction claim seems to demand its own judgment about how seriously to take "the effect went this way," and across a literature the sign reversals between high-profile originals and their replications can look like an unruly collection of one-off embarrassments or evidence that the phenomena themselves are simply unstable. Type S error compresses all of these onto a single mechanism and a single prospective scalar. The mechanism is the same significance-filter selection that drives magnitude inflation, evaluated at the sign instead of the size; the quantity that governs it is one number — Pr(sign correct | significant) — computable from the design before any data are collected. A scattered set of direction-of-effect anxieties collapses to one probability of pointing the wrong way.

What the analyst tracks reduces to essentially one thing: how large the true effect is relative to sampling noise, because that ratio alone sets how the two tails of the distribution contribute significant observations, and the qualitative verdict on a direction claim reads off it through a fixed regime split. A two-sided significance filter selects estimates whose absolute value clears a threshold; when the true effect is far from zero relative to the standard error, only the correct tail is reachably significant and the sign is safe, so Pr(sign correct | significant) sits near one and "we found a positive effect" can be believed. When the true effect is near zero relative to noise — the low-power regime — both tails clear the threshold with comparable probability, Pr(sign correct | significant) slides toward 0.5, and a significant result carries almost no information about which way the effect actually points; in extreme cases roughly one significant result in four points the wrong way. The verdict thus turns on a single boundary (true effect near versus far from zero relative to standard error), and the remedy the branch prescribes is specific: not better measurement of magnitude but enough power to push the opposite-tail mass toward negligible. The structure also fixes the downstream reading — a sign reversal between an original and its higher-powered replication is the predictable materialization of a Type S risk already present in the original design, so direction-reversals across a literature are read as a diagnostic of underpowered sign claims rather than as proof of an unstable phenomenon. A high-dimensional "can I trust the direction of each of these many contrasts" problem becomes a low-dimensional "compute one Pr(wrong sign) from the effect-to-noise ratio" problem with a single near-zero-versus-far-from-zero fork. The pair with Type M is the operational unit — exaggeration ratio and sign-error probability computed together from the same design — so one prospective calculation answers, jointly, whether a significant result's magnitude and its direction are both worth believing.

Abstract Reasoning

The defining move is conditioning the sign on the filter — asking not "is the estimate's direction right?" but "given that I will only report a significant estimate, how often does that significant estimate point the wrong way?" The analyst reasons from a two-sided significance gate to the geometry of which tails are reachable: the filter selects estimates whose absolute value clears a threshold, so when the true effect sits far from zero relative to sampling noise only the correct tail is reachably significant and the sign is safe, but when the true effect is near zero both tails clear the threshold with comparable probability. The characteristic inference runs from the effect-to-noise ratio to Pr(sign correct | significant): the closer the truth is to zero relative to the standard error, the more evenly the two tails contribute significant draws, and the closer that probability slides toward 0.5 — in extreme low-power cases roughly one significant result in four points the wrong way.

This converts an apparently safe claim into one with a computable failure probability available before data exist. The analyst computes Pr(wrong sign | significant) at design time from the assumed effect and the power, turning "we found a positive effect" into a statement carrying a known chance of being directionally false. The inference is from design parameters alone to whether a significant result will be trustworthy about direction, which licenses a prospective verdict: a study can be near-certain to reach significance yet provide almost no information about which way the effect points, so the analyst judges its informativeness about direction separately from its probability of detection.

A boundary-drawing move punctures the specific fallback the classical frame invites — "at least the direction is right" — by locating exactly the regime where it fails. The analyst reasons that this comfort holds when the true effect is far from zero relative to noise and collapses when it is near zero, so the move is to identify which side of that boundary a given contrast sits on before trusting its sign. Crucially this also fixes the remedy by diagnosis: because the danger is opposite-tail mass and not mismeasurement, the prescribed fix is not better measurement of magnitude but enough power to push the opposite-tail mass toward negligible — the inference runs from the cause of the sign risk to the one intervention that removes it, ruling out precision improvements that would not touch it.

The retrospective diagnostic move reads direction-reversals across a literature as signal rather than scandal: a sign flip between a high-profile original and its high-powered replication is inferred to be the predictable materialization of a Type S risk already present in the original design, not evidence that the phenomenon is inherently unstable nor that the original was fraudulent. The analyst reasons backward from observed reversals to underpowered sign claims in the originals. And because the same significance-filter selection drives both magnitude and direction failures, the framework licenses a joint move: compute the exaggeration ratio and Pr(wrong sign) together from one design, so a single prospective calculation answers whether a significant result's size and its direction are both worth believing — with the natural intervention being to prefer interval estimation and meta-analytic pooling over a dichotomous single-study sign claim whenever the effect-to-noise ratio puts the study in the danger regime.

Knowledge Transfer

Type S error is a named statistical bias and its prospective diagnostic quantity — the probability that a significant estimate points the wrong way, summarised by Pr(sign correct | significant) — not a causal mechanism in the world, so "mechanism within / metaphor beyond" applies only loosely. Within statistics and the significance-testing sciences it transfers as full mechanism, and what carries is the whole apparatus: the condition-the-sign-on-the-filter move, the two-tail reachability geometry, the prospective Pr(wrong sign) computed at design time from an assumed effect and the power, the near-zero-versus-far-from-zero regime fork, the puncturing of the "at least the direction is right" fallback, the power-up (not better-measurement) remedy, and the sign-reversal-on-replication retrospective signature. The mechanism recurs in identical form across fields that share the substrate of estimation from a two-sided significance-filtered sample with true effects near zero relative to noise: the behavioural-science replication crisis (priming and embodied-cognition reversals between originals and high-powered replications), health-policy A/B tests (underpowered studies reporting a harmful or beneficial direction when the truth is opposite), genome-wide association studies (top discovery-stage SNP hits flipping sign at replication), and election-poll subgroup analyses (small-cell subgroup swings reversing direction across consecutive surveys). It is the sign-sibling of Type M, always evaluated jointly with it — one prospective calculation answering whether a significant result's magnitude and its direction are both worth believing. The transfer is literal because the substrate — a noisy estimator near zero, filtered by a two-sided significance threshold — is held fixed.

Beyond the significance-testing scaffolding the situation is the shared-abstract-mechanism case, with the same exact parent Type M inherits: the cross-domain structural pattern is the winner_s_curseselection on a noisy estimator produces a biased conditional distribution that, when the true value is near zero, has non-trivial mass on the opposite tail, so the selected extremum can carry the wrong sign. That pattern recurs wherever noisy estimates are selected near zero, with selection_bias as the broader parent (here the selector is specifically a two-sided NHST filter). But that recurrence is winner's-curse-near-zero travelling, not Type S: strip the frequentist hypothesis-testing scaffolding and the significance gate, and the residual content is selection-on-noisy-estimates with an opposite-tail flip. What does not travel under Type S's own steam are the inference-domain particulars that make it Type S — the specific quantity Pr(wrong sign | significant), the specific two-sided NHST filter, and the design-stage retrospective inside frequentist testing; and note the boundary to its cousin regression_to_the_mean, which shrinks extreme follow-ups but does not in general flip the sign of true effects far from zero, whereas Type S flips signs precisely in the near-zero regime.

So the cross-domain lesson — select the extremum of noisy estimates near zero and you can get not just an exaggerated magnitude but a reversed direction — should be carried by winner_s_curse (and the broader selection_bias family), not by "Type S error," whose contribution is to operationalise that flip for the binary significance filter inside frequentist hypothesis testing. The portable structural reasoning belongs to the parent; the teachable, computable, replication-crisis-relevant operational content — Pr(wrong sign | significant), the joint Type M/Type S design-stage calculation, the NHST-conditioned sign diagnosis — is what Type S uniquely contributes and what stays in statistics and experimental design (see Structural Core vs. Domain Accent).

Examples

Canonical

Gelman and Carlin's 2014 "Beyond Power Calculations" (Perspectives on Psychological Science) introduced the prospective Type S/Type M design analysis. Take their extreme illustration: a study powered at only 6% to detect a standardized true effect of d = 0.05 — a genuine effect that is tiny relative to sampling noise. Because the two-sided significance gate admits only estimates whose absolute value clears the threshold, and the true effect sits so close to zero, both tails of the sampling distribution reach significance with comparable probability. Computing Pr(sign correct | significant) under these design parameters gives roughly 0.76 — so about one significant result in four points the wrong way, and (its Type M partner) the surviving significant estimates exaggerate the magnitude several-fold. The calculation uses only the assumed effect and the power; no data are needed.

Mapped back: The d = 0.05 truth in a 6%-power design is the near-zero noisy estimator; the two-sided p < .05 gate is the two-sided significance filter, and its reachability from both tails is the two-tail reachability. The 0.76 figure is the sign-error probability Pr(sign correct | significant), computed at design time; the ~24% wrong-way outcome is the conditional sign-flip. Reporting it beside the exaggeration ratio is the joint pairing with Type M.

Applied / In Practice

Genome-wide association studies operationalize the Type S danger as standing practice. Each candidate SNP has a true effect that is typically minuscule relative to per-marker sampling noise, and hundreds of thousands of markers pass through a significance filter — precisely the near-zero-estimator-through-a-gate regime. The field learned that discovery-stage "top hits" not only exaggerate their effect sizes but can carry the wrong allele direction, and it responded not with more precise genotyping but with an independent replication cohort requirement and a stringent genome-wide significance threshold (commonly p < 5×10⁻⁸) that raises power against the multiplicity of tests. A discovery-stage association that reverses sign in the replication cohort is read as an underpowered original, not as an unstable biological effect.

Mapped back: The tiny per-SNP effects amid genotyping noise are the near-zero noisy estimator; the genome-wide threshold is the two-sided significance filter. Requiring replication plus a stricter threshold — rather than better measurement — is the power-not-precision remedy. A hit that flips allele direction at replication is the sign-reversal signature, read as a materialized Type S risk in the original.

Structural Tensions

T1: Detection versus direction (significance answers a question about existence, not about sign). The classical frame poses "is there an effect?" and treats a significant result as evidence that there is one; Type S insists that clearing a two-sided filter and correctly identifying which way the effect points are separate achievements. A study can be near-certain to reach significance and still be almost uninformative about direction, because when the truth is near zero both tails clear the gate. The "at least the direction is right" fallback fuses the two questions and fails exactly where power is lowest. The tension is that the very result researchers treat as the payoff — a significant estimate — is silent on the property they most want from it, and the two only coincide when the true effect is far from zero relative to noise. Diagnostic: Is the study's high probability of reaching significance being read as high probability that its sign is correct, when those are different quantities?

T2: A design-time verdict versus its own missing input (Pr(wrong sign) needs the effect you do not know). Type S's signal virtue is that Pr(sign correct | significant) is computable before any data exist, from the assumed effect and the power — a prospective check that damns an uninformative design in advance. But the computation's key input is the true effect size, which is precisely what the study is being run to learn; the analyst must posit the near-zero magnitude whose plausibility is the whole question. The tool is at its most valuable exactly in the low-power regime where the assumed effect is least knowable, so the honest use is a sensitivity sweep across plausible small effects, not a single confident number. The prospective power is real, but it rests on an assumption of the same class as the one it is auditing. Diagnostic: Is the Pr(wrong sign) figure driven by a defensible range of assumed effects, or by a single guessed magnitude smuggled in as if known?

T3: Power versus precision (the fix is more of the thing intuition does not reach for). Confronted with a shaky direction claim, the reflexive corrective is to measure the effect more carefully — better instruments, tighter magnitude estimates. Type S locates the danger elsewhere: the risk is opposite-tail mass under low power, so the only intervention that removes it is enough power to push that mass toward negligible, and refining magnitude precision without raising power leaves the sign risk untouched. The tension is that the remedy is costly (larger samples, replication cohorts, stricter thresholds) and counter-intuitive, aimed at the estimator's reachability geometry rather than at the apparent problem of measurement quality — so effort spent on the intuitive fix is effort that does not touch the failure. Diagnostic: Does the proposed remedy raise power against the opposite tail, or merely sharpen magnitude measurement that leaves the sign-flip probability where it was?

T4: Sign error versus magnitude error (two verdicts on one significant result, computed together but distinct). Type S and Type M share the conditional-on-significance machinery and are calculated jointly from one design, yet they deliver different judgments: Type M says the reported magnitude is inflated (right direction, exaggerated size) and is the common case; Type S says the direction itself is reversed and is the rarer, sharper alarm confined to the near-zero regime. The tension is that reporting them together can blur which is biting — a result may be badly exaggerated but sign-safe, or modestly sized but sign-unstable — and treating "the estimate is untrustworthy" as one undifferentiated failure loses the distinction between a size that needs discounting and a direction that cannot be believed at all. Diagnostic: For this significant result, is the live risk an inflated magnitude with a safe sign, or a genuinely reversible direction — and does the effect-to-noise ratio put it near enough to zero for the latter?

T5: Reversal as materialized design risk versus reversal as genuine instability or misconduct (reading direction-flips correctly). Type S reframes a sign flip between a high-profile original and its high-powered replication as the predictable materialization of a risk already present in the underpowered original — not a scandal, not proof the phenomenon is capricious. That reading is exhaustively fair to the original researchers and dissolves a great deal of misplaced accusation. But the same charitable frame can over-exculpate: some direction-reversals really do signal an unstable or context-dependent effect, or a compromised design, and attributing every flip to mere low power risks treating a substantively fragile finding as a sound one that was simply under-resourced. The frame is a corrective against panic and blame, but it can also launder genuine instability as a power problem. Diagnostic: Was the original design demonstrably underpowered in the near-zero regime — making the flip predictable from power alone — or does the reversal survive that explanation and point to a real instability?

T6: Autonomy versus reduction (its own named NHST diagnostic or the significance-filter instance of the winner's curse). Type S is a named statistical construct with proprietary, teachable content — the quantity Pr(sign correct | significant), the two-sided p-value gate, the design-stage joint calculation with Type M, the replication-crisis vocabulary. Yet its portable structural core is not proprietary: strip the frequentist scaffolding and what remains is winner_s_curse — selection on a noisy estimator produces a biased conditional distribution that, when the truth is near zero, retains non-trivial opposite-tail mass so the selected extremum can carry the wrong sign — under the broader selection_bias family. That pattern travels wherever noisy estimates are selected near zero, with or without a p-value. The tension is between a standalone construct that earns its own operational apparatus and the recognition that the cross-domain lesson (select near-zero noisy estimates and you can get a reversed sign) already belongs to its parent. Diagnostic: Resolve toward winner_s_curse / selection_bias when carrying the flip to any selection-on-noise setting; toward Type S when operationalizing it for a specific two-sided significance filter inside frequentist hypothesis testing.

Structural–Framed Character

Type S error sits in the mixed band of the spectrum, on the same footing as its magnitude-sibling Type M: a genuinely mathematical selection mechanism inseparable from the human methodological practice of null-hypothesis significance testing. The criteria split. Two point structural. Its evaluative_weight is essentially nil: the construct names a selection geometry (two-tail reachability under low power), not a verdict, and is emphatic that a sign reversal is "not evidence of fraud or an inherently unstable phenomenon" but the predictable materialization of a design risk — mechanism-talk, not condemnation. And on import_vs_recognize, within the significance-testing sciences the mechanism transfers as recognition of the same object — behavioural-science reversals, health-policy A/B tests, GWAS allele-direction flips, and election-poll subgroup swings are recognized as the identical two-tail-reachability structure, only the subject matter changing.

Three criteria point framed and hold it mid-spectrum. It is human_practice_bound in the specific NHST sense: the construct presupposes a two-sided significance gate, a p-value, and a power calculation, and dissolves without them; strip the frequentist scaffolding and what remains is not "Type S error" but a bare selection effect. Its institutional_origin is a methodological artifact: the quantity Pr(sign correct | significant), the design-stage retrospective calculation, the two-sided p-value gate, and the joint pairing with Type M are constructs of the frequentist-hypothesis-testing framework (Gelman and Carlin's design analysis), not facts of nature. And vocab_travels fails: significance filter, power, conditional-on-significance, Pr(wrong sign), sign-reversal signature are pinned to the significance-testing substrate.

The portable structural skeleton is single and mathematical: selection on a noisy estimator produces a biased conditional distribution that, when the true value is near zero, retains non-trivial mass on the opposite tail, so the selected extremum can carry the wrong sign. That skeleton is exactly what Type S error instantiates from its parent prime winner_s_curse (its near-zero, wrong-sign face), with selection_bias as the broader parent and regression_to_the_mean as the cousin from which it differs precisely by flipping signs in the near-zero regime rather than merely shrinking: the cross-domain reach — any setting where noisy estimates are selected near zero — belongs to those umbrellas, whose instances are co-instances of selection-on-noisy-estimates, not applications of Type S. The named construct's distinctive content — Pr(sign correct | significant), the two-sided NHST filter, the design-stage joint Type M/Type S calculation — is precisely the home-bound cargo that does not lift. Its character: an evaluatively neutral, recognized-within-statistics operationalization of the winner's-curse sign-flip for the two-sided significance filter, structural in the select-the-near-zero-noisy-extremum skeleton it borrows from winner_s_curse but bound to its home domain by the NHST scaffolding that gives it its operational form, leaving it mixed rather than a free-floating prime.

Structural Core vs. Domain Accent

This section decides why Type S error is a domain-specific abstraction and not a prime — the sign-error twin of the Type M case, sharing its exact parent but keyed to the near-zero, wrong-direction regime.

What is skeletal (could lift toward a cross-domain prime). Strip the significance-testing scaffolding and a thin, mathematical relational structure survives: selection on a noisy estimator produces a biased conditional distribution, and when the true value sits near zero relative to the noise, that distribution retains non-trivial mass on the opposite tail, so the selected extremum can carry the wrong sign. Stated abstractly this is winner_s_curse — Type S is its near-zero, wrong-sign face — with selection_bias as the broader parent (here the selector is a two-sided NHST filter) and regression_to_the_mean as the cousin from which it differs precisely: regression shrinks extreme follow-ups but does not in general flip the sign of a true effect far from zero, whereas Type S flips signs specifically in the near-zero regime. This skeleton is genuinely substrate-portable: it governs any setting where noisy estimates are selected near zero — with or without a p-value. It is the core Type S shares, not what makes it distinctive.

What is domain-bound. Everything that makes the construct Type S error in particular is null-hypothesis-significance-testing machinery that does not survive extraction. The gate is specifically a two-sided p-value significance filter; the noisy estimator is an effect-size sampling distribution near zero relative to sampling noise; the diagnostic scalar is Pr(sign correct | significant), computed prospectively at design time from the effect-to-noise ratio and power; the reframe is stated against Type I/II decision error; the remedy is power, not precision (push opposite-tail mass toward negligible); the retrospective tell is a sign reversal between an original and its higher-powered replication; and the operational unit is the joint Type M/Type S design-stage calculation. The decisive test the entry supplies: an auctioneer whose winning bid is near his valuation, or any selector without a p-value, power, and a two-sided significance gate, is not committing a "Type S error" even when the same wrong-sign geometry is present. Remove the NHST scaffolding and what remains is bare selection-on-noisy-estimates-near-zero, a looser thing that no longer wears this name.

Why this does not clear the prime bar. A prime's vocabulary travels and its cross-domain transfer is recognition of the same mechanism, not analogy. Type S's transfer is bimodal. Within statistics and the significance-testing sciences the mechanism travels in identical form by genuine recognition — the condition-the-sign-on-the-filter move, the two-tail reachability geometry, the prospective Pr(wrong sign), the near-zero-versus-far fork, the power-not-precision remedy, and the sign-reversal-on-replication signature are the same objects across the behavioural-science replication crisis, health-policy A/B tests, GWAS allele-direction flips, and election-poll subgroup swings, because the substrate (a noisy estimator near zero, filtered by a two-sided significance threshold) is held fixed; and it is always evaluated jointly with its magnitude-sibling Type M. Beyond the significance-testing scaffolding the named construct does not travel; what recurs is the winner's curse near zero, not Type S. So when the bare structural lesson is needed elsewhere — select the extremum of noisy estimates near zero and you can get not just an exaggerated magnitude but a reversed direction — it is already carried, in general substrate-neutral form, by winner_s_curse (with selection_bias and regression_to_the_mean in the family), whose instances are co-instances of selection-on-noisy-estimates rather than exports of Type S. The cross-domain reach belongs to those parents; Type S's distinctive content — Pr(sign correct | significant), the two-sided NHST filter, the design-stage joint Type M/Type S calculation — is exactly the home-bound cargo that should stay in statistics and experimental design. Type S error clears the domain-specific bar comfortably for the significance-testing sciences, but its only substrate-spanning content is the winner's-curse sign-flip the parent prime already carries; its own contribution is to operationalize that flip for the two-sided significance filter.

Relationships to Other Abstractions

Local relationship map for Type S ErrorParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Type S ErrorDOMAINPrime abstraction: Effect Size — is part ofEffect SizePRIMEPrime abstraction: Statistical Power — is part ofStatisticalPowerPRIMEPrime abstraction: Statistical Significance (p-Value) — is part ofStatistical Sig…PRIMEPrime abstraction: Selection on Noisy Estimates — is a kind ofSelection onNoisy EstimatesPRIME

Current abstraction Type S Error Domain-specific

Parents (4) — more general patterns this builds on

  • Type S Error is a kind of Selection on Noisy Estimates Prime

    Type S is the near-zero two-sided-threshold species in which the selected estimate can enter from the tail opposite the true effect and reverse its sign.

  • Type S Error is part of Effect Size Prime

    Type S Error contains an estimated and true Effect Size whose signed directions are compared after significance selection.

  • Type S Error is part of Statistical Power Prime

    Type S Error contains Statistical Power because the effect-to-noise detection regime determines whether the opposite tail remains reachable.

  • Type S Error is part of Statistical Significance (p-Value) Prime

    Type S Error contains a two-sided statistical-significance gate whose opposite-tail admissions can carry the wrong sign.

Hierarchy paths (23) — routes to 8 parentless roots

Not to Be Confused With

  • Type M (magnitude) error. The magnitude-sibling, computed jointly from the same design — Type S says the direction is reversed, Type M says the direction is right but the size is exaggerated (the exaggeration ratio). They share the conditional-on-significance machinery, but a wrong sign and an inflated magnitude are distinct verdicts on one significant result; Type M is the common failure, Type S the rarer, sharper alarm confined to the near-zero regime. Tell: is the estimate pointing the wrong way (Type S), or the right way but too large (Type M)?
  • Type I / Type II error. The classical pair framed around the null-versus-alternative decision — Type I falsely declares an effect, Type II misses a real one. Type S presupposes the same scaffolding but asks a direction question the reject/fail-to-reject binary never poses: given a significant result, how often does it point the wrong way? A Type S error occurs on a real (near-zero) effect whose sign is reversed, not on a true null. Tell: is the question whether to reject the null at all (Type I/II), or whether a significant estimate's direction is trustworthy (Type S)?
  • Regression to the mean. The cousin that shrinks extreme follow-up measurements toward the average but does not, in general, flip the sign of a true effect far from zero. Type S flips signs precisely in the near-zero regime, where both tails of the filtered distribution are reachable — a different trigger and a different outcome (reversal, not mere shrinkage). Tell: does the follow-up merely moderate toward the mean without changing direction (regression to the mean), or can it reverse sign because the true effect sits near zero relative to noise (Type S)?
  • Publication bias. The field-level selection favoring significant/positive results across the literature. Type S is a within-design property — the probability a single underpowered study's significant estimate is directionally wrong — though the two compound in a literature. Publication bias is about which studies appear; Type S is about whether a given significant study's sign can be believed. Tell: is the concern that the literature over-represents significant findings (publication bias), or that an individual near-zero significant estimate points the wrong way (Type S)?
  • winner_s_curse / selection_bias (parent primes). The substrate-neutral skeleton Type S instantiates — selection on a noisy estimator produces a biased conditional distribution that, when the truth is near zero, retains opposite-tail mass, so the selected extremum can carry the wrong sign. This is what travels (any near-zero selection-on-noise setting), while Pr(wrong sign) and the two-sided NHST gate stay home. It is the umbrella, not a peer confusable. Tell: is the lesson the generic select-near-zero-noisy-estimates-can-reverse-sign pattern (the parents), or the specific two-sided-significance-filtered sign-error probability computed at design time (the named construct)? (Treated fully in a later section.)

Neighborhood in Abstraction Space

Type S Error sits in a sparse region of the domain-specific corpus (79th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Statistical Bias & Sampling Artifacts (6 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12