Skip to content

Small-Study Effects

The meta-analytic pattern in which smaller studies report systematically larger effects than larger ones, producing funnel-plot asymmetry that inflates the pooled estimate — a shared symptom of several biases, not a diagnosis of any one cause.

Core Idea

Small-study effects is the meta-analytic pattern in which smaller studies on a given question systematically report larger effect sizes than larger studies on the same question, producing a characteristic asymmetry in the funnel plot — a scatterplot of effect-size estimates against their precision (inverse standard error) — that under unbiased sampling should be symmetric around the true effect. The asymmetry inflates the meta-analytic point estimate toward the small-study end. Multiple upstream mechanisms can produce the observable pattern: publication bias (small null studies are less likely to be published or indexed, leaving only the positive small studies visible), outcome-reporting bias (selectively reporting favourable outcomes after data collection), quality differences correlated with sample size, and genuine clinical heterogeneity that places small trials in populations where effects are larger. These causes are not mutually exclusive and cannot be separated by the plot alone — the funnel asymmetry is the shared symptom, not a diagnosis of cause. The detection and partial correction of the pattern relies on a specific toolkit: visual funnel-plot inspection, Egger's regression test for funnel asymmetry, the trim-and-fill procedure that imputes missing small null studies, and sensitivity analyses restricting the meta-analysis to large studies only.

Structural Signature

Sig role-phrases:

  • the pooled study population — a body of K effect-size estimates on one question, each a study with its own sample size and standard error
  • the precision axis — study precision (inverse standard error), against which effect sizes are arrayed; the spread that makes the funnel a funnel
  • the size-dependent asymmetry — the observable: smaller studies report systematically larger effects, breaking the symmetry that unbiased sampling would produce
  • the inflated aggregate — the meta-analytic point estimate dragged toward the small-study end, overstating the truth
  • the multiplicity of upstream causes — publication bias, outcome-reporting bias, quality-by-size differences, genuine clinical heterogeneity, any of which can generate the same asymmetry
  • the detection-and-correction toolkit — funnel inspection, Egger's regression test, trim-and-fill imputation of missing small null studies, large-studies-only sensitivity analysis
  • the bounded-verdict discipline (what it deliberately discards) — the asymmetry is a shared symptom, not a diagnosis of cause, and the methods only partially correct; the honest output is suspicion plus a robustness-bounded re-estimate, never a debunking or a confident causal attribution

What It Is Not

  • Not a synonym for publication bias. Selective publication of small null studies is one upstream cause, but outcome-reporting bias, quality differences that track sample size, and genuine clinical heterogeneity can each produce the same funnel asymmetry. Small-study effects names the observable pattern; publication bias is one mechanism that may underlie it, and the plot alone cannot tell which is at work.
  • Not a diagnosis of cause. Funnel asymmetry is a shared symptom, not a fingerprint. Reading the plot as evidence that selective publication specifically inflated the estimate over-reads it; the honest conclusion is that something size-dependent is thinning the small null studies, licensing suspicion and sensitivity analysis but not a confident causal verdict.
  • Not proof the treatment does not work. When a robustness check collapses the pooled estimate to non-significance, the warranted claim is that the published-literature estimate of the effect is biased upward, not that the effect is absent. The pattern impeaches the size of the headline number, not the existence of the underlying effect.
  • Not a clean recovery of the true effect. Egger's test, trim-and-fill, and large-studies-only sensitivity analyses detect and partially correct the distortion; they impute or exclude under assumptions and cannot retrieve the genuinely unobserved studies. The output is a bounded re-estimate, not the unbiased truth.
  • Not evidence of fraud or incompetence. The asymmetry can arise from entirely good-faith processes — journals' preference for positive findings, or small trials run in higher-effect populations — with no fabrication and no methodological error. It is a property of how the literature was filtered, not an indictment of any individual study.

Scope of Application

Small-study effects lives within research synthesis — across the fields that pool effect-size estimates by meta-analysis; its reach is bounded by that practice, since each habitat needs a study-level unit with a measurable standard error to read precision off. (The size-and-selection analogues in fund tables or training-run scatters belong to the parent primes — selection on significance, survivorship bias, the winner's curse — not here.)

  • Meta-analysis of clinical trials — the home turf: funnel plots, Egger's regression test, and trim-and-fill all exist to detect and partially correct the size-dependent asymmetry in pooled trial evidence.
  • Education-research syntheses — small interventions routinely outshine their scaled-up replications, the small-study end inflating the pooled estimate of "what works."
  • Psychology and the decline effect — striking early findings attenuate as larger replications accumulate; small-study effects is one structural reason the published effect shrinks with sample size.
  • Pharmacoepidemiology and regulatory evidence-pooling — assessment of small-study bias is an expected step when pooling drug-effect evidence for regulatory submission, the asymmetry flagged before a pooled estimate is trusted.

Clarity

Naming small-study effects forces a separation that a naive meta-analysis silently collapses: the true underlying effect versus the distribution of reported effects as a function of study size. Pooling all studies with inverse-variance weights treats the literature as a transparent sample of the truth; the concept insists instead that what reaches the analyst is a filtered, size-dependent population, and that the filter predictably tilts the pooled estimate toward the small-study end. With the term in hand, an asymmetric funnel plot stops being a curiosity and becomes a structured warning — the meta-analyst now asks not "what is the pooled effect?" but "is my pooled effect inflated by whatever process is thinning the small null studies from view?"

The deeper clarity is that it cleanly distinguishes symptom from cause. Funnel asymmetry is one observable; publication bias, outcome-reporting bias, quality differences that track sample size, and genuine clinical heterogeneity are several distinct mechanisms that can each generate it. Holding "small-study effects" as the name for the pattern — and refusing to read the plot as a diagnosis of any one cause — keeps the analyst honest about what the evidence can and cannot support: the asymmetry licenses suspicion and sensitivity analysis (Egger's test, trim-and-fill, restricting to large trials), but not a confident attribution to selective publication alone. The sharper question it enables is therefore conditional and disciplined: given that something is making small and large studies disagree, how robust is the conclusion to the studies I may not be seeing?

Manages Complexity

A meta-analyst surveying a body of trials confronts a multiplicity of distinct worries about why the published record might misrepresent the truth: small null studies may never have been published or indexed; favourable outcomes may have been selectively reported after data collection; smaller and larger trials may differ in methodological quality; and the populations enrolled in small trials may genuinely differ from those in large ones. Each of these is a separate causal story with its own literature, and chasing them one at a time — auditing the file drawer, reconstructing reporting decisions, scoring study quality, modelling clinical heterogeneity — is an open-ended and largely unanswerable investigation, because the unpublished and unreported studies are by definition not in hand. Small-study effects compresses that thicket onto a single observable. All four upstream mechanisms, however different in kind, share one downstream symptom: smaller studies report systematically larger effects, which shows up as asymmetry in the funnel plot of effect size against precision — a scatter that should be symmetric around the true effect under unbiased sampling. The analyst stops trying to enumerate and adjudicate the causes and instead reads one geometric feature of one plot, with the asymmetry standing in for the entire family of distortions at once.

The compression is disciplined precisely by what it refuses to deliver, and that refusal is itself a simplification. Because the asymmetry is a shared symptom rather than a fingerprint of any one cause, the concept holds pattern and mechanism firmly apart: the funnel plot diagnoses that something is making small and large studies disagree, but it cannot attribute that something to publication bias rather than heterogeneity or quality, and the analyst who keeps this straight is spared the error of reading a confident causal verdict off a plot that cannot support one. What the single observable does support is a small, fixed toolkit applied in sequence rather than a bespoke investigation per study: visual funnel inspection, Egger's regression test to quantify the asymmetry, trim-and-fill to impute the missing small null studies and re-estimate, and a sensitivity analysis restricting the pool to large studies alone. The whole question of trustworthiness then reduces to tracking two things — the magnitude of the asymmetry, and how much the pooled estimate moves when the suspected small-study inflation is corrected or excluded — and reading off a clean conditional verdict: if the estimate is robust to dropping or imputing the small studies, the conclusion stands; if it collapses, the published-literature effect is inflated and the honest claim is suspicion plus a bounded re-estimate, not a debunking. Instead of treating the literature as a transparent window and re-deriving, study by study, every way it might be filtered, the analyst tracks one asymmetry and one robustness check, converting a high-dimensional census of unobservable biases into a low-dimensional, decidable structural diagnosis.

Abstract Reasoning

The core move is a symptom-to-suspicion inference that deliberately stops short of cause. The analyst reads one observable — smaller studies reporting systematically larger effects, visible as funnel-plot asymmetry against precision — and infers that something is thinning the small null studies from view and tilting the pooled estimate upward, without inferring which something. The characteristic inference runs from "small and large studies disagree in a size-dependent way" to "the published record is a filtered, size-dependent sample of the truth, not a transparent window onto it." The discipline of the move is its refusal to over-read: because publication bias, outcome-reporting bias, quality differences correlated with sample size, and genuine clinical heterogeneity can each produce the same asymmetry, the analyst treats the plot as a shared symptom and resists attributing it to selective publication alone. The move yields suspicion plus a mandate for sensitivity analysis, never a debunking.

The decisive interventionist-and-diagnostic move is the robustness check that converts suspicion into a bounded verdict: re-estimate the pooled effect with the suspected small-study inflation removed and read off how much it moves. The analyst reasons forward from a correction — trim-and-fill imputing the missing small null studies, or a sensitivity analysis restricting the pool to large trials alone — to a prediction about the conclusion's stability: if the estimate is robust to dropping or imputing the small studies, the finding stands; if it collapses to non-significance, the published-literature effect was inflated by the filter. The inference runs from the gap between the all-studies estimate and the large-studies-only (or imputed) estimate to how much of the headline effect was an artifact of size-dependent selection — and crucially the honest conclusion when it collapses is not "the treatment does not work" but "the published estimate of how well it works is biased upward."

A boundary-drawing move governs whether the diagnosis is even attemptable and how far it reaches. The analyst reasons from the range of sample sizes and the count of studies to whether asymmetry can be assessed at all — too few studies, or too narrow a precision range, and the funnel is uninterpretable — and from the nature of the available toolkit to the limit of what can be concluded: the methods detect and partially correct the pattern but cannot recover the truly unobserved studies, so the verdict is always a bounded re-estimate under assumptions, not a clean recovery of the unbiased effect.

Underwriting all of these is a reframing move that treats the literature itself as a measurement instrument with its own distortions. Rather than reasoning about study-level results as direct readings of the truth, the analyst reasons about the mapping from study-level results to the evidence-level estimate as a measurement question — asking how the size-dependent filter deforms that mapping — so the inference target shifts from "what is the effect?" to "how does the process that selects which studies I see bias my estimate of the effect, and how robust is my conclusion to the studies I may not be seeing?"

Knowledge Transfer

Within research synthesis the concept transfers as mechanism, and transfers cleanly across the fields that practise meta-analysis. Wherever a body of effect-size estimates is pooled across studies of varying sample size — meta-analysis of clinical trials (the home turf), education-research syntheses where small interventions outshine their scaled-up replications, psychology where striking early findings attenuate as larger replications accumulate (the "decline effect" as one of its structural causes), and pharmacoepidemiology and regulatory evidence-pooling where assessment of small-study bias is an expected step — the same object recurs and the same toolkit applies untranslated. The funnel plot of effect size against precision, Egger's regression test for asymmetry, the trim-and-fill imputation of missing small null studies, and the large-studies-only sensitivity analysis carry intact, because in every such field the unit is a study, the symptom is the size-dependent disagreement, and the discipline is the same: read the asymmetry as suspicion plus a bounded re-estimate, never as a debunking, and never as a confident verdict on which upstream cause is at work. The vocabulary — funnel asymmetry, precision, pooled estimate, publication bias, heterogeneity, robustness to dropping the small studies — travels without loss because the substrate is shared; what moves is not an analogy to evidence synthesis but evidence synthesis itself, applied to a different literature.

Beyond meta-analysis the picture is a shared abstract mechanism, not a travelling concept. What recurs across domains is the general pattern the small-study effect instantiates: a population of estimates filtered by selection on significance (or on success/survival) over-represents the extreme, low-precision tail and tilts the aggregate away from the truth. That pattern genuinely co-occurs elsewhere — performance-attribution analyses that examine only the funds that survived; scaling-law claims inferred from the handful of large training runs that were reported because they worked; any compilation that aggregates over outcomes pre-screened for being notable. But the cross-domain lesson there is carried by the parent mechanisms — selection on significance, survivorship bias, and the winner's curse (extreme estimates from small samples over-represented at the favourable tail) — each of which is substrate-independent and bears the weight on its own. The home-bound cargo is everything that makes small-study effects specifically itself: the funnel-plot geometry, the requirement of a study-level unit with a measurable standard error, the precision axis, and the named correction procedures, none of which has a referent where there are no "studies" and no "sample size" to read precision off. So invoking "small-study effects" for a fund table or a training-run scatter is analogy — it borrows the size-and-selection shape while dropping the funnel, the standard error, and the trim-and-fill machinery that give the original its operational bite. The honest move is to let the cross-domain reach ride on the parent primes the effect instantiates, and to reserve the name and its diagnostics for the meta-analytic substrate where they literally apply (see Structural Core vs. Domain Accent).

Examples

Canonical

Intravenous magnesium after acute myocardial infarction is the textbook case, used by Egger and colleagues (1997) to introduce their regression test. Through the early 1990s, a series of small randomized trials, and meta-analyses pooling them, suggested that magnesium sharply reduced post-heart-attack mortality — an apparently strong, cheap, life-saving intervention. The funnel plot of those trials, however, was markedly asymmetric: the smallest trials reported the largest benefits, while precision was low. When the mega-trial ISIS-4 then randomized over 58,000 patients, it found no mortality benefit from magnesium at all. The small-study end had inflated the pooled estimate; the large, precise trial landed near no effect, the outcome the asymmetry had warned about.

Mapped back: The magnesium trials are the pooled study population arrayed on the precision axis; the smallest trials showing the biggest benefit is the size-dependent asymmetry, dragging the meta-analysis into the inflated aggregate. ISIS-4's null is what the large-studies-only reading of the detection-and-correction toolkit anticipates. Honouring the bounded-verdict discipline, the asymmetry warranted suspicion of the pooled number, later confirmed by the mega-trial rather than by the plot alone.

Applied / In Practice

Turner and colleagues (2008) exposed how the published antidepressant literature overstated efficacy, using the FDA trial registry as a check on the published record. Of the registered trials, the published literature made nearly all appear positive, whereas by the FDA's own analyses only about half were positive; and the effect sizes in the published versions were inflated, on average by roughly a third, relative to the complete registered set. Negative and questionable trials were disproportionately unpublished or recast — a size-and-selection filter of exactly the kind small-study-effects diagnostics are built to flag. Meta-analysts now routinely screen antidepressant and other drug syntheses with funnel plots and asymmetry tests, treating a skewed funnel as a prompt to seek registry data rather than to trust the published pool.

Mapped back: Published versus registered trials form the pooled study population whose filtering produces the inflated aggregate. Selective (non-)publication is one member of the multiplicity of upstream causes; the registry comparison and funnel screening are the detection-and-correction toolkit. The verdict honoured the bounded-verdict discipline: the drugs were not declared ineffective, only their published effect sizes shown biased upward — a bounded re-estimate, not a debunking.

Structural Tensions

T1: Symptom-without-cause discipline versus the cause-specific remedy the analyst must choose. Refusing to read a confident cause off funnel asymmetry is the concept's central honesty — the plot cannot distinguish publication bias from reporting bias from quality-by-size from real heterogeneity. But the remedies diverge entirely by cause: if the filter is selective publication, the fix is to hunt registry data and re-pool; if it is genuine clinical heterogeneity, there is nothing to correct because the small-trial effect is really different in its population. The disciplined "we cannot say which" leaves the analyst unable to choose between seeking the missing studies and accepting the heterogeneity as substantive. The tension is that the intellectual honesty which forbids naming a cause is in direct conflict with the practical necessity of acting on one — the same plot that licenses only suspicion is consulted to justify a specific correction. Diagnostic: Does the chosen correction presuppose a cause (missing null studies) that the funnel asymmetry alone cannot actually establish?

T2: The correction as remedy versus the correction as its own bias. Trim-and-fill and large-studies-only sensitivity analyses convert suspicion into a bounded re-estimate — but each is itself a modelling procedure with assumptions that can mislead. Trim-and-fill imputes the missing studies as a symmetric mirror of the observed ones, which is exactly wrong if the true mechanism is not symmetric selection; large-studies-only discards real information and can trade small-study bias for reduced precision and a different sampling frame. The instrument built to detect a bias applies a further bias to correct it. The tension is that there is no assumption-free correction: the meta-analyst either trusts the raw pool (known to be inflated) or applies a fix whose own assumptions may be violated in the very cases where the distortion is worst. Diagnostic: Are the correction's own assumptions (symmetric missingness, exchangeability of large trials) plausible here, or is the fix importing a distortion of its own?

T3: A flag for bias versus a signal of real effect-modification. One of the legitimate upstream causes is that small trials are genuinely run in populations where the effect is larger — in which case the size-dependent disagreement is not an artefact to be corrected but a true feature of the world. The same asymmetry that warns of a filtered literature can be the fingerprint of real heterogeneity, and "correcting" it by down-weighting or imputing away the small studies would then erase a genuine effect-modification. The toolkit cannot tell the two apart from the plot. The tension is that the pattern the concept treats as a warning of distortion is, in a real fraction of cases, an accurate report of a difference that ought to be modelled rather than removed — so the correction risks flattening a truth. Diagnostic: Is the small-study end inflated by a filter, or reporting a genuinely larger effect in the populations small trials enroll?

T4: The diagnostic's demands versus the literatures that most need it. Funnel-plot asymmetry can only be assessed when there are enough studies spanning a wide enough precision range; with a handful of trials or a narrow spread, the funnel is uninterpretable and Egger's test underpowered. But a small literature is exactly where one or two inflated small studies do the most damage to the pooled estimate — the risk is highest precisely where the tool is weakest. The concept promises to convert an unbounded worry into one readable geometric feature, yet that feature becomes readable only once the literature is large, which is when any single small study matters least. The tension is that the diagnostic's reliability and the danger it guards against move in opposite directions with the number of studies. Diagnostic: Are there enough studies across a wide enough precision range to interpret this funnel at all, or is the asymmetry being read off a literature too thin to support it?

T5: The bounded re-estimate versus its downstream hardening into "the corrected truth." The concept insists its output is suspicion plus a robustness-bounded re-estimate under assumptions, never recovery of the unbiased effect — the unpublished studies are unobservable by definition. But a trim-and-fill-adjusted number, once computed, tends to circulate downstream as the corrected estimate, its inherent incompleteness forgotten, and guidelines or decisions treat it as the truth the raw pool obscured. The very act of producing a specific adjusted figure invites it to be read as more than the assumption-laden partial correction it is. The tension is that honesty about non-recoverability lives in the method's caveats while the method's product is a concrete number that behaves, in use, exactly like a recovered truth. Diagnostic: Is the adjusted estimate being carried forward as a bounded, assumption-dependent re-estimate, or as if it had recovered the effect the filtered literature hid?

T6: Autonomy versus reduction (a meta-analytic pattern or the selection/survivorship/winner's-curse parents). "Small-study effects" is a fully specified meta-analytic construct with home-bound cargo — the funnel-plot geometry, the study-level unit with a measurable standard error, the precision axis, Egger's test, trim-and-fill — and within research synthesis it transfers intact as mechanism across clinical, education, psychology, and pharmacoepidemiology syntheses. But the portable pattern it instantiates, a population of estimates filtered by selection on significance over-represents the low-precision extreme tail and tilts the aggregate, is carried by the parents selection on significance, survivorship_bias, and the winner's curse — each substrate-independent and load-bearing on its own. Invoking "small-study effects" for a surviving-funds table or a training-run scatter is analogy that borrows the size-and-selection shape while dropping the funnel and the standard error. The tension is that the cross-domain reach belongs to the parents while the funnel-and-trim-and-fill machinery stays in meta-analysis. Diagnostic: Resolve toward selection/survivorship/winner's-curse when carrying the lesson to any selection-filtered estimate population; toward small-study effects when a body of study-level effect sizes with standard errors is being pooled in situ.

Structural–Framed Character

Small-study effects sits at the mixed position on the structural–framed spectrum — an evaluatively neutral statistical pattern whose named form is bound to research-synthesis practice, wrapped around a substrate-independent selection mechanism. The criteria split. Evaluative_weight is essentially nil: funnel asymmetry is a geometric feature, not a verdict — the concept is scrupulous that the pattern is "a shared symptom, not a diagnosis of cause," and it impeaches the size of a number, not the honesty of anyone, so it convicts nothing. That neutrality is a strong structural mark. But three criteria pull framed at the named level. Human_practice_bound is real: the named concept requires a study-level unit with a measurable standard error and a pooled literature — it lives in the practice of meta-analysis and in the way a literature was filtered by publication and reporting, which are human research-and-publishing practices; remove the practice of pooling studies and there is no funnel to read. Institutional_origin is likewise pronounced for the apparatus: the funnel plot, Egger's regression test, and trim-and-fill are named methods of a statistical tradition, not structures found in nature. Vocab_travels fails outside research synthesis: funnel asymmetry, precision axis, pooled estimate, trim-and-fill lose their referents where there are no "studies." Yet import_vs_recognize leans structural for the parent: the underlying mechanism recurs as genuine co-instances — surviving-funds tables, reported-training-run scatters, any compilation pre-screened for notability — so what recurs is recognized, not analogized; only the named concept lifted whole (calling a fund table "small-study effects") is analogy.

The portable structural skeleton is a population of estimates filtered by selection on significance (or on success/survival) over-represents the extreme, low-precision tail and tilts the aggregate away from the truth. That skeleton is substrate-independent and is exactly what small-study effects instantiates from its umbrella primes — selection on significance, survivorship_bias, and the winner's curse (extreme estimates from small samples over-represented at the favourable tail). The cross-domain reach belongs to those parents, each load-bearing on its own, while the funnel-plot geometry, the standard-error precision axis, and the named correction procedures that make "small-study effects" the specific meta-analytic construct stay pinned to research synthesis. Its character: an evaluatively neutral statistical pattern whose underlying selection mechanism is genuinely structural and travels under selection on significance / survivorship_bias / winner's curse, but whose named form — funnel, standard error, trim-and-fill — is bound to the human practice and institutional methods of meta-analysis, leaving it mixed rather than a free-floating prime.

Structural Core vs. Domain Accent

This section decides why small-study effects is a domain-specific abstraction and not a prime, and it carries the case for its domain-specificity — there is no separate section for that.

What is skeletal (could lift toward a cross-domain prime). Strip the meta-analytic apparatus and a substrate-independent structure survives: a population of estimates filtered by selection on significance (or on success/survival) over-represents the extreme, low-precision tail, so the aggregate is tilted away from the truth toward the favourable end. The pieces that travel are abstract — a pool of noisy estimates, a filter keyed to how favourable (or how significant, or how surviving) each estimate is, and a resulting aggregate biased because the low-information extremes are disproportionately the ones that got through. That skeleton is genuinely portable and recurs as co-instances, not metaphors: a performance table showing only the funds that survived, a scaling-law claim read off the handful of large training runs that were reported because they worked, any compilation aggregating over outcomes pre-screened for being notable. It is the core small-study effects shares, not what makes it distinctive.

What is domain-bound. Almost all the operative content is research-synthesis furniture and none of it survives extraction intact: the funnel-plot geometry (effect size against precision, symmetric only under unbiased sampling); the requirement of a study-level unit with a measurable standard error off which precision is read; the precision axis itself; the named detection-and-correction toolkit — visual funnel inspection, Egger's regression test for asymmetry, trim-and-fill imputation of missing small null studies, and the large-studies-only sensitivity analysis; and the symptom-not-cause discipline that holds the funnel asymmetry apart from its several possible upstream mechanisms (publication bias, outcome-reporting bias, quality-by-size, genuine clinical heterogeneity). These are the worked vocabulary, the instruments, and the empirical cases — magnesium-after-MI, the antidepressant registry comparison — that the discipline actually studies. The decisive test: remove the study-level unit with a standard error and there is no funnel to plot, no precision axis to array against, and nothing for trim-and-fill to impute — "small-study effects" simply has no referent where there are no "studies" and no "sample size."

Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose cross-domain transfer is recognition of the same mechanism, not analogy. Small-study effects' transfer is bimodal. Within research synthesis it travels intact — across clinical-trial meta-analysis, education-research syntheses, the psychology decline effect, and pharmacoepidemiological evidence-pooling the funnel plot, Egger's test, trim-and-fill, and the large-studies-only check apply untranslated, because the unit is always a study, the symptom always the size-dependent disagreement; what moves is not an analogy to evidence synthesis but evidence synthesis itself applied to a different literature. Beyond it, invoking "small-study effects" for a surviving-funds table or a reported-training-run scatter is analogy: it borrows the size-and-selection shape while dropping the funnel, the standard error, and the trim-and-fill machinery that give the original its operational bite. And when the bare structural lesson is needed cross-domain, it is already carried, in more general and load-bearing form, by the parents the effect instantiates — selection on significance, survivorship_bias, and the winner's curse (extreme estimates from small samples over-represented at the favourable tail), each substrate-independent on its own. The cross-domain reach belongs to those parents; the funnel-plot geometry, the standard-error precision axis, and the named correction procedures that make "small-study effects" the specific meta-analytic construct are domain baggage that should stay in research synthesis.

Relationships to Other Abstractions

Local relationship map for Small-Study EffectsParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Small-Study EffectsDOMAINPrime abstraction: Effect Size — is part ofEffect SizePRIMEPrime abstraction: Selection on Noisy Estimates — is part of, conditionalSelection onNoisy EstimatesPRIMEDomain-specific abstraction: Funnel Plot Asymmetry — is a decomposition ofFunnel PlotAsymmetryDOMAIN

Current abstraction Small-Study Effects Domain-specific

Parents (2) — more general patterns this builds on

  • Small-Study Effects is part of Effect Size Prime

    Small-Study Effects contains effect-size estimates as the magnitude coordinate whose systematic relationship with study precision defines the pattern.

  • Small-Study Effects is part of, conditional Selection on Noisy Estimates Prime

    In significance-selected literatures, Small-Study Effects contains selection on noisy estimates because low-precision studies become visible only at an extreme tail.

    Condition / exception Genuine heterogeneity and quality-by-size differences can create the same observable without estimate-dependent selection.

Children (1) — more specific cases that build on this

  • Funnel Plot Asymmetry Domain-specific is a decomposition of Small-Study Effects

    Removing the funnel-plot encoding and named diagnostic apparatus leaves the size-dependent effect-estimate pattern named by Small-Study Effects.

Not to Be Confused With

  • Publication bias. The selective non-publication of small null studies — one of the several upstream mechanisms that can produce small-study effects, not a synonym for the pattern. Small-study effects names the observable (size-dependent asymmetry); publication bias is one cause that may underlie it, alongside outcome-reporting bias, quality-by-size, and heterogeneity. Tell: are you naming what the funnel shows (small-study effects) or a specific reason the small nulls are missing (publication bias)? The plot cannot tell you it was publication bias specifically.
  • (Clinical) heterogeneity. Genuine variation in the true effect across studies — often because small trials enroll different populations where the effect is really larger. It is both another upstream cause of the asymmetry and a rival reading of it: where heterogeneity is real, the small-study end is reporting a true effect-modification, not an artifact to correct. Tell: is the size-dependent difference a filter on which studies are seen (bias to correct) or a real difference in the effect across the populations studied (heterogeneity to model)? The funnel alone cannot separate them.
  • Winner's curse. The parent phenomenon that extreme estimates from small/underpowered samples are over-represented at the favourable tail, so the "winners" overstate the truth. This is the substrate-general mechanism small-study effects instantiates in meta-analysis. Tell: is a study-level funnel of effect-size-vs-precision being read (small-study effects) or the general principle that low-power discoveries are inflated (winner's curse)? The latter travels to any noisy-estimate setting; the former needs studies and standard errors. (Treated fully in a later section.)
  • Survivorship bias. The parent bias of drawing conclusions only from the units that "survived" a selection filter (surviving funds, completed trials). Small-study effects is its meta-analytic cousin where the filter is size-and-significance on published studies. Tell: is the filtered population a set of pooled effect-size estimates with precisions (small-study effects) or any surviving-subset whose invisible failures bias the aggregate (survivorship bias)? (Treated fully in a later section.)
  • Decline effect. The observation that striking early findings shrink as larger replications accumulate. Small-study effects is one structural cause of the decline effect, not the same thing: the decline effect is the temporal shrinkage of the reported effect; small-study effects is the size-dependent asymmetry that partly produces it. Tell: is the claim about effects fading over successive studies (decline effect) or about small studies systematically overstating at any one time (small-study effects)?
  • Funnel plot / Egger's test / trim-and-fill (the tools). These are the instruments for detecting and partially correcting small-study effects, not the pattern itself. The funnel plot displays it; Egger's test quantifies its asymmetry; trim-and-fill imputes the missing nulls. Tell: are you naming the phenomenon (small-study effects) or a method used to detect or adjust for it (the toolkit)? Confusing a trim-and-fill-adjusted number with "the true effect" is exactly the over-reading the concept warns against.

Neighborhood in Abstraction Space

Small-Study Effects sits in a sparse region of the domain-specific corpus (70th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Publication Bias & Research Artifacts (5 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12