File Drawer Problem¶
Recognize that studies with null results disproportionately go unpublished while significant ones enter the literature, so any synthesis treating the published record as the full population of conducted research systematically overestimates effect sizes toward the filter.
Core Idea¶
The file drawer problem is the meta-research phenomenon, named by Robert Rosenthal in 1979, in which studies with null, negative, or unsurprising results disproportionately go unpublished — sitting in the researcher's file drawer — while studies with statistically significant or publication-worthy results enter the literature, making the published record a systematically biased sample of the conducted research. The mechanism has three load-bearing parts: a population of conducted studies with some distribution of true effect sizes and sampling variation around them; a publication filter that selects studies for the visible record based on a criterion — statistical significance, novelty, directional alignment with prevailing expectations — that is correlated with the effect size itself; and a downstream synthesis (meta-analysis, narrative review, guideline development, policy decisions) that treats the filtered visible record as if it were the full population. The consequence is that any downstream synthesis operating on the filtered record systematically overestimates effect sizes, understates sampling uncertainty, and overstates the reliability of findings — sometimes by large margins. The structural payoff Rosenthal demonstrated is that what is absent from the record drives the conclusions drawn from the record: a substantial mass of null results filed away and never published can make a weak, marginal, or non-existent effect look robust and replicable. The principal institutional responses — pre-registration of trials before data collection (eliminating the filter prospectively), registered reports (decoupling publication decisions from results), trial-result reporting mandates for pharmaceutical sponsors, and meta-analytic correction methods (funnel plots, trim-and-fill, p-curve, Egger regression, selection models) — are all designed to either eliminate the filter or estimate and correct for the hidden mass it has already produced.
Structural Signature¶
Sig role-phrases:
- the conducted-studies population — all studies actually run, with their distribution of true effects and sampling variation around them
- the publication filter — the selection criterion (statistical significance, novelty, directional alignment) that admits studies to the visible record and is correlated with the effect size itself
- the published sample — the filtered visible record, biased relative to the conducted population by an amount proportional to the filter's correlation with effect size
- the downstream synthesis — the meta-analysis, review, guideline, or policy decision that mistakes the filtered record for the full population
- the hidden mass — the filed-away null/negative results whose absence drives the conclusions, making a weak or non-existent effect look robust
- the signed pull — because the criterion correlates positively with effect, the bias on the synthesised estimate is always toward zero; the corrected effect is bounded below the naive one, direction known before magnitude
- the funnel fingerprint — the filter's detectable signature in effect-versus-precision asymmetry (small imprecise studies survive only when they land large), readable without auditing any study's internals
- the prevent-or-correct fork — a filter not yet operated is eliminated prospectively (pre-registration, registered reports, reporting mandates); one already operated can only be estimated and corrected after the fact (trim-and-fill, p-curve, Egger, selection models)
What It Is Not¶
- Not a flaw inside any individual study. The defect lives at the publication-decision stage, downstream of data collection, in which results were allowed to become visible — not in how any one study was conducted. Every trial in the record can be impeccably designed and the pooled conclusion still wrong, because the bias is a property of the record's composition, not of its members.
- Not p-hacking or HARKing. Those are within-study analytic flexibility — flexible analysis, hypothesis re-statement — that corrupt results before they are written up; the file drawer is selection among completed studies at the publication gate. The two combine to produce most of the replication-crisis distortion, but the file drawer concerns which finished studies become visible, not how any was analyzed.
- Not an individual researcher's misconduct. The "problem" is the system-level distortion produced by the aggregation of many ordinary choices through journals' and reviewers' filtering, not any one author's bad faith. Filing a null result is a rational, blameless response to publication incentives; the bias emerges at the level of the literature, not the person.
- Not publication bias in general. It is the specific unpublished-null-results case — the canonical sub-mechanism of the broader publication-bias umbrella, which also covers selective outcome reporting, language bias, and citation bias. The file drawer names one particular filter (significance/novelty admitting some studies and filing others), not the whole family of ways the record can be skewed.
- Not cherry-picking. Cherry-picking is a rhetorical, within-argument move — a writer selectively citing favourable evidence; the file drawer is a structural feature of the literature itself, arising from distributed publication decisions rather than from any advocate's selection. One is a choice in an argument, the other a property of the corpus.
- Not a bias of unknown direction. Because the admission criterion correlates positively with the effect, the pull on the synthesised estimate is always toward zero — effects are inflated, uncertainty understated. The corrected effect is bounded below the naive one, so the direction of the correction is known in advance even before its magnitude is estimated.
Scope of Application¶
The file drawer problem lives across the subfields of meta-research, evidence synthesis, and scientific reform — the substrate of science as a publication-mediated knowledge system; its reach is within that domain, and the deeper selection-filter mechanism it instantiates (a filter on visibility biases the visible sample) belongs to the parent primes selection_bias, survivorship_bias, and selection_effect, of which it is the publication-pipeline case. It is the connective tissue of a cluster of publication-pathology concepts, recurring across the habitats below.
- Meta-analysis methodology — file-drawer correction as a routine diagnostic step: funnel-plot asymmetry, trim-and-fill, p-curve, Egger regression, and selection models estimating the hidden mass and the corrected effect.
- Clinical-trial registries — registration before data collection (clinicaltrials.gov, EU CTR) instituted precisely to make the conducted population visible regardless of outcome, eliminating the filter prospectively.
- The replication crisis — the file drawer as one proximate mechanism behind over-stated, non-replicating effects in psychology, biomedicine, and economics (the inflation of the published over the conducted literature).
- Drug-trial regulation — mandates that sponsors report all trial results, targeting the selective publication of trials favourable to a sponsored drug.
- Pre-registration and registered reports — editorial models that decide publication before results exist, decoupling the publication decision from the outcome and structurally removing the filter for accepted work.
Clarity¶
Naming the file drawer problem moves the locus of suspicion outside the individual study. The familiar threats to a finding — confounding, measurement error, sampling bias within a study — all ask whether each published result is sound. The file drawer makes a different object inspectable: the record itself, the published sample, treated as a measurement that can be biased independently of the quality of any study in it. Every individual trial can be impeccably conducted and the pooled conclusion still wrong, because the defect lives in which results were allowed to become visible, not in how any one of them was produced. That is the distinction the label crystallises — selection at the publication-decision stage, downstream of data collection, as opposed to bias inside the data.
It also converts a vague unease — "can I trust this literature?" — into a sharp, answerable question: is the visible record a fair sample of the conducted research, or a filtered one? Once posed that way, the absent results become something to reason about rather than ignore: their hidden mass can be estimated (funnel-plot asymmetry, trim-and-fill, p-curve, selection models) or pre-empted (registration before data collection, registered reports that decide publication before results exist). The sharper question a practitioner can now ask of any meta-analysis or review is not just "what do the studies show?" but "what would the studies that were run and never published have shown, and how much would they pull this estimate toward zero?"
Manages Complexity¶
The sprawl this tames is the diffuse, hard-to-act-on worry that hangs over any body of published evidence: can I trust this literature? Stated that way the question has no handle — it seems to demand re-auditing every study in the corpus for confounding, measurement error, and analytic flexibility, an unbounded task, and even a corpus of individually impeccable studies leaves the worry unresolved because the suspicion is about the corpus as a whole, not any member of it. The file drawer problem compresses that worry by supplying a single three-part model of how the corpus came to exist: a population of conducted studies with some distribution of true effects and sampling noise; a publication filter that admits studies on a criterion correlated with the effect itself; and a downstream synthesis that mistakes the filtered visible record for the full population. Once the corpus is seen through this model, "is the literature trustworthy?" collapses to one structured question with a definite object — is the visible record a fair sample of the conducted research, or a filtered one? — and the unbounded study-by-study audit is replaced by reasoning about a single missing mass.
What the analyst then tracks is a small, fixed set of quantities rather than the full content of every study. The load-bearing one is the hidden mass of filed-away results and the direction and size of the pull it exerts on the synthesised estimate — always toward zero, by an amount set by how tightly the filter's admission criterion correlates with effect size. The bias does not require modelling each study's internals; it reads off the relationship between the visible studies' effect sizes and their precision, because the filter's fingerprint is that small, imprecise studies survive only when they happen to land large, while small null studies are filed. So the analyst inspects effect-versus-precision asymmetry (the funnel plot) and runs a handful of named estimators over it — trim-and-fill, p-curve, Egger regression, selection models — to put a number on the hidden mass and on the corrected effect. A high-dimensional "re-evaluate the whole literature" problem becomes the estimation of one latent quantity from a low-dimensional summary the visible record already exposes.
The branch structure the concept supplies routes both diagnosis and remedy off the filter. Diagnostically, the first cut is the locus of bias: is the defect inside the studies (confounding, measurement error, sampling bias within a study) or at the publication-decision stage, downstream of data collection, in which results were allowed to become visible? The file drawer isolates the second, making the published record itself an object that can be biased independently of the soundness of any study in it — so a pooled conclusion can be wrong even when every trial is impeccable, and the analyst knows to interrogate the record's composition rather than re-litigate the studies. The remedy branch then forks on timing relative to the filter: a filter not yet operated can be eliminated prospectively — pre-registration before data collection, registered reports that decide publication before results exist, mandatory trial-result reporting — whereas a filter that has already produced its hidden mass can only be estimated and corrected after the fact with the meta-analytic toolkit. The analyst reads the appropriate intervention off where the literature sits relative to that fork. And the same model identifies which literatures are most exposed without case-by-case investigation — many small studies plus an easy significance threshold plus weak registration norms maximise the filter's grip, while large pre-registered portfolios blunt it. So in place of an unbounded "audit everything" worry, the analyst holds a three-part generative model, one latent quantity (the hidden mass and its zero-ward pull), a low-dimensional diagnostic (effect-versus-precision asymmetry), and a two-way diagnosis/remedy branch structure — and reads off whether the record is filtered, by how much the pooled effect is inflated, where the defect lives, and whether to prevent the filter or correct for it. A fuzzy question about an entire literature becomes the estimation of a single missing mass with a definite branch structure.
Abstract Reasoning¶
The first characteristic move is diagnostic from a fingerprint the visible record itself exposes. The filter leaves a detectable signature — small, imprecise studies survive publication only when they happen to land large, while small null studies are filed — so the analyst reads effect size against precision (the funnel plot) and infers the hidden mass from the asymmetry, without auditing any study's internals. The move reasons FROM "the small published studies cluster on the positive side and the small-null region is sparse" TO "a publication filter has removed a mass of null results, and the pooled estimate is inflated toward the filter." The named estimators — trim-and-fill, p-curve, Egger regression, selection models — put a number on that inference, imputing the missing studies and reporting the corrected effect; when trim-and-fill restores twelve null studies and the effect falls from d = 0.45 to d = 0.18, the diagnosis is that the original consensus was a filter artefact, and a registry-based estimate near d = 0.15 confirms it.
The second move is interventionist, forked on timing relative to the filter. The model says the bias is produced by a single filter, so the analyst reasons about when that filter can be reached. FROM "the filter has not yet operated on this literature" TO "eliminate it prospectively — pre-register before data collection, adopt registered reports that decide publication before results exist, mandate trial-result reporting — so the conducted population becomes visible regardless of outcome." FROM "the filter has already produced its hidden mass" TO "it cannot be removed, only estimated and corrected after the fact with the meta-analytic toolkit." Each fork is a prediction: prospective elimination should compress published effect sizes toward the conducted-population truth as registration infrastructure matures (a prediction partially borne out in clinical-trial registries after mandates), while post-hoc correction recovers the corrected effect but cannot manufacture the studies that were never run.
The third move is boundary-drawing on the locus of the defect. The first cut separates bias inside the studies — confounding, measurement error, within-study sampling bias, all properties of how each result was produced — from bias at the publication-decision stage, downstream of data collection, a property of which results were allowed to become visible. The file drawer isolates the second, making the published record itself an object that can be biased independently of the soundness of any study in it — so the analyst draws the boundary that lets a pooled conclusion be wrong even when every trial is impeccable, and knows to interrogate the record's composition rather than re-litigate the studies. The concept also draws a vulnerability boundary that predicts exposure without case-by-case investigation: many small studies, plus an easy significance threshold, plus weak registration norms and strong directional editorial preference, maximise the filter's grip, while large pre-registered portfolios and multi-site consortia blunt it — so the analyst can rank literatures by expected inflation before opening any of them. And the move carries a signed prediction: because the admission criterion correlates positively with the effect, the pull on the synthesised estimate is always toward zero, by an amount set by how tightly the filter tracks effect size — so the corrected effect is bounded below the naive one, and the direction of the correction is known in advance even before its magnitude is estimated.
Knowledge Transfer¶
Within meta-research, evidence synthesis, and scientific reform the file drawer problem transfers as mechanism, and it is the connective tissue of a whole cluster of publication-pathology concepts. The same three-part model — a conducted-studies population, a publication filter correlated with effect size, and a downstream synthesis that mistakes the filtered record for the full one — and the same diagnostic-and-remedy apparatus carry intact across meta-analysis methodology (funnel-plot asymmetry, trim-and-fill, p-curve, Egger regression, selection models as the routine correction toolkit), clinical-trial registries (registration before data collection instituted precisely to make the conducted population visible regardless of outcome), the replication crisis (the file drawer as one proximate mechanism behind over-stated, non-replicating effects in psychology, biomedicine, and economics), drug-trial regulation (mandates that sponsors report all results, targeting selective publication of favourable trials), and pre-registration / registered reports (editorial models that decouple the publication decision from the result and so eliminate the filter prospectively for accepted work). It sits in a family of sibling pathologies — publication bias (the umbrella, of which file drawer is the canonical unpublished-null case), p-hacking and HARKing (within-study analytic flexibility, which combine with the file drawer to produce most of the replication-crisis distortion), the garden of forking paths, citation bias — and the diagnostics, the signed prediction (the pull is always toward zero), and the prevent-versus-correct fork all transfer across this cluster without modification, because the substrate is constant: science as a publication-mediated knowledge system.
Beyond that substrate the report is clean (B). The deeper structural pattern the file drawer instantiates — a selection filter on what becomes visible biases the visible record relative to the underlying population — genuinely recurs across domains and is already carried at the prime level by selection_bias, survivorship_bias, and selection_effect; indeed the file drawer is precisely survivorship bias applied to the population of conducted studies under publication selection, the exact same mechanism that produces Wald's returning-aircraft fallacy and countless other named cases in other fields. So the cross-domain lesson — that an unobserved filter on visibility silently distorts any inference drawn from the survivors — belongs to those parents, and the file drawer is one instantiation among many, alongside aviator-survivorship and the rest. What stays home-bound is everything that makes it the file drawer specifically: the scientific-publication pipeline (journals, reviewers, editorial novelty/significance criteria), the conducted-research-population framing, and the named meta-analytic and reform machinery (funnel plots, trim-and-fill, p-curve, fail-safe N, registries, registered reports), none of which has a referent outside a publication-mediated knowledge system. Strip the vocabulary — publication, journal, study, meta-analysis — and what remains is "a filter on visibility biases the visible sample," which simply is selection bias / survivorship bias at the prime level. The boundary to mark is therefore that the cross-substrate reach is genuine shared mechanism (carry it via the selection-filter parent), while "the file drawer problem" is the meta-research instantiation that adds the publication-pipeline substrate and its correction apparatus on top. See Structural Core vs. Domain Accent.
Examples¶
Canonical¶
Turner and colleagues' 2008 study in the New England Journal of Medicine is the definitive demonstration. They obtained the FDA registration records for 74 trials of twelve antidepressants — trials registered before their results were known — and compared them with what reached the published literature. By the FDA's assessment, 38 of the 74 were positive and 36 were negative or questionable. In the journals the picture was transformed: 37 of the 38 positive trials were published, whereas of the 36 non-positive trials, 22 were never published and 11 more appeared framed as positive, leaving only 3 published as negative. A reader of the literature would see roughly 94% of trials as positive; the full registered record showed about 51%. Pooled effect sizes were inflated by roughly a third.
Mapped back: The 74 registered trials are the conducted-studies population; the "positive gets published" selection is the publication filter, correlated with the result itself. The journal record showing 94% positive is the published sample, and the 22 vanished null trials are the hidden mass. The one-third inflation of pooled effects is the signed pull — always toward a larger effect — that any downstream synthesis reading only the journals would inherit.
Applied / In Practice¶
The institutional remedy is prospective elimination of the filter through trial registration. Following reforms including the 2005 ICMJE requirement that journals publish only trials registered at inception, and the U.S. FDA Amendments Act of 2007 mandating registration and results reporting on ClinicalTrials.gov, the conducted population of trials became visible regardless of outcome. Registering a trial and its primary outcome before data collection means a null result can no longer quietly vanish into a file drawer: the trial's existence is on the public record, and a later reviewer can see which registered trials never reported and chase them down. Registered Reports — an editorial format in which journals accept a study on its design before results exist — extend the same logic to the wider behavioral sciences.
Mapped back: This is the prevent side of the prevent-or-correct fork: because the filter has not yet operated on a newly registered trial, it is eliminated prospectively rather than corrected after the fact. Registration makes the conducted-studies population visible regardless of outcome, so no hidden mass of filed nulls can form, and a downstream synthesis can no longer mistake a filtered record for the full one.
Structural Tensions¶
T1: Record composition versus study quality (two loci of scrutiny that compete for attention). The concept's signature move is to relocate suspicion from inside the study to the composition of the record: a pooled conclusion can be wrong even when every trial in it is impeccably conducted, because the defect lives at the publication-decision stage, not in the data. This is genuinely liberating — it makes the literature itself an inspectable object — but it also cuts the other way. An analyst absorbed in modelling the hidden mass can neglect that the visible studies may themselves be p-hacked or confounded, and the two pathologies compound: within-study analytic flexibility inflates the individual effects that a filtered record then over-represents. The tension is that record-composition bias and within-study bias are distinct defects requiring distinct fixes, yet they interact multiplicatively, so attending to either alone under-diagnoses the distortion. Correcting the funnel while trusting flawed survivors, or scrutinizing survivors while ignoring the filter, each leaves half the problem standing. Diagnostic: Is the pooled estimate being wrong attributable to which studies became visible, to how the visible ones were produced, or to both compounding — and does the chosen remedy address the operative locus?
T2: Reading the filter from the survivors versus inferring the unobservable (asymmetry has other causes). The elegant diagnostic is that the filter leaves a fingerprint the visible record already exposes — small imprecise studies survive only when they land large — so effect-versus-precision asymmetry (the funnel plot) licenses an inference to the hidden mass without auditing any study. But the correction methods impute studies that were never observed, and the asymmetry they read is not uniquely produced by publication selection: genuine small-study effects, between-study heterogeneity, and differences in methodological quality across study sizes can mimic the same funnel signature. Trim-and-fill and Egger regression estimate a latent quantity from the survivors under model assumptions that can fail, and no method can recover studies that were never run rather than merely filed. The tension is that the fingerprint's readability is the concept's power and its trap: the same asymmetry that reveals a filter can be manufactured by mechanisms that have nothing to do with publication, so a confident correction can misattribute ordinary heterogeneity to a file drawer. Diagnostic: Is the funnel asymmetry here the fingerprint of a publication filter, or could genuine small-study effects or heterogeneity produce the same signature without any selective filing?
T3: Prevent versus correct (a fork where each side buys what the other cannot). The remedy branch forks on timing relative to the filter: a filter not yet operated can be eliminated prospectively (pre-registration, registered reports, reporting mandates), while one already operated can only be estimated and corrected after the fact. The tension is that neither side dominates. Prevention makes the conducted population visible regardless of outcome, but it works only going forward, imposes upfront rigidity (locking in designs and analyses before data), and does nothing for the vast existing literature. Post-hoc correction salvages the record we already have, but it can only estimate the missing mass, never manufacture the studies that were never run, and it inherits all the model-dependence of T2. The prospective fix is clean but slow and prospective-only; the retrospective fix is available now but fundamentally lossy. An analyst facing a literature must decide which regime it sits in and accept that the tool matched to that regime forgoes what the other offers. Diagnostic: Has the filter already operated on this literature (only estimate-and-correct is available) or not yet (prevention is possible) — and is the chosen tool's characteristic limitation acceptable for the decision at stake?
T4: The signed pull as a strength versus its dependence on a filter that selects on magnitude. A prized feature is that the bias has a known direction: because the admission criterion correlates positively with effect size, the pull on the synthesised estimate is always toward zero, so the corrected effect is bounded below the naive one and the direction of correction is known before its magnitude is estimated. This is real analytic leverage. But the guarantee rests entirely on the filter selecting for large/significant effects, and that assumption is not universal: literatures where a surprising null is itself publishable, where directional editorial preference favors a particular sign rather than mere significance, or where confirmation of an established effect is the novelty criterion can invert or scramble the pull. The tension is that the concept's cleanest prediction — always toward zero — is a property of one common filter shape, not of the file drawer as such, and treating the signed pull as automatic risks miscorrecting a literature whose filter does not track magnitude in the assumed way. Diagnostic: Does this literature's publication filter actually select on effect magnitude/significance (pull toward zero holds), or on direction, surprise, or confirmation (the sign of the bias is no longer guaranteed)?
T5: System-level distortion versus blameless individual choices (a problem with no culprit). The file drawer is a property of the literature, not of any person: filing a null result is a rational, blameless response to publication incentives, and the distortion emerges only at the level of the corpus through the aggregation of many ordinary editorial and authorial decisions. This diagnosis is precise and important — it locates the fault in incentives and structure rather than misconduct — but it creates a genuine tension for remedy and accountability. Because no individual is at fault, there is no one to hold responsible, and the fix must reshape the system (registries, editorial formats, mandates) against the grain of every participant's locally rational behavior. Moralizing individual researchers misdiagnoses the mechanism; but purely structural framing can also let genuinely selective, results-driven filing off the hook by folding it into "blameless aggregation." The tension is that the same aggregation story that correctly exonerates the individual can obscure where a choice really was strategic rather than blameless. Diagnostic: Is the missing result absent through the ordinary, blameless operation of publication incentives, or through a choice selective enough that the aggregation story is masking it?
T6: Autonomy versus reduction (a named meta-research problem or the publication-pipeline case of selection bias). The file drawer problem is a named meta-research phenomenon with proprietary cargo — the conducted-studies-population framing, the scientific-publication pipeline, and the named correction and reform machinery (funnel plots, trim-and-fill, p-curve, fail-safe N, registries, registered reports) — and within meta-research it transfers as mechanism across meta-analysis, the replication crisis, drug-trial regulation, and pre-registration. But its deeper structure is not proprietary: it is survivorship_bias applied to the population of conducted studies under a publication filter, the same mechanism as Wald's returning-aircraft fallacy, and it is carried at the prime level by selection_bias and selection_effect. Strip the vocabulary — publication, journal, study, meta-analysis — and what remains is "a filter on visibility biases the visible sample," which simply is selection bias. The tension is between a named problem that earns its own detailed study and reform apparatus inside a publication-mediated knowledge system, and the recognition that its cross-domain lesson — an unobserved filter on visibility silently distorts inference from the survivors — already belongs to its parents. Diagnostic: Resolve toward the parents (selection bias, survivorship bias, selection effect) when carrying the lesson to any visibility-filtered inference; toward the named problem when diagnosing a scientific literature with its publication pipeline and correction toolkit in play.
Structural–Framed Character¶
The file drawer problem sits at mixed on the structural–framed spectrum — a named phenomenon bound to a human publication institution, but one whose underlying structure is a textbook substrate-neutral prime (selection/survivorship bias), giving it strong structural pull. The criteria split. On evaluative weight it is low and, by the entry's own insistence, deliberately de-moralized: "the problem" is a system-level distortion with no culprit, filing a null result is a "rational, blameless response to publication incentives," and the defect is explicitly "not an individual researcher's misconduct" — so the content is a neutral diagnosis of a compositional bias in a corpus, not a verdict on any agent. On human-practice-bound it leans framed: the construct presupposes "science as a publication-mediated knowledge system" — journals, reviewers, editorial significance/novelty criteria — and "the file drawer problem" as such dissolves without that publishing institution. But the mechanism it instantiates is fully substrate-neutral and runs wherever an unobserved filter selects what becomes visible, including non-institutional cases like Wald's returning aircraft, which pulls the underlying pattern firmly back toward the structural side. On institutional origin it is framed: the conducted-studies-population framing and the whole correction-and-reform apparatus (funnel plots, trim-and-fill, p-curve, fail-safe N, registries, registered reports) are artifacts of the meta-research discipline and the publication pipeline, not facts of nature.
On vocab-travels it is framed: publication, journal, study, meta-analysis, registration have no referent outside a publication-mediated knowledge system, and stripping them leaves only "a filter on visibility biases the visible sample." But on import-vs-recognize it is one of the strongest case-(B) instances in the corpus: the entry states outright that the file drawer is survivorship bias applied to the conducted-studies population, the identical mechanism as countless named cases in other fields, so the reuse across substrates is recognition of the same mechanism, not metaphor.
The portable structural skeleton is therefore an established prime family: a selection filter on what becomes visible biases the visible record relative to the underlying population — selection_bias, survivorship_bias, and selection_effect (one mechanism under three catalog names). That skeleton is maximally substrate-portable, and it is exactly what the file drawer instantiates from its umbrella, not what makes "the file drawer problem" itself travel: the cross-domain reach belongs to the selection/survivorship parents, while the publication-pipeline substrate, the conducted-research framing, and the meta-analytic correction toolkit are the accent that stays home. Its character: a de-moralized, low-evaluative-weight meta-research phenomenon whose distinctive apparatus is publication-system furniture, but which is pulled well into mixed territory because the visibility-filter mechanism it instantiates is one of the cleanest and most widely recognized cross-domain structures there is — structural in that borrowed skeleton, framed in its publication-institution specifics.
Structural Core vs. Domain Accent¶
This section decides why the file drawer problem is a domain-specific abstraction and not a prime — and it is one of the clearest cases, because its underlying structure is an established substrate-neutral prime family that other fields already name for themselves.
What is skeletal (could lift toward a cross-domain prime). Strip the publication pipeline and a textbook relational structure survives: a selection filter on what becomes visible biases the visible record relative to the underlying population, so any inference drawn from the survivors is systematically distorted in the direction of the filter. The portable pieces are abstract — an underlying population, a visibility filter whose admission criterion is correlated with the quantity of interest, and a downstream inference that mistakes the filtered survivors for the whole. That skeleton is maximally substrate-portable and recurs across domains that share no publishing machinery whatever — it is the exact mechanism of Wald's returning-aircraft fallacy and countless other named cases. Precisely because it recurs, it is carried by the parents the file drawer instantiates — survivorship_bias (the file drawer is survivorship bias applied to the population of conducted studies), and the broader selection_bias and selection_effect (one mechanism under three catalog names). But that is the core the file drawer shares, not what makes it distinctive.
What is domain-bound. What makes this specifically the file drawer problem is meta-research-and-publication furniture and none of it survives extraction. Its worked content presupposes science as a publication-mediated knowledge system: the scientific-publication pipeline (journals, reviewers, editorial novelty/significance criteria), the conducted-research-population framing, the signed pull toward zero, and the named diagnostic-and-reform machinery (funnel plots, trim-and-fill, p-curve, Egger regression, fail-safe N, selection models, clinical-trial registries, registered reports, reporting mandates). The empirical cases (Turner et al.'s 74 antidepressant trials, the ICMJE/FDAAA registration reforms) are drawn from it. The decisive test: strip the vocabulary — publication, journal, study, meta-analysis, registration — and what remains is "a filter on visibility biases the visible sample," which simply is selection/survivorship bias at the prime level, with none of the correction toolkit having a referent outside a publication system. The publication-pipeline substrate and its apparatus are the accent, and they stay home.
Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose cross-domain transfer is recognition of the same mechanism, not analogy — and the file drawer is one of the cleanest cases where that transfer is recognition, not metaphor, precisely because the shared mechanism is already a catalogued prime family. Its transfer is bimodal. Within meta-research, evidence synthesis, and scientific reform it moves intact as mechanism — the three-part model, the funnel-fingerprint diagnostic, the signed pull, and the prevent-or-correct fork carry without modification across meta-analysis methodology, the replication crisis, drug-trial regulation, and pre-registration, because the substrate is constant. Beyond that substrate the mechanism still recurs — but as co-instances of the parent selection/survivorship family, which each domain exhibits in its own terms (Wald's aircraft, the visible-winners of any survivorship trap), not by importing "the file drawer problem." So the cross-domain reach belongs to selection_bias, survivorship_bias, and selection_effect, whose substrate-neutral form is what actually generalizes; "the file drawer problem," as named, is the meta-research instantiation that adds the publication pipeline and its correction apparatus on top — baggage that should stay home.
Relationships to Other Abstractions¶
Current abstraction File Drawer Problem Domain-specific
Parents (1) — more general patterns this builds on
-
File Drawer Problem is a kind of Publication Bias Domain-specific
Drawer Problem is the whole-study null-suppression species of the broader results-dependent publication and reporting filter.Publication Bias supplies the genus: A scientific record becomes systematically unrepresentative when the probability that a study, result, or outcome becomes publicly available depends on its direction, magnitude, statistical significance, novelty, or sponsor-favoredness. File Drawer Problem preserves that general structure while adding its differentia: Recognize that studies with null results disproportionately go unpublished while significant ones enter the literature, so any synthesis treating the published record as the full population of conducted research systematically overestimates effect sizes toward the filter. The parent can occur without those added commitments, whereas removing the parent structure leaves no basis for classifying the child as this subtype. That asymmetry establishes subsumption rather than mere association.
Hierarchy paths (12) — routes to 6 parentless roots
- File Drawer Problem → Publication Bias → Selection Bias → Bias
- File Drawer Problem → Publication Bias → Selection on Noisy Estimates → Selection Bias → Bias
- File Drawer Problem → Publication Bias → Selection Bias → Statistical Inference → Inductive Reasoning
- File Drawer Problem → Publication Bias → Selection Bias → Statistical Inference → Uncertainty
- File Drawer Problem → Publication Bias → Selection Bias → Vantage-Induced Omission → Viewpoint
- File Drawer Problem → Publication Bias → Selection on Noisy Estimates → Selection Bias → Statistical Inference → Inductive Reasoning
- File Drawer Problem → Publication Bias → Selection on Noisy Estimates → Selection Bias → Statistical Inference → Uncertainty
- File Drawer Problem → Publication Bias → Selection on Noisy Estimates → Selection Bias → Vantage-Induced Omission → Viewpoint
- File Drawer Problem → Publication Bias → Selection Bias → Statistical Inference → Probability → Measure → Set and Membership
- File Drawer Problem → Publication Bias → Selection Bias → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
- File Drawer Problem → Publication Bias → Selection on Noisy Estimates → Selection Bias → Statistical Inference → Probability → Measure → Set and Membership
- File Drawer Problem → Publication Bias → Selection on Noisy Estimates → Selection Bias → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
Not to Be Confused With¶
-
Selective outcome reporting. Reporting only the favourable outcomes or analyses within a study that was published, while burying the rest. It is a sibling under the same publication-bias umbrella, but it operates inside a visible study, whereas the file drawer suppresses whole completed studies at the publication gate. Tell: is the missing evidence a discarded outcome within a published paper (selective outcome reporting), or an entire null study that never became visible at all (file drawer)?
-
Small-study effects / between-study heterogeneity. Genuine reasons small studies show systematically different (often larger) effects — a real dose-response with sample size, or true variation in effect across populations — which produce the same funnel-plot asymmetry the file drawer does (the entry's own T2). The asymmetry is a fingerprint the filter shares with innocent causes, so a confident file-drawer correction can misattribute ordinary heterogeneity to selective filing. Tell: could the effect-versus-precision asymmetry arise from a true small-study effect or population heterogeneity, or does it specifically trace a significance filter that removed null studies?
-
Winner's curse (significance-threshold inflation). The statistical fact that, among studies clearing a significance threshold, the estimated effect overshoots the true one because only large sampling draws pass — an inflation arising within the observed studies from selecting on significance. The file drawer is the complementary distortion at the level of which studies exist in the record at all; both inflate pooled effects and are easily merged. Tell: is the inflation in the surviving estimates' magnitudes given the threshold they passed (winner's curse), or in the biased composition of the visible sample from unpublished nulls (file drawer)?
-
The replication crisis. The broad empirical finding that many published effects fail to replicate. The file drawer is one proximate mechanism behind it, not the crisis itself — a part-to-whole relation — and it co-produces the distortion together with p-hacking and HARKing. Tell: the replication crisis names the outcome (effects not holding up) across many causes; the file drawer names one cause (selective non-publication of nulls) that inflates the record.
-
The selection/survivorship parents (
selection_bias,survivorship_bias,selection_effect). The substrate-neutral prime family the file drawer instantiates — a filter on visibility biasing the visible sample — of which it is the publication-pipeline case, the exact mechanism of Wald's returning aircraft. These carry the cross-domain lesson; "file drawer problem" adds the journals-and-registries apparatus. Tell: strip publication, journal, and meta-analysis and what remains — a filter on visibility distorts inference from the survivors — is these parents at work, not this named problem. (Treated fully in the sections above.)
Neighborhood in Abstraction Space¶
File Drawer Problem sits in a moderately populated region (59th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Publication Bias & Research Artifacts (5 abstractions)
Nearest neighbors
- Funnel Plot Asymmetry — 0.91
- Small-Study Effects — 0.86
- Type M Error — 0.85
- Cherry Picking — 0.82
- Outbreak Underascertainment — 0.82
Computed from structural-signature embeddings · 2026-07-12