Skip to content

Haldane's Sieve

Explain why new beneficial mutations that are dominant fix far more often than recessive ones — because while rare an allele sits almost only in heterozygotes, so selection sees a dominant from its first copy but is blind to a masked recessive, which usually drifts to loss before homozygotes form.

Core Idea

Haldane's sieve is the population-genetic result, derived by J. B. S. Haldane in 1927, that new beneficial mutations are far more likely to escape early stochastic loss and ultimately fix in a diploid population when they are dominant than when they are recessive. The mechanism is one of visibility: while a beneficial allele is rare, it exists almost exclusively in heterozygotes. If the allele is dominant, its fitness advantage is expressed immediately in those heterozygotes, and natural selection pushes the allele upward in frequency from the moment it first appears. If the allele is recessive, its beneficial effect is masked in heterozygotes — the phenotype is indistinguishable from the wild type — and selection cannot act on it. The allele drifts as though neutral, and the vast majority of such recessive alleles are eliminated by genetic drift before they ever rise to the frequency at which homozygotes form in appreciable numbers and selection can finally distinguish carriers. The fixation probability of a new recessive beneficial allele is thus much lower than that of a dominant one with the same selection coefficient.

The sieve metaphor captures this asymmetric loss: the population acts as a filter that retains dominant beneficial mutations and discards most recessive ones before they can spread. The consequence for the observable genetic architecture of adaptation is systematic: the alleles that actually fix and produce the signature of a selective sweep in population-genomic data are enriched for dominant phenotypic effects relative to the underlying distribution of beneficial mutations generated by the mutation process. The empirical record of adaptation is not a random sample of advantageous mutations; it is a dominance-biased sample produced by the sieve operating on that input.

A critical boundary condition determines when the sieve is active and when it is bypassed. For adaptation from new mutation — a beneficial allele arising once and spreading from a single initial copy — the sieve applies fully. For adaptation from standing variation — a previously neutral or slightly deleterious allele that was already present at some frequency in the population when the environment changed and made it beneficial — the sieve is weaker or absent, because the allele starts at a frequency where homozygotes already exist and selection is already acting. This distinction between hard sweeps from new mutation and soft sweeps from standing variation, a major organising contrast in modern molecular population genetics, has Haldane's sieve as one of its theoretical pillars. The dominance distribution of alleles involved in hard sweeps should differ from that of alleles involved in soft sweeps, and observed cases of rapid adaptation — antibiotic resistance arising from new mutation versus spreading from a pre-existing resistant minority — reflect this structure.

The sieve's strength scales with population size. In very small populations, genetic drift overwhelms selection for all but the most strongly selected alleles, and the distinction between dominant and recessive barely matters — both are likely to be lost. In large populations, selection is more effective even on heterozygotes, but a recessive beneficial allele must still drift across the frequency range where homozygotes are rare before selection takes hold. Intermediate population sizes produce the clearest sieve action. In conservation genetics, this scaling matters: small populations under selection pressure have reduced access to recessive beneficial variation from new mutation, compounding the already-limited genetic diversity available for adaptation.

Structural Signature

Sig role-phrases:

  • the diploid population — a sexually-reproducing population of effective size N_e in which new alleles begin at frequency ~½N_e, almost entirely in heterozygotes while rare
  • the new beneficial allele — an advantageous variant whose fate (loss or fixation) is in question
  • the dominance coefficient — h, governing how much of the allele's benefit is expressed in heterozygotes and thus how visible it is to selection while rare
  • the visibility-while-rare — effective selection on the rare allele scaling with h·s, so a dominant is seen from its first copy and a recessive is masked
  • the drift loss of the masked — genetic drift eliminating most recessive beneficials before they ever reach the frequency at which homozygotes form and selection can act
  • the homozygote-forming regime shift — the point where frequency rises enough for homozygotes to be common, selection acts on full s, and the bias switches off
  • the dominance-biased fixed record — the engineered output: the alleles that actually sweep are enriched for dominant effects relative to the mutational input, a filtered not random sample
  • the new-mutation-versus-standing-variation boundary — the condition that decides whether the sieve bites (hard sweep) or is bypassed (soft sweep, homozygotes already present), with population size setting the sieve's strength

What It Is Not

  • Not a claim that recessive beneficial alleles cannot fix. The sieve makes recessive fixation much less probable from new mutation, not impossible. A recessive can fix readily when it starts from standing variation, or when migration or inbreeding raises its frequency past the range where homozygotes are rare so selection can finally act. The result is a probability bias, not a prohibition.
  • Not selection eliminating recessives. Most recessive beneficials are lost to genetic drift, not removed by selection. While the allele is rare and masked in heterozygotes, selection is blind to it — it drifts as though neutral — so the loss is the random walk reaching zero, not selection acting against an allele it cannot even see.
  • Not always operative. The sieve bites only while the favourable allele is rare and homozygotes are absent — the new-mutation, hard-sweep regime. It weakens or vanishes for adaptation from standing variation (soft sweeps), where the allele already sits at appreciable frequency, homozygotes exist, and selection is already acting. Whether the sieve applies is itself the hard-versus-soft-sweep dividing line.
  • Not a random sample of beneficial mutations. The alleles that actually fix and leave a sweep signature are a dominance-biased sample, enriched for dominant effects relative to the underlying mutational input. The empirical record of adaptation is the output of the filter, not a faithful draw from what the mutation process generated — reading it as representative of beneficial mutation is exactly the error the sieve corrects.
  • Not the general "must be expressed to be selectable" intuition. That broad selection-visibility pattern — a selector samples only what is expressed, biasing the surviving record — is the portable parent and recurs in idea selection, sparse-feature learning, and market dynamics. Haldane's sieve is the diploid-genetics quantitative implementation: dominance coefficient, h·s heterozygous selection, the ½N_e initial frequency, the fixation-probability scaling. That machinery does not travel; only the bare visibility intuition does.

Scope of Application

Haldane's sieve lives within the population-genetics and adaptation subfields of biology, ranging over diploid (and partially polyploid) sexually-reproducing populations under selection; its reach is bounded by that domain, because the load-bearing content is irreducibly quantitative — the dominance coefficient, h·s heterozygous selection, the ½N_e initial frequency, the fixation-probability scaling. The broader selection-visibility intuition that recurs in idea or market selection belongs to a candidate parent, not the sieve. Within the domain it operates across these contexts.

  • Adaptive evolution theory — the baseline expectation that the dominance distribution of alleles fixed during recent adaptation differs from the mutational-input distribution, skewed toward dominants.
  • Crop genetics — explains why deliberately introgressed dominant resistance genes spread faster than recessive ones via single crosses, and informs marker-assisted selection and the use of inbreeding to manufacture homozygotes for recessive deployment.
  • Antibiotic and pesticide resistance — resistance alleles observed spreading in wild populations are dominance-biased on the target trait, with recessive resistance more often arising from standing variation than from new mutation.
  • Conservation genetics — bears on which adaptive responses are accessible to small populations, where strong drift makes the sieve more severe and fresh recessive beneficial variation harder to retain.
  • Hard-versus-soft-sweep theory — supplies the theoretical reason hard sweeps (from new mutation, sieved) should differ in dominance composition from soft sweeps (from standing variation, bypassing the sieve), a central organizing contrast in molecular population genetics.

Clarity

The sieve dissolves an apparent puzzle in adaptive evolution: the dominance composition of alleles that actually fix during recent adaptation looks systematically unlike the distribution of beneficial mutations the mutation process generates — too many dominants. Naming the sieve shows this needs no special biological mechanism. It is the expected consequence of a single fact about diploids: while a beneficial allele is rare it sits almost only in heterozygotes, so selection can act on it from the first copy if it is dominant but is blind to it if it is recessive, and most recessives drift to loss before homozygotes ever form. The clarifying move is to separate the mutational input (what arises) from the fixed record (what spreads) and to identify the population itself as the filter between them — so the genomic signature of adaptation is understood as a dominance-biased sample, not a random draw of advantageous mutations.

The sieve's second clarifying contribution is to make the boundary condition crisp, and with it one of the field's central organising contrasts. Because the bias operates only while the favourable allele is rare and homozygotes absent, it applies in full to adaptation from new mutation (a single new copy spreading) but weakens or vanishes for adaptation from standing variation (an allele already present at appreciable frequency when the environment shifts, where homozygotes exist and selection is already acting). That distinction is exactly the theoretical pillar under the hard-sweep versus soft-sweep contrast in molecular population genetics, and naming it lets a researcher ask sharper questions than "did this allele sweep?": is the resistance allele dominant or recessive on the selected trait, did it arise once or pre-exist, and therefore should its dominance match the hard-sweep expectation or the softer standing-variation one? It also makes the population-size dependence reasonable rather than mysterious — drift swamps the distinction in tiny populations, selection reaches even heterozygotes in very large ones, and the sieve bites hardest in between — which in turn frames the conservation question of whether a small population under new selection pressure even has access to recessive beneficial variation arising fresh.

Manages Complexity

The fate of a new beneficial allele is governed by a tangle of interacting quantities — the dominance coefficient, the selection coefficient, effective population size, mutation rate, and initial frequency all jointly determine its fixation probability through stochastic dynamics no analyst wants to integrate by hand for every allele. The sieve compresses that joint dependency into one qualitative principle: while the allele is rare it sits almost only in heterozygotes, so selection sees it from the first copy if it is dominant and is blind to it if it is recessive, and most recessives drift to loss before homozygotes form. From that single statement the consequential prediction falls out — the fixed record of adaptation is dominance-biased relative to the mutational input — so a researcher reads the expected genetic architecture of a sweep off the dominance coefficient rather than simulating each allele's trajectory. The compression also turns the multi-parameter problem into a small set of readable regime switches. One boundary condition (is the allele rare-from-new-mutation, or already at appreciable frequency in standing variation?) decides whether the sieve bites at all, which is exactly the dividing line between the hard-sweep and soft-sweep cases; and one further parameter, population size, fixes the strength — drift swamps the dominant/recessive distinction in tiny populations, selection reaches even heterozygotes in very large ones, and the sieve acts most cleanly in between. So instead of tracking five interacting variables through a fixation calculation, the analyst reasons from dominance, the new-mutation-versus-standing-variation boundary, and population size, and reads the qualitative outcome — which alleles pass the filter, what the swept record will look like, whether a small population can even access fresh recessive variation — off that compact set.

Abstract Reasoning

Haldane's sieve licenses a set of inferences that all run on its three working parameters — the dominance coefficient of the beneficial allele, the new-mutation-versus-standing-variation boundary, and effective population size.

Diagnostic. The sieve infers a hidden property of the mutational process from the observable record of adaptation, and the move runs from a swept genomic signature back to the dominance of the alleles that produced it. Because the population retains dominants and discards most recessives before they ever rise to where homozygotes form, the fixed record is a dominance-biased sample: the alleles caught in selective sweeps are enriched for dominant phenotypic effects relative to the underlying distribution of beneficial mutations. So an excess of dominants among recently swept loci is read not as a fact about what mutation generated but as the fingerprint of the sieve operating on that input — and conversely, a recessive allele observed to have fixed is diagnostic, inferred to have arrived either from standing variation or through a frequency-raising route (migration, inbreeding) that let it cross the homozygote-forming range under selection's eye. The type of sweep is read off the same logic: a recessive selected allele behind a sweep points to standing variation, a dominant one is consistent with the hard-sweep-from-new-mutation expectation.

Interventionist. The mechanism specifies exactly which lever moves the outcome — manipulate the frequency at which an allele meets selection and you bypass the sieve. The operative move for deploying a recessive beneficial allele is to raise its frequency past the range where homozygotes are rare before relying on selection: inbreeding to manufacture homozygotes, or introducing the allele simultaneously at many sites to lift local frequency, each predicts that a recessive that would otherwise drift to loss will instead respond to selection and spread. A dominant allele needs no such intervention; the prediction is that it climbs from its first copy, which is the rationale under which dominant resistance genes are favoured for single-cross introgression in breeding. The dominance coefficient is itself the dial: the effective strength of selection on a rare allele scales with how much of its benefit is expressed in heterozygotes, so an intervention that increases heterozygous expression predicts a higher fixation probability.

Boundary-drawing. The sieve forces a regime decision the bare observation of adaptation cannot settle: is this allele rare-from-new-mutation, where the sieve applies in full, or already at appreciable frequency in standing variation, where it weakens or vanishes because homozygotes already exist and selection is already acting? That boundary is exactly the hard-sweep versus soft-sweep dividing line, and the analyst must place the case on one side before predicting whether dominance should bias its outcome. A second boundary is set by population size: drift swamps the dominant/recessive distinction in very small populations (both are likely lost), selection reaches even heterozygotes in very large ones, and the sieve bites most cleanly at intermediate sizes — so the analyst assigns the population to a regime in which the sieve is inert, severe, or maximally discriminating before reading anything off the dominance coefficient.

Predictive / order-of-events. The sieve predicts a difference between two records that would otherwise be conflated: the dominance distribution of alleles fixed by hard sweeps should be shifted toward dominance relative to that of alleles fixed from standing variation, so a comparison of the two sweep classes is predicted to reveal the sieve's signature directly. It also makes a conservation-relevant forward prediction: a small population newly under selection has reduced access to recessive beneficial variation arising fresh, because such variants drift to loss before selection can rescue them, so its adaptive response is predicted to draw disproportionately on dominant new mutations or on whatever standing variation it already carries — the order being that drift acts first on the rare recessive, foreclosing it, before selection ever gets the chance to act.

Knowledge Transfer

Within population genetics Haldane's sieve transfers as mechanism across all diploid (and partially polyploid) sexually-reproducing organisms, because the cargo is one quantitative result — while a beneficial allele is rare it sits almost only in heterozygotes, so effective selection on it scales with h·s and recessives drift to loss before homozygotes form. From adaptive-evolution theory it carries directly into crop genetics (dominant resistance genes spread faster than recessives via single-cross introgression; marker-assisted selection design), into antibiotic and pesticide resistance (alleles observed spreading in the wild are dominance-biased on the target trait, with recessive resistance more often arising from standing variation), into conservation genetics (small populations have reduced access to fresh recessive beneficial variation), and into the hard-versus-soft-sweep literature, for which it supplies the theoretical reason the two sweep classes differ in dominance composition. Across all of these the apparatus carries without translation — the dominance-biased fixed record, the new-mutation-versus-standing-variation boundary that decides whether the sieve bites, and the population-size scaling that sets its strength. Its core mechanism — trait expression in heterozygotes mediated by dominance, with selection visibility scaling with expressed phenotype — even extends within biology to adjacent masked-at-low-frequency settings: modifier alleles in epistatic networks, and sex-linked alleles in homogametic carriers, where the same visibility logic governs which rare variants selection can rescue.

Beyond diploid populations under selection the named result does not travel, and the honest split is between loose analogy and a genuinely portable parent. The broader structural intuition — a variant must be expressed to be selectable, so rare-frequency masking biases the set of variants that ever become common (visibility-mediated sampling of a selective process) — does recur in selectionist arguments outside biology: idea selection (an idea unexpressed cannot be picked up), machine learning of sparse features, market selection of rare strategies. But these are (A) loose analogies, and they must be marked so, because the load-bearing content of the sieve is irreducibly quantitative and biological — the dominance coefficient, the h·s heterozygous selection, the ½N_e initial frequency, the fixation-probability scaling, the homozygote-forming threshold — and none of that machinery transfers; only the bare visibility intuition does. The honest characterization is the (B) one: what genuinely recurs cross-domain is that more general selection-visibility pattern — the idea that a selector samples only what is expressed, biasing the surviving record relative to the generating process — and that pattern (a candidate parent for selection-visibility, riding alongside genetic_drift as the loss process and dominance as the masking relation) is what a cross-domain lesson should carry. "Haldane's sieve," as named, keeps the population-genetic cargo that stays home; the exportable insight is the visibility-mediated sampling it instantiates, named at the right generality by the parent rather than by the diploid-genetics result. See Structural Core vs. Domain Accent.

Examples

Canonical

The result is Haldane's own, derived in his 1927 paper on the mathematical theory of natural and artificial selection. Treating a new advantageous mutation as beginning from a single copy and following its lineage as a branching process, Haldane showed that a dominant (or additive) beneficial allele with selective advantage s escapes early chance loss and fixes with probability of roughly 2s — so an allele conferring a 1% advantage fixes about 2% of the time, despite the steep odds against its lone first copy. A fully recessive allele of the same advantage is a different matter: while rare it lives almost entirely in heterozygotes, where its benefit is masked and selection cannot see it, so its fixation probability is far smaller, scaling with a much weaker function of s and population size rather than with s directly. That gap between the two probabilities is the sieve.

Mapped back: Haldane's single-copy start is the new beneficial allele in the diploid population at frequency ~½N_e; the contrast between the dominant's ≈2s fixation chance and the recessive's far lower one is set entirely by the dominance coefficient through the visibility-while-rare term. The recessive's masked benefit in heterozygotes and consequent drift loss of the masked is exactly what the low fixation probability quantifies, so the surviving alleles form the dominance-biased fixed record.

Applied / In Practice

The peppered moth Biston betularia is the textbook natural demonstration of the sieve's retained-dominant side. Before industrialization the pale typica form predominated and the dark melanic carbonaria form, conferred by a dominant allele, was rare. As soot darkened tree bark across industrial Britain through the nineteenth century, the melanic form gained a strong survival advantage against visual predation by birds, and its frequency climbed toward near-fixation in polluted regions within a few decades — among the fastest documented episodes of natural selection in the wild. Because carbonaria is dominant, its advantage was expressed in the very first heterozygous carriers, so selection could grip the allele from the moment it was rare rather than waiting for scarce homozygotes to form.

Mapped back: The moth population is the diploid population and the melanic variant is the new beneficial allele with a dominant dominance coefficient. Its expression in heterozygotes is the visibility-while-rare that let selection act from the first copies — sparing it the drift loss of the masked that a recessive melanic allele of equal advantage would likely have suffered. Its rise to near-fixation is the sieve retaining a dominant, the observable dominance-biased fixed record in action.

Structural Tensions

T1: Probability bias versus prohibition (how sharp is the sieve). The word "sieve" invites the reading that recessive beneficials are simply blocked — but the result is a probability skew, not a gate. A recessive with the same selection coefficient fixes far less often from new mutation, yet it fixes readily from standing variation, and migration or inbreeding can raise its frequency past the homozygote-forming range so selection grips it. The tension is that the mechanism's evocative filter imagery overstates its own action: read as a prohibition it wrongly rules out documented recessive sweeps and the whole soft-sweep route; read as merely "somewhat less likely" it understates how severe the drift-loss bias is for a rare masked allele in an intermediate-sized population. The honest content sits between an absolute barrier and a mild tilt, and which end one leans toward changes the inference. Diagnostic: Is the claim here that recessive fixation is rare-from-new-mutation, or the stronger and false claim that it is precluded?

T2: Retention by selection versus loss by drift (a selective filter that works by non-selection). The sieve is a selective phenomenon, yet its discarding arm is not selection at all: recessive beneficials are lost to genetic drift while selection is blind to them, masked in heterozygotes. So the "filter" retains dominants by selection but discards recessives by the random walk reaching zero — two different forces wearing one metaphor. The tension matters for intervention: the intuitive fix "let selection weed them in" is incoherent for the recessive arm, because selection cannot see a masked allele no matter how strong s is; the only lever is to change the frequency at which the allele meets selection, not the strength of selection itself. Misreading the loss as selective points effort at the wrong dial. Diagnostic: Is the recessive's loss here attributed to selection acting against it (wrong) or to drift eliminating an allele selection cannot yet see (right)?

T3: Sieve active versus sieve bypassed (the origin regime and its inference circularity). Whether the sieve bites at all turns on a single boundary — new mutation (rare, homozygotes absent, sieve full) versus standing variation (already at frequency, homozygotes present, sieve weak or absent) — which is exactly the hard-sweep/soft-sweep divide. But this creates a working circularity: the analyst wants to use the dominance of a swept allele to infer whether it arose fresh or pre-existed, while the sieve's applicability depends on that very origin. Dominance and origin are read off each other, so a recessive behind a sweep implicates standing variation only if one has independently excluded a frequency-raising route. The tension is that the sieve's central boundary condition is often the unknown it is being enlisted to resolve, and population size adds a second regime layer (inert in tiny, discriminating at intermediate, weak in huge) that must be fixed before dominance says anything. Diagnostic: Has the allele's origin (new mutation versus standing variation) and the population-size regime been fixed independently, or is dominance being used to infer the very regime that decides whether dominance is informative?

T4: Filtered record versus mutational input (what a sweep census measures). The alleles that fix and leave sweep signatures are a dominance-biased sample, and the sieve's chief diagnostic value is precisely that the fixed record is not a faithful draw of the beneficial mutations the mutation process generated. This cuts two ways. Used correctly, an excess of dominants among swept loci is read as the sieve's fingerprint, not as evidence that mutation mostly produces dominants. Used carelessly, the same census is taken at face value as the distribution of beneficial mutations, importing the filter's bias into a claim about mutational input. The tension is that the observable (swept alleles) and the quantity of interest (generated beneficial mutations) are separated by exactly the sieve, so any inference about mutation from fixation data must first undo the filter it is looking through. Diagnostic: Is the dominance skew in this swept sample being read as the sieve's output, or mistakenly as the mutational input's shape?

T5: Autonomy versus reduction (a quantitative population-genetics result or the selection-visibility parent). Haldane's sieve is a specific, load-bearing population-genetics result whose cargo is irreducibly quantitative — the dominance coefficient, h·s heterozygous selection, the ½N_e initial frequency, the fixation-probability scaling — and within diploid sexually-reproducing populations it transfers as full mechanism. But its exportable intuition is thin: a variant must be expressed to be selectable, so rare-frequency masking biases the set of variants that become common is the parent selection-visibility pattern, and cross-domain echoes (idea selection, sparse-feature learning, market strategies) are loose analogies that borrow only that bare intuition, none of the machinery. The tension is between a named genetic result that anchors adaptation theory and the recognition that its cross-substrate lesson is carried by the general selection-visibility pattern, not by the sieve's diploid quantitative core. Diagnostic: Resolve toward the selection-visibility parent (with drift as the loss process) when the substrate is not a diploid population under selection; toward Haldane's sieve when dominance, h·s, and fixation probability are actually in play.

Structural–Framed Character

Haldane's sieve sits toward the structural end of the spectrum but stops short of the pole — best read as mixed-structural, closely parallel to how isostasy is characterized: a genuine relational mechanism wearing heavy population-genetic vocabulary. Four of the five criteria give it strong structural credentials, and only the vocabulary criterion pulls it back toward its home domain.

Its evaluative weight is nil. A dominant allele climbing from its first copy while a masked recessive drifts to loss is neither good nor bad — the sieve renders no verdict, praises and blames nothing, and names a filtering fact rather than convicting a move the way "ad hominem" or a labelled fallacy does. Human-practice-bound it is not, in the strongest sense: the mechanism runs observer-free — the peppered moth's melanic allele swept toward fixation in industrial Britain with no geneticist present, and recessive beneficials drift to loss whether or not anyone models them; strip away every population geneticist and diploids still filter their beneficial mutations by dominance. Institutional origin is none: Haldane in 1927 described a consequence of how selection acts on rare alleles in diploids, he did not legislate it — it is a fact about branching-process dynamics and heterozygous masking, not an artifact of a survey, agency, or convention. And within its proper range cross-domain reuse is recognition, not import: moving from crop resistance genes to antibiotic resistance to conservation genetics to the hard-versus-soft-sweep contrast, the same mechanism is recognized intact — the dominance-biased fixed record, the new-mutation-versus-standing-variation boundary, the population-size scaling all keep their content, and even the extension to modifier alleles and sex-linked carriers is recognition of the identical visibility logic, not analogy.

What keeps it off the structural pole is vocab_travels, which it fails decisively. The operative vocabulary is irreducibly population-genetic — dominance coefficient, h·s heterozygous selection, the ½N_e initial frequency, fixation-probability scaling, the homozygote-forming threshold — and none of it floats free of diploid sexually-reproducing substrates. Beyond that substrate, "selection can only pick what is expressed" is borrowed by idea selection, sparse-feature learning, and market dynamics as loose analogy that renames every component and keeps none of the quantitative machinery; there the transfer is metaphor, and the import_vs_recognize mark flips from recognition (inside biology) to import-by-analogy (outside it).

The portable structural skeleton is visibility-mediated sampling of a selective process — a selector acts only on what is expressed, so masking at low frequency biases the surviving record relative to the generating distribution, with a stochastic loss process (drift) removing the unexpressed variants before the selector can grip them. That skeleton is genuinely substrate-portable, and it is exactly what the sieve instantiates from its umbrella — the candidate selection-visibility parent (riding alongside genetic drift as the loss process and dominance as the masking relation) — not what makes "Haldane's sieve" itself travel: the cross-domain reach belongs to that general selection-visibility pattern, while the diploid-genetics quantitative implementation stays home. Its character: a structural, evaluatively neutral, recognized-in-nature filtering mechanism whose skeleton is the portable selection-visibility pattern it instantiates, but whose distinctive content is stated in population-genetic vocabulary that pins it to its home domain, leaving it mixed-structural rather than a free-floating prime.

Structural Core vs. Domain Accent

This section decides why Haldane's sieve is a domain-specific abstraction and not a prime, and it carries the case for its domain-specificity — there is no separate section for that.

What is skeletal (could lift toward a cross-domain prime). Strip the population genetics and a thin relational structure survives: a selector acts only on variants that are expressed, so masking at low frequency filters most of the unexpressed candidates out — by a stochastic loss process — before the selector can grip them, and the surviving record is therefore biased relative to the generating distribution. The pieces that travel are abstract: a generator producing candidates, an expression gate that decides which candidates are visible to a selector while still rare, a random loss process that culls the invisible ones, and a resulting sample skewed toward the visible type. That skeleton is genuinely substrate-portable — a selector samples only what is expressed — which is exactly why it recurs as the candidate selection-visibility parent the sieve instantiates, riding alongside genetic_drift as the loss process. But it is the core the sieve shares, not what makes the sieve distinctive; the same bare intuition surfaces in idea selection, sparse-feature learning, and market dynamics without any of the sieve's content.

What is domain-bound. Almost all the load-bearing content is population-genetics furniture and none of it survives extraction intact: the diploid, sexually-reproducing substrate on which "heterozygote" and "homozygote" even mean anything; the dominance coefficient h that sets how much benefit is expressed while rare; the h·s heterozygous selection term; the ½N_e initial frequency of a new mutation; the homozygote-forming threshold at which the bias switches off; the fixation-probability scaling (the ≈2s result for a dominant); and the new-mutation-versus-standing-variation boundary that doubles as the hard-sweep/soft-sweep divide, with population size tuning the sieve's strength. The decisive test: remove the diploid substrate — take away the heterozygote in which a beneficial allele can be masked — and there is no sieve at all, only the loose "you have to be expressed to be selected" resemblance. Masking-by-dominance, the very mechanism that makes the record dominance-biased, has no referent off diploid genetics.

Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose cross-domain transfer is recognition of the same mechanism, not analogy. The sieve's transfer is bimodal. Within diploid, sexually-reproducing populations it travels intact — from crop resistance genes to antibiotic and pesticide resistance to conservation genetics to the hard-versus-soft-sweep contrast, and even to modifier alleles and sex-linked carriers, the dominance coefficient, the h·s selection, and the population-size scaling all keep their content and the mechanism is recognized, not borrowed. Beyond that substrate it travels only by analogy: "an idea unexpressed cannot be selected," market selection of rare strategies, sparse-feature learning each rename every component and keep none of the quantitative machinery. And when the bare structural lesson is needed cross-domain, it is already supplied in more general form by the parent it instantiates — the selection-visibility pattern (a selector samples only what is expressed, biasing the surviving record), with genetic_drift carrying the stochastic loss. The cross-domain reach belongs to that parent; "Haldane's sieve," as named, carries diploid-genetic baggage — dominance, heterozygous masking, fixation probability — that does not and should not travel.

Relationships to Other Abstractions

Local relationship map for Haldane's SieveParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Haldane's SieveDOMAINPrime abstraction: Selection-Visibility Gate — is a kind ofSelection-Visib…PRIME

Current abstraction Haldane's Sieve Domain-specific

Parents (1) — more general patterns this builds on

  • Haldane's Sieve is a kind of Selection-Visibility Gate Prime

    Haldane's sieve is selection visibility specialized to a new beneficial allele whose heterozygous expression scales with dominance while it is rare.

Hierarchy path (1) — routes to 1 parentless root

Not to Be Confused With

  • Genetic drift. The random change in allele frequency from finite-population sampling, with no directional cause. Drift is the loss process the sieve rides on — it is what culls the masked recessive before homozygotes form — but drift alone is undirected and acts on all rare alleles equally, whereas the sieve is the dominance-biased outcome that emerges when drift interacts with heterozygous masking under selection. Tell: is the claim just "rare alleles are lost by chance" (drift, no dominance dependence), or "recessive beneficials are lost far more often than equally-advantageous dominants" (the sieve, which needs dominance and selection on top of drift)?

  • Selective sweep / hard sweep. The observable genomic signature — reduced diversity around a locus — left when a beneficial allele rises to fixation. The sweep is the record; the sieve is the filter that biases which alleles produce it, so that swept loci are enriched for dominant effects relative to the mutational input. A sweep is a datum; the sieve is the explanation for the dominance skew across many such data. Tell: are you describing the diversity trough at one fixed locus (sweep), or the systematic over-representation of dominants among the alleles that sweep (sieve)?

  • Adaptation from standing variation / soft sweep. Rapid adaptation from an allele already present at appreciable frequency when the environment shifted, spreading from multiple copies. This is precisely the regime where the sieve is bypassed — homozygotes already exist and selection is already acting, so a recessive faces no masking-then-drift gauntlet. It is the contrast case that defines the sieve's boundary, not an instance of it. Tell: did the favourable allele arise as a single new copy (sieve bites) or pre-exist at frequency before becoming beneficial (soft sweep, sieve inert)?

  • Haldane's dilemma (cost of selection). A different Haldane result — the ceiling on how many loci can be substituted per generation given the reproductive "cost" of each selective replacement. It shares only the author's name and the substitution setting; it says nothing about dominance or the visibility of rare alleles, and concerns the rate/number of substitutions a population can afford, not which dominance class of beneficial fixes. Tell: is the question how many beneficial alleles can sweep in a given time (the dilemma), or why dominants sweep more readily than recessives (the sieve)?

  • Purging of deleterious recessives / inbreeding depression. The removal of harmful recessive alleles once inbreeding or drift raises their frequency enough to expose them in homozygotes, where selection can finally act against them. This is the deleterious mirror of the sieve: there selection removes recessives once exposed, whereas the sieve concerns beneficial recessives that drift to loss before exposure. Both turn on the fact that recessives are invisible in heterozygotes, but they run in opposite fitness directions. Tell: is selection culling a harmful recessive newly exposed in homozygotes (purging), or failing to promote a helpful recessive still masked in heterozygotes (the sieve)?

  • The selection-visibility pattern (umbrella). The broad, substrate-neutral intuition that a selector acts only on what is expressed, so masking biases the surviving record relative to what was generated — recurring in idea selection, sparse-feature learning, and market dynamics. Haldane's sieve instantiates this pattern with diploid-genetic machinery (dominance, h·s, ½N_e, fixation probability); the umbrella is what travels cross-domain, the sieve is the population-genetics implementation that stays home. Tell: strip away heterozygotes, dominance, and the fixation calculation and what remains is bare "you must be expressed to be selected" — at which point you are invoking the general pattern, not the sieve. (Treated fully in a later section.)

Neighborhood in Abstraction Space

Haldane's Sieve sits in a crowded region of the domain-specific corpus (31st percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Population Genetics & Kin Selection (10 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12