Lincoln Index¶
A two-sample overlap estimate of an unobserved population size under stated capture assumptions.
Core Idea¶
The Lincoln index estimates a total that neither sample alone reveals. In the standard two-occasion Lincoln–Petersen form, E1 distinct population members are first observed and identifiable, E2 are observed on a later occasion, and S of those E2 match the first set. The intuitive equation E1/N ≈ S/E2 gives N̂ ≈ E1E2/S for S>0. It is a model-based estimate, not a direct count. Identity matching, a stable target frame, mixing and comparable detection make the ratio interpretable; low or zero overlap weakens or defeats this simple form.
The named method originated in wildlife abundance work but the overlap relation can be reused for species lists, vocabularies and other ascertainable populations when their detection processes are justified. Lincoln's 1930 waterfowl circular tabulated real banding returns and proposed a related proportion calculation; it should not be redescribed as the exact 40/50/10 classroom recapture example. The statistical abstraction is a strict Estimation child: unknown N, incomplete observed sets, explicit assumptions and uncertainty. It is not automatically a Measurement child merely because the historical name calls it an index; the result is inferred from partial observations.
Structural Signature¶
Sig role-phrases:
- bounded target population — Declares the population and time frame whose total number is unknown. It is constitutive. Counterfactual: Counts from different regions or turnover periods cannot be merged as though they refer to one closed target.
- first observed set — Records E1 distinct target members and their identity for later overlap recognition. It is constitutive. Counterfactual: An unidentifiable aggregate count cannot reveal which items recur in sample two.
- second observed set — Records E2 distinct target members under a second sampling occasion or genuinely comparable detection process. It is constitutive. Counterfactual: A duplicate copy of the first list is not a fresh detection sample.
- recognized overlap — Counts S members observed in both sets and drives the ratio E1 E2 / S when S is positive. It is constitutive. Counterfactual: If S=0 the simple ratio is undefined; a low S is unstable.
- sampling-model check — Qualifies closure, mixing, identity retention and comparable detection; the formula is not self-validating. It is boundary. Counterfactual: Strong heterogeneity or list dependence can bias the estimated total.
What It Is Not¶
- Not a census. Unseen members are inferred rather than counted individually.
- Not any two totals. Matched identities and an overlap are required.
- Not valid at zero overlap. The simple quotient cannot divide by zero.
- Not assumption-free. Dependence and heterogeneous catchability can bias the output.
- Closest near-miss. Two lists with the same raw totals but no way to match repeated individuals are the closest excluded neighbor: E1 and E2 exist, yet S cannot be identified.
Scope of Application¶
- Wildlife population estimation. Use marked recaptures with closure and detectability assessed.
- Species richness. Compare observer lists only when shared species are identifiable and effort is modeled.
- Vocabulary sampling. Infer an unseen type inventory cautiously from list overlap and word-frequency skew.
- Historical methods analysis. Distinguish Lincoln's band-return proposal from the later textbook two-occasion formula.
Clarity¶
Name the one target population, each observed set and their common identities. Forty, fifty and ten matches give 200 under the simple model, not a directly witnessed count. Two unrelated totals with no matchable items are the nearest miss. A tiny or zero overlap makes the answer unstable or undefined, and biased capture can remain even when arithmetic is clean.
Manages Complexity¶
Two sample sizes and one overlap compress many individual capture histories into an estimate of unseen membership. That makes a hard counting problem tractable while retaining a visible diagnostic: the overlap fraction. The compression loses differences in individual detectability and population turnover. An analyst must reopen those conditions before moving from a quotient to an ecological or other substantive claim.
Abstract Reasoning¶
- Bound the population and interval before counting anything.
- Record unique first-sample and second-sample identities and establish which reappear.
- Check closure, mark recognition, mixing, independence structure and detection heterogeneity.
- Compute the quotient only with positive S, and report sample counts alongside the estimate.
- Explain uncertainty and avoid treating the estimated N as a complete observed census.
Knowledge Transfer¶
The overlap estimator transfers from animal capture to species and vocabulary ascertainment only when both lists refer to one bounded universe and matching/detection assumptions are credible. Lincoln's duck return rate cannot be moved unchanged to word-frequency data, whose sampling is highly uneven. Estimation is the portable parent; the two-sample overlap model is the specialist mechanism and does not describe arbitrary unknown-value inference.
Examples¶
Canonical¶
In one closed survey frame, a first occasion identifies 40 distinct animals, a second identifies 50, and 10 of the second set match marked first-set members. The simple overlap estimate is 40×50/10 = 200 target animals. The worked arithmetic does not itself prove closure, mixing or equal capture probability; changing the ten matches to zero would make this simple quotient undefined.
Mapped back: bounded target population → one closed animal population during the two occasions; first observed set → 40 identified first captures; second observed set → 50 identified second captures; recognized overlap → 10 recognized recaptures, giving estimate 200; sampling-model check → closure, mark retention and comparable detection are explicit assumptions.
Applied / In Practice¶
Lincoln's 1930 USDA circular uses actual 1920–1926 banding and first-season return records for ducks, totaling 17,449 banded and 2,083 returned in its table, to demonstrate a return-percentage route toward estimating continental waterfowl abundance. This is an attested data-based use of marked/returned overlap logic; it is not identical to the toy two-occasion E1E2/S example, and Lincoln calls the abundance method preliminary because hunting totals and sampling bias matter.
Mapped back: bounded target population → North American wild ducks in specified seasons; first observed set → recorded banded ducks; second observed set → first-season hunting-return records, a special observation process; recognized overlap → bands returned from members of the banded set; sampling-model check → station timing, sample scale and external kill total remain caveats.
Structural Tensions¶
T1 — Simple Inverse-Overlap Arithmetic versus Heterogeneous Detection. The product-overlap quotient is easy to compute, but rare species, trap response or unequal catchability can alter the observed overlap without a corresponding change in true N. The index gains utility from simplicity only if its capture model is checked.
Diagnostic: Could one type of target be systematically missed on either occasion?
T2 — Observed Return Fraction versus Population Abundance Claim. Lincoln's banding table provides real return counts; turning those fractions into a continent-wide abundance also needs hunting totals and representativeness. Treating a tabulated return rate as an observed full population would confuse evidence with inference.
Diagnostic: What additional target-total information and sampling assumptions connect the observed fraction to N?
Structural–Framed Character¶
Lincoln index is structural-leaning: the set-overlap relation is mathematical, while the population frame and detectability judgments are field-dependent. Evaluative weight: a numerical estimate is not automatically accurate. Human-practice-bound: mark recognition and sampling protocols are operational choices. Institutional origin: Lincoln's wildlife history names the index but does not restrict all later list uses. Vocabulary travels: overlap and estimate transfer, whereas capture and species are habitat-specific. Import versus recognize: two species lists with a justified common universe can instantiate it; two unrelated spreadsheet totals cannot.
The verified portable parent is Estimation; an overlap-based population estimator is a narrower, possibly future-prime candidate pattern. Its character: a conditional inverse-overlap inference whose simple arithmetic depends on hard sampling assumptions.
Structural Core vs. Domain Accent¶
The ratio is simple; the sampling model is not.
What is skeletal. Partial evidence about an unknown is converted into a usable estimate, preserving assumption and uncertainty disclosure. That is the Estimation parent.
What is domain-bound. Two membership samples of a bounded population and their positive identity-level overlap drive E1E2/S. Wildlife banding, species lists and vocabularies each require different capture/detection checks.
Why this does not clear the prime bar. Many estimates use regression, expert judgment or sensors rather than recapture overlap. Conversely, a numeric overlap of arbitrary lists need not estimate a meaningful population. The whole Lincoln identity is a specialist statistical method, not estimation in general.
Instantiates / Related Primes¶
This entry is a kind of Estimation.
-
Strict parent — estimation. The method derives unknown N from incomplete detection evidence under explicit assumptions.
-
Related — mark–recapture. The animal two-occasion method is a central habitat, while list-overlap variants need their own detection analysis.
-
Related — measurement. Sample membership is observed, but the population total is inferred rather than directly measured on a scale.
Relationships to Other Abstractions¶
Current abstraction Lincoln Index Domain-specific
Parents (1) — more general patterns this builds on
-
Lincoln Index is a kind of Estimation Prime
The Lincoln index estimates unknown population size from two incomplete samples and their overlap under a declared detection model.Estimation derives an unknown usable value from incomplete, noisy or indirect evidence with assumptions and uncertainty. Every valid Lincoln-index use estimates unknown N from two observed membership sets and their overlap, conditional on closure and detection assumptions. Estimation has many other methods with no recapture sets; hence this is strict upward subsumption. Measurement is not retained as a parent because its exact signature distinguishes upstream instrument-to-scale observation from downstream estimation/inference, which is the Lincoln index's constitutive act.
Hierarchy path (1) — routes to 1 parentless root
- Lincoln Index → Estimation → Approximation → Representation → Abstraction
Neighborhood in Abstraction Space¶
Lincoln Index sits in a moderately populated region (43rd percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Domain-Specific Indicators & Measurement Methods (26 abstractions)
Nearest neighbors
- Inferential Error — 0.88
- Estimation of Covariance Matrices — 0.87
- Bartlett's theorem — 0.87
- Open-access citation advantage — 0.87
- Cooperativity — 0.87
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Complete census. Tell: Was every member observed, or is unseen membership inferred?
- Two unmatched list lengths. Tell: Can shared identities be counted as S?
- Zero-overlap quotient. Tell: Is the denominator positive?
- Unconditional population fact. Tell: Are closure and detectability assumptions stated?
References¶
- USGS, Statistical inference from capture–recapture experiments (2008): https://pubs.usgs.gov/of/2008/1363/pdf/OF08-1363_508.pdf
- Frederick C. Lincoln, USDA Circular 118, Calculating waterfowl abundance on the basis of banding returns (1930): https://ia801906.us.archive.org/9/items/calculatingwater118linc/calculatingwater118linc.pdf
- Frozen Wikipedia discovery revision: https://en.wikipedia.org/wiki/Lincoln_index (revision 1325641448).