Skip to content

Counterfactual Origin and Omitted-Founder Probe

Structured review — instantiates Founding Population Composition and Drift Management

Tests how plausible alternative gates or omitted founders could have changed descendant composition, capabilities, and vulnerabilities.

The founders you actually admitted are visible; the ones you didn't are the whole point of this review. Counterfactual Origin and Omitted-Founder Probe asks a single disciplined question — if a plausible excluded source had crossed the gate, or if the gate had been drawn differently, how would today's descendant population differ? — and answers it by re-running the founder effect on the counterfactual seed instead of the real one. Its defining move is that it changes nothing in the world: it is an analytic thought-experiment that perturbs the founding gate on paper, propagates the alternative composition forward through the amplification model, and reads the resulting descendant population against the population's declared purpose. Where a running population tells you what its origin did produce, this probe estimates what a nearby origin would have produced — surfacing the correlated vulnerabilities and missing capabilities that the actual gate quietly baked in.

Example

A team maintaining a large language model reviews the seed corpus their base model was pretrained on years earlier. The corpus was assembled under deadline from three sources that happened to be cheap to license, all skewed toward formal written English. The model now generates synthetic data used to train its successors, so any origin skew is being copied forward. Rather than wait for downstream failures, they run a Counterfactual Origin and Omitted-Founder Probe. They name a plausible omitted founder — a large corpus of transcribed spoken dialogue that was available at the time but excluded on cost — and ask what would have changed had it crossed the gate.

Propagating that alternative seed through their model of how corpus composition amplifies (spoken-dialogue patterns would have been reinforced through each synthetic-data generation), they estimate the counterfactual descendant: a model markedly stronger at conversational repair and disfluency, and less brittle on colloquial input — exactly the capability gap real users now report. Judged against the model's stated purpose (a general assistant), the probe concludes the omission was not neutral but a load-bearing origin choice. Crucially, the probe delivers no fix; it delivers a finding — that conversational competence is an origin-level vulnerability — which then justifies commissioning an actual data-acquisition program.

How it works

  • Name plausible counterfactual founders. Enumerate sources that could realistically have crossed the gate but didn't, and alternative gate rules that were genuinely on the table — not fantasy inputs, only near-misses.
  • Redraw the gate on paper. Specify exactly how the admission rule would change to let each counterfactual founder in, holding everything else fixed.
  • Re-propagate. Push the alternative founding composition through the amplification model to estimate the counterfactual descendant population.
  • Read against purpose. Compare the counterfactual descendants with the actual population relative to the declared viability or target reference — capabilities gained or lost, vulnerabilities added or removed.
  • Report load-bearing omissions. Flag which exclusions materially changed the outcome, so real interventions can target those rather than cosmetic gaps.

Tuning parameters

  • Counterfactual plausibility band — how far from the actual gate the alternatives may sit. Wide bands surface more but risk fantasy; narrow bands stay credible but may miss the omission that mattered.
  • Number of counterfactuals — one focused probe versus a fan of alternatives. More scenarios map the origin's sensitivity but dilute depth.
  • Propagation horizon — how many descendant generations forward to simulate. Longer horizons expose compounding effects but accumulate model uncertainty.
  • Reference stringency — how demanding the target the counterfactuals are judged against; a strict reference makes more omissions look load-bearing.

When it helps, and when it misleads

Its strength is that it makes the invisible founders arguable: it turns "we used what we had" into an explicit estimate of what that choice cost, and it isolates which exclusions actually drove today's vulnerabilities so that scarce corrective effort targets origin-level causes rather than symptoms. It is at its best exactly when a population looks fine on its own terms but you suspect its terms were set by an accident of the gate.

Its failure mode is that counterfactuals are unfalsifiable in principle — you can never observe the population that wasn't founded, which is the fundamental problem of causal inference[1]. That invites two abuses: constructing a flattering counterfactual to argue the current composition was optimal, or a damning one to justify a predetermined intervention. The tidy narrative of "what might have been" is easy to run backwards to support a conclusion already chosen. The guarding discipline is to fix the plausible counterfactual set and the reference before propagating, and to report the estimate as a bounded argument about origin sensitivity, never as a measured fact.

How it implements the components

  • founding_gate_definition — it operates on the gate directly, articulating the actual admission rule precisely enough to redraw it counterfactually.
  • founder_effect_amplification_model — it consumes and exercises this model to propagate each alternative seed forward into a counterfactual descendant population.
  • target_population_or_viability_reference — it judges the counterfactual descendants against the declared purpose, which is what makes an omission "load-bearing" rather than merely different.

It does not measure the actual founders' realized influence — that is founding_composition_and_contribution_map, built by Effective Founder Contribution Analysis. And unlike its nearest twin Replicate-Foundation Experiment, it does not implement replicate_foundation_comparator: this probe reasons about unbuilt origins on paper, while the experiment actually stands up parallel founder sets and measures their divergence.

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: Counterfactual Origin and Omitted-Founder Probe operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it tests how plausible alternative gates or omitted founders could have changed descendant composition, capabilities, and vulnerabilities.

Independent corroboration: The frozen evidence defines Counterfactual Origin and Omitted-Founder Probe as 'Tests how plausible alternative gates or omitted founders could have changed descendant composition, capabilities, and vulnerabilities', so its operative form is Assessment, Review & Assurance.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Biology & Ecology

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Population genetics and ecology cohered founder effects as durable composition shifts caused by the traits of a small initial population.

Related originating lineages:

Review resolution: Both reviewers agree on the biological founder-effect primary and statistical counterfactual lineage. Organizational demography is retained because the entry explicitly extends founder composition to institutions, making the cross-domain probe synthetic.

Attribution caveat: The probe transfers a biological founder-effect model to institutional and designed populations.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Reconciled after independent review; medium confidence.

References

[1] The fundamental problem of causal inference (Holland, 1986): for any unit you can observe only one of the potential outcomes, never both the factual and the counterfactual. It is why an omitted-founder probe yields a modeled estimate rather than a measurement, and why its counterfactual set must be fixed in advance to stay honest. withdrawn registry