Skip to content

External Validity Check

Transfer threat assessment — instantiates Generalization Validation

Compares the conditions that produced a pattern against the conditions where it is meant to be used, and turns each mismatch into a scoped boundary or a demand for fresh evidence.

Version
v2 · 2026-08-28 · History
Mechanism #
3459
Type
Transfer Threat Assessment
Form family
Assessment, Review & Assurance
Solution family
Evidence, Inference & Validation
Problem family
Uncertainty, Evidence & Inference Failure
Problem subfamily
Sampling, Selection, Missingness & Generalization
Origin domain
Statistics & Experimental Design
Also from
Medicine & Healthcare
Instantiates
Generalization Validation

External Validity Check works before — or alongside — any fresh test, by reasoning about conditions rather than scoring cases. It lays the origin context of a pattern (who, where, when, under what mechanism it worked) beside the intended target context and asks, mismatch by mismatch, whether the causal story that made the pattern work in the first place is still present in the second. Its distinctive contribution is that it produces a map of transfer threats and an explicit boundary statement — "this should hold here, is doubtful there, and almost certainly breaks under these conditions" — rather than a pass/fail number. It is the mechanism that decides where fresh validation even needs to be aimed, and it can down-scope a claim on the strength of an argument alone when a target condition plainly violates the pattern's operating assumptions.

Example

A development team has strong evidence that a microfinance program raised small-business incomes in rural Bangladesh, and a ministry in a different country wants to adopt it nationally. Before spending on a pilot, the team runs an External Validity Check. They line up the conditions that made the original result work — dense village lending networks that enforced repayment socially, a population of would-be entrepreneurs already running informal stalls, scarce competing credit — against the target: a more urban population, weaker peer-enforcement norms, and banks already offering small loans. Each contrast is a transfer threat. The social-collateral mechanism, which did much of the original work, is largely absent in the target. The check does not fabricate a new result; it argues, from the mechanism, that the effect is likely to be smaller and concentrated in the more rural target districts. The output is a boundary statement: "promising in rural districts with existing peer-lending groups; unproven and mechanistically doubtful in urban areas with bank competition." That statement narrows the initial national ambition and tells the eventual pilot exactly which districts to test.

How it works

  • Reconstruct the origin mechanism. State not just that the pattern worked but why — the conditions and causal pathway doing the work.
  • Specify the target. Name the population, setting, period, and use case where the pattern is now meant to apply, concretely enough that mismatches become visible.
  • Enumerate the contrasts. For each dimension where origin and target differ, ask whether the difference threatens the mechanism, is neutral, or even helps.
  • Convert threats into boundaries or evidence demands. A benign mismatch relaxes; a mechanism-breaking one either narrows the claim outright or becomes a targeted question for a later test.

The discipline it formalizes is the classical worry about external validity — whether a result generalizes beyond the exact conditions that produced it.[1]

Tuning parameters

  • Dimension set — how many conditions (population, incentives, infrastructure, norms, scale, time) are examined. More dimensions catch subtler threats but risk a checklist that no one acts on.
  • Mechanism specificity — how sharply the origin's causal pathway is stated. A vague "it worked because it's a good program" cannot surface which mismatches matter; a crisp mechanism can.
  • Threat-to-action rule — whether a flagged threat triggers immediate down-scoping, a demand for evidence, or merely a note. Too lax and the check becomes theater; too strict and every difference blocks reuse.
  • Boundary granularity — how finely the resulting scope is carved (whole target vs. sub-segments). Finer boundaries transfer more of the value but demand more reasoning.

When it helps, and when it misleads

Its strength is cheapness and foresight: it can kill a doomed transfer, or aim an expensive pilot at exactly the doubtful segment, using structured reasoning before a single new case is collected. It is the mechanism that keeps validation relevant — pointed at the conditions that actually threaten the claim.

Its failure mode is staying a checklist. A list of noted differences that never changes the plan is validation theater: it launders a decision already made. It is also only as good as the origin mechanism it assumes; if the team is wrong about why the pattern worked, they will judge the wrong mismatches as important. And reasoning cannot substitute for a test — a plausible transfer argument is a hypothesis, not evidence. The guarding discipline is to require every serious threat to change something (the boundary, the target, or the next test to run) and to treat the check as scoping fresh validation, never replacing it.

How it implements the components

  • generalization_target — specifying the target context concretely is the check's first act; the whole assessment is a comparison against it.
  • transfer_boundary_statement — its primary output is exactly the "holds here, doubtful there, breaks under these conditions" boundary map.
  • scope_revision — a mechanism-breaking mismatch lets the check narrow the claim by argument, before any test, when the origin conditions are plainly absent.

It does not measure a pattern's stability by perturbing inputs and assumptions with a challenge_case_set — that is Robustness Check, its nearest twin, which differs by empirically stressing the pattern rather than reasoning about condition mismatch a priori. Nor does it set a numeric performance_threshold or gather a validation_case_set — those independent-evidence steps belong to Pilot Replication and Train/Test Split.

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: External Validity Check operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it compares the conditions that produced a pattern against the conditions where it is meant to be used, and turns each mismatch into a scoped boundary or a demand for fresh evidence.

Independent corroboration: The frozen evidence defines External Validity Check as 'Compares the conditions that produced a pattern against the conditions where it is meant to be used, and turns each mismatch into a scoped boundary or a demand for fresh evidence', so its operative form is Assessment, Review & Assurance.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Statistics & Experimental Design

Origin pattern: Single lineage

Present-day reach: Universal

Rationale: External validity and generalization across populations, settings, treatments, and times are canonical concerns of experimental design.

Related originating lineages:

  • Medicine & Healthcare — Clinical applicability and population-context comparison materially developed practical transfer-threat assessment.

Review resolution: Both reviewers agree that statistics_experimental_design is primary. I retain medicine_healthcare only as formative origin lineages; single_lineage is appropriate because the alternate domains informed practice without constituting independent ownership. Reach is universal because the structure is portable across essentially any domain with the stated problem, an applicability judgment kept separate from provenance. Encyclopedia synthesis is false because the artifact is already established enough that encyclopedia-specific synthesis is not required. No unresolved historical ambiguity remains after reconciling the secondary fields.

Review outcome: Reconciled after independent review; high confidence.

Notes

The check is most valuable as the first mechanism in a chain: run it, and the pilots, splits, and holdouts that follow can concentrate their limited budget on the segments the reasoning flagged as doubtful rather than re-confirming the segments that were never in question.

References

[1] Shadish, W. R., T. D. Cook, and D. T. Campbell. Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin (2002). Defines external validity as whether a causal relationship holds across variations in persons, settings, treatments, and outcomes. registry