Skip to content

Internal validity

The property of an empirical study that warrants its causal claim within its own sample and setting — whether the observed intervention-outcome association is genuinely produced by the intervention rather than by confounders, selection, or bias — established by ruling out a closed catalog of named threats.

Core Idea

Internal validity is the property of an empirical study that warrants the causal conclusion it draws within its own studied sample and setting — the degree to which the observed association between intervention and outcome is genuinely produced by the intervention rather than by confounders, selection artifacts, measurement biases, procedural contamination, or chance. The question it answers is strictly local: given the design we ran, on the participants we ran it on, did we actually establish that the intervention caused the change in outcome we are reporting? It is analytically separable from, and logically prior to, external validity — the question of whether that causal claim generalizes beyond the studied sample and setting — because a causal claim that is internally invalid cannot be made more credible by widening generalization.

The operative framework for internal validity, developed by Campbell and Stanley (1963) and elaborated by Cook and Campbell (1979), is a finite enumerable catalog of named threats: history (events co-occurring with the intervention that could account for the change), maturation (natural development of participants over time), testing effects (practice or sensitization from repeated measurement), instrumentation (changes in measurement tools or raters), regression to the mean (extreme scorers returning toward average on re-test), selection (pre-existing differences between treatment and control groups), attrition (differential dropout that leaves non-comparable groups), and several interaction effects among these. Each threat names a specific mechanism by which apparent intervention effects can be spurious. A study's internal validity is established by systematically ruling out each threat — either by design controls (randomization blocks selection; blinding blocks observer bias; pre-registration blocks outcome switching) or by post-hoc analysis that demonstrates the threat is implausible on the data. The threat catalog reduces the unlimited universe of possible spurious-effect stories to a tractable, checkable list, converting vague skepticism into a structured audit with specific design responses for each threat.

Structural Signature

Sig role-phrases:

  • the local causal claim — the assertion that the intervention produced the observed change within this sample and setting, the thing to be warranted
  • the threat catalog — the closed, enumerable list of mechanism-categories by which an apparent effect could be spurious: history, maturation, testing, instrumentation, regression to the mean, selection, attrition, and their interactions
  • the design controls — the structural features (randomization, blinding, blocking, pre-registration) that block specific threats by construction
  • the analysis controls — the statistical techniques (matching, adjustment) that block threats post hoc, contingent on the analyst having measured the right covariates
  • the three-valued per-threat status — the audit move assigning each threat one of: blocked by design, checkable post hoc, or residual concern
  • the design-beats-analysis ranking — the security ordering whereby a design-blocked threat earns a firmer warrant than an analytically-blocked one
  • the warrant read-out — the overall verdict summed across the catalog: all design-blocked is strongest, analytically-handled is weaker, unblocked threats remain as explicit discounts on the headline number
  • the priority-over-external-validity ordering — the constraint that the audit runs first, since an internally invalid claim cannot be rescued by generalizing it more widely

What It Is Not

  • Not external validity. Internal validity is the strictly local question — did the intervention cause the change in this sample and setting — separable from and logically prior to whether the effect generalizes. The two vary independently in any combination, and a claim that is internally invalid cannot be made credible by widening generalization; conflating them is what leads an analyst to "fix" a broken causal warrant by recruiting a more diverse sample.
  • Not fixable by a bigger or more diverse sample. Because the threats are upstream of generalization, an internally invalid causal claim is not rescued by running the study on more people or more varied populations — that addresses external validity, the wrong problem. The remedy is to repair the design (block the operating threat), not to widen the reach; a confidently misestimated effect generalized broadly is more harmful, not less.
  • Not measurement or construct validity. Those concern whether an operational measure captures the construct it purports to. Internal validity concerns the causal-inference warrant of the design — whether the observed association reflects the intervention rather than a spurious mechanism. A study can measure its construct impeccably and still be internally invalid, and vice versa.
  • Not a global "is this study any good?" verdict. Internal validity is one specific property — the local causal warrant — not overall study quality, which also turns on external validity, statistical-conclusion validity, construct validity, and measurement. It answers exactly one question, and lumping it with the others loses the separability that lets each be assessed and remedied on its own.
  • Not reproducibility or replicability. Reproducibility is other analysts getting the same answer from the same data; replicability is other studies getting the same answer from new data. Internal validity is about whether a single study's own design warrants its own causal claim — a property that holds or fails before any question of repeating it arises.
  • Not synonymous with randomization, nor with confounding control alone. Randomization is the firmest way to block selection and several other threats, but it is one technique, not the property itself; quasi-experimental and analytic controls can establish internal validity less securely. And confounding is only one threat on a closed catalog (history, maturation, testing, instrumentation, regression to the mean, selection, attrition); ruling it out is necessary but not sufficient for the warrant.

Scope of Application

Internal validity lives across the empirical-research fields that share the substrate of controlled inquiry on quantitatively measured outcomes; its reach is within that one methodological domain, the Campbell-Stanley / Cook-Campbell apparatus restaging identically across them. The thin underlying problem (genuine cause versus spurious co-occurrence) recurs in engineering, intelligence, and history, but each has its own substrate-specific threat list — the named catalog itself does not travel, and the portable lesson belongs to causal_inference and validation.

  • Experimental psychology and education research — the textbook home: the Cook-Campbell threat taxonomy underwrites research-methods training and quasi-experimental evaluation against the threat list.
  • Clinical trials — randomization, blinding, intention-to-treat, and pre-specified outcomes are each techniques for blocking specific named threats.
  • Program and policy evaluation — difference-in-differences, regression discontinuity, synthetic controls, and matching prioritise internal validity in observational settings.
  • Software A/B testing and product experimentation — sample-ratio mismatch, novelty effects, network interference, and funnel survivorship are the threat catalog reformulated for digital experiments.
  • Marketing and pricing experimentation — holdout groups, geo-experiments, and synthetic control centre internal validity in commercial decision-making.

Clarity

Internal validity separates two questions that practitioners routinely fuse into a single sense of "is this study any good?": did the intervention cause the observed change in this sample? and will the same effect appear in another sample, setting, or time? Naming the first as internal validity and the second as external validity makes them independently assessable, and reveals that they can vary in any combination — a tightly controlled laboratory experiment on undergraduates can be high internal and low external; a sprawling naturalistic field study can be the reverse. The sharper move the distinction licenses is ordering: because a causal claim that is internally invalid cannot be rescued by generalizing it more widely, internal validity is logically prior, and an analyst who confuses the two may try to fix a broken causal warrant by recruiting a more diverse sample — exactly the wrong remedy.

Its second clarifying service is to convert diffuse skepticism into a structured audit. Before the threat catalog, an objection to a causal claim is either accepted on faith or rejected by unargued doubt; the named threats — history, maturation, testing, instrumentation, regression to the mean, selection, attrition — give a reviewer a finite, checkable list, each naming a specific mechanism by which an apparent effect could be spurious and each carrying a standard design or analytic response. The question a researcher can now ask of any design is not the unbounded "could anything else explain this?" but the tractable "for each threat on the list, is it blocked by design, checkable post hoc, or a residual discount on the headline number?" The catalog also sharpens a distinction within the remedies: a threat blocked by design (randomization severing selection) is more secure than one blocked by analysis (matching or adjustment), because the latter rests on the analyst having measured the right covariates — a difference the framework makes explicit rather than leaving to intuition.

Manages Complexity

The objection that hangs over every causal claim from an empirical study — "maybe something other than the intervention produced this result" — is, taken literally, infinite. The space of alternative stories that could explain an observed association is unbounded: any co-occurring event, any pre-existing difference between groups, any drift in the instrument, any quirk of who stayed and who dropped out could in principle be the real cause, and a researcher facing that space with no structure can neither enumerate the objections nor know when enough have been answered. Internal validity compresses that unbounded space onto a finite, enumerable catalog. The Campbell-Stanley / Cook-Campbell threats — history, maturation, testing, instrumentation, regression to the mean, selection, attrition, and the named interactions among them — are not a list of every possible confounder by name, which would be impossible, but a closed list of the categories of mechanism by which a spurious effect can arise. The unlimited "could anything else explain this?" becomes the tractable "for each of these named mechanisms, is it operating here?" A sprawling, indefinable skepticism collapses to a checklist of fixed length.

What the analyst tracks is then one status per threat, and the verdict on the study reads off those statuses through a fixed branch structure. For each named threat the question is exactly three-valued: is it blocked by design, checkable post hoc, or a residual concern? — and the catalog attaches to each threat a standard set of design responses, so the classification is routine rather than invented anew for each study (randomization blocks selection; blinding blocks observer bias; pre-registration blocks outcome switching). The branches differ in security in a way the framework makes explicit: a threat blocked by design is firmer than one blocked by analysis, because the analytic remedy rests on the analyst having measured the right covariates while the design remedy does not — so the same threat resolves to a stronger or weaker leg of the branch depending on how it was handled. Summing across the catalog yields the study's overall internal-validity verdict: all threats blocked by design is the strongest causal warrant, threats handled only analytically is weaker, unblocked threats remain as explicit discounts on the headline number. And the framework fixes an ordering that keeps the compression from being misapplied — because an internally invalid claim cannot be rescued by generalizing it more widely, this audit is logically prior to any question of external validity, so the analyst runs the threat checklist first and only asks about generalization once the local causal warrant survives it. A high-dimensional "imagine every way this could be spurious" problem becomes a low-dimensional "walk a fixed list, assign each threat one of three statuses, read off the warrant" problem.

Abstract Reasoning

The defining move is a catalog-driven audit — reasoning from an observed intervention-outcome association to the closed list of named mechanisms that could have produced it spuriously, taken one at a time. Rather than entertain the unbounded "could anything else explain this?", the analyst walks the Campbell-Stanley / Cook-Campbell threats — history, maturation, testing, instrumentation, regression to the mean, selection, attrition, and their interactions — and for each asks whether it is operating here. The inference runs from a feature of the design (a single pre-post group; voluntary enrollment; extreme-scorer recruitment; repeated testing) to the specific threat it activates: extreme scorers selected at baseline → regression to the mean will mimic improvement; voluntary participation → selection may have stacked the groups before treatment; differential dropout → attrition may have left non-comparable arms. The move converts diffuse skepticism into a finite checklist where the analyst can know when enough objections have been answered, because the list has fixed length.

Each threat then drives a three-valued classification with a built-in security ranking — the move that turns the audit into a verdict. For every named threat the analyst assigns one of three statuses (blocked by design, checkable post hoc, or residual concern) and reads the study's overall causal warrant off the sum: all threats blocked by design is the strongest warrant; threats handled only analytically is weaker; unblocked threats remain as explicit discounts on the headline number. Crucially the move distinguishes blocked by design from blocked by analysis and reasons about which is firmer: randomization severs selection without relying on anything the analyst measured, whereas matching or covariate adjustment severs it only if the analyst measured the right covariates — so the same threat resolves to a stronger or weaker leg depending on how it was handled, and the inference runs from the mechanism of control to how much trust the resulting warrant earns.

A foundational ordering move governs when the audit runs at all: because an internally invalid causal claim cannot be rescued by generalizing it more widely, internal validity is logically prior to external validity, so the analyst reasons that the threat checklist must be cleared first and questions of generalization deferred until the local causal warrant survives. The characteristic error this move prevents is reaching for a more diverse sample to repair a broken causal claim — exactly the wrong remedy, since the threats are upstream of generalization. The inference is: this study's local warrant is in doubt → widening the population cannot help → fix the design, not the reach.

The interventionist face of the framework runs at design time and predicts effects on the warrant before the study is run. Reasoning forward from a contemplated design choice to the threats it would neutralize, the analyst predicts that randomizing assignment will block selection and attrition by construction, that blinding will block observer bias, that pre-registration will block outcome switching — and chooses the design that converts the most threats from residual concerns into design-blocked status. The inference runs from "if I add this control, which threats does it sever?" to a design whose causal warrant is secured in advance rather than defended after the fact — and where full design control is infeasible, the same reasoning ranks the fallbacks, preferring a discontinuity-based control that can be substantively defended over an analytic adjustment that rests on covariate completeness.

Knowledge Transfer

Within quantitative empirical research internal validity transfers as full apparatus, and what carries is the whole Campbell-Stanley / Cook-Campbell machinery: the internal/external separation and its priority ordering, the closed threat catalog (history, maturation, testing, instrumentation, regression to the mean, selection, attrition, and their interactions), the three-valued per-threat status (design-blocked / checkable / residual), and the design-beats-analysis security ranking. The framework restages identically across fields that share the substrate of controlled inquiry on quantitatively measured outcomes: experimental psychology and education research (its textbook home, where the threat list underwrites methods training and quasi-experimental evaluation), clinical trials (where randomization, blinding, intention-to-treat, and pre-specified outcomes are each techniques for blocking specific threats), program and policy evaluation (where difference-in-differences, regression discontinuity, synthetic controls, and matching prioritise internal validity in observational settings), software A/B testing (where sample-ratio mismatch, novelty effects, network interference, and survivorship in funnel metrics are the threat catalog reformulated for digital experiments), and marketing and pricing experimentation (holdout groups, geo-experiments, synthetic control). Across all of these the threat names, the design responses, and the audit procedure are the same objects; only the experimental setting changes. The transfer is literal because the substrate — an empirical study estimating a causal effect within its own sample — is held fixed.

Beyond the empirical-research substrate, the honest report is that the named apparatus does not travel, even though the underlying problem recurs. The substrate-independent insight is thin — apparent effects can be spurious; enumerate and rule out the alternative explanations — and it does appear elsewhere: engineering root-cause analysis, intelligence assessment, and historical causation analysis all face the genuine-cause-versus-spurious-co-occurrence problem. But those domains do not inherit the Cook-Campbell catalog; each has evolved its own substrate-specific threat list — fault-tree analysis in engineering, structured analytic techniques and cognitive-bias checks in intelligence, counterfactual reasoning in history — addressing the same general problem with different machinery. So the cross-domain appearance of "internal validity" is reinstantiation by analogy at the level of spirit, not the transfer of the apparatus as mechanism: invoking "internal-validity threats" for a historical argument borrows the audit posture while the history, maturation, and regression-to-the-mean threats and their randomization/blinding remedies stay behind, because there is no sample, no control arm, and no repeated measurement for them to attach to.

Where a genuine cross-domain lesson is wanted, it should be carried by the general primes internal validity operationalises, not by the validity-typology terminology: causal_inference and intervention (the recognition that severing or controlling a treatment's incoming influences is what licenses a causal claim), validation (the broader confirm-it-actually-works criterion, of which internal validity is the causal-warrant special case), and the falsificationist discipline of ruling out alternative explanations. The portable insight — do not credit a causal claim until the named ways it could be spurious have been blocked, and fix the design rather than widening the sample when the local warrant is in doubt — belongs to those parents. The Cook-Campbell threat catalog itself is best surfaced as an archetype-level checklist within research design: an excellent domain-specific instrument whose threat names and design responses are portable across quantitative-empirical fields but whose force does not extend past that substrate (see Structural Core vs. Domain Accent).

Examples

Canonical

The framework's textbook demonstration is the one-group pretest–posttest design and its vulnerability, laid out by Campbell and Stanley (1963). Suppose a school screens students, enrolls the 50 who scored lowest on a reading test into a tutoring program, and re-tests them afterward; their average rises, and the school credits the tutoring. Internal validity asks whether the tutoring caused the rise, and the threat catalog immediately flags candidates the enthusiastic reading misses. Regression to the mean: extreme low scorers, selected precisely for their extremity, drift upward on re-test even with no treatment. Maturation: the children simply grew as readers. History: a co-occurring change, such as a new curriculum, could account for it. Testing: practice on the first test lifts the second. With no control group, none of these is blocked, so the design cannot warrant the causal claim.

Mapped back: "Tutoring raised these scores" is the local causal claim. Regression to the mean, maturation, history, and testing are entries in the threat catalog that the design leaves live. The absence of a control group means no design controls sever them, so the warrant read-out is at its weakest — the threats remain as unblocked concerns, not discounts on a credible number but reasons the number is uninterpretable.

Applied / In Practice

The 1954 Salk polio vaccine field trial, directed by Thomas Francis Jr., is a landmark case of internal-validity reasoning shaping design. Two designs ran side by side. In the "observed control" areas, second-graders were offered the vaccine and compared with unvaccinated first- and third-graders — but because vaccination was voluntary, the vaccinated group self-selected (volunteers tended to come from higher-income families that, as it happened, carried higher paralytic-polio risk), threatening the comparison with selection bias. In the other areas, children were randomly assigned to vaccine or a saline placebo, double-blinded so that neither families nor the physicians diagnosing polio knew who received which. Randomization severed selection by construction; blinding severed observer bias in diagnosis. The randomized placebo-controlled arm therefore carried the firmer causal warrant, and its result — a clear reduction in paralytic polio — is the one history trusts.

Mapped back: "The vaccine prevents paralytic polio" is the local causal claim. The volunteer self-selection and the risk of biased diagnosis are the selection and observer-bias entries in the threat catalog. Randomization and blinding are design controls that block them by construction, whereas the observed-control comparison relied on weaker footing — a vivid instance of the design-beats-analysis ranking and the warrant read-out that made the randomized arm the trustworthy one.

Structural Tensions

T1: Internal-external priority versus real-world irrelevance (a warrant so local it may not be worth having). The framework's ordering move is that internal validity is logically prior — an internally invalid claim cannot be rescued by generalizing it — so the analyst clears the threat checklist first and defers reach. Sound as far as it goes, but priority in logic is not priority in value: a study can be maximally internally valid on a design so artificial (undergraduates in a lab, a stripped-down task) that its airtight local warrant licenses a causal claim no one has reason to care about. Maximizing internal validity can trade against the external validity that makes the finding matter, and the priority rule offers no counsel on how much local rigor to buy at what cost in relevance. The very ordering that protects the causal warrant can license a rigorous answer to a question stripped of consequence. Diagnostic: Is this study's local causal warrant purchased at a cost in realism that leaves the established effect irrelevant to any setting anyone acts in?

T2: Closed catalog versus the unlisted threat (a finite list that says "enough" can say it too soon). The catalog's whole service is converting the unbounded "could anything else explain this?" into a fixed list where the analyst knows when enough objections are answered. That closure is the compression. But the universe of spurious-effect mechanisms is not actually finite, and a threat with no entry — a novel interference, a substrate-specific artifact the 1963 taxonomy never anticipated — passes the audit invisibly precisely because the checklist reports "complete." The confidence that the list confers is confidence that it is exhaustive, and that confidence is exactly what blinds a reviewer to the confounder outside it. The closure that makes the audit terminable also makes it possible to certify a study clean against a list that omitted its actual flaw. Diagnostic: Does this design activate a spurious-effect mechanism the standard catalog has no slot for, which walking the list will therefore never surface?

T3: Design-blocked security versus its hidden assumptions (randomization severs threats but not unconditionally). The framework ranks design-blocked above analytically-blocked because randomization severs selection without relying on measured covariates, whereas adjustment works only if the right covariates were measured. Real and important. But "blocked by design" carries its own unstated conditions — randomization blocks selection only if allocation was concealed and compliance held; blinding blocks observer bias only if the blind was not broken; intention-to-treat holds only under acceptable attrition. Reading a design-blocked status as unconditionally firm can hide the fact that a broken blind or differential dropout has quietly demoted it to analytically-contingent. The security ranking that rightly prefers design over analysis can lull the analyst into treating a design control as self-guaranteeing when its own preconditions have failed. Diagnostic: Did the design control actually hold in execution — concealment, blind, compliance intact — or has a downstream failure silently converted it into an assumption-laden analytic remedy?

T4: Threat-by-threat audit versus interaction effects (a serial checklist against a joint problem). Walking the catalog one threat at a time, assigning each a three-valued status, is what makes the audit tractable and terminable. But threats interact — the catalog itself names selection-by-maturation and other joint effects — and a serial, item-by-item pass can clear each threat individually while missing the way two of them combine to produce a spurious effect neither creates alone. Attrition that is differential because of maturation, selection that biases which history a group is exposed to: these live between the checklist items, not in them. The decomposition that makes internal validity assessable one mechanism at a time can dissolve exactly the interaction structure that generates the hardest confounds. Diagnostic: Have the threats been checked only individually, or has the joint interaction of two live threats — the effect neither produces alone — been examined?

T5: Structured audit versus checklist ritualism (converting skepticism into a list can hollow the skepticism). Before the catalog, an objection was accepted on faith or rejected by unargued doubt; the named threats turn diffuse skepticism into a checkable procedure with standard responses. That is the framework's gift. But a procedure invites ritual: a researcher can recite "randomized, therefore selection blocked; blinded, therefore observer bias blocked" and produce a clean-looking audit without the substantive judgment about whether, in this particular study, the mechanism is genuinely inoperative. The catalog's routineness — the very thing that makes the classification reproducible rather than reinvented per study — is what lets it be performed as box-ticking, substituting the appearance of an audit for the adversarial thought the audit was meant to institutionalize. Diagnostic: Is each threat's status the product of genuine engagement with how it could operate here, or a reflexive citation of the design feature nominally associated with blocking it?

T6: Autonomy versus reduction (its own methodological property or the empirical-research instance of its parents). "Internal validity" is a named, canonically studied property with proprietary apparatus — the Campbell-Stanley / Cook-Campbell threat catalog, the three-valued status audit, the design-beats-analysis ranking, the priority-over-external-validity ordering — and that whole machinery transfers as mechanism across the empirical-research substrate (psychology, clinical trials, policy evaluation, A/B testing), where only the setting changes. But beyond that substrate the named apparatus does not travel: engineering root-cause, intelligence assessment, and historical causation face the same genuine-cause-versus-spurious-co-occurrence problem with their own threat lists (fault trees, structured analytic techniques, counterfactuals), inheriting none of the Cook-Campbell catalog. What actually carries cross-domain is the general parents it operationalizes — causal_inference and intervention (controlling a treatment's incoming influences licenses the causal claim), validation (the broader confirm-it-works criterion), and the falsificationist discipline of ruling out alternatives. The tension is between a standalone methodological property that repays its own catalog and the recognition that its portable cargo already belongs to causal inference and validation. Diagnostic: Resolve toward the parents when asking what carries outside quantitative-empirical research; toward "internal validity" when auditing whether a specific study's design warrants its own causal claim.

Structural–Framed Character

Internal validity sits at mixed, sharing the distinctive profile of its causal-inference-methodology siblings (the instrumental variable, inter-annotator agreement): a formally structured, evaluatively light epistemic instrument with no worldly instance, which is what caps it short of structure. On evaluative_weight it is near-structural: it names a property — whether a design warrants its causal claim — assessed by a neutral catalog audit, and while it is a quality judgment it renders no moral praise or blame, only a technical warrant read-out. On human_practice_bound it is firmly framed in the epistemic sense: there is no internal validity running in nature — it is a property of an empirical study, presupposing a design, a sample, an intervention, and measured outcomes, and it dissolves the instant that research apparatus is removed. On institutional_origin it is framed: the Campbell-Stanley / Cook-Campbell threat catalog, the three-valued status audit, and the design-beats-analysis ranking are artifacts of research methodology, though the causal-inference discipline they serve is not. On vocab_travels it is mixed-to-low: the whole apparatus restages identically across psychology, clinical trials, policy evaluation, and A/B testing (one substrate, many settings), but beyond quantitative-empirical research the named catalog does not travel — engineering, intelligence, and history each grow their own threat lists — and only the general discipline lifts. On import_vs_recognize it patterns as method-port within its substrate and mere spirit-level analogy beyond it: invoking "internal-validity threats" for a historical argument borrows the audit posture while history, maturation, and regression-to-the-mean stay behind for want of a sample and a control arm.

The portable structural skeleton is the causal-warrant discipline — do not credit an apparent intervention-outcome effect until the named ways it could be spurious have been ruled out, preferring control by design over post-hoc adjustment, and fixing the design rather than widening the sample when the local warrant is in doubt. That skeleton is what internal validity operationalizes and instantiates from its parent primes — causal_inference and intervention (controlling a treatment's incoming influences is what licenses the claim) and validation (the broader confirm-it-actually-works criterion, of which internal validity is the causal-warrant special case), sharpened by the falsificationist rule-out-alternatives posture — and it is those parents that carry any cross-domain lesson; the domain-accented specifics (the closed Cook-Campbell threat catalog, the three-valued per-threat audit, the design-beats-analysis security ranking, the priority-over-external-validity ordering) stay within empirical research and do not lift. Its character: an evaluatively light methodological property with no instance in the world — an epistemic audit constituted by empirical-research practice — structural only in the causal-inference-and-validation discipline it instantiates, its named threat catalog staying inside quantitative research.

Structural Core vs. Domain Accent

This section decides why internal validity is a domain-specific abstraction and not a prime — a case, like its causal-inference-methodology siblings, where a portable causal-warrant discipline sits under a closed, tradition-specific threat catalog that is the property's whole distinctive content.

What is skeletal (could lift toward a cross-domain prime). Strip the research apparatus and a portable discipline survives: do not credit an apparent cause-effect association until the named ways it could be spurious have been ruled out; prefer control by design over post-hoc adjustment; and when the local warrant is in doubt, repair the design rather than widen the reach. The portable pieces are abstract — a candidate causal claim, an enumeration of alternative explanations, and a rule-out procedure that prefers severing a confound by construction to adjusting for it afterward. This discipline is genuinely substrate-portable, recurring wherever genuine cause must be told from spurious co-occurrence — and it is what the entry names as its parents: causal_inference and intervention (controlling a treatment's incoming influences is what licenses the claim), validation (of which internal validity is the causal-warrant special case), sharpened by the falsificationist rule-out-alternatives posture. But this discipline is the core internal validity shares with every rule-out method, not what makes the property distinctive.

What is domain-bound. Everything with proprietary content is research-methodology furniture that does not survive extraction. The closed Campbell-Stanley / Cook-Campbell threat catalog (history, maturation, testing, instrumentation, regression to the mean, selection, attrition, and their interactions); the three-valued per-threat status audit (design-blocked / checkable / residual); the design-beats-analysis security ranking; and the priority-over-external-validity ordering all presuppose an empirical study with a design, a sample, an intervention, and measured outcomes. The decisive test: there is no internal validity running in the world to recognize — remove the research apparatus and the property has nothing to attach to. And the named catalog does not even travel across the general rule-out problem: engineering root-cause, intelligence assessment, and historical causation face the same genuine-cause-versus-spurious problem but grow their own threat lists (fault trees, structured analytic techniques, counterfactuals), inheriting none of the Cook-Campbell threats — because there is no sample, no control arm, and no repeated measurement for history, maturation, and regression-to-the-mean to attach to.

Why this does not clear the prime bar. A prime's vocabulary travels and its transfer is recognition of the same mechanism, not analogy. Internal validity's transfer is bimodal. Within quantitative empirical research the whole apparatus restages identically — the internal/external separation, the threat catalog, the three-valued audit, and the design-beats-analysis ranking carry intact across psychology, clinical trials, policy evaluation, A/B testing, and marketing experimentation, because the substrate (a study estimating a causal effect within its own sample) is held fixed. Beyond that substrate the named apparatus does not travel: invoking "internal-validity threats" for a historical argument borrows the audit posture while the specific threats and their randomization/blinding remedies stay behind — reinstantiation by analogy at the level of spirit, not mechanism transfer. That is the prime-bar verdict: when a genuine cross-domain lesson is wanted, it is already carried, in more general form, by the parents the property operationalizes — causal_inference, intervention, and validation. The cross-domain reach belongs to those parents; "internal validity," as named, carries the Cook-Campbell catalog, the status audit, and the security ranking that stay within empirical research — which is exactly what places it as a domain-specific abstraction (best surfaced as an archetype-level checklist within research design) rather than a prime.

Relationships to Other Abstractions

Local relationship map for Internal validityParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Internal validityDOMAINDomain-specific abstraction: Causal Inference — presupposesCausal InferenceDOMAINPrime abstraction: Experimental Design — presupposes, typicalExperimentalDesignPRIMEPrime abstraction: Intervention — presupposes, typicalInterventionPRIMEDomain-specific abstraction: External Validity — presupposesExternalValidityDOMAIN

Current abstraction Internal validity Domain-specific

Parents (3) — more general patterns this builds on

  • Internal validity presupposes Causal Inference Domain-specific

    Internal Validity presupposes a Causal Inference whose in-setting warrant it audits against confounding, selection, history, attrition, and other threats.

  • Internal validity presupposes, typical Experimental Design Prime

    Internal validity typically presupposes experimental design because its strongest threat controls are built into assignment and measurement before analysis.

  • Internal validity presupposes, typical Intervention Prime

    Internal validity typically audits whether an externally assigned treatment produced the outcome after its normal incoming causes were controlled.

Children (1) — more specific cases that build on this

  • External Validity Domain-specific presupposes Internal validity

    A positive transport warrant presupposes that the in-sample causal effect being transported is locally warranted rather than an artifact.

Hierarchy paths (9) — routes to 8 parentless roots

Not to Be Confused With

  • External validity. The paired concept — whether a causal claim generalizes beyond the studied sample, setting, and time. Internal validity is the strictly local question (did the intervention cause the change here) and is logically prior: an internally invalid claim cannot be rescued by generalizing it more widely. The two vary independently in any combination. Tell: is the doubt about whether the effect is real in this study (internal) or whether it carries to other populations/settings (external) — and note that widening the sample fixes only the latter?

  • Construct / measurement validity. Whether an operational measure actually captures the construct it purports to (does this scale measure "anxiety"?). Internal validity concerns the causal-inference warrant of the design, not the fidelity of the measures. A study can measure its construct impeccably and still be internally invalid, and vice versa. Tell: is the concern whether the instrument captures the intended construct (construct/measurement) or whether the observed association reflects the intervention rather than a spurious mechanism (internal)?

  • Statistical-conclusion validity. The sibling validity type concerning whether the statistical inference itself is sound — adequate power, correct test assumptions, no p-hacking — i.e. whether there is a real covariation to explain. Internal validity presupposes a real association and asks whether the intervention caused it rather than a confound. Tell: is the doubt about whether the effect is statistically real (statistical-conclusion) or, granting it is real, whether it is causally attributable to the intervention (internal)?

  • Reproducibility / replicability. Reproducibility is other analysts getting the same answer from the same data; replicability is other studies getting the same answer from new data. Internal validity is whether a single study's own design warrants its own causal claim — a property that holds or fails before any question of repeating arises. Tell: is the concern repeating the result (reproducibility/replicability) or the standalone causal warrant of one study (internal)?

  • Randomization (a technique, not the property). Randomization is the firmest way to block selection and several other threats, but it is one design tool, not the property itself — quasi-experimental and analytic controls can establish internal validity less securely. Likewise, confounding control is ruling out one threat on a closed catalog, necessary but not sufficient. Tell: is the referent a method for blocking a threat (randomization, confounding control) or the overall warrant summed across the whole threat catalog (internal validity)?

  • Causal inference / validation (the parents). The substrate-neutral primes internal validity operationalizes — that controlling a treatment's incoming influences licenses a causal claim (causal_inference / intervention), and the broader confirm-it-actually-works criterion (validation), of which internal validity is the causal-warrant special case. These carry the cross-domain lesson (rule out the alternatives; fix the design) that the named Cook-Campbell catalog does not. Tell: outside quantitative-empirical research what travels is this parent discipline, treated more fully in a later section — the threat catalog and status audit stay home.

Neighborhood in Abstraction Space

Internal validity sits in a sparse region of the domain-specific corpus (78th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (309 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12