Skip to content

Endogeneity

The condition in which a regressor is correlated with a model's error term — through confounding, simultaneity, or measurement error — so OLS coefficients are biased and inconsistent for the causal effect, collapsing the coefficient's causal reading while leaving its predictive one intact.

Core Idea

Endogeneity is the condition in which a regressor — an explanatory variable on the right-hand side of a statistical model — is correlated with the model's error term, so that the ordinary least squares estimator (and most regression-style estimators) produce coefficient estimates that are biased and inconsistent for the causal effect of that variable on the outcome. The condition arises through three primary structural mechanisms. Omitted-variable confounding (the most common form): an unobserved variable causes both the regressor and the outcome, so variation in the regressor is partly variation in the confound, and the estimated coefficient picks up the confound's effect in addition to any causal effect of the regressor itself. Simultaneity (reverse causation or mutual determination): the outcome and the regressor jointly determine each other — price and quantity in a demand-supply system, police presence and crime rates, health and income — so the regressor is not causally upstream of the outcome but co-determined with it, and no single regression coefficient captures either causal direction cleanly. Measurement error in the regressor: when a regressor is measured with error, the estimated coefficient is attenuated toward zero (attenuation bias) because the mismeasured variable has noise that is by definition uncorrelated with the true value but correlated with the regression error.

What the three mechanisms share is a violation of exogeneity — the assumption that the regressor can be treated as if it were set by the researcher independently of everything else that affects the outcome. When exogeneity holds, the regression coefficient has a causal interpretation: it is the expected change in the outcome when the regressor is changed while holding the error constant (i.e., while holding everything else constant). When endogeneity is present, that interpretation collapses. In econometrics and causal inference, the diagnosis of endogeneity organises the entire toolkit of quasi-experimental remedies: instrumental variables (find a variable that shifts the endogenous regressor but has no direct effect on the outcome and is uncorrelated with the confound, and use that variation alone to estimate the causal effect); regression discontinuity (find a threshold in an assignment variable at which treatment changes sharply, producing local quasi-randomisation); difference-in-differences (use parallel-trends variation across groups and time to cancel both unit-level confounders and common time trends); panel fixed effects (remove time-invariant unit heterogeneity by differencing within units over time). The choice among these remedies depends on which source of endogeneity is operative and what identifying variation the data or institutional setting makes available.

Structural Signature

Sig role-phrases:

  • the model and regressor — an outcome specified as a function of right-hand-side explanatory variables plus an error term, with one regressor under causal scrutiny
  • the exogeneity assumption — the requirement that the regressor be uncorrelated with the error term, i.e. treatable as if set independently of everything else affecting the outcome
  • the exogeneity violation — the defining condition: the regressor is correlated with the error, so the assumption fails
  • the three sources — the structural mechanisms that break exogeneity: omitted-variable confounding (a shared unobserved cause), simultaneity (regressor and outcome co-determine each other), and measurement error in the regressor
  • the biased-and-inconsistent consequence — the engineered fact that under the violation OLS coefficients no longer estimate the causal effect, with bias direction set by the source (attenuation for measurement error, confound-contamination for omitted variables, co-determination for simultaneity)
  • the prediction/causation split — the boundary that an endogenous regressor can still predict in-sample while being worthless as a causal estimate
  • the remedy menu — the source-indexed quasi-experimental fixes that restore exogeneity locally: instrumental variables, regression discontinuity, difference-in-differences, panel fixed effects
  • the burden-shift — the discipline the construct imposes: the routine challenge becomes "why should this regressor be treated as exogenous?" rather than "is the coefficient significant?"

What It Is Not

  • Not the slogan "correlation is not causation." That coarse maxim says only that the two can diverge; endogeneity says exactly how and why in a given model — the regressor is correlated with the error term — and names the operative source (confounding, simultaneity, or measurement error). It is the precise mechanism the slogan gestures at, not a restatement of it.
  • Not synonymous with confounding. Omitted-variable confounding is the most common source of endogeneity, but only one of three: simultaneity (regressor and outcome co-determine each other) and measurement error in the regressor also break exogeneity without any unobserved common cause. Confounding is one way in; endogeneity is the umbrella condition (the regressor is correlated with the error), which several distinct structures can produce.
  • Not a feature of the world. Endogeneity is a property of a model-and-data pairing — a regressor is endogenous in a model when that model's exogeneity assumption fails for it. The world has feedback, common causes, and measurement error; "endogeneity" names the modeller's problem when those features violate an assumption, not a thing happening out in the system being studied.
  • Not a problem for prediction. An endogenous regressor can forecast the outcome perfectly well in-sample; what collapses is the causal reading — the coefficient as the effect of intervening on the regressor. A variable excellent for prediction and worthless for causal estimation is not a paradox; the two purposes sit on opposite sides of the exogeneity condition.
  • Not the macroeconomic "endogenous variable." In general-equilibrium and systems modelling, "endogenous" versus "exogenous" denotes the model's closure choice — which variables are determined inside the model versus taken as given — independent of any error term or estimation. That sense collides on the word but is methodologically distinct from the regression condition this entry describes.
  • Not fixable by adding controls or collecting more data. Once exogeneity fails, a bigger sample yields a more precise estimate of the wrong number (it remains inconsistent), and more covariates help only if they close the specific back-door path. The remedy is identifying variation — an instrument, a discontinuity, a parallel-trends design, a within-unit difference — selected by which source broke exogeneity, not more of the same regression.

Scope of Application

Because endogeneity is a diagnostic condition on a model-and-data pairing with an attached remedy-menu, not a causal mechanism in the world, it is not bounded by a single subject-matter domain: it applies wherever causal effects are estimated from a model whose right-hand-side variables may be entangled with the outcome. The fields below are real applications of the identical apparatus — the same exogeneity condition, the same source classification (omitted-variable confounding, simultaneity, measurement error, selection), the same remedy menu (IV/2SLS, RDD, DiD, fixed effects) and detection tests (Hausman, Sargan-Hansen) — not metaphor; the boundary is that precondition holding versus the general-equilibrium "endogenous variable" homonym (a model-closure choice), which collides on the word but is a different concept.

  • Labour economics — the textbook case: returns to schooling and the unobserved-ability confound on the Mincer equation.
  • Industrial organisation — price-quantity simultaneity in demand estimation.
  • Development economics — programme placement endogenous to baseline characteristics.
  • Macroeconomics — policy variables that respond to the economic conditions they are meant to explain.
  • Epidemiology — confounding-by-indication, selection-into-treatment, and reverse causation, now organised under target-trial emulation.
  • Political science and public policy — treatment-effect estimation with endogenous take-up; IV, RDD, and DiD policy evaluations.
  • Sociology and education — peer-effects estimation and Manski's reflection problem; school-choice and network-influence studies.
  • Finance — capital-structure, governance, and executive-compensation studies of simultaneous determination.

Clarity

Naming endogeneity gives a field a single word for the precise reason a regression coefficient can be a fine description yet a false cause. Without it, the gap between "the data show X moves with Y" and "X causes Y" is patched over by the coarse slogan that correlation is not causation; endogeneity replaces the slogan with a mechanism — the regressor is correlated with the error term — that says exactly how and why the two diverge in a given model. That sharpens the central distinction of empirical work: descriptive prediction (where an endogenous variable can still forecast the outcome perfectly well in-sample) versus causal estimation (where the same variable cannot be read as the effect of intervening on it). A coefficient that is excellent for one purpose and worthless for the other is no longer a paradox once the failure of exogeneity is named.

The construct's second clarifying act is to collect several superficially unrelated pathologies — an unobserved common cause, mutual determination of regressor and outcome, a mismeasured regressor — under one umbrella, the violation of exogeneity, while still keeping them separable enough to diagnose individually. This matters because the umbrella is what makes the toolkit legible: instrumental variables, regression discontinuity, difference-in-differences, and panel fixed effects are no longer a grab-bag of techniques but answers to a single question asked in different institutional circumstances — "what restores exogeneity here, and which source of its failure am I confronting?" The label thereby reorganises applied practice around a routine challenge it places on every causal claim: not "is the coefficient significant?" but "why should this regressor be treated as exogenous?" — a burden of proof that the analyst must now discharge explicitly, naming the threat (confounding, simultaneity, or attenuation) and the identifying variation that defeats it, rather than letting the regression's tidy output stand in for a causal warrant it has not earned.

Manages Complexity

Applied empirical work faces a double sprawl: a long list of distinct ways a regression can mislead about causation — an unobserved common cause, the outcome feeding back on the regressor, a mismeasured regressor, outcome-dependent sampling — and a parallel grab-bag of repair techniques — instrumental variables, regression discontinuity, difference-in-differences, panel fixed effects, matching, synthetic control. Endogeneity compresses both into a single hinge. On the diagnostic side, every one of those superficially unrelated pathologies is collected under one condition: the regressor is correlated with the error term, i.e. exogeneity fails — so the analyst stops carrying a catalogue of named biases and instead tracks one binary property of the regressor (is it correlated with the error?) plus, when it is, which of three sources is operative (omitted-variable confounding, simultaneity, or measurement error). On the remedy side, the same condition turns the toolkit from a grab-bag into an organised menu: each technique is an answer to one question — "what restores exogeneity here?" — selected by the diagnosed source and the identifying variation the data or institutional setting supplies. The map from one tracked condition to a structured menu is what makes the otherwise-bewildering choice of empirical strategy tractable.

The qualitative verdict on any coefficient then reads off that single condition without re-deriving the case. Exogeneity holds, and the regression coefficient carries its causal interpretation — the expected change in the outcome from moving the regressor while holding everything else constant; exogeneity fails, and the same coefficient is biased and inconsistent for the causal effect, however well it still describes or predicts in-sample. That branch is the central compression: a coefficient excellent for prediction yet worthless for causal estimation stops being a paradox once it is located on the exogeneity branch. The construct also makes the direction of the error read off the source — attenuation toward zero for measurement error, contamination by the confound's effect for omitted variables, co-determination for simultaneity — so the analyst anticipates the sign and shape of the bias rather than discovering it. Above all it reduces the entire burden of a causal claim to one explicit question the analyst must discharge — not "is the coefficient significant?" but "why should this regressor be treated as exogenous, and if it is not, which threat and which identifying variation defeats it?" A high-dimensional design space of biases and methods thus collapses to a single exogeneity condition, a three-way source classification, and a remedy menu indexed by them — from which the causal warrant, the bias direction, and the appropriate quasi-experimental design all read off.

Abstract Reasoning

Endogeneity licenses reasoning that turns on a single condition — is the regressor correlated with the error term? — and routes diagnosis, bias-prediction, and remedy off that one question. The defining boundary-drawing move separates description from causation by a mechanism rather than a slogan: a coefficient can forecast the outcome perfectly well in-sample yet be biased and inconsistent for the effect of intervening on the regressor, and the dividing line is exactly whether exogeneity holds. Reasoning FROM "is this regressor correlated with everything else that affects the outcome" TO "can this coefficient be read as a cause or only as a description" is what dissolves the apparent paradox of a variable excellent for prediction and worthless for causal estimation — they sit on opposite sides of the exogeneity branch.

A diagnostic move classifies which of three structural sources has broken exogeneity, because the source determines both the bias and the fix. Omitted-variable confounding (an unobserved common cause of regressor and outcome), simultaneity (the outcome and regressor co-determine each other — price and quantity, police and crime, health and income), or measurement error in the regressor — the reasoner identifies the operative source from the setting. Reasoning FROM the data-generating structure TO "this is confounding / simultaneity / attenuation" is the move that converts a vague worry that "correlation isn't causation" into a named, located threat.

That classification powers a predictive move on the bias itself: the reasoner anticipates the sign and shape of the error before estimating, rather than discovering it after. Measurement error predicts attenuation toward zero; omitted variables predict contamination by the confound's effect (sign and size set by the confound's correlations); simultaneity predicts co-determination that no single coefficient captures cleanly. Reasoning FROM "which source is operative" TO "which direction the estimate is wrong" lets the analyst bound the damage and sometimes sign it even when it cannot be removed.

The interventionist move maps the diagnosed source to a remedy that restores exogeneity locally, turning a grab-bag of techniques into an organized menu. Instrumental variables when a source of variation shifts the endogenous regressor but has no direct path to the outcome and is uncorrelated with the confound; regression discontinuity when a sharp threshold in an assignment variable locally randomizes treatment; difference-in-differences when parallel-trends variation across groups and time cancels unit-level confounders and common trends; panel fixed effects when the confounder is time-invariant and can be differenced out within units. Reasoning FROM "which source broke exogeneity, and what identifying variation does the setting supply" TO "which quasi-experimental design recovers the effect" is the move that selects an empirical strategy by the threat rather than by habit.

Underlying all of these is a burden-shifting move that reorganizes what an analyst must prove. The routine challenge on any causal claim becomes not "is the coefficient significant?" but "why should this regressor be treated as exogenous, and if it is not, which threat and which identifying variation defeats it?" Reasoning FROM "a regression produced a tidy coefficient" TO "that output is not a causal warrant until exogeneity is argued" is what forces the identifying assumption into the open as something to be defended rather than assumed.

Knowledge Transfer

Endogeneity is a property of a model-and-data pairing, not a causal mechanism in the world — a name for the modeller's problem when feedback, common causes, or measurement error in the data-generating process violate a model's exogeneity assumption — so "mechanism within / metaphor beyond" applies only loosely. What transfers within statistics is the diagnostic condition and its remedy-menu, literally, to any field that estimates causal effects from a model with right-hand-side variables that may be entangled with the outcome. The precondition (a regressor possibly correlated with the error term) is the same regardless of subject matter, so the apparatus restages identically across labour economics (returns to schooling and the ability confound), industrial organisation (price-quantity simultaneity in demand estimation), development (programme placement endogenous to baseline characteristics), macroeconomics (policy variables responding to economic conditions), epidemiology (confounding-by-indication, selection-into-treatment, reverse causation — now under target-trial emulation), political science and public policy (endogenous take-up in treatment-effect estimation), sociology and education (peer effects and Manski's reflection problem), and finance (capital-structure and governance studies of simultaneous determination). In each, the same exogeneity condition, the same three-to-four source classification (omitted-variable confounding, simultaneity, measurement error, selection), the same bias-direction predictions (attenuation, confound-contamination, co-determination), the same remedy menu (IV/2SLS, regression discontinuity, difference-in-differences, panel fixed effects, synthetic control, matching), and the same detection tests (Hausman, Wu-Hausman, Sargan-Hansen) carry over without translation. This is wide instrument-reach within the causal-inference substrate.

The boundary to mark has two distinct edges. First, instrument-reach versus over-reading: the burden the construct imposes — "why should this regressor be treated as exogenous?" — is exactly a guard against the over-read of taking a tidy regression coefficient as a causal warrant it has not earned; an endogenous variable can be an excellent in-sample predictor while being worthless as a causal estimate, and conflating the two (now a live issue for machine-learning models evaluated only for predictive accuracy but used causally) is the central over-read endogeneity names. Second, a genuine homonym to keep separate: in general-equilibrium macroeconomics and systems modelling, "endogenous variable" / "exogenous variable" means something methodologically different — the model's closure choice of which variables are determined inside the model versus taken as given — independent of any statistical estimation or error term. That sense collides on the word but is not the regression condition this entry describes, and importing one for the other would be a category error.

Where a genuinely cross-domain lesson is wanted, it is not unique to "endogeneity" and should be carried by the general structural primes that produce the condition rather than by the econometric label. When biologists, ecologists, or systems theorists speak of "endogenous" versus "exogenous" drivers, the structural carrier is one of those primes with "endogenous" borrowed as terminology: feedback supplies the simultaneity case (regressor and outcome co-determine each other), confounding/common_cause supplies the omitted-variable case (a shared cause opens a back-door path — the same fact Pearl's do-calculus and back-door criterion formalise), selection_bias supplies the outcome-dependent-sampling case, and measurement_error supplies the noisy-proxy/attenuation case. The portable insight — a variable you are using to explain an outcome may itself be entangled with that outcome, breaking the causal reading — belongs to those entanglement primes. What stays home-bound is everything that makes this construct econometric: the exogeneity assumption as a regression object, the biased-and-inconsistent OLS result, the IV/2SLS/RDD/DiD/FE remedy apparatus, and the specification tests built to detect or correct it (see Structural Core vs. Domain Accent).

Examples

Canonical

The textbook case is estimating the return to education — how much an extra year of schooling raises earnings. Regress log earnings on years of schooling by OLS and the coefficient is contaminated: unobserved ability plausibly raises both schooling and earnings, so the estimate mixes the causal effect of school with the effect of the ability it is proxying — omitted-variable confounding, exogeneity violated. Angrist and Krueger's 1991 study is the canonical remedy. They exploited compulsory-schooling laws: children born earlier in the year reach the legal dropout age at a lower grade, so quarter of birth shifts years of completed schooling but has no plausible direct effect on earnings and is uncorrelated with ability. Using only the schooling variation driven by quarter of birth — an instrumental-variables estimate — recovers a return to education close to the OLS figure, suggesting ability bias was smaller than feared.

Mapped back: Earnings-on-schooling is the model and regressor; treating schooling as if independently set is the exogeneity assumption. Ability driving both is the exogeneity violation via the omitted-variable member of the three sources. Quarter of birth is the instrument from the remedy menu, and that OLS still predicts earnings while misstating the causal effect is the prediction/causation split.

Applied / In Practice

A concrete field deployment is Card and Krueger's 1994 study of the minimum wage. The causal question — does raising the minimum wage cut employment? — is beset by endogeneity: states raise minimum wages under particular economic conditions that also drive employment, so a naive before-after or cross-state comparison confounds the policy with those conditions. Card and Krueger used a difference-in-differences design around a real event: in 1992 New Jersey raised its minimum wage while neighboring Pennsylvania did not. Surveying fast-food restaurants in both states before and after, they compared the change in New Jersey employment to the change in Pennsylvania, differencing out both fixed state differences and the common regional trend. The identifying variation is the policy change itself, with Pennsylvania serving as the counterfactual for what New Jersey would have done absent the increase.

Mapped back: Minimum-wage policy correlating with the economic conditions that also move employment is the exogeneity violation (a confounding/simultaneity source among the three sources). Difference-in-differences is the design chosen from the remedy menu, its identifying variation the NJ-versus-PA policy contrast. Choosing it by naming the threat and the counterfactual that defeats it, rather than reporting a raw regression, is the burden-shift in practice.

Structural Tensions

T1: Prediction versus causation (one coefficient, two incompatible readings). The same estimated coefficient can be an excellent in-sample predictor and a worthless causal estimate, and endogeneity is exactly what pries the two apart. When exogeneity fails, the regressor still carries information that forecasts the outcome — nothing about attenuation or confounding stops a variable from correlating with what you are trying to predict — but the coefficient no longer measures the effect of intervening on that variable. The tension is that the tidy regression output serves both purposes on its face while satisfying only one, so a number that is fit-for-purpose under a forecasting goal is disqualified under a causal goal with no change in the number itself. This is now acute for machine-learning models tuned solely for predictive accuracy and then read causally: the very optimization that maximizes prediction is indifferent to the exogeneity the causal reading requires. Diagnostic: Is this coefficient being used to forecast the outcome (endogeneity may be harmless) or to estimate the effect of changing the regressor (endogeneity is fatal)?

T2: The umbrella condition versus its separable sources (lump to route the toolkit, split to sign the bias). Endogeneity's power is that it collects confounding, simultaneity, and measurement error under one condition — the regressor is correlated with the error — so the analyst tracks a single binary instead of a catalogue of named biases, and the whole remedy toolkit becomes answers to one question. But the umbrella cannot be worked at the umbrella level: which remedy applies, and which direction the bias runs, depend entirely on which source broke exogeneity — attenuation toward zero for measurement error, contamination by the confound for omitted variables, co-determination for simultaneity. The tension is that the construct must simultaneously lump (to make the design space legible and the burden uniform) and split (to diagnose, sign the bias, and select the fix), so an analyst who rests at "it's endogenous" has a slogan, not a strategy, while one who only ever names the specific bias loses the unifying question that organizes the menu. Diagnostic: Has the endogeneity been resolved to its operative source (confounding, simultaneity, or measurement error), or left at the umbrella level where no remedy or bias-direction is determined?

T3: Locally restored exogeneity versus precision and external validity (the clean estimate is a narrow one). Each remedy on the menu restores exogeneity locally — using only the identifying variation the instrument, threshold, or parallel-trends design supplies. That local recovery is precisely what buys the causal warrant, and precisely what costs generality: an IV estimate is a local average treatment effect for the compliers the instrument moves, an RDD estimate holds only near the threshold, and both throw away most of the sample's variation, so they are typically far less precise than the biased OLS they replace. The tension is that defeating endogeneity means confining the estimate to a slice of variation that is credibly exogenous, and that slice is often small, atypical, and noisy — so the analyst trades a precise, general, biased number for an unbiased number that is imprecise and may not generalize beyond the compliers or the threshold. A "clean" identification can answer a narrower question than the one originally asked. Diagnostic: Does the identifying variation recover the effect for the population and margin the causal question is about, or only for a local subgroup whose estimate cannot be read as the general effect?

T4: The burden-shift versus the assumptions it imports (one untestable exogeneity traded for another). Endogeneity disciplines practice by shifting the burden from "is the coefficient significant?" to "why should this regressor be treated as exogenous?" — forcing the identifying assumption into the open. But the remedies do not verify exogeneity; they relocate the untestable assumption. IV replaces "the regressor is uncorrelated with the error" with "the instrument affects the outcome only through the regressor and is uncorrelated with the confound" — an exclusion restriction that is itself generally untestable. DiD substitutes the parallel-trends assumption; RDD substitutes continuity at the threshold. The tension is that the burden-shift makes the causal claim honest by naming its assumption, yet the named assumption is usually no more checkable from the data than the exogeneity it stands in for — so the discipline is real but the certainty is not, and a badly chosen instrument can smuggle back the very endogeneity it was meant to defeat while looking rigorous. Diagnostic: Is the remedy's identifying assumption (exclusion, parallel trends, continuity) more defensible on institutional grounds than the original exogeneity assumption, or merely less familiar?

T5: A property of the model-and-data versus a feature of the world (why more data cannot fix it). Endogeneity is not something happening in the system being studied; it is a property of a model-and-data pairing — a regressor is endogenous in a model when that model's exogeneity assumption fails for it. The world has feedback, common causes, and measurement error; endogeneity names the modeller's problem when those features violate an assumption. This framing cuts two ways. It is empowering: the fix is in the design and the identifying variation, not in changing the world, and the same data-generating world yields endogeneity or not depending on which question and closure the analyst adopts. But it also defeats the intuitive remedy — because the problem is in the model's assumption rather than in a data shortage, a bigger sample only yields a more precise estimate of the wrong number (it stays inconsistent), and adding controls helps only if they close the specific back-door path. The tension is that the trouble feels like it lives in the data yet is dissolved only by changing the model or the design. Diagnostic: Is the proposed fix collecting more of the same data or adding arbitrary controls (which cannot restore consistency), or supplying identifying variation that changes what the model can be treated as knowing?

T6: Autonomy versus reduction (an econometric condition or the entanglement primes that produce it). "Endogeneity" carries genuinely home-bound cargo — the exogeneity assumption as a regression object, the biased-and-inconsistent OLS result, the IV/2SLS/RDD/DiD/FE remedy apparatus, the Hausman and Sargan-Hansen specification tests — none of which transfers as named outside econometric estimation. What travels cross-domain is the set of general structural primes that produce the condition: feedback for the simultaneity case, confounding/common_cause for the omitted-variable case (the same back-door path Pearl's do-calculus formalises), selection_bias for outcome-dependent sampling, and measurement_error for the noisy-proxy/attenuation case. The tension is between a standalone econometric construct that earns its own toolkit and tests, and the recognition that its portable insight — a variable you are using to explain an outcome may itself be entangled with that outcome, breaking the causal reading — belongs to those entanglement primes. Compounding it, "endogenous variable" is a homonym in general-equilibrium modelling (a model-closure choice), so borrowing the econometric label outside its home risks both reducing to the wrong parent and colliding with a different concept. Diagnostic: Resolve toward the entanglement primes (feedback, confounding/common_cause, selection_bias, measurement_error) when carrying the lesson to another field; toward "endogeneity" when diagnosing a specific regression's causal warrant — and check the word is not being used in the closure-choice sense.

Structural–Framed Character

Endogeneity sits at the framed pole of the structural–framed spectrum, and it gets there by a route worth stating precisely: it is definitionally a property of a model, not of the world. On human-practice-bound it scores at the maximum, and the entry is emphatic — "Not a feature of the world," "a property of a model-and-data pairing," "endogeneity names the modeller's problem." Strip away the researcher who specifies a regression, treats a regressor as if independently set, and imposes an exogeneity assumption, and there is no endogeneity at all — only feedback, common causes, and measurement error, which are real world-features but different things. The condition is constituted by the modelling practice and dissolves the instant that practice is removed, exactly the way ad hominem dissolves once the argumentation practice is removed. On evaluative weight it reads framed in the diagnostic register: the construct names a defect of an estimator — "biased and inconsistent," a causal reading that "collapses," a coefficient "worthless" for causal estimation — so calling a regressor endogenous convicts the estimate rather than neutrally naming a mechanism. Institutional origin is pronounced: the exogeneity assumption as a regression object, the biased-and-inconsistent OLS result, the IV/2SLS/RDD/DiD/FE remedy menu, and the Hausman and Sargan-Hansen tests are all furniture of the econometrics/causal-inference tradition, not form read off nature. Vocab-travels scores low: regressor, error term, OLS, exogeneity, attenuation bias, instrument all presuppose the estimation substrate and rename or vanish off it — and "endogenous variable" even collides with a distinct general-equilibrium homonym. And import-vs-recognize patterns as import-by-analogy: the entry is explicit that outside econometrics the lesson rides on the entanglement primes, with "endogenous" borrowed as terminology.

The portable structural skeleton is a variable used to explain an outcome is itself entangled with that outcome, so the causal reading breaks — and here the case genuinely needs more than one parent, because endogeneity is an umbrella that instantiates a small family of distinct entanglement primes: feedback for the simultaneity source, confounding/common_cause for the omitted-variable source, selection_bias for outcome-dependent sampling, and measurement_error for attenuation. Those parents each recur across substrates as genuine world-mechanisms, and that is where the cross-domain reach lives — but "endogeneity," as named, does not travel: it is the model-relative diagnosis those world-mechanisms trigger when they violate a regression's exogeneity assumption, and its remedy apparatus and specification tests stay home in econometrics. Its character: a model-constituted, defect-diagnosing econometric condition, structural only in the entanglement skeleton it instantiates from the several world-level primes (feedback, confounding, selection_bias, measurement_error) that actually produce it.

Structural Core vs. Domain Accent

This section decides why endogeneity is a domain-specific abstraction and not a prime — a case sharpened by the fact that endogeneity is a model-relative diagnosis, so the sorting is between the world-level entanglement primes that produce it and the econometric estimation apparatus that names it.

What is skeletal (could lift toward a cross-domain prime). Strip the regression away and a thin relational structure survives: a variable used to explain an outcome is itself entangled with that outcome, so treating it as an independent cause misreads the relationship. But — crucially — this skeleton is not one thing; endogeneity is an umbrella over a small family of distinct entanglement structures, each of which is a genuine world-mechanism and a separate prime. Simultaneity (the explanatory variable and the outcome co-determine each other) is feedback. Omitted-variable confounding (a shared unobserved cause opens a back-door path — the fact Pearl's do-calculus formalizes) is confounding/common_cause. Outcome-dependent sampling is selection_bias. A noisy proxy standing in for the true variable is measurement_error. These parents each recur across substrates as real world-mechanisms — a biologist's or ecologist's "endogenous vs. exogenous driver" is one of them with the word borrowed as terminology — and they are where the portable insight genuinely lives. This is the core endogeneity shares with (indeed, is assembled from), not what makes the named construct distinctive.

What is domain-bound. Almost everything that makes the construct endogeneity in particular is econometric furniture that does not survive extraction — and notably, endogeneity is not even a feature of the world but a property of a model-and-data pairing. The setup is specific: an outcome regressed on right-hand-side variables plus an error term, an exogeneity assumption (the regressor uncorrelated with the error) treated as a regression object, and the engineered consequence that when it fails the OLS estimator is biased and inconsistent for the causal effect, with a source-indexed bias direction (attenuation for measurement error, confound-contamination for omitted variables, co-determination for simultaneity). The remedy apparatus is specific — the quasi-experimental menu (instrumental variables / 2SLS, regression discontinuity, difference-in-differences, panel fixed effects) that restores exogeneity locally — and so are the detection tests (Hausman, Wu-Hausman, Sargan-Hansen). These are the worked vocabulary and instruments the discipline actually operates. The decisive test: remove the model with its error term and exogeneity assumption and there is no endogeneity — only feedback, common causes, and measurement error, which are real but different things happening in the world, not the modeller's diagnosis. (A separate hazard: "endogenous variable" is a homonym in general-equilibrium modelling, denoting a model-closure choice with no error term at all — colliding on the word while being methodologically distinct.)

Why this does not clear the prime bar. A prime's vocabulary travels and its cross-domain transfer is recognition of the same mechanism, not analogy. Endogeneity's transfer is bimodal, split along an unusual seam. Within the causal-inference substrate it travels intact as a diagnostic instrument — across labour economics, industrial organization, development, macroeconomics, epidemiology, political science, sociology, and finance, the identical exogeneity condition, source classification, bias-direction predictions, remedy menu, and specification tests restage without translation, because each field supplies the same precondition (a regressor possibly correlated with the error term); this is genuine mechanism-reach, not metaphor. Beyond estimation it does not travel as "endogeneity" at all: when biologists or systems theorists want the lesson, the structural carrier is one of the world-level primes with "endogenous" borrowed as a label, and importing the econometric apparatus would both reduce to the wrong parent and risk colliding with the closure-choice homonym. And when the bare structural lesson is needed cross-domain — a variable you use to explain an outcome may itself be entangled with it, breaking the causal reading — it is already carried, in more general and world-level form, by the primes endogeneity instantiates: feedback, confounding/common_cause, selection_bias, and measurement_error. The cross-domain reach belongs to those parents; "endogeneity," as named — the exogeneity assumption as a regression object, the biased-and-inconsistent OLS result, the IV/RDD/DiD/FE apparatus, the specification tests — is econometric furniture that should stay home, which is why it clears the domain-specific bar for causal inference but not the prime bar.

Relationships to Other Abstractions

Local relationship map for EndogeneityParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.EndogeneityDOMAINDomain-specific abstraction: Regression — presupposesRegressionDOMAINDomain-specific abstraction: Attenuation Bias — presupposesAttenuation BiasDOMAINDomain-specific abstraction: Instrumental variable — presupposesInstrumentalvariableDOMAINDomain-specific abstraction: Omitted Variable Bias — presupposesOmittedVariable BiasDOMAIN

Current abstraction Endogeneity Domain-specific

Parents (1) — more general patterns this builds on

  • Endogeneity presupposes Regression Domain-specific

    Endogeneity presupposes a regression-style model whose regressor, error term, and exogeneity assumption make the violation definable.

Children (3) — more specific cases that build on this

  • Attenuation Bias Domain-specific presupposes Endogeneity

    Classical regressor-measurement attenuation presupposes the endogeneity created when the noisy observed regressor correlates with the composite error.

  • Instrumental variable Domain-specific presupposes Endogeneity

    Instrumental-variable identification presupposes an endogenous treatment whose correlation with the model error destroys the naive causal coefficient.

  • Omitted Variable Bias Domain-specific presupposes Endogeneity

    Omitted-variable bias presupposes endogeneity because its two-condition gate makes an included regressor correlated with the regression error.

Not to Be Confused With

  • Confounding / common cause. An unobserved variable that causes both the regressor and the outcome, opening a back-door path (Pearl's formalization). It is the most common source of endogeneity, but only one of three — a part, not the whole. Tell: is a shared unobserved cause the specific mechanism (confounding), or the umbrella condition "regressor correlated with the error" that confounding, simultaneity, or measurement error can each produce (endogeneity)?
  • Simultaneity / reverse causation (feedback). The regressor and outcome jointly determine each other (price and quantity, health and income). Another source of endogeneity, carried world-side by feedback. Tell: do the two variables co-determine each other (simultaneity), or is there a distinct upstream shared cause (confounding) — either of which makes the regressor endogenous?
  • Measurement error / attenuation bias. A regressor observed with noise, biasing its coefficient toward zero. A third source of endogeneity — distinct from confounding and simultaneity, and requiring no unobserved common cause. Tell: is the regressor mismeasured relative to its true value (measurement error), or correctly measured but entangled with the outcome by another route (the other sources)?
  • "Correlation is not causation" (the slogan). The coarse maxim that association and cause can diverge. Endogeneity is the precise mechanism the slogan gestures at — the regressor is correlated with the error term — plus a named source and a matched remedy. Tell: is it only the observation that the two can differ (slogan), or a specification of how and why in a given model with a fix indexed to the source (endogeneity)?
  • Endogenous variable (general-equilibrium homonym). In systems and macro modeling, "endogenous" vs. "exogenous" denotes a model-closure choice — which variables are solved inside the model versus taken as given — with no error term and no estimation involved. It collides on the word but is a different concept. Tell: is the term about which variables the model determines internally (closure choice), or about a regressor correlated with a regression's error term (endogeneity)?
  • Selection bias. Outcome-dependent sampling that distorts the estimated relationship. A near-sibling that some treatments fold into endogeneity's sources and that the world-level prime selection_bias carries. Tell: does the distortion arise from how the sample was selected on the outcome (selection bias), or from the regressor's entanglement with the error in a full-sample model (the other endogeneity sources)?
  • The entanglement parents (umbrella). The world-level primes endogeneity instantiates — feedback, confounding/common_cause, selection_bias, measurement_error — which carry the portable insight ("a variable used to explain an outcome may itself be entangled with it") across domains. Endogeneity is the model-relative diagnosis those world-mechanisms trigger in a regression. Tell: is the phenomenon a real mechanism in the system (these parents), or the modeller's diagnosis when it violates a regression's exogeneity assumption (endogeneity)? (Treated more fully in Structural Core vs. Domain Accent.)

Neighborhood in Abstraction Space

Endogeneity sits in a sparse region of the domain-specific corpus (82nd percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (309 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12