Skip to content

Omitted Variable Bias

Correct for the distortion in a regression coefficient when a left-out variable both causes the outcome and correlates with an included regressor, so the estimate absorbs the omitted effect as the signable product of two relationships.

Core Idea

Omitted variable bias is the systematic distortion in OLS regression coefficients when the model excludes a variable that both determines the outcome and correlates with an included regressor. The included coefficient then absorbs part of the omitted effect: the estimate equals the true direct coefficient plus the product β₂·δ. The bias direction — inflated or attenuated — reads off the signs of those two relationships from domain knowledge alone, without data on the missing variable.

Scope of Application

Omitted variable bias lives wherever regression-based inference is drawn from observational data, bounded by that formalism since every habitat runs the same OLS machinery.

  • Econometrics — the canonical ability-bias problem in returns-to-schooling.
  • Epidemiology — the same object as confounding adjustment, with E-value sensitivity tools.
  • Education and policy evaluation — selection on unobservables in non-randomised programmes.
  • Political science and sociology — the endogeneity worry in cross-unit regressions.
  • A/B testing and ML causal inference — covariate imbalance when randomisation fails.

Clarity

Naming the bias turns "can I trust this coefficient?" into a two-pronged test: a left-out variable is harmless unless it both influences the outcome and correlates with a regressor. It separates omissions that bias the estimate from those that merely cost precision, and makes the bias direction legible from two known signs, reframing instruments and sensitivity analyses accordingly.

Manages Complexity

An unbounded worry about infinitely many unmeasured variables collapses to a two-condition gate, partitioning omissions into biasing, precision-costing, and irrelevant. Once a variable clears the gate, a regression's worth of worry reduces to tracking two sign bits and reading their product β₂·δ, which also organizes the remedy space around the same small parameter set.

Abstract Reasoning

The concept licenses signing the bias from domain knowledge with no data on the missing variable, a two-condition boundary test partitioning every omission, an interventionist move that reads the remedy off which term of β₂·δ it neutralizes, and a sensitivity move that bounds how strong a confounder must be to overturn the conclusion.

Knowledge Transfer

Within regression-based inference the concept transfers literally as mechanism across econometrics, epidemiology, education, political science, and ML causal inference, because the shared formalism carries the β₂·δ decomposition, the sign diagnostic, and the remedy menu untranslated. Beyond regression the structure recurs in other machinery, but that cross-domain weight belongs to the parent prime confounding (with hidden_variable and selection_bias nearby), of which OVB is the regression rendering.

Relationships to Other Abstractions

Local relationship map for Omitted Variable BiasParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Omitted Variable BiasDOMAINDomain-specific abstraction: Endogeneity — presupposesEndogeneityDOMAINPrime abstraction: Confounding — is a decomposition of, conditionalConfoundingPRIME

Current abstraction Omitted Variable Bias Domain-specific

Parents (2) — more general patterns this builds on

  • Omitted Variable Bias presupposes Endogeneity Domain-specific

    Omitted-variable bias presupposes endogeneity because its two-condition gate makes an included regressor correlated with the regression error.

  • Omitted Variable Bias is a decomposition of, conditional Confounding Prime

    Omitted-variable bias exposes confounding when the omitted determinant is a pre-treatment common cause of the included regressor and outcome; correlation arising otherwise is excluded.

Hierarchy paths (17) — routes to 10 parentless roots

Neighborhood in Abstraction Space

Omitted Variable Bias sits in a sparse region of the domain-specific corpus (97th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (309 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12