Skip to content

Truncated Regression Model

Infer a population response–covariate relation from records admitted only when the response falls inside a known region, by conditioning on that inclusion.

Version
v2 · 2026-10-03 · History
Domain-specific #
13678
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomains
Regression Models, Sample Selection → Experimental Design & Statistics
Aliases
Truncated regression, Outcome-truncated regression

Core Idea

A truncated regression model addresses a sample in which a unit is recorded only when its outcome \(Y\) falls inside a declared region. The goal is to infer a population relationship between \(Y\) and covariates \(X\), while the analyst actually observes selected pairs \((x,y)\) with \(y\in A\). For a model density \(f(y\mid x;\theta)\) and inclusion region \(A\), the density of an observed outcome is

\[ f(y\mid x,Y\in A;\theta) =\frac{f(y\mid x;\theta)} {P_\theta(Y\in A\mid X=x)},\qquad y\in A. \]

The denominator is not an arbitrary correction constant: it is the model-implied probability of admission for that covariate value. In the common normal linear case, the Stata reference manual writes this likelihood explicitly for lower and upper truncation; maximizing it uses the selected sample while accounting for the missing outcome tails under its assumptions.[1][2]

Outcome-dependent admission generally alters the regression-error distribution among included cases. Fitting ordinary least squares only to them can therefore target the selected sample rather than the intended population coefficients. This is a general risk, not a claim that every special parameter configuration yields bias. Correct likelihood conditioning also does not repair a wrong population density, an unknown threshold or selection on additional unmodeled variables.[1]

Structural Signature

  1. Population relation: a response \(Y\), covariates \(X\) and target parameter \(\theta\) describe the relation before selection.
  2. Known inclusion rule: a unit is in the analyzed sample only if \(Y\in A\), for example \(a<Y<b\). Bounds can be one-sided or covariate-dependent if explicitly modeled.[1]
  3. Selected observations: the standard dataset contains full observed \(x,y\) pairs only for included units. An auxiliary frame could reveal counts or covariates for excluded units, so “nothing is known” is not a universal statement.
  4. Population conditional density: \(f(y\mid x;\theta)\) represents outcomes before truncation; a normal linear specification is one important case, not the definition.
  5. Admission probability: \(P_\theta(Y\in A\mid X=x)\) normalizes the observed density and usually depends on both \(x\) and \(\theta\).
  6. Inference target: the analyst distinguishes fit to the selected distribution from inference about the population relation.[1]

Condensed: population regression + outcome-based record exclusion + conditional observed density = truncated regression.

Sig role-phrases: pre-selection population → defines inferential target; covariates \(X\) → index conditional mean/density; outcome \(Y\) → determines admission; truncation region → excludes whole analysis records; admission probability → normalizes each observed density; likelihood → combines included-record contributions under assumptions.

What It Is Not

  • Not censored regression. With censoring, units outside a boundary remain in the sample but their outcomes are reported at a limit or as partial information; truncation omits their records from the classical analyzed sample.[3]
  • Not merely a truncated normal distribution. That is a marginal distribution; the regression model also links the outcome and covariates through \(\theta\).
  • Not every incomplete dataset. Missingness unrelated to outcome region, selection on a covariate alone, and survey nonresponse have different mechanisms.
  • Not automatically corrected by adding a threshold indicator to ordinary least squares. The observed error distribution has changed; the likelihood accounts for inclusion probability under a specified model.[1]
  • Not evidence that no information of any kind exists for excluded units. An external sampling frame or administrative register may hold their counts or characteristics, even if their outcomes are absent from this sample.
  • Not a universally unbiased answer from maximum likelihood. The model and inclusion rule must be defensible; misspecification can remain consequential.

Scope of Application

A constructed income-ceiling sample illustrates upper truncation. Imagine observing earnings \(Y\) and schooling \(X\) only for people whose earnings fall below eligibility ceiling \(b\). A population model for earnings conditional on schooling must account for \(P(Y<b\mid X=x)\) when using the admitted records. This is an illustrative model, not a claim that a particular real benefit program supplies exactly these data.

A constructed business-register threshold illustrates lower truncation. If a register records only firms whose measured output exceeds \(a\), the conditional density for an observed firm's output is \(f(y\mid x)/P(Y>a\mid X=x)\). The covariates and selection rule must be those of the actual register; an unmodeled selection policy would change the problem.

Normal linear truncated regression is a common parametric implementation. Other density families or semiparametric approaches are possible, but each needs its own identification and estimation argument. Amemiya's original research studied truncated-normal regression; modern work continues to study computational and statistical estimation under truncation.[2][4]

Clarity

The central distinction is whether a boundary changes the recorded value of a retained unit or determines whether that unit appears at all. In a censored sample, one might see every firm's covariates and know that output is above a cap without seeing its exact value. In a truncated sample, those above-cap firms are absent from the standard analysis file. This difference changes the likelihood.[3]

The normalizer clarifies why ordinary regression on selected cases is not usually population regression. Even if \(f(y\mid x)\) is normal with mean \(x^\top\beta\), the distribution after \(Y\in A\) is the same density divided by an inclusion probability that changes with \(x\). Its selected conditional mean need not equal \(x^\top\beta\).[1]

Dividing by an admission probability is exact within the declared model but cannot validate that model against the real sampling process. This is an evidence boundary, not another design tradeoff. Diagnostic: are the threshold, outcome law and any additional selection variables supported by data and design?[1][4]

Manages Complexity

The conditional-density formula compresses selection into one explicit probability term. It separates two questions: what is the population model, and which potential records become visible? The normalizer is computed from those two declarations rather than guessed from the shape of the observed sample.

That compression also exposes failure modes. If \(A\) is misreported, the denominator is wrong. If the population error distribution is misspecified, maximum likelihood may be precise about the wrong model. If outside-unit counts or covariates are known, a richer likelihood may use them instead of discarding that information.[1][3]

Abstract Reasoning

First identify the target population and which outcome values permit inclusion. Confirm whether excluded cases disappear or remain with censored values. Then specify a population conditional density \(f(y\mid x;\theta)\). For each included observation, divide by its inclusion probability under the same model, and estimate \(\theta\) from the product of these conditional densities (or a justified alternative method). Check whether threshold, covariates and density assumptions match the sampling process.[1]

For a lower threshold \(a\) in a normal linear model with mean \(x^\top\beta\) and standard deviation \(\sigma\), the normalizer is \(1-\Phi((a-x^\top\beta)/\sigma)\). For an upper threshold \(b\), it is \(\Phi((b-x^\top\beta)/\sigma)\). This explicit \(x\)-dependence is why simply trimming the data and using the untruncated likelihood loses information about the selection mechanism.[1]

Knowledge Transfer

The income-ceiling and firm-size examples differ in substantive domain but share the same statistical roles: population response, covariates, selection threshold, admitted records and conditional likelihood. The side of the threshold changes only the admission probability, not the logic of conditioning.

The live Statistical Model node is the accepted strict genus: this entry specifies a family of conditional observed-data laws with outcome-admission normalization. Regression is a closer subject-matter neighbor but is not the accepted edge here. Truncated Normal Distribution is a building block for one parametric instance, not a parent of every truncated regression model.

Examples

Constructed upper-truncated earnings record

Take a deliberately constructed normal model with \(x=(1,2)\), \(\beta=(0,0.5)\), so the pre-selection mean is \(\mu=x^\top\beta=1\), and \(\sigma=1\). An income file includes only outcomes \(Y<b=2\). A retained record has \(y=1\). Its inclusion probability is \(\Phi((2-1)/1)=\Phi(1)\approx0.84134\). Its untruncated normal density is \(\phi((1-1)/1)/1=\phi(0)\approx0.39894\), so its conditional density contribution is \(0.39894/0.84134\approx0.47417\), not \(0.39894\). These are chosen numbers, not an observed policy evaluation; the target is the \(Y\)-on-\(X\) relation before records above the ceiling were omitted.[1]

Mapped back: outcome = constructed earnings \(y=1\); covariates = \(x=(1,2)\); population mean = \(1\); upper truncation = \(b=2\); inclusion probability = \(0.84134\); observed-record likelihood density = \(0.47417\); target = pre-selection regression.

Stata's lower-truncated labor-hours demonstration

The Stata manual's worked labor-supply demonstration begins with a 250-record subsample: 150 women have positive market-work hours and 100 are nonmarket laborers. Its truncated-regression command models work hours among the 150 positive-hour observations with lower limit \(a=0\), reporting “100 obs. truncated.” In contrast, its illustrated Tobit treatment of censoring uses all 250 records, 150 uncensored. This is a source-located software example, not an independently verified population study. Each positive-hour record contributes its normal density divided by \(P(Y>0\mid X=x)\); the fitted slope is meant to describe a pre-selection latent/population relation under the normal model, not just an ordinary least-squares line among the positive-hour records.[1]

Mapped back: outcome = market-work hours; covariates = child counts, age and education in the manual; threshold = \(Y>0\); truncated analysis = 150 positive-hour records, with 100 excluded from that fit; normalizer = \(P(Y>0\mid X=x)\); contrast = Tobit retains all 250 under a different observation model.

Censored near miss

Suppose every firm appears, but values above \(b\) are displayed as “\(b\) or greater.” Those firms are not removed. Their contribution uses a probability mass for a censored event, not the density of a truncated sample. Calling the data truncated would mis-specify the information actually observed.[3]

Structural Tensions

Ease of fitting the visible sample versus the pre-selection target. Ordinary regression on admitted records is easier to specify and may answer a valid question about the selected group, but generally does not recover the stated pre-selection relation under outcome truncation. Conditional-likelihood modeling addresses that target under assumptions, at the cost of a stronger outcome-law and admission-rule commitment; in Stata's example the ordinary and truncated fits even return different coefficients. Diagnostic: is the target the observed selected group or the pre-selection relation, and which assumptions justify the latter?[1]

Structural–Framed Character

This occupies a mixed structural–framed position. The conditional-density identity \(f(y\mid x,Y\in A)=f(y\mid x)/P(Y\in A\mid x)\) is mathematical once a population law and inclusion region are supplied; whether those inputs describe a real file is empirical. Its evaluative weight depends on the target: an easy line through selected records can be appropriate for the selected group yet misleading for a pre-selection population. Human and institutional sampling practices—eligibility ceilings, registries or an analyst's decision to retain only positive work hours—create the record boundary. Statistical and software terminology carries “truncated” into applications, but the word travels correctly only when outside-region units are absent from the analyzed sample. Applying it to retained threshold-coded observations imports a label that belongs to censoring. Its character: a mathematically exact conditional-likelihood construction whose population interpretation is dependent on a documented selection practice and credible distributional assumptions.

Structural Core vs. Domain Accent

The portable skeleton is conditioning an underlying model on an inclusion event; that general statistical pattern might warrant a future prime-level inquiry but is not asserted as an existing strict parent. The domain-bound mechanism here ties selection specifically to a regression outcome, keeps covariates and a pre-selection parameter target, and normalizes each observed-record density by its covariate-specific admission probability. The named entry fails the prime bar because generic missing data, a marginal truncated distribution and retained censored observations do not all preserve that outcome-truncated regression likelihood. Statistical Model is the staged broad genus; the selection correction is the distinguishing specialized mechanism.

This entry is a kind of Statistical Model.

It keeps a conditional probability law, as any Statistical Model does, while adding outcome truncation and an inclusion-normalized likelihood.

Regression is a near genus that may later prove a more direct parent, and Truncated Normal Distribution is a possible component in a normal instance.

Relationships to Other Abstractions

Local relationship map for Truncated Regression ModelParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.TruncatedRegression ModelDOMAINDomain-specific abstraction: Statistical Model — is a kind ofStatisticalModelDOMAIN

Current abstraction Truncated Regression Model Domain-specific

Parents (1) — more general patterns this builds on

  • Truncated Regression Model is a kind of Statistical Model Domain-specific

    An outcome-truncated regression law is a specialized statistical model.

Hierarchy paths (6) — routes to 4 parentless roots

Neighborhood in Abstraction Space

Truncated Regression Model sits in a sparse region of the domain-specific corpus (96th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (2551 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

Censoring retains bounded or partially known outcomes for units in the file; truncation removes records outside the selection region. A truncated-normal marginal law is not a full regression. Ordinary least squares on admitted records usually targets the wrong population relation under outcome selection; model-corrected likelihood still requires correct assumptions. Excluded units need not be wholly unknowable if auxiliary data exist.[1][3]

References

[1] Stata, truncreg — Truncated regression, Methods and formulas, first-party conditional likelihood specification. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n

[2] Amemiya, “Regression Analysis when the Dependent Variable Is Truncated Normal,” Econometrica 41 (1973), original paper identified; full text not checked. registry ↩a ↩b

[3] Stata, User's Guide, section on censored and truncated regression, first-party distinction between absent and retained-boundary cases. registry ↩a ↩b ↩c ↩d ↩e

[4] “Computationally and Statistically Efficient Truncated Regression”, original research abstract. registry ↩a ↩b