Least Trimmed Squares¶
A robust regression estimator that minimizes the sum of the h smallest squared residuals, reselecting the retained cases for each trial fit.
Core Idea¶
Least trimmed squares (LTS) is a robust regression estimator defined by an ordered-residual objective. For a candidate parameter vector \(\beta\), calculate every residual \(r_i(\beta)=y_i-x_i^T\beta\), order their squares as \(r^2_{(1)}(\beta)\leq\cdots\leq r^2_{(n)}(\beta)\), and minimize
over \(\beta\). The retained \(h\) cases can change as \(\beta\) changes; they are not a fixed list discarded before fitting. For \(h=n\), this is ordinary least squares (OLS). With a suitably chosen smaller \(h\), it can resist contamination that would pull OLS toward a few extreme observations.[1]
The specific value of \(h\) matters. In the original simple-line exposition, \(h=\lfloor n/2\rfloor+1\); a later appendix gives \(h=\lfloor n/2\rfloor+\lfloor(p+1)/2\rfloor\) for \(p\) fitted parameters in its general-position treatment. High breakdown approaching one half is a conditional property of that design, not a blanket guarantee that any \(h\), data geometry, or numerical implementation tolerates \(n-h\) arbitrary bad cases. A trimming rule identifies a fit to a majority structure; it does not certify that excluded observations are errors.[1]
Structural Signature¶
Sig role-phrases: regression cases and model — trial-dependent squared residuals — retained count \(h\) — ordered trimmed objective — global minimizer — downstream case interpretation.
- Regression cases and model. Observed \((x_i,y_i)\) and a parameterized prediction \(x_i^T\beta\) define what is being fitted. The linear form is the original setting; applying the same objective to another regression function requires rechecking its statistical properties.[1]
- Trial-dependent squared residuals. Every proposed \(\beta\) gives a new vector of discrepancies. Their squares are ordered anew; sorting once using an OLS pilot does not define LTS.[1]
- Retained count \(h\). It fixes how many residual squares contribute. Small \(h\) emphasizes a best-fitting subset; \(h=n\) removes trimming. Feasible high-breakdown choices also depend on model dimension and identifiability.[1]
- Trimmed objective and minimizer. The estimator is the parameter value attaining the smallest \(Q_h\), not any arbitrary result of one subset search or one computational iteration. Approximate FAST-LTS computation is a means of searching for it, not a different definition.[2]
- Diagnostic interpretation. Large residuals under the selected fit may be inspected for error, subgroup, leverage or valid exception. This is an important use of LTS, not an extra term in \(Q_h\).[3]
What It Is Not¶
LTS is not least median of squares: that method minimizes a median/order statistic rather than a sum of the \(h\) smallest squared residuals. It is not least absolute deviations, which sums absolute residuals rather than reselecting a trimmed squared-residual subset. It is not OLS after deleting cases chosen once from an OLS diagnostic, because LTS reorders residuals for every trial \(\beta\). It is not the later FAST-LTS algorithm or its concentration step; an implementation may approximate the global minimizer. Nor does a high LTS residual automatically prove that a record is erroneous or expendable.[1][3][2]
Scope of Application¶
The original actuarial article considers simple linear relations in life-insurance, pension, health-insurance and inflation data, using LTS where grossly deviant cases can dominate OLS. A later original econometric study applies LTS diagnostics to multivariable national-growth regressions. Both instantiate the same ordered-residual objective, although the observations, predictors and inferential questions differ.[1][3]
The method presupposes a regression model whose residuals meaningfully compare observations. Breakdown statements require their stated model dimension, admissible \(h\), equivariance and general-position conditions. LTS cannot by itself fix an incorrect functional form, identify a causal effect, diagnose why a point differs, or guarantee that the majority data are generated by the model. Reweighting, scale estimation, hypothesis tests and subsequent OLS refits are optional downstream procedures with separate assumptions, not part of \(Q_h\).[1][3]
Clarity¶
The word trimmed can suggest that observations are discarded at the outset. LTS instead chooses the lowest \(h\) squared residuals within each trial fit. A data point may be retained for one \(\beta\) and excluded for another. That distinction explains why the estimator is a global optimization problem over both coefficients and a changing candidate subset, and why simply refitting after a preliminary outlier screen is not identical.[1]
It also separates two different statements: an LTS coefficient estimate and a subsequent judgment about suspicious cases. In the original growth reanalysis, a very high standardized LTS residual for Zambia prompted a separate OLS comparison without that country. The latter was an analytic follow-up, not the definition of the LTS fit or proof that the observation was invalid.[3]
Manages Complexity¶
OLS gives each squared residual a place in its objective. LTS compresses the fitting decision to the best-fitting \(h\) cases under each trial model, reducing the influence of observations that are grossly discordant with the candidate majority pattern. This creates a usable common language for actuarial and econometric datasets: model, residual ranking, \(h\), minimized trimmed sum and post-fit inspection.[1][3]
The compression has costs. The ordered subset changes discontinuously when residual ranks cross, making global search harder than the quadratic OLS problem. The fit may neglect a scientifically meaningful minority. FAST-LTS was developed precisely because exact LTS search became impractical for larger datasets; its reported large-data results may be approximate. Therefore computational output should not be silently equated with the mathematical global optimum.[2]
Abstract Reasoning¶
To identify a legitimate LTS analysis, first state \(n\), the predictors, response, parameter dimension and coverage \(h\). Then, for each candidate coefficient vector, compute all squared residuals, order them, and sum only the smallest \(h\). Minimize this objective, report how it was computed and whether the reported value is exact or approximate. Finally inspect high-residual or high-leverage cases against substantive knowledge rather than calling all trimmed cases mistakes.[1][3][2]
This sequence licenses a limited inference: if a small number of observations exert large influence on OLS but do not control an LTS fit under an appropriate \(h\), the two estimates reveal sensitivity to contamination or model heterogeneity. It does not tell which estimate describes the desired population until observation provenance, model form and inferential target are examined.[3]
Knowledge Transfer¶
The actuarial and growth-regression settings share literal statistical roles. Both specify a regression relationship; both calculate squared residuals at each trial fit; both retain an \(h\)-case subset through ordering; both compare the resulting majority fit with the behavior of OLS. What changes is the data-generating context and the consequence of calling a case atypical: an insurance record, a country-year growth pattern and a survey measurement error need different substantive review.[1][3]
Outside regression, a loose analogy to “ignore the worst few observations” is not LTS. Without a meaningful residual function, declared coverage, and optimization over its ordered squares, the mathematical guarantee and method name do not transfer. Even within regression, an approximation algorithm and a reweighted second-stage fit should be labeled separately from the underlying LTS estimator.[2]
Examples¶
Actuarial line fitting. Rousseeuw, Daniels and Leroy set out \(y=ax+b+\text{error}\) for insurance-related paired observations, including life-insurance, pension and health-insurance contexts. Where exceptional cases can bend the OLS line, their LTS criterion seeks the line with the lowest sum of \(h=\lfloor n/2\rfloor+1\) squared residuals in the simple-line formulation. The example shows a method for fitting a majority relation, not a rule for discarding every exceptional policy or claim.[1] Mapped back: regression cases and model = actuarial input/output pairs under a line; trial-dependent squared residuals = vertical squared discrepancies for each line; retained count = the declared \(h\); ordered trimmed objective = sum of its \(h\) smallest squared discrepancies; global minimizer = preferred LTS line; downstream case interpretation = investigate large residuals as possible errors, different populations or valid events.
Cross-country growth regression. Zaman, Rousseeuw and Orhan revisit a 61-country 1960–1985 growth model with equipment and non-equipment investment among its predictors. They use LTS residuals to inspect the majority fit and report Zambia as having a very high standardized LTS residual. Their subsequent OLS comparison with that country omitted changes some reported coefficient inferences; it is a separate follow-up and cannot be substituted for the LTS objective itself.[3] Mapped back: regression cases and model = countries, growth response and economic predictors; trial-dependent squared residuals = growth deviations under each coefficient vector; retained count = an \(h\) chosen for high-breakdown fitting; ordered trimmed objective = \(h\) smallest squared growth residuals; global minimizer = LTS coefficient vector; downstream case interpretation = substantive review of Zambia and a separately labeled OLS sensitivity comparison.
Near miss. Fit OLS, permanently delete the \(n-h\) largest OLS residuals, and refit OLS once. That pipeline uses a fixed exclusion list from a preliminary model. Unless the resulting fit happens also to minimize the re-ranked \(Q_h\) globally, it is not the LTS estimator.[1]
Structural Tensions¶
Contamination resistance versus clean-model information. Choosing a smaller \(h\) limits how many cases enter each candidate fit and can protect against gross errors, but it also gives up information from observations that would have been valid under a clean homogeneous model. A larger \(h\) retains more data and can improve precision while making contamination more influential. Neither pole can be optimized without a model of plausible contamination and an efficiency target. Diagnostic: Given parameter dimension and credible contamination, which admissible \(h\) keeps a stable majority fit without throwing away needed clean-model precision?[1]
Majority-pattern fidelity versus exceptional-case fidelity. A trimmed fit can expose a common relation that unusual cases obscure; insisting all cases shape the coefficients can mask that relation. Yet the unusual cases might be genuine and central to the scientific or actuarial question, so treating them as waste can conceal an important second process. Diagnostic: Does the large-residual case reflect a recording problem, a different population, or a valid event requiring its own model and report?[1][3]
Structural–Framed Character¶
LTS lies toward the structural end as a defined estimator, but its applicability is framed by statistical modeling choices. Vocabulary travel: the same order-statistic objective is recognizable across actuarial and economic regression, but not in any generic act of ignoring bad cases. Evaluative weight: robustness is useful only relative to a specified target and clean-model efficiency; trimming is not intrinsically better. Institutional origin: robust-statistics research developed the method as an answer to OLS sensitivity. Human-practice dependence: choosing predictors, \(h\), error interpretation and computational algorithm is analyst-dependent, whereas the stated \(Q_h\) is a mathematical predicate once those are fixed. Import versus recognition: using the name outside regression imports residual and optimization assumptions that may fail. Its character: a domain-specific statistical method with a formally repeatable trimmed-loss core and context-sensitive interpretation.[1][3]
Structural Core vs. Domain Accent¶
The portable skeleton is selectively limit influence to preserve a result under perturbation. Live Robustness captures that wider idea; it does not specify regression coefficients, squared residuals, order statistics or \(h\). The domain accent is exactly the candidate-dependent \(h\)-smallest squared-residual optimization and its statistical breakdown/efficiency analysis. Remove those features and one may have another robust method, but not LTS. Thus the named entry does not clear the prime bar despite exemplifying a prime robustness idea.[1]
Within the domain, the live Robust Regression identity is a proposed strict parent: it explicitly includes high-breakdown LTS among methods designed to keep a limited amount of contamination from dominating the fit. The present entry is a specified estimator, not a duplicate of that family. Prime Robustness is a broader conceptual relation, not an asserted strict parent edge here.
Instantiates / Related Primes¶
This entry is a kind of Robust Regression.
DAG parent: Robust Regression, pending global DAG review. The broader abstraction explicitly names LTS as one high-breakdown member; LTS supplies the additional ordered squared-residual objective. Related prime: Robustness describes stability under disturbance, which LTS operationalizes for a declared contamination setting, but its broad signature does not alone identify the regression estimator. Declined nearby nodes: Linear least squares shares quadratic residuals but generally minimizes the full sum; Least absolute deviations changes loss rather than using the LTS order-statistic sum. No canonical edge is changed.
Relationships to Other Abstractions¶
Current abstraction Least Trimmed Squares Domain-specific
Parents (1) — more general patterns this builds on
-
Least Trimmed Squares is a kind of Robust Regression Domain-specific
LTS is a specified high-breakdown robust regression estimator.Live Robust Regression explicitly treats LTS as a high-breakdown member of its regression family. LTS adds a particular h-smallest-squared-residual objective and coverage choice.
Hierarchy paths (13) — routes to 7 parentless roots
- Least Trimmed Squares → Robust Regression → Regression → Signal Extraction
- Least Trimmed Squares → Robust Regression → Regression → Function (Mapping)
- Least Trimmed Squares → Robust Regression → Regression → Statistical Inference → Inductive Reasoning
- Least Trimmed Squares → Robust Regression → Regression → Statistical Inference → Uncertainty
- Least Trimmed Squares → Robust Regression → Regression → Distributional Assumption → Assumption → Epistemic Mode Of A Proposition
- Least Trimmed Squares → Robust Regression → Regression → Distributional Assumption → Statistical Inference → Inductive Reasoning
- Least Trimmed Squares → Robust Regression → Regression → Distributional Assumption → Statistical Inference → Uncertainty
- Least Trimmed Squares → Robust Regression → Regression → Distributional Assumption → Probability → Measure → Set and Membership
- Least Trimmed Squares → Robust Regression → Regression → Statistical Inference → Probability → Measure → Set and Membership
- Least Trimmed Squares → Robust Regression → Regression → Distributional Assumption → Probability → Measure → Aggregation → Micro Macro Linkage
- Least Trimmed Squares → Robust Regression → Regression → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
- Least Trimmed Squares → Robust Regression → Regression → Distributional Assumption → Statistical Inference → Probability → Measure → Set and Membership
- Least Trimmed Squares → Robust Regression → Regression → Distributional Assumption → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
Neighborhood in Abstraction Space¶
Least Trimmed Squares sits in a sparse region of the domain-specific corpus (74th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Causal Inference & Regression Modeling (15 abstractions)
Nearest neighbors
- Residual Sum of Squares — 0.86
- Regression — 0.84
- Nonlinear Least Squares — 0.83
- FWL theorem — 0.82
- Hat matrix — 0.82
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
OLS: all \(n\) residual squares contribute. Least median of squares: one middle squared residual is minimized, not the sum of \(h\) smallest. Least absolute deviations: absolute residuals replace squares without LTS trimming. Fixed outlier deletion plus OLS: the discarded list does not vary during optimization. FAST-LTS: a search algorithm that can approximate the global LTS objective, not the objective itself. Outlier diagnosis: a high residual flags a question, not a conclusion that a case is erroneous.[1][3][2]
References¶
[1] Peter J. Rousseeuw, B. Daniels and A. Leroy, “Applying robust regression to insurance,” Insurance: Mathematics and Economics 3 (1984), 67–72; equations (1)–(4), actuarial applications, and Appendix equation (5) and breakdown discussion. https://wis.kuleuven.be/stat/robust/papers/publications-1984/rousseeuwdanielsleroy-robustregressioninsurance-im.pdf registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r ↩s
[2] Peter J. Rousseeuw and Katrien Van Driessen, “Computing LTS Regression for Large Data Sets,” Data Mining and Knowledge Discovery 12 (2006), 29–45, original-author abstract only; full text not inspected. https://doi.org/10.1007/s10618-005-0024-4 registry ↩a ↩b ↩c ↩d ↩e ↩f
[3] Asad Zaman, Peter J. Rousseeuw and Mehmet Orhan, “Econometric applications of high-breakdown robust regression techniques,” Economics Letters 71 (2001), 1–8, §2–3 and Table 1. https://wis.kuleuven.be/stat/robust/papers/2001/zamanrousseeuworhan-econometricapplications-econom.pdf registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m