Fraction of variance unexplained¶
A regression-fit statistic equal to the proportion of dependent-variable variance left unexplained by the model's predictions.
Core Idea¶
Fraction of variance unexplained is a regression-fit statistic equal to the proportion of dependent-variable variance left unexplained by the model's predictions. [1]
The fraction of variance unexplained is the residual variation divided by total variation for the same response and evaluation sample. Under ordinary least squares with an intercept and standard sum-of-squares decomposition, FVU=SSE/SST=1−R². Outside those conditions, the equality and even the baseline denominator require explicit definition.
Its operative boundary is not supplied by the name alone. Preserve this identity: A regression-fit statistic equal to the proportion of dependent-variable variance left unexplained by the model's predictions. Validity boundary: The numerator and denominator must use compatible residual and total variance definitions for the same regressand; generic prediction error is insufficient. The entry therefore captures a reusable specialist role structure rather than a topic label, a single historical instance, or a loose analogy.
Structural Signature¶
Sig role-phrases:
- the response observations — the dependent-variable values being evaluated
- the model predictions — fitted or out-of-sample estimates for those observations
- the residuals — observed minus predicted values
- the residual variation — SSE or a compatible residual-variance estimate
- the baseline center — usually the evaluation-sample mean of the response
- the total variation — SST or compatible variance about that baseline
- the ratio — residual variation divided by total variation
- the evaluation regime — training, validation, weighted, grouped, or population context
Recognition test. A case qualifies only when the analyst can map the declared the response observations, the model predictions, the residuals, the residual variation, the baseline center and preserve the specialist validity conditions. Shared vocabulary, a similar output, or a generic instance of one parent relation is insufficient.
What It Is Not¶
- Not the residual variance alone. FVU normalizes residual variation by total response variation.
- Not always exactly one minus R-squared. The identity depends on matched definitions and decomposition conditions.
- Not mean squared error. MSE is unnormalized and changes with outcome scale.
- Not causal unexplained variance. Residual association does not identify causal omissions.
- Not automatically bounded between zero and one. Out-of-sample or no-intercept predictions can be worse than the mean baseline.
Scope of Application¶
The abstraction recurs literally within regression and prediction settings where residual error is compared with a same-sample mean baseline. The following habitats preserve the same recognition machinery; they are not invitations to extend the name metaphorically.
- Ordinary least squares. SSE/SST complements the standard R² with an intercept.
- Out-of-sample evaluation. FVU compares predictions to the held-out mean baseline and can exceed one.
- Weighted regression. both residual and total sums must use the same weights.
- Signal modeling. unmodeled signal energy is normalized by total centered energy.
- Model comparison. lower FVU indicates better squared-error fit on a common response sample.
Clarity¶
State centering, weights, degrees-of-freedom convention, evaluation sample, and treatment of missing values. Never combine training residual variance with test total variance. Interpret 'unexplained' predictively; it includes noise, misspecification, and sampling error and is not proof of absent causes.
A practical identification audit begins with the typed roles rather than the title: establish the response observations, verify the model predictions, then test the remaining conditions and exclusions. If the case retains only the portable skeleton described below, it should be named through a parent abstraction rather than as Fraction of variance unexplained.
Manages Complexity¶
The ratio makes squared prediction error scale-free relative to a simple mean benchmark. Sum-of-squares decomposition localizes what a fitted linear model captures and what remains, while boundary cases warn when the model underperforms baseline.
The compression remains accountable because each simplification has a named failure condition. Disagreement can be localized to a missing role, an invalid assumption, an ambiguous measurement, or a neighboring abstraction instead of being hidden inside an unanalyzed label.
Abstract Reasoning¶
R1. Fix one response vector, prediction vector, weights, and evaluation sample. R2. Compute residuals and their compatible sum or variance. R3. Compute total variation about the explicitly chosen baseline with the same weights. R4. Take the ratio and establish whether the OLS decomposition licenses 1−R². R5. Report uncertainty and avoid causal language unless supported by a separate design.
These moves separate definition, derivation, measurement, and interpretation. A formal consequence does not by itself prove that an observed case instantiates the abstraction, while an observed resemblance does not relax the formal or institutional recognition conditions.
Knowledge Transfer¶
The statistic transfers literally across squared-error regression evaluations with compatible numerator and denominator. Measurement and decomposition are parents; generic model error or unexplained narrative variation is not FVU.
The transfer boundary is explicit: DOMAIN-SPECIFIC PASS / PRIME FAIL: The statistic is computed across regression models and datasets by comparing residual variation with total variation. Literal recognition retains the specialist vocabulary and validity conditions of regression analysis; outside that setting only broader parent operations transfer. The safe move beyond the home habitat is to carry the applicable parent relation and leave the specialist name behind unless every defining role remains literal.
Examples¶
Canonical: OLS with an intercept¶
For a least-squares fit on its training sample, residuals are orthogonal to the fitted values and the response is centered at its mean, so SST=SSR+SSE. The FVU is SSE/SST and equals 1−R² under these definitions. [1]
Mapped back: the response observations; the model predictions; the residuals; the residual variation; the total variation; the ratio.
Applied / In Practice: worse-than-baseline test predictions¶
On held-out data, compute SSE from fixed model predictions and SST around the held-out response mean. If SSE exceeds SST, FVU is greater than one, correctly indicating worse squared error than predicting the test mean. [2]
Mapped back: the baseline center; the total variation; the ratio; the evaluation regime.
Structural Tensions¶
T1: Simple complement vs definition dependence. The mnemonic 1−R² conceals intercept, centering, and sample assumptions. Diagnostic: Do numerator and denominator support the decomposition?
T2: Training fit vs generalization. In-sample projection identities do not guarantee held-out performance. Diagnostic: Which evaluation regime is reported?
T3: Scale-free ratio vs low-variance outcomes. A small denominator makes the ratio unstable or uninformative. Diagnostic: Is total variation materially nonzero?
T4: Predictive residual vs causal explanation. Unexplained statistical variation has many causal and measurement sources. Diagnostic: Is causal language warranted?
T5: Weighting consistency vs convenience. Different weights silently change the estimand. Diagnostic: Are identical units and weights used throughout?
T6: Domain autonomy vs prime reduction. Measurement and Decomposition omit the specialist objects, constraints, and validity tests named above. Diagnostic: Would retaining only the portable parent pattern still satisfy the recognition test?
Structural–Framed Character¶
The five-criterion aggregate is 0.15 (structural). The judgment is criterion-specific:
- Vocabulary travels — low (0.25). The complete vocabulary remains tied to the typed roles in the Structural Signature.
- Evaluative weight — low (0.00). Application carries the stated degree of normative or interpretive judgment beyond structural recognition.
- Institutional origin — low (0.25). The abstraction depends to this degree on a scholarly, technical, legal, or social convention.
- Human-practice bound — low (0.00). Recognition depends to this degree on organized practice, language, measurement, or institutional action.
- Import versus recognize — low (0.25). Beyond its home habitat, use of the full name increasingly becomes analogy rather than literal recognition.
The portable skeleton is residual variation is normalized by total baseline variation to measure the share left after prediction. The named abstraction remains structural because that skeleton alone does not supply its specialist objects, constraints, or tests.
Structural Core vs. Domain Accent¶
Structural core: Residual variation is normalized by total baseline variation to measure the share left after prediction.
Domain accent: Regressands, predictions, residual sums of squares, total sums of squares, mean baselines, r-squared, intercepts, weights, and test evaluation.
Why it does not clear the prime bar: Measurement and decomposition travel; FVU is the regression-specific residual-to-total variance ratio. Generalization therefore routes through parent abstractions; preserving the specialist name requires the full accent.
Instantiates / Related Primes¶
- Measurement (
prime:measurement). The statistic quantifies model error relative to response variability. - Decomposition (
prime:decomposition). Under OLS conditions total variation splits into explained and residual components.
These are prose placement proposals only. They create no dag_edges; endpoint, redundancy, and cycle checks are recorded separately in the bundle's placement memo.
Relationships to Other Abstractions¶
Current abstraction Fraction of variance unexplained Domain-specific
Parents (1) — more general patterns this builds on
-
Fraction of variance unexplained is a kind of Measurement Prime
Measurement (
prime:measurement).The statistic quantifies model error relative to response variability.
Hierarchy path (1) — routes to 1 parentless root
- Fraction of variance unexplained → Measurement
Neighborhood in Abstraction Space¶
Fraction of variance unexplained sits in a sparse region of the domain-specific corpus (71st percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Statistical Adjustment & Estimation Effects (14 abstractions)
Nearest neighbors
- Variance function — 0.86
- Regression — 0.85
- Nonlinear Least Squares — 0.85
- Kriging — 0.84
- Dependent and independent variables — 0.84
Computed from structural-signature embeddings · 2026-09-08
Not to Be Confused With¶
- R-squared. the complementary explained fraction under standard definitions. Tell: Are conditions for one minus FVU satisfied?
- Residual variance. the numerator scale alone. Tell: Is it normalized by total variation?
- Mean squared error. average squared residual. Tell: Is there a response-variance denominator?
- Unexplained sum of squares. SSE before normalization. Tell: Is an absolute or fractional quantity reported?
- Noise variance. irreducible conditional variability. Tell: Do residuals also include model bias and estimation error?
References¶
[1] Norman R. Draper and Harry Smith, Applied Regression Analysis, 3rd ed., Wiley, 1998. registry ↩a ↩b
[2] Trevor Hastie, Robert Tibshirani, and Jerome Friedman, The Elements of Statistical Learning, 2nd ed., Springer, 2009. registry ↩