Skip to content

Tensions in Practice: Preserving variation in tension with preserving class separation

Four invented labeled points

Each point has coordinates u and v. Two derived coordinates are x = (u + v)/2 and y = (u − v)/2. Across these four points, x varies between −10 and 10; y varies between −1 and 1. Keeping x preserves the largest variation but merges a Plus and Minus point at each retained value. Keeping y separates the declared classes but discards the much larger variation along x.

Keep dominant variation

Retain the strongest variation without using class labels.

Keep the target distinction

Retain the coordinate that separates the declared classes.

Why these aims pull against each other

A single retained coordinate cannot preserve both independent directions. The objective decides which differences survive.

Compare the arrangements

Keep dominant variation

Project onto x. Its four equally weighted values have variance 100; y has variance 1. These orthogonal directions have zero covariance.

Keep the largest-variance axis
u, vKeptClass
P111, 910Plus
P29, 1110Minus
P3−9, −11−10Plus
P4−11, −9−10Minus
What it protects
Large-scale variation remains available without consulting the labels.
What it costs
Opposite classes collapse to the same x values.
When it fits
Fits a descriptive task whose priority is dominant variation rather than these labels.

Illustration note: This is a rescaled principal axis of the declared data; no claim is made about generalization to new data.

Keep the class distinction

Project onto y. Plus has y = 1 and Minus has y = −1.

Keep the class axis
u, vKeptClass
P111, 91Plus
P29, 11−1Minus
P3−9, −111Plus
P4−11, −9−1Minus
What it protects
The two declared classes remain separable in one dimension.
What it costs
Variation along x is lost, and the objective depends on the selected labels.
When it fits
Fits this label-focused task when preserving x is less important.

Illustration note: Both coordinates are linear combinations of u and v, not merely selecting an original feature. Separability here is not predictive validation.

What this illustration does—and does not—establish

The source supplies the structural tension; the invented example makes one relation inspectable. Costs and conditions are part of each arrangement, not exceptions to a universal recommendation.

  • Coordinates, equal weights and labels are invented. No fitted model performance is reported.
  • The calculation concerns these four points. Both projections discard a direction; neither is lossless.
  • The x/y scaling is explicitly defined; changing measurement units can change a variance objective.

Source entries

Dimensionality Reduction

Prime · Source of the tension

The canonical tension motivates this comparison. The setting, finite values and arrangements are declared editorial illustrations, not measured findings.

Variance Preservation versus Task-Relevance

PCA preserves variance, but variance is not always the structure that matters for downstream tasks. For classification, LDA (which maximizes class separability) may outperform PCA; for clustering, methods that preserve local structure (UMAP, t-SNE) may be preferable. The tension is that an unsupervised objective like variance is task-agnostic, but tasks have their own structure-preservation needs.

Read the source section

The source operation

Dimensionality reduction transforms high-dimensional data into a lower-dimensional representation that preserves the structural properties most important for downstream tasks — variance, pairwise distances, neighborhood relationships, or predictive information — while discarding redundant, noisy, or low-information dimensions

Read the source section