Tensions in Practice: Preserving variation in tension with preserving class separation¶
Four invented labeled points
Each point has coordinates u and v. Two derived coordinates are x = (u + v)/2 and y = (u − v)/2. Across these four points, x varies between −10 and 10; y varies between −1 and 1. Keeping x preserves the largest variation but merges a Plus and Minus point at each retained value. Keeping y separates the declared classes but discards the much larger variation along x.
Keep dominant variation
Retain the strongest variation without using class labels.
Keep the target distinction
Retain the coordinate that separates the declared classes.
Why these aims pull against each other
A single retained coordinate cannot preserve both independent directions. The objective decides which differences survive.
Choose an arrangement to see what changes and what remains difficult.
The same four input points and class labels remain fixed. Only the kept coordinate changes; equal kept values make the lost distinction visible.
What this choice protects
What it costs
When it fits
Compare the arrangements
Keep dominant variation
Project onto x. Its four equally weighted values have variance 100; y has variance 1. These orthogonal directions have zero covariance.
| u, v | Kept | Class | |
|---|---|---|---|
| P1 | 11, 9 | 10 | Plus |
| P2 | 9, 11 | 10 | Minus |
| P3 | −9, −11 | −10 | Plus |
| P4 | −11, −9 | −10 | Minus |
- What it protects
- Large-scale variation remains available without consulting the labels.
- What it costs
- Opposite classes collapse to the same x values.
- When it fits
- Fits a descriptive task whose priority is dominant variation rather than these labels.
Illustration note: This is a rescaled principal axis of the declared data; no claim is made about generalization to new data.
Keep the class distinction
Project onto y. Plus has y = 1 and Minus has y = −1.
| u, v | Kept | Class | |
|---|---|---|---|
| P1 | 11, 9 | 1 | Plus |
| P2 | 9, 11 | −1 | Minus |
| P3 | −9, −11 | 1 | Plus |
| P4 | −11, −9 | −1 | Minus |
- What it protects
- The two declared classes remain separable in one dimension.
- What it costs
- Variation along x is lost, and the objective depends on the selected labels.
- When it fits
- Fits this label-focused task when preserving x is less important.
Illustration note: Both coordinates are linear combinations of u and v, not merely selecting an original feature. Separability here is not predictive validation.
What this illustration does—and does not—establish
The source supplies the structural tension; the invented example makes one relation inspectable. Costs and conditions are part of each arrangement, not exceptions to a universal recommendation.
- Coordinates, equal weights and labels are invented. No fitted model performance is reported.
- The calculation concerns these four points. Both projections discard a direction; neither is lossless.
- The x/y scaling is explicitly defined; changing measurement units can change a variance objective.
Source entries
Dimensionality Reduction
The canonical tension motivates this comparison. The setting, finite values and arrangements are declared editorial illustrations, not measured findings.
Variance Preservation versus Task-Relevance
PCA preserves variance, but variance is not always the structure that matters for downstream tasks. For classification, LDA (which maximizes class separability) may outperform PCA; for clustering, methods that preserve local structure (UMAP, t-SNE) may be preferable. The tension is that an unsupervised objective like variance is task-agnostic, but tasks have their own structure-preservation needs.
The source operation
Dimensionality reduction transforms high-dimensional data into a lower-dimensional representation that preserves the structural properties most important for downstream tasks — variance, pairwise distances, neighborhood relationships, or predictive information — while discarding redundant, noisy, or low-information dimensions