Tensions in Practice: Sparse coefficients in tension with distributed shrinkage¶
Two coefficients in an invented orthogonal fitting problem
A two-coefficient fit has target (3, 1), and its fit cost is half the squared distance from that target. Add either |β1| + |β2|, called an L1 penalty, or (β1² + β2²)/2, an L2 penalty. The first objective is minimized at (2, 0); the second at (1.5, 0.5). Both charge for large coefficients, yet one removes the second coefficient while the other retains both.
Favor a sparse fit
Allow weak coefficients to become exactly zero.
Spread the shrinkage
Reduce both coefficients without selecting one away.
Why these aims pull against each other
Penalty shape changes the kind of solution favored; adjusting a weight is a separate choice and cannot be treated as a neutral notion of simplicity.
Choose an arrangement to see what changes and what remains difficult.
Rows retain the same coefficient targets. The selected coefficient values expose exact sparsity versus shrinkage of both coordinates.
What this choice protects
What it costs
When it fits
Compare the arrangements
Penalize absolute size
Minimize the stated fit cost plus |β1| + |β2|. Each positive target is reduced by one, with zero as the stopping point.
| Fit target | Chosen | Is zero? | |
|---|---|---|---|
| β1 | 3 | 2 | No |
| β2 | 1 | 0 | Yes |
- What it protects
- The second coordinate drops out of this fit.
- What it costs
- A real weak contribution could be removed; the larger coefficient remains biased downward.
- When it fits
- Fits a justified sparse-signal hypothesis, with predictive performance checked separately.
Illustration note: At (2, 0), fit cost is 1 and penalty is 2. The formula is a soft charge, not a forbidden-coordinate rule.
Penalize squared size
Minimize the same fit cost plus (β1² + β2²)/2. The optimum halves each target.
| Fit target | Chosen | Is zero? | |
|---|---|---|---|
| β1 | 3 | 1.5 | No |
| β2 | 1 | 0.5 | No |
- What it protects
- Both possible contributors remain in the fitted model.
- What it costs
- The result is not sparse and both coefficients are biased downward.
- When it fits
- Fits a justified distributed-signal hypothesis when removing a weak contributor is undesirable.
Illustration note: At (1.5, 0.5), fit cost is 1.25 and penalty is 1.25. Objective totals across different penalties are not a ranking of predictive quality.
What this illustration does—and does not—establish
The source supplies the structural tension; the invented example makes one relation inspectable. Costs and conditions are part of each arrangement, not exceptions to a universal recommendation.
- The orthogonal quadratic loss and weights are declared toy choices, not tuned recommendations.
- The weights have different meanings under different penalties; equal numerical coefficients would not make the strengths equivalent.
- No true coefficients or unseen outcomes are given. Neither penalty is demonstrated to generalize better.
Source entries
Regularization
The canonical tension motivates this comparison. The setting, finite values and arrangements are declared editorial illustrations, not measured findings.
Choice of Norm versus Choice of Strength (Two Independent Knobs)
Regularization has two separable degrees of freedom: which complexity measure to penalize (the norm's shape) and how hard (the weight). They do different work — L2 shrinks uniformly, L1 sparsifies, total-variation preserves edges — and a well-tuned weight on the wrong norm favors the wrong kind of solution.
The source operation
Regularization is the structural move of adding a penalty on the complexity — the roughness, the norm, the deviation-from-prior — of a candidate solution to a fitting or optimization procedure, so that the solution chosen is one that trades data-fit against complexity according to an explicit, tunable weight.