Skip to content

Tensions in Practice: Sparse coefficients in tension with distributed shrinkage

Two coefficients in an invented orthogonal fitting problem

A two-coefficient fit has target (3, 1), and its fit cost is half the squared distance from that target. Add either |β1| + |β2|, called an L1 penalty, or (β1² + β2²)/2, an L2 penalty. The first objective is minimized at (2, 0); the second at (1.5, 0.5). Both charge for large coefficients, yet one removes the second coefficient while the other retains both.

Favor a sparse fit

Allow weak coefficients to become exactly zero.

Spread the shrinkage

Reduce both coefficients without selecting one away.

Why these aims pull against each other

Penalty shape changes the kind of solution favored; adjusting a weight is a separate choice and cannot be treated as a neutral notion of simplicity.

Compare the arrangements

Penalize absolute size

Minimize the stated fit cost plus |β1| + |β2|. Each positive target is reduced by one, with zero as the stopping point.

Absolute-value penalty
Fit targetChosenIs zero?
β132No
β210Yes
What it protects
The second coordinate drops out of this fit.
What it costs
A real weak contribution could be removed; the larger coefficient remains biased downward.
When it fits
Fits a justified sparse-signal hypothesis, with predictive performance checked separately.

Illustration note: At (2, 0), fit cost is 1 and penalty is 2. The formula is a soft charge, not a forbidden-coordinate rule.

Penalize squared size

Minimize the same fit cost plus (β1² + β2²)/2. The optimum halves each target.

Squared-value penalty
Fit targetChosenIs zero?
β131.5No
β210.5No
What it protects
Both possible contributors remain in the fitted model.
What it costs
The result is not sparse and both coefficients are biased downward.
When it fits
Fits a justified distributed-signal hypothesis when removing a weak contributor is undesirable.

Illustration note: At (1.5, 0.5), fit cost is 1.25 and penalty is 1.25. Objective totals across different penalties are not a ranking of predictive quality.

What this illustration does—and does not—establish

The source supplies the structural tension; the invented example makes one relation inspectable. Costs and conditions are part of each arrangement, not exceptions to a universal recommendation.

  • The orthogonal quadratic loss and weights are declared toy choices, not tuned recommendations.
  • The weights have different meanings under different penalties; equal numerical coefficients would not make the strengths equivalent.
  • No true coefficients or unseen outcomes are given. Neither penalty is demonstrated to generalize better.

Source entries

Regularization

Prime · Source of the tension

The canonical tension motivates this comparison. The setting, finite values and arrangements are declared editorial illustrations, not measured findings.

Choice of Norm versus Choice of Strength (Two Independent Knobs)

Regularization has two separable degrees of freedom: which complexity measure to penalize (the norm's shape) and how hard (the weight). They do different work — L2 shrinks uniformly, L1 sparsifies, total-variation preserves edges — and a well-tuned weight on the wrong norm favors the wrong kind of solution.

Read the source section

The source operation

Regularization is the structural move of adding a penalty on the complexity — the roughness, the norm, the deviation-from-prior — of a candidate solution to a fitting or optimization procedure, so that the solution chosen is one that trades data-fit against complexity according to an explicit, tunable weight.

Read the source section