Tensions in Practice: Tradable complexity in tension with an inviolable limit¶
Two invented fitted-model candidates
Two models have invented fitting errors and complexity scores: S has error 6 and complexity 0; R has error 0 and complexity 2. Charging one error-unit per complexity unit makes their totals 6 and 2, so R wins. A hard complexity limit of 1 excludes R regardless of its fit. The numbers reveal a difference between pricing a departure and forbidding it.
Trade fit against complexity
Allow a more complex candidate when its fit benefit justifies the chosen price.
Enforce a hard boundary
Keep every selected candidate within a stated non-negotiable limit.
Why these aims pull against each other
A penalty cannot promise the hard limit; a hard limit cannot use a fit improvement to justify exceeding it.
Choose an arrangement to see what changes and what remains difficult.
Arrows express the declared relations, not measured effect sizes. Examples and quantities are illustrative.
What this choice protects
What it costs
When it fits
Compare the arrangements
Charge for complexity
Minimize fitting error plus 1 times the stated complexity score.
- What it protects
- R can pay the penalty and retain its fit advantage.
- What it costs
- The selected model can exceed a desired size limit; the weight needs justification.
- When it fits
- The complexity preference is tradable, with predictive usefulness assessed independently.
Illustration note: The toy illustrates the mechanism only; it does not select a valid generalization weight or show held-out performance.
Exclude excess complexity
Require complexity at most 1, then compare fitting error among eligible candidates.
- What it protects
- The selected model obeys the declared boundary.
- What it costs
- A candidate with much better fit can be excluded.
- When it fits
- The limit is genuinely mandatory, such as an established representational budget, rather than merely a preference.
Illustration note: This counter-arrangement is a constraint, not another form of the source’s soft regularization.
What this illustration does—and does not—establish
The source supplies the stated tension; the selected arrangements are bounded editorial illustrations. Costs and conditions remain part of the comparison.
- All numbers are illustrative and the candidate set contains only S and R.
- Lower training error or lower penalized objective does not by itself establish better unseen-data performance.
- Changing the price and changing the allowable set are different operations; neither is universally preferable.
Source entries
Regularization
This source passage supplies the contextual tension. The concrete arrangements and schematic examples are editorial illustrations, not measured findings.
Soft Penalty versus Hard Constraint (the Boundary That Defines the Prime)
T3 — Soft Penalty versus Hard Constraint (the Boundary That Defines the Prime). Regularization is a tradable penalty, not a ban; a hard constraint forbids candidates outright and is a different structural object. The boundary is load-bearing because the frame's compression stays honest only by excluding untunable rules. The failure mode is calling any rule "regularization" — a flat prohibition, an inviolable limit — and importing the bias-variance intuitions that only apply to tradable penalties. The diagnostic is to ask whether a candidate can buy its way past the penalty at a price: if departure is forbidden rather than charged, it is a constraint, not regularization, and the tunable-weight reasoning does not transfer to it.
The source operation
Regularization is the structural move of adding a penalty on the complexity — the roughness, the norm, the deviation-from-prior — of a candidate solution to a fitting or optimization procedure, so that the solution chosen is one that trades data-fit against complexity according to an explicit, tunable weight.