Jensen's Inequality¶
For a convex function, the function of a mean is no greater than the mean of the function's values, under the required domain and expectation conditions.
Core Idea¶
Jensen's inequality compares two orders of operation on the same weighted or random input. If \(X\) has a defined mean in a convex domain and \(f\) is convex there, with the relevant expectations meaningful, then
For a concave function the direction reverses. In the two-point case, this is the chord condition \(f(\theta x+(1-\theta)y)\leq\theta f(x)+(1-\theta)f(y)\) for \(0\leq\theta\leq1\); an expectation extends the same comparison to a probability-weighted collection. The theorem concerns average first, transform second versus transform first, average second. It does not say that every nonlinear function has the same sign of gap.[1][2]
The Jensen gap \(E[f(X)]-f(E[X])\) is nonnegative in the convex case, but its magnitude is not universally a curvature-times-variance formula. Equality is not limited to constant \(X\) or globally affine \(f\): an ordinary convex function can be affine along values that carry the probability mass. Under Duchi's stricter convexity-at-the-mean condition, equality implies \(X=E[X]\) almost surely. These distinctions matter when a weak bound is promoted to a strict one.[2]
Structural Signature¶
Sig role-phrases: mean-bearing input → convex map on admissible domain → two operation orders → directed comparison.
- Mean-bearing input. A probability law or normalized nonnegative weights supply \(E[X]\); \(X\) and \(f(X)\) must be meaningful under the stated domain and integrability conditions. Negative coefficients or an undefined mean do not fit the ordinary expectation theorem.[1][2]
- Convex map on the relevant domain. The graph lies below its chords for the mixtures actually used. A concave map yields the reversed inequality, while mere nonlinearity provides no fixed direction.[1]
- Two operation orders. The same \(f\) and law appear in \(f(E[X])\) and \(E[f(X)]\). Changing the distribution between sides or substituting an unrelated average would not instantiate the theorem.[2]
- Directed comparison. Convexity bounds transform-of-mean from above by mean-of-transform. It is the order relation, not a measured gap or a claimed equality condition, that completes the identity.[1]
For finite weighted inputs the roles are identical with \(E\) replaced by a normalized weighted sum. Variance approximations, risk premia and entropy bounds are consequences or applications; none is needed to state this inequality.[1][3][4]
What It Is Not¶
- Not convexity itself. Live prime Convexity is the chord/mixture property of a set or function. Jensen's inequality is the derived mean–transformation comparison when that property is applied to weighted inputs.[1]
- Not a claim that every nonlinearity increases an average. Concavity reverses the sign; a function lacking the relevant global shape can change direction across distributions. The function's domain and the input support must be checked.[1]
- Not necessarily a strict gap. If \(f\) is affine over the relevant supported mixtures, equality can hold for nonconstant \(X\). The strict-convexity condition permits the stronger constant-input equality criterion.[2]
- Not a monetary risk premium. A concave utility comparison is measured in utility units. MIT defines the certainty equivalent by applying the inverse increasing utility function; only then can an expected-wealth-minus-certainty-equivalent amount be expressed in wealth units.[3]
- Not Jensen's alpha. The live finance statistic evaluates abnormal investment performance against a benchmark model; it does not test mean/transform order under convexity.
Scope of Application¶
The theorem applies to finite convex combinations and to expectation under a probability law when the random input, domain, and transformed expectation satisfy the needed conditions. Stanford's EE364a slides state both the two-point and expectation forms; Duchi's exercises make the mean-domain and strict equality qualifications explicit.[1][2]
In expected-utility analysis, MIT considers a lottery of wealth outcomes and an increasing concave utility function \(u\). Jensen gives \(E[u(X)]\leq u(E[X])\). This is one mathematical route to comparing a lottery with certain expected wealth under the specified preference representation. The source then defines a certainty equivalent \(u^{-1}(E[u(X)])\); the extra inverse-utility step matters, and the theorem alone is not individual financial advice.[3]
In information theory, Cover and Thomas prove relative entropy nonnegative. On the positive support of density \(p\), concavity of \(\log\) gives \(E_p[\log(q/p)]\leq\log E_p[q/p]=\log\int_{p>0}q\leq0\) where ratios/logs are finite. Thus \(D(p\Vert q)=E_p[-\log(q/p)]\geq0\). If \(q\) is zero on positive \(p\) mass, the divergence is infinite by the usual extended-value convention, not a finite Jensen calculation. This support distinction is an operating boundary, not a defect of the theorem.[4][2]
Clarity¶
The inequality resolves a common ambiguity: the response to an average input and the average response need not be equal. Identify the law or weights, the function, and whether it is convex or concave across the relevant range. Merely observing that an output varies with input does not determine which side is larger.[1]
Keep the units and equality claim visible. A Jensen utility gap is not wealth, and a KL divergence can be extended-valued when supports mismatch. A zero gap does not prove the input is constant unless strict-convexity conditions rule out affine behavior on the supported values.[2][3][4]
Manages Complexity¶
A potentially long sum or integral of nonlinear values receives one directional bound through the function's chord property and the weights' normalization. The analyst need not evaluate the full expectation to learn that the convex transform of the mean is a lower bound. This is powerful but deliberately coarse: the theorem supplies direction, not an exact error estimate.[1]
The same compact rule replaces different domain-specific calculations: MIT's concave-utility comparison and Cover–Thomas's log-ratio proof use different carriers and interpretations, yet both reduce to weighted convex/concave averaging. The reduction is invalid if one forgets that the probabilities, domain and finite/extended-value conventions differ between those examples.[3][4]
Abstract Reasoning¶
First establish a probability law or normalized weights and a convex domain containing the supported inputs and their mean. Then establish convexity of \(f\) on that domain and verify that \(E[f(X)]\) is meaningful. Only then conclude \(f(E[X])\leq E[f(X)]\). For a concave \(u\), apply the same argument to \(-u\) to reverse the comparison.[1][2]
If a stronger conclusion is needed, test the support and strictness. Under strict convexity at the mean in Duchi's setting, equality forces an almost-sure constant input; without that condition, a flat/affine segment may absorb variation without a gap. In the KL proof, separately check where the reference density vanishes before writing finite log ratios.[2][4]
Knowledge Transfer¶
Literal transfer from utility to information theory consists of keeping the same abstract comparison roles: a weighted input, a convex or concave map, and two orders of averaging. Utility uses wealth and a concave preference representation; relative entropy uses a log-ratio under one distribution. Neither interpretation transfers wholesale to the other.[3][4]
Live prime Convexity provides the more portable chord property. Jensen's named theorem remains a specific probabilistic/weighted inequality; applying it in several fields does not make it identical to all mixture-preserving phenomena. Its staged root status reflects a live DAG quality issue, not a claim that Convexity is irrelevant.[1]
Examples¶
MIT wealth lottery. Take a lottery \(F\) with finite expected wealth and an increasing concave \(u\) on the outcome range. Mapped back: mean-bearing input = the normalized lottery law and \(E_F[X]\); convex map on admissible domain = \(-u\) (or concave \(u\) with reversed direction); two operation orders = \(u(E_F X)\) and \(E_Fu(X)\) for the same \(F\); directed comparison = \(E_Fu(X)\leq u(E_FX)\). MIT's certainty equivalent requires the additional inverse \(u^{-1}\) and is not simply the utility gap.[3]
Cover–Thomas relative entropy. Let \(p,q\) be probability densities and work on positive \(p\) support where the finite log ratio is defined. Mapped back: mean-bearing input = expectation under \(p\) of \(q/p\); convex map = \(-\log\) on positive ratios; two operation orders = \(-\log E_p(q/p)\) and \(E_p[-\log(q/p)]\); directed comparison = \(D(p\Vert q)\geq-\log\int_{p>0}q\geq0\). If \(q=0\) on positive \(p\) mass, treat \(D\) as infinite instead of pretending the finite proof applies unchanged.[4]
Boundary: Jensen's alpha. The finance statistic may have a name and an average, but its defining comparison is abnormal performance relative to a model, not the convex mean/transform inequality.
Structural Tensions¶
Broad weak bound versus strict equality diagnosis. Convexity supplies an inequality over a large class of distributions, but it may be equality on an affine supported region. Adding strict convexity narrows applicability yet permits a sharper constant-input conclusion. Leaning too broad overclaims a positive gap; leaning too narrow discards valid weak bounds. Diagnostic: Is the function strictly convex at the relevant mean, or flat over supported mixtures?[2]
Utility comparison versus wealth interpretation. Direct Jensen reasoning is short and assumption-light once expected utility is specified, but its difference has utility units. An inverse-utility certainty equivalent yields a wealth comparison but requires monotonic invertibility and the preference model. The first is not automatically the second. Diagnostic: Was the certainty equivalent actually computed, or was a utility gap relabeled as money?[3]
Finite ratio algebra versus singular support. The KL proof is compact when positive \(p\) support has valid finite \(q/p\) ratios. Ignoring zeros in \(q\) can make the logarithm undefined; handling them as an extended infinite divergence is sound but precludes the same finite algebra on those points. Diagnostic: Is \(q>0\) wherever \(p>0\), or is an extended-value case present?[4]
Structural–Framed Character¶
Evaluative weight: The inequality is a mathematical order relation, not a value judgment. Risk aversion is one interpretation supplied by a preference model; the theorem itself does not decide whether risk is good or bad.[1][3]
Human-practice dependence: Choosing a lottery, likelihood model or function is an analyst's practice, but the theorem's validity follows from the weights and convexity once specified. A particular finance or coding community does not create the relation.[1][4]
Institutional origin: Jensen's name and disciplinary teaching are historical labels; no institution makes a candidate function convex or a probability law normalized. Usage conventions can vary while the proof obligation remains.[1]
Vocabulary travel: The theorem travels literally between economics and information theory because the same mean/transform relation is verified, though “risk premium” and “relative entropy” do not travel as synonyms. This is technical mathematical reach rather than unrestricted cross-domain prime breadth.[3][4]
Import versus recognition: In a new setting the inequality must be imported with its hypotheses—convexity on the support, valid weights, and expectations. Seeing a curved response and an average is insufficient without those checks.[2]
Its character: predominantly structural within probability and convex analysis, but domain-specific as the named weighted-average theorem. Live prime Convexity carries the more portable skeleton; a DAG edge awaits repair of that prime's current Optimization ancestry.[1]
Structural Core vs. Domain Accent¶
Skeletal relation: Live prime Convexity states chord dominance/mixture geometry, which supplies the theorem's indispensable mathematical shape. A composition/presupposes relationship is semantically plausible.
Domain-bound mechanism: Jensen's rule adds a normalized weighted input or probability law, a convex/concave function on the relevant domain, and a comparison of mean-before-transform against transform-before-mean. Utility's wealth outcomes and KL's density ratios are domain accents; neither is a role required in every instance.[3][4]
Why not prime: The named inequality's exact identity needs mathematical averaging and convex function structure. Broad metaphors such as “nonlinearity matters” do not preserve its assumptions or direction. The reusable shape belongs to Convexity; the theorem remains a specialized rigorous consequence even though it is useful across fields.[1][2]
Instantiates / Related Primes¶
Convexity is the live conceptual prerequisite: its Core Idea explicitly describes chord dominance and Jensen-style averaging. This entry is not a kind of convexity—one is an inequality theorem, the other a property of a set/function—so strict subsumption would be wrong. A composition/presupposes edge might be warranted after repair of Convexity's current edge to Optimization; for now the staged graph has no asserted parent and preserves the issue in metadata.
Expected Utility is one application framework, not a necessary parent for a theorem also used in information theory. Live Jensen's alpha is a lexical surname neighbor only, not a mathematical subtype.
Neighborhood in Abstraction Space¶
Jensen's Inequality sits in a sparse region of the domain-specific corpus (70th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Foundations of Probability & Inference (29 abstractions)
Nearest neighbors
- Invex Function — 0.84
- Maharam Algebra — 0.84
- Closed Linear Operator — 0.83
- Riemann–Liouville integral — 0.83
- Monotone Likelihood Ratio Property — 0.83
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
Convexity: Check whether the claim merely says a function's graph lies below its chords or whether it compares expectation orders. The former is a property; the latter is Jensen's inequality.[1]
A curvature approximation: An estimate proportional to variance can be useful under smoothness and small-spread conditions, but Jensen itself supplies only the direction of the gap. It works for nonsmooth convex functions as well.[2]
Risk premium: A monetary premium requires a certainty-equivalent calculation from an increasing invertible utility function. The raw inequality compares utility values.[3]
Gibbs/KL inequality: Relative-entropy nonnegativity is a particular application derived with log ratios and support rules, not the general Jensen statement.[4]
References¶
[1] Stephen Boyd and Lieven Vandenberghe, Convex Optimization I, Stanford EE364a lecture slides, Convex Functions 3–12, PDF p. 49 (two-point and expectation forms). registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r
[2] John Duchi, Exercises for Theory of Statistics, Stanford Stats300b (Winter 2021), §2 Questions 2.1(e)–2.2, PDF p. 7 (general and strict-equality hypotheses; KL exercise). registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n
[3] Alexander Wolitzky, Lecture 9: Attitudes toward Risk, MIT 14.121 (Fall 2015), PDF pp. 3–6 (concave utility and certainty equivalents). registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l
[4] Thomas M. Cover and Joy A. Thomas, “Determinant Inequalities via Information Theory”, SIAM Journal on Matrix Analysis and Applications 9(3) (1988), §2 Lemma 1, PDF pp. 1–2 (relative-entropy nonnegativity via Jensen). registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l