Skip to content

Cramér's Theorem (Large Deviations)

An i.i.d.-sample-mean large-deviation principle whose exponential rate is the convex conjugate of the one-observation log moment-generating function.

Version
v1 · 2026-10-03 · History
Domain-specific #
13105
Domain group
Formal Sciences
Origin domain
Mathematics
Subdomains
Probability Theory, Large Deviations → Mathematics
Aliases
Cramér's theorem for empirical means, Cramér large-deviation theorem

Core Idea

Cramér's theorem gives the large-deviation principle for averages \(\overline X_n\) of i.i.d. real observations when the one-observation log moment-generating function \(\Lambda(t)=\log E[e^{tX_1}]\) is finite near zero. The exponential rate is its convex conjugate \(I(x)=\sup_t\{tx-\Lambda(t)\}\). For a Borel set of possible averages, the lower logarithmic probability bound uses the set's interior and the upper bound uses its closure. A simple limit \(n^{-1}\log P(\overline X_n\ge a)\to-I(a)\) follows for a threshold above the mean when \(I\) is finite and continuous around that threshold; it is not an unconditional formula for every event.[ref-02a40cce71a5][ref-d820259d1eab]

Scope of Application

Independent Bernoulli trials give an empirical success fraction with binary-relative-entropy rate. Independent standard Gaussian observations give a sample-average rate \(I(x)=x^2/2\). Both satisfy the theorem's moment condition, but the numerical rates and finite-sample probabilities differ. A fixed-horizon sum of modeled claims or arrivals could qualify if its assumptions are checked; ultimate ruin or all-time overflow is a different event and needs a separate theorem. Cramér's theorem concerns a leading logarithmic rate, not the exact finite-\(n\) tail or its prefactor.[ref-02a40cce71a5][ref-55b77c843ad1]

Clarity

State the summand law, verify i.i.d. observations and a log MGF finite near zero, construct its convex conjugate, and identify the event for \(\overline X_n\). Then distinguish the theorem's open/closed-set bounds from a one-line equality: the equality requires the relevant interior and closure rate infima to coincide. A Chernoff calculation alone is only an upper bound; the matching lower bound is the theorem's extra force. Here “exponential rate” means \(n^{-1}\log P\) convergence, not probability divided by \(e^{-nI}\) tending to one.[ref-02a40cce71a5][ref-d820259d1eab][^ref-55b77c843ad1]

Manages Complexity

The theorem turns an entire sequence of increasingly rare empirical-mean events into a deterministic rate function derived from a single observation's distribution. The event's least-cost point controls its exponent when the bounds meet. This compression makes unlike distributions comparable, yet deliberately omits finite-sample prefactors and warns that irregular event boundaries can prevent a single limit.[ref-02a40cce71a5][ref-d820259d1eab]

Abstract Reasoning

The reusable procedure is: identify i.i.d. real summands; check an exponential moment near zero; form \(\Lambda\) and \(I=\Lambda^*\); apply lower and upper large-deviation bounds to the actual event; only then simplify to a threshold rate if the boundary is regular. For Bernoulli\((p)\) success fractions above \(q\) with \(p<q<1\), this yields \(I(q)=q\log(q/p)+(1-q)\log((1-q)/(1-p))\). For a standard Gaussian average above \(b>0\), it yields \(b^2/2\); a separate Gaussian calculation reveals an additional order-\(n^{-1/2}\) prefactor.[ref-02a40cce71a5][ref-55b77c843ad1]

Knowledge Transfer

The construction transfers between discrete and continuous i.i.d. laws because the roles remain the same: observation law, empirical mean, log MGF, convex-conjugate rate and event bounds. It does not transfer automatically to dependent samples, heavy-tailed mechanisms, first-passage events, or a named application merely because they feature rare events. The live Convex Conjugate entry supplies the proposed compositional DAG parent; the theorem adds the domain-specific probability hypotheses and asymptotic conclusion. Neither the broader Large Deviations Theory field nor a one-sided concentration bound duplicates this identity.[ref-02a40cce71a5][ref-d820259d1eab]

[^ref-02a40cce71a5]: Timo Seppäläinen, Translation Invariant Exclusion Processes (2005), Appendix A.7, printed pp. 192–193, Theorem A.9 and Exercise A.9(a). Author-text PDF. [^ref-d820259d1eab]: MIT 6.265/15.070J, Lecture 3: Large Deviations Theory. Cramer's Theorem (2013), printed pp. 1–3. Official course PDF. [^ref-55b77c843ad1]: Jonathan Goodman, Variance Reduction (2005), §4.1, printed pp. 6–7. Author-course PDF.

Relationships to Other Abstractions

Local relationship map for Cramér's Theorem (Large Deviations)Parents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Cramér's Theorem(Large Deviations)DOMAINDomain-specific abstraction: Convex conjugate — presupposesConvex conjugateDOMAIN

Current abstraction Cramér's Theorem (Large Deviations) Domain-specific

Parents (1) — more general patterns this builds on

  • Cramér's Theorem (Large Deviations) presupposes Convex conjugate Domain-specific

    Cramér's rate is defined as the convex conjugate of the one-observation log MGF.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Cramér's Theorem (Large Deviations) sits in a moderately populated region (50th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.

Family — Foundations of Probability & Inference (29 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08