Skip to content

Cross-entropy

The expected negative log probability that a model distribution q assigns to outcomes actually drawn from a distribution p.

Version
v1 · 2026-09-28 · History
Domain-specific #
8788
Domain group
Formal Sciences
Origin domain
Information Theory

Core Idea

Cross-entropy is the expected negative log probability that a model or coding distribution q assigns to outcomes drawn from p: \(H(p,q)=-\mathbb{E}_p[\log q]\). It equals \(H(p)+D_{KL}(p\|q)\), separating intrinsic uncertainty from mismatch, and its units depend on the logarithm base. Its coding interpretation is the expected message length when data follow p but the code is optimized for q.

How would you explain it like I'm…

Average Surprise Score

Imagine a friend guesses what your favorite snacks are, and they're surprised whenever you pick something they thought was unlikely. Cross-entropy is how surprised they are on average, over all the snacks you really pick. If their guesses match what you actually like, they're surprised less.

Average Surprise Score

Cross-entropy measures how well a guess about how often things happen matches what really happens. Suppose you really pick apples most days, but a computer model thinks you almost never do. Every time you pick an apple, the model gets a big "surprise score." Cross-entropy averages that surprise score over what really happens. Even a perfect guess still has some surprise if what happens is naturally random, but a worse guess adds extra surprise on top. That's why computers learn by trying to make this number smaller.

Expected Log-Loss Under a Model

Cross-entropy compares two probability distributions: p, how the data actually behaves, and q, what a model predicts. It's H(p,q) = the average, over outcomes drawn from p, of minus the log of the probability q gives that outcome. So outcomes are weighted by how often they really happen, but their "cost" comes from the model. Using log base 2 gives the answer in bits; the natural log gives nats. In coding terms, it's the average message length if you use a code designed for q to send data that actually follows p. It equals the entropy of p, the unavoidable uncertainty, plus the KL divergence, the extra cost of using the wrong model. In practice, you estimate it by averaging minus log q over real data, and minimizing that is the same as maximizing the model's likelihood, which is why it's the standard loss for classification.

 

Cross-entropy between a data distribution p and a model distribution q is H(p,q) = -E_p[log q(X)]: outcomes are weighted by their frequency under p but charged a cost determined by the probability q assigns them. The logarithm base sets the unit, bits for base 2 and nats for the natural log. Its coding interpretation is the expected length of messages when data follow p but the code is optimized for q. The decomposition H(p,q) = H(p) + D_KL(p||q) separates the irreducible uncertainty of p, its entropy, from the excess cost of the mismatch between p and q, the Kullback-Leibler divergence. Consequently H(p,q) is minimized over q exactly when q = p, where it equals H(p), not zero. When p is unknown, the empirical average of -log q(x_i) over observed data estimates cross-entropy. Under standard independent sampling, minimizing that sample average is equivalent to maximizing the model's likelihood, which is why cross-entropy, also called log loss, is the standard training objective for classification.

Scope of Application

The functional applies wherever a probability model q is evaluated against outcomes distributed as p or sampled from it. The functional applies wherever probability predictions are scored logarithmically against representative outcomes.

  • Source coding. It measures expected length under a mismatched code.
  • Language modeling. Average token surprise evaluates model predictions on held-out text.
  • Classification. One-hot labels and predicted class probabilities yield log loss.
  • Maximum likelihood. Minimizing empirical cross-entropy maximizes fitted probability of observed data.
  • Distribution fitting. With fixed p, reducing cross-entropy also reduces forward KL divergence.

Clarity

Write the ordered pair and formula, identify which distribution generates data, state the log base, and check q’s support wherever p is positive. Distinguish population expectation from its finite-sample estimate and cross-entropy from the overloaded joint-entropy notation. The closest near miss sets the boundary: KL divergence is the nearest formulaic near miss: DKL(p∥q)=H(p,q)−H(p), so the two differ by p’s entropy when p is fixed.

Manages Complexity

Cross-entropy compresses an entire pattern of predictive probabilities into one expected logarithmic cost while retaining severe penalties for confident errors. The entropy-plus-divergence decomposition separates intrinsic data uncertainty from model mismatch and makes coding, likelihood, and prediction views mutually translatable. The central fit to observed events–probability calibration tradeoff is this: Log loss rewards correct ranking but also strongly evaluates confidence. A second finite sample–population expectation tension matters because Empirical cross-entropy is noisy and can be biased by sample selection.

Abstract Reasoning

Use three linked moves: align p and q on the same event space and verify absolute-continuity or support requirements; convert each q-probability to a negative log cost with a declared base; average under p or an independent sample representing p. As a collapse test, the case exits when the expectation uses the wrong distribution, q is not a valid density on p’s support, or the logarithmic score is replaced. A fourth check is to use H(p)+DKL(p∥q) to separate uncertainty and mismatch.

Knowledge Transfer

The functional transfers literally across coding, statistics, and machine learning whenever calibrated probability assignments are scored logarithmically against data. ‘Cross-entropy’ used for a software loss should still expose its label and prediction distributions. Expectation and logarithmic scoring carry broader structure, but no canonical parent is asserted. No canonical parent prime is currently asserted; broader structural comparisons remain related-prime analogies until separately adjudicated in the DAG. p supplies the weighting measure.

Neighborhood in Abstraction Space

Cross-entropy sits in a moderately populated region (59th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.

Family — Empirical Measurement & Statistical Inference Methods (50 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08