Skip to content

Cross-entropy

The expected negative log probability that a model distribution q assigns to outcomes actually drawn from a distribution p.

Version
v1 · 2026-09-28 · History
Domain-specific #
8788
Domain group
Formal Sciences
Origin domain
Information Theory

Core Idea

Cross-entropy between a data distribution \(p\) and a model distribution \(q\) is \(H(p,q)=-\mathbb{E}_p[\log q(X)]\). Outcomes are weighted by how often they occur under p, while their cost is determined by the probability q assigns them. The logarithm’s base sets units: base 2 gives bits and the natural logarithm gives nats.

Its coding interpretation is the expected message length when data follow p but the code is optimized for q. The identity \(H(p,q)=H(p)+D_{KL}(p\|q)\) separates irreducible uncertainty in p from the excess mismatch cost of using q.

When p is unknown, the empirical average of \(-\log q(x_i)\) over data estimates cross-entropy. Minimizing that sample quantity is equivalent to maximizing model likelihood under standard independent sampling, which explains cross-entropy or log loss in classification.

How would you explain it like I'm…

Average Surprise Score

Imagine a friend guesses what your favorite snacks are, and they're surprised whenever you pick something they thought was unlikely. Cross-entropy is how surprised they are on average, over all the snacks you really pick. If their guesses match what you actually like, they're surprised less.

Average Surprise Score

Cross-entropy measures how well a guess about how often things happen matches what really happens. Suppose you really pick apples most days, but a computer model thinks you almost never do. Every time you pick an apple, the model gets a big "surprise score." Cross-entropy averages that surprise score over what really happens. Even a perfect guess still has some surprise if what happens is naturally random, but a worse guess adds extra surprise on top. That's why computers learn by trying to make this number smaller.

Expected Log-Loss Under a Model

Cross-entropy compares two probability distributions: p, how the data actually behaves, and q, what a model predicts. It's H(p,q) = the average, over outcomes drawn from p, of minus the log of the probability q gives that outcome. So outcomes are weighted by how often they really happen, but their "cost" comes from the model. Using log base 2 gives the answer in bits; the natural log gives nats. In coding terms, it's the average message length if you use a code designed for q to send data that actually follows p. It equals the entropy of p, the unavoidable uncertainty, plus the KL divergence, the extra cost of using the wrong model. In practice, you estimate it by averaging minus log q over real data, and minimizing that is the same as maximizing the model's likelihood, which is why it's the standard loss for classification.

 

Cross-entropy between a data distribution p and a model distribution q is H(p,q) = -E_p[log q(X)]: outcomes are weighted by their frequency under p but charged a cost determined by the probability q assigns them. The logarithm base sets the unit, bits for base 2 and nats for the natural log. Its coding interpretation is the expected length of messages when data follow p but the code is optimized for q. The decomposition H(p,q) = H(p) + D_KL(p||q) separates the irreducible uncertainty of p, its entropy, from the excess cost of the mismatch between p and q, the Kullback-Leibler divergence. Consequently H(p,q) is minimized over q exactly when q = p, where it equals H(p), not zero. When p is unknown, the empirical average of -log q(x_i) over observed data estimates cross-entropy. Under standard independent sampling, minimizing that sample average is equivalent to maximizing the model's likelihood, which is why cross-entropy, also called log loss, is the standard training objective for classification.

Structural Signature

Sig role-phrases:

  • Data distribution p. Supplies the outcome frequencies and the expectation measure. Constitutive reference for what actually occurs. If altered: Taking expectation under q computes a different quantity.
  • Model or coding distribution q. Assigns probabilities and therefore code lengths or predictive penalties to outcomes. Constitutive evaluated representation. If altered: If q assigns zero probability to a p-possible outcome, the cost diverges.
  • Negative log score. Turns q-probability into surprise or code length with units set by the log base. Identity-bearing local cost. If altered: Replacing log score with another loss produces another scoring functional.
  • Expectation or sample average. Aggregates local costs under p or empirical observations. Identity-bearing global functional and estimator. If altered: A single unaveraged loss is an observation-level log loss, not the population cross-entropy alone.

What It Is Not

  • Not entropy alone. H(p) uses p for both expectation and code probabilities.
  • Not KL divergence. KL subtracts H(p) and is zero, rather than H(p), when p=q.
  • Not joint entropy. The notation H(p,q) can be overloaded, but cross-entropy evaluates one distribution with another.
  • Not any error rate. Log loss evaluates assigned probabilities, not only the final class label.

Scope of Application

The functional applies wherever a probability model q is evaluated against outcomes distributed as p or sampled from it.

  • Source coding. It measures expected length under a mismatched code.
  • Language modeling. Average token surprise evaluates model predictions on held-out text.
  • Classification. One-hot labels and predicted class probabilities yield log loss.
  • Maximum likelihood. Minimizing empirical cross-entropy maximizes fitted probability of observed data.
  • Distribution fitting. With fixed p, reducing cross-entropy also reduces forward KL divergence.

Clarity

Write the ordered pair and formula, identify which distribution generates data, state the log base, and check q’s support wherever p is positive. Distinguish population expectation from its finite-sample estimate and cross-entropy from the overloaded joint-entropy notation.

Manages Complexity

Cross-entropy compresses an entire pattern of predictive probabilities into one expected logarithmic cost while retaining severe penalties for confident errors. The entropy-plus-divergence decomposition separates intrinsic data uncertainty from model mismatch and makes coding, likelihood, and prediction views mutually translatable.

Abstract Reasoning

  1. Align p and q on the same event space and verify absolute-continuity or support requirements.
  2. Convert each q-probability to a negative log cost with a declared base.
  3. Average under p or an independent sample representing p.
  4. Use H(p)+DKL(p∥q) to separate uncertainty and mismatch.
  5. Interpret optimization carefully when p or q, rather than the other, is held fixed.

Knowledge Transfer

The functional transfers literally across coding, statistics, and machine learning whenever calibrated probability assignments are scored logarithmically against data. ‘Cross-entropy’ used for a software loss should still expose its label and prediction distributions. Expectation and logarithmic scoring carry broader structure, but no canonical parent is asserted.

Examples

Canonical

If p gives true symbol frequencies and q supplies code probabilities, average −log2 q(x) under p is the expected bits per symbol for the mismatched code.

Mapped back: data distribution p → true symbol frequencies; model or coding distribution q → code-implied probabilities; negative log score → bit length −log2 q; expectation or sample average → p-weighted mean.

Applied / In Practice

For one-hot classification labels, the per-example loss is the negative log probability assigned to the observed class, averaged over the test set.

Mapped back: data distribution p → empirical label distribution; model or coding distribution q → predicted class probabilities; negative log score → −log probability of the observed label; expectation or sample average → mean test log loss.

Structural Tensions

T1: fit to observed events vs. probability calibration. Log loss rewards correct ranking but also strongly evaluates confidence. Diagnostic: Are improvements due to more accurate labels, better calibration, or both?

T2: finite sample vs. population expectation. Empirical cross-entropy is noisy and can be biased by sample selection. Diagnostic: Does the evaluation sample represent p independently of training?

T3: zero probability vs. infinite penalty. A q-zero event with positive p mass makes population cost infinite. Diagnostic: What smoothing or support model prevents impossible confident exclusions?

Structural–Framed Character

Cross-entropy is strongly structural. Evaluative weight: it evaluates probability models using a proper logarithmic cost. Human-practice-bound: event modeling and log base are chosen, while the functional is mathematical. Institutional origin: information theory and statistics stabilize the definition. Vocabulary travels: coding, likelihood, and log loss are equivalent views under stated conditions. Import versus recognize: literal transfer requires distributions and logarithmic scoring. Its character: an ordered two-distribution expectation decomposing uncertainty and mismatch.

Structural Core vs. Domain Accent

Skeletal core. Outcomes from one source are scored by another representation and the local penalties are averaged under the source.

Domain-bound accent. Sources and representations are probability distributions, penalties are negative log probabilities, and entropy and KL divergence supply the decomposition.

Why not prime. Evaluation and mismatch are portable, but cross-entropy is a specific information-theoretic functional with probability and logarithmic structure.

  • Expectation. p supplies the weighting measure.
  • Measurement. The result summarizes average predictive or coding cost.
  • Divergence. KL is the excess above p’s entropy.
  • The root placement remains.

Neighborhood in Abstraction Space

Cross-entropy sits in a moderately populated region (59th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.

Family — Empirical Measurement & Statistical Inference Methods (50 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • Entropy. Tell: Set q=p; otherwise cross-entropy includes model mismatch.
  • KL divergence. Tell: Subtract H(p) from cross-entropy to obtain forward KL.
  • Joint entropy. Tell: Check whether H(p,q) denotes two distributions or two random variables.
  • Misclassification rate. Tell: Cross-entropy uses the entire predicted probability, not only the chosen class.

References

  • Frozen Wikipedia discovery revision: https://en.wikipedia.org/wiki/Cross-entropy (revision 1367589881).
  • Preserved source candidate: https://arxiv.org/pdf/2304.07288.pdf
  • Preserved source candidate: https://books.google.com/books?id=jDrp4QEGioMC&dq=%22logarithmic+loss%22+%22log+loss%22&pg=PA82
  • Preserved source candidate: https://scikit-learn.org/1.7/modules/generated/sklearn.metrics.log_loss.html
  • Preserved source candidate: https://link.springer.com/article/10.1007/s10479-005-5724-z

The frozen Wikipedia revision is discovery provenance. The retained source set was reviewed for identity, formal or operational relation, and scope. The encyclopedia's structural synthesis is bounded to those claims; a thin authority surface is recorded as a nonblocking source-strengthening repair rather than concealed.