Cross-entropy¶
The expected negative log probability that a model distribution q assigns to outcomes actually drawn from a distribution p.
Core Idea¶
Cross-entropy between a data distribution \(p\) and a model distribution \(q\) is \(H(p,q)=-\mathbb{E}_p[\log q(X)]\). Outcomes are weighted by how often they occur under p, while their cost is determined by the probability q assigns them. The logarithm’s base sets units: base 2 gives bits and the natural logarithm gives nats.
Its coding interpretation is the expected message length when data follow p but the code is optimized for q. The identity \(H(p,q)=H(p)+D_{KL}(p\|q)\) separates irreducible uncertainty in p from the excess mismatch cost of using q.
When p is unknown, the empirical average of \(-\log q(x_i)\) over data estimates cross-entropy. Minimizing that sample quantity is equivalent to maximizing model likelihood under standard independent sampling, which explains cross-entropy or log loss in classification.
How would you explain it like I'm…
Average Surprise Score
Average Surprise Score
Expected Log-Loss Under a Model
Structural Signature¶
Sig role-phrases:
- Data distribution p. Supplies the outcome frequencies and the expectation measure. Constitutive reference for what actually occurs. If altered: Taking expectation under q computes a different quantity.
- Model or coding distribution q. Assigns probabilities and therefore code lengths or predictive penalties to outcomes. Constitutive evaluated representation. If altered: If q assigns zero probability to a p-possible outcome, the cost diverges.
- Negative log score. Turns q-probability into surprise or code length with units set by the log base. Identity-bearing local cost. If altered: Replacing log score with another loss produces another scoring functional.
- Expectation or sample average. Aggregates local costs under p or empirical observations. Identity-bearing global functional and estimator. If altered: A single unaveraged loss is an observation-level log loss, not the population cross-entropy alone.
What It Is Not¶
- Not entropy alone. H(p) uses p for both expectation and code probabilities.
- Not KL divergence. KL subtracts H(p) and is zero, rather than H(p), when p=q.
- Not joint entropy. The notation H(p,q) can be overloaded, but cross-entropy evaluates one distribution with another.
- Not any error rate. Log loss evaluates assigned probabilities, not only the final class label.
Scope of Application¶
The functional applies wherever a probability model q is evaluated against outcomes distributed as p or sampled from it.
- Source coding. It measures expected length under a mismatched code.
- Language modeling. Average token surprise evaluates model predictions on held-out text.
- Classification. One-hot labels and predicted class probabilities yield log loss.
- Maximum likelihood. Minimizing empirical cross-entropy maximizes fitted probability of observed data.
- Distribution fitting. With fixed p, reducing cross-entropy also reduces forward KL divergence.
Clarity¶
Write the ordered pair and formula, identify which distribution generates data, state the log base, and check q’s support wherever p is positive. Distinguish population expectation from its finite-sample estimate and cross-entropy from the overloaded joint-entropy notation.
Manages Complexity¶
Cross-entropy compresses an entire pattern of predictive probabilities into one expected logarithmic cost while retaining severe penalties for confident errors. The entropy-plus-divergence decomposition separates intrinsic data uncertainty from model mismatch and makes coding, likelihood, and prediction views mutually translatable.
Abstract Reasoning¶
- Align p and q on the same event space and verify absolute-continuity or support requirements.
- Convert each q-probability to a negative log cost with a declared base.
- Average under p or an independent sample representing p.
- Use H(p)+DKL(p∥q) to separate uncertainty and mismatch.
- Interpret optimization carefully when p or q, rather than the other, is held fixed.
Knowledge Transfer¶
The functional transfers literally across coding, statistics, and machine learning whenever calibrated probability assignments are scored logarithmically against data. ‘Cross-entropy’ used for a software loss should still expose its label and prediction distributions. Expectation and logarithmic scoring carry broader structure, but no canonical parent is asserted.
Examples¶
Canonical¶
If p gives true symbol frequencies and q supplies code probabilities, average −log2 q(x) under p is the expected bits per symbol for the mismatched code.
Mapped back: data distribution p → true symbol frequencies; model or coding distribution q → code-implied probabilities; negative log score → bit length −log2 q; expectation or sample average → p-weighted mean.
Applied / In Practice¶
For one-hot classification labels, the per-example loss is the negative log probability assigned to the observed class, averaged over the test set.
Mapped back: data distribution p → empirical label distribution; model or coding distribution q → predicted class probabilities; negative log score → −log probability of the observed label; expectation or sample average → mean test log loss.
Structural Tensions¶
T1: fit to observed events vs. probability calibration. Log loss rewards correct ranking but also strongly evaluates confidence. Diagnostic: Are improvements due to more accurate labels, better calibration, or both?
T2: finite sample vs. population expectation. Empirical cross-entropy is noisy and can be biased by sample selection. Diagnostic: Does the evaluation sample represent p independently of training?
T3: zero probability vs. infinite penalty. A q-zero event with positive p mass makes population cost infinite. Diagnostic: What smoothing or support model prevents impossible confident exclusions?
Structural–Framed Character¶
Cross-entropy is strongly structural. Evaluative weight: it evaluates probability models using a proper logarithmic cost. Human-practice-bound: event modeling and log base are chosen, while the functional is mathematical. Institutional origin: information theory and statistics stabilize the definition. Vocabulary travels: coding, likelihood, and log loss are equivalent views under stated conditions. Import versus recognize: literal transfer requires distributions and logarithmic scoring. Its character: an ordered two-distribution expectation decomposing uncertainty and mismatch.
Structural Core vs. Domain Accent¶
Skeletal core. Outcomes from one source are scored by another representation and the local penalties are averaged under the source.
Domain-bound accent. Sources and representations are probability distributions, penalties are negative log probabilities, and entropy and KL divergence supply the decomposition.
Why not prime. Evaluation and mismatch are portable, but cross-entropy is a specific information-theoretic functional with probability and logarithmic structure.
Instantiates / Related Primes¶
- Expectation. p supplies the weighting measure.
- Measurement. The result summarizes average predictive or coding cost.
- Divergence. KL is the excess above p’s entropy.
- The root placement remains.
Neighborhood in Abstraction Space¶
Cross-entropy sits in a moderately populated region (59th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Empirical Measurement & Statistical Inference Methods (50 abstractions)
Nearest neighbors
- Bayesian Programming — 0.86
- M-Estimator — 0.85
- Risk Score — 0.85
- Uncertainty analysis — 0.85
- Bootstrapping populations — 0.84
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Entropy. Tell: Set q=p; otherwise cross-entropy includes model mismatch.
- KL divergence. Tell: Subtract H(p) from cross-entropy to obtain forward KL.
- Joint entropy. Tell: Check whether H(p,q) denotes two distributions or two random variables.
- Misclassification rate. Tell: Cross-entropy uses the entire predicted probability, not only the chosen class.
References¶
- Frozen Wikipedia discovery revision: https://en.wikipedia.org/wiki/Cross-entropy (revision 1367589881).
- Preserved source candidate: https://arxiv.org/pdf/2304.07288.pdf
- Preserved source candidate: https://books.google.com/books?id=jDrp4QEGioMC&dq=%22logarithmic+loss%22+%22log+loss%22&pg=PA82
- Preserved source candidate: https://scikit-learn.org/1.7/modules/generated/sklearn.metrics.log_loss.html
- Preserved source candidate: https://link.springer.com/article/10.1007/s10479-005-5724-z
The frozen Wikipedia revision is discovery provenance. The retained source set was reviewed for identity, formal or operational relation, and scope. The encyclopedia's structural synthesis is bounded to those claims; a thin authority surface is recorded as a nonblocking source-strengthening repair rather than concealed.