Skip to content

Information Entropy

The probability-weighted average of logarithmic surprise across the outcomes of a discrete random variable.

Version
v1 · 2026-10-03 · History
Domain-specific #
13330
Domain group
Formal Sciences
Origin domain
Information Theory
Subdomains
Discrete Information Theory, Probability → Information Theory
Aliases
Shannon Entropy

Core Idea

Information entropy is the average uncertainty of a finite discrete random variable. For outcomes \(x\) with probabilities \(p(x)\), it averages each outcome's logarithmic surprise:

\[H_b(X)=-\sum_x p(x)\log_b p(x)=\mathbb{E}[-\log_b p(X)],\]

where \(b>1\) and \(0\log_b0\) is taken as zero. Base two measures entropy in bits. The value is zero when one outcome is certain and highest, \(\log_b m\), when \(m\) outcomes are equally likely. It belongs to the probability distribution, not to any one observed outcome.[^ref-99a36c426d12]

The same formula can summarize uncertainty about a source's next symbol or diversity of taxon identities in an ecological community. Its numerical calculation transfers, but the interpretation and the premises for downstream claims change.[ref-99a36c426d12][ref-299144d3cd60]

Scope of Application

For a memoryless symbol source, \(H(X)\) is its entropy per symbol and is connected to the limiting performance of suitable lossless coding. For a source whose symbols depend on earlier state, a one-symbol marginal \(H(X_t)\) need not equal the process's entropy rate; Shannon treats the finite-state source rate using state-conditional uncertainty.[^ref-99a36c426d12]

For ecological alpha diversity, \(p(x)\) can be the true relative abundance of taxon \(x\) and \(X\) the taxon of a randomly selected individual. The resulting Shannon index reflects both richness and evenness. Counts observed in a sample estimate, but do not automatically equal, the community's true abundance distribution. The index does not by itself measure relatedness among taxa.[^ref-299144d3cd60]

Clarity

Entropy distinguishes a distribution's expected surprise from the surprise of one rare event. It also forces the analyst to state which outcomes, probabilities and logarithm base are used. Without those choices, saying a source or community has “high entropy” is incomplete.[^ref-99a36c426d12]

It is not automatically conditional entropy, cross-entropy, a dependent process's entropy rate or thermodynamic entropy. Those quantities have additional or different carriers and assumptions; formula resemblance or shared vocabulary does not make them identical.[^ref-99a36c426d12]

Manages Complexity

One number compresses an entire probability vector into average logarithmic uncertainty. That makes comparisons concise, but it discards which outcomes carry probability mass: distinct distributions can have the same entropy. Ecological use also requires care because a count-based estimate can be affected by rare or unobserved taxa and the sampling model.[ref-99a36c426d12][ref-299144d3cd60]

Abstract Reasoning

Specify a finite outcome variable, assign a probability distribution, calculate \(-\log_b p(x)\) for each positive-probability outcome, and average these values under that same distribution. Compare results only after aligning outcome definitions and log units. For coding claims, additionally establish the source and lossless-coding premises; for an ecological claim, establish the community and abundance/sampling model.[ref-99a36c426d12][ref-299144d3cd60]

For example, a memoryless source with symbol probabilities \((3/4,1/4)\) has \(H_2\approx0.811\) bits per symbol, below the one-bit maximum for two equally likely symbols. A two-taxon community with the same true abundance shares has the same numerical Shannon index, but this does not mean the organisms transmit coded messages. If a balanced binary process alternates deterministically, its one-position marginal entropy is one bit while its long-run entropy rate is zero after the phase is known.[ref-99a36c426d12][ref-299144d3cd60]

Knowledge Transfer

The weighted-log functional transfers literally wherever a finite discrete outcome law is specified. Live Expected Value supplies its probability-weighted averaging operation and is the proposed composition parent, but the logarithmic score is the information-entropy residual. The communication-theory coding consequence does not transfer automatically to ecological diversity. Nor does the name establish equivalence with the live thermodynamic-entropy prime.

[^ref-99a36c426d12]: Claude E. Shannon, “A Mathematical Theory of Communication”, Bell System Technical Journal 27 (1948), Harvard-hosted reprint, §§6–7 and §9. [^ref-299144d3cd60]: Amy D. Willis and Bryan D. Martin, “Estimating Diversity in Networked Ecological Communities”, Biostatistics 23, no. 1 (2022): 207–222, §§1 and 2.1.1.

Relationships to Other Abstractions

Local relationship map for Information EntropyParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Information EntropyDOMAINPrime abstraction: Expected Value — presupposesExpected ValuePRIME

Current abstraction Information Entropy Domain-specific

Parents (1) — more general patterns this builds on

  • Information Entropy presupposes Expected Value Prime

    Shannon entropy presupposes probability-weighted expectation of outcome-level logarithmic surprise.

Hierarchy paths (3) — routes to 2 parentless roots

Neighborhood in Abstraction Space

Information Entropy sits in a sparse region of the domain-specific corpus (72nd percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Foundations of Probability & Inference (29 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08