Skip to content

Adjusted mutual information

By adopting a hypergeometric model of randomness, it can be shown that the expected mutual information between two random clusterings is.

Version
v1 · 2026-09-28 · History
Domain-specific #
7884
Domain group
Interdisciplinary & Synthetic
Origin domain
Data Science & Analytics
Subdomains
Clustering Evaluation, Machine Learning → Data Science & Analytics

Core Idea

Adjusted mutual information is treated here as the recurring computer science and information systems identity summarized by this source-grounded definition: By adopting a hypergeometric model of randomness, it can be shown that the expected mutual information between two random clusterings is. In probability theory and information theory, adjusted mutual information, a variation of mutual information may be used for comparing clusterings. It corrects the effect of agreement solely due to chance between clusterings, similar to the way the adjusted rand index corrects the Rand index.

How would you explain it like I'm…

Matching Minus Luck

Two kids each sort the same pile of toys into groups. You want to know how much their groupings match. But some matching would happen just by luck, so adjusted mutual information takes away the lucky part and only counts the real agreement.

Luck-Corrected Group Matching

Imagine two people sorting the same set of photos into piles, each person in their own way. Mutual information is a score for how much knowing one person's piles tells you about the other's. The trouble is that even random piles would match a little by chance. Adjusted mutual information subtracts the amount of matching you would expect from random sorting, so a high score really means the two sortings agree. Each photo goes in exactly one pile.

Chance-Corrected Clustering Agreement

Adjusted mutual information (AMI) is a way to compare two clusterings, meaning two ways of dividing the same items into non-overlapping groups. It starts from mutual information, which measures how much one grouping tells you about the other, computed from a table counting how many items land in each pair of groups. Raw mutual information is biased upward because two random clusterings still share some information by chance. AMI corrects this by subtracting the expected mutual information between random clusterings, calculated under a hypergeometric model of randomness, similar to how the adjusted Rand index corrects the Rand index. After this adjustment the score is no longer a true distance metric.

 

Adjusted mutual information is a chance-corrected measure for comparing two partitions U and V of the same objects. Cluster overlap is summarized in an R×C contingency table whose entry n_ij counts objects shared by clusters U_i and V_j, from which mutual information is computed. Under a hypergeometric model of randomness, the expected mutual information between two random clusterings can be derived, and AMI adjusts the observed mutual information by removing this expected chance agreement. This parallels the way the adjusted Rand index corrects the Rand index. It is closely related to variation of information: applying the analogous adjustment to VI makes it equivalent to AMI. The adjusted measure is no longer a metric, and the construction assumes hard (pairwise disjoint) clusters.

Scope of Application

  • Documented setting. In probability theory and information theory, adjusted mutual information, a variation of mutual information may be used for comparing clusterings.

  • Mutual information of two partitions. Given a set S of N elements S={s1, s2,\ldots sN} , consider two partitions of S, namely U={U1, U2,\ldots, UR} with R clusters, and V={V1, V2,\ldots.

  • Mutual information of two partitions. It is presumed here that the partitions are so-called hard clusters; the partitions are pairwise disjoint.

  • Mutual information of two partitions. Suppose an object is picked at random from S; the probability that the object falls into cluster Ui is.

  • Mutual information of two partitions. H(U) is non-negative and takes the value 0 only when there is no uncertainty determining an object's cluster membership, i.e., when there is only one cluster.

Clarity

A clear use of Adjusted mutual information names the carrier, the operative relation, and the conditions under which the source treats the identity as present. The minimal definition is By adopting a hypergeometric model of randomness, it can be shown that the expected mutual information between two random clusterings is.

Manages Complexity

Adjusted mutual information compresses multiple computer science and information systems details into a stable diagnostic relation. The source shows both the central mechanism—it quantifies the information shared by the two clusterings and thus can be employed as a clustering similarity measure.—and the practical consequence—h(U) is non-negative and takes the value 0 only when there is no uncertainty determining an object's cluster membership, i.e., when there.

Abstract Reasoning

  1. Type the carrier. Identify the computer science and information systems entities to which the claim applies.
  2. State the relation. Use the source-grounded identity: By adopting a hypergeometric model of randomness, it can be shown that the expected mutual information between two random clusterings is.
  3. Check operation and conditions. Given a set S of N elements S={s1, s2,\ldots sN} , consider two partitions of S, namely U={U1, U2,\ldots, UR} with R clusters, and V={V1, V2,\ldots, VC} with C clusters. 4.

Knowledge Transfer

Within the home domain. Knowledge about Adjusted mutual information transfers literally when a new case preserves the same carrier type, relation, and recognition test. In probability theory and information theory, adjusted mutual information, a variation of mutual information may be used for comparing clusterings. Given a set S of N elements S={s1, s2,\ldots sN} , consider two partitions of S, namely U={U1, U2,\ldots, UR} with R clusters, and V={V1, V2,\ldots, VC} with C clusters. Beyond the home domain. No canonical parent is asserted for Adjusted mutual information.

Neighborhood in Abstraction Space

Adjusted mutual information sits in a sparse region of the domain-specific corpus (64th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Clustering Methods & Validity Measures (14 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08