Clustering Methods & Validity Measures¶
← Back to Domain-Specific Families
Abstractions about grouping data and validating the resulting groupings, covering clustering algorithms (k-means, complete-linkage clustering, medoids), cluster-quality and similarity measures (silhouette, Davies-Bouldin index, Rand index, adjusted mutual information), and related structures like locality-sensitive hashing and exponential trees.
14 abstractions in this family — domain-specific abstractions that sit near one another in structural-signature space (k-means over structural-signature embeddings). Each is shown with its short description.
- Adjusted mutual information — By adopting a hypergeometric model of randomness, it can be shown that the expected mutual information between two random clusterings is.
- Cluster algebra — A commutative algebra generated from overlapping algebraically independent clusters by iterated birational mutations governed by exchange matrices or quivers.
- Complete-linkage clustering — Build an agglomerative hierarchy by defining intercluster distance as the farthest cross-cluster pair and repeatedly merging the pair with the smallest such maximum.
- Davies–Bouldin index — An internal cluster-validity score averaging, over clusters, the worst ratio of combined within-cluster scatter to between-centroid separation, with lower values indicating better compactness-separation tradeoff.
- Determining the number of clusters in a data set — The model-selection problem of choosing a clustering resolution or number k that balances within-cluster fit, separation, stability, complexity, domain meaning, and intended use.
- Disjunct matrix — A binary nonadaptive group-testing design in which every column has a row containing 1 while any chosen set of at most d other columns all contain 0.
- Exponential tree — A search-tree structure whose branching factors shrink doubly exponentially with depth, storing keys at leaves and auxiliary predecessor structures at internal nodes.
- Hopkins statistic — A nearest-neighbor statistic comparing observed data with uniform reference points to assess spatial cluster tendency.
- K-means clustering — An optimization method that partitions vectors into k groups by assigning each observation to its nearest centroid and minimizing total within-cluster squared Euclidean distance.
- Locality-sensitive hashing — A randomized indexing method using hash families whose collision probability increases with similarity under a target distance measure.
- Medoid — An observed member of a dataset or cluster minimizing total dissimilarity to the other members, used as a robust representative when an arithmetic centroid is unavailable or inappropriate.
- Rand index — A pair-counting similarity measure for two partitions that counts element pairs on which both clusterings agree.
- Random indexing — An incremental dimensionality-reduction method that assigns sparse random index vectors to items and accumulates their contextual vectors, approximating high-dimensional distributional geometry in fixed space.
- Silhouette (clustering) — An internal cluster-validation score comparing each observation’s mean within-cluster dissimilarity with its smallest mean dissimilarity to another cluster.