Skip to content

Clustering Methods & Validity Measures

← Back to Domain-Specific Families

Abstractions about grouping data and validating the resulting groupings, covering clustering algorithms (k-means, complete-linkage clustering, medoids), cluster-quality and similarity measures (silhouette, Davies-Bouldin index, Rand index, adjusted mutual information), and related structures like locality-sensitive hashing and exponential trees.

14 abstractions in this family — domain-specific abstractions that sit near one another in structural-signature space (k-means over structural-signature embeddings). Each is shown with its short description.

  • Adjusted mutual information — By adopting a hypergeometric model of randomness, it can be shown that the expected mutual information between two random clusterings is.
  • Cluster algebra — A commutative algebra generated from overlapping algebraically independent clusters by iterated birational mutations governed by exchange matrices or quivers.
  • Complete-linkage clustering — Build an agglomerative hierarchy by defining intercluster distance as the farthest cross-cluster pair and repeatedly merging the pair with the smallest such maximum.
  • Davies–Bouldin index — An internal cluster-validity score averaging, over clusters, the worst ratio of combined within-cluster scatter to between-centroid separation, with lower values indicating better compactness-separation tradeoff.
  • Determining the number of clusters in a data set — The model-selection problem of choosing a clustering resolution or number k that balances within-cluster fit, separation, stability, complexity, domain meaning, and intended use.
  • Disjunct matrix — A binary nonadaptive group-testing design in which every column has a row containing 1 while any chosen set of at most d other columns all contain 0.
  • Exponential tree — A search-tree structure whose branching factors shrink doubly exponentially with depth, storing keys at leaves and auxiliary predecessor structures at internal nodes.
  • Hopkins statistic — A nearest-neighbor statistic comparing observed data with uniform reference points to assess spatial cluster tendency.
  • K-means clustering — An optimization method that partitions vectors into k groups by assigning each observation to its nearest centroid and minimizing total within-cluster squared Euclidean distance.
  • Locality-sensitive hashing — A randomized indexing method using hash families whose collision probability increases with similarity under a target distance measure.
  • Medoid — An observed member of a dataset or cluster minimizing total dissimilarity to the other members, used as a robust representative when an arithmetic centroid is unavailable or inappropriate.
  • Rand index — A pair-counting similarity measure for two partitions that counts element pairs on which both clusterings agree.
  • Random indexing — An incremental dimensionality-reduction method that assigns sparse random index vectors to items and accumulates their contextual vectors, approximating high-dimensional distributional geometry in fixed space.
  • Silhouette (clustering) — An internal cluster-validation score comparing each observation’s mean within-cluster dissimilarity with its smallest mean dissimilarity to another cluster.