Skip to content

Cophenetic correlation

In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points.

Version
v1 · 2026-09-28 · History
Domain-specific #
8721
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomains
Biostatistics, Cluster Analysis → Experimental Design & Statistics

Core Idea

Cophenetic correlation is treated here as the recurring crossdomainmodelsstructuresrepresentations identity summarized by this source-grounded definition: In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points. In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points.

How would you explain it like I'm…

Does the Tree Tell the Truth?

Imagine you draw a family-tree picture to show which of your toys are most alike. Cophenetic correlation is a score that checks how well your tree picture keeps the real 'how alike' amounts. A high score means the picture tells the truth about who is close to whom.

Tree-Match Score

Scientists often group things into a tree, where very similar things join low down and less similar groups join higher up. This tree is called a dendrogram. But making a tree can squash or stretch how far apart things really were. Cophenetic correlation is a number that checks how well the tree keeps the real distances. You compare, for every pair of things, how far apart they really are with how high up the tree they first join together. If the two match up closely, the score is high and the tree is a faithful summary.

Dendrogram Distance Fidelity

Cophenetic correlation (the cophenetic correlation coefficient) measures how faithfully a dendrogram — the tree produced by hierarchical clustering — preserves the pairwise distances in the original data. For any two data points, the cophenetic distance is the height of the tree node where they are first joined. The coefficient is the correlation between the original pairwise distances and these cophenetic distances, taken over all pairs. A value close to 1 means the tree is a faithful summary; a lower value means the clustering has distorted the real distances. It is used most in biostatistics, for example to check clustering models of DNA sequences or taxonomies, but applies wherever data come in clumps.

 

The cophenetic correlation coefficient measures how faithfully a dendrogram produced by hierarchical clustering preserves the pairwise distances among the original, unmodeled data points. For each pair (i, j), x(i, j) is the original distance and t(i, j) is the cophenetic distance, the height of the node at which i and j are first merged in the dendrogram. The coefficient c is the correlation between the x(i, j) and t(i, j) values over all pairs, computed from their deviations about their respective means. Values near 1 indicate that the tree closely reflects the original distance structure. It is widely used in biostatistics to evaluate cluster-based models of DNA sequences and taxonomic trees, applies wherever data naturally clump, and has been proposed as a test for nested clusters. It evaluates the fit of a given tree; it is not itself a clustering method.

Scope of Application

  • Calculating the cophenetic correlation coefficient. Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close".

  • Documented setting. Although it has been most widely applied in the field of biostatistics (typically to assess cluster-based models of DNA sequences, or other taxonomic models), it can also be used in other.

  • Calculating the cophenetic correlation coefficient. x(i,j) = |Xi-Xj| , the Euclidean distance between the ith and jth observations.

  • Calculating the cophenetic correlation coefficient. t(i,j) , the dendrogrammatic distance between the model points Ti and Tj .

  • Calculating the cophenetic correlation coefficient. This distance is the height of the node at which these two points are first joined together.

Clarity

A clear use of Cophenetic correlation names the carrier, the operative relation, and the conditions under which the source treats the identity as present. The minimal definition is In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points.

Manages Complexity

Cophenetic correlation compresses multiple crossdomainmodelsstructuresrepresentations details into a stable diagnostic relation. The source shows both the central mechanism—x(i,j) = |Xi-Xj| , the Euclidean distance between the ith and jth observations.—and the practical consequence—it is possible to calculate the cophenetic correlation in R using the dendextend R package. This compression makes cases comparable while leaving parameters, conventions, exceptions, and evidential quality explicit.

Abstract Reasoning

  1. Type the carrier. Identify the crossdomainmodelsstructuresrepresentations entities to which the claim applies.
  2. State the relation. Use the source-grounded identity: In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points.
  3. Check operation and conditions. t(i,j) , the dendrogrammatic distance between the model points Ti and Tj .
  4. Demand recognition evidence.

Knowledge Transfer

Within the home domain. Knowledge about Cophenetic correlation transfers literally when a new case preserves the same carrier type, relation, and recognition test. Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree. Although it has been most widely applied in the field of biostatistics (typically to assess cluster-based models of DNA sequences, or other taxonomic models), it can also be used in other fields of inquiry where raw data tend to.

Relationships to Other Abstractions

Local relationship map for Cophenetic correlationParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.CopheneticcorrelationDOMAINPrime abstraction: Correlation — is a kind ofCorrelationPRIME

Current abstraction Cophenetic correlation Domain-specific

Parents (1) — more general patterns this builds on

  • Cophenetic correlation is a kind of Correlation Prime

    Cophenetic correlation is a correlation between original pairwise distances and dendrogram-induced distances.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Cophenetic correlation sits in a moderately populated region (40th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.

Family — Data Structures & Graph Variants (17 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08