Skip to content

Cophenetic correlation

In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points.

Version
v1 · 2026-09-28 · History
Domain-specific #
8721
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomains
Biostatistics, Cluster Analysis → Experimental Design & Statistics

Core Idea

Cophenetic correlation is treated here as the recurring cross_domain_models_structures_representations identity summarized by this source-grounded definition: In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points.

In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points. Although it has been most widely applied in the field of biostatistics (typically to assess cluster-based models of DNA sequences, or other taxonomic models), it can also be used in other fields of inquiry where raw data tend to occur in clumps, or clusters. This coefficient has also been proposed for use as a test for nested clusters.

This distance is the height of the node at which these two points are first joined together. Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree. Then, letting \bar{x} be the average of the x(i, j), and letting \bar{t} be the average of the t(i, j), the cophenetic correlation coefficient c is given by.

For Cophenetic correlation, the abstraction is narrower than the article's general subject matter: a positive case must preserve In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points. Retaining only the name, a familiar example, or a downstream effect is insufficient. The specialist roles and tests remain anchored in cross_domain_models_structures_representations, which is why this identity is domain-specific rather than prime.

How would you explain it like I'm…

Does the Tree Tell the Truth?

Imagine you draw a family-tree picture to show which of your toys are most alike. Cophenetic correlation is a score that checks how well your tree picture keeps the real 'how alike' amounts. A high score means the picture tells the truth about who is close to whom.

Tree-Match Score

Scientists often group things into a tree, where very similar things join low down and less similar groups join higher up. This tree is called a dendrogram. But making a tree can squash or stretch how far apart things really were. Cophenetic correlation is a number that checks how well the tree keeps the real distances. You compare, for every pair of things, how far apart they really are with how high up the tree they first join together. If the two match up closely, the score is high and the tree is a faithful summary.

Dendrogram Distance Fidelity

Cophenetic correlation (the cophenetic correlation coefficient) measures how faithfully a dendrogram — the tree produced by hierarchical clustering — preserves the pairwise distances in the original data. For any two data points, the cophenetic distance is the height of the tree node where they are first joined. The coefficient is the correlation between the original pairwise distances and these cophenetic distances, taken over all pairs. A value close to 1 means the tree is a faithful summary; a lower value means the clustering has distorted the real distances. It is used most in biostatistics, for example to check clustering models of DNA sequences or taxonomies, but applies wherever data come in clumps.

 

The cophenetic correlation coefficient measures how faithfully a dendrogram produced by hierarchical clustering preserves the pairwise distances among the original, unmodeled data points. For each pair (i, j), x(i, j) is the original distance and t(i, j) is the cophenetic distance, the height of the node at which i and j are first merged in the dendrogram. The coefficient c is the correlation between the x(i, j) and t(i, j) values over all pairs, computed from their deviations about their respective means. Values near 1 indicate that the tree closely reflects the original distance structure. It is widely used in biostatistics to evaluate cluster-based models of DNA sequences and taxonomic trees, applies wherever data naturally clump, and has been proposed as a test for nested clusters. It evaluates the fit of a given tree; it is not itself a clustering method.

Structural Signature

Sig role-phrases:

  • Defining carrier — Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree.
  • Constitutive relation — x(i,j) = |X_i-X_j| , the Euclidean distance between the ith and jth observations.
  • Operating condition — t(i,j) , the dendrogrammatic distance between the model points T_i and T_j .
  • Recognition evidence — This distance is the height of the node at which these two points are first joined together.
  • Admissible variation — Then, letting \bar{x} be the average of the x(i, j), and letting \bar{t} be the average of the t(i, j), the cophenetic correlation coefficient c is given by.
  • Characteristic consequence — It is possible to calculate the cophenetic correlation in R using the dendextend R package.
  • Failure boundary — In Python, the SciPy package also has an implementation.

What It Is Not

  • Not the whole field of cross_domain_models_structures_representations. The node requires the specific identity stated by In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points.
  • Not an over-broad reading. Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree.
  • Not an over-broad reading. x(i,j) = |X_i-X_j| , the Euclidean distance between the ith and jth observations.
  • Not an over-broad reading. t(i,j) , the dendrogrammatic distance between the model points T_i and T_j .
  • Not automatically Polychoric correlation. Retrieval proximity does not establish equivalence; the two identities must be compared by carrier, operation, and failure boundary.

Scope of Application

Cophenetic correlation applies literally inside cross_domain_models_structures_representations wherever the source-defined carrier and relation can be established. Its documented habitats include:

  • Calculating the cophenetic correlation coefficient. Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree.
  • Documented setting. Although it has been most widely applied in the field of biostatistics (typically to assess cluster-based models of DNA sequences, or other taxonomic models), it can also be used in other fields of inquiry where raw data tend to occur in clumps, or clusters.
  • Calculating the cophenetic correlation coefficient. x(i,j) = |X_i-X_j| , the Euclidean distance between the ith and jth observations.
  • Calculating the cophenetic correlation coefficient. t(i,j) , the dendrogrammatic distance between the model points T_i and T_j .
  • Calculating the cophenetic correlation coefficient. This distance is the height of the node at which these two points are first joined together.
  • Calculating the cophenetic correlation coefficient. Then, letting \bar{x} be the average of the x(i, j), and letting \bar{t} be the average of the t(i, j), the cophenetic correlation coefficient c is given by.

Outside cross_domain_models_structures_representations, the name should be retained only when these same operational conditions survive; otherwise the comparison belongs to the broader parent Pattern or should be marked as analogy.

Clarity

A clear use of Cophenetic correlation names the carrier, the operative relation, and the conditions under which the source treats the identity as present. The minimal definition is In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points. The strongest recognition evidence in the frozen account is: This distance is the height of the node at which these two points are first joined together. A report should distinguish that evidence from a proxy, consequence, or common implementation. It should also state the qualification Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree. so that a reader can reproduce the classification rather than infer it from topical resemblance.

Manages Complexity

Cophenetic correlation compresses multiple cross_domain_models_structures_representations details into a stable diagnostic relation. The source shows both the central mechanism—x(i,j) = |X_i-X_j| , the Euclidean distance between the ith and jth observations.—and the practical consequence—it is possible to calculate the cophenetic correlation in R using the dendextend R package. This compression makes cases comparable while leaving parameters, conventions, exceptions, and evidential quality explicit. It is lossy by design: local history and implementation details may be omitted only when they do not alter the defining relation.

Abstract Reasoning

  1. Type the carrier. Identify the cross_domain_models_structures_representations entities to which the claim applies.
  2. State the relation. Use the source-grounded identity: In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points.
  3. Check operation and conditions. t(i,j) , the dendrogrammatic distance between the model points T_i and T_j .
  4. Demand recognition evidence. This distance is the height of the node at which these two points are first joined together.
  5. Test variation. Change an implementation or setting while preserving then, letting \bar{x} be the average of the x(i, j), and letting \bar{t} be the average of the t(i, j), the cophenetic correlation coefficient c is given by.
  6. Run the collapse test. Remove the defining operation; if the label still seems equally apt, only a topic or correlate was retained.
  7. Reduce cautiously. When the specialist conditions cannot be carried, route the residual comparison to Pattern.

Knowledge Transfer

Within the home domain. Knowledge about Cophenetic correlation transfers literally when a new case preserves the same carrier type, relation, and recognition test. Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree. Although it has been most widely applied in the field of biostatistics (typically to assess cluster-based models of DNA sequences, or other taxonomic models), it can also be used in other fields of inquiry where raw data tend to occur in clumps, or clusters.

Beyond the home domain. No canonical parent is asserted for Cophenetic correlation. An outside case receives the specialist name only when the same typed roles and rejection conditions can be filled literally; otherwise the comparison remains an analogy pending later graph densification.

Examples

Canonical

Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree. This case is canonical because it supplies a concrete carrier and lets the defining relation be checked rather than merely named.

Mapped back: carrier → the entities in the documented case; operation → In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points; recognition evidence → This distance is the height of the node at which these two points are first joined together

Applied / In Practice

x(i,j) = |X_i-X_j| , the Euclidean distance between the ith and jth observations. The applied case shows how the identity is used under a second setting or qualification while keeping the same operative relation.

Mapped back: changed setting → Calculating the cophenetic correlation coefficient; invariant → In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points; boundary → the case exits the class when suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree

Structural Tensions

T1 — Stable identity versus admissible variation. Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.

Diagnostic: Which changes preserve the defining relation, and which replace it?

T2 — Recognition versus proxy. x(i,j) = |X_i-X_j| , the Euclidean distance between the ith and jth observations. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.

Diagnostic: Does the cited evidence establish the identity or only a correlated sign?

T3 — Definition versus implementation. t(i,j) , the dendrogrammatic distance between the model points T_i and T_j . The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.

Diagnostic: Is the observed implementation constitutive, optional, or merely common?

T4 — Scope versus overextension. This distance is the height of the node at which these two points are first joined together. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.

Diagnostic: Can every claimed application fill the same typed roles without metaphor?

T5 — Transfer versus domain accent. Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.

Diagnostic: Does the receiving case instantiate Cophenetic correlation literally, co-instantiate Pattern, or only resemble it?

T6 — Autonomy versus reduction. x(i,j) = |X_i-X_j| , the Euclidean distance between the ith and jth observations. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.

Diagnostic: What does Cophenetic correlation distinguish that the broader parent Pattern leaves together?

Structural–Framed Character

Cophenetic correlation is mixed or framed-leaning. Its structural side is the repeatable organization summarized by In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points. Its framed side is the cross_domain_models_structures_representations vocabulary that fixes the carrier, evidence, exceptions, and admissible transformations.

Evaluative weight: the identity can be stated descriptively even when applications carry practical stakes. Human-practice dependence: the source-grounded carrier determines whether the relation exists independently or is constituted by a practice. Institutional origin: disciplinary conventions stabilize the name and test. Vocabulary portability: t(i,j) , the dendrogrammatic distance between the model points T_i and T_j . Import versus recognition: literal transfer requires the same mechanism; shape alone is analogy.

Its portable skeleton is Pattern. Its character: a recurring specialist identity whose thin organization can be abstracted, while its operational meaning remains domain-bound.

Structural Core vs. Domain Accent

What is skeletal. In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points. The stable skeleton is the typed relation expressed in that definition and the entry's recognition and collapse tests. The source identifies these operative conditions: Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree. x(i,j) = |Xi-Xj| , the Euclidean distance between the ith and jth observations. It further constrains recognition and variation through: t(i,j) , the dendrogrammatic distance between the model points Ti and Tj . This distance is the height of the node at which these two points are first joined together.

What is domain-bound. cross domain models structures representations supplies the operative entities, technical vocabulary, warrants, and exceptions that make Cophenetic correlation literal. Its documented scope includes the condition that Suppose that the original data {X i } have been modeled using a cluster method to produce a dendrogram {T i }; that is, a simplified model in which data that are "close" have been grouped into a hierarchical tree. Another bounded application condition is that Although it has been most widely applied in the field of biostatistics (typically to assess cluster-based models of DNA sequences, or other taxonomic models), it can also be used in other fields of inquiry where raw data tend to occur in clumps, or clusters. These are not decorative examples; they determine which carrier and evidence can fill the abstraction's roles.

Why no parent is asserted. Removing those specialist details does not currently yield one live catalog node that is a necessary genus for every instance. The entry is therefore approved as unparented rather than attached by topical resemblance. Its collapse evidence remains specific—Then, letting \bar{x} be the average of the x(i, j), and letting \bar{t} be the average of the t(i, j), the cophenetic correlation coefficient c is given by.—and future graph densification may discover a defensible relation only if it preserves that boundary.

This entry is a kind of Correlation.

  • Approved unparented node. No current live node supplies a defensible necessary genus or structural prerequisite for Cophenetic correlation. The reviewed identity is: In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points. The accelerated suggestion was declined because topical or lexical similarity does not establish hierarchy; the node is admitted without a parent pending later graph densification.
  • Related reasoning operations. Evidence, representation, comparison, classification, transformation, or evaluation may participate in particular cases, but participation does not make any one of them a necessary parent of every instance.

Relationships to Other Abstractions

Local relationship map for Cophenetic correlationParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.CopheneticcorrelationDOMAINPrime abstraction: Correlation — is a kind ofCorrelationPRIME

Current abstraction Cophenetic correlation Domain-specific

Parents (1) — more general patterns this builds on

  • Cophenetic correlation is a kind of Correlation Prime

    Cophenetic correlation is a correlation between original pairwise distances and dendrogram-induced distances.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Cophenetic correlation sits in a moderately populated region (40th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.

Family — Data Structures & Graph Variants (17 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • Pattern. The parent omits the specialist differentia. Tell: Can the case establish In statistics, and especially in biostatistics, cophenetic correlation (more precisely, the cophenetic correlation coefficient) is a measure of how faithfully a dendrogram preserves the pairwise distances between the original unmodeled data points?
  • Polychoric correlation. An estimate of the correlation between two latent normally distributed continuous variables inferred from their observed ordinal categories through threshold models. Tell: Which entry's carrier, operation, and failure condition are satisfied?
  • Pearson correlation coefficient. The unitless covariance of two variables divided by the product of their standard deviations, measuring linear association from minus one to one. Tell: Which entry's carrier, operation, and failure condition are satisfied?
  • Cophonicity. A vibrational-overlap metric measuring how strongly two atomic species contribute together within a selected frequency range. Tell: Which entry's carrier, operation, and failure condition are satisfied?
  • A measurement, proxy, or consequence. Those may provide evidence without being the identity. Tell: Would Cophenetic correlation remain present if the detector or downstream effect changed?
  • A metaphorical analogue. A similar shape outside cross_domain_models_structures_representations lacks the specialist mechanism. Tell: Do the native roles transfer literally, or only the parent Pattern?

References

  • Frozen Wikipedia discovery revision: https://en.wikipedia.org/wiki/Cophenetic_correlation (revision 1353474049).
  • Preserved source candidate: http://www.osti.gov/bridge/servlets/purl/9576-lcvvCD/webviewable/9576.pdf
  • Preserved source candidate: http://life.bio.sunysb.edu/ee/rohlf/reprints/RohlfFisher_1968.pdf
  • Preserved source candidate: http://www.mathworks.com/access/helpdesk/help/toolbox/stats/index.html?/access/helpdesk/help/toolbox/stats/cophenet.html
  • Preserved source candidate: https://cran.r-project.org/web/packages/dendextend/vignettes/dendextend.html
  • Preserved source candidate: https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.cluster.hierarchy.cophenet.html
  • Preserved source candidate: https://www.mathworks.com/help/stats/cophenet.html
  • Preserved source candidate: http://people.revoledu.com/kardi/tutorial/Clustering/index.html
  • Preserved source candidate: https://stackoverflow.com/questions/5639794/in-r-how-can-i-plot-a-similarity-matrix-like-a-block-graph-after-clustering-d

The frozen Wikipedia revision is discovery provenance. The retained source set was reviewed for identity, formal or operational relation, and scope. The encyclopedia's structural synthesis is bounded to those claims; a thin authority surface is recorded as a nonblocking source-strengthening repair rather than concealed.