Skip to content

Canonical correlation

In statistics, canonical-correlation analysis (CCA), also called canonical variates analysis, is a way of inferring information from cross-covariance matrices.

Version
v1 · 2026-09-28 · History
Domain-specific #
8323
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomain
Multivariate Analysis → Experimental Design & Statistics

Core Idea

Canonical correlation is treated here as the recurring computerscienceandinformation identity summarized by this source-grounded definition: In statistics, canonical-correlation analysis (CCA), also called canonical variates analysis, is a way of inferring information from cross-covariance matrices. In statistics, canonical-correlation analysis (CCA), also called canonical variates analysis, is a way of inferring information from cross-covariance matrices. If we have two vectors X = (X 1 , ..., X n ) and Y = (Y 1 , ..., Y m ) of random variables, and there are correlations among the variables, then canonical-correlation analysis will find linear combinations of X and Y that have a maximum correlation.

How would you explain it like I'm…

The Best-Matching Mix

Imagine you have two groups of facts about the same kids, like how tall and how heavy they are, and how fast they run and how far they jump. Canonical correlation looks for a way to mix the facts in each group so that the two mixes go up and down together as much as possible. That shows how the two groups are connected.

Linking Two Sets of Measurements

Canonical correlation analysis, or CCA, is a statistics tool for finding how two groups of measurements are related. Say you have several measurements in group X and several in group Y for the same people. CCA looks for a weighted mix of the X measurements and a weighted mix of the Y measurements that line up with each other as strongly as possible, meaning they have the highest correlation. Those mixes are called canonical variates. It was introduced by Harold Hotelling in 1936, and it is used a lot in statistics and in machine learning that combines different kinds of data.

Maximally Correlated Linear Combinations

Canonical-correlation analysis (CCA), also called canonical variates analysis, is a statistical method for studying the relationship between two sets of variables by working with their cross-covariance matrices. Given two random vectors X = (X_1, ..., X_n) and Y = (Y_1, ..., Y_m) with correlations among their variables, CCA finds linear combinations of the X variables and of the Y variables that have the maximum possible correlation with each other. These pairs of linear combinations summarize how the two sets are related. It is broad enough that, as Knapp noted, virtually all common parametric significance tests can be treated as special cases of it. Harold Hotelling introduced it in 1936, although Camille Jordan had published the underlying mathematics, in terms of angles between flat subspaces, in 1875. It is important in multivariate statistics and multi-view learning, and has extensions such as probabilistic, sparse, multi-view, and deep CCA.

 

Canonical-correlation analysis (CCA), also known as canonical variates analysis, infers information from cross-covariance matrices. For random vectors X = (X_1, ..., X_n) and Y = (Y_1, ..., Y_m) with correlations among the variables, CCA finds linear combinations aᵀX and bᵀY that have maximum correlation with each other; these are the first pair of canonical variates, and their correlation is the first canonical correlation. The analysis is driven by the within-set covariance matrices of X and Y and the cross-covariance matrix between them. Knapp observed that virtually all commonly encountered parametric significance tests can be treated as special cases of CCA, which he called the general procedure for investigating relationships between two sets of variables. Harold Hotelling introduced the method in 1936, while Camille Jordan had published the mathematical concept in 1875 in the context of angles between flats. CCA is central in multivariate statistics and multi-view learning, with extensions including probabilistic CCA, sparse CCA, multi-view CCA, deep CCA, and DeepGeoCCA. Notation in the literature is sometimes inconsistent, so conventions should be checked when comparing sources.

Scope of Application

  • Population CCA definition via correlations. In practice, we would estimate the covariance matrix based on sampled data from X and Y (i.e. from a pair of data matrices).

  • Implementation. R as the standard function cancor and several other packages, including candisc, CCA and vegan.

  • SPSS as macro CanCorr shipped with the main software. The cosine function is ill-conditioned for small angles, leading to very inaccurate computation of highly correlated principal vectors in finite precision computer arithmetic.

  • Implementation. It is available as a function in.

  • Hypothesis testing. Each row can be tested for significance with the following method.

Clarity

A clear use of Canonical correlation names the carrier, the operative relation, and the conditions under which the source treats the identity as present. The minimal definition is In statistics, canonical-correlation analysis (CCA), also called canonical variates analysis, is a way of inferring information from cross-covariance matrices.

Manages Complexity

Canonical correlation compresses multiple computerscienceandinformation details into a stable diagnostic relation. The source shows both the central mechanism—cCA can be computed using singular value decomposition on a correlation matrix.—and the practical consequence—the definition of the canonical variables U and V is then equivalent to the definition of principal vectors for the pair of subspaces spanned by the entries of X and Y with respect to this inner.

Abstract Reasoning

  1. Type the carrier. Identify the computerscienceandinformation entities to which the claim applies.
  2. State the relation. Use the source-grounded identity: In statistics, canonical-correlation analysis (CCA), also called canonical variates analysis, is a way of inferring information from cross-covariance matrices.
  3. Check operation and conditions. Visualization of the results of canonical correlation is usually through bar plots of the coefficients of the two sets of variables for the pairs of canonical variates showing significant correlation.
  4. Demand recognition evidence.

Knowledge Transfer

Within the home domain. Knowledge about Canonical correlation transfers literally when a new case preserves the same carrier type, relation, and recognition test. In practice, we would estimate the covariance matrix based on sampled data from X and Y (i.e. from a pair of data matrices). R as the standard function cancor and several other packages, including candisc, CCA and vegan. Beyond the home domain. No canonical parent is asserted for Canonical correlation.

Neighborhood in Abstraction Space

Canonical correlation sits in a sparse region of the domain-specific corpus (67th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Multivariate & Spectral Signal Analysis (10 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08