Skip to content

Canberra Distance

A coordinatewise distance between real vectors defined by summing |p_i−q_i|/(|p_i|+|q_i|), with a zero contribution when both coordinates are zero, thereby emphasizing relative differences near zero.

Version
v1 · 2026-09-28 · History
Domain-specific #
8321
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomains
Multivariate Analysis, Dissimilarity Measures → Experimental Design & Statistics
Aliases
Canberra Metric, Canberra Dissimilarity

Core Idea

Canberra distance asks how large each coordinate difference is relative to the total magnitude present at that coordinate, then adds those relative discrepancies. This makes it less dominated by large-valued features than raw L1 distance.

The same feature also creates fragility near zero. Joint-zero, missing-value, measurement-floor, scaling, and optional normalization rules belong in every reproducible result.

How would you explain it like I'm…

The Fair-Size Difference Score

Two friends compare their collections: one has 2 marbles and 4 stickers, the other has 3 marbles and 400 stickers. Canberra distance looks at each kind of thing on its own and asks how big the difference is compared to how many there are in total, then adds those up. That way a huge pile of one thing doesn't drown out the rest.

Relative-Difference Distance

Canberra distance is a way to measure how different two lists of numbers are. For each position in the lists, you take the difference between the two numbers and divide it by their total size. So a difference of 1 between 1 and 2 counts a lot, but a difference of 1 between 100 and 101 counts very little. Then you add up all these relative differences. This keeps big-valued items from dominating, but it gets touchy when numbers are near zero, so you have to decide ahead of time how to handle zeros and missing values.

Relative Per-Coordinate L1 Distance

Canberra distance measures the dissimilarity between two vectors by comparing each coordinate's difference with the total magnitude at that coordinate, then summing: for each coordinate i it uses |x_i − y_i| / (|x_i| + |y_i|). Because each term is a relative discrepancy, large-valued features do not dominate the way they do in ordinary L1 (Manhattan) distance, which just sums raw absolute differences. The same property makes it fragile near zero: when both values are tiny, a small absolute difference can produce a large term, and when both are exactly zero the term is undefined. So any reproducible result has to state its rules for joint zeros, missing values, measurement floors (values below what can be detected), scaling, and any optional normalization.

 

Canberra distance between vectors x and y is a sum of coordinatewise relative discrepancies, Σ_i |x_i − y_i| / (|x_i| + |y_i|), in which each coordinate difference is scaled by the total magnitude present at that coordinate. Compared with raw L1 distance, this weighting prevents features with large values from dominating the total, since each term is bounded and reflects proportional rather than absolute disagreement. The same normalization causes instability near zero: small absolute differences between small values yield large terms, and the term is undefined when both coordinates are zero. Measurement noise near a detection floor can therefore drive the distance. Because results depend heavily on these edge cases, a reproducible analysis must specify the treatment of joint zeros, missing values, measurement floors, scaling of inputs, and any optional normalization of the total, for example dividing by the number of coordinates used.

Scope of Application

  • Pattern recognition. Compares sparse feature vectors.
  • Rank and profile comparison. Measures relative coordinate disagreement.
  • Cybersecurity. Detects anomalous feature profiles.
  • Ecology and microbiome research. Compares abundance vectors with careful zero handling.

Clarity

Report formula, feature alignment, units and scaling, zero convention, missingness, negative-value handling, detection limits, optional n−Z normalization, weights, preprocessing, and whether the output is raw sum or average. Inclusion test: Require coordinate-aligned vectors and the Canberra coordinate ratio with explicit treatment of joint zeros and any final normalization. Exclusion test: Exclude ordinary Manhattan distance, Bray–Curtis dissimilarity applied to totals, cosine distance, ad hoc percent difference, and software output whose zero/missing conventions are unknown. Nearest boundary: Manhattan distance sums absolute differences without dividing by local magnitude; Canberra distance therefore weights relative discrepancy most strongly near zero. Exit condition: The result becomes misleading when coordinates have incompatible meanings, measurement noise dominates near zero, missingness is encoded as zero, or variant normalization differs across comparisons. Common misclassifications: It is not unnormalized Manhattan distance. A missing value is not automatically zero. Near-zero noise can contribute strongly. Raw and averaged Canberra variants should not be mixed. Nearest named distinctions: Manhattan distance: Sums unnormalized absolute coordinate differences. Bray–Curtis dissimilarity: Normalizes the total absolute difference by total abundance for nonnegative data. Cosine distance: Compares vector direction rather than coordinatewise relative differences. Mean absolute percentage error: Uses one selected reference value in each denominator.

Manages Complexity

A simple per-feature normalization converts absolute geometry into relative geometry. Sparse high-dimensional data make edge conventions and measurement floors as influential as the central formula.

Abstract Reasoning

  1. Align vectors to identical coordinate definitions.
  2. Separate true zero, missing, censored, and below-detection values.
  3. Compute each normalized absolute difference under one zero convention.
  4. Aggregate using the declared raw or normalized variant.
  5. Test sensitivity to scaling, near-zero noise, feature selection, and alternative distances.

Knowledge Transfer

The formula transfers across numeric vectors, but meaning does not: coordinate semantics, zero, detection limits, and compositional structure must be rebuilt. A metric useful for one sparse domain can distort another.

Relationships to Other Abstractions

Local relationship map for Canberra DistanceParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Canberra DistanceDOMAINPrime abstraction: Comparison — presupposesComparisonPRIME

Current abstraction Canberra Distance Domain-specific

Parents (1) — more general patterns this builds on

  • Canberra Distance presupposes Comparison Prime

    Canberra Distance presupposes Comparison: the parent's defining role is necessary to the child's frozen mechanism or criterion.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Canberra Distance sits in a crowded region of the domain-specific corpus (29th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Matrices, Measures & Numeric Structures (30 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08