Canberra Distance¶
A coordinatewise distance between real vectors defined by summing |p_i−q_i|/(|p_i|+|q_i|), with a zero contribution when both coordinates are zero, thereby emphasizing relative differences near zero.
Core Idea¶
Canberra distance asks how large each coordinate difference is relative to the total magnitude present at that coordinate, then adds those relative discrepancies. This makes it less dominated by large-valued features than raw L1 distance.
The same feature also creates fragility near zero. Joint-zero, missing-value, measurement-floor, scaling, and optional normalization rules belong in every reproducible result.
How would you explain it like I'm…
The Fair-Size Difference Score
Relative-Difference Distance
Relative Per-Coordinate L1 Distance
Scope of Application¶
- Pattern recognition. Compares sparse feature vectors.
- Rank and profile comparison. Measures relative coordinate disagreement.
- Cybersecurity. Detects anomalous feature profiles.
- Ecology and microbiome research. Compares abundance vectors with careful zero handling.
Clarity¶
Report formula, feature alignment, units and scaling, zero convention, missingness, negative-value handling, detection limits, optional n−Z normalization, weights, preprocessing, and whether the output is raw sum or average. Inclusion test: Require coordinate-aligned vectors and the Canberra coordinate ratio with explicit treatment of joint zeros and any final normalization. Exclusion test: Exclude ordinary Manhattan distance, Bray–Curtis dissimilarity applied to totals, cosine distance, ad hoc percent difference, and software output whose zero/missing conventions are unknown. Nearest boundary: Manhattan distance sums absolute differences without dividing by local magnitude; Canberra distance therefore weights relative discrepancy most strongly near zero. Exit condition: The result becomes misleading when coordinates have incompatible meanings, measurement noise dominates near zero, missingness is encoded as zero, or variant normalization differs across comparisons. Common misclassifications: It is not unnormalized Manhattan distance. A missing value is not automatically zero. Near-zero noise can contribute strongly. Raw and averaged Canberra variants should not be mixed. Nearest named distinctions: Manhattan distance: Sums unnormalized absolute coordinate differences. Bray–Curtis dissimilarity: Normalizes the total absolute difference by total abundance for nonnegative data. Cosine distance: Compares vector direction rather than coordinatewise relative differences. Mean absolute percentage error: Uses one selected reference value in each denominator.
Manages Complexity¶
A simple per-feature normalization converts absolute geometry into relative geometry. Sparse high-dimensional data make edge conventions and measurement floors as influential as the central formula.
Abstract Reasoning¶
- Align vectors to identical coordinate definitions.
- Separate true zero, missing, censored, and below-detection values.
- Compute each normalized absolute difference under one zero convention.
- Aggregate using the declared raw or normalized variant.
- Test sensitivity to scaling, near-zero noise, feature selection, and alternative distances.
Knowledge Transfer¶
The formula transfers across numeric vectors, but meaning does not: coordinate semantics, zero, detection limits, and compositional structure must be rebuilt. A metric useful for one sparse domain can distort another.
Relationships to Other Abstractions¶
Current abstraction Canberra Distance Domain-specific
Parents (1) — more general patterns this builds on
-
Canberra Distance presupposes Comparison Prime
Canberra Distance presupposes Comparison: the parent's defining role is necessary to the child's frozen mechanism or criterion.
Hierarchy path (1) — routes to 1 parentless root
- Canberra Distance → Comparison → Self Checking
Neighborhood in Abstraction Space¶
Canberra Distance sits in a crowded region of the domain-specific corpus (29th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.
Family — Matrices, Measures & Numeric Structures (30 abstractions)
Nearest neighbors
- Log-Sum Inequality — 0.91
- Distance Matrix — 0.89
- Probability Density Function — 0.89
- Complex Affine Space — 0.89
- Grey Relational Analysis — 0.88
Computed from structural-signature embeddings · 2026-10-08