Tensions in Practice: Native magnitude in tension with standardized magnitude¶
Two invented populations measured on the same point scale
Group A has a mean difference of 10 points and a within-group reference standard deviation of 5. Group B has a mean difference of 20 points and a reference standard deviation of 20. B’s raw difference is larger. Dividing by the declared spread gives 2 standard deviations for A and 1 for B, reversing that magnitude comparison. Standardization changes the question from points gained to gain relative to variation.
Preserve the native unit
Compare actual point differences on the shared scale.
Compare relative to spread
Express the difference against each population’s variability.
Why these aims pull against each other
A common standardized unit facilitates some comparisons but embeds the selected denominator. It cannot replace the substantive meaning of the original point scale.
Choose an arrangement to see what changes and what remains difficult.
Raw differences and spreads stay fixed. The primary reported value changes units, and highlighted text identifies the larger value only under that declared comparison.
What this choice protects
What it costs
When it fits
Compare the arrangements
Lead with points
Use 10 and 20 points as the primary magnitude comparison; retain spreads in the table for transparency.
| Raw gain | Spread | Main value | |
|---|---|---|---|
| Group A | 10 | 5 | 10 points |
| Group B | 20 | 20 | 20 points |
- What it protects
- The larger change on the common point scale remains visible.
- What it costs
- The primary number does not show whether that change is large relative to each population’s spread.
- When it fits
- Fits decisions whose consequences attach to points on this same meaningful scale.
Illustration note: B has the larger raw difference, but no cost, uncertainty or causal benefit is inferred.
Lead with standard deviations
Divide each raw difference by its stipulated reference standard deviation.
| Raw gain | Spread | Main value | |
|---|---|---|---|
| Group A | 10 | 5 | 2 SD |
| Group B | 20 | 20 | 1 SD |
- What it protects
- The relative size of the difference within each population becomes directly comparable under the declared convention.
- What it costs
- The interpretation depends on different denominators and can obscure what a point means in practice.
- When it fits
- Fits a question explicitly about relative separation rather than total points gained.
Illustration note: A’s 2 versus B’s 1 is not evidence that A has a larger raw or more valuable effect. The divisor is shown rather than hidden.
What this illustration does—and does not—establish
The source supplies the structural tension; the invented example makes one relation inspectable. Costs and conditions are part of each arrangement, not exceptions to a universal recommendation.
- These are exact invented population quantities, not sample estimates; no uncertainty interval or significance claim is implied.
- The chosen reference standard deviation is part of the definition. Alternative standardizers can change the comparison.
- Neither difference is established as causal, and standardized size alone does not determine practical importance.
Source entries
Effect Size
The canonical tension motivates this comparison. The setting, finite values and arrangements are declared editorial illustrations, not measured findings.
Standardized versus raw/unstandardized metrics
Standardized effect sizes (Cohen's d, Pearson's r, η², Cramér's V) enable direct comparison across studies with different outcome scales, sample sizes, and populations, making them essential for meta-analysis. Raw unstandardized effect sizes (percent-lift, minutes-saved, dollars-earned, lives-saved) preserve substantive units that decision-makers and end-users understand intuitively. The tension is that standardization enables synthesis but obscures interpretability; raw metrics enable interpretation but lack cross-study comparability.
The source operation
Effect size quantifies the magnitude of a relationship or difference — the size of an observed effect in substantive, interpretable units — independent of sample size and separately from statistical significance.