Tensions in Practice: Current-reference reporting in tension with attribution across revisions¶
Four fixed benchmark items and two reference versions
The old labels are (0,0,1,1); the revised labels are (0,1,1,1). The old model predicts (0,0,1,0); the new predicts (0,1,1,1). Comparing each model with its contemporary labels shows a rise from 3/4 to 4/4 matches. Cross-scoring reveals that the new model also scores 3/4 on the old labels, while the old model scores 2/4 on the new labels. Model and reference changes interact, so the diagonal increase has no unique context-free attribution.
Report against the operative standard
Give a compact score under the reference used at that time.
Separate co-changing contributions
Retain enough cross-version scores to inspect reference dependence.
Why these aims pull against each other
Cross-version scoring requires preserved model outputs, old labels and a meaningful common item set. Contemporary scores are cheaper to maintain but cannot alone attribute a historical change.
Choose an arrangement to see what changes and what remains difficult.
Rows hold model versions; columns hold reference-label versions. Off-diagonal scores reveal the dependence hidden by comparing only the two contemporary cells.
What this choice protects
What it costs
When it fits
Compare the arrangements
Keep contemporary scores
Retain old-model/old-label and new-model/new-label results, explicitly tagged with their versions.
| Old labels | New labels | |
|---|---|---|
| Old model | 3/4 | Not scored |
| New model | Not scored | 4/4 |
- What it protects
- Each model has a compact report under its operative reference.
- What it costs
- The diagonal change alone cannot establish how much the model improved under an unchanged target.
- When it fits
- Fits current-version conformance reporting when historical attribution is not claimed.
Illustration note: Different label versions are not silently treated as the same truth; the report makes the changed standard visible.
Add the cross-version scores
Score both retained prediction vectors against both reference vectors.
| Old labels | New labels | |
|---|---|---|
| Old model | 3/4 | 2/4 |
| New model | 3/4 | 4/4 |
- What it protects
- The effect of changing models can be compared under each fixed reference.
- What it costs
- Old outputs and reference definitions must remain available and comparable; a four-cell report is more demanding to maintain and explain.
- When it fits
- Fits a meaningful shared benchmark where investigating a historical score change matters.
Illustration note: The model difference is 0 under old labels and 2/4 under new labels; the reference difference is −1/4 for the old model and +1/4 for the new. No unique additive attribution is implied.
What this illustration does—and does not—establish
The source supplies the structural tension; the invented example makes one relation inspectable. Costs and conditions are part of each arrangement, not exceptions to a universal recommendation.
- All four items and label/prediction vectors are invented. Matches measure agreement with the chosen labels, not proven real-world truth.
- A deliberately revised normative standard can legitimately change; freezing every reference is not the goal.
- Different item sets, unavailable old outputs or changed label meanings can prevent a valid bridge. This finite complete cross-score does not solve those problems.
Source entries
Reference Standard Decay
The canonical tension motivates this comparison. The setting, finite values and arrangements are declared editorial illustrations, not measured findings.
Apparatus Improvement versus Reference Movement Confound (Coupling)
T5 — Apparatus Improvement versus Reference Movement Confound (Coupling). When both the apparatus and the reference change, a performance delta reflects both contributions, and they cannot be separated without explicit reference-vintage analysis. The failure mode is crediting the apparatus for a gain that was actually the reference moving toward it (or blaming it for a loss that was the reference moving away). Diagnostic: ask whether the apparatus and reference changed in the same interval; if so, score the *new* apparatus against the *old* reference (and vice versa) to decompose the delta — attributing the whole change to either side without this parallel scoring is a confound, not a measurement.
The source operation
Reference standard decay is the structural failure in which a measuring, scoring, or validating apparatus depends on a reference standard — a contingent artifact treated as fixed-and-correct for the purpose of evaluating something else — and that reference itself silently drifts on its own clock, while the apparatus continues to score against the now-decayed reference and reports numbers as if nothing had changed.
Fixed Reference Assumed versus Reference Meant to Move (Scopal)
The whole intervention catalogue presupposes the reference is *supposed* to be fixed; where it is intentionally revisable by definition — style guides, fashion norms, evolving taste, deliberately updated policy targets — tracking the moving reference is the feature, not the failure.