Skip to content

Tensions in Practice: Current-reference reporting in tension with attribution across revisions

Four fixed benchmark items and two reference versions

The old labels are (0,0,1,1); the revised labels are (0,1,1,1). The old model predicts (0,0,1,0); the new predicts (0,1,1,1). Comparing each model with its contemporary labels shows a rise from 3/4 to 4/4 matches. Cross-scoring reveals that the new model also scores 3/4 on the old labels, while the old model scores 2/4 on the new labels. Model and reference changes interact, so the diagonal increase has no unique context-free attribution.

Report against the operative standard

Give a compact score under the reference used at that time.

Separate co-changing contributions

Retain enough cross-version scores to inspect reference dependence.

Why these aims pull against each other

Cross-version scoring requires preserved model outputs, old labels and a meaningful common item set. Contemporary scores are cheaper to maintain but cannot alone attribute a historical change.

Compare the arrangements

Keep contemporary scores

Retain old-model/old-label and new-model/new-label results, explicitly tagged with their versions.

Score each model in its own reference version
Old labelsNew labels
Old model3/4Not scored
New modelNot scored4/4
What it protects
Each model has a compact report under its operative reference.
What it costs
The diagonal change alone cannot establish how much the model improved under an unchanged target.
When it fits
Fits current-version conformance reporting when historical attribution is not claimed.

Illustration note: Different label versions are not silently treated as the same truth; the report makes the changed standard visible.

Add the cross-version scores

Score both retained prediction vectors against both reference vectors.

Score both models against both references
Old labelsNew labels
Old model3/42/4
New model3/44/4
What it protects
The effect of changing models can be compared under each fixed reference.
What it costs
Old outputs and reference definitions must remain available and comparable; a four-cell report is more demanding to maintain and explain.
When it fits
Fits a meaningful shared benchmark where investigating a historical score change matters.

Illustration note: The model difference is 0 under old labels and 2/4 under new labels; the reference difference is −1/4 for the old model and +1/4 for the new. No unique additive attribution is implied.

What this illustration does—and does not—establish

The source supplies the structural tension; the invented example makes one relation inspectable. Costs and conditions are part of each arrangement, not exceptions to a universal recommendation.

  • All four items and label/prediction vectors are invented. Matches measure agreement with the chosen labels, not proven real-world truth.
  • A deliberately revised normative standard can legitimately change; freezing every reference is not the goal.
  • Different item sets, unavailable old outputs or changed label meanings can prevent a valid bridge. This finite complete cross-score does not solve those problems.

Source entries

Reference Standard Decay

Prime · Source of the tension

The canonical tension motivates this comparison. The setting, finite values and arrangements are declared editorial illustrations, not measured findings.

Apparatus Improvement versus Reference Movement Confound (Coupling)

T5 — Apparatus Improvement versus Reference Movement Confound (Coupling). When both the apparatus and the reference change, a performance delta reflects both contributions, and they cannot be separated without explicit reference-vintage analysis. The failure mode is crediting the apparatus for a gain that was actually the reference moving toward it (or blaming it for a loss that was the reference moving away). Diagnostic: ask whether the apparatus and reference changed in the same interval; if so, score the *new* apparatus against the *old* reference (and vice versa) to decompose the delta — attributing the whole change to either side without this parallel scoring is a confound, not a measurement.

Read the source section

The source operation

Reference standard decay is the structural failure in which a measuring, scoring, or validating apparatus depends on a reference standard — a contingent artifact treated as fixed-and-correct for the purpose of evaluating something else — and that reference itself silently drifts on its own clock, while the apparatus continues to score against the now-decayed reference and reports numbers as if nothing had changed.

Read the source section

Fixed Reference Assumed versus Reference Meant to Move (Scopal)

The whole intervention catalogue presupposes the reference is *supposed* to be fixed; where it is intentionally revisable by definition — style guides, fashion norms, evolving taste, deliberately updated policy targets — tracking the moving reference is the feature, not the failure.

Read the source section