Skip to content

Relevant-Difference Challenge Test

Test or assessment — instantiates Fairness-Standard Selection and Reconciliation

Requires differential treatment to connect to purpose and survive proxy, counterexample, and necessity review.

Version
v1 · 2026-08-24 · History
Mechanism #
7362
Type
Test or Assessment
Form family
Assessment, Review & Assurance
Solution family
Comparison & Evaluation
Problem family
Goal, Value & Purpose Misalignment
Problem subfamily
Normative Standard & Weighting Choice
Origin domain
Law & Governance
Also from
Philosophy
Instantiates
Fairness-Standard Selection and Reconciliation

The Relevant-Difference Challenge Test interrogates a single proposed distinction — "we treat A differently from B because of feature F" — and forces F to earn its keep. F must name the standard and purpose it serves, survive a proxy check (is it really tracking something illegitimate?), survive a counterexample (a case where F holds but the purpose fails), and survive a necessity check (is there a narrower feature that serves the same purpose with less collateral differentiation?). It evaluates the justification of one differentiation — the logic of a classification — not the spread of outcomes across a population and not the durability of a whole rule under attack. The burden always rests on whoever proposes the difference, and it rises when the trait is immutable or the harm severe.

Example

An auto insurer prices premiums partly on occupation, charging some occupations roughly 20% more, and a regulator applies the test to that factor. Purpose: the insurer claims occupation predicts claim risk, so the legitimate purpose is pricing to expected loss. Proxy check: does occupation actually predict crashes, or is it standing in for income, neighborhood, or a protected trait? The insurer is asked to show the actuarial link net of those correlates — and the signal shrinks sharply once territory and mileage are controlled. Counterexample: two drivers in a "low-risk" occupation with atrocious individual records still receive the favorable occupation credit, showing the factor overrides the very risk it claims to predict. Necessity: a narrower feature — the driver's own record and annual mileage — serves the predictive purpose directly, with far less collateral differentiation, so occupation fails necessity.

The verdict is not a premium. The test does not set anyone's rate; it strips out a distinction that cannot connect to purpose, leaving the insurer to price on the narrower, better-justified features that survived.

How it works

  • One distinction at a time. The unit of analysis is a single "treat differently because of F," not a whole scoring model.
  • Demand an explicit purpose tied to a named candidate standard — "F serves this principle's legitimate aim."
  • Run four gates: purpose-link, proxy/laundering check, counterexample search, and necessity/narrower-alternative. Each is pass or fail.
  • Put the burden on the proponent, raised for immutable or protected traits and for severe, hard-to-remedy harms.
  • Emit keep / narrow / reject, with the reason recorded for each surviving or defeated distinction.

Tuning parameters

  • Burden of proof — how strong the required purpose-link evidence must be; raise it for immutable traits and severe consequences, but set too high it defeats even legitimate differences.
  • Proxy sensitivity — how aggressively the feature is decomposed against suspect correlates before its link is accepted.
  • Counterexample depth — how hard the review hunts for purpose-failing cases; deeper search finds more laundered distinctions but costs effort.
  • Necessity strictness — whether a merely-narrower alternative defeats the feature or only a strictly-dominating one.

When it helps, and when it misleads

Its strength is dismantling proxy differentiation and disciplining "we've always distinguished by X": a feature that cannot state a purpose, survive its counterexamples, and beat a narrower alternative is exposed as decoration or laundering. It is the mechanism that keeps categories from smuggling inherited disadvantage into a "difference."

Its failure mode is two-edged. Applied without restraint it can challenge every distinction into paralysis; applied credulously it accepts a manufactured just-so purpose story that passes the purpose gate on paper while doing other work underneath. The classic misuse is asymmetric rigor — demanding near-impossible proof for a disfavored difference while waving a favored one through the same gates. The discipline that keeps it honest is narrow tailoring[n1] applied symmetrically: the same four gates, at the same strictness, for every proposed difference, favored or not.

How it implements the components

  • comparability_and_relevant_difference_model — it is this model in action: defining like cases and testing each proposed difference for purpose-link, proxy risk, counterexamples, and narrower alternatives.
  • candidate_fairness_standard_set — it forces every difference to name which candidate standard's purpose it serves, filtering the set down to distinctions with a legitimate principle behind them.

It does not measure how burdens and errors spread across parties and distribution tails (distributional_outcome_and_revision_monitor — that's Distributional Impact and Tail Audit, its nearest twin: that audit measures the spread of outcomes, while this test judges the legitimacy of one distinction) or adversarially probe a rule for gaming, baseline shift, and exception capture (consistent_application_exception_and_reason_rule, baseline_counterfactual_and_burden_definition — that's Fairness-Metric and Exception Stress Test).

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: Relevant-Difference Challenge Test operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it requires differential treatment to connect to purpose and survive proxy, counterexample, and necessity review.

Independent corroboration: The frozen evidence defines Relevant-Difference Challenge Test as 'Requires differential treatment to connect to purpose and survive proxy, counterexample, and necessity review', so its operative form is Assessment, Review & Assurance.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Law & Governance

Origin pattern: Convergent development

Present-day reach: Multi-domain

Rationale: Purpose, proxy, counterexample, and necessity review closely follows legal scrutiny of differential treatment.

Related originating lineages:

  • Philosophy — Normative equality theory materially supplies the idea that only relevant differences justify unequal treatment.

Review resolution: Both blind reviewers agree that law_governance is the primary historical origin. Explicit reconciliation of alternate origin disagreement, origin mode disagreement adopts reviewer_a's evidence: Purpose, proxy, counterexample, and necessity review closely follows legal scrutiny of differential treatment. The selected record uses alternates=philosophy, origin_mode=convergent, and domain_reach=multi_domain; the other review proposed alternates=public_administration_policy, origin_mode=single_lineage, and domain_reach=multi_domain. The selected combination better preserves the mechanism-specific formative lineages and calibrated scope; broader present-day use is not treated as proof of additional historical origin.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Reconciled after independent review; high confidence.

Notes

[n1] Narrow tailoring — the requirement that a differentiating rule use the least-differentiating means that still serves its legitimate purpose. A feature survives only if no narrower feature achieves the same aim with less collateral differentiation.