Proxy-Relevance Audit¶
Test / assessment — instantiates Relevance-Substitution Detection and Correction
Checks whether a proxy metric is target-faithful enough to support the weight of the update it is being asked to carry.
Proxy-Relevance Audit is aimed at one specific kind of substitute signal: a measurable proxy — a score, index, rate, or metric — that has quietly come to stand in for a target it only imperfectly tracks. The audit sets the fidelity bar the decision actually requires, tests how faithfully the proxy tracks the target under current conditions, and then caps the update the proxy is allowed to drive to the weight its fidelity earns. Its organizing idea is load versus fidelity: a proxy can be good enough to inform a low-stakes nudge and nowhere near good enough to carry a high-stakes ranking, so the question is never "is the proxy valid?" in the abstract but "is it faithful enough for this weight?" What makes it distinct is this coupling of a required standard to a permitted weight for a metric; it does not type the evidential relation of an arbitrary active cue in the abstract, and it does not merely record a correction after the fact — it decides how much a measure may be trusted and enforces that ceiling.
Example¶
A digital-marketing team has been treating click-through rate as the metric of ad effectiveness, and a campaign with a high CTR is about to be scaled fivefold on that basis. The Proxy-Relevance Audit intervenes. First it sets the standard: the target is incremental purchases, and to justify a 5× budget increase the proxy must track incremental purchase behavior faithfully, not just attention. Then it tests fidelity: pulling the funnel data, the team finds this campaign's clicks convert to purchases at a fraction of the baseline rate — the creative is attracting curiosity clicks, not buyers, so CTR and the purchase target have come apart for exactly this campaign. The proxy is faithful enough to compare headline appeal but far too weak to carry a spend decision. The audit caps CTR's role accordingly: it may inform creative iteration, but the scale-up decision is reweighted onto conversion and incremental-lift evidence, and the 5× is put on hold pending a proper holdout test.
How it works¶
- Fix the required fidelity. State the target the proxy stands for and how faithfully it must track that target to bear the specific decision weight at stake — the bar is set by the load, not by the proxy.
- Test target fidelity. Examine how well the proxy actually tracks the target under present conditions: correlation, coverage, and especially whether the two have decoupled for this case.
- Compare fidelity to load. Judge the measured fidelity against the required bar for the weight being placed on it.
- Cap the weight. Set a ceiling on how much the proxy may drive the update, so a metric that passes for a light decision is not allowed to carry a heavy one.
Tuning parameters¶
- Fidelity bar height — how faithful the proxy must be for a given decision weight. A high bar prevents over-trusting weak metrics but can stall decisions that must run on imperfect proxies; a low bar keeps things moving but readmits substitution.
- Decoupling sensitivity — how alert the test is to the proxy and target coming apart in the specific case versus on average. High sensitivity catches local breakdowns but can over-reject a broadly sound proxy; low sensitivity trusts the historical relationship too far.
- Weight-cap severity — how hard the ceiling bites when fidelity falls short. A severe cap contains gaming and drift; a soft cap preserves the proxy's usefulness at the cost of some leakage.
- Re-audit cadence — one-time versus recurring. Proxies decay and get gamed, so recurring audits catch drift but cost ongoing effort.
When it helps, and when it misleads¶
Its strength is refusing the false binary of trusting or banning a metric: by tying a permitted weight to a measured fidelity, it lets useful-but-imperfect proxies do the work they can bear and no more. It is the working defense against Goodhart's law — that a measure adopted as a target degrades as the proxy is optimized apart from the thing it was meant to track.[1] Capping the proxy's weight to its fidelity is what keeps a convenient number from silently becoming the goal.
Its failure mode is demanding a fidelity that no available proxy can meet and thereby stalling into paralysis, or — the opposite — rubber-stamping a proxy because it correlated with the target once, historically, while ignoring that it has since decoupled. The classic misuse is standard-shopping: choosing whichever fidelity definition lets the favored metric clear the bar. The guarding discipline is to fix the fidelity standard from the decision's stakes before looking at how the proxy performs, to test for present-case decoupling rather than trusting the average relationship, and to route a proxy that cannot meet the bar to reconstructed direct evidence or an explicit "insufficient" state rather than using it anyway.
How it implements the components¶
relevance_standard— the required-fidelity bar, set from the decision's stakes, is the standard the proxy must meet.relevance_test— the target-fidelity test measures how faithfully the proxy tracks the target under current conditions.correction_and_reweighting_rule— the weight cap is the rule that limits how much the proxy may drive the update.
It does not name a cue's psychological_activation_pathway or register an arbitrary active cue in a substitute_signal_register and type its relation in the abstract — that general single-cue relation-typing belongs to its close cousin Cue-Validity Audit; this assessment is specific to a measurable proxy standing in for a target.
Related¶
- Instantiates: Relevance-Substitution Detection and Correction — provides the fidelity-versus-load assessment that decides how much a proxy metric may be trusted.
- Sibling mechanisms: Affect–Evidence Split Prompt · Red Herring Filter Checklist · Analogy Mapping Table · Question–Evidence Matrix · Decision-Basis Disclosure · Reweighted Update Log · Cue-Validity Audit · Blinded or Masked Review
Editorial Notes¶
Form Classification¶
Form family: Assessment, Review & Assurance
Rationale: Proxy-Relevance Audit operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it checks whether a proxy metric is target-faithful enough to support the weight of the update it is being asked to carry.
Independent corroboration: The frozen evidence defines Proxy-Relevance Audit as 'Checks whether a proxy metric is target-faithful enough to support the weight of the update it is being asked to carry', so its operative form is Assessment, Review & Assurance.
Review outcome: Independent reviewer agreement; high confidence.
Origin Attribution¶
Primary origin: Statistics & Experimental Design
Origin pattern: Cross-disciplinary synthesis
Present-day reach: Universal
Rationale: Testing whether a proxy validly bears on a target construct is rooted in statistical measurement and construct-validity practice.
Related originating lineages:
- Philosophy — Epistemology supplies the normative distinction between relevant evidence and substitution of an easier question.
- Psychology — Psychometrics materially developed construct, criterion, and convergent validity for indirect measures.
Review resolution: Both blind reviewers agree on statistics_experimental_design as the primary origin. Explicit reconciliation resolves alternate_origin_disagreement, origin_mode_disagreement, domain_reach_disagreement. The merged alternate lineages retain only domains the reviewers identified as materially formative; domain_reach=universal records later applicability separately from origin breadth.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Reconciled after independent review; high confidence.
References¶
[1] Goodhart, Charles A. E. "Problems of Monetary Management: The U.K. Experience". In Papers in Monetary Economics, Reserve Bank of Australia, 1–20, 1975. States that an observed statistical regularity tends to collapse when pressure is placed on it for control purposes. registry ↩