Skip to content

Metric Gaming Review

Review and audit mechanism — instantiates Goal Congruence Alignment

A recurring audit that asks whether a metric improved because the real outcome improved, or because someone found a cheaper way to move the number.

Metric Gaming Review is the recurring investigation that asks one question of an improving number: did the outcome actually get better, or did someone find a cheaper path to the metric? Its defining move is detection after the fact — it takes a metric that has already moved and looks for the gap between the measured gain and the real gain, hunting for the specific mechanism by which the number was lifted without the outcome following, and for the costs that were quietly shifted elsewhere to lift it. It does not design measures, set rewards, or display data; it inspects live measures for corruption, treating a suspiciously good number as a hypothesis to be tested rather than a success to be celebrated. Where design-time guardrails fence off the anticipated ways to cheat, this review catches the ones that emerged in the wild.

Example

A utility's customer-service operation is measured heavily on average handle time — how long agents spend per call — and the number has been improving for two quarters. Leadership is pleased; the gaming review is skeptical, and runs the audit that turns the number into a question.

It starts by looking for the mechanism of improvement rather than trusting the trend. It finds it: handle time dropped because agents learned to end calls fast — transferring hard cases to another queue, closing tickets as "resolved" before confirming the fix, and steering callers to hang up and try the website. The review's diagnosis is that measured improvement has decoupled from the real outcome, first-contact resolution, which has actually fallen. Then it traces where the cost went — the register of shifted burden: repeat-call volume is up in a different team's queue, complaint escalations have risen, and the "saved" agent-minutes reappeared as customer hours spent calling back. The audit's product is not a new metric and not a punishment; it is a documented finding — the handle-time gain is a gaming artifact, and here is the outcome it hollowed and the burden it exported — handed to the mechanisms that will repair the measure and the reward.

How it works

  • Treat the good number as a hypothesis. Start from an improving metric and ask what else could explain the improvement besides the outcome getting better.
  • Find the gaming mechanism. Look for the specific cheaper path — reclassification, cherry-picking, symptom-fixing, threshold-timing, definition-bending — that could move the measure without moving the outcome.
  • Check the metric against a harder-to-game witness. Compare the measure to an independent signal of the real outcome; a divergence between them is the gaming signature.
  • Register the shifted costs. Trace where the burden went — another team's queue, a later time period, a third party — and log it, because gaming usually exports a cost rather than eliminating it.

What distinguishes it is that it is detective and periodic: it produces findings about a live measure's integrity, not a redesigned measure, a reward, or a dashboard.

Tuning parameters

  • Review cadence — how often the audit runs. Frequent catches gaming before it entrenches but burns scrutiny and can feel like surveillance; rare lets distortions compound.
  • Suspicion threshold — how anomalous a trend must look before it triggers a deep dive. Low thresholds catch subtle gaming but raise false alarms and reviewer load; high ones conserve effort but miss slow corruption.
  • Witness independence — how independent the comparison signal is from the metric under review. A truly independent witness (an outcome the gamer can't also game) makes detection sharp; a correlated one can be fooled by the same trick.
  • Investigation depth — how far the audit traces shifted costs across teams and time. Deeper finds exported burden the surface hides but costs investigator time and can strain relationships.
  • Response coupling — whether findings feed automatically into metric repair and incentive change, or sit in a report. Tight coupling makes the review consequential; loose coupling risks audit-as-theatre.

When it helps, and when it misleads

Its strength is that it protects the invariant the whole archetype rests on: measured improvement must stay causally connected to real improvement. It is the standing immune response to Goodhart-style corruption, catching the gaming that design-time guardrails did not anticipate.

Its failure mode is captured by Campbell's law — the more a quantitative indicator is used for high-stakes decisions, the more it will be gamed and the more it will corrupt the process it monitors — which applies to the review's own findings if they harden into a new target to dodge.[n1] The classic misuse is running the review as a blame hunt: framed as catching cheaters, it drives concealment and defensiveness, and people get better at hiding gaming rather than stopping it. It can also over-fire, mistaking a legitimate efficiency for a trick. The guarding discipline is to aim the review at the structure that made gaming rational rather than at individuals, to hand findings to metric repair and incentive redesign rather than to punishment, and to keep an independent witness of the real outcome so the audit is judging against something the gamer could not also bend.

How it implements the components

  • misalignment_diagnosis — its core act is diagnosing the gap between a measure's improvement and the real outcome, naming the specific gaming mechanism by which the number moved without the goal following.
  • externality_register — it traces and logs where the metric's gains were actually paid for — another team's queue, a later period, a third party — documenting the burden the gaming exported rather than eliminated.

It detects and diagnoses a live measure; it does not build or repair the metric or its bounding countermetric (metric_redesign, countermetric_guardrail) — that is Shared Metric Design, its nearest neighbor, which constructs gauges while this review inspects them — and it holds no standing display of outcomes (system_outcome_monitor), which is System Outcome Dashboard.

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: Metric Gaming Review operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it a recurring audit that asks whether a metric improved because the real outcome improved, or because someone found a cheaper way to move the number.

Independent corroboration: The frozen evidence defines Metric Gaming Review as 'A recurring audit that asks whether a metric improved because the real outcome improved, or because someone found a cheaper way to move the number', so its operative form is Assessment, Review & Assurance.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Economics & Finance

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Universal

Rationale: Strategic response to measured incentives follows economic principal-agent and Goodhart-style reasoning.

Related originating lineages:

Review resolution: Both independent reviews place the primary provenance in economics_finance. The queued differences (alternate_origin_disagreement, domain_reach_disagreement) concern secondary metadata, not primary lineage. The final retains organizational_management, public_administration_policy, statistics_experimental_design only where a reviewer supplied a formative-lineage rationale; downstream use or broad applicability by itself is not treated as origin. origin_mode=cross_disciplinary_synthesis because the supplied rationales identify formative contributions that are composed in the mechanism's present form. domain_reach=universal records established application breadth separately from provenance. confidence=high preserves the more cautious evidence assessment. encyclopedia_synthesis=false records whether either reviewer identified deliberate corpus-level composition.

Review outcome: Reconciled after independent review; high confidence.

Notes

The review's power is that it looks backward at what already happened, which is also its limit: it can only catch gaming that has already occurred and left a trace. It pairs naturally with the design-time countermetric guardrail of Shared Metric Design and Incentive Redesign — the guardrail blocks the anticipated cheats before they happen, and this review catches the unanticipated ones after they do.

[n1] Campbell's law — formulated by social scientist Donald T. Campbell: "The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor." It is the reason metrics used for high-stakes decisions need a standing gaming review, and the reason that review must not become just another target to dodge.