Skip to content

Campbell's Law

When consequential social decisions depend on a quantitative indicator, measured actors game it and distort the process the indicator was meant to monitor.

Core Idea

Campbell's Law describes what happens when an authority uses a quantitative social indicator to make consequential decisions about the people or institutions being measured. The indicator stops serving only as a passive description and becomes something actors must produce. They redirect effort toward improving the number, exploit classification and reporting rules, or change participation in the measured process. The indicator becomes more vulnerable to corruption, and the social process it was meant to monitor is distorted by the attempt to govern through it.[1]

The distinction from Goodhart's Law is taxonomic rather than merely rhetorical. Goodhart supplies the substrate-portable mechanism: optimization pressure on a proxy degrades its relationship with the target. Campbell is the social-governance species. It additionally requires a decision-making authority, a quantitative social indicator, consequential use of that indicator, measured actors capable of strategic response, and distortion of the governed social activity. Remove those institutional roles and the remaining optimizer-versus-proxy pattern is Goodhart's Law, not Campbell's Law.

Structural Signature

the social objectivethe quantitative social indicatorthe authority using it for consequential decisionsthe measured actorsstrategic adaptation and corruption pressuredistortion of the monitored social process

The pattern is present when each of the following holds:

  • A social objective. An authority cares about a difficult-to-observe property such as learning, health, public safety, research quality, or service performance.
  • A quantitative social indicator. A count, score, rate, rank, or threshold is adopted because it is thought to summarize that objective.
  • Consequential decision use. Funding, pay, sanctions, promotion, eligibility, reputation, closure, or another material decision depends on the indicator.
  • Measured actors with agency. People or organizations subject to the decision can change behavior, reporting, classification, selection, or participation in response.
  • Corruption pressure. Some indicator-improving actions do not improve the objective, and consequential use makes those actions comparatively attractive.
  • Process distortion. Adaptation changes not only what is recorded but how the monitored institution operates: curricula narrow, patients are selected, complaints are discouraged, or research programs reorganize around countable output.

Composed, these make the use of an indicator an intervention in the social process that generates it. A Campbell diagnosis is incomplete if it shows only proxy degradation without identifying who governs through the indicator, who adapts, and how the governed activity is deformed.

What It Is Not

  • Not competition. competition is rivalry over a scarce prize; Campbell's law is the corruption of a proxy once a stake is attached — agents game the measure, not each other, and the failure is proxy-target detachment, not who wins.
  • Not Goodhart's Law in general. Every Campbell case is a Goodhart case, but Goodhart also covers mechanical reward hacking and other non-social optimizers. Those cases lack Campbell's authority, social indicator, measured actors, and institution-distorting consequences.
  • Not cost-benefit analysis. cost_benefit_analysis weighs outcomes against costs; Campbell's law is about what happens when a measurement used in a decision becomes an active target of optimization, distorting the measure itself.
  • Not measurement disturbance. observer_effect and measurement_and_disturbance are non-strategic perturbations from the act of measuring; Campbell's law requires an adaptive agent responding to a stake, gaming its own fate under the measure.
  • Not adverse selection or moral hazard. adverse_selection and moral_hazard are hidden-type and hidden-action problems under information asymmetry; Campbell's law is proxy gaming under high stakes, which can occur with full information about who is doing the gaming.
  • Not increasing returns. increasing_returns describes self-reinforcing advantage; Campbell's law describes self-defeating measurement — the more a staked metric is optimized, the less it tracks its target.
  • Common misclassification. Diagnosing "Goodhart/Campbell" wherever a metric drifts. If nothing adapts its behavior to its own fate under the measure — the drift is non-strategic noise or honest measurement error — the decouple/harden/triangulate cures do not apply; ordinary error-correction does.

Scope of Application

  • Education — when test scores determine pay, funding, or promotion, instruction reshapes to the test (curriculum narrowing, teaching-to-the-test, score inflation, cheating scandals); the score rises while the underlying construct does not.[2]
  • Healthcare — surgical mortality league tables incentivize patient selection over better surgery; readmission penalties incentivize reclassification over better discharge.[3]
  • Policing — accountability for reported crime numbers produces reclassification and discouraged reporting, with measured crime diverging from victimization surveys.[4]
  • Corporate KPIs — quarterly-revenue targets generate channel stuffing; engagement targets generate engagement-bait; acquisition-cost targets generate fraudulent attribution.
  • Academic publishing — citation counts and impact factors used for hiring generate citation cartels, salami-slicing, and p-hacking.[5]
  • Economic policy — when an inflation measure becomes the policy target, the underlying construct detaches from the measured index as the basket is reshaped.
  • Platform growth metrics — engagement metrics used as success criteria drive product changes (autoplay, infinite scroll, notification spam) at the cost of the user value the metric proxied.[6]

Across these settings the institutional domain changes, but the governance substrate remains: an authority attaches decisions to a quantitative social indicator, measured actors adapt, and the attempt to manage through the number changes the process being managed.

Clarity

Campbell's clarifying move is the four-way separation of social objective, indicator, decision use, and measured actors, which everyday talk about "metric gaming" runs together. The indicator may have been informative before it governed anyone. Once funding, sanctions, promotion, or reputation depend on it, the people who generate the record face a different problem and the record is produced under a different regime.

A second clarification is causal. The damage is not limited to a less trustworthy number. High-stakes use can redirect instruction, patient selection, complaint recording, publication strategy, or service delivery. Campbell's Law therefore asks both whether the indicator still measures the objective and what the accountability regime has done to the underlying institution.

Manages Complexity

The pattern lets a policy or organizational designer predict, before attaching consequences, which responses a social indicator will invite. A large and otherwise bewildering space of "how will this go wrong?" questions becomes a structured audit: who is measured, which decisions depend on the number, which cheap actions move it, which valuable activities those actions displace, and which independent observations remain outside the reporting chain.

This also separates three intervention surfaces that are easily confused: redesign the indicator, change the consequences attached to it, or protect the social process from being reorganized around what is countable. The distinction matters because a technically improved measure can still deform practice when used as a single high-stakes decision rule.

Abstract Reasoning

The load-bearing reasoning is reflexive governance. A social indicator is generated by the process it describes; consequentially using the indicator changes the incentives and classifications inside that process; the altered process then generates a different indicator. Evidence that the number tracked the objective under observation does not automatically transport into the regime in which the number allocates reward or punishment.

The domain-specific interrogation is therefore: What decisions will this indicator control? Who can alter the indicator or the population it describes? Which unmeasured activities will become less attractive? What observation remains independent after the accountability loop closes? The portable signal-versus-target logic underneath these questions belongs to Goodhart's Law; the authority–indicator–actor loop and institutional deformation are Campbell's additional structure.

Knowledge Transfer

The structure carries interventions across social institutions. Separate management from evaluation: do not ask the same visible, consequential indicator to direct behavior and provide an independent account of whether the behavior worked. Triangulate across reporting chains: combine measures whose errors and gaming channels differ. Audit the objective directly: use independent surveys, blind site visits, sampled low-stakes assessments, or outcome reviews that measured actors cannot easily anticipate. Inspect selection and classification: ask who disappears from the denominator, which cases are relabeled, and which hard cases are avoided. Distribute consequences: avoid letting one threshold or ranking decide too much.

These interventions transfer literally among schools, hospitals, police departments, research institutions, public agencies, and firms because each contains authorities, quantified social performance, consequential decisions, and actors who can reorganize practice around the indicator. Reward-function design in machine learning may recommend analogous techniques, but that transfer is carried by Goodhart's broader prime rather than by Campbell's institutional identity.

Examples

Formal/abstract

Suppose an authority wants to improve a difficult-to-observe social objective \(T\), selects an indicator \(M\) because \(M\) correlated with \(T\) before consequential use, and then makes a decision \(D\) depend heavily on \(M\). The measured actors choose behavior from a set containing both value-producing actions \(V\), which improve \(T\) and usually \(M\), and indicator-producing actions \(G\), which improve \(M\) without improving \(T\). Consequential use changes the actors' payoff ordering: actions in \(G\) that were formerly pointless become attractive, while valuable but unmeasured actions can be displaced. As effort shifts from \(V\) toward \(G\), the \(M\)-\(T\) relationship weakens and the composition of the social process changes.

Mapped back: \(T\) is the social objective, \(M\) the quantitative social indicator, \(D\) its decision use, the people or organizations subject to \(D\) are the measured actors, \(G\) is the corruption channel, and displacement of \(V\) is the process distortion that distinguishes Campbell's institutional species from bare proxy degradation.

Applied/industry

High-stakes standardized testing in education is Campbell's Law in its home substrate, with the full institutional structure visible. The social objective is student learning; the indicator is the test score; the decision use is teacher evaluation, school funding, intervention, or closure. Teachers and administrators can raise scores through genuine learning, but also through curriculum narrowing, rehearsal of tested formats, reclassification or exclusion of low-scoring students, and in documented scandals, answer-changing. Scores can rise while independent low-stakes assessments show that gains do not transfer.[7]

The damage is dual. The score becomes a less trustworthy indicator of learning, and the educational process itself is reorganized around tested content and threshold management. The intervention menu follows from those two failures: preserve an independent low-stakes assessment for tracking, audit exclusions and classifications, combine evidence whose gaming channels differ, and reduce the amount any single score decides. The same social-governance pattern appears in healthcare when mortality rankings encourage patient selection and coding changes rather than better care.[3]

Mapped back: High-stakes testing realizes the abstraction end-to-end — learning as the social objective, test scores as the indicator, accountability decisions as the consequential use, teachers and administrators as measured actors, curriculum narrowing and exclusion as corruption channels, and altered instruction as distortion of the monitored social process.

Structural Tensions

T1 — Stake intensity versus detachment speed (sign/direction). Campbell's Law predicts heavier stakes accelerate corruption pressure — but a stake light enough to avoid gaming may be too weak to drive any behavior, leaving the indicator inert. The failure mode is the dilemma's far horn: removing stakes to preserve diagnostic value while losing the management leverage the stake existed to provide. Diagnostic: ask whether the indicator is meant to manage or to track; a single visible number asked to do both will either fail to motivate or be gamed, and the resolution is usually to separate the functions.

T2 — Cheap-to-move versus value-producing paths (measurement). Residual diagnostic power depends on how many cheap indicator-moving actions also improve the social objective — but the dangerous response is often the one the designer did not imagine, such as reclassification, denominator management, or selective intake. The failure mode is auditing the obvious moves, certifying the indicator safe, and being blindsided by a creative institutional response. Diagnostic: assume measured actors know operational loopholes the designer does not; if integrity depends on enumerating every exploit, use independent target audits rather than confidence in the list.

T3 — Multi-metric robustness versus combined gaming (scalar). Coupling a staked metric to orthogonal measures raises the cost of faking, but a sufficiently adaptive agent gamed against the bundle will find moves that satisfy all of them jointly, and more metrics also dilute focus and multiply gaming surface. The failure mode is adding measures faster than they add genuine orthogonality, producing a dashboard that looks robust but shares a common exploit. Diagnostic: ask whether a single behavior can move all coupled metrics in the desired direction; if such a joint path exists and is cheap, the metrics are not orthogonal and the triangulation is illusory.

T4 — Measure refresh versus longitudinal comparability (temporal). Rotating or redesigning metrics faster than agents adapt defeats gaming, but each refresh breaks the time series, so the system loses the ability to compare across periods and to detect slow real trends. The failure mode is churning metrics so often that gaming is suppressed while no measure survives long enough to reveal whether the target is actually improving. Diagnostic: weigh the gaming half-life against the horizon over which you need to track real change; if refresh is faster than the target moves, you have traded corruption for blindness and can no longer tell improvement from noise.

T5 — Strategic gaming versus honest measurement disturbance (scopal). Campbell's corruption is driven by an adaptive agent responding to a stake; it is a different failure from physical or statistical measurement disturbance, where the act of measuring perturbs the system without strategic intent. The failure mode is misdiagnosing a drifting metric as gaming (and adding anti-gaming audits) when the drift is non-strategic, or vice versa — treating a genuine exploit as innocent measurement error. Diagnostic: ask whether something adapts its behavior to its own fate under the measure; only then is it Campbell's law, and only then do the decouple/harden/triangulate interventions apply rather than ordinary measurement-error correction.

T6 — Detaching the stake versus governability (scopal/sign). "Stop staking the measure" is the most effective cure, but a system with no consequential metric loses its handle on the agents entirely, and management reverts to unaccountable judgment that may be worse than a partly-gamed number. The failure mode is purging all staked metrics in the name of integrity and producing an organization that can no longer steer, reward, or detect failure. Diagnostic: before detaching a stake, ask what governs behavior in its absence; if the answer is "nothing observable," the gamed metric may still be more accountable than its removal, and the better move is to harden or triangulate rather than abandon measurement.

Structural–Framed Character

Campbell's Law is strongly framed on the structural–framed spectrum. Its defining nouns are not incidental examples: an authority, a quantitative social indicator, consequential administrative use, measured people or organizations, and distortion of a governed institution. Institutional origin, human-practice boundedness, and import-versus-recognize therefore sit on the framed pole. The target–indicator distinction remains formally expressible, and the law need not morally condemn the measured actors, so vocabulary and evaluative weight are less completely framed. The resulting aggregate of 0.8 describes an abstraction whose explanatory value is real but whose identity depends on social-governance furniture.

Structural Core vs. Domain Accent

The structural core is the Goodhart mechanism: a proxy that was informative under one regime loses fidelity when optimization pressure is applied to it. That core travels to machine learning, biological selection, and other non-institutional substrates without Campbell's Law.

Campbell's domain accent is identity-bearing rather than decorative. The proxy must be a quantitative social indicator used by an authority for consequential social decisions. People or organizations subject to those decisions adapt, producing corruption pressure on the indicator and changing the social process being monitored. The abstraction transfers literally across education, healthcare, policing, public administration, firms, and academic governance because those settings preserve all of these roles. Its composite substrate-independence score is therefore 2 / 5: broad within the social-institutional domain, formally clear, and empirically well supported, but not an independent cross-substrate prime.

  • Composite substrate independence — 2 / 5
  • Domain breadth — 3 / 5
  • Structural abstraction — 3 / 5
  • Transfer evidence — 3 / 5
  • Goodhart's Law is the strict parent. Campbell's Law adds the authority–social-indicator–measured-actor–process-distortion differentia to Goodhart's proxy-collapse mechanism.
  • Proxy-Target Divergence is inherited through Goodhart's Law. A direct Campbell edge would flatten the nearer taxonomic parent.
  • Measurement, Incentive, and Accountability supply important roles, but none by itself contains the complete Campbell mechanism.
  • Observer Effect remains a contrast: measurement can perturb a system without actors strategically responding to consequential use of an indicator.

Relationships to Other Abstractions

Local relationship map for Campbell's LawParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Campbell's LawDOMAINPrime abstraction: Goodhart's Law — is a kind ofGoodhart's LawPRIME

Current abstraction Campbell's Law Domain-specific

Parents (1) — more general patterns this builds on

  • Campbell's Law is a kind of Goodhart's Law Prime

    Campbell's Law is Goodhart's Law specialized to consequential use of quantitative social indicators, where measured actors game the indicator and distort the social process it governs.

Not to Be Confused With

The most important boundary is with Goodhart's Law. Goodhart is the strict parent, not a duplicate: it requires a proxy placed under optimization pressure and covers mechanical reward hacking or intentless selection. Campbell requires the narrower social-governance configuration—a quantitative social indicator used for consequential decisions, measured human or institutional actors, corruption pressure, and distortion of the monitored social process. Every Campbell case is Goodhart; not every Goodhart case is Campbell.

Campbell's Law is also not Competition. Competition is rivalry among agents over a scarce prize. Campbell concerns governance through an indicator; a sole provider can manipulate a readmission measure without competing with anyone. Contest design and indicator-governance design therefore address different mechanisms.

Finally, Campbell's Law is not Observer Effect or generic measurement disturbance. Those can occur when measurement physically or statistically perturbs a system without anyone responding to a decision rule. Campbell requires measured actors adapting to consequential use of a social indicator. If no one changes behavior, reporting, classification, selection, or participation because of what the indicator decides, ordinary measurement error or disturbance is the nearer diagnosis.

References

[1] Campbell, Donald T. "Assessing the Impact of Planned Social Change." Evaluation and Program Planning, vol. 2, no. 1 (1979): 67–90. Original statement of Campbell's law: the more a quantitative social indicator is used for decision-making, the more it is subject to corruption pressures and the more it distorts what it monitors.

[2] Nichols, Sharon L., and David C. Berliner. Collateral Damage: How High-Stakes Testing Corrupts America's Schools. Cambridge: Harvard Education Press, 2007. Documents curriculum narrowing, teaching to the test, score inflation, and cheating under high-stakes testing.

[3] Bevan, Gwyn, and Christopher Hood. "What's Measured Is What Matters: Targets and Gaming in the English Public Health Care System." Public Administration, vol. 84, no. 3 (2006): 517–538. Documents reclassification and patient-selection gaming of staked health-care and public-sector targets.

[4] Eterno, John A., and Eli B. Silverman. The Crime Numbers Game: Management by Manipulation. Boca Raton: CRC Press, 2012. Documents crime-statistic reclassification and discouraged reporting under accountability for reported crime numbers.

[5] Biagioli, Mario. "Watch Out for Cheats in Citation Game." Nature, vol. 535 (2016): 201. Describes citation cartels, salami-slicing, and gaming of citation and impact metrics used for hiring.

[6] Vivrekar, Devangi. Persuasive Design Techniques in the Attention Economy: User Awareness, Theory, and Ethics. Master's thesis, Stanford University, 2018. Documents engagement-maximizing persuasive design (infinite scroll, autoplay, notification prompts) on platforms that profit by maximizing time-on-site, at the cost of the user value the metric proxied.

[7] Koretz, Daniel. The Testing Charade: Pretending to Make Schools Better. Chicago: University of Chicago Press, 2017. Shows score gains under high-stakes testing failing to transfer to independent low-stakes assessments (e.g., NAEP), the detachment invariant observed directly.