Skip to content

Forecast Impact Audit

Post-release audit — instantiates Reflexive Forecast Impact Governance

Examines, after release, how a forecast actually moved behavior — comparing the reaction that occurred against the reaction that was modeled, and testing whether anyone gamed it — to tell a self-defeating forecast apart from a merely wrong one.

Once a forecast is out and the outcome is in, a raw accuracy number hides why. Forecast Impact Audit reconstructs the causal path: it takes the pre-release model of how the audience was expected to react, checks it against how the audience actually behaved, and runs an adversarial check for actors who bent the outcome by gaming the forecast rather than responding to it honestly. Its defining move is that it evaluates the forecast on its behavioral footprint — separating "the model of the world was wrong" from "belief in the forecast changed the world," including the case where someone engineered that change for gain.

Example

A grocery chain's internal model forecasts short supply of a staple, and the projection leaks to customers. Shelves empty within a day — not because supply actually fell, but because the forecast triggered stockpiling: a self-defeating forecast that manufactured its own shortage. A Forecast Impact Audit, run afterward, refuses the easy reading of "forecast of shortage → shortage → accurate." It compares the modeled reaction (a gradual demand uptick) against the observed one (a same-day run) and flags the divergence as forecast-induced. Its adversarial check then asks whether any actor — a reseller, an arbitrageur — amplified the signal to profit from the run. The verdict reframes an apparently accurate forecast as a governance failure: the disclosure caused the shortage, and the fix belongs in how the next such forecast is released, not in the supply model.

How it works

  • Refit the reaction model to realized data. Where observed behavior diverged from the pre-release model, quantify the gap and test whether the forecast's release explains it.
  • Run the adversarial check. Look for actors whose behavior only makes sense as exploitation of the forecast — front-running, arbitrage, panic amplification.
  • Classify the forecast. Wrong, right-but-inert, self-defeating-harmful, self-defeating-beneficial, or gamed.
  • Feed it forward. Route the finding into impact memory and into the release process, so the next disclosure of this type is designed around the reaction it will provoke.

Tuning parameters

  • Reaction-model fidelity — a coarse before/after comparison versus a structural model of the response. Higher fidelity isolates the forecast's effect but costs data and time.
  • Attribution window — how long after release behavior is attributed to the forecast. Too short misses slow reactions; too long absorbs unrelated noise.
  • Adversarial depth — a surface scan for obvious front-running versus a deep look for coordinated gaming. Depth catches sophisticated exploitation but risks reading normal behavior as manipulation.
  • Classification threshold — how large a modeled-versus-observed gap must be before a forecast is called "self-defeating" rather than merely noisy.

When it helps, and when it misleads

Its strength is that it catches the corruption a raw accuracy score cannot: Goodhart's law[n1] in action, where a forecast used as a signal becomes a target and stops measuring what it once measured. It turns "was the forecast accurate?" into the far more useful "what did releasing it do?", and it is the evidence base for redesigning disclosure.

Its failure mode is that after-the-fact causal attribution is genuinely hard: you can over-attribute — blaming the forecast for behavior with other causes — or under-attribute and miss a diffuse reaction. Adversarial checks can drift into seeing manipulation everywhere. The classic misuse is running the audit to assign blame after a bad outcome rather than to learn — an audit built to exonerate the forecaster will pick the attribution that acquits. The discipline is to pre-commit the attribution window and classification thresholds, and to treat the audit as input to redesign rather than a verdict on a person.

How it implements the components

Forecast Impact Audit realizes the realized-reaction side of the archetype — the components that can only be filled after the forecast has been acted on:

  • audience_reaction_model — it re-fits and validates this model against what the audience actually did; the pre-release version comes from siblings, but the realized fit is the audit's own product.
  • adversarial_reaction_check — it performs the post-release version: detecting actors who gamed the released forecast rather than responding to it.

It does not build the no-reaction counterfactual baseline — Avoided-Loss Counterfactual Review does — and it consumes, rather than owns, the live monitoring signal (Post-Release Behavior Dashboard) and the pre-release gaming simulation (Strategic Gaming Stress Test).

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: Forecast Impact Audit operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it examines, after release, how a forecast actually moved behavior — comparing the reaction that occurred against the reaction that was modeled, and testing whether anyone gamed it — to tell a self-defeating forecast apart from a merely wrong one.

Independent corroboration: The frozen evidence defines Forecast Impact Audit as 'Examines, after release, how a forecast actually moved behavior — comparing the reaction that occurred against the reaction that was modeled, and testing whether anyone gamed it — to tell a self-defeating forecast apart from a merely wrong one', so its operative form is Assessment, Review & Assurance.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Economics & Finance

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Economics is primary because expectations and published forecasts can alter prices, behavior, and the outcomes they predict. Foresight, feedback systems, behavioral economics, and technology governance contribute impact assessment; the named audit is a cross-disciplinary encyclopedia synthesis.

Related originating lineages:

Review resolution: Economics is primary because expectations and published forecasts can alter prices, behavior, and the outcomes they predict. Foresight, feedback systems, behavioral economics, and technology governance contribute impact assessment; the named audit is a cross-disciplinary encyclopedia synthesis.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Researched adjudication after independent review; high confidence.

Sources consulted:

Notes

The audit is only as sharp as the recorded pre-release reaction model it compares against. Without a captured expectation of how the audience was supposed to react — from the Reaction Channel Premortem or the release log — the audit has no baseline to detect divergence from, and collapses into anecdote. Its value depends on that expectation being written down before release.

[n1] "When a measure becomes a target, it ceases to be a good measure" — Charles Goodhart's law. A forecast that actors treat as a target to hit or beat stops describing the world and starts describing their gaming of it, which is exactly what the adversarial check is built to expose.