Skip to content

Outcome Monitoring Review

A monitoring workflow — instantiates Self-Fulfilling Prophecy Interruption

Tracks over time whether an interruption actually changed outcomes and treatment — and whether the loop simply moved to a subtler channel — so optimism does not replace one untested story with another.

An interruption can feel successful and still have failed, or worse, spawned a new distortion. Outcome Monitoring Review is the standing feedback loop of the archetype itself: after other mechanisms have cut a channel, it tracks over time whether outcomes and treatment gaps actually changed, whether confirmation is still being manufactured, and whether the loop simply relocated to a subtler channel. Its defining move is that it is downstream of every intervention and specifically hunts for the loop hopping channels — it is the reversibility guard that keeps an optimistic redesign from hardening into a second untested story. It measures; it does not intervene, and it explicitly closes back to revision when the evidence disappoints.

Example

A company redesigned onboarding for hires flagged "slow to ramp," expecting fairer outcomes. Six months on, the monitoring review does three things at once. It checks the outcome: did the flagged hires' 90-day performance and one-year retention converge with everyone else's, and did the flag stop predicting exits? It checks for a manufactured-confirmation residue: is any remaining gap still being produced by treatment? And it runs a new-distortion hypothesis by asking the flagged hires themselves how the change landed. The finding: retention improved, but flagged hires report being "handled with kid gloves" and under-challenged — the loop shifted from neglect to over-protection. Because the review caught the relocated channel, it triggers a further tweak rather than declaring victory.

How it works

Its distinguishing discipline is recurring, multi-signal verification aimed at channel-hopping:

  • Track the outcome over time — pre/post comparison of the results the interruption was meant to change.
  • Re-check the treatment gap — whether unequal treatment actually narrowed, and whether any surviving confirmation is still produced by it.
  • Ask the affected people — surface distortions the metrics miss and the intervener cannot see.
  • Watch for relocation — specifically test whether the loop moved to a subtler channel, and close back to revision if outcomes stall or harms appear.

Tuning parameters

  • Metrics and cadence — which outcomes are tracked and how often; too rare misses relocation, too frequent reads noise as signal.
  • Relocation signals — what counts as evidence the loop moved (over-help, new avoidance, a shifted metric); naming these in advance is what makes them detectable.
  • Weight on lived experience — how much stakeholder-reported experience counts against the quantitative outcome.
  • Revise/stop triggers — predefined thresholds at which the intervention is adjusted, reversed, or escalated.

When it helps, and when it misleads

Its strength is that it keeps the whole archetype evidence-based and reversible: it is the only mechanism that can catch a prophecy loop that has merely changed disguise, and it stops a hopeful redesign from being banked as proven.

It misleads when it fixates on a headline metric and misses the subtler relocated channel — or when it reads "no change yet" as failure and pulls a working intervention too early. The classic trap is Goodhart's law:[1] once a monitored number becomes the target, it can be moved without the underlying loop changing at all. The discipline is to pair quantitative outcomes with the stakeholder check, give interventions time before judging, and predefine the revise/stop triggers rather than rationalizing after the fact.

How it implements the components

  • outcome_monitor — its core: it tracks whether the changed interaction pattern actually changed outcomes without introducing a new self-confirming story.
  • reinforcing_outcome — it re-interrogates any surviving confirmation, asking whether it is still being produced by expectation-driven treatment.
  • stakeholder_impact_check — it asks affected people whether the treatment is now experienced as exclusion, over-surveillance, over-help, or neglect.

It does not build the initial loop map (expectation_map, Expectation Audit), perform any interruption (e.g. channel_interruption_choice, Blind Review), or back-test predictive accuracy net of confounds (base_rate_and_confound_check, Expectation Calibration Review).

Notes

It looks backward and forward differently than its review cousins. Expectation Calibration Review asks whether a past prediction was accurate net of its effects; this tracks whether a current intervention is working and staying honest. Its most important and most neglected job is the relocation check — a loop that has moved to a subtler channel will otherwise read as a success.

References

[1] Goodhart's law — "when a measure becomes a target, it ceases to be a good measure." Cited here as the monitor's central failure mode: a loop can be made to satisfy a tracked metric while continuing unchanged underneath, which is why the stakeholder check and the relocation test are not optional.