Skip to content

Post-Outcome Recalibration Review

Review loop — instantiates Heuristic Calibration and Confidence Judgment

A scheduled loop that ingests newly-resolved outcomes, re-fits the heuristic's confidence controls, and reassigns owned follow-up as the world drifts.

A Post-Outcome Recalibration Review is the longitudinal loop that keeps a heuristic's confidence honest over time. Calibration is not a one-time fit: the environment drifts, adversaries adapt, populations shift, and a correction that was right last year silently goes stale. This review closes the loop — it waits for judgments to resolve, gathers the freshly-observed outcomes, compares them against the confidence that was claimed, and re-fits the controls (caps, thresholds, corrections) so that trust tracks current reliability rather than the reliability of a bygone regime. Its defining property is that it is outcome-triggered and owned: it fires when new resolutions accumulate or a regime change is suspected, it belongs to a named owner accountable for keeping calibration current, and it produces updated boundaries, not just an observation. Where a static audit describes the past, this loop reaches forward and adjusts the machinery.

Example

An intelligence shop runs geopolitical forecasts — "70% chance this border tension de-escalates within 90 days" — that feed real resource decisions. Quarterly, once a batch of forecasts has resolved, a designated calibration owner convenes a recalibration review. This isn't a one-off audit; it's the standing loop. The team pulls every forecast whose window has closed, scores confidence against what actually happened, and looks specifically for drift: are the forecasters, well-calibrated last year, now overconfident on a particular region since a new government took power? They find exactly that — calls about one country have decayed since the regime change, while the rest hold. The review's product is a set of updated controls: a tighter confidence cap for that region, a new boundary condition noting the regime shift, and an owner assigned to watch it. Next quarter the loop runs again and checks whether the cap did its job. The discipline mirrors large forecasting tournaments, where scoring resolved predictions and feeding the result back is what separates improving forecasters from static ones.[n1]

How it works

The loop runs on a cadence tied to when judgments actually resolve — there is no point reviewing forecasts whose windows are still open. Each cycle: collect the newly-resolved outcomes with their original claimed confidence; score calibration on that fresh slice; and — the move that distinguishes this from a one-shot analysis — compare against prior cycles to detect drift, decay, or a regime break. Findings are converted into updated controls: re-fit caps, adjusted escalation thresholds, revised boundary conditions marking where the environment has changed. A named owner is accountable for the re-fit landing and for watching the flagged boundaries until the next cycle confirms or clears them. The loop is deliberately periodic and forward-acting: its output is a changed system plus a watch-list, and its success is measured next cycle by whether the previous cycle's changes worked.

Tuning parameters

  • Review cadence — how often the loop fires. Frequent cycles catch drift early but often review too few freshly-resolved cases to say anything; sparse cycles have solid samples but let the heuristic run stale between them.
  • Resolution-lag handling — how the loop treats judgments whose outcomes take a long time to arrive. Waiting for full resolution is accurate but slow to detect drift; using early proxies is timely but noisier.
  • Drift-detection sensitivity — how large a shift between cycles counts as real decay versus sampling wobble. Sensitive settings react fast but chase noise; insensitive ones are stable but slow to catch a regime break.
  • Ownership scope — how much authority the calibration owner has to re-fit controls unilaterally versus recommend. Broad authority is fast; narrow authority is safer but can let known drift persist while approval is sought.

When it helps, and when it misleads

Its strength is that it is the only mechanism here that fights staleness: it keeps calibration matched to the present world, catches decay before it compounds, and — through named ownership — prevents calibration from becoming everyone's job and therefore no one's. It turns resolved outcomes into a standing engine of improvement rather than a one-time report.

Its failure mode is chasing noise across cycles: with thin per-cycle samples, the loop re-fits controls to random swings, thrashing the system and destabilizing the very confidence it means to steady. A subtler trap is hindsight bias in the review room — once the outcome is known, the original confident-wrong call feels obviously wrong, and the review "corrects" for a pattern that was not predictable ex ante. The classic misuse is the no-action recalibration: the loop dutifully scores outcomes and files a report, but nobody owns re-fitting the controls, so the finding never changes behavior. The guarding discipline is to require a minimum resolved sample before re-fitting, to judge calls against what was knowable at claim time, and to make the owner accountable for a concrete control change each cycle or an explicit "hold, insufficient evidence."

How it implements the components

  • feedback_collection_loop — it is the loop: a recurring intake of newly-resolved outcomes fed back to update the heuristic's calibration.
  • calibration_owner — it names and empowers the person accountable for the re-fit landing and for watching flagged boundaries between cycles.
  • boundary_condition_register — each cycle updates the register with newly-detected regime shifts and decay zones where confidence must now be lowered.

It does not, by itself, capture each forecast and its confidence at claim time — that prospective record is the heuristic_track_record_evidence built by Prediction Journal, which this loop consumes. Nor does it compile a single standing numeric transform on every claim — that is Calibration Adjustment Rule; this review is the periodic event that decides when such a rule must be re-fit, not the always-on rule itself.

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: Post-Outcome Recalibration Review operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it a scheduled loop that ingests newly-resolved outcomes, re-fits the heuristic's confidence controls, and reassigns owned follow-up as the world drifts.

Independent corroboration: The frozen evidence defines Post-Outcome Recalibration Review as 'A scheduled loop that ingests newly-resolved outcomes, re-fits the heuristic's confidence controls, and reassigns owned follow-up as the world drifts', so its operative form is Assessment, Review & Assurance.

Nearest alternative: Intervention, Treatment & Transformation — Post-Outcome Recalibration Review includes features of a direct treatment or transformation applied to a target to change its state or condition, but its defining operation is a bounded evaluation of existing evidence or work that produces a finding or disposition.

Review outcome: Independent reviewer agreement; medium confidence.

Origin Attribution

Primary origin: Psychology

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Recalibrating judgment confidence against newly resolved outcomes derives from psychological forecasting and decision research.

Related originating lineages:

Review resolution: Light authoritative-source research resolves the primary-origin disagreement in favor of psychology. PubMed: Context, Feedback, and Confidence Calibration directly documents the defining practice or theory described in the selected origin rationale. Other domains are retained only where the blind reviews identify material co-development or translation; broad application is recorded separately as domain_reach=multi_domain, while origin_mode=cross_disciplinary_synthesis describes the relationship among origin lineages.

Attribution caveat: The boundary with statistics experimental design is substantive because that tradition materially developed or translated part of the mechanism; the cited provenance places the defining form in psychology.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Researched adjudication after independent review; high confidence.

Sources consulted:

Notes

[n1] The Good Judgment Project, led by Philip Tetlock, ran large forecasting tournaments in which participants' resolved predictions were scored and the feedback fed back into their practice. The finding that systematic outcome feedback measurably improves calibration is the empirical warrant for running recalibration as a standing loop rather than a one-time exercise.