Skip to content

Post-Decision Calibration Review

Workflow — instantiates Bounded-Rationality Decision Design

Compares forecast, confidence, process, and outcome to recalibrate future thresholds and methods.

A Post-Decision Calibration Review is the periodic look backward over many decisions that have already run their course, comparing what the process predicted — and how confident it was — against what actually happened, and feeding the gap back into the design's dials. The idea that makes it this mechanism and not a sibling is that it operates on the aggregate track record, not the live case: by the time it runs, every decision it examines is closed and its outcome known. It is not trying to get any one decision right; it is trying to make the next hundred decisions better by retuning the thresholds and method assignments that produced this batch. Its unit of work is a population of past forecasts, and its output is a corrected setting.

Example

A regional weather forecasting office issues a probability of precipitation with every daily forecast — "70% chance of rain tomorrow." Any single forecast is impossible to grade: it either rained or it didn't. The Post-Decision Calibration Review is the quarterly workflow that grades them in bulk. Analysts pull every forecast that carried a "70%" tag over the last season and check the base rate: did it rain on close to 70% of those days? If it rained on only 55% of them, the office's 70% is systematically overconfident, and the review says so with a calibration curve and a Brier score[1] that quantifies the gap. It then does the thing that makes it worth running: it changes the process going forward — nudging the guidance the forecasters lean on, and tightening the confidence band at which a forecast auto-publishes versus gets a second look. No individual forecast is re-litigated; the machinery that generates confidence is recalibrated against its own history.

How it works

The review works by (1) assembling a batch of closed decisions along with the forecast, confidence, and method each one recorded at the time; (2) scoring predicted against realized outcomes as a distribution — not case by case, but as calibration and error rates across the batch; (3) diagnosing the pattern (systematic over- or under-confidence, a method that underperforms on a segment, a threshold set in the wrong place); and (4) writing the correction back into the design's settings and confidence bands. Its distinctive move is the comparison of stated confidence to realized frequency: a decision can be individually "right" and still reveal a badly calibrated process, and a well-calibrated process still gets individual calls wrong. The review is indifferent to the former and hunts the latter.

Tuning parameters

  • Review cadence and batch size — how often and over how many decisions; larger batches give statistical power but slow the feedback and can average over a regime change.
  • Scoring rule — which proper score or calibration metric grades the forecasts; different rules penalize over- and under-confidence differently.
  • Segmentation — whether calibration is checked overall or split by case type, method, or affected group; finer cuts catch a method that fails on one segment but cost sample size.
  • Correction aggressiveness — how hard a detected miscalibration moves the thresholds; over-correcting chases noise, under-correcting lets a known bias persist.
  • Look-back window — how much history counts as current; too long and a stale regime pollutes the estimate, too short and the signal is noise.

When it helps, and when it misleads

Its strength is that it is the only sibling that closes the loop on the design itself — without it, thresholds and confidence levels are set once and drift forever, and no one learns whether the process's "80% sure" means anything. It turns outcomes into evidence that updates the machinery.

Its characteristic failure is hindsight bias:[2] once the outcome is known, it looks inevitable, and reviewers grade the decision as obvious-in-retrospect rather than grading the process against the information available at the time. That collapses calibration into blame and teaches the process to be timid rather than accurate. A related misuse is over-fitting to a small or unrepresentative batch — "recalibrating" on noise and injecting a swing the next batch will have to undo. The guarding discipline is to score against what was knowable when the decision was made, carry sample-size honesty into every correction, and move thresholds only when the pattern survives the noise.

How it implements the components

  • outcome_feedback_and_calibration_loop — this is the mechanism's spine: it is the loop that compares outcomes to expectations and updates the design.
  • uncertainty_and_error_budget — by scoring stated confidence against realized frequency, it measures whether the process's error budget is honest and resets it where it is not.
  • escalation_threshold — its corrections land, in part, on the confidence and stakes thresholds that decide when a case escalates, tightening or loosening them from evidence.

It does NOT act within a single live decision, and it does not hold the method portfolio or reversibility profile a live decision runs on — deciding provisionally now and verifying that same case later is Two-Stage Review, which owns method_portfolio and stakes_and_reversibility_profile. This review touches only the closed batch.

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: Post-Decision Calibration Review operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it compares forecast, confidence, process, and outcome to recalibrate future thresholds and methods.

Independent corroboration: The frozen evidence defines Post-Decision Calibration Review as 'Compares forecast, confidence, process, and outcome to recalibrate future thresholds and methods', so its operative form is Assessment, Review & Assurance.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Psychology

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Comparing confidence and forecast with outcomes derives from judgment and decision-making psychology.

Related originating lineages:

Review resolution: Both blind reviewers agree that psychology is the primary origin. Reconciliation resolves alternate origin disagreement. Formative alternate lineages are retained as organizational_management, statistics_experimental_design, behavioral_economics; later breadth of use is recorded separately as domain_reach=multi_domain, while origin_mode=cross_disciplinary_synthesis describes the relationship among origin lineages.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Reconciled after independent review; medium confidence.

References

[1] Brier, Glenn W. "Verification of Forecasts Expressed in Terms of Probability". Monthly Weather Review 78(1): 1–3, 1950. Introduces a quadratic score comparing probability forecasts with realized categorical outcomes; it does not introduce a calibration curve or isolate the 70%-versus-55% calibration gap. registry

[2] Fischhoff, Baruch. "Hindsight ≠ Foresight: The Effect of Outcome Knowledge on Judgment Under Uncertainty". Journal of Experimental Psychology: Human Perception and Performance 1(3): 288–299, 1975. Shows that outcome knowledge inflates perceived prior predictability and what people believe could have been known beforehand. registry