Skip to content

Shadow Review Board

Standing review body — instantiates Structural Filter Intersection Audit

A standing independent panel that re-adjudicates a running stream of filtered-out outputs under its own declared criteria, revealing what the live filter set would have passed had the judgment been someone else's.

A Shadow Review Board is a standing body, deliberately independent of the production and approval chain, that takes a running stream of the outputs the real filters rejected or transformed and re-judges them under its own explicit criteria — producing a parallel "shadow" verdict set. Its defining move is that it is neither sampling nor describing the rejects (that is Rejected-Item Sampling), nor waiting to hear appeals brought by aggrieved producers (that is Appeal and Exception Review): it re-adjudicates proactively, on its own initiative, so that the gap between the live filter's verdicts and the board's becomes a standing, measurable estimate of the filter set's structural bias. Its output is a maintained shadow corpus: the same inputs, a different judgment, tracked over time.

Example

An intelligence agency worries that its finished assessments cluster around a house consensus and that dissenting reads get filtered out before they reach policymakers. It stands up a shadow review board — an independent team outside the main analytic line — that periodically takes the raw material behind a sample of assessments, including the leads and sources the main line dropped, and produces its own read under stated criteria. Where the main product said "unlikely," the board sometimes says "plausible, under-examined." A single divergence proves nothing. But a persistent, patterned divergence — the board keeps resurrecting a class of hypotheses the live filter reliably kills — is precisely the structural signal the audit needs. The historical "Team B" competitive-analysis exercise is the real precedent;[1] here it is institutionalized as a standing shadow stream rather than a one-off.

How it works

  • Stand up a panel insulated from the production and approval chain — a separate reporting line, mixed with outside membership.
  • Feed it a running stream of filtered-out and transformed outputs together with their original inputs.
  • Have it re-adjudicate under its own explicit, declared criteria — not to overturn the live decision, but to generate a parallel verdict.
  • Maintain the divergences as a shadow corpus: a standing holdout of "what a different judgment would have surfaced."

Tuning parameters

  • Independence depth — how insulated the board is in reporting line, funding, and membership; more independence means more signal and more friction with the main line.
  • Criteria stance — whether the board judges by the same stated rules (testing whether the live filter applies its own rules consistently) or by different declared values (testing the rules themselves); these answer different questions.
  • Input completeness — whether the board receives only what the live filter saw, or also the dropped leads and sources; the latter is more revealing and more expensive.
  • Bindingness — advisory shadow verdicts versus power to force a re-review; teeth raise both the stakes and the capture risk.
  • Cadence and sampling — continuous versus periodic review, and which slice of the rejects the board sees.

When it helps, and when it misleads

Its strength is producing a live counterfactual — "what would have surfaced under a different judgment" — that no purely descriptive mechanism can, with a verdict-divergence that is concrete and trackable. It is the institutional memory of the roads not taken.

A shadow board is only as good as its independence, and it is a prime capture target: absorbed into the main culture, it converges on the same verdicts and launders the filter behind a veneer of independent blessing. It can also over-correct, treating its own contrarian reads as inherently truer and manufacturing false dissent, and a board handed responsibility without real authority becomes decorative. The discipline is to protect independence structurally — rotation, outside seats, separate reporting — require the board to declare its criteria, and watch whether its divergence from the live line is genuine signal or performance.

How it implements the components

  • shadow_corpus_or_holdout_stream — its product: the maintained parallel stream of independently re-adjudicated outputs, and the running record of where its verdicts diverge from the live filter's.

It re-judges rejects but does not gather or sample them in the first place — rejected_or_transformed_output_sample is Rejected-Item Sampling's — hear producer- or subject-initiated appeals of specific decisions (appeal_exception_public_reason_channel, Appeal and Exception Review), or install the standing independence_and_diversity_controls on the main filters themselves (Filter Rotation or External Challenge).

  • Instantiates: Structural Filter Intersection Audit — it is the independent judgment that turns "what the filters removed" into "what a different judgment would have kept."
  • Consumes: Rejected-Item Sampling supplies the stream of filtered-out outputs the board re-adjudicates.
  • Sibling mechanisms: Rejected-Item Sampling · Appeal and Exception Review · Filter Rotation or External Challenge · Filter Independence Check · Intersection Matrix · Producer Pressure Survey

Notes

The board's entire value lives in its independence, and its health check is divergence itself. If the shadow verdicts stop diverging from the live line, one of two very different things is true — the filter is genuinely fine, or the board has been captured and now shares its culture — and the mechanism is only worth running if you have kept enough independence to tell which.

References

[1] "Team B" — a group of outside analysts commissioned in 1976 to independently reassess intelligence the main analytic line had produced, as a competitive check on consensus estimates. It is the real precedent for an independent re-adjudication body; the broader practice is competitive analysis or red-teaming. Its own contested reputation is itself the caution about panels that are contrarian by design.