Skip to content

Consistency Audit

Audit — instantiates Symmetry-Based Fairness

Measures decisions already made across reviewers, units, and time to surface unexplained variation — treatment that shifted when only the decider or the date changed, not the case.

Version
v1 · 2026-08-24 · History
Mechanism #
1806
Type
Audit
Form family
Assessment, Review & Assurance
Solution family
Comparison & Evaluation
Problem family
Exclusion, Inequality & Distributional Harm
Problem subfamily
Distributive Allocation & Equal-Treatment Harm
Origin domain
Law & Governance
Also from
Statistics & Experimental Design
Instantiates
Symmetry-Based Fairness

Consistency Audit looks backward at decisions a system has already produced and asks whether treatment tracks the case or the circumstance. It scopes a population of comparable decisions, groups them by a factor that should be irrelevant — which reviewer, which office, which month — and measures whether outcomes vary along that factor once case facts are held constant. Its defining move is empirical and retrospective: it does not read the rule to see whether the rule is fair (that is a rule test), it reads the outcomes to see whether the practice is fair. A policy can be flawless on paper and still produce wildly different treatment depending on who happened to handle the case, and only an audit of realized decisions can catch that.

Example

An immigration court assigns asylum claims to judges essentially at random, so two claimants fleeing the same country under similar circumstances may draw different judges. A Consistency Audit scopes the case set — asylum claims from one country of origin, heard in one court over one two-year window — pulls the full record of dispositions, and computes each judge's grant rate. The spread is the finding: some judges grant a large share of these claims and others grant almost none, on caseloads that were randomized to be comparable. This is the phenomenon the Refugee Roulette study made famous.[1] The audit does not stop at the spread; it runs a relevant-difference test, checking whether the variation is explained by legitimate case-mix differences (representation, corroborating evidence) or survives them. What survives is unexplained variation — treatment determined by the luck of the assignment rather than the merits of the claim.

How it works

  • Scope the case set. Fix the population of decisions that are supposed to be comparable, and publish the scope so no one can later swap in a flattering comparison group.
  • Assemble the record. Pull the preserved facts, decisions, and outcomes for those cases — the audit is only as good as the trail it reads.
  • Group by the suspect factor. Compute treatment (grant rate, sanction severity, approval time) within each reviewer, unit, or period.
  • Control and test. Adjust for stated case criteria and check whether variation persists; the residual, unexplained variation is the result — not the raw spread.

Tuning parameters

  • Grouping factor — whether you slice by reviewer, unit, geography, or time. Each exposes a different kind of drift; pick the axis you suspect and be aware disparities hide along the axes you didn't slice.
  • Scope breadth — one team versus an entire agency. Wider scope improves statistical power and catches systemic drift but blurs legitimate local context.
  • Control set — which case facts you adjust for before calling variation "unexplained." Too few controls indict legitimate differences; too many launder real bias into an approved covariate.
  • Significance threshold — how large and stable a gap must be to count. Loose thresholds chase noise; strict ones excuse real but modest disparities.
  • Cadence — continuous monitoring for high-volume automated flows, periodic review for slower human ones.

When it helps, and when it misleads

Its strength is making drift countable: it converts "reviewers seem inconsistent" into a measured spread that a governance body can act on, and it is the only mechanism here that can catch unfairness produced entirely in application rather than in the rule.

Its sharpest failure mode is that low variation is not the same as fairness. A system in which every reviewer is equally biased shows perfect consistency and total injustice — the audit sees agreement, not correctness. Confounding cuts both ways: an unmeasured relevant difference can masquerade as bias, and a measured "control" can quietly absorb the very bias you were hunting. The classic misuse is declaring a process fair because reviewers agree with each other, mistaking consensus for accuracy. The guarding discipline is to pair the audit with an independent quality or ground-truth check, and to pre-specify the case scope and controls before looking, so the analysis can't be steered toward a comfortable answer. The underlying pathology — unwanted variability in judgments that should be identical — is what the literature on decision noise is about.[2]

How it implements the components

Consistency Audit realizes the realized-outcome face of the archetype — the components that live in the record of decisions rather than in the rule:

  • case_set_scope — its first act is fixing and publishing which decisions are being compared, which is what guards the audit against reference-class manipulation.
  • audit_trail — it reads and tests the preserved trail of case facts, judgments, and outcomes; the audit is that trail put to use.
  • relevant_difference_test — it statistically checks whether observed variation tracks a legitimate difference or is genuinely unexplained.

It does not inspect the rule itself (equivalence_rule, equal_treatment_rule, value_priority_statement) — that is [Policy Symmetry Test], which convicts a policy a priori where the audit weighs outcomes a posteriori; nor does it record the reason for an individual departure (exception_rationale) — that is [Equal-Treatment Checklist].

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: Measures decisions already made across reviewers, units, and time to surface unexplained variation — treatment that shifted when only the decider or the date changed, not the case, making its operative form a bounded evaluation of existing evidence or work that produces a finding or disposition.

Independent corroboration: The frozen evidence defines Consistency Audit as 'Measures decisions already made across reviewers, units, and time to surface unexplained variation — treatment that shifted when only the decider or the date changed, not the case', so its operative form is Assessment, Review & Assurance.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Law & Governance

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Administrative-law quality review cohered retrospective tests of whether comparable cases receive materially different outcomes across adjudicators, offices, or time; statistical controls and variance analysis supply the relevant-difference test.

Related originating lineages:

Review resolution: Georgetown's Refugee Roulette study uses large adjudication datasets and regression analyses to expose outcome disparities associated with which decision-maker received a case. GAO independently recommends systematic study of adjudicative decisions and recognizes the need to distinguish expected fact-sensitive variation from systemic inconsistency. These directly instantiate the source mechanism's retrospective comparison of comparable legal decisions, making law and governance primary and statistical design formative; the cross-domain audit framing is an encyclopedia synthesis.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Researched adjudication after independent review; high confidence.

Sources consulted:

References

[1] Refugee Roulette (Ramji-Nogales, Schoenholtz & Schrag, 2007) documented that U.S. asylum grant rates varied dramatically depending on which adjudicator heard a claim — even within the same court, on caseloads assigned essentially at random. It is a canonical demonstration of unexplained variation across decision-makers. registry

[2] Noise — the unwanted variability in judgments that ought to be identical — is the subject of Kahneman, Sibony & Sunstein's Noise: A Flaw in Human Judgment (2021). A consistency audit is, in effect, a noise audit aimed at fairness rather than accuracy. registry