Skip to content

False-Capture Audit

Audit — instantiates Bycatch-Aware Selective Intervention Design

An arm's-length review that samples what the selector actually caught, sorts true target from non-target, and reports a false-capture rate the operator can't self-certify away.

A dashboard can only chart the bycatch someone chose to feed it, and the operator scoring its own catch has every reason to feed it flatteringly. False-Capture Audit breaks that loop: an independent reviewer draws a sample of what the selector actually captured, classifies each item as true target or false capture against a trusted ground truth, and reports the realized non-target rate. Its defining feature is independence plus ground truth — it measures the selector's real-world specificity after deployment, not the number the operator would like to report. Where the Bycatch Rate Dashboard displays the trend, this audit is what makes the trend trustworthy in the first place.

Example

A border agency's automated risk system flags travelers for secondary inspection. The target is genuine threats; the bycatch is ordinary travelers pulled aside on a bad proxy. The agency's own stats say the system is "highly accurate" — but accuracy quoted against flagged cases says nothing about how many flags were wrong.

The audit is run by a separate oversight unit: it pulls a random sample of secondary inspections, independently adjudicates each against the actual outcome (was there anything to find?), and reports the share that were false captures — broken out by traveler group. Setup to outcome: the headline "accuracy" survives, but the audit surfaces that among one nationality the false-capture rate is several times the baseline — a realized specificity gap invisible in the operator's own reporting.[1] The number is credible precisely because the people who own the target metric did not produce it.

How it works

  • Independence is the point. The audit is run by a party that does not own the target score, so it cannot quietly grade its own homework.
  • Ground truth, not proxies. Sampled captures are adjudicated against a trusted determination of what they truly were, which is what turns "flagged" into "right or wrong."
  • Realized, in-production specificity. It measures how the selector actually discriminates once deployed — distinct from a design-time sweep, because field conditions drift from the lab.
  • Sliced by class. The false-capture rate is reported per non-target class, so a harm concentrated in one group is not averaged into an acceptable-looking whole.

Tuning parameters

  • Sample size and frequency — how many captures are adjudicated, how often. Larger, more frequent samples tighten the estimate but cost scarce independent-review time.
  • Ground-truth standard — how authoritative the "true class" determination is. A stronger standard yields a more trustworthy rate but is slower and more expensive to obtain.
  • Independence distance — internal-but-separate team versus fully external auditor. Greater distance resists capture but knows the domain less well.
  • Stratification — random sampling versus over-sampling suspected-harm classes. Stratifying finds concentrated false captures faster but needs care to reconstruct the true overall rate.
  • Adjudication blinding — whether reviewers see the selector's original verdict. Blinding removes anchoring bias at the cost of some efficiency.

When it helps, and when it misleads

Its strength is credibility: an arm's-length, ground-truthed rate is the one number an operator cannot wave away, and it converts "we're accurate" into an auditable claim with a real denominator. It is what lets every downstream mechanism — dashboards, stop rules, compensation — rest on a figure that has been checked rather than asserted.

Its failure modes cluster around the sample and the auditor. A sample too small or drawn only from easy cases yields a confident but wrong rate; a "ground truth" that is itself the selector's output measures nothing. And the classic misuse is a captured audit — nominally independent but run by, or reporting to, the team whose metric it grades, which drifts toward confirming rather than testing. The discipline that guards against this is real independence, a genuinely external ground truth, adequate and representative sampling, and publishing the confidence around the rate rather than a lone reassuring point.

How it implements the components

  • independent_non_target_audit — the mechanism is this component: an arm's-length sampling-and-adjudication of actual captures that produces a false-capture rate the operator cannot self-certify.
  • selector_specificity_profile — it measures the realized, in-production specificity (the share of captures that were truly target); the Selectivity Window Test profiles this prospectively, and the audit confirms it after deployment.

It does not display the trend over time — that's the Bycatch Rate Dashboard; it does not set the acceptable level (that's the Bycatch Tolerance Stop Rule); and it does not retune the selector to close the gap it finds (Selector Retuning Cycle).

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: False-Capture Audit operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it an arm's-length review that samples what the selector actually caught, sorts true target from non-target, and reports a false-capture rate the operator can't self-certify away.

Independent corroboration: The frozen evidence defines False-Capture Audit as 'An arm's-length review that samples what the selector actually caught, sorts true target from non-target, and reports a false-capture rate the operator can't self-certify away', so its operative form is Assessment, Review & Assurance.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Statistics & Experimental Design

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Statistical classification and experimental evaluation formalized false-positive rates, specificity, representative samples, and class-disaggregated error measurement.

Related originating lineages:

Review resolution: NIST AI risk guidance explicitly requires false-positive and false-negative measures with realistic, representative test sets. Ecological bycatch and audit independence shape the mechanism's capture-and-review form.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Researched adjudication after independent review; high confidence.

Sources consulted:

References

[1] Grother, P., Ngan, M., & Hanaoka, K. Face Recognition Vendor Test Part 3: Demographic Effects. NISTIR 8280. National Institute of Standards and Technology (2019). Reports large country-of-birth differentials in false-positive face-recognition rates and warns that aggregate accuracy summaries can hide those operational demographic gaps. registry