Case Selection Bias Audit¶
Audit — instantiates Structured Comparative Case Design
Interrogates how the cases were chosen — above all whether they were picked because they already show the outcome — and demands the negative cases the choice left out.
The most common way a comparison lies is not in its measurement but in its guest list. The Case Selection Bias Audit is performed after cases are chosen and interrogates the selection procedure for bias: were cases picked because they already display the outcome (selection on the dependent variable)? Are there negative cases — where the supposed cause was present but the outcome did not follow, or absent yet the outcome appeared anyway? Does the sample lean on one corner of the case universe because it was easy to reach? Its distinguishing move is to demand variation on the outcome, so that a design assembled only from winners is forced to admit the losers that could disconfirm it.
Example¶
A strategy team drafts "the six habits of breakout products" by studying six blockbuster launches and extracting what they shared. The audit asks the uncomfortable question: every case was chosen because it succeeded, so any habit common to all six might be equally common among the flops nobody studied. It names the missing cell — comparable products that had the same habits and failed — and requires those negative cases to enter before any habit may be called a cause. It also checks whether "breakout" was defined after the cases were in hand, a target that moved to fit them.
The output is a short memo: each selection-bias threat named, the specific negative cases that must be added to test it, and a per-claim verdict on which of the original six habits cannot survive without them. The study is not condemned wholesale; its conclusions are graded by their exposure to the flaw.
How it works¶
- Reconstruct the actual selection rule — often implicit behind a "convenience" or "illustrative" sample — and ask what it correlates with.
- Test for selection on the outcome: is there real variance on the dependent variable, or were only high-outcome cases admitted?
- Demand the missing cells — the negative and disconfirming cases the rule excluded, without which cause and coincidence look identical.
- Grade exposure per claim rather than passing or failing the study as a whole.
Tuning parameters¶
- Bias catalogue scope — which biases you screen for (selection-on-outcome, survivorship, availability, researcher access); broader is safer but slower.
- Negative-case demand — how many disconfirming cases the audit requires before a claim is allowed to stand.
- Selection-rule reconstruction depth — how hard you dig for the unstated rule behind a sample that calls itself convenient.
- Verdict granularity — a whole-study flag versus a per-claim exposure rating.
When it helps, and when it misleads¶
Its strength is that it catches the most frequent and most fatal flaw in case work — cherry-picking, and especially selecting on the outcome and then generalising — before the conclusions harden into received wisdom. Its failure mode is over-correction: an audit can dismiss a legitimately purposive sample as "biased" when the study never claimed representativeness, and it can only flag biases in its own catalogue. The classic misuse is wielding it selectively — auditing findings one dislikes to death while waving through the same flaw in congenial work. The discipline that keeps it honest is to judge the selection against the study's stated inferential goal: the sin is selecting on the outcome and then over-claiming, and the remedy is negative cases.[1]
How it implements the components¶
case_selection_rationale— the audit's object: it exposes and stress-tests the (often unstated) reason each case was admitted.negative_and_deviant_case_rule— its central remedy: it enforces the rule that cases lacking the outcome must be included, so the design can be disconfirmed rather than merely illustrated.
It does not define the eligible case universe the sample is judged against (that's Case Universe Sampling Frame), nor test how much the finding wobbles when cases are swapped in and out (that's sensitivity-to-case-set analysis); it audits the selection that already happened.
Related¶
- Instantiates: Structured Comparative Case Design — guards the selection stage against bias.
- Consumes: Case Universe Sampling Frame — the defined universe is the benchmark the realised selection is judged against.
- Sibling mechanisms: Case Universe Sampling Frame · Sensitivity to Case Set Analysis · Deviant Case Follow-Up Protocol · Matched Case Pairing Protocol · Cross-Case Evidence Matrix Tool
References¶
[1] Selection on the dependent variable — choosing cases by their outcome (e.g. only successes) and then explaining that outcome, which cannot separate a genuine cause from a trait the unstudied non-cases also share. The corrective is to include cases that vary on the outcome — a staple caution of Barbara Geddes and of King, Keohane & Verba's Designing Social Inquiry. ↩