Selection Bias Correction¶
Diagnose how entry, participation, survival, visibility, or analytic inclusion made observed cases differ from a target population, then repair the evidence or bound the claim.
Essence¶
Selection Bias Correction repairs evidence after nonrandom gates have already determined who or what becomes observable. The pattern begins from a crucial asymmetry: the analyst can see selected cases, but the decision concerns a broader target. Participation, survival, reporting, access, recording, publication, and analytic inclusion can each make the visible cases systematically different even when measurement inside the visible set is flawless.
The archetype therefore treats selection as an intervention pathway to reconstruct, not a generic warning label. It declares the target, maps each gate, characterizes observed and excluded cases, estimates or bounds inclusion, identifies which claims are distorted, chooses correction or recollection, protects against non-overlap and unstable weights, tests unobserved selection, validates against independent evidence, and preserves residual limits. Its product is not a magically unbiased number; it is target-relevant evidence whose empirical support, assumptions, and remaining blind spots are governable.
Canonical formula: target_population + observed_sample + selection_pathway + inclusion_model + correction_or_recollection + overlap_guard + sensitivity_bounds + validation -> target-relevant_evidence_with_honest_residual_uncertainty
When to Use This Archetype¶
Use this archetype when the evidence already exists and the inclusion process can no longer be redesigned from scratch. Typical signals include voluntary response related to satisfaction, attrition related to outcomes, learning only from survivors, access barriers that shape who appears, reporting systems that miss vulnerable cases, or analyses conditioned on a gate influenced by multiple causes. The candidate is especially valuable when a credible frame, auxiliary data, follow-up study, administrative linkage, or sensitivity range can illuminate unseen cases.
Do not use it as a synonym for every biased analysis. If sampling is still prospective, design a representative sample. If values are missing within an otherwise appropriate sample, select a missingness-aware estimator. If combining groups produces reversal, correct aggregation. If a common cause links treatment and outcome, control confounding. Where no target, overlap, auxiliary evidence, or credible assumption can be established, the correct intervention may be to narrow the claim, collect new evidence, or decline point estimation.
Structural Problem¶
Selection bias is generated by a pathway, not merely by unequal counts. A target unit must be eligible, reachable, invited, willing, persistent, visible, recorded, and retained in analysis. Variables that affect passage may also affect the outcome, and conditioning on passage can create associations that do not exist in the target. Survivors can make strategies look safer; respondents can make services look more satisfactory; published successes can make interventions look more reliable; and convenient records can redirect resources away from invisible need.
The observed sample alone often cannot reveal the distortion because the missing cases are absent by definition. Correction therefore depends on frame data, auxiliary variables, follow-up evidence, causal assumptions, and explicit bounds. A second structural danger is positivity failure: some target regions have no plausible observed analog. Mathematical weighting cannot create information there. A mature response distinguishes imbalance that can be corrected, uncertainty that can be bounded, and absence that requires recollection or narrower claims.
Intervention Logic¶
Begin by fixing the target population, estimand, time window, and decision. Profile the observed sample without assuming it defines the target. Build an operational and causal map of eligibility, access, contact, participation, survival, reporting, recording, and analytic inclusion. For every gate, ask who controls it, what predicts passage, how it relates to outcomes, and what evidence exists about excluded cases. Diagnose which margins, associations, effects, or priorities are materially distorted.
Choose the least assumption-dependent viable response. Known design probabilities may support weighting; target margins may support raking; a frame may support propensity modeling; nonresponse may require follow-up; survivorship may require case reconstruction; poor overlap may require new collection or a narrower target. Publish weight stability and effective sample size, test unobserved selection, compare alternative methods, validate against evidence not used to fit the correction, and version the pathway as access and behavior change. The action sequence closes only when claim scope and residual uncertainty travel with the corrected result.
Key Components¶
| Component | Description |
|---|---|
| Target Population and Claim ↗ | Defines the population, process, or case universe to which the corrected inference or decision is intended to apply. Operationally, this component must be represented in a form that decision makers can inspect and update. State the estimand or operational claim, time window, units, and decision use before choosing a correction. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by anchoring every diagnosis and correction to a declared target rather than the visible cases. It is required because correction can optimize the observed sample while remaining irrelevant to the intended population Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Observed Sample Profile ↗ | Describes who or what is observed, with relevant attributes, outcomes, exposure histories, and observation channels. Operationally, this component must be represented in a form that decision makers can inspect and update. Preserve provenance, timing, duplicate handling, and the difference between absence and non-observation. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by providing the empirical object whose selection history must be explained. It is required because the analysis treats a convenience set as though it were self-defining Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Selection Pathway Map ↗ | Traces eligibility, access, invitation, participation, survival, reporting, recording, and analytic inclusion gates. Operationally, this component must be represented in a form that decision makers can inspect and update. Represent gates in causal and operational order, including who controls them and which variables influence passage. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by making the distortion-producing pathway inspectable. It is required because teams correct demographics while missing the actual gate that selected outcomes or survivors Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Inclusion Probability Model ↗ | Estimates or bounds each target unit's probability of appearing in the observed analytic set. Operationally, this component must be represented in a form that decision makers can inspect and update. Use design probabilities when known and model-based probabilities only with explicit assumptions, diagnostics, and positivity checks. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by connecting selection mechanisms to correction weights or resampling priorities. It is required because correction strength becomes arbitrary and extreme weights silently dominate Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Excluded and Unseen Case Profile ↗ | Characterizes omitted, unreachable, nonresponding, attrited, censored, or otherwise invisible cases using available frame and auxiliary evidence. Operationally, this component must be represented in a form that decision makers can inspect and update. Distinguish known exclusions from structurally unobservable cases and do not invent attributes unsupported by evidence. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by showing what the observed sample cannot reveal by itself. It is required because the correction assumes missing cases resemble observed cases without testing that assumption Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Selection Distortion Diagnosis ↗ | Determines which distributions, associations, effects, risks, or priorities are changed by the selection pathway. Operationally, this component must be represented in a form that decision makers can inspect and update. Compare observed and target margins, simulate gate effects, and distinguish harmless imbalance from decision-relevant distortion. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by preventing blanket correction when selection is present but immaterial to the claim. It is required because adjustment adds variance and model dependence without reducing relevant bias Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Identification Assumption Register ↗ | Records the conditions under which the selected data can identify or bound the target quantity. Operationally, this component must be represented in a form that decision makers can inspect and update. Name exchangeability, missing-at-random, transportability, positivity, exclusion, and measurement assumptions in plain language. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by separating empirical evidence from assumptions required by the correction. It is required because a mathematically complete estimate is presented as assumption-free Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Correction Strategy and Estimand ↗ | Selects the adjustment, augmentation, resampling, reconstruction, or bounded-claim strategy appropriate to the pathway and target. Operationally, this component must be represented in a form that decision makers can inspect and update. Choose the cheapest defensible correction and explain whether it changes weights, cases, model, target, or claim scope. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by turning diagnosis into a governed remedial intervention. It is required because several incompatible corrections are mixed without a coherent estimand Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Weight Stability and Overlap Guard ↗ | Controls extreme weights, sparse strata, non-overlap, and extrapolation beyond supported target regions. Operationally, this component must be represented in a form that decision makers can inspect and update. Publish effective sample size, weight distribution, trimming rules, and the target population lost when overlap fails. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by keeping mathematical correction from manufacturing unsupported information. It is required because a few cases carry most of the result or excluded regions are silently extrapolated Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Sensitivity and Partial-Identification Layer ↗ | Tests how conclusions change under plausible unobserved selection and reports bounds when point correction is not identified. Operationally, this component must be represented in a form that decision makers can inspect and update. Vary gate effects, outcome differences, and dependence structures using domain-credible ranges rather than cosmetic perturbations. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by making irreducible selection uncertainty visible. It is required because one favored selection model is treated as the only possible data-generating process Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Post-Correction Validation ↗ | Evaluates whether corrected data reproduce held-out target margins, outcomes, or external benchmarks without introducing new distortions. Operationally, this component must be represented in a form that decision makers can inspect and update. Use negative controls, calibration targets not used in fitting, subgroup checks, and alternative correction methods. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by testing the correction as an intervention rather than accepting it because it ran. It is required because weighted balance is mistaken for truth despite poor outcome or external validity Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
| Claim Limitation and Monitoring ↗ | Constrains communication to supported populations and establishes monitoring for changing selection pathways. Operationally, this component must be represented in a form that decision makers can inspect and update. Version the selection map, document residual blind spots, and trigger redesign or recollection when gates or participation change. Its owner, evidence source, confidence, and update cadence should be explicit; otherwise the component becomes a decorative label rather than load-bearing structure. The component connects to the archetype by keeping corrected conclusions aligned with evolving access and inclusion conditions. It is required because a one-time correction becomes permanent authority after the pathway drifts Reviewers should test it with ordinary, boundary, missing-data, and adversarial cases instead of assuming that a plausible description will behave correctly in use. |
Common Mechanisms¶
| Mechanism | Description |
|---|---|
| Inverse-Probability Weighting ↗ | Weights observed cases by the inverse of an estimated or known inclusion probability to reconstruct a target distribution. Use it when selection probabilities are estimable and overlap is adequate Implementation should expose inputs, thresholds or transformation rules, expected outputs, responsible actors, and evidence of performance. Inspect positivity, stabilize or trim extreme weights, and report effective sample size. This mechanism is not the archetype by itself. It instantiates the broader intervention only when connected to the complete component set, monitoring, exception handling, and revision logic. Weighting is one correction method; it does not supply the pathway map, assumptions, validation, and claim governance. |
| Post-Stratification and Raking ↗ | Adjusts observed margins to known target totals across selected auxiliary dimensions. Use it when reliable population margins exist and selection is explainable through those dimensions Implementation should expose inputs, thresholds or transformation rules, expected outputs, responsible actors, and evidence of performance. Do not claim correction for unmeasured outcome-related selection, and test sparse cells. This mechanism is not the archetype by itself. It instantiates the broader intervention only when connected to the complete component set, monitoring, exception handling, and revision logic. Margin calibration implements part of the strategy but is not the complete archetype. |
| Selection-Propensity Model ↗ | Models the probability of observation or participation from frame, administrative, or contact variables. Use it when a frame includes observed and unobserved units with common predictors Implementation should expose inputs, thresholds or transformation rules, expected outputs, responsible actors, and evidence of performance. Separate prediction accuracy from causal adequacy and avoid using post-selection colliders carelessly. This mechanism is not the archetype by itself. It instantiates the broader intervention only when connected to the complete component set, monitoring, exception handling, and revision logic. A propensity model becomes useful only inside a declared target and identification argument. |
| Sample-Selection Outcome Model ↗ | Jointly models selection and outcome processes when their unobserved determinants may be correlated. Use it when outcomes are observed only after a nonrandom gate and credible identifying structure exists Implementation should expose inputs, thresholds or transformation rules, expected outputs, responsible actors, and evidence of performance. State exclusion restrictions, distributional assumptions, and sensitivity to misspecification. This mechanism is not the archetype by itself. It instantiates the broader intervention only when connected to the complete component set, monitoring, exception handling, and revision logic. The model cannot replace evidence for why selection occurs or justify weak instruments. |
| Targeted Oversampling and Recontact ↗ | Collects additional cases from underrepresented, high-uncertainty, or high-leverage target strata. Use it when correction from existing data would depend on extreme extrapolation Implementation should expose inputs, thresholds or transformation rules, expected outputs, responsible actors, and evidence of performance. Preserve consent, access, burden, and comparability; do not coerce hard-to-reach groups. This mechanism is not the archetype by itself. It instantiates the broader intervention only when connected to the complete component set, monitoring, exception handling, and revision logic. New collection is one remediation route and may be preferable to elaborate modeling. |
| Nonresponse Follow-Up Study ↗ | Samples nonrespondents or late responders to estimate how participation relates to relevant outcomes. Use it when survey, service, or audit nonresponse may be informative Implementation should expose inputs, thresholds or transformation rules, expected outputs, responsible actors, and evidence of performance. Use multiple contact modes and record differential reachability without treating late response as perfect truth. This mechanism is not the archetype by itself. It instantiates the broader intervention only when connected to the complete component set, monitoring, exception handling, and revision logic. The study supplies missing pathway evidence but still requires integration and limits. |
| Survivorship Reconstruction ↗ | Reintroduces failed, exited, censored, dissolved, or otherwise missing historical cases into analysis. Use it when visible current cases are selected by persistence or successful completion Implementation should expose inputs, thresholds or transformation rules, expected outputs, responsible actors, and evidence of performance. Define failure and exit consistently, recover denominators, and distinguish censoring from true absence. This mechanism is not the archetype by itself. It instantiates the broader intervention only when connected to the complete component set, monitoring, exception handling, and revision logic. Reconstruction addresses the survival gate but not every selection pathway. |
| Causal Selection Diagram ↗ | Represents variables that affect inclusion and outcomes to reveal conditioning, collider, and transportability risks. Use it when several plausible selection mechanisms compete or adjustment could open biasing paths Implementation should expose inputs, thresholds or transformation rules, expected outputs, responsible actors, and evidence of performance. Treat arrows as assumptions to challenge with experts and data, not as decorative certainty. This mechanism is not the archetype by itself. It instantiates the broader intervention only when connected to the complete component set, monitoring, exception handling, and revision logic. The diagram diagnoses structure; correction also needs an estimand and operational method. |
| Transportability Reweighting ↗ | Reweights source data toward a destination population using covariate distributions and effect-modification assumptions. Use it when evidence comes from one setting but decisions concern another Implementation should expose inputs, thresholds or transformation rules, expected outputs, responsible actors, and evidence of performance. Check support, measurement equivalence, effect modifiers, and destination-specific constraints. This mechanism is not the archetype by itself. It instantiates the broader intervention only when connected to the complete component set, monitoring, exception handling, and revision logic. Transport is a selection correction only when the source-to-target pathway is explicit. |
| Unobserved-Selection Sensitivity Analysis ↗ | Varies the prevalence and outcome influence of unmeasured selection drivers to bound conclusions. Use it when point identification depends on untestable selection assumptions Implementation should expose inputs, thresholds or transformation rules, expected outputs, responsible actors, and evidence of performance. Use interpretable parameters and report threshold values that would reverse the decision. This mechanism is not the archetype by itself. It instantiates the broader intervention only when connected to the complete component set, monitoring, exception handling, and revision logic. Sensitivity analysis does not repair data; it governs uncertainty and claim scope. |
Parameter / Tuning Dimensions¶
Target Breadth¶
Sets how far the intended population extends beyond observed support. Low settings create an unnecessarily narrow but defensible claim; high settings create unsupported extrapolation into target regions with no overlap. Tune against frame coverage, decision scope, and overlap diagnostics and record the rationale.
Correction Strength¶
Controls the influence of weights, modeled outcomes, augmentation, or reconstructed cases. Low settings create residual selection bias; high settings create variance inflation and domination by modeled assumptions. Tune against effective sample size, external calibration, and bias-variance analysis and record the rationale.
Weight Trimming¶
Sets caps or stabilization rules for low inclusion probabilities. Low settings create unstable estimates driven by a few cases; high settings create reintroducing bias and silently narrowing the target. Tune against weight distribution, overlap, loss function, and sensitivity curves and record the rationale.
Unobserved-Selection Sensitivity¶
Sets the range of plausible hidden gate effects considered. Low settings create false confidence from cosmetic perturbations; high settings create bounds so broad that the analysis cannot guide action. Tune against follow-up data, domain knowledge, negative controls, and comparable studies and record the rationale.
Recollection Intensity¶
Determines investment in follow-up, oversampling, frame repair, or new collection. Low settings create continued dependence on extrapolation; high settings create cost, delay, privacy burden, and possible coercion. Tune against value of information and which additional cases most reduce decision uncertainty and record the rationale.
Invariants to Preserve¶
Declared Target¶
Every correction is tied to an explicit target population and estimand. Preserve it by versioning the target and checking that data, weights, and validation correspond to it A violation is indicated by the same corrected estimate is reused for incompatible populations
Pathway Before Method¶
Correction follows an explicit selection-pathway diagnosis rather than an estimator chosen by habit. Preserve it by reviewing each gate, its causes, and its relation to outcomes before modeling A violation is indicated by teams apply demographic weights without explaining participation or survival
Overlap Honesty¶
No correction claims information where target cases have no supported analog in observed data. Preserve it by publishing support diagnostics and narrowing or recollecting when positivity fails A violation is indicated by extreme weights or model-only predictions dominate unsupported strata
Assumption Visibility¶
Empirical facts, identifying assumptions, and sensitivity choices remain distinguishable. Preserve it by maintaining an assumption register and alternative analyses A violation is indicated by a model output is communicated as direct observation
Residual-Uncertainty Disclosure¶
Correction never erases uncertainty from unobserved selection or weak validation. Preserve it by reporting bounds, effective sample size, and unsupported regions with the result A violation is indicated by precision increases after correction without supporting information
Independent Validation¶
At least some target evidence not used to fit the correction tests its consequences. Preserve it by holding out benchmarks, follow-up samples, or future outcomes A violation is indicated by success is defined only as balance on fitting variables
Target Outcomes¶
The immediate outcome is evidence whose relationship to a declared target is explicit and testable. Decision makers can see which gates selected the observed cases, which correction assumptions are empirical or hypothetical, where overlap supports recovery, and where only bounds or recollection are defensible. Corrected estimates should improve target-margin calibration, external prediction, subgroup coverage, and decision stability without hiding variance inflation.
Secondary outcomes include recovery of failed and invisible cases, better prioritization of outreach and collection, reduced survivorship storytelling, fairer representation of hard-to-reach groups, and clearer separation between data limitations and substantive absence. The system should become capable of detecting pathway drift, such as changing response incentives or platform churn, and revising the correction rather than treating yesterday's weights as permanent truth.
Tradeoffs¶
Bias versus Variance¶
Strong reweighting can reduce systematic distortion while sharply reducing effective information. The design should not pretend this tension disappears. Use stabilized estimators, recollection, multiple corrections, and decision-focused loss analysis and make the selected position visible to affected actors.
Breadth versus Support¶
A broad target improves relevance but may include regions absent from observed data. The design should not pretend this tension disappears. Use overlap diagnostics, target refinement, partial identification, and explicit unsupported strata and make the selected position visible to affected actors.
Modeling versus Recollection¶
Model correction is fast but assumption-dependent; new data are costly but can restore support. The design should not pretend this tension disappears. Use value-of-information analysis targeted at the gates and strata driving uncertainty and make the selected position visible to affected actors.
Visibility versus Privacy¶
Diagnosing selection often requires attributes and pathway records that create privacy and burden risks. The design should not pretend this tension disappears. Use data minimization, controlled linkage, consent, aggregation, and independent governance and make the selected position visible to affected actors.
Failure Modes¶
Wrong Target¶
The correction represents an administratively convenient frame rather than the population implicated by the decision. Detect it through stakeholder and estimand review reveals excluded target groups Respond by redefine the target and repeat pathway, overlap, and validation work
Incomplete Pathway Map¶
Analysts model response but omit earlier access, eligibility, survival, reporting, or analytic gates. Detect it through new data sources reveal selection before or after the modeled gate Respond by extend the causal and operational map and reassess identification
Extreme-Weight Domination¶
A small number of low-probability cases determine the corrected result. Detect it through low effective sample size, high leverage, and unstable leave-one-out estimates Respond by recollect, trim with target disclosure, use augmentation, or narrow the claim
Unmeasured Selection¶
Participation or survival depends on outcome-related variables absent from the correction data. Detect it through follow-up studies, negative controls, or sensitivity thresholds contradict the fitted model Respond by collect pathway evidence, bound results, or stop point identification
Collider Adjustment Error¶
Conditioning or controlling within the selected set opens a spurious association. Detect it through causal selection diagrams and alternative conditioning sets produce reversals Respond by revise the graph, estimand, and correction strategy
Correction Without Validation¶
Weighted balance is accepted although corrected outcomes fail external checks. Detect it through held-out margins, follow-up cases, or prospective outcomes remain miscalibrated Respond by reject or revise the model and communicate residual bias
Neighbor Distinctions¶
Representative Sampling Design¶
Prospectively constructs a sample whose inclusion process is designed before observation.
Boundary rule: Use Representative Sampling Design when the frame and selection plan can still be designed; use Selection Bias Correction when a nonrandom selection pathway has already produced the observed data and must be diagnosed, repaired, bounded, or supplemented.
Hybrid cases should assign each design obligation to its proper owner rather than using Selection Bias Correction as an umbrella for all adjacent work.
Aggregation Bias Detection and Correction¶
Detects distortion caused by combining heterogeneous groups or levels.
Boundary rule: Use it when aggregation changes an estimate despite adequate inclusion; use Selection Bias Correction when entry, participation, survival, visibility, or analytic inclusion makes observed cases systematically unlike the target.
Hybrid cases should assign each design obligation to its proper owner rather than using Selection Bias Correction as an umbrella for all adjacent work.
Missingness-Aware Estimator Selection¶
Chooses estimation procedures appropriate to missing-data patterns.
Boundary rule: Use it when incomplete variables within a defined sample are central; use Selection Bias Correction when entire cases or pathways into observation are nonrandom and target-population recovery is the design obligation.
Hybrid cases should assign each design obligation to its proper owner rather than using Selection Bias Correction as an umbrella for all adjacent work.
Confounder Control¶
Controls common causes of exposure and outcome in causal estimation.
Boundary rule: Use it for treatment-outcome backdoor paths; use Selection Bias Correction for conditioning on observation, participation, survival, or other selection gates, including collider bias created by analytic inclusion.
Hybrid cases should assign each design obligation to its proper owner rather than using Selection Bias Correction as an umbrella for all adjacent work.
Generalization Validation¶
Tests whether a finding or model transfers beyond its development setting.
Boundary rule: Use it to evaluate external performance broadly; use Selection Bias Correction when a specific selection pathway must be modeled and actively corrected before the claim can represent the target.
Hybrid cases should assign each design obligation to its proper owner rather than using Selection Bias Correction as an umbrella for all adjacent work.
Bias-Specific Decision Audit¶
Audits a decision process for a named cognitive, procedural, or distributional bias.
Boundary rule: Use it when the decision rule or actor behavior is the intervention target; use Selection Bias Correction when inclusion into the evidence base is the distortion-producing mechanism.
Hybrid cases should assign each design obligation to its proper owner rather than using Selection Bias Correction as an umbrella for all adjacent work.
Cross-Domain Examples¶
Clinical Evidence¶
Investigators combine trial and registry frames to model which patients entered a treatment study, reweight toward eligible patients, test overlap, and bound effects for poorly represented severity groups.
Why it fits: treatment evidence already comes from a selected participant pathway and the target is broader than enrolled patients The design is evaluated by weight stability, target-margin recovery, external outcome calibration, and sensitivity to unmeasured participation
Workplace Safety¶
A firm reconstructs injuries among contractors and departed workers omitted from employee surveys, follows a sample of nonrespondents, and revises risk estimates and controls.
Why it fits: employment status and reporting access select which harms become visible The design is evaluated by recovered denominator, reporting-propensity differences, corrected risk ranking, and policy effect
Product Analytics¶
A platform corrects satisfaction findings drawn only from active users by modeling churn and survey response, adding exit interviews, and limiting claims for unreachable accounts.
Why it fits: survival and voluntary response jointly select the observed experience The design is evaluated by held-out churn prediction, recontact evidence, effective sample size, and decision stability
Education¶
A program evaluates outcomes after accounting for school opt-in, student attrition, and test participation, oversampling missing contexts and reporting bounds where overlap fails.
Why it fits: multiple gates select both institutions and measured students The design is evaluated by coverage of target schools, attrition reconstruction, sensitivity bounds, and subgroup calibration
Public Consultation¶
A city maps access, invitation, language, attendance, and comment-submission gates, supplements the record with targeted outreach, and weights only where frame evidence supports it.
Why it fits: visible comments systematically overrepresent actors able and motivated to participate The design is evaluated by target-margin coverage, participation burden, policy-priority changes, and residual blind spots
Organizational Learning¶
A company reconstructs project lessons including canceled and failed initiatives rather than learning only from surviving launches, then adjusts its portfolio guidance.
Why it fits: survivorship selects the cases available for retrospective inference The design is evaluated by failed-case recovery, changed success-factor estimates, and prospective prediction
Non-Examples¶
Prospective Stratified Sample¶
Researchers still control the frame and draw strata before data collection. It fails the archetype boundary because the task is to design inclusion rather than remediate an already selected evidence set Route the case to Representative Sampling Design when that is the actual design object.
Item-Level Imputation¶
A valid sample has sporadic missing answers and analysts impute them under a documented missingness model. It fails the archetype boundary because no case-level inclusion pathway or target-population distortion is central Route the case to Missingness-Aware Estimator Selection when that is the actual design object.
Confounding Adjustment¶
Treatment and outcome share a common cause, but observation is not conditioned on a selection gate. It fails the archetype boundary because the causal problem is confounding rather than selection into the analytic set Route the case to Confounder Control when that is the actual design object.
Aggregation Reversal¶
A relationship reverses after combining groups despite adequate representation within each group. It fails the archetype boundary because group aggregation, not selection into observation, produces the distortion Route the case to Aggregation Bias Detection and Correction when that is the actual design object.
Related Abstractions¶
Abstractions this archetype builds on — directly (a source ingredient) or as a related pattern. Links follow the typed catalog namespace.
Built directly on (3)
- Causality: Cause-effect relationships.
- Sampling (Representativeness): Representative subset selection.
- Selection Bias: Skewed sampling.
Also references 9 related abstractions
- Boundary: Defines system limits.
- Classification: Sorting entities into discrete categories by explicit rules, turning unbounded variation into a finite, reusable map for downstream reasoning and action.
- Constraint: Limits possibilities to guide outcomes.
- Fairness: Judging whether an allocation or procedure treats comparable parties impartially according to a defensible standard, given that multiple such standards can conflict.
- Feedback: Outputs influence inputs.
- Inductive Reasoning: Specific to general inference.
- Missingness
- Observability: Infer internal state externally.
- Uncertainty: Incomplete knowledge.
Variants¶
Narrower or domain-specific specializations that share this archetype's core structure. Recognized variants are established; candidate variants are provisional.
Nonresponse Bias Correction
Corrects distortion when participation or response depends on characteristics related to the target outcome.
Survivorship Bias Correction
Restores failed, exited, censored, or dissolved cases omitted by observing only survivors.
Access and Visibility Bias Correction
Corrects evidence that overrepresents people or events easiest to reach, record, publish, or observe.
Collider Selection-Bias Correction
Repairs associations distorted by conditioning on a common effect that controls inclusion.