Experimental Comparison & Hypothesis-Test Design¶
← Back to Uncertainty, Evidence & Inference Failure
Treatment, control, assignment, blinding, power, and evidence thresholds are insufficiently designed to support the intended comparison.
91 mechanisms across 9 solution archetypes. This is a recurring problem pattern within Uncertainty, Evidence & Inference Failure; the mechanisms below inherit it from the primary archetype they instantiate.
Because this set contains more than 30 mechanisms, it is divided by form family—the concrete kind of thing a practitioner deploys, enacts, maintains, or convenes. This is a browsing subdivision only; it does not change the inherited problem classification. Click a form below to jump to its fully visible section.
| Form family | Mechanisms | Description |
|---|---|---|
| Analysis, Modeling & Optimization | 17 | A calculation, model, estimator, diagnostic, comparison, simulation, or optimization that transforms inputs into an inference, prediction, recommendation, or formal result. |
| Assessment, Review & Assurance | 16 | A bounded evaluation of existing evidence, work, compliance, or readiness that produces a finding, approval, correction, or disposition. |
| Control, Automation & Runtime | 3 | A state-dependent executable mechanism that senses, triggers, schedules, filters, throttles, routes, or actuates during operation. |
| Decision, Gate & Allocation | 2 | A bounded selection, disposition, routing, admission, prioritization, matching, or allocation among eligible alternatives. |
| Experiment, Test & Rehearsal | 20 | An active probe, controlled variation, simulated condition, or practiced execution used to generate evidence or readiness. |
| Interface, Display & Cue | 1 | A user-facing perceptual surface or interactive affordance that presents status, options, warnings, prompts, or controls. |
| Monitoring, Sensing & Alerting | 2 | Ongoing or repeated observation of actual state that emits measurements, indicators, dashboards, surveillance signals, or alerts. |
| Organization, Role & Governance | 1 | An enduring actor, authority, body, program, service, pooled capacity, or institutional arrangement whose mandate, membership, resources, or continuity is operative. |
| Protocol, Workflow & Routine | 4 | A repeatable ordered sequence of actions, handoffs, states, or escalation steps, including procedures, runbooks, routines, recovery sequences, and lifecycle workflows. |
| Record, Log & Register | 3 | A durable, usually accumulating account of actual events, decisions, custody, exceptions, or state transitions whose value depends on history, provenance, or accountability. |
| Representation, Specification & Plan | 11 | A non-executable information artifact that externalizes understood, desired, or future structure, including maps, matrices, templates, checklists, specifications, reports, plans, and schedules. |
| Rule, Policy & Commitment | 9 | A standing constraint, permission, default, threshold, quota, obligation, right, or conditional action rule governing future behavior. |
| Structure, Architecture & Configuration | 2 | An enduring physical, digital, spatial, material, or organizational topology, partition, boundary, component arrangement, or configured state. |
Analysis, Modeling & Optimization¶
A calculation, model, estimator, diagnostic, comparison, simulation, or optimization that transforms inputs into an inference, prediction, recommendation, or formal result.
17 mechanisms · View full form family
- Block-Adjusted Effect Estimator — Combines the within-block treatment contrasts into a single effect estimate using prespecified weights and block-aware uncertainty, so the analysis matches the way units were actually assigned.
- Bonferroni-Like Correction — Stiffens each test's significance bar in proportion to how many tests share the family, so that clearing it stays hard even after many simultaneous attempts.
- Closed-Form Power Calculation — Solves the sample-size or power equation analytically, returning required N or expected power for a standard, well-characterized test in a single evaluation.
- Configurational Comparison Truth Table — Sorts cases by which combination of conditions each one has, and reads off which combinations — not which single factors — go with the outcome.
- Counterfactual Contrast Memo — Argues one case's causal claim by spelling out what would have happened absent the cause, anchored to a closely matched case where the cause was in fact missing.
- False Discovery Rate Control — Ranks a whole family of results and draws the significance line to hold the expected share of false discoveries below a chosen rate, trading a little purity for far more power.
- Minimum Detectable Effect Table — Reverses the sample-size question — for a design whose size is already fixed by budget or population, tabulates the smallest effect it can detect at the target power.
- Most-Different Systems Design — Compares cases that differ in almost every way yet share the same outcome, so the one condition they all hold in common becomes the candidate cause.
- Most-Similar Systems Design — Compares cases held alike on their background conditions but differing in outcome, so the handful of remaining differences becomes the short list of candidate causes.
- Multiverse Analysis Report — Runs the analysis across every defensible analytic choice at once and shows the whole spread of results, exposing whether the headline depends on one lucky path.
- Operating Characteristic Curve — Plots detection probability across the full range of plausible true effects, replacing a single power number with the whole sensitivity profile of the design.
- Power Sensitivity Grid — Recomputes power across a grid of alternative variance, attrition, and compliance assumptions to expose designs that only clear the bar under optimistic inputs.
- Rival Explanation Elimination Table — Lays every candidate explanation for an outcome side by side and rules each out by the evidence it would predict but the cases do not show.
- Sensitivity to Case-Set Analysis — Re-runs the comparison while dropping, swapping, or adding cases, to see whether the conclusion survives the particular set of cases that happened to be chosen.
- Simulation-Based Power Analysis — Estimates power for a complex or nonstandard design by repeatedly generating synthetic datasets under an assumed effect and running the actual planned analysis on each.
- Standardized Mean Difference Table — Reports each baseline covariate's between-group gap on a unit-free standardized scale, so imbalance is judged against a fixed threshold rather than a sample-size-sensitive p-value.
- Within-Block Randomization Inference — Tests the treatment effect by re-enacting only the assignment permutations the actual blocked randomization could have produced, deriving p-values and intervals from the design itself rather than a distributional model.
Assessment, Review & Assurance¶
A bounded evaluation of existing evidence, work, compliance, or readiness that produces a finding, approval, correction, or disposition.
16 mechanisms · View full form family
- A/B Test Interpretation Protocol — Reads a live randomized experiment against a pre-declared primary metric and launch criteria, turning the measured difference into a ship, hold, or iterate decision.
- Benchmark Deduplication Scan — Searches the training and development corpus for copies or restatements of the evaluation benchmark, so a memorised answer can't masquerade as a solved problem.
- Blind Integrity Questionnaire — A questionnaire that asks masked roles what condition they believe they encountered and why.
- Blinded Outcome Adjudication — A procedure in which evaluators judge outcomes from evidence packets that omit condition or source identity.
- Case Selection Bias Audit — Interrogates how the cases were chosen — above all whether they were picked because they already show the outcome — and demands the negative cases the choice left out.
- Control Condition Fidelity Checklist — An item-by-item verification that the control arm, as actually delivered, matched its specification — that the intended differences were present and the required equivalences held.
- Equivalence or Noninferiority Test — Implements a variant where the goal is to show sufficiently small difference or no unacceptable loss rather than superiority.
- Feature Availability Audit — Walks every candidate input and asks whether its value would truly have been known at decision time, cataloguing the fields that would not.
- Inspection Pass/Fail Test — Applies predefined criteria to classify an item, process, or condition as acceptable or unacceptable.
- Label Proxy Screen — Scans every candidate feature for the tell-tale signature of a target proxy — a column that is suspiciously predictive because it is really a downstream trace of the outcome — and files the suspects for confirmation.
- Measurement Equivalence Audit — Checks that each variable denotes the same construct and is measured the same way in every case before any cross-case difference is trusted.
- Null Hypothesis Significance Test — Implements a default-versus-alternative comparison with a formal statistical threshold under specified assumptions.
- Quality Acceptance Test — Uses predefined acceptance criteria to decide whether a product, batch, process, or deliverable meets a required standard.
- Randomization Integrity Audit — A forensic check that the assignment actually recorded in the data matches the intended randomization — right allocation ratio, right sequence, no overrides or broken linkage.
- Stratified Balance Check — Verifies covariate balance within each stratum, block, cluster, or site — at the true unit of assignment — instead of trusting a pooled comparison that can hide local imbalance.
- Within-Case Process Tracing — Follows the causal chain inside a single case step by step, testing whether the proposed mechanism actually left the traces it should have.
Control, Automation & Runtime¶
A state-dependent executable mechanism that senses, triggers, schedules, filters, throttles, routes, or actuates during operation.
3 mechanisms · View full form family
- As-Of Join Rule — Joins each record only to the feature values that were already knowable as of that record's decision timestamp, so no later information leaks into a training row.
- Central Randomization and Masking Service — A central system that generates random assignments and releases only the coded, role-appropriate information each site needs to act.
- Covariate-Adaptive Randomization — Adjusts each unit's assignment probability as enrollment proceeds to minimize the running imbalance across many prognostic covariates, without pre-defining fixed strata.
Decision, Gate & Allocation¶
A bounded selection, disposition, routing, admission, prioritization, matching, or allocation among eligible alternatives.
2 mechanisms · View full form family
- Permuted-Block Sequence — Generates randomized treatment sequences in short fixed-length blocks so the allocation ratio stays near-balanced throughout enrollment, at the cost of making late-in-block assignments guessable.
- Sequential Review Gate — Re-evaluates evidence at predefined milestones while controlling how interim findings change action.
Experiment, Test & Rehearsal¶
An active probe, controlled variation, simulated condition, or practiced execution used to generate evidence or readiness.
20 mechanisms · View full form family
- Attention Control Script — A scripted contact routine that gives the control group the same amount of human attention and time as the treatment, minus the active ingredient, so a positive result cannot be credited to attention alone.
- Cluster or Site Blocking — Blocks whole clusters — sites, classrooms, batches, communities — that are the actual unit of assignment, then compares treatments within groups of comparable clusters.
- Confirmatory Follow-Up — Turns one promising exploratory lead into a single pre-specified confirmatory test, so a pattern found by searching must earn its status on fresh ground.
- Double-Blind Trial Protocol — A protocol that masks both recipients and delivery personnel from knowing active versus comparator assignment.
- Falsification Protocol — Specifies what evidence would count against a favored claim before the evidence is sought.
- Fresh Holdout Retest — Re-scores the frozen model on newly collected or freshly sealed cases the moment its old holdout is suspected of contamination, measuring how much of the reported skill survives.
- Holdout Validation — Seals a slice of the evidence away untouched during all discovery, then judges the selected finding once against that fresh partition.
- Incomplete-Block Design — Assigns only a connected subset of the treatments to each block when a block cannot hold them all, arranging the overlaps so every treatment comparison is still recoverable somewhere.
- Leakage Ablation Test — Removes a suspected leak pathway, refits, and reads the drop in performance — a collapse convicts the pathway and its size is the leak's severity, while the leak-free score is the honest number to expect in deployment.
- Matched-Pair Randomization — Forms pairs of maximally similar units and randomizes treatment within each pair, so every comparison is between two units already alike on what predicts the outcome.
- Nested Cross-Validation — Wraps model selection in an inner cross-validation loop nested inside an outer one, so hyperparameters and model choices are never tuned on the same data used to report performance.
- Pilot Variance Estimation — Runs a small pilot to measure the variance, baseline rate, and dropout that every power calculation depends on, replacing guessed nuisance parameters with data.
- Placebo or Sham Procedure — An inert but convincingly treatment-like stimulus — a dummy pill, a fake procedure — that reproduces the ritual and expectancy of the treatment while delivering none of the active mechanism, so the specific effect can be separated from the placebo response.
- Randomized Complete-Block Design — Places every treatment condition once inside each block, so all comparisons are made within homogeneous blocks and between-block nuisance variation is removed from the contrast.
- Replication Case Sampling Cycle — Adds new cases in deliberate rounds — some expected to repeat the result, some expected to overturn it — to map where a finding holds and where it stops.
- Replication Study — Re-runs the finding from scratch in independent hands to see whether it survives outside the conditions and choices that first produced it.
- Stratified Randomization Schedule — Divides units into categorical strata defined by a few strong pretreatment predictors and runs a separate randomization inside each stratum, forcing balance on those factors by construction.
- Time, Batch, Run, or Location Block — Treats operational conditions — production runs, time periods, machines, rooms, fields, operators — as blocks, so treatments are compared within the same run and drift between runs stays out of the contrast.
- Time-Based Holdout — Splits data by time rather than at random — training on everything before a cutoff and evaluating only on what came after — so a model meant to predict the future is graded on a genuine future it never saw.
- Waitlist Control Schedule — A timed-access plan in which control participants receive the intervention after a defined delay, creating an early-versus-delayed contrast while guaranteeing eventual access — with outcomes measured before the wait ends.
Interface, Display & Cue¶
A user-facing perceptual surface or interactive affordance that presents status, options, warnings, prompts, or controls.
1 mechanism · View full form family
- Covariate Balance Plot — A figure — often a Love plot — that arrays every covariate's standardized imbalance against a tolerance reference line, before and after any adjustment, so the whole balance picture reads at a glance.
Monitoring, Sensing & Alerting¶
Ongoing or repeated observation of actual state that emits measurements, indicators, dashboards, surveillance signals, or alerts.
2 mechanisms · View full form family
- Automated A/B Balance Dashboard — A live monitoring surface that continuously checks the assignment split and baseline balance of a running online experiment and alarms the moment traffic allocation breaks.
- Duplicate and Near-Duplicate Scan — Hunts for the same or nearly-identical cases sitting on both sides of a split — the overlap that quietly turns memorisation into apparent generalisation.
Organization, Role & Governance¶
An enduring actor, authority, body, program, service, pooled capacity, or institutional arrangement whose mandate, membership, resources, or continuity is operative.
1 mechanism · View full form family
- Comparative Case Review Panel — A standing panel that stress-tests the cross-case interpretation with domain and stakeholder members, and records why each reading was accepted, revised, or sent back.
Protocol, Workflow & Routine¶
A repeatable ordered sequence of actions, handoffs, states, or escalation steps, including procedures, runbooks, routines, recovery sequences, and lifecycle workflows.
4 mechanisms · View full form family
- Deviant Case Follow-Up Protocol — Governs what to do with a case that breaks the cross-case pattern — re-investigate it before deciding whether it is error, omission, or a genuine limit on the theory.
- Emergency Unblinding Procedure — A controlled pathway that reveals one participant's assignment when safety requires it, under authorization and with a permanent record.
- Entity-Grouped Split — Partitions train and test by the underlying entity — patient, speaker, site, household, lineage — so no single entity has rows on both sides of the boundary.
- Matched Case Pairing Protocol — Builds one-to-one case pairs matched on background factors, so within each pair only the factor of interest is left free to vary.
Record, Log & Register¶
A durable, usually accumulating account of actual events, decisions, custody, exceptions, or state transitions whose value depends on history, provenance, or accountability.
3 mechanisms · View full form family
- Claim Registry — A living ledger of every attempted claim — its status, owner, and follow-up burden — so selective memory can't erase the failed tries that made a discovery look surprising.
- Contamination Monitoring Log — A running record kept during execution that captures every instance of crossover, spillover, and drift so the tested contrast can be reported as what actually happened, not what was planned.
- Holdout Access Log — Records every query, submission, and human view of protected evaluation material, so exposure is metered and a spent or peeked-at holdout stops being trusted as fresh evidence.
Representation, Specification & Plan¶
A non-executable information artifact that externalizes understood, desired, or future structure, including maps, matrices, templates, checklists, specifications, reports, plans, and schedules.
11 mechanisms · View full form family
- Balance Exception Report — A focused write-up of only the covariates that breached tolerance — the breach, the decided response, and the independent reviewer's sign-off — kept with the study record.
- Baseline Characteristics Table — The arm-by-arm 'Table 1' that enumerates a frozen set of pre-treatment covariates and displays their distribution across study groups as the published balance record.
- Blinded Data Analysis Plan — Pre-specifies every analytic decision before condition labels are revealed, so analyst discretion cannot be steered toward the favored result.
- Case Universe Sampling Frame — Fixes the population of cases the study could have chosen — the boundary, the unit, and the eligibility rule — before any case is picked.
- Comparative Historical Timeline — Lines up the sequence of events across cases on one shared clock so you can see whether the supposed cause actually came before the effect in each.
- Cross-Case Evidence Matrix Tool — Assembles a cases-by-variables grid — one row per case, one column per factor — filled with comparably-coded, sourced values so patterns can be read across cases.
- External Control Justification Memo — A written case for using patients or data from outside the current study — historical cohorts, registries, natural-history data — as the comparator, filtering them for comparability and bounding what the borrowed contrast can claim.
- Masked Label Codebook — A protected mapping between neutral labels and true conditions.
- Scientific Claim Evaluation Template — Prompts analysts to state claim, default, alternative, evidence, assumptions, thresholds, error costs, and interpretation limits.
- Standard-Care Comparator Specification — Defines a single, prescribed best-current-practice regimen as the active comparator arm, so the study answers the adoption question — does the new option improve on the real alternative — rather than beating a strawman.
- Usual-Care Inventory Form — A structured survey that documents what 'usual care' actually contains — service by service, site by site — so the black-box comparator is described rather than assumed, and re-checked when practice shifts.
Rule, Policy & Commitment¶
A standing constraint, permission, default, threshold, quota, obligation, right, or conditional action rule governing future behavior.
9 mechanisms · View full form family
- Alpha-Spending Plan — Treats the total false-positive budget as a currency spent in pre-planned fractions across repeated interim looks, so peeking at accumulating data never inflates the error rate.
- Control Arm Protocol — The master operating document for the comparator arm — what control units receive, are denied, are told, and are measured on, plus how deviations are handled — so the control is reproducible rather than a label.
- Decision Threshold Rule — Operationalizes the evidence threshold as a cut point, burden, gate, or standard that changes action status.
- Legal Burden-of-Proof Analog — Uses a formal presumption and evidentiary burden to protect against costly false judgments.
- Metric Hierarchy — Ranks metrics into primary, secondary, and exploratory tiers before the data land, so a disappointing primary can't be quietly swapped for a flattering secondary.
- Pre-Analysis Power Statement — Records the target effect, error budget, frame, and interpretation boundaries before data collection, turning power calibration into a pre-committed design contract.
- Preprocessing Fit-on-Training-Only — Requires every fitted transform — scalers, imputers, encoders, vectorizers, feature selectors, resamplers — to learn its parameters from the training partition alone, then apply unchanged to validation and test.
- Preregistration — Timestamps the hypotheses, primary outcome, and analysis plan before the data exist, so what counts as confirmatory is fixed in advance rather than chosen after.
- Prespecified Adjusted Estimation Plan — A pre-registered rule that fixes, before any outcome is seen, which baseline covariates the effect estimate will adjust for and how — so adjustment corrects imbalance without becoming a fishing license.
Structure, Architecture & Configuration¶
An enduring physical, digital, spatial, material, or organizational topology, partition, boundary, component arrangement, or configured state.
2 mechanisms · View full form family
- Sham or Placebo Control — An inactive or alternative comparator designed to preserve credibility and mask active-condition identity.
- Single-Blind Participant Masking — Masks the recipient alone from knowing which condition they received, while implementers stay informed.