Skip to content

Experimental Comparison & Hypothesis-Test Design

← Back to Uncertainty, Evidence & Inference Failure

Treatment, control, assignment, blinding, power, and evidence thresholds are insufficiently designed to support the intended comparison.

91 mechanisms across 9 solution archetypes. This is a recurring problem pattern within Uncertainty, Evidence & Inference Failure; the mechanisms below inherit it from the primary archetype they instantiate.

Because this set contains more than 30 mechanisms, it is divided by form family—the concrete kind of thing a practitioner deploys, enacts, maintains, or convenes. This is a browsing subdivision only; it does not change the inherited problem classification. Click a form below to jump to its fully visible section.

Form familyMechanismsDescription
Analysis, Modeling & Optimization17A calculation, model, estimator, diagnostic, comparison, simulation, or optimization that transforms inputs into an inference, prediction, recommendation, or formal result.
Assessment, Review & Assurance16A bounded evaluation of existing evidence, work, compliance, or readiness that produces a finding, approval, correction, or disposition.
Control, Automation & Runtime3A state-dependent executable mechanism that senses, triggers, schedules, filters, throttles, routes, or actuates during operation.
Decision, Gate & Allocation2A bounded selection, disposition, routing, admission, prioritization, matching, or allocation among eligible alternatives.
Experiment, Test & Rehearsal20An active probe, controlled variation, simulated condition, or practiced execution used to generate evidence or readiness.
Interface, Display & Cue1A user-facing perceptual surface or interactive affordance that presents status, options, warnings, prompts, or controls.
Monitoring, Sensing & Alerting2Ongoing or repeated observation of actual state that emits measurements, indicators, dashboards, surveillance signals, or alerts.
Organization, Role & Governance1An enduring actor, authority, body, program, service, pooled capacity, or institutional arrangement whose mandate, membership, resources, or continuity is operative.
Protocol, Workflow & Routine4A repeatable ordered sequence of actions, handoffs, states, or escalation steps, including procedures, runbooks, routines, recovery sequences, and lifecycle workflows.
Record, Log & Register3A durable, usually accumulating account of actual events, decisions, custody, exceptions, or state transitions whose value depends on history, provenance, or accountability.
Representation, Specification & Plan11A non-executable information artifact that externalizes understood, desired, or future structure, including maps, matrices, templates, checklists, specifications, reports, plans, and schedules.
Rule, Policy & Commitment9A standing constraint, permission, default, threshold, quota, obligation, right, or conditional action rule governing future behavior.
Structure, Architecture & Configuration2An enduring physical, digital, spatial, material, or organizational topology, partition, boundary, component arrangement, or configured state.

Analysis, Modeling & Optimization

A calculation, model, estimator, diagnostic, comparison, simulation, or optimization that transforms inputs into an inference, prediction, recommendation, or formal result.

17 mechanisms · View full form family

  • Block-Adjusted Effect Estimator — Combines the within-block treatment contrasts into a single effect estimate using prespecified weights and block-aware uncertainty, so the analysis matches the way units were actually assigned.
  • Bonferroni-Like Correction — Stiffens each test's significance bar in proportion to how many tests share the family, so that clearing it stays hard even after many simultaneous attempts.
  • Closed-Form Power Calculation — Solves the sample-size or power equation analytically, returning required N or expected power for a standard, well-characterized test in a single evaluation.
  • Configurational Comparison Truth Table — Sorts cases by which combination of conditions each one has, and reads off which combinations — not which single factors — go with the outcome.
  • Counterfactual Contrast Memo — Argues one case's causal claim by spelling out what would have happened absent the cause, anchored to a closely matched case where the cause was in fact missing.
  • False Discovery Rate Control — Ranks a whole family of results and draws the significance line to hold the expected share of false discoveries below a chosen rate, trading a little purity for far more power.
  • Minimum Detectable Effect Table — Reverses the sample-size question — for a design whose size is already fixed by budget or population, tabulates the smallest effect it can detect at the target power.
  • Most-Different Systems Design — Compares cases that differ in almost every way yet share the same outcome, so the one condition they all hold in common becomes the candidate cause.
  • Most-Similar Systems Design — Compares cases held alike on their background conditions but differing in outcome, so the handful of remaining differences becomes the short list of candidate causes.
  • Multiverse Analysis Report — Runs the analysis across every defensible analytic choice at once and shows the whole spread of results, exposing whether the headline depends on one lucky path.
  • Operating Characteristic Curve — Plots detection probability across the full range of plausible true effects, replacing a single power number with the whole sensitivity profile of the design.
  • Power Sensitivity Grid — Recomputes power across a grid of alternative variance, attrition, and compliance assumptions to expose designs that only clear the bar under optimistic inputs.
  • Rival Explanation Elimination Table — Lays every candidate explanation for an outcome side by side and rules each out by the evidence it would predict but the cases do not show.
  • Sensitivity to Case-Set Analysis — Re-runs the comparison while dropping, swapping, or adding cases, to see whether the conclusion survives the particular set of cases that happened to be chosen.
  • Simulation-Based Power Analysis — Estimates power for a complex or nonstandard design by repeatedly generating synthetic datasets under an assumed effect and running the actual planned analysis on each.
  • Standardized Mean Difference Table — Reports each baseline covariate's between-group gap on a unit-free standardized scale, so imbalance is judged against a fixed threshold rather than a sample-size-sensitive p-value.
  • Within-Block Randomization Inference — Tests the treatment effect by re-enacting only the assignment permutations the actual blocked randomization could have produced, deriving p-values and intervals from the design itself rather than a distributional model.

Assessment, Review & Assurance

A bounded evaluation of existing evidence, work, compliance, or readiness that produces a finding, approval, correction, or disposition.

16 mechanisms · View full form family

  • A/B Test Interpretation Protocol — Reads a live randomized experiment against a pre-declared primary metric and launch criteria, turning the measured difference into a ship, hold, or iterate decision.
  • Benchmark Deduplication Scan — Searches the training and development corpus for copies or restatements of the evaluation benchmark, so a memorised answer can't masquerade as a solved problem.
  • Blind Integrity Questionnaire — A questionnaire that asks masked roles what condition they believe they encountered and why.
  • Blinded Outcome Adjudication — A procedure in which evaluators judge outcomes from evidence packets that omit condition or source identity.
  • Case Selection Bias Audit — Interrogates how the cases were chosen — above all whether they were picked because they already show the outcome — and demands the negative cases the choice left out.
  • Control Condition Fidelity Checklist — An item-by-item verification that the control arm, as actually delivered, matched its specification — that the intended differences were present and the required equivalences held.
  • Equivalence or Noninferiority Test — Implements a variant where the goal is to show sufficiently small difference or no unacceptable loss rather than superiority.
  • Feature Availability Audit — Walks every candidate input and asks whether its value would truly have been known at decision time, cataloguing the fields that would not.
  • Inspection Pass/Fail Test — Applies predefined criteria to classify an item, process, or condition as acceptable or unacceptable.
  • Label Proxy Screen — Scans every candidate feature for the tell-tale signature of a target proxy — a column that is suspiciously predictive because it is really a downstream trace of the outcome — and files the suspects for confirmation.
  • Measurement Equivalence Audit — Checks that each variable denotes the same construct and is measured the same way in every case before any cross-case difference is trusted.
  • Null Hypothesis Significance Test — Implements a default-versus-alternative comparison with a formal statistical threshold under specified assumptions.
  • Quality Acceptance Test — Uses predefined acceptance criteria to decide whether a product, batch, process, or deliverable meets a required standard.
  • Randomization Integrity Audit — A forensic check that the assignment actually recorded in the data matches the intended randomization — right allocation ratio, right sequence, no overrides or broken linkage.
  • Stratified Balance Check — Verifies covariate balance within each stratum, block, cluster, or site — at the true unit of assignment — instead of trusting a pooled comparison that can hide local imbalance.
  • Within-Case Process Tracing — Follows the causal chain inside a single case step by step, testing whether the proposed mechanism actually left the traces it should have.

Control, Automation & Runtime

A state-dependent executable mechanism that senses, triggers, schedules, filters, throttles, routes, or actuates during operation.

3 mechanisms · View full form family

  • As-Of Join Rule — Joins each record only to the feature values that were already knowable as of that record's decision timestamp, so no later information leaks into a training row.
  • Central Randomization and Masking Service — A central system that generates random assignments and releases only the coded, role-appropriate information each site needs to act.
  • Covariate-Adaptive Randomization — Adjusts each unit's assignment probability as enrollment proceeds to minimize the running imbalance across many prognostic covariates, without pre-defining fixed strata.

Decision, Gate & Allocation

A bounded selection, disposition, routing, admission, prioritization, matching, or allocation among eligible alternatives.

2 mechanisms · View full form family

  • Permuted-Block Sequence — Generates randomized treatment sequences in short fixed-length blocks so the allocation ratio stays near-balanced throughout enrollment, at the cost of making late-in-block assignments guessable.
  • Sequential Review Gate — Re-evaluates evidence at predefined milestones while controlling how interim findings change action.

Experiment, Test & Rehearsal

An active probe, controlled variation, simulated condition, or practiced execution used to generate evidence or readiness.

20 mechanisms · View full form family

  • Attention Control Script — A scripted contact routine that gives the control group the same amount of human attention and time as the treatment, minus the active ingredient, so a positive result cannot be credited to attention alone.
  • Cluster or Site Blocking — Blocks whole clusters — sites, classrooms, batches, communities — that are the actual unit of assignment, then compares treatments within groups of comparable clusters.
  • Confirmatory Follow-Up — Turns one promising exploratory lead into a single pre-specified confirmatory test, so a pattern found by searching must earn its status on fresh ground.
  • Double-Blind Trial Protocol — A protocol that masks both recipients and delivery personnel from knowing active versus comparator assignment.
  • Falsification Protocol — Specifies what evidence would count against a favored claim before the evidence is sought.
  • Fresh Holdout Retest — Re-scores the frozen model on newly collected or freshly sealed cases the moment its old holdout is suspected of contamination, measuring how much of the reported skill survives.
  • Holdout Validation — Seals a slice of the evidence away untouched during all discovery, then judges the selected finding once against that fresh partition.
  • Incomplete-Block Design — Assigns only a connected subset of the treatments to each block when a block cannot hold them all, arranging the overlaps so every treatment comparison is still recoverable somewhere.
  • Leakage Ablation Test — Removes a suspected leak pathway, refits, and reads the drop in performance — a collapse convicts the pathway and its size is the leak's severity, while the leak-free score is the honest number to expect in deployment.
  • Matched-Pair Randomization — Forms pairs of maximally similar units and randomizes treatment within each pair, so every comparison is between two units already alike on what predicts the outcome.
  • Nested Cross-Validation — Wraps model selection in an inner cross-validation loop nested inside an outer one, so hyperparameters and model choices are never tuned on the same data used to report performance.
  • Pilot Variance Estimation — Runs a small pilot to measure the variance, baseline rate, and dropout that every power calculation depends on, replacing guessed nuisance parameters with data.
  • Placebo or Sham Procedure — An inert but convincingly treatment-like stimulus — a dummy pill, a fake procedure — that reproduces the ritual and expectancy of the treatment while delivering none of the active mechanism, so the specific effect can be separated from the placebo response.
  • Randomized Complete-Block Design — Places every treatment condition once inside each block, so all comparisons are made within homogeneous blocks and between-block nuisance variation is removed from the contrast.
  • Replication Case Sampling Cycle — Adds new cases in deliberate rounds — some expected to repeat the result, some expected to overturn it — to map where a finding holds and where it stops.
  • Replication Study — Re-runs the finding from scratch in independent hands to see whether it survives outside the conditions and choices that first produced it.
  • Stratified Randomization Schedule — Divides units into categorical strata defined by a few strong pretreatment predictors and runs a separate randomization inside each stratum, forcing balance on those factors by construction.
  • Time, Batch, Run, or Location Block — Treats operational conditions — production runs, time periods, machines, rooms, fields, operators — as blocks, so treatments are compared within the same run and drift between runs stays out of the contrast.
  • Time-Based Holdout — Splits data by time rather than at random — training on everything before a cutoff and evaluating only on what came after — so a model meant to predict the future is graded on a genuine future it never saw.
  • Waitlist Control Schedule — A timed-access plan in which control participants receive the intervention after a defined delay, creating an early-versus-delayed contrast while guaranteeing eventual access — with outcomes measured before the wait ends.

Interface, Display & Cue

A user-facing perceptual surface or interactive affordance that presents status, options, warnings, prompts, or controls.

1 mechanism · View full form family

  • Covariate Balance Plot — A figure — often a Love plot — that arrays every covariate's standardized imbalance against a tolerance reference line, before and after any adjustment, so the whole balance picture reads at a glance.

Monitoring, Sensing & Alerting

Ongoing or repeated observation of actual state that emits measurements, indicators, dashboards, surveillance signals, or alerts.

2 mechanisms · View full form family

  • Automated A/B Balance Dashboard — A live monitoring surface that continuously checks the assignment split and baseline balance of a running online experiment and alarms the moment traffic allocation breaks.
  • Duplicate and Near-Duplicate Scan — Hunts for the same or nearly-identical cases sitting on both sides of a split — the overlap that quietly turns memorisation into apparent generalisation.

Organization, Role & Governance

An enduring actor, authority, body, program, service, pooled capacity, or institutional arrangement whose mandate, membership, resources, or continuity is operative.

1 mechanism · View full form family

  • Comparative Case Review Panel — A standing panel that stress-tests the cross-case interpretation with domain and stakeholder members, and records why each reading was accepted, revised, or sent back.

Protocol, Workflow & Routine

A repeatable ordered sequence of actions, handoffs, states, or escalation steps, including procedures, runbooks, routines, recovery sequences, and lifecycle workflows.

4 mechanisms · View full form family

  • Deviant Case Follow-Up Protocol — Governs what to do with a case that breaks the cross-case pattern — re-investigate it before deciding whether it is error, omission, or a genuine limit on the theory.
  • Emergency Unblinding Procedure — A controlled pathway that reveals one participant's assignment when safety requires it, under authorization and with a permanent record.
  • Entity-Grouped Split — Partitions train and test by the underlying entity — patient, speaker, site, household, lineage — so no single entity has rows on both sides of the boundary.
  • Matched Case Pairing Protocol — Builds one-to-one case pairs matched on background factors, so within each pair only the factor of interest is left free to vary.

Record, Log & Register

A durable, usually accumulating account of actual events, decisions, custody, exceptions, or state transitions whose value depends on history, provenance, or accountability.

3 mechanisms · View full form family

  • Claim Registry — A living ledger of every attempted claim — its status, owner, and follow-up burden — so selective memory can't erase the failed tries that made a discovery look surprising.
  • Contamination Monitoring Log — A running record kept during execution that captures every instance of crossover, spillover, and drift so the tested contrast can be reported as what actually happened, not what was planned.
  • Holdout Access Log — Records every query, submission, and human view of protected evaluation material, so exposure is metered and a spent or peeked-at holdout stops being trusted as fresh evidence.

Representation, Specification & Plan

A non-executable information artifact that externalizes understood, desired, or future structure, including maps, matrices, templates, checklists, specifications, reports, plans, and schedules.

11 mechanisms · View full form family

  • Balance Exception Report — A focused write-up of only the covariates that breached tolerance — the breach, the decided response, and the independent reviewer's sign-off — kept with the study record.
  • Baseline Characteristics Table — The arm-by-arm 'Table 1' that enumerates a frozen set of pre-treatment covariates and displays their distribution across study groups as the published balance record.
  • Blinded Data Analysis Plan — Pre-specifies every analytic decision before condition labels are revealed, so analyst discretion cannot be steered toward the favored result.
  • Case Universe Sampling Frame — Fixes the population of cases the study could have chosen — the boundary, the unit, and the eligibility rule — before any case is picked.
  • Comparative Historical Timeline — Lines up the sequence of events across cases on one shared clock so you can see whether the supposed cause actually came before the effect in each.
  • Cross-Case Evidence Matrix Tool — Assembles a cases-by-variables grid — one row per case, one column per factor — filled with comparably-coded, sourced values so patterns can be read across cases.
  • External Control Justification Memo — A written case for using patients or data from outside the current study — historical cohorts, registries, natural-history data — as the comparator, filtering them for comparability and bounding what the borrowed contrast can claim.
  • Masked Label Codebook — A protected mapping between neutral labels and true conditions.
  • Scientific Claim Evaluation Template — Prompts analysts to state claim, default, alternative, evidence, assumptions, thresholds, error costs, and interpretation limits.
  • Standard-Care Comparator Specification — Defines a single, prescribed best-current-practice regimen as the active comparator arm, so the study answers the adoption question — does the new option improve on the real alternative — rather than beating a strawman.
  • Usual-Care Inventory Form — A structured survey that documents what 'usual care' actually contains — service by service, site by site — so the black-box comparator is described rather than assumed, and re-checked when practice shifts.

Rule, Policy & Commitment

A standing constraint, permission, default, threshold, quota, obligation, right, or conditional action rule governing future behavior.

9 mechanisms · View full form family

  • Alpha-Spending Plan — Treats the total false-positive budget as a currency spent in pre-planned fractions across repeated interim looks, so peeking at accumulating data never inflates the error rate.
  • Control Arm Protocol — The master operating document for the comparator arm — what control units receive, are denied, are told, and are measured on, plus how deviations are handled — so the control is reproducible rather than a label.
  • Decision Threshold Rule — Operationalizes the evidence threshold as a cut point, burden, gate, or standard that changes action status.
  • Legal Burden-of-Proof Analog — Uses a formal presumption and evidentiary burden to protect against costly false judgments.
  • Metric Hierarchy — Ranks metrics into primary, secondary, and exploratory tiers before the data land, so a disappointing primary can't be quietly swapped for a flattering secondary.
  • Pre-Analysis Power Statement — Records the target effect, error budget, frame, and interpretation boundaries before data collection, turning power calibration into a pre-committed design contract.
  • Preprocessing Fit-on-Training-Only — Requires every fitted transform — scalers, imputers, encoders, vectorizers, feature selectors, resamplers — to learn its parameters from the training partition alone, then apply unchanged to validation and test.
  • Preregistration — Timestamps the hypotheses, primary outcome, and analysis plan before the data exist, so what counts as confirmatory is fixed in advance rather than chosen after.
  • Prespecified Adjusted Estimation Plan — A pre-registered rule that fixes, before any outcome is seen, which baseline covariates the effect estimate will adjust for and how — so adjustment corrects imbalance without becoming a fishing license.

Structure, Architecture & Configuration

An enduring physical, digital, spatial, material, or organizational topology, partition, boundary, component arrangement, or configured state.

2 mechanisms · View full form family

  • Sham or Placebo Control — An inactive or alternative comparator designed to preserve credibility and mask active-condition identity.
  • Single-Blind Participant Masking — Masks the recipient alone from knowing which condition they received, while implementers stay informed.