Skip to content

Inference Bias & Multiple Testing

← Back to Domain-Specific Families

Abstractions about false discovery control, selective evidence, post hoc hypotheses, diagnostic discrimination, expectancy, and reasoning biases in statistical inference.

10 abstractions in this family — domain-specific abstractions that sit near one another in structural-signature space (k-means over structural-signature embeddings). Each is shown with its short description.

  • Benjamini–Hochberg Procedure — A step-up rule that sorts m p-values and rejects through the largest rank k where p(k) ≤ (k/m)·α, bounding the false discovery rate — the expected proportion of false rejections among discoveries — rather than the probability of any false positive.
  • Cherry Picking — Selectively presenting confirming evidence while suppressing disconfirming evidence from the same available population, so the offered sample gives an impression the full distribution would not support — locating the dishonesty in the selection process, not the individual data points.
  • Congruence Bias — The reasoning bias of designing only tests that could confirm a favored hypothesis rather than tests that discriminate it from equally plausible rivals — a failure of test design, not evidence evaluation, that leaves a probe locally valid but globally non-diagnostic because its likelihood ratio sits near one.
  • Conjunction Fallacy — Detect a probability-judgment error by watching for a more detailed scenario being rated more probable than the simpler scenario it is a strict subset of — the signature of resemblance quietly standing in for probability.
  • HARKing (Hypothesizing After the Results are Known) — The research practice of building a hypothesis by inspecting already-collected data and then presenting it as if it had been specified in advance, silently inflating the reported false-positive rate because the test's independence assumption is violated.
  • Null Ritual — The institutionalised practice of mechanically executing a null hypothesis significance test — nil-null, p-value, p < .05 verdict — severed from the alternatives, priors, effect sizes, and decision context inference requires, yet retaining full editorial authority as if it had not been.
  • Optimality criterion — An objective measure used to compare candidate statistical models for a hypothesis and designate the model with the best criterion value.
  • Receiver Operating Characteristic — Sweep a binary classifier's decision threshold across its full score range to trace every achievable sensitivity-versus-false-positive-rate tradeoff at once, factoring detection into orthogonal discriminability (the curve's height) and criterion (where the threshold sits) coordinates.
  • Subject-Expectancy Effect — Treat a participant's beliefs about an intervention as a genuine driver of their reported and physiological responses, so any unblinded measurement confounds the manipulation with anticipation until blinding, placebo control, or a balanced-placebo design pulls the two apart.
  • Wason Selection Task — A four-case conditional-rule probe asks which partially observed cases must be inspected to find the uniquely falsifying conjunction, exposing how logic, interpretation, and context shape human selections.