Statistics & Research Methods¶
43 domain-specific abstractions whose origin domain is Statistics & Research Methods.
- Anscombe's Quartet — Four hand-built bivariate datasets that share nearly identical low-order summary statistics yet have radically different scatter geometries, demonstrating that summaries are lossy projections and that one must visualize before granting any parametric summary authority over the data.
- Atomistic Fallacy — The inferential error of reading a within-individual relationship directly onto a group or population, ignoring the contextual variance operating only at the group level — the mirror image of the ecological fallacy.
- Attenuation Bias — The systematic shrinkage of an OLS regression coefficient toward zero caused by classical random noise in the regressor — the estimate equals the true slope times the reliability ratio, a known-sign distortion invertible by dividing out that ratio or instrumenting.
- Benford's Law — Score a dataset's honesty by checking whether its leading digits follow the fixed logarithmic curve log₁₀((d+1)/d) — about 30% start with 1, only 5% with 9 — that scale-spanning multiplicative data must obey.
- Benjamini–Hochberg Procedure — A step-up rule that sorts m p-values and rejects through the largest rank k where p(k) ≤ (k/m)·α, bounding the false discovery rate — the expected proportion of false rejections among discoveries — rather than the probability of any false positive.
- Bonferroni Correction — Control the family-wise probability of any false rejection across m tests by comparing each p-value with alpha/m, or equivalently multiplying each p-value by m, without requiring independence.
- Carryover Effect — The validity threat in crossover and within-subject designs where residual influence from an earlier treatment persists into a later measurement window, biasing the contrast — its magnitude set by the unit's relaxation time against the inter-treatment gap.
- Causal Inference — Infer the effect of changing X on Y from data by fixing a causal estimand and defending an identification design or assumption that separates that effect from noncausal association, then quantify its uncertainty and scope.
- Cohort Effect — An observed outcome difference associated with membership in groups defined by a shared birth or entry interval, kept distinct from aging, period shocks, and any unproven causal account of the cohort contrast.
- Collinearity Inflation — The pathological ballooning of individual regression-coefficient variance when predictors carry overlapping information — a near-singular predictor cross-product matrix that degrades per-predictor attribution while leaving joint prediction untouched, unfixable by more data.
- Difference-in-Differences — Estimate a causal effect from observational data by subtracting the control group's before-after change from the treatment group's, netting out time-invariant unit confounders and common time trends — valid only if parallel trends holds.
- Ecological Correlation — A correlation computed on aggregated group-level units that need not equal — and can reverse the sign of — the individual-level correlation, so it warrants no claim about the individuals inside the groups.
- Ecological Inference Problem — Recover individual-level joint distributions from group-level marginal totals, a many-to-one inverse problem where the data alone only pin the answer to the Duncan-Davis bounds and any tighter estimate rests on an explicit, contestable identifying assumption.
- Endogeneity — The condition in which a regressor is correlated with a model's error term — through confounding, simultaneity, or measurement error — so OLS coefficients are biased and inconsistent for the causal effect, collapsing the coefficient's causal reading while leaving its predictive one intact.
- External Validity — The warrant by which an effect estimated in one study setting can be expected to hold in a target setting outside it — holding conditional on every effect-modifying feature that differs between the two being matched or adjusted.
- False Discovery Rate — The expected proportion of false rejections among all rejected hypotheses, conventionally V/max(R,1), used as an at-scale error criterion that accepts a controlled fraction of false discoveries in exchange for power.
- File Drawer Problem — Recognize that studies with null results disproportionately go unpublished while significant ones enter the literature, so any synthesis treating the published record as the full population of conducted research systematically overestimates effect sizes toward the filter.
- Funnel Plot Asymmetry — Plot each study's effect against its precision and read a departure from the symmetric inverted-funnel expected under unbiased sampling — a gap where small null studies should be — as the visual fingerprint of a publication filter, licensing scrutiny against a fixed set of causes rather than a verdict.
- Gold-Standard Erosion — Recognize that a model scored against a mutable reference label can show stable metrics while its real validity silently degrades, because the answer key — not the model — has drifted away from the construct it once operationalized.
- HARKing (Hypothesizing After the Results are Known) — The research practice of building a hypothesis by inspecting already-collected data and then presenting it as if it had been specified in advance, silently inflating the reported false-positive rate because the test's independence assumption is violated.
- Instrumental variable — Recover the causal effect of a confounded treatment by finding a quantity Z that moves the treatment, reaches the outcome only through it, and is independent of the confounders — then reading the effect off the ratio of Z's reduced-form to first-stage effects, importing randomization the analyst never performed.
- Inter-Annotator Agreement — Measure whether independent raters applying the same coding scheme to the same items converge, using a statistic that subtracts the agreement expected by chance from the marginal label distribution.
- Internal validity — The property of an empirical study that warrants its causal claim within its own sample and setting — whether the observed intervention-outcome association is genuinely produced by the intervention rather than by confounders, selection, or bias — established by ruling out a closed catalog of named threats.
- Jeffreys-Lindley Paradox — The result that a frequentist significance test and a Bayesian posterior-odds comparison of the same data against the same point null can reach opposite verdicts, with the disagreement growing without bound as sample size increases — because the two answer different questions.
- Look-Elsewhere Effect — Discount an exciting best-of-many find by the size of the search that produced it — converting a local p-value at one scanned peak into a global p-value asking whether any peak this extreme would occur anywhere, via the trials factor.
- Lord's Paradox — Show that two arithmetically correct analyses of the same pre-post data — raw change scores versus baseline adjustment — can reach opposite verdicts about an effect, because adjustment is a causal-modeling choice and the two answer different questions depending on whether baseline is itself caused by group membership.
- Monty Hall problem — A worked three-door puzzle in which switching wins ⅔ of the time because the host's reveal was constrained by what he knew — drilling the move of updating on the protocol that produced an observation, not on its bare surface.
- Natural Experiment — A design that borrows the RCT's identification logic from a real-world process — a policy, boundary, or lottery — judged plausibly as-good-as-random, where the as-if-random assumption must be substantively defended rather than guaranteed by protocol.
- Null Ritual — The institutionalised practice of mechanically executing a null hypothesis significance test — nil-null, p-value, p < .05 verdict — severed from the alternatives, priors, effect sizes, and decision context inference requires, yet retaining full editorial authority as if it had not been.
- Omitted Variable Bias — Correct for the distortion in a regression coefficient when a left-out variable both causes the outcome and correlates with an included regressor, so the estimate absorbs the omitted effect as the signable product of two relationships.
- Publication Bias — A scientific record becomes systematically unrepresentative when the probability that a study, result, or outcome becomes publicly available depends on its direction, magnitude, statistical significance, novelty, or sponsor-favoredness.
- Receiver Operating Characteristic — Sweep a binary classifier's decision threshold across its full score range to trace every achievable sensitivity-versus-false-positive-rate tradeoff at once, factoring detection into orthogonal discriminability (the curve's height) and criterion (where the threshold sits) coordinates.
- Regression Discontinuity Design — Recover a causal effect from a threshold rule by comparing units just above and just below a sharp cutoff on a continuous running variable, where they are comparable in expectation, so any jump in the outcome at exactly the cutoff is attributable to the treatment rather than to selection.
- Reliability Paradox — Explain why tasks with robust group-level effects (Stroop, IAT) can be useless for ranking individuals: the design minimized within-subjects error for group power without guaranteeing the between-subjects variance that reliability, true-score over total variance, requires.
- Selection on Observables — Assume that, conditional on a named set of measured covariates, treatment assignment is independent of potential outcomes — so within each covariate stratum treated and untreated units are exchangeable and adjustment recovers the causal effect.
- Small-Study Effects — The meta-analytic pattern in which smaller studies report systematically larger effects than larger ones, producing funnel-plot asymmetry that inflates the pooled estimate — a shared symptom of several biases, not a diagnosis of any one cause.
- Snowball Sampling — Recruit an unenumerable population by seeding a few participants and having each nominate others along their social ties, substituting relational proximity for random selection and buying access at the cost of representativeness.
- Standard of Care — Use the currently accepted reference practice as the single dynamic baseline against which both efficacy (is a treatment better than what we already do?) and accountability (did a clinician meet what a reasonable body of practitioners would have done?) are measured by deviation.
- Stein's Paradox — The result that estimating three or more means each by its own sample mean is inadmissible under total squared-error loss — a shrinkage estimator pulling each toward a common reference achieves strictly lower joint error for every true parameter vector, however unrelated the quantities.
- Surrogate Endpoint Problem — The failure that arises when a trial's biomarker surrogate diverges from the clinical endpoint it stands in for — because the intervention acts through off-pathway mechanisms the surrogate cannot see — so individual-level correlation does not license an intervention-level claim.
- Type M Error — Quantify how much a significant effect's reported magnitude is exaggerated by the significance filter under low power, via the exaggeration ratio — the expected significant estimate divided by the true effect — computable from the design before any data exist.
- Type S Error — Quantify the risk that a statistically significant estimate points the wrong way by computing, before data collection, the probability that a two-sided significance filter is cleared from the opposite tail when the true effect is near zero relative to noise.
- Underfitting — The failure mode where a model's hypothesis class is too restrictive to capture the structure genuinely present in the data — high bias, with training and test error both elevated and close together — curable only by a richer functional form, not by more data or regularization.