Statistics & Research Methods¶
← Back to Domain-Specific Abstractions by Domain
131 domain-specific abstractions whose origin domain is Statistics & Research Methods.
- 68–95–99.7 rule — The normal-distribution rule that about 68%, 95%, and 99.7% of probability lies within one, two, and three standard deviations of the mean.
- Anderson–Darling test — Test a sample’s agreement with a specified continuous distribution by integrating squared empirical-CDF deviations with extra weight in the tails.
- Anscombe's Quartet — Four hand-built bivariate datasets that share nearly identical low-order summary statistics yet have radically different scatter geometries, demonstrating that summaries are lossy projections and that one must visualize before granting any parametric summary authority over the data.
- Asymptotic theory (statistics) — The large-sample framework that studies limiting distributions, consistency and efficiency of estimators and tests as sample size tends to infinity.
- Atomistic Fallacy — The inferential error of reading a within-individual relationship directly onto a group or population, ignoring the contextual variance operating only at the group level — the mirror image of the ecological fallacy.
- Attenuation Bias — The systematic shrinkage of an OLS regression coefficient toward zero caused by classical random noise in the regressor — the estimate equals the true slope times the reliability ratio, a known-sign distortion invertible by dividing out that ratio or instrumenting.
- Benford's Law — Score a dataset's honesty by checking whether its leading digits follow the fixed logarithmic curve log₁₀((d+1)/d) — about 30% start with 1, only 5% with 9 — that scale-spanning multiplicative data must obey.
- Benjamini–Hochberg Procedure — A step-up rule that sorts m p-values and rejects through the largest rank k where p(k) ≤ (k/m)·α, bounding the false discovery rate — the expected proportion of false rejections among discoveries — rather than the probability of any false positive.
- Best linear unbiased prediction — The minimum-mean-square-error predictor among estimators linear in observations and unbiased for a target random effect under a specified linear mixed model.
- Binomial regression — Model a binomial response by linking each observation's success probability to predictors, keeping trial denominators, link choice, variance assumptions, and overdispersion diagnostics explicit.
- Bonferroni Correction — Control the family-wise probability of any false rejection across m tests by comparing each p-value with alpha/m, or equivalently multiplying each p-value by m, without requiring independence.
- Brown–Forsythe test — A robust test of equality of group variances obtained by applying one-way ANOVA to absolute deviations from group medians.
- C-chart — Monitor the count of nonconformities in constant-size inspection units against Poisson-based center and control limits.
- Carryover Effect — The validity threat in crossover and within-subject designs where residual influence from an earlier treatment persists into a later measurement window, biasing the contrast — its magnitude set by the unit's relaxation time against the inter-treatment gap.
- Causal Inference — Infer the effect of changing X on Y from data by fixing a causal estimand and defending an identification design or assumption that separates that effect from noncausal association, then quantify its uncertainty and scope.
- Chauvenet's criterion — Flag a single extreme observation when, under a fitted normal-error model, the expected number of sample observations at least as far from the mean is below one half.
- Coefficient of variation — A dimensionless relative-dispersion statistic equal to standard deviation divided by mean, interpreted only where the measurement scale and nonzero mean make the ratio meaningful.
- Cohort Effect — An observed outcome difference associated with membership in groups defined by a shared birth or entry interval, kept distinct from aging, period shocks, and any unproven causal account of the cohort contrast.
- Collinearity Inflation — The pathological ballooning of individual regression-coefficient variance when predictors carry overlapping information — a near-singular predictor cross-product matrix that degrades per-predictor attribution while leaving joint prediction untouched, unfixable by more data.
- Concentration parameter — A distribution-family parameter controlling how tightly probability mass clusters around a direction, center or base distribution without necessarily changing that center.
- Correlation ratio — An effect-size measure equal to the square root of between-category variance divided by total variance, detecting nonlinear mean association.
- Count data — Observations taking nonnegative integer values because they record event or object counts rather than ranks or arbitrary numeric labels.
- Deviance Information Criterion — Compare Bayesian hierarchical models by adding posterior mean deviance to an effective-complexity penalty derived from posterior deviance, using quantities readily estimated from MCMC draws.
- Difference-in-Differences — Estimate a causal effect from observational data by subtracting the control group's before-after change from the treatment group's, netting out time-invariant unit confounders and common time trends — valid only if parallel trends holds.
- Ecological Correlation — A correlation computed on aggregated group-level units that need not equal — and can reverse the sign of — the individual-level correlation, so it warrants no claim about the individuals inside the groups.
- Ecological Inference Problem — Recover individual-level joint distributions from group-level marginal totals, a many-to-one inverse problem where the data alone only pin the answer to the Duncan-Davis bounds and any tighter estimate rests on an explicit, contestable identifying assumption.
- Empirical Bayes method — Estimate a shared prior distribution or its hyperparameters from the same ensemble of observations and then perform Bayesian-style shrinkage or posterior inference conditional on that estimate.
- Empirical likelihood — A nonparametric likelihood method that assigns probabilities to observed sample points and maximizes their product subject to estimating-equation constraints.
- Empirical probability — An event-probability estimate given by its observed relative frequency in a finite sample of trials.
- Endogeneity — The condition in which a regressor is correlated with a model's error term — through confounding, simultaneity, or measurement error — so OLS coefficients are biased and inconsistent for the causal effect, collapsing the coefficient's causal reading while leaving its predictive one intact.
- External Validity — The warrant by which an effect estimated in one study setting can be expected to hold in a target setting outside it — holding conditional on every effect-modifying feature that differs between the two being matched or adjusted.
- Extreme value theory — A branch of statistics modeling the limiting behavior and tail risk of unusually large or small observations, especially block maxima and threshold exceedances.
- Factor Analysis — A latent-variable statistical model explains covariance among observed variables through fewer common factors, variable-specific loadings, and residual variation while making rotational and identification choices explicit.
- False confidence theorem — Show that a continuous data-dependent additive probability distribution can, for some false assertion, assign arbitrarily high belief with high sampling probability, motivating assertion-wise validity checks.
- False Discovery Rate — The expected proportion of false rejections among all rejected hypotheses, conventionally V/max(R,1), used as an at-scale error criterion that accepts a controlled fraction of false discoveries in exchange for power.
- Family-Wise Error Rate — The probability that a declared family of simultaneous hypothesis tests contains at least one false rejection, with weak or strong control determined by which configurations of true nulls are covered.
- File Drawer Problem — Recognize that studies with null results disproportionately go unpublished while significant ones enter the literature, so any synthesis treating the published record as the full population of conducted research systematically overestimates effect sizes toward the filter.
- Fisher Consistency — A population-level calibration property requiring an estimator or decision rule, viewed as a functional, to recover the target parameter or Bayes-optimal action when applied to the true data-generating distribution.
- Floor Effect — An instrument or scale compresses distinct low-end target states at its minimum, erasing downward discrimination and attenuating observed differences or change.
- Focused Information Criterion — Select a candidate statistical model by the estimated risk of its estimator for a declared focus parameter, allowing the preferred model to change when the inferential target changes.
- Folded-t and half-t distributions — Nonnegative distributions obtained by taking the absolute value of a Student-t variate, with the half-t arising from a centered symmetric t distribution restricted or folded at zero.
- Formation Matrix — The inverse expected or observed likelihood-information matrix expresses local parameter dispersion for covariance bounds, standard errors, and likelihood asymptotics.
- Fowlkes–Mallows Index — Compare two partitions by the geometric mean of pairwise co-membership precision and recall, rewarding pairs clustered together by both while excluding true-negative pairs from the score.
- Fraction of variance unexplained — A regression-fit statistic equal to the proportion of dependent-variable variance left unexplained by the model's predictions.
- Funnel Plot Asymmetry — Plot each study's effect against its precision and read a departure from the symmetric inverted-funnel expected under unbiased sampling — a gap where small null studies should be — as the visual fingerprint of a publication filter, licensing scrutiny against a fixed set of causes rather than a verdict.
- Generalized Randomized Block Design — A randomized block experiment with within-block replication of every treatment, enabling treatment-by-block interaction to be separated from experimental error.
- Geometric standard deviation — A dimensionless multiplicative spread factor obtained by exponentiating the standard deviation of logarithms.
- Gold-Standard Erosion — Recognize that a model scored against a mutable reference label can show stable metrics while its real validity silently degrades, because the answer key — not the model — has drifted away from the construct it once operationalized.
- Gower's Distance — Compare mixed-type records by converting each available feature to a bounded type-appropriate similarity or dissimilarity, then taking a weighted pairwise average with missingness and binary-presence rules in the denominator.
- HARKing (Hypothesizing After the Results are Known) — The research practice of building a hypothesis by inspecting already-collected data and then presenting it as if it had been specified in advance, silently inflating the reported false-positive rate because the test's independence assumption is violated.
- Higher-order statistics — Statistics based on third- or higher-order moments, cumulants or spectra that characterize distributional shape and nonlinear dependence beyond mean and covariance.
- Homoscedasticity and heteroscedasticity — Distinguish statistical models whose disturbance variance is constant across the conditioning space from models whose variance changes with predictors, fitted values, time, or another declared index.
- Instrumental variable — Recover the causal effect of a confounded treatment by finding a quantity Z that moves the treatment, reaches the outcome only through it, and is independent of the confounders — then reading the effect off the ratio of Z's reduced-form to first-stage effects, importing randomization the analyst never performed.
- Inter-Annotator Agreement — Measure whether independent raters applying the same coding scheme to the same items converge, using a statistic that subtracts the agreement expected by chance from the marginal label distribution.
- Internal validity — The property of an empirical study that warrants its causal claim within its own sample and setting — whether the observed intervention-outcome association is genuinely produced by the intervention rather than by confounders, selection, or bias — established by ruling out a closed catalog of named threats.
- Inverse probability weighting — An estimation method that weights observed units by the inverse probability of their observed sampling, treatment or response status to reconstruct a target population or intervention distribution.
- Jeffreys-Lindley Paradox — The result that a frequentist significance test and a Bayesian posterior-odds comparison of the same data against the same point null can reach opposite verdicts, with the disagreement growing without bound as sample size increases — because the two answer different questions.
- Kaplan–Meier estimator — A nonparametric product-limit estimator of a survival function from observed event times in the presence of right-censoring.
- Kernel smoother — A nonparametric estimator that predicts a function by distance-weighted averaging of nearby observations.
- Lack-of-Fit Sum of Squares — Decompose regression residual variation at replicated predictor settings into irreducible within-setting pure error and systematic discrepancy between fitted values and setting means.
- Least absolute deviations — Fit a model by minimizing the sum of absolute residuals, yielding median-centered robustness to large response outliers while retaining leverage and identifiability boundaries.
- Linear Discriminant Analysis — A supervised linear projection and classifier that separates labeled classes relative to their within-class covariance.
- Location parameter — A distribution parameter whose change translates the probability law along its sample space without changing its shape.
- Look-Elsewhere Effect — Discount an exciting best-of-many find by the size of the search that produced it — converting a local p-value at one scanned peak into a global p-value asking whether any peak this extreme would occur anywhere, via the trials factor.
- Lord's Paradox — Show that two arithmetically correct analyses of the same pre-post data — raw change scores versus baseline adjustment — can reach opposite verdicts about an effect, because adjustment is a causal-modeling choice and the two answer different questions depending on whether baseline is itself caused by group membership.
- Maximum likelihood estimation — Parameter estimation by selecting the model value that makes the observed data most likely under a specified statistical family.
- Measurement Invariance — Establish that an instrument relates latent construct values to observed responses by the same measurement rule across specified groups, occasions, or conditions before interpreting their score differences.
- Method of Moments — A parameter-estimation procedure that equates selected model moments to empirical moments and solves the resulting identifying equations.
- Midhinge — Summarize distributional location by averaging the first and third quartiles, placing the center halfway between the hinges while keeping the interquartile spread analytically separate.
- Monty Hall problem — A worked three-door puzzle in which switching wins ⅔ of the time because the host's reveal was constrained by what he knew — drilling the move of updating on the protocol that produced an observation, not on its bare surface.
- Natural Experiment — A design that borrows the RCT's identification logic from a real-world process — a policy, boundary, or lottery — judged plausibly as-good-as-random, where the as-if-random assumption must be substantively defended rather than guaranteed by protocol.
- Nonlinear Least Squares — Estimate parameters that enter a model nonlinearly by minimizing a residual sum of squares, usually through initialization-sensitive local iterations built from the residual Jacobian.
- Normal probability plot — A quantile plot comparing ordered observations with expected normal quantiles so approximate normality appears linear and systematic departures reveal skew, tails, mixtures or outliers.
- Normal-exponential-gamma distribution — A heavy-tailed continuous location-scale-shape distribution obtained through a normal variance mixture whose variance follows an exponential-gamma hierarchy.
- Np-chart — Monitor a fixed-size sequence of samples by plotting each sample's count of nonconforming units against binomial center and control limits, separating common-cause fluctuation from special-cause signals.
- Nuisance parameter — A model parameter not itself of inferential interest but necessary to account for when estimating or testing the target parameter.
- Null distribution — Represent the sampling distribution of a declared test statistic under the null hypothesis and sampling scheme used to calibrate tail probabilities, critical values, and type-I error.
- Null Ritual — The institutionalised practice of mechanically executing a null hypothesis significance test — nil-null, p-value, p < .05 verdict — severed from the alternatives, priors, effect sizes, and decision context inference requires, yet retaining full editorial authority as if it had not been.
- Omitted Variable Bias — Correct for the distortion in a regression coefficient when a left-out variable both causes the outcome and correlates with an included regressor, so the estimate absorbs the omitted effect as the signable product of two relationships.
- Optimality criterion — An objective measure used to compare candidate statistical models for a hypothesis and designate the model with the best criterion value.
- Particle Filter — Approximate a recursive hidden-state posterior with a weighted particle population that is propagated through a state model, corrected by observation likelihoods, and selectively resampled to control weight degeneracy.
- Pearson correlation coefficient — The unitless covariance of two variables divided by the product of their standard deviations, measuring linear association from minus one to one.
- Point-biserial correlation coefficient — The Pearson correlation between one continuous variable and a genuinely dichotomous variable, expressible through group means, proportions and overall standard deviation.
- Polychoric correlation — An estimate of the correlation between two latent normally distributed continuous variables inferred from their observed ordinal categories through threshold models.
- Portmanteau test — An omnibus hypothesis test designed to detect a broad family of departures from a well-specified null model rather than optimize power for one narrowly specified alternative.
- PRESS Statistic — The sum of squared leave-one-out prediction errors from a fitted regression model, computed by refitting without each case or through leverage-adjusted ordinary residuals.
- Publication Bias — A scientific record becomes systematically unrepresentative when the probability that a study, result, or outcome becomes publicly available depends on its direction, magnitude, statistical significance, novelty, or sponsor-favoredness.
- Quadrant Count Ratio — Center paired quantitative observations at their sample means, score same-side pairs as concordant and opposite-side pairs as discordant, and normalize the signed count difference by sample size to obtain a coarse association statistic in [-1, 1].
- Quantile normalization — Replace values by a shared rank-indexed reference so multiple samples have the same empirical marginal distribution while preserving within-sample rank order.
- Quantile–Quantile Plot — Pair corresponding quantiles from two distributions so reference-line alignment and systematic departures diagnose location, scale, shape, and tail disagreement.
- Quartile — One of the three cut points corresponding approximately to the 25th, 50th and 75th percentiles, dividing ordered data or a distribution into four equal-probability parts.
- Random Digit Dialing — Build a probability-oriented telephone survey frame by generating numbers within eligible numbering blocks, screening reached numbers, and weighting the resulting sample for selection and response.
- Randomness Test — Challenge a sequence against a specified stochastic null using a pattern-sensitive statistic and calibrated rejection rule, while treating a pass only as failure to detect the tested departures.
- Receiver Operating Characteristic — Sweep a binary classifier's decision threshold across its full score range to trace every achievable sensitivity-versus-false-positive-rate tradeoff at once, factoring detection into orthogonal discriminability (the curve's height) and criterion (where the threshold sits) coordinates.
- Regression analysis — A family of statistical methods for estimating conditional relationships between an outcome and one or more predictors, supporting explanation, adjustment and prediction under explicit model assumptions.
- Regression Discontinuity Design — Recover a causal effect from a threshold rule by comparing units just above and just below a sharp cutoff on a continuous running variable, where they are comparable in expectation, so any jump in the outcome at exactly the cutoff is attributable to the treatment rather than to selection.
- Reliability Paradox — Explain why tasks with robust group-level effects (Stroop, IAT) can be useless for ranking individuals: the design minimized within-subjects error for group power without guaranteeing the between-subjects variance that reliability, true-score over total variance, requires.
- Rodger's method — A post-hoc multiple-comparison framework selecting orthogonal contrasts while controlling the expected proportion of false rejection decisions.
- Sampling error — The difference between a sample statistic and the corresponding population parameter caused by observing only a sample.
- Scoring Rule — Evaluate a probabilistic forecast after its outcome by mapping the report–outcome pair to a numeric loss or reward, with propriety governing whether truthful distributions are optimal in expectation.
- Selection on Observables — Assume that, conditional on a named set of measured covariates, treatment assignment is independent of potential outcomes — so within each covariate stratum treated and untreated units are exchangeable and adjustment recovers the causal effect.
- Sign test — Test a paired-difference or one-sample median null by reducing non-tied observations to positive and negative signs and evaluating the positive count against its exact binomial distribution under a declared null probability, usually one half.
- Small-Study Effects — The meta-analytic pattern in which smaller studies report systematically larger effects than larger ones, producing funnel-plot asymmetry that inflates the pooled estimate — a shared symptom of several biases, not a diagnosis of any one cause.
- Smoothing — A scale-setting operator suppresses local or high-frequency variation in observed data to estimate a smoother component, exchanging variance and roughness for bias, lost resolution, and boundary dependence.
- Snowball Sampling — Recruit an unenumerable population by seeding a few participants and having each nominate others along their social ties, substituting relational proximity for random selection and buying access at the cost of representativeness.
- Standard of Care — Use the currently accepted reference practice as the single dynamic baseline against which both efficacy (is a treatment better than what we already do?) and accountability (did a clinician meet what a reasonable body of practitioners would have done?) are measured by deviation.
- Standard score — A normalized value equal to an observation’s deviation from a reference mean divided by the reference standard deviation.
- Statistic — A measurable function of the observed sample alone, with no dependence on unknown population parameters, used to summarize data or support estimation and testing.
- Statistical Contrast — Encode a comparison among statistical means or parameters as a zero-sum linear estimand, propagate its sampling variance, and—when design-weighted orthogonality holds—decompose uncorrelated directions under the declared design.
- Statistical Model — Represent possible observable data by a declared sample space and family of candidate probability laws—often indexed by parameters and assumptions—so estimation, testing, prediction, and uncertainty statements have an explicit conditional basis.
- Stein's Paradox — The result that estimating three or more means each by its own sample mean is inadmissible under total squared-error loss — a shrinkage estimator pulling each toward a common reference achieves strictly lower joint error for every true parameter vector, however unrelated the quantities.
- Stepwise regression — An automated regression-model selection procedure that iteratively adds, removes or exchanges predictors according to a prespecified statistical criterion.
- Student's t-Test — A family of mean-inference procedures that divides an observed mean or mean difference by its estimated standard error and evaluates the resulting statistic against a Student t distribution whose degrees of freedom account for estimating variance from the sample.
- Studentization — Dividing a sample statistic by a sample-based estimate of its standard deviation.
- Studentized Range — The range across several normal quantities divided by an independent estimate of their common standard deviation, producing a q-distribution whose group-count and degrees-of-freedom quantiles support simultaneous pairwise mean comparisons.
- Suppressor variable — A predictor whose inclusion improves another predictor's criterion-relevant signal by accounting for variance that is irrelevant, oppositely signed, or otherwise obscuring in the reduced model.
- Surrogate Endpoint Problem — The failure that arises when a trial's biomarker surrogate diverges from the clinical endpoint it stands in for — because the intervention acts through off-pathway mechanisms the surrogate cannot see — so individual-level correlation does not license an intervention-level claim.
- Two-way analysis of variance — An analysis-of-variance model estimating two categorical factors’ main effects and their interaction on a continuous response.
- Type M Error — Quantify how much a significant effect's reported magnitude is exaggerated by the significance filter under low power, via the exaggeration ratio — the expected significant estimate divided by the true effect — computable from the design before any data exist.
- Type S Error — Quantify the risk that a statistically significant estimate points the wrong way by computing, before data collection, the probability that a two-sided significance filter is cleared from the opposite tail when the true effect is near zero relative to noise.
- Underfitting — The failure mode where a model's hypothesis class is too restrictive to capture the structure genuinely present in the data — high bias, with training and test error both elevated and close together — curable only by a richer functional form, not by more data or regularization.
- Unevenly spaced time series — A time series represented by observation-time and value pairs whose successive observation intervals are not constant.
- Universal Hypothesis Testing — A goodness-of-fit testing problem that compares one fully specified null distribution with the unrestricted alternative of every other distribution, seeking level-controlled tests that remain consistent or error-exponent optimal without modeling a particular alternative.
- Variance — The expected squared deviation of a random variable from its mean, measuring dispersion in squared units and equaling its second central moment.
- Variance function — A smooth function expressing the conditional variance of a random quantity as a function of its mean.
- Variational Bayesian Methods — Bayesian inference methods that choose a tractable distribution from a declared family by optimizing an evidence bound or divergence to approximate an intractable posterior.
- Variational Message Passing — Compile mean-field variational Bayes updates into local exchanges of moments and natural-parameter contributions on a probabilistic graph, iteratively increasing an evidence lower bound without claiming exact posterior recovery.
- Variogram — A geostatistical lag function measuring expected squared differences between field values, used to model spatial dependence, anisotropy, nugget, range, and kriging weights.
- Violin Plot — A statistical distribution graphic that mirrors a kernel-density estimate around an axis, usually combining shape with median, quartiles, box-plot summaries, or raw observations.
- Yule–Simon Distribution — A one-parameter distribution on positive integers with beta-function probability mass and a power-law tail, associated with cumulative-advantage frequency models.
- Z-test — A hypothesis test whose null statistic follows, exactly or approximately, a standard normal distribution after centering and scaling by a known or consistently estimated standard error.