Descriptive Statistics & Correlation Measures¶
← Back to Domain-Specific Families
Abstractions about quantifying spread, association and distributional shape in data, spanning dispersion measures (variance, coefficient of variation, geometric standard deviation), correlation coefficients (Pearson, Kendall, polychoric), hypothesis tests (F-test, Z-test, sign test), and summary statistics like quartiles.
39 abstractions in this family — domain-specific abstractions that sit near one another in structural-signature space (k-means over structural-signature embeddings). Each is shown with its short description.
- 68–95–99.7 rule — The normal-distribution rule that about 68%, 95%, and 99.7% of probability lies within one, two, and three standard deviations of the mean.
- Anderson–Darling test — Test a sample’s agreement with a specified continuous distribution by integrating squared empirical-CDF deviations with extra weight in the tails.
- Asymptotic theory (statistics) — The large-sample framework that studies limiting distributions, consistency and efficiency of estimators and tests as sample size tends to infinity.
- Balanced repeated replication — A replicate-weight variance estimator for complex surveys that repeatedly selects one primary sampling unit from each paired stratum according to a balanced sign matrix.
- Barnardisation — A statistical-disclosure-control method that pseudo-randomly perturbs nonzero interior table counts by plus one, zero or minus one according to a fixed probability rule before recomputing totals.
- Biweight midcorrelation — A robust correlation measure that centers each variable at its median and downweights observations far from the median using a redescending biweight.
- Chauvenet's criterion — Flag a single extreme observation when, under a fitted normal-error model, the expected number of sample observations at least as far from the mean is below one half.
- Coefficient of variation — A dimensionless relative-dispersion statistic equal to standard deviation divided by mean, interpreted only where the measurement scale and nonzero mean make the ratio meaningful.
- Concentration parameter — A distribution-family parameter controlling how tightly probability mass clusters around a direction, center or base distribution without necessarily changing that center.
- Correlation ratio — An effect-size measure equal to the square root of between-category variance divided by total variance, detecting nonlinear mean association.
- Correspondence analysis — A dimension-reduction and visualization method for contingency tables using chi-square geometry to jointly map row and column profiles.
- Empirical Bayes method — Estimate a shared prior distribution or its hyperparameters from the same ensemble of observations and then perform Bayesian-style shrinkage or posterior inference conditional on that estimate.
- F-test of equality of variances — A parametric hypothesis test that compares two independent normal-population variances using the ratio of their sample variances.
- Five-number summary — Compress a univariate ordered dataset into minimum, first quartile, median, third quartile, and maximum, exposing center, spread, skew, and extremes while remaining dependent on the chosen quartile convention.
- Generalized entropy index — A parameterized family of decomposable inequality measures computed from powers or logarithms of each observation's ratio to the population mean.
- Geometric standard deviation — A dimensionless multiplicative spread factor obtained by exponentiating the standard deviation of logarithms.
- Higher-order statistics — Statistics based on third- or higher-order moments, cumulants or spectra that characterize distributional shape and nonlinear dependence beyond mean and covariance.
- Hoover index — An inequality measure equal to the share of total income or another resource that would need redistribution to achieve equal per-capita shares.
- Inverse probability weighting — An estimation method that weights observed units by the inverse probability of their observed sampling, treatment or response status to reconstruct a target population or intervention distribution.
- Kendall rank correlation coefficient — A rank-association statistic based on the excess of concordant over discordant observation pairs.
- L-estimator — An estimator formed as a linear combination of sample order statistics.
- Mean integrated squared error — The expected integrated squared difference between a functional estimator and its unknown target, commonly used as global density-estimation risk.
- Medcouple — A robust median-based statistic of univariate skewness formed from paired observations on opposite sides of the sample median.
- Median — A central cut of ordered data or a distribution with at least half the observations or probability at or below it and at least half at or above it.
- Midhinge — Summarize distributional location by averaging the first and third quartiles, placing the center halfway between the hinges while keeping the interquartile spread analytically separate.
- Nemenyi test — A rank-based post-hoc multiple-comparison procedure that identifies pairs of treatments whose average ranks differ beyond a familywise-error-controlled critical distance after repeated-block comparison.
- Normal probability plot — A quantile plot comparing ordered observations with expected normal quantiles so approximate normality appears linear and systematic departures reveal skew, tails, mixtures or outliers.
- Outliers ratio — A legacy objective-video-quality metric reporting the fraction of model predictions lying outside a declared tolerance interval around subjective mean-opinion scores.
- Pearson correlation coefficient — The unitless covariance of two variables divided by the product of their standard deviations, measuring linear association from minus one to one.
- Point-biserial correlation coefficient — The Pearson correlation between one continuous variable and a genuinely dichotomous variable, expressible through group means, proportions and overall standard deviation.
- Polychoric correlation — An estimate of the correlation between two latent normally distributed continuous variables inferred from their observed ordinal categories through threshold models.
- Quartile — One of the three cut points corresponding approximately to the 25th, 50th and 75th percentiles, dividing ordered data or a distribution into four equal-probability parts.
- Sign test — Test a paired-difference or one-sample median null by reducing non-tied observations to positive and negative signs and evaluating the positive count against its exact binomial distribution under a declared null probability, usually one half.
- Standard score — A normalized value equal to an observation’s deviation from a reference mean divided by the reference standard deviation.
- Statistical thinking — A mode of reasoning that treats outcomes as products of interconnected processes containing variation and uses data, uncertainty and context to guide learning and improvement.
- Studentization — Dividing a sample statistic by a sample-based estimate of its standard deviation.
- Two-way analysis of variance — An analysis-of-variance model estimating two categorical factors’ main effects and their interaction on a continuous response.
- Variance — The expected squared deviation of a random variable from its mean, measuring dispersion in squared units and equaling its second central moment.
- Z-test — A hypothesis test whose null statistic follows, exactly or approximately, a standard normal distribution after centering and scaling by a known or consistently estimated standard error.