Skip to content

Statistical Model

Represent possible observable data by a declared sample space and family of candidate probability laws—often indexed by parameters and assumptions—so estimation, testing, prediction, and uncertainty statements have an explicit conditional basis.

Version
v3 · 2026-09-06 · History
Domain-specific #
2849
Origin domain
statistics
Subdomain
foundations of statistical inference

Core Idea

A statistical model is an explicit family of probability laws proposed for observable data on a declared sample space. In compact notation, a model may be written as \(\mathcal P={P_\theta:\theta\in\Theta}\), where each \(P_\theta\) is a probability distribution on the same measurable observation space, \(\Theta\) is a parameter space, and the map \(\theta\mapsto P_\theta\) says how parameters determine candidate laws. McCullagh describes the conventional foundation as a set of probability distributions on a sample space and distinguishes a parameterized model as a parameter set together with a map into that space of distributions.[1] The parameterization is common but not universal: nonparametric models can be infinite-dimensional or described directly as restricted sets of laws.

The family is a disciplined claim about what data could have arisen and with what probabilities. It can encode independence, exchangeability, response distributions, link functions, time dependence, censoring, covariate effects, measurement error, or other restrictions. For dominated families it determines a density- or mass-based likelihood for observed data. Locally dominated or mutually absolutely continuous comparisons can instead use likelihood ratios, while more general statistical models may compare probability measures directly without a universal common density. Each formulation conditions the model-based estimates, intervals, tests, predictions, or posterior calculations built from it. University of Bath notes make the separation explicit: the model specifies sample space, parameter space, and distribution family; the observation is then used to infer a parameter or predict another random variable.[2]

The statistical model is not the truth by definition. Its law family may omit the real data-generating process. Model checking, residual analysis, sensitivity analysis, robust inference, and comparison with alternatives probe that gap. Box's account of scientific statistics treats modeling as an iterative confrontation between theory and practice, with parsimonious models revised in response to inadequacy rather than regarded as literal reality.[3]

Nor is a model identical to one fitted result. Fitting selects or weighs particular members of the candidate family—perhaps \(\widehat\theta\), a posterior over \(\theta\), or a predictive mixture—but sampling and parameter uncertainty are meaningful only because the larger family remains in view. A displayed regression line without a response distribution and dependence assumptions is a systematic component, not yet a complete statistical model for observable outcomes.

The abstraction is domain-specific. Sample spaces, probability measures, distribution families, likelihoods, parameters, identifiability, sampling assumptions, and model criticism transfer literally across statistics, but they do not travel intact to ordinary conceptual, mechanical, or visual models. Their portable skeleton is Representation; the probability-law family is the statistical accent.

Structural Signature

Sig role-phrases:

  • the observational domain — statistical units, covariates, design, time points, or other conditions specifying what observations the model addresses
  • the measurable sample space — the possible data records or response configurations and the events to which probabilities are assigned
  • the candidate law family — a nonempty set of probability distributions admitted as possible observable-data laws
  • the parameterization or structural index — a parameter space and map into the law family when the model is indexed, including finite- or infinite-dimensional structure
  • the modeling assumptions — restrictions such as independence, dependence form, distribution family, link, stationarity, exchangeability, censoring, or measurement process
  • the observed data (use-interface role) — a realized point in the sample space against which members of the family may be evaluated, fitted, checked, or compared; it is not required for the law family to exist
  • the inferential readouts (use-interface role) — likelihood ratios, estimates, tests, intervals, predictions, posterior quantities, or decisions that can be derived conditionally from the model but are not mandatory components of it
  • the adequacy and identifiability diagnostics (use-interface role) — checks for mismatch between family and data and for parameter distinctions that fail to produce distinct observable laws; these assess a model rather than constitute its declared family

Three layers must be kept separate. The model image is the set of admitted distributions. A parameterization is a coordinate system or map used to describe that set. A fitted model is a selected member, estimate, or posterior-weighted construction based on data. Two parameterizations can generate the same model image; one parameterization can be nonidentifiable if different parameter points yield the same \(P_\theta\); and a fitted member does not replace the family from which its uncertainty was derived.[1] A declared law family is already a statistical model before any observation arrives, inferential procedure is selected, or adequacy diagnostic is run; those are interfaces for using and assessing it.

The recognition test is probabilistic completeness relative to the declared observations: can the proposed object assign a joint probability law to the modeled data configurations under every admitted parameter or structural state? A conditional mean equation alone does not answer this unless variance, dependence, or a suitable semiparametric estimating framework supplies the remaining inferential contract. Conversely, a singleton family is formally possible, though it leaves no unknown law within the model to infer.

What It Is Not

  • Not the observed dataset. Data are a realized point in the sample space. They can support, contradict, fit, or select a model but are not the family of laws.
  • Not one probability distribution by default. A distribution is one candidate law; a statistical model is normally a set or indexed family of such laws. A fitted \(P_{\widehat\theta}\) is one member selected from that family.
  • Not merely a regression equation. \(E(Y\mid X)=X\beta\) states a mean structure. A complete normal linear model also specifies an error law, variance/covariance, and conditional or design assumptions; other frameworks must declare what replaces them.
  • Not a fitting or inference algorithm. Maximum likelihood, least squares, MCMC, variational inference, and bootstrap procedures operate on or approximate consequences of a model. Their update rules are not the model's law family.
  • Not a scientific theory or causal mechanism automatically. A model may be descriptive, predictive, mechanistic, or causal. A conditional association family does not become causal merely because its coefficients are interpretable.
  • Not necessarily parametric. Fixed finite-dimensional \(\Theta\) defines a parametric model; semiparametric and nonparametric models retain law-family restrictions without that finite coordinate form.
  • Not necessarily identifiable or correct. These are properties to investigate. Distinct parameters may induce the same observable distribution, and the true law may lie outside the family.
  • Not necessarily Bayesian. A prior over parameters extends a sampling model into a Bayesian model. Frequentist and likelihood-based models require no prior, while Bayesian models may add hierarchical layers and latent variables.
  • Not a likelihood function alone. In a dominated family, likelihood is a chosen density or mass evaluated at observed data and viewed as a function of parameters; likelihood-ratio or local-domination formulations cover some families without one global reference measure, and fully general models may require direct measure-based comparison. Any of these evidential surfaces can omit information needed to reconstruct the entire admitted law family.

Scope of Application

Statistical Model is the common formal object across statistical inference wherever observable uncertainty is represented by a restricted family of laws.

Classical parametric sampling. Bernoulli, binomial, Poisson, normal, gamma, and other families describe repeated outcomes, counts, measurements, waiting times, or lifetimes through finite-dimensional parameters. The model includes the sample structure and assumptions such as independence and identical distribution, not only the named marginal distribution.

Regression and generalized linear models. A design or covariate domain enters a conditional response family. Linear predictors, link functions, response distributions, variance functions, and dependence assumptions jointly define what outcomes are admitted. McCullagh uses the normal linear and generalized-linear cases to show that a mean representation alone must be connected to a family of distributions on the sample space.[1]

Time series and stochastic processes. ARMA, state-space, point-process, survival, and spatial models define joint laws over ordered or indexed observations. Marginal distributions alone are insufficient because temporal, spatial, or event-history dependence is load-bearing.

Hierarchical and Bayesian models. Sampling distributions, latent-variable laws, and priors form a joint model whose conditioning yields posterior and predictive distributions. The general identity remains a family or joint law on declared variables; prior choice is an added Bayesian component rather than a universal requirement.

Nonparametric and semiparametric inference. Smoothness classes, shape constraints, moment restrictions, proportional-hazards structure, and infinite-dimensional density families demonstrate that “model” is broader than a formula with a few coefficients. Van der Vaart's asymptotic framework treats estimators and tests over parameterized experiments whose complexity may grow or be infinite-dimensional.[4]

Survey, missing-data, measurement, and causal analyses. Sampling designs, response mechanisms, error processes, assignment mechanisms, and potential-outcome restrictions must be modeled when inferential claims depend on them. A causal estimand additionally needs identification assumptions; a statistical outcome model alone does not supply causality.

Machine learning with probabilistic semantics. Probabilistic classifiers, generative models, Gaussian processes, and calibrated predictive distributions instantiate the node when they define candidate laws over observable data. A deterministic scoring rule without a declared probabilistic interpretation is a predictive function, not automatically a statistical model.

The scope ends where probability-law semantics disappear. A scale architectural model, a business canvas, or a conceptual diagram is a representation and may be called a model, but it is not a statistical model unless it supplies a family of probability distributions for defined observations.

Clarity

The abstraction turns the vague phrase “we assumed a model” into a checklist. What is the observational unit? What outcomes are possible? Which joint distributions are admitted? Which restrictions encode dependence or design? Which parameters index distinct laws? Which readout is conditional on those commitments? This makes hidden assumptions inspectable rather than letting a fitted curve stand in for a generative specification.

It also separates sampling uncertainty from model uncertainty and misspecification. A confidence interval can be exactly calibrated inside \(\mathcal P\) while the actual process lies outside \(\mathcal P\). More observations shrink parameter uncertainty only when the sampling process accumulates relevant information and the model satisfies suitable identifiability and regularity conditions; even then, they need not repair a missing dependence, wrong tail, omitted mixture, or invalid sampling assumption. The remedy for the latter is criticism, expansion, replacement, or robustification of the family.

The family-versus-fit distinction clarifies apparently contradictory statements. “The model predicts 0.7” may refer to a plug-in fitted member. “The model permits probabilities from 0 to 1” refers to the family. “The model is uncertain” may mean parameter posterior spread, selection among families, or residual stochasticity. Naming the layer prevents these uncertainties from being combined casually.

A practical diagnostic is to demand a generative reading: under a proposed parameter value and design, how could a complete dataset be simulated? If no probability law or equivalent stochastic specification exists, the object may be a useful score or structural equation but its statistical-model status is incomplete.

Manages Complexity

The space of all probability laws on a realistic sample space is too large to infer from finite data without restriction. A statistical model manages this underdetermination by declaring a smaller family. A two-parameter normal family replaces an arbitrary continuous distribution with mean and variance; a regression family replaces unrelated response laws at every covariate value with a shared predictor; a Markov model replaces an unrestricted path law with local transition structure.

This compression creates leverage. Likelihoods can rank candidate parameter points; sufficient statistics may summarize data; standard errors and tests can be derived; predictions can integrate parameter and observation uncertainty; simulations can expose consequences. Parameterization makes systematic variation navigable, while independence or conditional-independence structure can turn an intractable joint law into products of local terms.

The same compression creates fragility. Strong restrictions improve efficiency when approximately correct but concentrate error when wrong. A narrow family may yield precise, reproducible answers to the wrong question. A flexible family can absorb more shapes but require more data, weaken identifiability, and complicate computation. Model management is therefore not a one-time selection but a loop: specify, fit, check, compare, revise, and report what remains conditional.[3]

Failure can be localized by role. Poor residual shape implicates the response family. Patterned residual dependence implicates independence assumptions. Unstable or flat likelihood directions implicate identifiability. Good in-sample fit with poor prediction implicates complexity or distribution shift. Sensitivity to priors implicates weak data information relative to parameterization. The structural signature tells the analyst which part of the family to change.

Abstract Reasoning

The model-family view licenses several predictions.

Restriction prediction. If two models admit different law families, an estimator or test valid under the smaller family need not retain its efficiency or calibration under the larger one. Greater precision usually comes from stronger commitments.

Identifiability prediction. If \(P_{\theta_1}=P_{\theta_2}\) for distinct parameter points, no amount of observable data generated under the model can distinguish them by likelihood alone. Repair requires reparameterization, added measurements, stronger restrictions, or accepting a set-valued target.

Sample-size prediction. When observations contribute accumulating information and the inferential procedure and model meet appropriate regularity conditions, increasing sample size can concentrate inference near a closest admitted law. If the true law is outside the family, however, that concentration can make a misspecified conclusion more confident rather than more correct; without those information, procedure, and regularity conditions, concentration is not guaranteed at all. Check adequacy separately from precision.

Parameterization prediction. Smooth one-to-one reparameterization can change numerical coordinates, priors, and optimization behavior without changing the model image. Claims about the model should be invariant where possible; claims about a particular coefficient need a parameter meaning and identifiability argument.

Design prediction. Changing which units, covariates, or time points are observed can change the sample space and what predictions are licensed. McCullagh emphasizes that prediction requires an extension whose domain includes future units or time points; a family defined only for the observed array may not support the proposed extrapolation.[1]

Model-comparison prediction. Goodness of fit is relative: a flexible family cannot be judged only by training fit because it can mimic noise, and two families may make indistinguishable predictions in the observed region while diverging under intervention or extrapolation. Compare using the readout actually needed and probe the region where their laws differ.

These predictions license interventions: expand a tail family, add dependence, collect identifying variables, simplify redundant parameters, use robust procedures, average over plausible models, or narrow the domain of the claim. Each repair modifies the declared law family rather than treating every problem as an optimization failure.

Knowledge Transfer

Within statistics, the sample-space-plus-law-family identity transfers intact across parametric, semiparametric, nonparametric, Bayesian, frequentist, regression, survival, time-series, spatial, and causal work. Analysts can translate roles even when notation changes: observable domain, admitted joint laws, parameter or structural index, assumptions, realized data, inferential target, and adequacy checks.

The main transfer hazard is mistaking a familiar family for a universal default. Independent Gaussian errors, proportional hazards, stationarity, or exchangeability may be natural in one setting and indefensible in another. The structural form transfers; the substantive restrictions must be re-earned.

Knowledge about identifiability, model checking, robustness, and conditional claims travels especially well. A flat direction in a mixture model and collinearity in a regression are both failures of observable parameter distinction. Residual autocorrelation and clustered survey dependence both reveal a missing joint-law structure. Posterior sensitivity and weak frequentist identification both signal that the data do not sharply distinguish admitted laws.

Outside statistics, the word model transfers but this mechanism often does not. A physical simulation becomes a statistical model only when combined with a stochastic observation/error law or an ensemble of candidate probabilistic outputs. A mental model or process diagram remains a Representation. The cross-domain skeleton belongs to that prime; the family-of-probability-laws contract stays home.

Examples

Canonical

For \(n\) coin tosses, let the observable data be the ordered binary sequence \(Y=(Y_1,\ldots,Y_n)\in\{0,1\}^n\). The model assumes conditional independence and a common success probability \(p\in[0,1]\):

\[ P_p(Y=y)=p^{\sum y_i}(1-p)^{n-\sum y_i}. \]

The sample space contains all \(2^n\) sequences; the candidate family is \({P_p:p\in[0,1]}\); and \(p\) parameterizes distinct laws when (n>0). Observing \(k=\sum Y_i\) successes yields likelihood proportional to \(p^k(1-p)^{n-k}\), from which one may estimate, test, or predict. The assumption of identical independent tosses is part of the model. A sequence with learning, drift, or dependence may lie outside it even though each observation remains binary.[2]

Mapped back: toss indices define the observational domain; \({0,1}^n\) is the sample space; the Bernoulli product laws form the candidate family; \(p\) is the parameterization; identical independence is the modeling assumption; the observed sequence is the data; likelihood and prediction are inferential readouts; and run patterns or time drift activate adequacy checks.

Applied / In Practice

A medical study models binary 30-day readmission \(Y_i\) conditional on recorded covariates \(x_i\) using

\[ P(Y_i=1\mid x_i,\beta)=\operatorname{logit}^{-1}(x_i^{\mathsf T}\beta), \]

with conditional independence across patients given the design. The sample space is the vector of binary outcomes for the enrolled patients; the law family varies with \(\beta\); the logit link and additive linear predictor restrict how covariates change risk. Fitting chooses \(\widehat\beta\), but the model remains the whole family needed for standard errors and predictive uncertainty. If patients are clustered by hospital, the independence assumption may understate uncertainty; if an omitted severity measure affects both covariates and readmission, coefficients need not be causal. Remedies modify the family or claim—random effects, clustered covariance, additional variables, or a narrower predictive interpretation—not merely the optimizer.

Mapped back: enrolled patients and covariates form the observational domain/design; binary outcome vectors are the sample space; logistic conditional laws are the family; \(\beta\) is the index; link, linearity, and conditional independence are assumptions; outcomes are observations; fitted risks and intervals are readouts; clustering and residual calibration are adequacy diagnostics.

Structural Tensions

T1: Parsimony versus adequacy. A small family is interpretable and estimable; a richer family can represent more real variation. Diagnostic: which observed or decision-relevant patterns are ruled out by the simplifying restrictions, and have those exclusions been checked?

T2: Efficiency versus robustness. Strong distributional assumptions can yield narrow intervals when correct and severe distortion when wrong. Diagnostic: how do conclusions change under a wider family or robust procedure?

T3: Identifiability versus flexibility. Latent components, nonlinearities, and interactions can improve fit while creating observationally equivalent parameter settings. Diagnostic: do distinct parameter values induce detectably different laws on the actual design?

T4: Interpretability versus predictive performance. Simple coefficients support explanation; complex ensembles or latent models may predict better. Diagnostic: is the required readout a stable causal/structural interpretation or accurate prediction within a bounded domain?

T5: Conditional certainty versus model uncertainty. Standard intervals condition on a selected family, while several families may remain plausible. Diagnostic: does the reported uncertainty include only parameter variation inside one model or also selection and misspecification uncertainty?

T6: Fit versus extrapolation. Multiple models can fit observed data similarly and diverge for future units, extreme covariates, or interventions. Diagnostic: does the model's declared domain and structural extension actually contain the target of prediction?

T7: Bayesian completeness versus prior sensitivity. A prior produces a joint model and coherent posterior, but weak data can make conclusions depend heavily on prior structure. Diagnostic: which posterior features are data-driven, and which move materially under defensible alternative priors?

T8: Stable family versus iterative criticism. Reproducible inference benefits from a fixed model, while scientific learning requires revision after mismatch. Diagnostic: are checks being used to improve future specification, or to repeatedly tune the current analysis without accounting for adaptation?

T9: Autonomy versus reduction. Statistical Model is a Representation composed of Probability Distributions, yet it owns the law-family, parameterization, likelihood, identifiability, and misspecification closure. Diagnostic: are candidate observable-data laws and model-conditional inference load-bearing? If not, resolve to Representation or another domain model rather than exporting this node.

Structural–Framed Character

Statistical Model is mixed-structural. Its formal object—a sample space and family of probability measures—is evaluatively neutral and mathematically precise. Once an observational domain and law family are declared, likelihood, identifiability, and many sampling consequences are structural rather than conventional.

It is nevertheless human-practice-bound in the instrument sense. Investigators decide what counts as an observation, which units enter, which restrictions are plausible, which parameters matter, and which mismatch is tolerable. Its institutional origin lies in mathematical statistics and scientific modeling. Its vocabulary travels literally across statistical disciplines but becomes imported metaphor outside probabilistic inference.

Cross-subfield recurrence is recognition: a survival analyst and a time-series analyst both specify joint law families, parameters, data, and conditional readouts. Cross-domain recurrence is usually only the broader Representation skeleton. Its character: a formally structural family-of-laws object whose construction, use, and adequacy remain framed by statistical design and scientific purpose.

Structural Core vs. Domain Accent

What is skeletal. A simplified representation restricts possible states of a target so reasoning becomes tractable. Parameters or structural choices index alternatives, and observed consequences can discriminate among them. This belongs to prime:representation and broader abstractions about assumption and formalization.

What is domain-bound. The represented alternatives are probability laws on a measurable sample space. Parameters map to distributions; observations produce likelihoods; dependence and sampling assumptions define joint laws; identifiability concerns observational equivalence; and model checking probes whether the true process lies outside the family. These are irreducibly statistical roles.

Why it does not clear the prime bar. A scale model, causal diagram, legal model, and mental model do not automatically assign probabilities to all modeled observations or license likelihood-based inference. Their modelhood is covered by Representation. Exporting this node would either drop its defining law-family contract or incorrectly require every model to be probabilistic.

Instantiates prime:representation. A statistical model stands in for possible observable-data-generating processes by retaining a restricted family of laws and suppressing irrelevant or unknown detail. It strictly specializes Representation with probability and inference machinery. This is a candidate direct strict-subsumption edge.

Contains domain_specific:probability_distribution as constitutive structure. Each admitted member is a probability distribution; the family or parameter map organizes those members into the model. One distribution does not supply the alternative-law and parameter uncertainty structure of the whole. This is a candidate composition/part_of edge.

Supports but is not subsumed by prime:statistical_inference. Inference is reasoning from observed data to parameters, populations, or predictions. It ordinarily presupposes a statistical model or a weaker design-based probability structure, whereas a model can be specified before inference occurs. Keep this as a strong downstream relation unless the structured DAG supports the inverse dependency direction.

prime:probability is declined as a direct edge because the live Probability Distribution node already decomposes to it. prime:formalization is related but too broad and redundant with Representation to improve the minimal parent set.

Relationships to Other Abstractions

Local relationship map for Statistical ModelParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Statistical ModelDOMAINDomain-specific abstraction: Probability Distribution — is part ofProbabilityDistributionDOMAINPrime abstraction: Representation — is a kind ofRepresentationPRIMEDomain-specific abstraction: Factor Analysis — is a kind ofFactor AnalysisDOMAINDomain-specific abstraction: Probabilistic Graphical Model — is a kind ofProbabilisticGraphical ModelDOMAIN

Current abstraction Statistical Model Domain-specific

Parents (2) — more general patterns this builds on

  • Statistical Model is a kind of Representation Prime

    Instantiates prime:representation. A statistical model stands in for possible observable-data-generating processes by retaining a restricted family of laws and suppressing irrelevant or unknown detail.

  • Statistical Model is part of Probability Distribution Domain-specific

    Instantiates prime:representation. A statistical model stands in for possible observable-data-generating processes by retaining a restricted family of laws and suppressing irrelevant or unknown detail.

Children (2) — more specific cases that build on this

  • Factor Analysis Domain-specific is a kind of Statistical Model

    The proposed parent is Statistical Model: factor analysis declares a family of probability/covariance structures indexed by loadings, factor covariances, and residual parameters, then supports estimation and fit assessment.

  • Probabilistic Graphical Model Domain-specific is a kind of Statistical Model

    The candidate is a strict specialization of domain_specific:statistical_model: it declares possible data variables and a family of joint laws, with graph structure restricting that family.

Hierarchy paths (6) — routes to 4 parentless roots

Neighborhood in Abstraction Space

Statistical Model sits in a sparse region of the domain-specific corpus (69th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (1565 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-09-08

Not to Be Confused With

  • Probability distribution. One complete law on an outcome space. Tell: is the object one \(P\), or a family/map of candidate laws supporting unknown structure?
  • Fitted model. A selected member, parameter estimate, or trained object. Tell: are uncertainty and alternatives still represented, or only the plug-in result?
  • Dataset. Realized observations. Tell: is the object a point in the sample space, or the law family under which points receive probabilities?
  • Regression equation. A systematic mean or predictor relation. Tell: are response distribution, variance, and dependence specified sufficiently for the claimed inference?
  • Statistical inference. The reasoning operation using data and probability structure. Tell: is the object the admitted family, or an estimate/test/prediction derived from it?
  • Likelihood. Model probabilities or densities evaluated at the observed data as a function of parameters. Tell: can the full data law be reconstructed, or only a proportional evidence surface?
  • Algorithm. A procedure for fitting, sampling, or predicting. Tell: is the concern which laws are possible, or how a computer calculates with them?
  • Scientific or causal model. A mechanism or structural explanation. Tell: does it specify a probability family for observables and the assumptions needed for the claimed statistical or causal readout?
  • Parametric model. A finite-dimensional subtype. Tell: is finite-dimensional \(\Theta\) constitutive, or does the candidate family include infinite-dimensional possibilities?
  • Bayesian model. A statistical model extended with priors and often latent/hierarchical laws. Tell: is a prior over unknowns required, or only the sampling-law family?

References

[1] McCullagh, P. “What Is a Statistical Model?” Annals of Statistics 30(5), 2002, 1225–1310. Defines conventional and parameterized models and develops design, domain-extension, and parameter-meaning requirements. registry ↩a ↩b ↩c ↩d

[2] University of Bath, APTS. “Statistical models,” Introductory Notes on Statistical Inference. Defines sample space, parameter space, distribution family, observations, estimators, and inference. registry ↩a ↩b

[3] Box, G. E. P. “Science and Statistics.” Journal of the American Statistical Association 71(356), 1976, 791–799. Supports parsimonious modeling and iterative confrontation of theory with practice and inadequacy. registry ↩a ↩b

[4] van der Vaart, A. W. Asymptotic Statistics. Cambridge University Press, 1998. Develops model-indexed asymptotic inference across parametric and broader statistical settings. registry