Anderson–Darling test¶
Test a sample’s agreement with a specified continuous distribution by integrating squared empirical-CDF deviations with extra weight in the tails.
Core Idea¶
The Anderson–Darling statistic is a Cramér–von Mises-type goodness-of-fit measure weighted by [F(x)(1−F(x))]^{-1}, emphasizing tail departures.[1] Data are transformed through the null CDF, ordered, and compared with uniform order-statistic expectations; logarithmic tail terms accumulate discrepancies into A². The abstraction is therefore identified by a declared carrier, a transformation or constraint over that carrier, and an invariant that tells an analyst whether the named structure is genuinely present.
The load-bearing residual is not the broad topic of statistics. It is the Anderson–Darling tail-weighted statistic and distribution-specific calibration. That residual remains recognizable when examples, notation, scale, or implementation change, but it disappears if Kolmogorov–Smirnov critical values are reused, fitted parameters are ignored, discrete data use continuous tables without adjustment, or a p-value is treated as fit magnitude. This gives the entry an operational identity rather than merely a historical label.
A useful analysis keeps three layers separate. The constitutive layer says what must be true: the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test. The evidential layer asks what observation or proof warrants the claim: state whether parameters are known or estimated, use the correct finite-sample calibration, handle ties and censoring, inspect effect size and plots, and avoid accepting the null from nonsignificance. The use layer asks what reasoning becomes available once the identity is established: testing distributional fit, detecting tail deviations, comparing samples, and validating model assumptions. Conflating the layers is the most common source of scope inflation.
Structural Signature¶
- Carrier: an ordered sample, a fully specified or fitted continuous null cumulative distribution, and the transformed empirical distribution
- Inputs or antecedent state: sample size, observations, null family and parameters, estimation method, Anderson–Darling statistic, critical-value calibration, ties, censoring, and significance rule
- Constitutive operation: Data are transformed through the null CDF, ordered, and compared with uniform order-statistic expectations; logarithmic tail terms accumulate discrepancies into A².
- Invariant: the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test
- Recognition test: state whether parameters are known or estimated, use the correct finite-sample calibration, handle ties and censoring, inspect effect size and plots, and avoid accepting the null from nonsignificance
- Output or consequence: testing distributional fit, detecting tail deviations, comparing samples, and validating model assumptions
- Failure boundary: Kolmogorov–Smirnov critical values are reused, fitted parameters are ignored, discrete data use continuous tables without adjustment, or a p-value is treated as fit magnitude
What It Is Not¶
- It is not the whole field of statistics. The field contains many questions and methods that do not instantiate Anderson–Darling test.
- It is not its most familiar example. Testing a fully specified continuous F maps observations to ordered uniform values and computes A² from their log tail probabilities. exhibits the structure, but the example is evidence for the abstraction rather than its definition.
- It is not the neighboring catalog concept Hypothesis Testing. Hypothesis testing is the broad Prime; Anderson–Darling fixes an empirical-CDF statistic, tail weight, and calibration regime.
- It is not a claim that every boundary case has one uncontested classification. a qualified variant may preserve the core while changing notation, parameterization, or implementation, so the constitutive condition must decide the boundary
- It is not an unrestricted metaphor for any process that seems similar. Outside statistics, the vocabulary and validity conditions do not transfer literally.
Scope of Application¶
Anderson–Darling test belongs to statistics and is useful where the analyst can specify an ordered sample, a fully specified or fitted continuous null cumulative distribution, and the transformed empirical distribution, then evaluate the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test. The scope is broad within that domain but bounded by the need for the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test. The entry records a descriptive analytical identity; practical use requires the governing domain's evidence, standards, and safety obligations.[2]
- Definition and recognition. Determine whether a proposed instance satisfies the constitutive conditions rather than merely sharing terminology.
- Construction or evolution. Track how sample size, observations, null family and parameters, estimation method, Anderson–Darling statistic, critical-value calibration, ties, censoring, and significance rule are converted, constrained, or organized by Data are transformed through the null CDF, ordered, and compared with uniform order-statistic expectations; logarithmic tail terms accumulate discrepancies into A²..
- Comparison. Compare instances using carrier, defining parameters, convention, scale, scope, evidence, limiting cases, and implementation, without treating convenience measures as the definition.
- Boundary analysis. Diagnose cases where a qualified variant may preserve the core while changing notation, parameterization, or implementation, so the constitutive condition must decide the boundary and state which convention or theorem controls the decision.
- Downstream reasoning. Use the established identity to support testing distributional fit, detecting tail deviations, comparing samples, and validating model assumptions while preserving the assumptions under which the inference is valid.
Clarity¶
The abstraction clarifies a crowded vocabulary by making the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test the center of the account. A claim should name the carrier, the governing operation or relation, the applicable assumptions, and the recognition test. A bare label is insufficient because the name Anderson–Darling test can be used for a formal identity, an implementation, or a neighboring result unless carrier and convention are stated. The disciplined statement is: given sample size, observations, null family and parameters, estimation method, Anderson–Darling statistic, critical-value calibration, ties, censoring, and significance rule, the structure counts as Anderson–Darling test exactly when the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test.
This format also separates identity from measurement. Empirical, computational, or documentary proxies support recognition only under declared validity and uncertainty assumptions; formal cases require proof rather than measurement. Measurements can be noisy, implementations can approximate, and proofs can use equivalent characterizations; none of those facts licenses changing the object being measured. When reports disagree, first check scope and convention, then data or proof, and only then interpret the disagreement as substantive.
Manages Complexity¶
Without the abstraction, an analyst must reason directly over many local details: the carrier roles, admissibility assumptions, competing conventions, derived invariants, boundary cases, and proof or validation obligations specific to Anderson–Darling test. Anderson–Darling test compresses them into the roles in the structural signature. That compression permits comparison across instances without erasing the variables that determine validity. It also exposes which details may be varied safely and which are constitutive.
The compression has a price. A single label can hide standard, generalized, restricted, approximate, computational, and historically variant formulations of Anderson–Darling test. Good use therefore carries a small declaration of assumptions alongside the name. The abstraction manages complexity when it reduces the state space of the question while keeping the failure boundary visible; it mismanages complexity when the label substitutes for that boundary analysis.
Abstract Reasoning¶
- Identify the carrier. State what the elements, states, objects, or observations are: an ordered sample, a fully specified or fitted continuous null cumulative distribution, and the transformed empirical distribution. Reject examples whose alleged carrier belongs to a different problem.
- Lock the constitutive rule. Express the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test independently of one notation or implementation. This step prevents the canonical example from becoming the definition.
- Derive consequences. From the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test, infer testing distributional fit, detecting tail deviations, comparing samples, and validating model assumptions. Record each assumption used so that a later change of setting does not silently preserve an invalid conclusion.
- Test adversarial cases. Examine a qualified variant may preserve the core while changing notation, parameterization, or implementation, so the constitutive condition must decide the boundary and a Q–Q plot is useful diagnostic evidence but is not the Anderson–Darling test statistic and calibration. A robust identity explains why the first is convention-sensitive and why the second is outside the class.
- Compare and refine. Use carrier, defining parameters, convention, scale, scope, evidence, limiting cases, and implementation to compare legitimate instances, and refine the model when discrepancies reflect hidden variation rather than failure of the abstraction itself.
Knowledge Transfer¶
Knowledge transfers strongly among subfields of statistics because they reuse an ordered sample, a fully specified or fitted continuous null cumulative distribution, and the transformed empirical distribution, Data are transformed through the null CDF, ordered, and compared with uniform order-statistic expectations; logarithmic tail terms accumulate discrepancies into A²., and state whether parameters are known or estimated, use the correct finite-sample calibration, handle ties and censoring, inspect effect size and plots, and avoid accepting the null from nonsignificance. A theorem, diagnostic, or modeling warning can travel when those roles remain literal. For example, the distinction between constitutive identity and a convenient observable transfers from Testing a fully specified continuous F maps observations to ordered uniform values and computes A² from their log tail probabilities. to A normality test estimates mean and variance and uses adjusted critical values rather than the distribution-free simple-null table..[3]
Transfer outside the home domain is weaker. The skeletal pattern—type a carrier, apply a constitutive relation, preserve its invariant, and derive only qualified consequences—may suggest an analogy, but the domain-specific mechanisms, admissible evidence, and consequences do not come along automatically. The safe transfer procedure maps each role explicitly, checks the invariant again, and refuses the name when only a superficial resemblance remains.
Examples¶
Canonical¶
Testing a fully specified continuous F maps observations to ordered uniform values and computes A² from their log tail probabilities. The denominator weighting makes comparable deviations near zero and one contribute more than central deviations. This example is canonical because every role can be inspected: the carrier is an ordered sample, a fully specified or fitted continuous null cumulative distribution, and the transformed empirical distribution; the operative rule is Data are transformed through the null CDF, ordered, and compared with uniform order-statistic expectations; logarithmic tail terms accumulate discrepancies into A².; the invariant is the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test; and the result supports testing distributional fit, detecting tail deviations, comparing samples, and validating model assumptions.[1] Changing incidental notation or scale leaves the structure intact, while removing the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test destroys the classification.
Mapped back: an ordered sample, a fully specified or fitted continuous null cumulative distribution, and the transformed empirical distribution → Data are transformed through the null CDF, ordered, and compared with uniform order-statistic expectations; logarithmic tail terms accumulate discrepancies into A². → the declared tail-weighted empirical-versus-null CDF statistic and matching null calibration determine the test → testing distributional fit, detecting tail deviations, comparing samples, and validating model assumptions
Applied / In Practice¶
A normality test estimates mean and variance and uses adjusted critical values rather than the distribution-free simple-null table. Parameter estimation changes the null distribution of the statistic and is part of the test identity. The applied case is not licensed merely by vocabulary. It qualifies because the same recognition test—state whether parameters are known or estimated, use the correct finite-sample calibration, handle ties and censoring, inspect effect size and plots, and avoid accepting the null from nonsignificance—can be run and because the same failure boundary—Kolmogorov–Smirnov critical values are reused, fitted parameters are ignored, discrete data use continuous tables without adjustment, or a p-value is treated as fit magnitude—remains meaningful.[2] The case also shows why practical outputs should report assumptions, resolution, and uncertainty instead of a naked label.
Mapped back: declared instance → recognition test → boundary check → qualified use
Structural Tensions¶
- T1: Axiomatic identity vs. operational recognition. The defining conditions may be exact while empirical or computational recognition is approximate. Neither pole can be removed without changing the analytical task. Diagnostic: Can the reviewer state both the exact condition and the evidence used to infer it?
- T2: Local roles vs. global consequence. The mechanism is enacted through local relations, but the abstraction is usually valued for a global classification or prediction. Neither pole can be removed without changing the analytical task. Diagnostic: Does the claimed global result actually follow from the declared local conditions?
- T3: Ideal form vs. finite representation. Theory states a clean invariant while data structures, measurements, or proofs expose only finite representations. Neither pole can be removed without changing the analytical task. Diagnostic: Would increasing resolution converge toward the same classification?
- T4: Canonical convention vs. legitimate variants. A standard formulation supports communication, while variants may preserve the same core under changed assumptions. Neither pole can be removed without changing the analytical task. Diagnostic: Which role is invariant across variants, and which convention-specific conclusion changes?
- T5: Compression vs. hidden assumptions. The name compresses a complex argument but can conceal prerequisites. Neither pole can be removed without changing the analytical task. Diagnostic: Can each downstream inference be traced to an explicit assumption?
- T6: Autonomous residual vs. reduction to catalog neighbors. The candidate uses broader structures but adds an identity-bearing residual. Neither pole can be removed without changing the analytical task. Diagnostic: After subtracting the proposed parent and named neighbors, does the constitutive residual still support independent diagnostics?
Structural–Framed Character¶
The entry is structurally mixed but domain-framed. Its portable skeleton is type a carrier, apply a constitutive relation, preserve its invariant, and derive only qualified consequences. Its identity-bearing terms—Anderson–Darling test, carrier, parameter, relation, invariant, boundary, evidence, and application—derive their meaning from statistics and cannot be replaced by generic systems language without losing the tests that distinguish valid from invalid instances.
This mixed character explains why the abstraction is reusable inside the domain yet does not meet the Prime bar. The structure organizes reasoning, but its claims still depend on domain-specific objects, evidence, and intervention semantics.
Structural Core vs. Domain Accent¶
The structural core consists of a carrier, Data are transformed through the null CDF, ordered, and compared with uniform order-statistic expectations; logarithmic tail terms accumulate discrepancies into A²., a recognition invariant, and a consequence. That skeleton may resemble patterns elsewhere, especially type a carrier, apply a constitutive relation, preserve its invariant, and derive only qualified consequences. The domain accent is not decorative: Anderson–Darling test, carrier, parameter, relation, invariant, boundary, evidence, and application determine what counts as an admissible carrier, a valid transition, and successful evidence.
The abstraction therefore remains domain-specific. A cross-domain reuse that preserves only words such as 'balance,' 'cut,' 'sequence,' 'loss,' or 'simulation' is metaphor. Literal transfer requires the original role structure and diagnostics, which in this case remain anchored in statistics.
Instantiates / Related Primes¶
The proposed strict upward parent is prime:hypothesis_testing_null_vs_alternative. The procedure literally compares sample evidence with a null distribution through a rejection rule; its tail-weighted statistic supplies the residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Anderson–Darling test adds domain-specific constraints.
The entry does not collapse into that parent because the Anderson–Darling tail-weighted statistic and distribution-specific calibration It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Anderson–Darling test. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge.
The prospective workspace queue contains one strict upward edge to prime:hypothesis_testing_null_vs_alternative. No live DAG mutation is authorized.
Relationships to Other Abstractions¶
Current abstraction Anderson–Darling test Domain-specific
Parents (1) — more general patterns this builds on
-
Anderson–Darling test is a kind of Hypothesis Testing (Null vs. Alternative) Prime
The proposed strict upward parent is
prime:hypothesis_testing_null_vs_alternative.The procedure literally compares sample evidence with a null distribution through a rejection rule; its tail-weighted statistic supplies the residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Anderson–Darling test adds domain-specific constraints. The entry does not collapse into that parent because the Anderson–Darling tail-weighted statistic and distribution-specific calibration It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Anderson–Darling test. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge toprime:hypothesis_testing_null_vs_alternative. No live DAG mutation is authorized.
Hierarchy paths (5) — routes to 5 parentless roots
- Anderson–Darling test → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Inductive Reasoning
- Anderson–Darling test → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Uncertainty
- Anderson–Darling test → Hypothesis Testing (Null vs. Alternative) → Verification → Evaluation → Comparison → Self Checking
- Anderson–Darling test → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Probability → Measure → Set and Membership
- Anderson–Darling test → Hypothesis Testing (Null vs. Alternative) → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
Neighborhood in Abstraction Space¶
Anderson–Darling test sits in a moderately populated region (55th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Statistical Dispersion & Testing (44 abstractions)
Nearest neighbors
- Z-test — 0.89
- Portmanteau test — 0.88
- Logrank test — 0.88
- Exact test — 0.87
- Normality test — 0.87
Computed from structural-signature embeddings · 2026-09-08
Not to Be Confused With¶
- Kolmogorov–Smirnov test. Uses the supremum CDF gap.
- Cramér–von Mises test. Uses an unweighted integrated squared gap.
- Shapiro–Wilk test. A specialized normality statistic.
- K-sample Anderson–Darling. Tests equality of several distributions.
- AIC. Compares fitted models rather than testing one distributional null in this way.
References¶
[1] T. W. Anderson and D. A. Darling, ‘Asymptotic Theory of Certain Goodness of Fit Criteria,’ Annals of Mathematical Statistics 23(2), 193–212 (1952), DOI 10.1214/aoms/1177729437. registry ↩a ↩b
[2] M. A. Stephens, ‘EDF Statistics for Goodness of Fit and Some Comparisons,’ Journal of the American Statistical Association 69, 730–737 (1974), DOI 10.1080/01621459.1974.10480196. registry ↩a ↩b
[3] T. W. Anderson and D. A. Darling, ‘A Test of Goodness of Fit,’ Journal of the American Statistical Association 49, 765–769 (1954), DOI 10.1080/01621459.1954.10501232. registry ↩