Skip to content

Chauvenet's criterion

Flag a single extreme observation when, under a fitted normal-error model, the expected number of sample observations at least as far from the mean is below one half.

Version
v2 · 2026-08-30 · History
Domain-specific #
1462
Origin domain
statistics
Subdomain
classical outlier rejection

Core Idea

Chauvenet's criterion marks an observation with standardized absolute deviation \(d=|x_i-\bar{x}|/s\) when the two-sided normal tail probability at least as extreme as \(d\), multiplied by sample size \(n\), is less than \(1/2\).[1][1] The fitted normal model converts distance from the sample center into a tail probability, multiplication by sample size converts that probability into an expected count of equally extreme observations, and a sample-size-dependent boundary selects values whose modeled expected occurrence is below one half.

Its autonomous residual is the normal-tail expected-count rule with the fixed one-half cutoff and sample-size-dependent standardized-deviation boundary, not generic sigma clipping, every unusual datum, a robust estimator, or a model-free license to discard inconvenient observations. The identity fails when normality is unjustified, the tail is one-sided without declaration, the same fixed sigma cutoff is used for every n, center and scale are recomputed opportunistically without a protocol, several observations are removed as though the single-outlier derivation applied unchanged, or rejection is treated as proof of error.

Recognition requires an analyst to verify the normal and independence assumptions, define center and scale estimators, compute the two-sided tail rather than one tail, preserve the sample-size dependence, identify at most the candidate values permitted by the chosen procedure, and separate flagging from an automatic scientific decision to delete. Once established, it supports providing a historically important quantitative screen for suspect measurements, exposing how outlier thresholds depend on sample size, comparing classical rejection rules, and motivating more robust or explicitly modeled alternatives without turning those uses into the definition.

Structural Signature

  • Carrier: a finite univariate sample provisionally modeled as independent observations from one normal population, with one suspected extreme observation and estimated center and scale
  • Inputs or antecedent state: sample size, sample mean, sample standard deviation, suspected deviation, two-sided normal tail probability, the one-half expected-count threshold, estimation convention, and any iteration rule
  • Constitutive operation: The fitted normal model converts distance from the sample center into a tail probability, multiplication by sample size converts that probability into an expected count of equally extreme observations, and a sample-size-dependent boundary selects values whose modeled expected occurrence is below one half
  • Invariant: one declared normal-error model, sample-dependent center and scale, a two-sided extremeness calculation, and the condition n times tail probability below one half jointly determine the flagged observation
  • Recognition test: verify the normal and independence assumptions, define center and scale estimators, compute the two-sided tail rather than one tail, preserve the sample-size dependence, identify at most the candidate values permitted by the chosen procedure, and separate flagging from an automatic scientific decision to delete
  • Output or consequence: providing a historically important quantitative screen for suspect measurements, exposing how outlier thresholds depend on sample size, comparing classical rejection rules, and motivating more robust or explicitly modeled alternatives
  • Failure boundary: normality is unjustified, the tail is one-sided without declaration, the same fixed sigma cutoff is used for every n, center and scale are recomputed opportunistically without a protocol, several observations are removed as though the single-outlier derivation applied unchanged, or rejection is treated as proof of error

What It Is Not

  • It is not the whole field of statistics; many objects in that field do not satisfy its constitutive rule.
  • It is not its canonical example. For a declared sample size, the criterion chooses the standard-normal cutoff whose combined tail area is one divided by twice the sample size, then compares the suspected absolute standardized residual with that cutoff. That is an instance, not a definition.
  • It is not Outlier Leverage. Outlier Leverage describes disproportionate influence of extreme observations on an aggregate. Chauvenet's criterion is a particular normal-model screening rule. Randomness Test and Hypothesis Testing are broader families and do not fix its expected-count threshold.
  • It is not an unrestricted metaphor. Historical and software presentations differ over biased or unbiased scale estimates, sequential repetition, small-sample corrections, and whether equality at the threshold is retained; these choices can change borderline classifications and must be declared

Scope of Application

Chauvenet's criterion applies when the analyst can specify a finite univariate sample provisionally modeled as independent observations from one normal population, with one suspected extreme observation and estimated center and scale and establish that one declared normal-error model, sample-dependent center and scale, a two-sided extremeness calculation, and the condition n times tail probability below one half jointly determine the flagged observation. The entry is descriptive statistical reference material, not an instruction to delete data. Any exclusion must be justified by the scientific model, protocol, provenance, and consequences for inference.[2]

  • Recognition. verify the normal and independence assumptions, define center and scale estimators, compute the two-sided tail rather than one tail, preserve the sample-size dependence, identify at most the candidate values permitted by the chosen procedure, and separate flagging from an automatic scientific decision to delete
  • Comparison. Compare legitimate instances through sample size, center estimator, scale estimator, tail convention, cutoff inequality, number of suspected points, iteration, normality, independence, contamination fraction, and substantive provenance.
  • Boundary. Historical and software presentations differ over biased or unbiased scale estimates, sequential repetition, small-sample corrections, and whether equality at the threshold is retained; these choices can change borderline classifications and must be declared
  • Use. Preserve every assumption when using the identity for providing a historically important quantitative screen for suspect measurements, exposing how outlier thresholds depend on sample size, comparing classical rejection rules, and motivating more robust or explicitly modeled alternatives.

Clarity

A clear claim names the carrier, governing rule, assumptions, and recognition test. This matters because criterion is often presented as an objective fact about bad data even though its verdict depends on a fitted normal model, estimator choices, sample size, and procedural convention. The disciplined statement is that the object counts as Chauvenet's criterion exactly when one declared normal-error model, sample-dependent center and scale, a two-sided extremeness calculation, and the condition n times tail probability below one half jointly determine the flagged observation

Identity and measurement remain separate. The suspected point helps determine the sample mean and standard deviation used to judge itself, making the classical rule nonrobust; sensitivity analyses and transparent retention of raw observations are essential context. Approximation or noisy evidence may weaken a classification without changing its definition.

Manages Complexity

The abstraction compresses historical table lookup, quantile-function implementations, one-pass and iterative uses, alternative scale conventions, robust descendants, and measurement-science applications into a stable carrier, rule, invariant, and failure boundary. It makes comparison tractable while retaining the variables that control validity.

Compression can hide assumptions. A responsible use therefore declares sample size, center estimator, scale estimator, tail convention, cutoff inequality, number of suspected points, iteration, normality, independence, contamination fraction, and substantive provenance and returns to the full diagnostic whenever a convention or boundary case changes.

Abstract Reasoning

  1. Type the carrier. Establish a finite univariate sample provisionally modeled as independent observations from one normal population, with one suspected extreme observation and estimated center and scale and reject examples from a different problem.
  2. Lock the rule. Express that one declared normal-error model, sample-dependent center and scale, a two-sided extremeness calculation, and the condition n times tail probability below one half jointly determine the flagged observation independently of one notation or implementation.
  3. Derive carefully. Infer providing a historically important quantitative screen for suspect measurements, exposing how outlier thresholds depend on sample size, comparing classical rejection rules, and motivating more robust or explicitly modeled alternatives only under the stated assumptions.
  4. Stress-test. Contrast the legitimate boundary case—Historical and software presentations differ over biased or unbiased scale estimates, sequential repetition, small-sample corrections, and whether equality at the threshold is retained; these choices can change borderline classifications and must be declared—with this counterexample: discarding every observation more than two sample standard deviations from the mean is not Chauvenet's criterion because the threshold does not adapt to sample size and may not use the one-half expected-count rule.

Knowledge Transfer

Transfer within statistics is strong when new cases preserve the same carrier, mechanism, and diagnostic. The move from For a declared sample size, the criterion chooses the standard-normal cutoff whose combined tail area is one divided by twice the sample size, then compares the suspected absolute standardized residual with that cutoff. to Measurement-science guidance can use the criterion as one diagnostic alongside provenance, instrument behavior, physical plausibility, residual structure, and robust analysis rather than as an unreviewed deletion command.[2] demonstrates that continuity.[3]

Outside the domain, only the skeleton—translate extremeness under a null model into an expected count across the number of opportunities and flag values below a rarity budget—travels automatically. The terms outlier, normal distribution, standardized residual, two-sided tail, expected count, sample mean, sample standard deviation, cutoff, rejection, and model assumption retain domain-specific meanings, so every role and inference must be revalidated.

Examples

Canonical

For a declared sample size, the criterion chooses the standard-normal cutoff whose combined tail area is one divided by twice the sample size, then compares the suspected absolute standardized residual with that cutoff. Larger samples require a more extreme standardized deviation because seeing at least one distant observation becomes less surprising as the number of opportunities grows; the rule is therefore not simply a universal two- or three-sigma threshold. It is canonical because the carrier, rule, invariant, and consequence are all inspectable.[1]

Mapped back: a finite univariate sample provisionally modeled as independent observations from one normal population, with one suspected extreme observation and estimated center and scale → The fitted normal model converts distance from the sample center into a tail probability, multiplication by sample size converts that probability into an expected count of equally extreme observations, and a sample-size-dependent boundary selects values whose modeled expected occurrence is below one half → one declared normal-error model, sample-dependent center and scale, a two-sided extremeness calculation, and the condition n times tail probability below one half jointly determine the flagged observation → providing a historically important quantitative screen for suspect measurements, exposing how outlier thresholds depend on sample size, comparing classical rejection rules, and motivating more robust or explicitly modeled alternatives

Applied / In Practice

Measurement-science guidance can use the criterion as one diagnostic alongside provenance, instrument behavior, physical plausibility, residual structure, and robust analysis rather than as an unreviewed deletion command.[2] A point produced by a known transcription or instrument fault has external grounds for exclusion, whereas a point from a heavy-tailed but valid process can be informative; the same numerical flag does not settle those cases. It qualifies only after the same diagnostic and failure boundary are checked.[2]

Mapped back: declared instance → recognition test → boundary check → qualified use

Structural Tensions

  • T1: Exact identity vs. practical recognition. The constitutive condition may be exact while evidence is indirect. Diagnostic: Can the reviewer state both the condition and the warrant?
  • T2: Canonical form vs. variants. historical table lookup, quantile-function implementations, one-pass and iterative uses, alternative scale conventions, robust descendants, and measurement-science applications can preserve or change the identity. Diagnostic: Which named role is invariant across the variants?
  • T3: Compression vs. hidden assumptions. The label is useful only while prerequisites remain visible. Diagnostic: Can each downstream inference be traced to a declared assumption?
  • T4: Autonomy vs. reduction. The candidate uses broader structures but claims the normal-tail expected-count rule with the fixed one-half cutoff and sample-size-dependent standardized-deviation boundary, not generic sigma clipping, every unusual datum, a robust estimator, or a model-free license to discard inconvenient observations. Diagnostic: Does that residual still support independent recognition after the parent and neighbors are subtracted?

Structural–Framed Character

The entry is structurally mixed but domain-framed. Its portable skeleton is translate extremeness under a null model into an expected count across the number of opportunities and flag values below a rarity budget; its identity-bearing terms are outlier, normal distribution, standardized residual, two-sided tail, expected count, sample mean, sample standard deviation, cutoff, rejection, and model assumption. Those terms determine admissible objects, evidence, and consequences inside statistics.

Structural Core vs. Domain Accent

The structural core is a carrier governed by The fitted normal model converts distance from the sample center into a tail probability, multiplication by sample size converts that probability into an expected count of equally extreme observations, and a sample-size-dependent boundary selects values whose modeled expected occurrence is below one half and tested by verify the normal and independence assumptions, define center and scale estimators, compute the two-sided tail rather than one tail, preserve the sample-size dependence, identify at most the candidate values permitted by the chosen procedure, and separate flagging from an automatic scientific decision to delete. The domain accent is constitutive rather than decorative, so an analogy that preserves only the skeleton is not another instance of Chauvenet's criterion.

The proposed strict upward parent is prime:statistical_inference. The criterion reasons from a finite noisy sample and a probabilistic population model to a qualified judgment about one observation. The Gaussian tail, sample-size adjustment, and one-half expected-count rule provide the domain-specific residual. The edge is proposal-only and points to a frozen prior-baseline Prime.

The entry does not collapse into the parent because the normal-tail expected-count rule with the fixed one-half cutoff and sample-size-dependent standardized-deviation boundary, not generic sigma clipping, every unusual datum, a robust estimator, or a model-free license to discard inconvenient observations A thematic neighbor is declined whenever it does not literally subsume that rule.

The prospective workspace queue contains one strict upward edge to prime:statistical_inference. No live DAG mutation is authorized.

Relationships to Other Abstractions

Local relationship map for Chauvenet's criterionParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Chauvenet's criterionDOMAINPrime abstraction: Statistical Inference — is a kind ofStatisticalInferencePRIME

Current abstraction Chauvenet's criterion Domain-specific

Parents (1) — more general patterns this builds on

  • Chauvenet's criterion is a kind of Statistical Inference Prime

    The proposed strict upward parent is prime:statistical_inference.

Hierarchy paths (4) — routes to 4 parentless roots

Neighborhood in Abstraction Space

Chauvenet's criterion sits in a moderately populated region (47th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.

Family — Bayesian Inference & Probabilistic Models (23 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-09-08

Not to Be Confused With

  • Peirce's criterion. A distinct classical rejection calculation designed to compare hypotheses about doubtful observations and capable of handling multiple suspected values.
  • Grubbs' test. A formal normal-theory outlier test with a chosen significance level and test-statistic distribution.
  • Sigma clipping. Repeated removal outside a fixed or selected sigma band, often without Chauvenet's sample-size cutoff.
  • Robust Chauvenet rejection. A later family replacing vulnerable center and scale estimates and adding staged procedures; it is not the nineteenth-century rule unchanged.

References

[1] William Chauvenet, A Manual of Spherical and Practical Astronomy, Volume II, J. B. Lippincott, 1863, discussion of rejection of doubtful observations. registry ↩a ↩b ↩c

[2] Vic Barnett and Toby Lewis, Outliers in Statistical Data, 3rd ed., Wiley, 1994, ISBN 978-0-471-93094-5. registry ↩a ↩b ↩c ↩d

[3] John R. Taylor, An Introduction to Error Analysis, 2nd ed., University Science Books, 1997, pp. 166-168, ISBN 978-0-935702-75-0. registry