Skip to content

Siegel–Tukey Test

The Siegel–Tukey test is a two-sample nonparametric rank procedure that detects relative dispersion by assigning ranks alternately from the pooled extremes toward the center.

Version
v1 · 2026-09-28 · History
Domain-specific #
7760
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomain
Nonparametric Statistics → Experimental Design & Statistics
Aliases
Siegel–Tukey rank test, Siegel–Tukey test for scale

Core Idea

The Siegel–Tukey test is a two-sample nonparametric procedure for testing whether one population is more dispersed than another when observations are at least ordinal and the samples are independent.[1] It converts the pooled observations into a special rank order that emphasizes distance from the center: ranks are assigned alternately to the smallest and largest remaining observations, proceeding inward.[2] An ordinary Wilcoxon rank-sum calculation on those dispersion ranks then tests whether one group occupies the extremes more often.[3]

The transformation is the identity's distinctive operation. If groups have similar location and scale, their labels should mix through both center and tails. If group A is more dispersed, its values tend to appear at both low and high extremes, receiving systematically different Siegel–Tukey ranks from a concentrated group. The test detects relative spread without calculating variance, but location differences can mimic or obscure dispersion differences, so equal or suitably aligned central tendency is a load-bearing assumption.[4]

The invariant is: two independent samples are pooled and ordered, alternate-extreme dispersion ranks are assigned by the defined convention, and a rank-sum null distribution evaluates whether one group is systematically more extreme. Use ordinary ascending ranks and the procedure becomes a location test. Use paired data or incomparable ordinal scales and the identity fails.

Structural Signature

Sig role-phrases:

  • independent samples — two unpaired groups are observed on the same ordinal or quantitative scale
  • dispersion hypothesis — the null compares group scale or spread, with a directional or two-sided alternative stated in advance
  • location condition — equal or suitably aligned central tendency is required for a clean scale interpretation
  • pooled order — all observations are combined and sorted while their group labels are retained
  • alternate-extreme scoring — ranks are assigned from alternating low and high extremes toward the pooled center under a fixed convention
  • tie rule — equal observations are handled by a declared scoring and reference-distribution procedure
  • group rank sum — each sample's dispersion scores are added to express its occupancy of the pooled extremes
  • rank-sum transform — the special rank sum is converted to its Wilcoxon or Mann–Whitney-type statistic
  • null reference law — exact or asymptotic sampling probabilities are selected for the group sizes, direction, and tie structure
  • inferential output — a p-value or rejection decision evaluates relative dispersion under the declared assumptions
  • magnitude limitation — the result does not estimate how much larger a variance or scale parameter is
  • procedure boundary — ordinary ascending ranks, paired observations, or an unaddressed location shift do not support the Siegel–Tukey dispersion interpretation

What It Is Not

  • Not the ordinary Mann–Whitney or Wilcoxon location test. It reuses rank-sum machinery only after alternate-extreme scoring; ordinary ascending ranks test systematic location ordering instead of dispersion.
  • Not a variance estimate. The procedure supplies evidence about relative spread under its scale model but does not state how many units larger a variance or scale parameter is.
  • Not a general omnibus test for any distributional difference. Skewness, mixtures, tail imbalance, and isolated outliers can influence the special ranks without being identified by the final statistic.
  • Not valid for paired samples by default. Its pooled ranking and null reference law require two independent groups; pairing changes the data structure and the applicable procedure.
  • Not assumption-free because it is nonparametric. Comparable ordering, independence, a defensible location condition, a declared tie rule, and a valid exact or asymptotic reference distribution remain necessary.
  • Not Levene's or Bartlett's test. Those procedures target variance homogeneity through different statistics and assumptions rather than Siegel–Tukey's alternating extreme-to-center ranks.
  • Not a multiple-testing or multi-group post-hoc correction. Benjamini–Hochberg controls false discoveries and Nemenyi compares groups after a rank test; neither performs this two-sample dispersion transformation.
  • Not proof that the group with the extreme ranks has a larger numerical variance. A small p-value rejects the stated null under the convention; location shift, ties, shape, and outliers must still be examined before interpreting the source of the difference.

Scope of Application

The Siegel–Tukey test applies to two independent samples on a common ordinal or quantitative scale when relative dispersion is the question, pooled ordering is meaningful, group locations are equal or defensibly aligned, and the alternate-extreme scoring and null calculation are fully specified.

  • Ordinal two-sample comparisons — ordered observations can support the procedure without numerical distance assumptions when ties and the common scale are handled explicitly.
  • Continuous two-sample comparisons — measurements from independent groups are pooled and ranked when a scale difference is sought without relying on a normal-theory variance test.
  • Equal-location dispersion studies — the cleanest habitat compares groups centered at the same location so occupancy of both pooled tails can be interpreted as relative spread.
  • Location-adjusted analysis — a defensible centering step can precede the test when the adjustment rule and its effect on the reference distribution are stated rather than silently imposed.
  • One-sided scale alternatives — a prespecified direction tests whether one named group is more dispersed under the chosen rank orientation.
  • Two-sided dispersion alternatives — rejection indicates unequal scale evidence without by itself identifying variance magnitude, tail balance, or the source of the difference.
  • Small-sample exact inference — exact rank-sum probabilities are used when group sizes and tie structure permit the required finite-sample null calculation.
  • Larger-sample asymptotic inference — approximations are used only with adequate sample size and documented continuity, variance, and tie corrections.
  • Tied observations — discrete or rounded measurements remain admissible only under a declared alternate-extreme score allocation and a compatible reference-law adjustment.
  • Nonparametric dispersion screening — the test can complement distribution plots and effect descriptions when parametric scale assumptions are doubtful but independence and ordering remain defensible.
  • Sensitivity diagnosis — rechecking location alignment and the influence of isolated extremes distinguishes broad dispersion evidence from a result driven by shift or a few observations.
  • Method boundary cases — paired, censored, strongly location-shifted, heavily tied, or radically shape-different samples require another procedure or a specifically justified adjustment rather than mechanical Siegel–Tukey use.

Clarity

A clear report names the null and alternative, whether the test is one- or two-sided, the rank convention, group sizes, ties, missing-data handling, exact or asymptotic p-value, and any location adjustment. “More variable” should specify which group and under what scale model.

Because implementations can reverse rank direction, raw rank sums alone are not interpretable without the convention. The p-value is evidence under the null, not the probability that equal dispersion is true.

Manages Complexity

The special rank transform reduces a two-sided tail pattern to a standard rank-sum problem. It avoids estimating high-order moments and remains meaningful on ordinal data.

Compression loses magnitude and shape detail. A significant result does not reveal whether spread comes from both tails, skewness, outliers, or mixture structure. The test manages the decision problem, not the full distributional diagnosis.

Abstract Reasoning

The diagnostic move goes from the pooled ordinal observations to an alternate-extreme rank scale, and from each group's rank sum to evidence about relative dispersion under the null distribution. Values near either tail receive similar leverage on the derived scale, so a group appearing at both extremes is distinguished from one concentrated near the pooled center. The inference is about relative spread under the test's exchangeability and location conditions, not a direct estimate of either population's variance.

Interventions expose confounding. Shifting one sample without changing its scale can move it into one tail and alter the dispersion ranks; aligning the locations asks whether the apparent scale difference persists. Removing or separately diagnosing a few extreme observations tests whether the result is broadly distributed or carried by isolated points. Ties, pairing, and unequal shapes mark regime boundaries because they change the rank assignment or reference law. A small p-value therefore licenses rejection of the stated equal-dispersion null under the declared convention, but it does not by itself identify whether tail balance, skewness, mixture structure, or outliers produced the difference.

Knowledge Transfer

Within nonparametric statistics, the literal Siegel–Tukey procedure transfers across subject matters and ordinal or continuous measurements when there are two independent samples, a defensible common ordering, and a dispersion question under appropriate location conditions. What carries is the pooled order, alternating extreme-to-center scores, group rank sum, and exact or asymptotic null calculation. Location alignment, tie handling, and sensitivity to isolated extremes are the diagnostics that determine whether the same output supports a dispersion interpretation.

Beyond this test, the honest reach is a mix of (B) a shared abstract mechanism, and (C) instrument or measure: other nonparametric procedures can redesign scores for a target property and then use exchangeability and rank-sum logic as an inferential instrument. The alternate-extreme Siegel–Tukey convention and its relative-dispersion claim remain home-bound. Outside statistical inference, “ranking extremes” is only (A) analogy. Transfer stops without independent samples, a pooled order, a declared scoring rule, and a valid reference distribution; a ranked list alone is not a statistical test.

Examples

Canonical

The frozen thirteen-observation calculation. The two independent samples are A = {33, 62, 84, 85, 88, 93, 97} and B = {4, 16, 48, 51, 66, 98}.[5] After pooling and sorting, the fixed alternate-extreme convention assigns ranks 5, 12, 11, 10, 7, 6, 3 to A and 1, 4, 8, 9, 13, 2 to B.[6] Thus W_A = 54 and W_B = 37; subtracting the minimum possible sums gives U_A = 54 − 7(8)/2 = 26 and U_B = 37 − 6(7)/2 = 16.[7] Under the Wilcoxon rank-sum reference law for group sizes 7 and 6, the frozen example reports p = 0.2669, providing little reason to reject equal dispersion.[8]

Mapped back: A and B are the independent samples, and relative spread supplies the dispersion hypothesis under the stated location condition. Sorting produces the pooled order; the listed scores implement alternate-extreme scoring; 54 and 37 instantiate the group rank sum role; and 26 and 16 result from the rank-sum transform. The group-size-specific Wilcoxon distribution is the null reference law, and p = 0.2669 supplies the inferential output without estimating a variance difference under the magnitude limitation.

Applied / In Practice

A location check before a dispersion comparison. An analyst has two independent groups measured on the same ordered scale and wants to test whether one is more variable. A plot and location summary first show that the groups have materially different centers.[9] Running alternate-extreme ranks immediately would mix location shift with tail occupancy, so the analyst either justifies a prespecified location-alignment procedure whose reference law remains valid or declines the Siegel–Tukey interpretation and chooses another analysis.[10] If the locations are defensibly aligned, the analyst then freezes the scoring direction, states the tie rule, computes the rank sum, selects an exact or asymptotic reference law appropriate to group size and ties, and reports the p-value alongside the distributions.

Mapped back: The two groups must satisfy independent samples and the common-scale requirement before their pooled order is meaningful. The preliminary median comparison tests the location condition; the prespecified tail direction supplies the dispersion hypothesis; the declared convention governs alternate-extreme scoring and the tie rule; and the resulting group rank sum is evaluated by the appropriate null reference law. Declining a scale conclusion when location remains confounded enforces the procedure boundary.

Structural Tensions

T1: Dispersion sensitivity versus location confounding. Alternate-extreme scores make observations in both pooled tails influential, which gives the test sensitivity to relative spread without estimating variance. A location shift can move one group disproportionately toward one tail and alter the same scores even when its dispersion is unchanged, so the feature that detects scale also creates a route for confounding. Aligning locations may clarify the target but can change the null calculation if imposed after inspection. Diagnostic: establish the location condition or a prespecified valid adjustment before interpreting the rank sum, and decline a scale conclusion when shifting the groups to a common center removes the apparent tail-occupancy difference.

T2: Distribution-free ranks versus exchangeability and independence. Replacing numerical distances with pooled order avoids a parametric variance model and permits ordinal observations. The null reference law still requires that group labels be exchangeable under the stated hypothesis, that observations be independent across the two samples, and that the same measurement ordering apply to both. Calling the test nonparametric can conceal those structural assumptions. Diagnostic: use the rank-sum reference law only when sampling, pairing, and measurement conditions make relabeling under the null legitimate; otherwise classify the calculation as lacking the inferential distribution required by the test.

T3: Ordinal robustness versus tie ambiguity. Rank-based scoring remains meaningful when only order, rather than calibrated numerical distance, is defensible. Discrete or rounded data can place many observations at the same value, where the fixed alternating sequence no longer determines a unique allocation of extreme-to-center scores and ordinary continuous-null probabilities may be wrong. An arbitrary tie break preserves computability at the cost of reproducibility. Diagnostic: declare the tie-scoring convention and use a compatible exact or adjusted reference distribution, treating materially different conclusions across defensible tie rules as an unresolved result rather than as robust scale evidence.

T4: Simple decision output versus unresolved distribution shape. Transforming the data and returning one p-value makes a two-sided tail pattern easy to evaluate across samples. The same rejection can be driven by balanced scale expansion, skewness, a mixture, unequal tails, or isolated observations, and the test does not estimate how much larger a scale parameter is. Replacing the test with a complete distributional description would lose its focused decision role; reporting only the decision overstates what it diagnoses. Diagnostic: pair the p-value with ordered-data or distribution summaries and sensitivity to isolated extremes, and describe the conclusion as relative-dispersion evidence only when the observed shape is compatible with that interpretation.

T5: Exact inference versus asymptotic convenience. An exact null calculation respects the finite group sizes, score convention, direction, and tie structure, which is especially valuable in small samples. Asymptotic approximations are easier to obtain for larger data sets but can misstate tail probabilities when samples are small, unbalanced, or heavily tied; exact enumeration can itself become impractical as the problem grows. Diagnostic: name the one- or two-sided alternative and the computation used, and accept an asymptotic p-value only when group size and tie corrections make its difference from the corresponding finite-sample law immaterial to the decision.

T6: Siegel–Tukey autonomy versus reduction to Evaluation. The exact parent Prime Evaluation strictly subsumes the test: every qualifying Siegel–Tukey procedure evaluates relative dispersion by mapping observations through an explicit criterion-bearing comparison to an inferential judgment. The named procedure remains in situ because it additionally requires two independent samples, pooled ordering, alternate-extreme scores, a group rank sum, and the matching null reference law. Reducing the test to Evaluation gains portable object–criterion–mapping structure but loses the scoring convention and dispersion-specific evidential chain; treating it as wholly autonomous hides that complete evaluative genus. Diagnostic: if the pooled alternate-extreme ranking and reference-law judgment are removed while a criterion-governed assessment remains, Evaluation survives but the Siegel–Tukey identity does not.

Structural–Framed Character

The Siegel–Tukey test is mixed-structural. Its pooled order, alternate-extreme scoring, rank-sum transform, and null reference law form a reproducible procedure, but the scale hypothesis, location condition, tail direction, tie rule, and decision convention frame what the output means. The smallest portable skeleton is Evaluation, which preserves a bounded object, criterion-bearing frame, relevant observations, traceable mapping, and evaluative result. That portable reach belongs to the Evaluation Prime; the Siegel–Tukey test remains the specific nonparametric dispersion procedure.

Its evaluative_weight is moderate because rejection or non-rejection depends on a chosen null, direction, and significance convention rather than residing in the samples alone. Its human_practice_bound character is moderate: the ranked data and null probabilities are formal, while analysts set the hypothesis, assumptions, and inferential use. Its institutional_origin is low to moderate because statistical practice standardizes the procedure without constituting the observations. Its vocab_travels result is partial: evaluation, ranking, and null-reference language carries, whereas alternate-extreme scores and the dispersion interpretation remain method-specific. Under import_vs_recognize, Evaluation can be recognized in any criterion-to-result mapping, but this test must be imported with two independent samples, pooled ordinal order, the prescribed scores, and a valid dispersion null.

Its character: mixed-structural because Evaluation owns the portable criterion-and-verdict skeleton while statistical conventions fix the test's exact procedure and interpretive scope.

Structural Core vs. Domain Accent

The Siegel–Tukey test is a domain-specific statistical abstraction rather than a prime; it is a strict specialization of Evaluation. Its complete named signature is two independent samples and a dispersion hypothesis → pooled order → alternate-extreme scoring and tie rule → group rank sum and rank-sum transform → exact or asymptotic null reference law → p-value or rejection judgment, with location, magnitude, and procedure boundaries.

What is skeletal (could lift toward a cross-domain prime). Evaluation owns a bounded object, criterion-bearing frame, relevant observations, an evaluator or rule-governed mapping, and a verdict, score, rank, or action-guiding result with a traceable route from inputs to judgment. That complete pattern survives in educational grading, engineering design review, and legal adjudication—three unrelated domains. Removing the test-specific accent therefore leaves a genuine Evaluation: observations are still interpreted against a declared criterion and mapped to an evaluative result.

What is domain-bound. Independent two-sample structure, a relative-dispersion hypothesis, pooled ordinal observations, the fixed alternating extreme-to-center rank convention, location alignment, tie handling, group rank sums, exchangeability, and the matching finite-sample or asymptotic reference law constitute the Siegel–Tukey test. They determine what its p-value means and distinguish it from ordinary location tests; Evaluation requires none of these statistical occupants.

Why this does not clear the prime bar. The procedure adds no second substrate-independent evaluation invariant; it fixes an evaluative mapping to one nonparametric dispersion design. Remove the criterion-governed mapping from observations to a judgment and the remainder is an ordered data set, not a test. Remove alternate-extreme scoring, two-group assumptions, location condition, and the null reference law and the residual is Evaluation rather than Siegel–Tukey. Strict subsumption therefore preserves the portable object–criterion–result architecture without turning the scoring convention into a universal abstraction.

This entry is a kind of Evaluation.

Instantiates — Evaluation (Evaluation). The bounded object is the relative dispersion of two independent populations as represented by their samples. The Siegel–Tukey procedure is the evaluator; the stated null, alternative, direction, location condition, independence condition, tie rule, and significance convention form the criterion-bearing frame. Pooled observations and retained group labels are the relevant features. Alternate-extreme scoring, the group rank sum, and the exact or asymptotic null reference law provide the traceable relational mapping from those features to a p-value and rejection or non-rejection judgment. The recorded scoring convention, group sizes, ties, and reference distribution make the route auditable. Remove the dispersion criterion, the alternate-extreme mapping, or the inferential result and what remains may be ranking, measurement, or description, but it is no longer this evaluation.

The strict relation preserves a statistical residual. Evaluation supplies the object–criterion–observation–mapping–judgment architecture; the named test additionally fixes two independent samples, center-to-extreme ranks, a rank-sum reference law, and the limitation that its result is not an estimate of variance magnitude. Ordinary ascending ranks therefore instantiate a different evaluation even though they reuse some machinery.

Relationships to Other Abstractions

Local relationship map for Siegel–Tukey TestParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Siegel–Tukey TestDOMAINPrime abstraction: Evaluation — is a kind ofEvaluationPRIME

Current abstraction Siegel–Tukey Test Domain-specific

Parents (1) — more general patterns this builds on

  • Siegel–Tukey Test is a kind of Evaluation Prime

    The bounded object is the relative dispersion of two independent populations as represented by their samples.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Siegel–Tukey Test sits in a sparse region of the domain-specific corpus (67th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (2551 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • Mann–Whitney–Wilcoxon rank-sum test. The ordinary rank-sum test assigns ascending ranks to detect location ordering, whereas Siegel–Tukey first ranks observations alternately from the two extremes toward the center to target dispersion. Tell: inspect the pooled scoring convention before applying the rank-sum statistic.
  • Levene's test. Levene's test compares dispersion through transformed deviations from a group center under its own statistic and assumptions, not alternate-extreme pooled ranks. Tell: determine whether the analysis ranks pooled extremes or analyzes within-group absolute deviations.
  • Bartlett's test. Bartlett's test is a parametric likelihood-based test of variance homogeneity and is sensitive to distributional assumptions unlike the ordinal Siegel–Tukey procedure. Tell: identify whether the input is sample variances under a parametric model or ordered pooled observations.
  • Variance estimate. A variance estimate reports the magnitude of spread in squared units, while the Siegel–Tukey test gives evidence about relative dispersion without estimating that magnitude. Tell: distinguish a scale estimate from a rank-based null-test result.
  • Omnibus distributional-equality test. An omnibus test seeks any distributional difference, whereas Siegel–Tukey is interpreted specifically through its dispersion ranks and location condition. Tell: ask whether the statistic can identify the source as relative spread or merely rejects equality for an unspecified shape difference.

References

[1] National Institute of Standards and Technology, Siegel–Tukey Test, Dataplot Reference Manual (accessed 2026-09-13). registry ↩

[2] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩

[3] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩

[4] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩

[5] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩

[6] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩

[7] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩

[8] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩

[9] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩

[10] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩