Skip to content

Tukey's Test of Additivity

Test a one-degree product-of-main-effects departure from additivity in a two-way response table, especially when each cell has one observation.

Version
v1 · 2026-10-03 · History
Domain-specific #
13679
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomains
Two Way Anova, Interaction Tests → Experimental Design & Statistics
Aliases
Tukey's one-degree-of-freedom test for nonadditivity, Tukey nonadditivity test

Core Idea

Tukey's test of additivity asks whether a two-way table departs from an additive row-plus-column mean structure in one particular way. With one response per row–column cell, a fully unrestricted interaction fits the table exactly and leaves no independent residual variance for an ordinary interaction F test. Tukey's construction instead allocates one degree of freedom to an interaction proportional to the product of the row and column main effects, leaving the remaining residual variation for comparison.[1][2][3]

For rows \(i=1,\ldots,a\) and columns \(j=1,\ldots,b\), write the restricted model as

\[ Y_{ij}=\mu+\alpha_i+\beta_j+\gamma\alpha_i\beta_j+\epsilon_{ij}. \]

The null is \(\gamma=0\). Fitted main effects from the additive model define a product direction; the test compares residual variation projected onto that direction with the variation still left over. In the usual complete unreplicated layout, the numerator has one degree of freedom and the remaining denominator has \(ab-a-b=(a-1)(b-1)-1\), assuming that number is positive. The familiar F calibration depends on the specified error model.[1][2]

The test's strength and limitation are the same restriction: it can make a particular nonadditivity testable in an unreplicated table, but does not examine every possible interaction shape. A rejection is evidence for this product-shaped departure under the assumptions, not a diagnosis of its substantive cause or a mandate to apply a particular response transformation. Nonrejection does not prove universal additivity.[3]

Structural Signature

  1. Two-factor response table: a quantitative response is indexed by two categorical factors, canonically with one observation per combination.[2]
  2. Additive baseline: the expected response is represented by a grand mean, row effect and column effect.
  3. Restricted alternative: a single coefficient \(\gamma\) multiplies the product of corresponding row and column effects; this is a one-dimensional direction within the broader interaction space.[1][3]
  4. Estimated contrast: additive-model fitted effects provide the product regressor, equivalent to the interaction contrast used for the one-degree sum of squares.[1][3]
  5. Residual comparator: the contrast's sum of squares is compared with the remaining mean square, requiring positive residual degrees of freedom and a defensible error distribution.[1][2]
  6. Bounded inference: a calibrated F test assesses \(\gamma=0\) against this alternative, not all ways the factors might interact.

Condensed: unreplicated two-way table + additive fit + product-shaped interaction contrast + residual F comparison = Tukey's additivity test.

Sig role-phrases: crossed response table → supplies one observation per cell; additive fit → supplies row and column effects; their product → defines the single nonadditivity direction; residual mean square → supplies comparator; F statistic → tests that direction under assumptions, not every interaction.

What It Is Not

  • Not the full two-way ANOVA interaction test. An unrestricted interaction in a one-observation-per-cell table consumes all remaining degrees of freedom; replicated designs can support a more general interaction assessment.[2]
  • Not Tukey's honestly significant difference procedure. That is a multiple-comparison method for group means, not this one-degree nonadditivity diagnostic.
  • Not the Siegel–Tukey test. That separate two-sample rank test concerns relative dispersion, not a row-by-column product interaction.
  • Not an omnibus proof of additivity. Failure to reject may reflect low power or an interaction orthogonal to the chosen product direction.[3]
  • Not automatically a response-transformation recipe. A transformation may be investigated after nonadditivity, but the test neither specifies which one nor guarantees it will restore an additive interpretation.
  • Not valid as a classical F comparison with no denominator degrees of freedom. A \(2\times2\) unreplicated table leaves \(ab-a-b=0\), so this construction cannot estimate its remaining error variance.[1]

Scope of Application

The canonical use is a balanced, complete table with two factors and one response per cell: for example, crop varieties crossed with blocks, or machines crossed with measurement conditions. These are settings where one wants to assess whether row and column effects combine additively, but there is no within-cell replication for a conventional unrestricted interaction F test. The test requires enough factor levels to leave residual degrees of freedom after its one-degree contrast.[1][2]

Montana State teaching notes demonstrate the calculation using a temperature-by-pressure table of impurity measurements. That example supplies an actual two-factor, one-observation-per-cell context rather than evidence that every temperature–pressure relation is product-shaped. A constructed variety-by-block table serves the same structural demonstration in agriculture.[2]

The method is especially sensitive to interaction aligned with products of main effects. Other nonadditivity tests are available when other structures are plausible; an original R Journal research article places Tukey's one-degree procedure among multiple alternatives for unreplicated two-way layouts.[3]

Clarity

With one observation per cell, the additive model has \(a+b-1\) free mean parameters under standard identification constraints. Of the \(ab\) observed responses, \((a-1)(b-1)\) directions remain after fitting it. Calling all of that variation “error” presumes additivity; assigning all of it to unrestricted interaction leaves nothing to estimate error. Tukey's test claims only one direction for interaction and compares it with the rest. Thus it is a targeted diagnostic rather than a magical recovery of unrestricted interaction from unreplicated data.[2]

The product direction matters. Suppose row effects are large only for some levels and column effects also vary. The restricted alternative asks whether the extra cell deviation scales with \(\alpha_i\beta_j\). A crossing pattern or latent subgroup pattern that is not aligned with that product can escape detection. The null therefore needs to be read as “no detectable product-shaped departure under this test,” not “there is no interaction whatsoever.”[1][3]

A rejection flags this product-shaped nonadditivity under the test assumptions; it does not select a transformation, causal mechanism or replacement model. Nonrejection likewise does not establish global additivity. This is an inference boundary rather than another intrinsic tension. Diagnostic: what additional design, residual pattern or substantive evidence would justify a revised model?

Manages Complexity

The full \(a\times b\) interaction has \((a-1)(b-1)\) degrees of freedom. A single product-shaped parameter compresses that space to one testable contrast and leaves \(ab-a-b\) for a residual estimate. That is a deliberate trade: a sharper, feasible comparison at the cost of narrower sensitivity.[2]

This compression also makes the source of evidence inspectable. Analysts can identify the fitted row and column effects, the product regressor, the one-degree sum of squares, the remaining mean square and the associated assumptions. If one main-effect vector is effectively zero, the product contrast can degenerate; the method does not conjure information from an uninformative design.

Abstract Reasoning

Begin by identifying the response, the two factors, their levels, replication and missing cells. If there is one observation in each complete cell, fit the additive model. Compute each row and column effect relative to the grand mean; form their products by cell. Project the additive residual pattern onto that product direction, then compare its one-degree sum of squares with the remaining residual mean square, provided \(ab-a-b>0\) and the error assumptions are credible. Test \(H_0:\gamma=0\) against the product-shaped alternative.[1][2]

The test should be interpreted alongside the experimental design and residual diagnostics. Independence, a suitable variance model and, for the stated exact F reference, an appropriate error distribution are not replaced by the algebra. A significant result can motivate investigation of substantive interaction mechanisms, other models or response scales; a nonsignificant result leaves open alternative interaction shapes.[1][3]

Knowledge Transfer

A variety-by-block yield table and a temperature-by-pressure impurity table differ in subject matter, but both instantiate two factors, one response per cell, additive main effects, a product-effect contrast and a residual comparison. The method transfers by these roles, not by assuming identical physical mechanisms.

In the live catalog, Statistical Test is the accepted strict parent because this procedure has a null, statistic, reference comparison and bounded conclusion. Two-Way Analysis of Variance is its model context; Interaction (statistics) is the broader phenomenon it probes. Siegel–Tukey Test is a lexical near miss, not a parent or alias.

Examples

Variety-by-block yields

Imagine \(a\) varieties each grown once in each of \(b\) blocks. The additive baseline attributes yield to variety and block effects. A product contrast asks whether the variety deviation is systematically amplified or damped in blocks with large block effects. If the F comparison is significant, the data support that restricted departure under the design assumptions; they do not alone reveal a biological mechanism. This is a constructed application of the test.[1]

Mapped back: rows = varieties; columns = blocks; response = yield; null = additive mean; alternative = one product-effect coefficient.

Temperature-by-pressure impurity

Montana State's teaching example crosses three temperature levels with five pressure levels, once per cell. The worked ANOVA allocates one degree of freedom to the nonadditivity term and seven to the remaining error. Its reported product-term p-value is approximately 0.566, which does not supply evidence for that particular nonadditivity in those data; it is not proof of a perfectly additive process.[2]

Mapped back: rows = temperature; columns = pressure; response = impurity; target = product-shaped departure; output = bounded F comparison.

Unreplicated \(2\times2\) near miss

Four cells with one observation each have one additive-model residual direction. Assigning that single direction to the Tukey product term leaves zero degrees of freedom for the denominator; the classical test statistic cannot be calibrated by the stated F comparison. More replication or a different justified design is needed.[1]

Structural Tensions

Feasible test versus incomplete interaction coverage. With one observation per cell, restricting interaction to the product-of-main-effects direction gives a one-degree test with residual degrees left for an F denominator; a full unrestricted interaction would exhaust the information needed for an independent error estimate. The benefit is a calculable targeted diagnostic; the cost is blindness to departures orthogonal to that chosen direction. Diagnostic: is product-of-main-effects interaction scientifically plausible for this table, and what other departures could remain invisible?[3]

Additive-model residuals versus genuine error. Treating the variation left after the one-degree contrast as error makes the test possible in an unreplicated table; if other interactions are present, that denominator can absorb systematic structure and distort the F comparison. Modeling more interaction would expose structure but would consume the very degrees of freedom used as error here. Diagnostic: what design knowledge or residual pattern supports treating the remaining variation as error rather than unmodeled systematic structure?[2]

Structural–Framed Character

This sits between structural and framed. Given a response table, additive fit, product-effect contrast and distributional assumptions, the one-degree statistic and F comparison are mathematical. Its evidential weight depends on whether the row/column factors, no-replication design and residual-error assumptions fit the study; a \(p=0.5660\) result in the cited impurity table means this test did not reject that product alternative, not that all interaction is absent. Human experimental practice determines the factors and whether replication was affordable, while statistics teaching and software have carried Tukey's name as a concise diagnostic for this constrained design. The vocabulary travels literally from agricultural blocks to industrial temperature/pressure cells only when the same crossed table and product-shaped alternative are present; importing “Tukey's test” to any generic ANOVA interaction test misidentifies the procedure. Its character: a targeted, assumption-framed statistical diagnostic with exact algebra but limited inferential reach, not a phenomenon or a causal explanation.[1][3][2]

Structural Core vs. Domain Accent

The portable skeleton is testing a restricted direction inside a larger alternative while using remaining variation as a comparator; Hypothesis Testing is the broad prime-level relation. The domain-bound mechanism is two-way ANOVA without replication, row and column effect estimates, their product as the one-degree interaction direction, and an F ratio using the residual degrees of freedom. The named entry fails the prime bar because most restricted tests do not use this product regressor, and many factorial designs have replication that permits a different interaction analysis. A generic one-degree contrast can instantiate the skeleton without becoming Tukey's additivity test.[1]

This entry is a kind of Statistical Test.

  • Hypothesis Testing (Null vs. Alternative): the test contrasts \(\gamma=0\) with a declared product-shaped alternative.
  • Factorial Design: rows and columns correspond to crossed factor levels, though the inference is narrowed by absent replication.
  • Projection: the one-degree component can be understood as a projection of additive residuals onto a specified product-effect direction.

These are conceptual relations; the only accepted strict upward DAG edge is to the live domain-specific Statistical Test node.

Relationships to Other Abstractions

Local relationship map for Tukey's Test of AdditivityParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Tukey's Testof AdditivityDOMAINDomain-specific abstraction: Statistical Test — is a kind ofStatistical TestDOMAIN

Current abstraction Tukey's Test of Additivity Domain-specific

Parents (1) — more general patterns this builds on

  • Tukey's Test of Additivity is a kind of Statistical Test Domain-specific

    Tukey's one-degree additivity procedure is a statistical test.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Tukey's Test of Additivity sits in a sparse region of the domain-specific corpus (99th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (2551 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

Tukey HSD compares means after ANOVA; Siegel–Tukey ranks two samples to investigate dispersion; unrestricted two-way interaction allows an entire interaction surface; Tukey's additivity test tests one product-shaped direction in that surface. Similar names and shared ANOVA vocabulary do not make the procedures interchangeable.

References

[1] Purdue University statistics lecture, “Some More Topics on Two-Way Models,” Tukey's Test for Additivity. Model, contrast and F degrees of freedom. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n

[2] Montana State University statistics teaching notes, §4.12, “Tukey's Test for Nonadditivity”. Unreplicated design, residual degrees of freedom and worked temperature–pressure example. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m

[3] The R Journal, “Exploring Interaction Effects in Two-Factor Studies using the hiddenf Package in R,” §6.1. Original research/software article on unreplicated interaction tests and Tukey's one-degree model. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j

[4] CRAN package documentation, additivityTests: tukey.test. Implementation and bibliographic reference to Tukey, J. W. (1949), “One Degree of Freedom for Non-Additivity,” Biometrics 5, 232–242; original full article not yet checked. registry