Skip to content

Log-Linear Analysis

Fit and compare expected-count models for categorical contingency tables by using log-scale interaction terms to express joint and conditional associations.

Version
v2 · 2026-10-03 · History
Domain-specific #
13398
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomains
Categorical Data Analysis, Contingency Tables → Experimental Design & Statistics
Aliases
Loglinear analysis, Log-linear modeling of contingency tables

Core Idea

Log-linear analysis studies associations among categorical variables by modeling the expected count in every cell of their contingency table. It represents the logarithm of an expected count as an overall term plus selected main-effect and interaction terms. Omitting an interaction can express an independence restriction; retaining terms permits specified dependencies. Two-way as well as larger tables qualify.[^ref-e111d434554e]

The analyst fits and compares such structures under a stated sampling model, then interprets the supported associations. Likelihood-ratio comparisons are common, but sparse data or impossible cell combinations can undermine routine chi-square approximations. This is a method of categorical inference, not simply taking the logarithm of each observed count.[ref-e111d434554e][ref-e111d434554e-2]

Scope of Application

Penn State uses the Berkeley graduate-admissions table—department, sex and decision—to compare joint and conditional association structures. Bickel and colleagues' original paper supplies the underlying admissions data, not the later course's log-linear fit.[ref-e111d434554e][ref-b5a4c845778e] Vinciotti and Wit use a much larger table of 69 General Social Survey answers to estimate a dependence structure with methods suited to extreme sparsity. Both cases preserve the same expected-cell and interaction logic, though their computational regimes differ.[^ref-6cb25e66e10e]

Clarity

The method distinguishes a relationship seen in a pooled table from one that persists after another categorical variable is represented. It also differs from logistic regression: a logit model identifies a response to explain conditionally, while a log-linear contingency model represents the joint categorical distribution. Some specifications have corresponding fit results, but the questions and parameterizations are not identical.[^ref-e111d434554e]

Manages Complexity

A multiway table has many cells and possible dependencies. A small set of interaction restrictions summarizes what the model treats as independent or associated. Simpler models are easier to interpret but may miss structure; saturated models fit the sample more closely but explain less parsimoniously. As dimensions rise, most possible cells may be empty, so the sample's support and estimator must be examined rather than applying a fixed expected-count cutoff.[ref-e111d434554e][ref-e111d434554e-2][^ref-6cb25e66e10e]

Abstract Reasoning

Identify the categorical variables and observed cells, state a candidate interaction structure, fit expected counts under the relevant sampling assumptions, and compare the model with a defensible alternative. Then ask which omitted term explains a meaningful misfit and whether uncertainty assessment remains valid for the table's sparsity. The result concerns statistical association, not a causal effect by itself.[ref-e111d434554e][ref-6cb25e66e10e]

Knowledge Transfer

The same method maps from a modest admissions table to a huge survey table: cells contain category-combination counts, interaction terms express dependence claims, and fit connects those claims to observations. The proposed DAG parent is live Statistical Inference, whose sample-to-model reasoning travels more broadly. Log-linear analysis itself stays domain-specific because its expected-cell machinery and sparse-table diagnostics require categorical count data.

[^ref-e111d434554e]: Pennsylvania State University, “STAT 504: Lesson 10 — Log-Linear Models,” Overview, §§10.1–10.2.6 and §10.3, including Berkeley Example 10.3. Indexed original course text checked; direct HTML fetch returned 502 during authoring. [^ref-6cb25e66e10e]: Veronica Vinciotti and Ernst C. Wit, “Loglinear modelling of huge contingency tables,” Statistics and Computing 36, article 209 (2026), Abstract, Introduction and §5. Original open-access study. [^ref-e111d434554e-2]: Pennsylvania State University, “STAT 504: Lesson 12 — Inference for Log-linear Models: Sparse Data,” sparse-data and zero-cell discussion; indexed text checked, direct HTML fetch unavailable. [^ref-b5a4c845778e]: P. J. Bickel, E. A. Hammel and J. W. O'Connell, “Sex bias in graduate admissions: data from Berkeley,” Science 187 (1975), 398–404, original dataset context only.

Relationships to Other Abstractions

Local relationship map for Log-Linear AnalysisParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Log-Linear AnalysisDOMAINPrime abstraction: Statistical Inference — is a kind ofStatisticalInferencePRIME

Current abstraction Log-Linear Analysis Domain-specific

Parents (1) — more general patterns this builds on

  • Log-Linear Analysis is a kind of Statistical Inference Prime

    Log-linear analysis infers categorical association structure from sampled cell counts under an explicit stochastic model.

Hierarchy paths (4) — routes to 4 parentless roots

Neighborhood in Abstraction Space

Log-Linear Analysis sits in a sparse region of the domain-specific corpus (78th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Causal Inference & Regression Modeling (15 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08