Gelman–Rubin Statistic¶
Compare between-chain with within-chain spread to flag incomplete mixing of iterative simulations.
Core Idea¶
The Gelman–Rubin potential scale reduction factor compares variability across several independently started iterative-simulation chains with variability within them, separately for each monitored scalar quantity. Large cross-chain disagreement can warn that the estimate may still change if simulation continues. Near-one agreement is limited evidence about these chains and this scalar, not a proof that the target distribution was reached or sampled accurately.[ref-c395e579b41e][ref-95d7791d9a7c]
The 1992 original used overdispersed starts, within-chain variance \(W\), between-chain mean variance \(B/n\), and a Student-\(t\)-adjusted square-root scale comparison. A common simpler classical expression is \(\sqrt{\widehat V/W}\) with \(\widehat V=(n-1)W/n+B/n\), but the frozen seed's unrooted \(\widehat V/W\) is not the original potential Scale reduction. Modern split, rank-normalized and folded versions are later modifications, not parts of Gelman and Rubin's 1992 formula.[ref-c395e579b41e][ref-95d7791d9a7c]
Scope of Application¶
The original paper monitored scalar estimands in a Bayesian random-effects mixture model for reaction-time data. More generally the diagnostic is used with multiple chains from MCMC or other iterative simulation. Vehtari and colleagues later illustrated that traditional split R-hat could miss a chain with a different spread or a heavy-tailed mismatch, motivating rank-normalized and folded-split checks. A package's reported R-hat should therefore be read with its formula version identified.[ref-c395e579b41e][ref-95d7791d9a7c]
Clarity¶
Report chain count, retained draws, scalar estimand, starting/warmup policy and estimator variant. \(W\) averages each chain's retained sample variance, while \(B\) scales dispersion of the chain means. With one chain, the between-chain component is unavailable. A high result flags a possible mixing problem. A low result does not rule out a mode missed by every chain, poor tail exploration, or imprecise Monte Carlo estimates; ESS and MCSE provide different checks.[ref-c395e579b41e][ref-95d7791d9a7c]
Manages Complexity¶
A single between/within comparison per scalar makes many simulation histories screenable. The original reaction-time analysis used it to decide which model summaries might still sharpen with longer runs. But the compression is lossy: chains can agree for a wrong reason, and raw-variance diagnostics can miss unequal scale or heavy-tail problems. Modern rank and fold changes address specific blind spots without making R-hat infallible.[ref-c395e579b41e][ref-95d7791d9a7c]
Abstract Reasoning¶
Choose a scalar function, compare its within-chain spread to cross-chain disagreement, and interpret a larger pooled-to-within scale as a signal that continued sampling may change inference. Then ask which failure the version can detect: unsplit location disagreement, split-chain drift, or modern folded/rank-normalized scale and tail mismatch. The numerical comparison is diagnostic evidence, never a logical implication of convergence.[ref-c395e579b41e][ref-95d7791d9a7c]
Knowledge Transfer¶
The roles transfer from one posterior model or iterative sampler to another if multiple comparable chains and the same scalar function exist. The named statistic is not the mathematical convergence property itself, nor the Monte Carlo sampler producing the chains.
[^ref-c395e579b41e]: Andrew Gelman and Donald B. Rubin, “Inference from Iterative Simulation Using Multiple Sequences”, Statistical Science 7(4) (1992), 457–472; original scan, §§1–2.4 and §4. [^ref-95d7791d9a7c]: Aki Vehtari et al., “Rank-Normalization, Folding, and Localization: An Improved \(\widehat R\) for Assessing Convergence of MCMC”, Bayesian Analysis 16(2) (2021), 667–718, §§1–4, especially Fig. 2 and §4.2.
Neighborhood in Abstraction Space¶
Gelman–Rubin Statistic sits in a sparse region of the domain-specific corpus (84th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Empirical Measurement & Statistical Inference Methods (50 abstractions)
Nearest neighbors
- Median Absolute Deviation — 0.83
- Laser Diffraction Analysis — 0.81
- Advanced Z-Transform — 0.81
- Statistical regularity — 0.81
- Complete mixing — 0.81
Computed from structural-signature embeddings · 2026-10-08