Slice Sampling¶
An MCMC method that adds a height beneath an unnormalized density and updates within its level-set slice using a transition that preserves the target distribution.
Core Idea¶
Slice sampling constructs a Markov chain for a target distribution whose nonnegative integrable density is known only up to a constant. Let \(f(x)\) be the unnormalized density and \(Z=\int f(x)\,dx\). Regard the region under its graph, \(U=\{(x,y):0<y<f(x)\}\), as a joint state space with uniform density. The marginal distribution of \(x\) is \(f(x)/Z\). From a current \(x\), draw a vertical height \(y\) uniformly between zero and \(f(x)\); the resulting horizontal slice is \(S_y=\{x:f(x)>y\}\).[1]
An ideal Gibbs-style update could then draw \(x\) uniformly from the entire slice \(S_y\). But this ideal conditional draw may be difficult: the slice can be wide, high-dimensional or disconnected. Neal's practical one-dimensional method instead constructs an interval around the current point, expands it by stepping out or doubling, and uses proposals plus shrinkage to define a transition that leaves the proper slice distribution invariant. It need not produce an independent uniform draw from the whole global slice at each step. That distinction is part of the method's correctness, not a minor implementation detail.[1]
The horizontal updates, combined with height draws, leave \(f/Z\) stationary when their conditions are met. Stationarity does not itself promise rapid mixing or access to every disconnected support component. Neal gives a concrete zero-density-gap case where ordinary stepping out is not irreducible.[1]
Structural Signature¶
Sig role-phrases: unnormalized target → auxiliary height → horizontal level set → slice-invariant update → correct marginal with reachability limits.
- Target shape \(f\). A nonnegative integrable function need only be evaluated up to a proportionality constant; membership in \(S_y\) follows from comparing \(f(x)\) to \(y\). The normalizing integral \(Z\) need not be computed to run the construction.[1]
- Current state and vertical height. The current \(x\) determines the available height interval \((0,f(x))\). Drawing \(y\) uniformly there couples the new level to the local density: a high-density point may yield high or low levels, while a low-density point only yields lower ones.[1]
- Slice \(S_y\). This is the full set above the threshold, potentially a union of separated intervals. It is not synonymous with one bracket chosen by a practical update.[1]
- Valid horizontal transition. The ideal independent conditional law is uniform on all of \(S_y\) when feasible. A practical kernel may instead take a correlated move, but must preserve the uniform conditional law. Neal's interval-selection and shrinkage rules serve this invariant.[1]
- Marginal target and convergence boundary. Projecting the correct joint stationary distribution onto \(x\) yields \(f/Z\). Finite-run accuracy still requires reachability and adequate mixing; the method cannot cross every support gap merely because its target is formally stationary.[1]
- Variant geometry. Elliptical slice sampling preserves the height-threshold idea but uses a Gaussian-prior ellipse and an angular bracket, not the scalar line bracket of Neal's basic method.[2]
What It Is Not¶
It is not a claim that every implementation samples uniformly and independently over the whole slice at each iteration. That is the ideal full conditional, not the generic practical transition. A stepping-out interval is constructed from the current point and may not contain all separated parts of \(S_y\). Shrinkage uses rejected proposals to restrict subsequent candidates. The resulting valid kernel can still leave the target invariant without being the ideal independent conditional draw.[1]
It is not simple rejection sampling of independent \(x\) values from a proposal envelope, nor is it ordinary Metropolis–Hastings with a fixed proposal and accept/reject ratio. Slice sampling uses an auxiliary height and membership above that height. It is not always tuning-free: scalar stepping out uses an initial width \(w\) and can cap expansion; poor choices can affect efficiency or even reachability in special supports.[1]
It is not guaranteed fast on multimodal or high-dimensional targets. Neal explicitly discusses slow moves between separated regions and possible nonirreducibility when density is zero in gaps. Elliptical slice sampling addresses a narrower latent-Gaussian structure and is not a generic replacement for every scalar or multivariate slice update.[1][2]
Scope of Application¶
For a univariate conditional density known up to scale, Neal's stepping-out/shrinkage method can update one parameter inside a larger Bayesian model. A width \(w\) gives an initial bracket scale; expansion can adapt to a wider local slice and shrinkage can make a too-wide bracket usable. The method is convenient where direct sampling from the full conditional is unavailable but density evaluations are cheap enough. However, an update that changes one coordinate at a time may mix slowly when variables are strongly dependent.[1]
The ideal auxiliary-variable formulation also supports multivariate slice kernels, but drawing uniformly from a high-dimensional slice is harder. Neal develops hyperrectangle and other update schemes precisely because a scalar interval procedure does not transfer automatically to many dimensions. The correctness of any new kernel is a question about preserving the appropriate slice distribution, not merely generating candidates above the threshold.[1]
Elliptical slice sampling is an important specialization for a latent vector with multivariate Gaussian prior and a likelihood. Murray, Adams and MacKay draw an auxiliary Gaussian direction, use a likelihood threshold and propose along the induced ellipse with angular-bracket shrinkage. Their method targets the posterior proportional to Gaussian prior times likelihood without the conventional scalar bracket-width tuning parameter. The Gaussian-prior assumption and likelihood-only threshold are constitutive to that variant.[2]
Clarity¶
Two different uses of “uniform” must be kept separate. The mathematical joint distribution is uniform under the graph of \(f\). Given a fixed \(y\), the ideal conditional \(x\mid y\) is uniform on the full \(S_y\). A practical bracket algorithm draws provisional points uniformly from its current interval, rejecting outside-slice points and shrinking the interval. These statements are not interchangeable; the last does not imply a fresh uniform sample over every global slice component.[1]
The height's numeric value scales with the arbitrary multiplier on \(f\): if \(f\) is multiplied by a positive constant, heights scale with it and the resulting \(x\) marginal remains the same. The algorithm therefore needs the target only up to normalization. By contrast, an invalid bracket rule cannot be excused by rescaling \(f\); detailed balance or another invariance argument is still required.[1]
Manages Complexity¶
The auxiliary height converts a difficult density-sampling problem into geometric questions about a level set: where is \(f(x)>y\), and how can a valid move be made within it? For scalar targets, stepping out and shrinkage let the bracket respond to local width rather than requiring a single proposal scale to match every part of the density. This can reduce hand-tuning compared with a fixed-width random walk, while retaining a visible conditional target.[1]
The compression hides costs if stated too casually. Each endpoint or rejected proposal needs a density evaluation; a broad interval may need many tests, while a narrow one may yield only local moves. Multiple modes or zero-density gaps can defeat exploration even when local slice transitions are valid. A useful report therefore states width and expansion rules, evaluations per accepted transition, and diagnostics of mixing or irreducibility rather than treating “adaptive” as a universal performance guarantee.[1]
Abstract Reasoning¶
First check that \(f\ge0\) and has a finite, nonzero integral on the intended state space. Define \(U\) and integrate over \(y\) at fixed \(x\): the vertical length is \(f(x)\), giving marginal \(f/Z\). This proves the target of the ideal joint construction. Next specify the actual horizontal transition. If it is ideal full-slice Gibbs, draw uniformly from all of \(S_y\); if it is stepping out and shrinkage, verify the interval-selection and acceptance rules that leave the slice law invariant. Never infer the latter merely because a proposed \(x'\) satisfies \(f(x')>y\).[1]
Finally inspect support and geometry. Can the chain move between separated high-density regions? Does coordinate-wise updating explore correlated variables? Does an elliptical Gaussian-prior method fit the actual model? These are distinct from proving one-step stationarity. In Neal's explicit example with two separated support intervals and fixed \(w=1\), a common bracket scheme fails irreducibility despite having the intended stationary law on its reachable pieces.[1][2]
Knowledge Transfer¶
The pattern transfers from a scalar Bayesian full conditional to a latent-Gaussian model: introduce an auxiliary threshold, define acceptable states above it, and use a reversible or otherwise slice-invariant update. But the horizontal geometry changes. Neal's scalar method searches an interval on a line; Murray and colleagues' elliptical method traverses a Gaussian-prior ellipse and thresholds only the likelihood. Copying a scalar bracket rule directly into the ellipse would not preserve the specialized algorithm's proof.[1][2]
The live Monte Carlo Simulation prime supplies the broad idea of using random samples for distributional calculation. Slice Sampling adds a particular under-graph augmentation and transition constraint. An unrelated “slice” in graphics or data partitioning shares vocabulary, not this MCMC identity.
Examples¶
Canonical: a Gaussian-shaped target at one height¶
Let \(f(x)=\exp(-x^2/2)\), with unknown normalizing constant ignored. For an illustrative height \(y=\exp(-1/2)\) and a current point inside the slice, \(f(x)>y\) precisely when \(-1<x<1\). An ideal horizontal conditional update at this fixed height is uniform on \((-1,1)\). If height is repeatedly redrawn uniformly below the current density and the horizontal conditional is updated validly, the stationary \(x\) marginal is proportional to \(\exp(-x^2/2)\). The exact chosen height is a mathematical illustration, not an event with positive probability under a continuous draw.[1]
Mapped back: target = Gaussian-shaped \(f\); height = illustrative \(e^{-1/2}\) below the current \(f(x)\); slice = connected interval \((-1,1)\); horizontal update = ideal full-slice conditional uniform draw for this example; marginal = normalized Gaussian shape over repeated valid updates; variant geometry = absent.
Applied: Neal's stepping-out and shrinkage update¶
Neal's Figure 1 starts from \(x_0\), draws \(y\) uniformly below \(f(x_0)\), randomly places a width-\(w\) interval around \(x_0\), expands its ends until they fall outside the selected slice, then proposes from the interval. Rejected outside-slice points shrink the bracket around \(x_0\) until a valid \(x_1\) is obtained. The algorithm's interval and acceptance rules, not an assumption that the interval equals the whole slice, support its invariant target. Disconnected support requires a separate reachability check.[1]
Mapped back: target = evaluable unnormalized \(f\); height = \(y\sim\operatorname{Uniform}(0,f(x_0))\); slice = \(\{x:f(x)>y\}\), possibly wider than one bracket; horizontal update = Neal's stepped-out interval and shrinkage kernel; marginal = \(f/Z\) when valid and adequately explored; variant geometry = one-dimensional bracket.
Applied: latent Gaussian posterior with an elliptical slice¶
Murray, Adams and MacKay consider a latent vector \(\mathbf f\) with multivariate Gaussian prior and likelihood \(L(\mathbf f)\). They draw an auxiliary Gaussian vector \(\nu\), set a likelihood threshold, and propose points \(\mathbf f\cos\theta+\nu\sin\theta\) along an ellipse; rejected angles shrink a bracket until the proposal passes the threshold. Their construction preserves the posterior proportional to Gaussian prior times likelihood. It is a slice sampler because the auxiliary threshold governs acceptance, but it does not sample a new vector uniformly from an arbitrary full Euclidean level set.[2]
Mapped back: target = Gaussian prior times likelihood; height = likelihood-based threshold conditional on current latent state; slice = accepted angles on a prior-preserving ellipse; horizontal update = angular bracket shrinkage; marginal = latent posterior; variant geometry = Gaussian-prior ellipse rather than a scalar interval.
Boundary: disconnected support¶
Neal describes a density uniform on \([0,0.2]\cup[1.5,1.6]\) with a fixed step-out width \(w=1\). That method can fail irreducibility: the chain need not pass from one support interval to the other. The full ideal slice includes both intervals, so it would be false to describe this practical update as drawing uniformly from that whole disconnected set.[1]
Structural Tensions¶
T1 — Larger moves versus evaluation cost. Expanding a bracket can expose a wider local slice and permit a farther move, but each expansion and rejected proposal consumes density evaluations; a narrow bracket costs less per update yet can leave the chain moving locally. Diagnostic: How much effective movement is obtained per density evaluation under the actual target and chosen width/expansion rule?[1]
T2 — Simple local transition versus global reachability. Stepping out and shrinkage are attractive because their local implementation is simple and adapts to width, but disconnected support or zero-density gaps can make some components unreachable. A globally reaching kernel may fix that but needs a more demanding correctness and efficiency argument. Diagnostic: Can the implemented chain demonstrably reach every support component relevant to the inference?[1]
Structural–Framed Character¶
Slice Sampling is structural with a methodological frame. The under-graph marginal and invariant-kernel conditions are mathematical; an implementation either satisfies them or does not. Evaluative weight enters when deciding acceptable computation time and effective sample size. Human-practice dependence enters through bracket width, expansion cap, coordinate representation and convergence diagnostics. Institutional origin in MCMC research names and packages these algorithms but does not determine their stationary laws.[1]
Vocabulary travel to elliptical latent-Gaussian sampling is justified by a proven auxiliary-threshold construction; travel to unrelated “slicing” metaphors is not. Import versus recognition means verifying the actual height, slice and invariant update before importing performance expectations from another target. Its character: a precise probability-sampling family whose practical validity and mixing depend on implementation details.
Structural Core vs. Domain Accent¶
The structural core is an auxiliary-variable representation: make a target marginal by adjoining a height beneath a graph, then alternate a vertical conditional with a horizontal transition preserving uniform mass on the selected level set. The live Monte Carlo Simulation prime is a broad parent but does not itself encode this under-graph mechanism. A fully general prime about auxiliary levels and invariant conditional kernels across unrelated domains would be an unadmitted future-prime question, not an automatic abstraction from this one statistical family.[1]
The domain accent is integrable densities, conditional probability, Markov-chain stationarity and practical algorithms such as interval stepping out, shrinkage or a Gaussian-prior ellipse. Remove the probabilistic height or the invariant slice update and this named MCMC identity disappears. The full-slice ideal should not be silently substituted for an interval-based kernel just because both can have the same stationary marginal.[1][2]
Instantiates / Related Primes¶
This entry is a kind of Monte Carlo Simulation.
The broader abstraction is live Monte Carlo Simulation. Slice Sampling uses a random Markov chain to represent a target distribution, adding the specific auxiliary-height and slice-invariant-transition machinery. The parent is broader; it includes many random simulation methods without level-set slices.
Umbrella sampling is a neighboring statistical-physics method that biases windows and reweights; it does not supply this construction's genus. Thompson Sampling uses posterior samples to select actions, a different decision procedure. Elliptical Slice Sampling is a specialized member of this method family with a Gaussian-prior carrier and angular geometry, even though its likelihood threshold differs from the simplest scalar full-density formulation.[2]
Relationships to Other Abstractions¶
Current abstraction Slice Sampling Domain-specific
Parents (1) — more general patterns this builds on
-
Slice Sampling is a kind of Monte Carlo Simulation Prime
Slice sampling is a Markov-chain Monte Carlo construction with an auxiliary height and slice-invariant update.The live Monte Carlo Simulation prime broadly uses randomized samples to represent a distribution or approximate an expectation. Slice Sampling is a strict MCMC specialization: it augments an unnormalized density with a vertical height, forms a level-set slice and uses a valid horizontal kernel targeting the correct marginal.
Hierarchy paths (4) — routes to 4 parentless roots
- Slice Sampling → Monte Carlo Simulation → Approximation → Representation → Abstraction
- Slice Sampling → Monte Carlo Simulation → Iteration
- Slice Sampling → Monte Carlo Simulation → Probability → Measure → Set and Membership
- Slice Sampling → Monte Carlo Simulation → Probability → Measure → Aggregation → Micro Macro Linkage
Neighborhood in Abstraction Space¶
Slice Sampling sits in a sparse region of the domain-specific corpus (81st percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Numerical Root-Finding & Quadrature Methods (7 abstractions)
Nearest neighbors
- Epigraph — 0.85
- Space-Filling Curve — 0.83
- Ridders' Method — 0.82
- Jensen's Inequality — 0.82
- Interval Contractor — 0.81
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Independent full-slice sampling in every update: an ideal conditional special case, not a universal property of practical bracket algorithms.[1]
- Uniform proposal within a bracket: provisional candidates may be uniform on the current bracket even though accepted states are not independent uniform draws from a disconnected global slice.
- Rejection sampling: independent proposals against an envelope, rather than an auxiliary-height Markov transition.
- Metropolis–Hastings alone: a different general acceptance mechanism; slice sampling uses height membership and valid slice transitions.
- Guaranteed irreducibility or fast mixing: Neal's zero-density-gap example rules this out.[1]
- Zero tuning for every variant: scalar stepping out has width and optional cap, while the cited elliptical method removes free tuning in its stated formulation.[1][2]
- Any data “slice”: a partition of a table or image need not be a level set of a target density.
References¶
[1] Radford M. Neal, “Slice Sampling”, Annals of Statistics 31(3), 2003, original paper and rejoinder, especially §3 pp.710–712, §4 pp.712–720 and rejoinder pp.759–760. The paper distinguishes ideal full-slice conditional sampling from invariant practical updates and gives the disconnected-support irreducibility counterexample. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r ↩s ↩t ↩u ↩v ↩w ↩x ↩y ↩z ↩27 ↩28 ↩29 ↩30 ↩31
[2] Iain Murray, Ryan Prescott Adams and David J. C. MacKay, “Elliptical Slice Sampling”, AISTATS 2010, original proceedings paper, §2 and Figure 2, pp.541–545; Gaussian-prior likelihood-threshold algorithm and angular-bracket shrinkage. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i