Skip to content

Dirichlet negative multinomial distribution

A multivariate count law formed by Dirichlet-mixing negative-multinomial category probabilities.

Core Idea

The Dirichlet negative multinomial distribution (DNM) is a compound multivariate count law. Start with one success category and several failure categories; a negative multinomial counts failures before a specified number of successes. Let the category-probability vector vary by unit according to a Dirichlet distribution, then integrate those probabilities out. The resulting law assigns joint probability to nonnegative failure-count vectors and can represent dependence and overdispersion. It is not the Dirichlet multinomial, which conditions on a fixed total number of trials.

Farewell and Farewell present the construction in equation (2.1), give an equivalent Gamma/Poisson random-effects account, and use a regression form on observed oncology trial-recruitment counts. The actual application follows 22 multidisciplinary teams over three or four six-month periods. In that longitudinal setting the categorical stopping story is a mathematical construction, not a literal account of hospital recruitment; the equivalent random-effects interpretation is more natural. The distribution exists with positive parameters, but its mean requires α0>1 and covariance α0>2.

How would you explain it like I'm…

Mystery Marble Bags

Imagine each child has a bag of marbles with a different mix of colors, and nobody knows each bag's mix. Each child pulls marbles until they get a set number of gold ones, counting the other colors along the way. The Dirichlet negative multinomial is a math rule for guessing how those color counts come out when every bag's mix is a mystery.

Counting Misses, Mixed Chances

The Dirichlet negative multinomial is a rule in statistics for predicting several counts at once. Picture drawing again and again, where each draw is either a 'success' or one of several kinds of 'failure', and you stop after a set number of successes; you then count each kind of failure. Now suppose the chances of each outcome are different for each person or team, and they vary in a random way. Averaging over all those possible chances gives this distribution. It lets the counts be more spread out and more linked together than simple models allow. It's different from a similar model where the total number of tries is fixed in advance.

Dirichlet-Mixed Negative Multinomial

The Dirichlet negative multinomial (DNM) distribution is a compound distribution for vectors of counts. Start with a negative multinomial: each trial results in either success or one of several failure categories, and you count failures of each type before reaching a fixed number of successes. Now let the vector of category probabilities differ from unit to unit, following a Dirichlet distribution, and average (integrate) over it. The result gives joint probabilities for vectors of nonnegative failure counts and can model overdispersion (more variability than the basic model) and dependence between categories. It differs from the Dirichlet multinomial, which fixes the total number of trials in advance. Farewell and Farewell used it with an equivalent Gamma/Poisson random-effects version to model cancer trial recruitment counts from 22 medical teams over several six-month periods, where the random-effects reading is more natural than the literal stopping story. Its mean exists only if the parameter α₀ is greater than 1, and its covariance only if α₀ is greater than 2.

 

The Dirichlet negative multinomial distribution is a compound multivariate count law obtained by mixing a negative multinomial over a Dirichlet distribution on its category-probability vector. The negative multinomial has one success category and several failure categories and counts failures of each type before a specified number of successes; letting the probability vector vary across units according to a Dirichlet and integrating it out yields a joint law on nonnegative failure-count vectors that captures overdispersion and inter-category dependence. It must be distinguished from the Dirichlet multinomial, which conditions on a fixed number of trials. Farewell and Farewell present the construction, give an equivalent Gamma/Poisson random-effects representation, and fit a regression form to oncology trial-recruitment counts from 22 multidisciplinary teams observed over three or four six-month periods. In that longitudinal application the categorical stopping story is a mathematical device, and the random-effects interpretation is the more natural one. The distribution exists for positive parameters, but its mean requires α₀ > 1 and its covariance α₀ > 2.

Structural Signature

Sig role-phrases:

  • Categorical trial outcomes — Each trial has one success class and m distinct failure classes. It is constitutive. Counterfactual: A single undifferentiated count lacks the multivariate category structure unless m=1 is treated as a special case.
  • Success stopping parameter — y0>0 specifies when counting stops rather than fixing a total number of all trials. It is constitutive. Counterfactual: Fixed-total sampling belongs to the Dirichlet multinomial construction instead.
  • Failure-count vector — Records nonnegative counts (y1,…,ym) accumulated before the stopping condition. It is constitutive. Counterfactual: Continuous measurements are outside this discrete count support.
  • Dirichlet probability mixing — Unit-level categorical probabilities vary according to positive α parameters and are integrated out. It is constitutive. Counterfactual: A single fixed probability vector gives an ordinary negative multinomial, not this compound law.
  • Joint probability law — Assigns normalized mass to each admissible failure-count vector, inducing correlation and overdispersion. It is constitutive. Counterfactual: A list of empirical counts without a probability model is not the distribution.
  • Moment and modeling regime — Distinguishes law existence from finite-mean/covariance conditions and from any regression use. It is central. Counterfactual: Using formulas for covariance when α0≤2 exceeds their existence condition.

What It Is Not

  • Not Dirichlet multinomial. That compound law fixes total trials, while DNM stops at a success parameter and allows variable failure totals.
  • Not ordinary negative multinomial. Fixed category probabilities lack the additional Dirichlet mixing.
  • Not any correlated count model. The probability law has a specific kernel and mixing structure.
  • Not necessarily finite-moment. The law can exist when mean or covariance does not.
  • Closest near-miss. The parameter y0 may be extended to positive real values mathematically; the literal trials story is an intuition, not a restriction on all analytic parameterizations.

Scope of Application

  • Biostatistics. Model correlated repeated counts from clinical research teams.
  • Quantitative marketing. Represent varying purchase counts across product categories where the generative assumptions fit.
  • Count regression. Use the hierarchical representation to separate within- and between-unit variation.
  • Probability theory. Study compound laws, marginals, moments, and heavy tails under parameter constraints.

Clarity

DNM is what results when a negative multinomial's category probabilities themselves vary according to a Dirichlet law. It models a vector of nonnegative counts, not a fixed-total allocation or just a histogram. The mathematical stopping-trials construction can support other applications, including repeated hospital counts, without asserting that patients are literally 'failures' in a trial.

Manages Complexity

The distribution packages a high-dimensional joint count law into a parameterized mixture that can express extra variation and dependence. The convenience comes with interpretive costs: the original success/failure story may be artificial for repeated measures, moment formulas require α0 thresholds, and fitted correlations depend on model assumptions. The alternative hierarchical representation makes these costs easier to inspect.

Abstract Reasoning

  1. Identify a vector of nonnegative category or repeated counts.
  2. Specify the negative-multinomial success-stopping kernel and positive parameters.
  3. Place a Dirichlet law on the categorical probability vector and marginalize it.
  4. Check that fixed-total or fixed-probability alternatives do not describe the same process.
  5. Verify α0 thresholds before reporting finite moments.
  6. When fitting data, state whether the stopping construction is literal or only an equivalent model representation.

Knowledge Transfer

The joint law can model shopping baskets, clinical repeated counts, and other overdispersed correlated vectors when the parameter and support assumptions hold. The mathematical identity transfers; calling every heterogeneous count vector DNM without checking fit, stopping/mixing structure, and moment regime does not.

Examples

Canonical

Farewell and Farewell §2 specify one success outcome and m failure outcomes in repeated categorical trials. Count the failure categories until y0 successes, assign the categorical probability vector a Dirichlet(α0,…,αm) law, then integrate that vector out of the negative-multinomial mass function. Their equation (2.1) is the resulting DNM joint mass on (y1,…,ym); y0>0 and αj>0 define the law, while stronger α0 thresholds govern finite moments.

Mapped back: Categorical trial outcomes → one success and m distinct failures; Success stopping parameter → positive y0 success count/analytic parameter; Failure-count vector → nonnegative (y1,…,ym); Dirichlet probability mixing → positive α0,…,αm probability-vector law; Joint probability law → equation (2.1) after integration; Moment and modeling regime → α0>1 for finite mean, >2 for finite covariance.

Applied / In Practice

Farewell and Farewell fit a DNM regression to observed clinical-trial recruitment counts: patients approached by 12 oncology multidisciplinary teams in three successive six-month periods and 10 further teams in four periods. They compare Poisson, negative-binomial, negative-multinomial, GEE, DNM, and GLMM fits and report DNM dispersion/covariate estimates. This is an actual statistical application, not proof that DNM is uniquely best for all recruitment data.

Mapped back: Categorical trial outcomes → repeated period-specific count categories in the adapted regression representation; Success stopping parameter → latent DNM parameterization, not literally observed successful patients; Failure-count vector → patients approached per period for each team; Dirichlet probability mixing → equivalent hierarchical random-effects/mixing formulation; Joint probability law → fitted DNM likelihood across each team's count vector; Moment and modeling regime → finite-moment regression parameters and model comparisons.

Structural Tensions

T1 — Fixed-Probability Simplicity versus Heterogeneity Fit. A single negative-multinomial probability vector is parsimonious but may miss unit-level variation represented by Dirichlet mixing.

Diagnostic: Is extra dispersion supported by the data?

T2 — Joint Dependence versus Interpretability. A multivariate DNM can fit correlated counts, while its original stopping-trials story may be unnatural for longitudinal hospital periods.

Diagnostic: Is an equivalent random-effects parameterization more interpretable?

T3 — Flexible Tails versus Moment Existence. Parameters that admit heavy overdispersion can also lack finite mean or covariance under the α0 thresholds.

Diagnostic: Which summaries does inference require?

Structural–Framed Character

The approved DAG parent is Probability Distribution: a normalized joint law over nonnegative count vectors. This child obtains that law by mixing a success-stopped negative-multinomial kernel over a Dirichlet category-probability vector; it is not the fixed-total Dirichlet multinomial.

Evaluative weight: Low in the law's identity; suitability for a dataset is a separate model-checking judgment. Human-practice-bound: Low formally, although analysts choose parameters and whether the generative assumptions match an application. Institutional origin: Statistical practice names the family; no institutional convention can replace its support and probability normalization. Vocabulary travels: The law can model differently labeled count categories when its stopping and mixing roles hold; mere correlated counts do not inherit the name. Import versus recognize: A new application can be recognized as DNM only when its law or justified model has those roles; calling heterogeneous counts DNM imports a model assumption.

Its character: A formal statistical species of probability distribution with transferable mathematics and strict generative conditions.

Structural Core vs. Domain Accent

Skeletal core. A specified probability law distributes mass over possible count vectors. Domain-bound accent. Negative-multinomial stopping, Dirichlet probability mixing, positive parameters, and moment thresholds define this named statistical family. Transfer boundary. Merely correlated empirical counts or a fixed-total Dirichlet multinomial do not instantiate the same generative law.

This entry is a kind of Probability Distribution.

  • Strict parent: Probability Distribution. The live catalog node is a normalized law over a structured outcome space; DNM narrows it to a named compound discrete family on nonnegative count vectors.

  • Neighbor: negative multinomial. It supplies the fixed-probability kernel, before Dirichlet mixing adds heterogeneity.

Relationships to Other Abstractions

Local relationship map for Dirichlet negative multinomial distributionParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Dirichlet negative m…DOMAINDomain-specific abstraction: Probability Distribution — is a kind ofProbabilityDistributionDOMAIN

Current abstraction Dirichlet negative multinomial distribution Domain-specific

Parents (1) — more general patterns this builds on

  • Dirichlet negative multinomial distribution is a kind of Probability Distribution Domain-specific

    DNM is a named joint probability distribution on nonnegative count vectors formed by a specific Dirichlet mixture.

Hierarchy paths (5) — routes to 3 parentless roots

Neighborhood in Abstraction Space

Dirichlet negative multinomial distribution sits in a moderately populated region (55th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.

Family — Domain-Specific Indicators & Measurement Methods (26 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • Dirichlet multinomial distribution. Tell: Fixed-total compound multinomial, not success-stopped negative multinomial.
  • Negative multinomial distribution. Tell: Fixed categorical probabilities rather than Dirichlet-mixed probabilities.
  • Beta negative binomial. Tell: One-failure-category special case rather than the general vector law.
  • Observed count table. Tell: Data to which a distribution may be fitted, not the law itself.

References