Dirichlet negative multinomial distribution¶
A multivariate count law formed by Dirichlet-mixing negative-multinomial category probabilities.
Core Idea¶
The Dirichlet negative multinomial distribution (DNM) is a compound multivariate count law. Start with one success category and several failure categories; a negative multinomial counts failures before a specified number of successes. Let the category-probability vector vary by unit according to a Dirichlet distribution, then integrate those probabilities out. The resulting law assigns joint probability to nonnegative failure-count vectors and can represent dependence and overdispersion. It is not the Dirichlet multinomial, which conditions on a fixed total number of trials.
Farewell and Farewell present the construction in equation (2.1), give an equivalent Gamma/Poisson random-effects account, and use a regression form on observed oncology trial-recruitment counts. The actual application follows 22 multidisciplinary teams over three or four six-month periods. In that longitudinal setting the categorical stopping story is a mathematical construction, not a literal account of hospital recruitment; the equivalent random-effects interpretation is more natural. The distribution exists with positive parameters, but its mean requires α0>1 and covariance α0>2.
How would you explain it like I'm…
Mystery Marble Bags
Counting Misses, Mixed Chances
Dirichlet-Mixed Negative Multinomial
Structural Signature¶
Sig role-phrases:
- Categorical trial outcomes — Each trial has one success class and m distinct failure classes. It is constitutive. Counterfactual: A single undifferentiated count lacks the multivariate category structure unless m=1 is treated as a special case.
- Success stopping parameter — y0>0 specifies when counting stops rather than fixing a total number of all trials. It is constitutive. Counterfactual: Fixed-total sampling belongs to the Dirichlet multinomial construction instead.
- Failure-count vector — Records nonnegative counts (y1,…,ym) accumulated before the stopping condition. It is constitutive. Counterfactual: Continuous measurements are outside this discrete count support.
- Dirichlet probability mixing — Unit-level categorical probabilities vary according to positive α parameters and are integrated out. It is constitutive. Counterfactual: A single fixed probability vector gives an ordinary negative multinomial, not this compound law.
- Joint probability law — Assigns normalized mass to each admissible failure-count vector, inducing correlation and overdispersion. It is constitutive. Counterfactual: A list of empirical counts without a probability model is not the distribution.
- Moment and modeling regime — Distinguishes law existence from finite-mean/covariance conditions and from any regression use. It is central. Counterfactual: Using formulas for covariance when α0≤2 exceeds their existence condition.
What It Is Not¶
- Not Dirichlet multinomial. That compound law fixes total trials, while DNM stops at a success parameter and allows variable failure totals.
- Not ordinary negative multinomial. Fixed category probabilities lack the additional Dirichlet mixing.
- Not any correlated count model. The probability law has a specific kernel and mixing structure.
- Not necessarily finite-moment. The law can exist when mean or covariance does not.
- Closest near-miss. The parameter y0 may be extended to positive real values mathematically; the literal trials story is an intuition, not a restriction on all analytic parameterizations.
Scope of Application¶
- Biostatistics. Model correlated repeated counts from clinical research teams.
- Quantitative marketing. Represent varying purchase counts across product categories where the generative assumptions fit.
- Count regression. Use the hierarchical representation to separate within- and between-unit variation.
- Probability theory. Study compound laws, marginals, moments, and heavy tails under parameter constraints.
Clarity¶
DNM is what results when a negative multinomial's category probabilities themselves vary according to a Dirichlet law. It models a vector of nonnegative counts, not a fixed-total allocation or just a histogram. The mathematical stopping-trials construction can support other applications, including repeated hospital counts, without asserting that patients are literally 'failures' in a trial.
Manages Complexity¶
The distribution packages a high-dimensional joint count law into a parameterized mixture that can express extra variation and dependence. The convenience comes with interpretive costs: the original success/failure story may be artificial for repeated measures, moment formulas require α0 thresholds, and fitted correlations depend on model assumptions. The alternative hierarchical representation makes these costs easier to inspect.
Abstract Reasoning¶
- Identify a vector of nonnegative category or repeated counts.
- Specify the negative-multinomial success-stopping kernel and positive parameters.
- Place a Dirichlet law on the categorical probability vector and marginalize it.
- Check that fixed-total or fixed-probability alternatives do not describe the same process.
- Verify α0 thresholds before reporting finite moments.
- When fitting data, state whether the stopping construction is literal or only an equivalent model representation.
Knowledge Transfer¶
The joint law can model shopping baskets, clinical repeated counts, and other overdispersed correlated vectors when the parameter and support assumptions hold. The mathematical identity transfers; calling every heterogeneous count vector DNM without checking fit, stopping/mixing structure, and moment regime does not.
Examples¶
Canonical¶
Farewell and Farewell §2 specify one success outcome and m failure outcomes in repeated categorical trials. Count the failure categories until y0 successes, assign the categorical probability vector a Dirichlet(α0,…,αm) law, then integrate that vector out of the negative-multinomial mass function. Their equation (2.1) is the resulting DNM joint mass on (y1,…,ym); y0>0 and αj>0 define the law, while stronger α0 thresholds govern finite moments.
Mapped back: Categorical trial outcomes → one success and m distinct failures; Success stopping parameter → positive y0 success count/analytic parameter; Failure-count vector → nonnegative (y1,…,ym); Dirichlet probability mixing → positive α0,…,αm probability-vector law; Joint probability law → equation (2.1) after integration; Moment and modeling regime → α0>1 for finite mean, >2 for finite covariance.
Applied / In Practice¶
Farewell and Farewell fit a DNM regression to observed clinical-trial recruitment counts: patients approached by 12 oncology multidisciplinary teams in three successive six-month periods and 10 further teams in four periods. They compare Poisson, negative-binomial, negative-multinomial, GEE, DNM, and GLMM fits and report DNM dispersion/covariate estimates. This is an actual statistical application, not proof that DNM is uniquely best for all recruitment data.
Mapped back: Categorical trial outcomes → repeated period-specific count categories in the adapted regression representation; Success stopping parameter → latent DNM parameterization, not literally observed successful patients; Failure-count vector → patients approached per period for each team; Dirichlet probability mixing → equivalent hierarchical random-effects/mixing formulation; Joint probability law → fitted DNM likelihood across each team's count vector; Moment and modeling regime → finite-moment regression parameters and model comparisons.
Structural Tensions¶
T1 — Fixed-Probability Simplicity versus Heterogeneity Fit. A single negative-multinomial probability vector is parsimonious but may miss unit-level variation represented by Dirichlet mixing.
Diagnostic: Is extra dispersion supported by the data?
T2 — Joint Dependence versus Interpretability. A multivariate DNM can fit correlated counts, while its original stopping-trials story may be unnatural for longitudinal hospital periods.
Diagnostic: Is an equivalent random-effects parameterization more interpretable?
T3 — Flexible Tails versus Moment Existence. Parameters that admit heavy overdispersion can also lack finite mean or covariance under the α0 thresholds.
Diagnostic: Which summaries does inference require?
Structural–Framed Character¶
The approved DAG parent is Probability Distribution: a normalized joint law over nonnegative count vectors. This child obtains that law by mixing a success-stopped negative-multinomial kernel over a Dirichlet category-probability vector; it is not the fixed-total Dirichlet multinomial.
Evaluative weight: Low in the law's identity; suitability for a dataset is a separate model-checking judgment. Human-practice-bound: Low formally, although analysts choose parameters and whether the generative assumptions match an application. Institutional origin: Statistical practice names the family; no institutional convention can replace its support and probability normalization. Vocabulary travels: The law can model differently labeled count categories when its stopping and mixing roles hold; mere correlated counts do not inherit the name. Import versus recognize: A new application can be recognized as DNM only when its law or justified model has those roles; calling heterogeneous counts DNM imports a model assumption.
Its character: A formal statistical species of probability distribution with transferable mathematics and strict generative conditions.
Structural Core vs. Domain Accent¶
Skeletal core. A specified probability law distributes mass over possible count vectors. Domain-bound accent. Negative-multinomial stopping, Dirichlet probability mixing, positive parameters, and moment thresholds define this named statistical family. Transfer boundary. Merely correlated empirical counts or a fixed-total Dirichlet multinomial do not instantiate the same generative law.
Instantiates / Related Primes¶
This entry is a kind of Probability Distribution.
-
Strict parent: Probability Distribution. The live catalog node is a normalized law over a structured outcome space; DNM narrows it to a named compound discrete family on nonnegative count vectors.
-
Neighbor: negative multinomial. It supplies the fixed-probability kernel, before Dirichlet mixing adds heterogeneity.
Relationships to Other Abstractions¶
Current abstraction Dirichlet negative multinomial distribution Domain-specific
Parents (1) — more general patterns this builds on
-
Dirichlet negative multinomial distribution is a kind of Probability Distribution Domain-specific
DNM is a named joint probability distribution on nonnegative count vectors formed by a specific Dirichlet mixture.The live Probability Distribution node requires an outcome space, normalized probability law, named parametric family, and generative interpretation. DNM has nonnegative integer-vector outcomes, a normalized joint mass function, positive parameters, and an explicit Dirichlet-mixed negative-multinomial construction. It is thus a strict member of that distribution genus; the success-stopping and mixing details narrow the child and are not required of the parent.
Hierarchy paths (5) — routes to 3 parentless roots
- Dirichlet negative multinomial distribution → Probability Distribution → Random Variable → Function (Mapping)
- Dirichlet negative multinomial distribution → Probability Distribution → Probability → Measure → Set and Membership
- Dirichlet negative multinomial distribution → Probability Distribution → Probability → Measure → Aggregation → Micro Macro Linkage
- Dirichlet negative multinomial distribution → Probability Distribution → Random Variable → Probability → Measure → Set and Membership
- Dirichlet negative multinomial distribution → Probability Distribution → Random Variable → Probability → Measure → Aggregation → Micro Macro Linkage
Neighborhood in Abstraction Space¶
Dirichlet negative multinomial distribution sits in a moderately populated region (55th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Domain-Specific Indicators & Measurement Methods (26 abstractions)
Nearest neighbors
- Inferential Error — 0.86
- Funnel Chart — 0.86
- Experiment (Probability Theory) — 0.85
- Probability matching — 0.85
- Dichotomous Statistical Thinking — 0.85
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Dirichlet multinomial distribution. Tell: Fixed-total compound multinomial, not success-stopped negative multinomial.
- Negative multinomial distribution. Tell: Fixed categorical probabilities rather than Dirichlet-mixed probabilities.
- Beta negative binomial. Tell: One-failure-category special case rather than the general vector law.
- Observed count table. Tell: Data to which a distribution may be fitted, not the law itself.
References¶
- D. M. Farewell and V. T. Farewell, Dirichlet negative multinomial regression for overdispersed correlated count data (2013) — primary §2 compound-law construction and moment conditions, §§3–4 regression interpretation, and §6 real oncology recruitment analysis.