Blackwell–Girshick Equation¶
A two-term identity separating the variance of an independent-count random sum into mark-size and count-uncertainty contributions.
Core Idea¶
The Blackwell–Girshick equation gives the variance of a sum whose number of terms is itself random. Let \(N\) take values in \(\{0,1,2,\ldots\}\); let \(X_1,X_2,\ldots\) be independent and identically distributed random variables independent of \(N\); assume \(N\) and \(X_1\) have finite second moments. With \(S=\sum_{i=1}^{N}X_i\) and \(S=0\) when \(N=0\), write \(\mu=\mathbb E[X_1]\) and \(\sigma^2=\operatorname{Var}(X_1)\). Then
The first term is variability of the marks within a given count, averaged over possible counts. The second is variability of the conditional mean \(N\mu\) as the count changes. Conditioning shows the distinction exactly: \(\mathbb E[S\mid N]=N\mu\) and \(\operatorname{Var}(S\mid N)=N\sigma^2\), so the law of total variance recombines them with no residual term under these hypotheses.[1][2]
This is a reusable equation family, not one numerical result. \(N\) may have many count distributions, and \(X_i\) may have many mark distributions, as long as the assumptions and moments hold. A Poisson count is a familiar special case: when \(N\sim\operatorname{Pois}(\lambda)\), mean and variance of \(N\) both equal \(\lambda\), giving \(\operatorname{Var}(S)=\lambda\mathbb E[X_1^2]\). Poisson is not part of the general equation's definition.[2][3]
Structural Signature¶
Sig role-phrases: random count — independent iid marks — indexed random sum — conditional mark variance — count-variance contribution — finite-moment and independence boundary.
- Random count \(N\): a nonnegative integer variable selects how many marks enter. Its mean weights average mark variance, while its variance controls the changing-total term.
- Independent iid marks \(X_i\): each has the same \(\mu\) and \(\sigma^2\), the marks are mutually independent, and the entire sequence is independent of \(N\). These are sufficient conditions for the simple conditional-moment identities, not a claim that no mathematical weakening exists.[1]
- Indexed random sum \(S\): the target is addition of \(X_1\) through \(X_N\), not a product, maximum, fixed-length average, or tail indicator. The empty sum is zero.
- Conditional mark-variance component: \(\mathbb E[N]\sigma^2\) comes from \(\mathbb E[\operatorname{Var}(S\mid N)]\). It can persist even if \(N\) is fixed.
- Count-variance component: \(\operatorname{Var}(N)\mu^2\) comes from \(\operatorname{Var}(\mathbb E[S\mid N])\). It vanishes if the count is fixed or the marks have mean zero.
- Moment and independence boundary: finite second moments keep the displayed variances defined; violating independence can add covariance or count-conditioned terms.[1][2]
The identity test is the conjunction of this carrier, these conditional moments, and the displayed recomposition. Merely seeing a random sum is not enough.
What It Is Not¶
- Not Wald's mean equation. \(\mathbb E[S]=\mathbb E[N]\mu\) concerns the first moment. The Blackwell–Girshick equation concerns the second central moment and includes an additional count-variance component.[3]
- Not restricted to compound Poisson processes. Those impose a Poisson count process; the present equality holds for eligible non-Poisson counts too. For Poisson \(N\), the formula simplifies to \(\lambda\mathbb E[X_1^2]\).[2]
- Not a universal formula for any random sum. Correlated marks or marks whose distribution changes with \(N\) can require extra terms. Cohen derives a wider formula and recovers this equality as a special case; that broader result is not an instance of the same two-term identity without its restrictions.[1]
- Not a probability of a large loss or long queue. Variance alone does not determine a quantile or exceedance probability; a subsequent approximation requires its own model and checks.[3]
- Not the law of total variance itself. That general conditioning identity supplies the proof; the named equation is its specialized evaluation for an independent count and iid marks.
The closest near-miss is \(S=\sum_{i=1}^{N}X_i\) with count-dependent severities. It has the same syntax, but then \(\mathbb E[S\mid N]\) need not equal \(N\mu\) for one fixed \(\mu\), so the second component cannot simply be \(\operatorname{Var}(N)\mu^2\).[1]
Scope of Application¶
The formula is useful where event count and event size can be modeled separately: aggregate claims, cumulative failure costs, inventory demand, or other compound quantities. The model must state a time or population window, what counts as an event, how the individual mark is measured, whether zero counts are possible, and whether the marks and count satisfy the required independence and finite-moment assumptions. Otherwise the decomposition may be arithmetically neat but scientifically misplaced.[3][4]
No Poisson hypothesis is needed for the general equality. If \(N\) is deterministically \(n\), the count term is zero and ordinary fixed-sum variance \(n\sigma^2\) remains. If marks equal a constant \(c\), mark variance is zero and \(S=cN\), so the remaining variance is \(c^2\operatorname{Var}(N)\). These limits show what each component actually measures.[2]
In data work, estimated count and mark moments may be uncertain or count-linked. The equation is an exact identity inside its stochastic model, not a guarantee that observed claims or failures were generated by that model. The model check and the variance calculation are separate tasks.
Clarity¶
The phrase “variance of a random total” hides two different sources of variability. A high total can arise because there are many events or because individual events are large. Conditioning on the count makes that split explicit without pretending the two terms are estimated by identical evidence. \(\mathbb E[N]\sigma^2\) has the units of the total squared; so does \(\operatorname{Var}(N)\mu^2\). Adding \(\operatorname{Var}(N)\sigma^2\) instead of \(\operatorname{Var}(N)\mu^2\) would confuse variability of event sizes with variability of the number of their means.[1][2]
The equation also clarifies a modeling boundary: it is not enough for every mark to have the same unconditional distribution if marks are correlated with the count. One must verify the conditional moments that make the two-term derivation work.
Manages Complexity¶
A compound distribution can be difficult to derive in full. If the immediate question is dispersion, the Blackwell–Girshick equation compresses the needed information to four quantities: \(\mathbb E[N]\), \(\operatorname{Var}(N)\), \(\mathbb E[X_1]\), and \(\operatorname{Var}(X_1)\). It thereby separates changes in event frequency from changes in mark severity. Daniel's health-claim example computes its aggregate variance without first deriving the full compound-Poisson density.[3]
That compression deliberately loses distributional shape. Different count and mark laws can agree on these moments but disagree in tail behavior. A variance calculation is a checkpoint for more detailed modeling, not a replacement for a tail model when the decision depends on extremes.[3]
Abstract Reasoning¶
First condition on \(N=n\). Independence and identical distribution give \(\mathbb E[S\mid N=n]=n\mu\) and \(\operatorname{Var}(S\mid N=n)=n\sigma^2\); the empty-sum convention makes these valid at \(n=0\). Next average conditional variance over \(N\), obtaining \(\mathbb E[N]\sigma^2\). Then vary the conditional mean \(N\mu\), obtaining \(\operatorname{Var}(N)\mu^2\). The law of total variance adds the two; no approximation or large-sample limit is involved.[2]
This derivation supplies a failure diagnostic. If mark means change with \(n\), the conditional mean becomes \(n\mu_n\); if marks co-vary, conditional variance contains their cross-covariances. Either change defeats the simple formula even though \(S\) remains a random sum. A replacement must recalculate those conditional moments, not merely relabel the old terms.[1]
Knowledge Transfer¶
Insurance and reliability engineering use unlike observables, yet the mapping is literal. For health claims, \(N\) is a count of filed claims and \(X_i\) is dollars per claim. For equipment reliability, \(N(t)\) counts failures and \(V_i\) is cost per failure. In both, the total is an indexed sum and the same two conditional-moment calculations apply when the stipulated independent-mark model is warranted.[3][4]
The safe transfer is therefore a proof obligation, not a metaphor: map count, marks, total, moments and independence; verify the units; then recompute both terms. Cohen's examples in traffic, ecology, trading and tornado claims explain why wider problems can look similar yet fail this transfer because mark size may depend on event count or marks may correlate.[1]
Examples¶
Ten days of health-insurance claims. In Daniel's worked example, claims arrive with Poisson rate 20 per day, so over ten days \(\mathbb E[N]=\operatorname{Var}(N)=200\). Independent exponential severities have mean $500$ dollars and variance \(500^2\) dollars squared. The mark component is \(200(500^2)=50{,}000{,}000\) dollars squared; the count component is another \(200(500^2)=50{,}000{,}000\). The total variance is $100{,}000{,}000$ dollars squared. The numerical balance of terms is special to these assumptions, not a universal half-and-half rule.[3] Mapped back: random count \(N\) = ten-day claims; independent iid marks \(X_i\) = individual exponential claim amounts; indexed random sum \(S\) = total claim dollars; conditional mark-variance component = first $50$ million; count-variance component = second $50$ million; moment and independence boundary = finite exponential moments with marks and count independent.
Accumulated equipment-failure costs. Vatn's compound homogeneous Poisson model counts failures by time \(t\) and adds iid costs \(V_i\). If its event rate is \(\lambda\), then \(\mathbb E[N(t)]=\operatorname{Var}(N(t))=\lambda t\). The Blackwell–Girshick equation gives \(\operatorname{Var}(Z(t))=\lambda t\operatorname{Var}(V)+\lambda t(\mathbb E[V])^2=\lambda t\mathbb E[V^2]\). This is a symbolic model calculation, not a claim that real failure costs are always independent.[4] Mapped back: random count \(N\) = equipment failures by \(t\); independent iid marks \(X_i\) = \(V_i\) costs; indexed random sum \(S\) = \(Z(t)\) accumulated cost; conditional mark-variance component = \(\lambda t\operatorname{Var}(V)\); count-variance component = \(\lambda t(\mathbb E[V])^2\); moment and independence boundary = finite cost second moment and independent count/marks.
Counterexample to automatic transfer. If high-event-count periods also have systematically larger marks, conditional mark means depend on \(N\). The same summation notation still describes the total, but the displayed two-term equality is not licensed. The difference is a model assumption, not a change of application label.[1]
Structural Tensions¶
Separable interpretability versus dependence realism. Independent count and marks yield a clean size-versus-number variance split; allowing count-linked severities or mutually correlated marks can better fit a real system, but then the clean equality is usually lost and additional conditional moments must be estimated. Diagnostic: Do mark means, variances and cross-correlations stay stable when the observed count changes? If not, use a model whose conditional calculation includes those dependencies.[1]
Tractable variance summary versus tail fidelity. The four-moment expression makes aggregate spread easy to compare across scenarios; it cannot identify an exceedance probability. A full compound distribution or separately justified approximation carries more tail information but needs more assumptions and computation. Diagnostic: Is the decision about squared dispersion, or about the chance of a threshold crossing? The latter cannot be read directly from this equation.[3]
Structural–Framed Character¶
The identity is structural within probability theory. Its count/mark roles transfer literally from insurance to engineering; the equality is nonnormative and does not rank outcomes as desirable; it can hold in a stochastic model whether or not a human chooses to use it; no institution or legal rule constitutes it; and recognition requires checking stated conditional moments rather than a verbal resemblance to a “random total.” Its character is structural, but domain-specific to random-sum probability calculus, not a free-standing prime principle of every kind of variability. A physical dataset may need statistical model checking before the mathematical structure is an appropriate representation.[1][4]
Structural Core vs. Domain Accent¶
The portable skeleton is an exact decomposition and recomposition: one whole quantity equals contributions differentiated by source. Live prime Decomposition already captures that generic move. The domain accent is the precise random-count sum and its conditioning: \(\mathbb E[N]\operatorname{Var}(X_1)\) plus \(\operatorname{Var}(N)(\mathbb E[X_1])^2\) under specified independence and finite moments. Remove the stochastic roles and one has only a generic two-part split, not this named equation.[1][2]
Even the law of total variance is broader. It becomes the Blackwell–Girshick form only after conditional mean and variance are evaluated for the independent-count iid-mark model. Conversely, replacing Poisson with another eligible count distribution preserves the named equation; Poisson is an application, not the accent.
Instantiates / Related Primes¶
This entry is a kind of Decomposition.
The proposed typed parent Decomposition is justified by the exact two-part reconstruction of \(\operatorname{Var}(S)\). Covariance is related because variance is self-covariance, but the present entry is a specialized equality for a random sum, not a general covariance operator. Aggregation describes formation of the sum, while the named equation concerns its variance; that thematic connection is not asserted as a second strict parent.
Relationships to Other Abstractions¶
Current abstraction Blackwell–Girshick Equation Domain-specific
Parents (1) — more general patterns this builds on
-
Blackwell–Girshick Equation is a kind of Decomposition Prime
A random-sum variance is exactly decomposed into mark and count components.Under the stated count/mark hypotheses, total variance is reconstructed without remainder from expected conditional variance and variance of the conditional mean. This instantiates the live Decomposition prime's whole-to-parts and recomposition signature. The two exact probability terms and finite-moment assumptions make the candidate strictly narrower. It is not asserted as a subtype of a Poisson process, generic variance measure, or central-limit theorem.
Hierarchy path (1) — routes to 1 parentless root
- Blackwell–Girshick Equation → Decomposition
Neighborhood in Abstraction Space¶
Blackwell–Girshick Equation sits in a sparse region of the domain-specific corpus (80th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Foundations of Probability & Inference (29 abstractions)
Nearest neighbors
- Kolmogorov's Three-Series Theorem — 0.84
- Yule–Simon Distribution — 0.83
- Average Order of an Arithmetic Function — 0.82
- Large Set (Combinatorics) — 0.82
- Random Variable — 0.82
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
Variance: the second central moment of any square-integrable random variable, not a particular formula for an indexed random sum. Compound Poisson process: a process class in which the count is Poisson; the equation is valid under broader count laws. Law of total covariance / variance: a conditioning theorem from which this case can be obtained; the named form inserts special conditional moments. Wald's equation: the companion mean identity. Central limit theorem: an approximation to distributional shape under further conditions, not the exact variance equality. A dependent random-sum generalization: may add covariance and count-mark dependence terms and must not inherit the simple two-term right-hand side.[1][2]
References¶
[1] Joel E. Cohen, “Sum of a Random Number of Correlated Random Variables that Depend on the Number of Summands”, The American Statistician 73(1) (2019), especially Introduction, §4.1 Eq. (13), and §5. This original research paper prints the classical equation and attributes it to Blackwell and Girshick (1947, theorem 2); it also derives a broader correlated/count-dependent result. The 1947 theorem text was not directly inspected here. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m
[2] Rick Durrett, Probability: Theory and Examples, §2.3 “Compound Poisson Processes” excerpt, random-sum variance theorem and proof. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i
[3] James W. Daniel, Poisson Processes actuarial notes, hosted by the Casualty Actuarial Society, Definition 1.8, Fact 1.11, Example 1.12 (PDF pp. 6–8). registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i
[4] Jørn Vatn, Counting Processes, NTNU course notes, “Compound HPPs” (PDF pp. 3–4). registry ↩a ↩b ↩c ↩d