Skip to content

Binomial Proportion Confidence Interval

Bound a common binary-trial success probability from a success count using a stated interval rule and repeated-sampling coverage.

Version
v1 · 2026-10-04 · History
Domain-specific #
13714
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomains
Interval Estimation, Categorical Data → Experimental Design & Statistics
Aliases
Binomial confidence interval, Confidence interval for a binomial proportion

Core Idea

A binomial proportion confidence interval is a rule that turns k successes in n independent trials with a common binary-outcome probability p into lower and upper bounds for that unknown p. The rule is set before interpreting the observed count and is judged by coverage: across repetitions of the same binomial sampling process with fixed p, how often would the computed interval contain it? The observed sample proportion is p̂=k/n, but it is not itself an uncertainty interval.[1][2]

The name covers several constructions rather than one universal formula. The Wald interval plugs p̂ into a normal standard error and is simple but can cover badly for small samples or boundary probabilities. The Wilson interval solves an approximate score-test inequality for possible p values and stays inside the parameter range. Clopper–Pearson obtains endpoints from binomial-tail probabilities and guarantees at least nominal coverage under its model, usually paying with conservatism. Calling the latter exact refers to its binomial-tail calibration, not exact equality to the nominal coverage at every p: the binomial count is discrete.[1][2]

This domain-specific family is a specialized interval-estimation practice. Its special carrier is the one-parameter binomial count; its results are invalid if a set of binary observations is treated as independent common-p trials when it is actually clustered, weighted, adaptively selected, or otherwise governed by a different design without adjustment. The live Confidence Intervals node is a close neighbor, but its current wording requires at-least-nominal coverage for every admitted method; the Wald member of this family does not meet that requirement at small samples.

Structural Signature

Sig role-phrases:

  • Binary trial population and estimand — Define what counts as a success and the common success probability p being estimated.
  • Sampling model — Fix n and assume independent, identically distributed Bernoulli outcomes so the success count is binomial; other designs require other interval rules.
  • Observed count — Record k successes, n-k failures and p̂=k/n, retaining the count rather than only a rounded percentage.[1]
  • Nominal level — State the target confidence level 1−α and whether the interval is one- or two-sided.[2]
  • Endpoint construction — Choose Wilson score, Clopper–Pearson, Wald, adjusted Wald or another specified rule; different rules can produce different bounds for the same k,n.[1]
  • Coverage assessment — Interpret calibration across hypothetical repeated samples at fixed p, not as a posterior probability that p lies in this realized interval.

The condensed signature is binary common-p count + declared confidence construction → random interval with stated repeated-sampling coverage behavior. The method choice is not incidental: the same count can produce materially different coverage and width.

What It Is Not

  • Not a probability distribution over p. A frequentist 95% construction is calibrated across repeated samples; it does not by itself assign a 95% posterior probability to the realized bounds.
  • Not the sample proportion alone. k/n is a point estimate, with no interval rule or coverage guarantee.
  • Not a single binomial test. A test assesses one specified p0; test inversion can generate an interval by considering many p0 values.[2]
  • Not exact 95% coverage at every p when labeled “exact.” Clopper–Pearson's at-least-nominal guarantee is conservative at many parameter values because counts are discrete.[1]
  • Not automatically valid for a complex survey or repeated observations from the same person. Binary labels alone do not establish independent common-p sampling.

Scope of Application

NIST's process-monitoring example samples twenty units from a continuous production line and observes four defectives. Under a binomial sampling model, its displayed two-sided 90% Clopper–Pearson calculations give \(p_L=0.071354\) and \(p_U=0.401029\), or approximately \((0.071,0.401)\) to three decimals. NIST's concluding sentence instead prints \((0.071,0.400)\); that upper endpoint is inconsistent with its displayed calculation, so this entry follows the calculation rather than repeating the apparent rounding error. This is a sample-to-parameter statement, not an assertion about the defect fraction in the next twenty units.[2]

The same structure applies to any fixed-size set of genuinely independent common-probability binary trials, whether “success” names a desired outcome or a defect. NIST also illustrates thirty binary observations with eight successes and reports a 95% Wilson interval of about (0.141827, 0.444480). If trials have heterogeneous risks, dependence, or survey weights, one needs a model or design-based method that reflects them rather than treating the pooled count as automatically binomial.[1]

Clarity

The interval makes uncertainty about a population/process probability visible while preserving the assumptions needed to interpret it. Saying “8 out of 30” does not identify whether an ordinary Wald, Wilson, or exact-tail construction was used. Reporting the method, level and sampling design makes two apparently conflicting intervals reconcilable rather than mysterious.[1]

It also clarifies exactness. Clopper–Pearson solves exact binomial-tail equations but its coverage need not equal the nominal level pointwise. Wilson solves a score-based approximate construction and can perform better on width and practical coverage in many settings. “Exact” is therefore not a synonym for best, shortest, or most truthful in every use.[1]

Manages Complexity

The binomial model compresses an entire fixed-size binary sample to a count k and sample size n for inference on one common p. That reduction enables reproducible endpoint calculations and a tractable comparison of procedures by coverage and interval length. Its economy is conditional: if trial probabilities differ or outcomes are dependent, the count no longer carries all relevant design information.

Different constructions then manage different sources of approximation. Wald gains algebraic simplicity by replacing the unknown variance with an observed estimate; Wilson keeps the candidate p0 inside the score-test variance; Clopper–Pearson uses binomial tails directly. The gain in one dimension—simplicity, coverage floor or width—can cost another.[1][2]

Abstract Reasoning

Start with the estimand and sampling mechanism, not a favorite formula. Verify binary outcomes, fixed n, approximate independence and a common probability. Record k,n,α, select an endpoint rule and ask what coverage claim it warrants. For a score interval, possible p0 values are retained when their standardized discrepancy from k/n is not too large under the null variance; the accepted set becomes the interval.[2]

Check boundary behavior before trusting a shortcut. At k=0, the Wald plug-in standard error is zero and can collapse misleadingly at the boundary; alternative constructions retain uncertainty that unobserved successes might still occur. For a design with clustering or unequal probabilities, this diagnostic is not enough: the model itself has changed, so one must leave the simple binomial-proportion interval framework or justify an adjusted procedure.[1]

Knowledge Transfer

The inferential pattern transfers literally among process defects, laboratory binary trials and other independently sampled two-outcome observations with a common p. It does not transfer merely because data are recorded as yes/no. A weighted poll, paired clinical outcomes or repeated visits from one user can be binary yet have a different sampling structure. The broader act of deriving a range estimate belongs to live Estimation; this identity adds binomial-count-specific constraints and constructions. Confidence Intervals remains a near neighbor whose present coverage-floor boundary excludes some of this family's admitted procedures.

Examples

Defect proportion in a production line

NIST describes twenty sampled units, four of them defective. Its displayed exact-tail calculations \(p_L=0.071354\) and \(p_U=0.401029\) yield a 90% interval approximately \((0.071,0.401)\) at three decimals. The source's final prose prints 0.400, a source-internal inconsistency with \(p_U\), not a distinct statistical method. Under the stated binomial assumptions, the interval procedure has its at-least-nominal repeated-sampling coverage meaning for the process defect probability; it does not certify the defect fraction of every future finite batch.[2]

Mapped back: trials = 20 sampled units; count = 4 defectives; parameter = common defect probability; construction = two-sided 90% exact-binomial tail inversion; calibration = binomial-model coverage.

Binary observations with Wilson bounds

NIST's illustrative dataset has thirty binary observations and eight successes. Its 95% Wilson score interval is approximately (0.142, 0.444). Here the score construction, not a raw Wald plug-in variance, determines the endpoints. If the thirty observations were instead heavily clustered or unequally weighted, this numerical interval would not automatically retain its advertised model-based meaning.[1]

Mapped back: trials = 30 binary observations under the common-p model; count = 8 successes; parameter = success probability; construction = 95% Wilson score inversion; calibration = approximate repeated-sampling coverage.

Structural Tensions

Coverage assurance versus interval width. Exact-tail bounds can guarantee at least the nominal coverage under the binomial model, but discreteness often makes them wider than a well-behaved approximate interval. Diagnostic: does the use case require a conservative lower coverage bound, or is shorter expected width with controlled approximation the priority?[1]

Formula simplicity versus boundary reliability. Wald is compact and familiar, but its behavior can be poor near zero/one or at small n; Wilson and exact-tail constructions handle those cases differently. Diagnostic: what bounds arise when k=0, k=n, or the expected count is small?[1][2]

Structural–Framed Character

The identity is structural but model-framed. The interval rules and binomial coverage probabilities are mathematically determined once the data-generating assumptions, nominal level and construction are fixed. Human practice chooses what counts as a trial or success and what trade-off is acceptable; those choices matter to application but do not rewrite the mathematics. The evaluative word confidence can invite an incorrect subjective-probability reading, so its frequentist interpretation must be stated. Across domains, the same Bernoulli sampling structure is recognized, not imported by metaphor. Its character: a formal inferential method family whose reliability depends on explicit sampling assumptions and coverage criteria.

Structural Core vs. Domain Accent

The more general pattern is a confidence interval: data map to a range with a stated repeated-sampling property. The irreducible domain accent here is one unknown common Bernoulli probability, a fixed binary-trial count, and the specific score, normal or exact-tail constructions that produce the bounds. Remove those, and one has general interval estimation, not a binomial-proportion interval. Manufacturing, trials and polling are potential settings, not constituent roles; their sampling designs determine whether the same mathematics literally applies.

This entry is a kind of Estimation.

DAG parent: Estimation by strict subsumption. Every admitted method maps binomial count evidence into a range estimate for unknown p; Estimation also covers other quantities and non-interval outputs. The live Confidence Intervals is related but is not a strict parent under its present universal coverage-floor wording: with n=1, true p=1/2, the Wald 95% interval is {0} or {1} and has zero actual coverage. The live Binomial test is related because score and exact-tail tests can be inverted to form intervals, but a single test yields a decision about one p0, not a range. Poisson binomial distribution concerns nonidentical Bernoulli probabilities and is an important boundary rather than a parent. A future change to the live Confidence Intervals definition requires separate curation before adding that edge.

Relationships to Other Abstractions

Local relationship map for Binomial Proportion Confidence IntervalParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Binomial ProportionConfidence IntervalDOMAINPrime abstraction: Estimation — is a kind ofEstimationPRIME

Current abstraction Binomial Proportion Confidence Interval Domain-specific

Parents (1) — more general patterns this builds on

  • Binomial Proportion Confidence Interval is a kind of Estimation Prime

    A binomial-proportion interval is a rule-based range estimate of an unknown probability.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Binomial Proportion Confidence Interval sits in a sparse region of the domain-specific corpus (83rd percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (2551 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • Wilson score interval: one member of this family, not the whole family.[2]
  • Clopper–Pearson interval: the exact-binomial-tail member, whose at-least-nominal coverage can be conservative.[1]
  • Bayesian credible interval: a different inferential interpretation that requires a prior and posterior distribution for p.
  • Binomial test: evaluates a hypothesized p0; inversion over many p0 values can construct bounds.
  • Complex-survey proportion interval: may require weights, design effects or dependence handling that the simple binomial model lacks.

References

[1] National Institute of Standards and Technology, “Proportion Confidence Interval,” Dataplot Reference Manual, methods 1–5 and the worked 30/8 example. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n

[2] NIST/SEMATECH, “7.2.4.1 Confidence intervals,” e-Handbook of Statistical Methods, Wilson score inversion and the 20/4 process-defect example. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j