Probability of Success¶
A model-based probability that a specified planned or ongoing study will meet its declared success criterion at a future assessment.
Core Idea¶
A study probability of success (PoS) is the model-based probability that a specified planned or ongoing study will meet a declared success rule at a future assessment. The target is the observed study result under its design: possible final data pass or fail the rule. The calculation needs a model for those future data, and its number is conditional on the stated information and assumptions. It is neither the realized result nor merely a probability that an underlying treatment effect is favorable.[ref-678eb869f67a][ref-badf0b583d11]
The same structure appears at two horizons. Saville's author-worked example predicts whether an ongoing single-arm study will meet its final evidence rule after more responses. Burke and colleagues use data from earlier Phase II trials to predict whether a possible new Phase III trial's finite-sample interval will show a favorable effect. The cases differ in evidence, model and decision rule; neither demonstrates that a treatment worked.[ref-7d3216c17f6b][ref-badf0b583d11]
Scope of Application¶
The admitted class includes forecasts made before a study starts and updates while it is ongoing, provided the future assessment and observed rule are explicit. A study can be evaluated under fixed-effect assumptions or by averaging over effect uncertainty; neither Bayesian posterior conditioning nor one frequentist threshold defines every member. The sources here establish clinical-trial design and monitoring cases, so this entry does not claim that a particular numeric procedure works across every research field.[ref-678eb869f67a][ref-7d3216c17f6b][^ref-badf0b583d11]
Within that scope, PoS can support a design conversation or an interim decision. Those are uses of the probability, not membership conditions. A high PoS can still be wrong when the model is misspecified or new data vary; it is not proof of efficacy, safety, approval or eventual enrollment.[ref-678eb869f67a][ref-badf0b583d11][^ref-a0ccdb300ae2]
Clarity¶
The roles are a specified study and future horizon, an observed final-data success event, stated information and design, a probability model for possible future observations, and an interpreted event probability. Each is needed to keep the number tied to a reproducible study forecast.[ref-678eb869f67a][ref-badf0b583d11]
To interpret a reported PoS, ask five questions: Which study and future readout? What exact final-data result counts as success? Which information and design were supplied? What model distributes the unobserved data? Which number was actually computed? If the last number concerns only a latent effect or a broad program event, rename it rather than silently treating it as this study-endpoint PoS.[ref-7d3216c17f6b][ref-badf0b583d11][^ref-a0ccdb300ae2]
A threshold must be read with its context. Saville's illustration sets final success as posterior probability of response rate above 0.5 exceeding 0.95, under its uniform beta prior and 100-person final sample. Burke instead asks whether a possible new trial's 95% Bayesian credible interval for its odds ratio lies wholly on the beneficial side of 1. Neither event definition can be imported into the other trial without changing the calculation.[ref-7d3216c17f6b][ref-badf0b583d11]
Manages Complexity¶
The single number summarizes possible future observations against a stated end rule. It helps compare a candidate sample size or an interim continuation policy without listing every possible final dataset. Yet compression hides the model, effect assumptions, endpoint, prior evidence and uncertainty sources. Keeping these visible makes clear why two PoS values can differ even when they concern the same treatment.[ref-678eb869f67a][ref-badf0b583d11]
Burke's example shows why the distinction matters: its estimated probability of a favorable true new-trial effect is 0.824, but finite new trials at 2,000 or 4,000 patients per arm have roughly 0.4 or 0.6 modeled probability of demonstrating benefit by the declared interval rule. A probability of benefit is not automatically a probability of showing benefit in a finite study.[^ref-badf0b583d11]
Abstract Reasoning¶
Let Y_future denote possible future study data, I the declared present information and assumptions, and S(Y_future) the final-data success event. Then PoS has the form P(S(Y_future) | I, model, design). The model must normalize probabilities over possible results; the event and horizon must be fixed enough to calculate. For an ongoing study, conditioning can use observed interim data. For a planned study, design assumptions may be declared before any interim observation.[ref-678eb869f67a][ref-7d3216c17f6b][^ref-badf0b583d11]
A counterfactual changes the target while holding the model fixed. In Saville's worked numbers, asking whether the response parameter now exceeds 0.5 produces 0.81; asking whether enough future responses will occur for final study success produces 0.54. In Burke's model, moving from 2,000 to 4,000 participants per arm changes the future sampling distribution and raises the modeled chance of an interval showing benefit from about 0.4 to 0.6. Those differences follow from event and design, not a universal monotonic effect of every design change.[ref-7d3216c17f6b][ref-badf0b583d11]
Knowledge Transfer¶
The role audit transfers between within-trial and future-trial prediction: identify the study, observed success event, information/design, future-data model and resulting event probability. Saville supplies an interim beta-binomial continuation; Burke combines evidence from nine real earlier Phase II trials with predicted variation in a possible new Phase III study. What transfers is this audit, not the beta prior, response threshold, odds-ratio interval or numerical success chance.[ref-7d3216c17f6b][ref-badf0b583d11]
The broader Probability prime supports probabilities of many future events. Study success is narrower because the event and uncertainty are shaped by research design and evaluative rules. A metaphorical “chance of success” for an unmodeled project does not acquire this trial-specific identity by using the phrase.[ref-678eb869f67a][ref-a0ccdb300ae2]
Example¶
Canonical: Saville's illustrative interim study¶
Saville's author lecture considers an illustrative one-arm study of 100 binary responses with a uniform Beta(1,1) prior. Final success means posterior probability that response rate exceeds 0.5 is above 0.95, which under this setup requires at least 59 total responses. At the first interim look, 12 of 20 responses have occurred, so at least 47 of the next 80 are needed. The beta-binomial predictive chance of reaching final success is 0.54, distinct from the current posterior probability 0.81 that the parameter exceeds 0.5. Mapped back: the 100-person study and final look are the study and horizon; the posterior-evidence threshold is the observed success event; 12/20, prior and remaining 80 are the information and design; beta-binomial continuation is the future-observation model; and 0.54 is the event probability. This is a methodological illustration, not a real observed patient cohort or efficacy claim.[ref-7d3216c17f6b][ref-678eb869f67a]
Applied: Burke's possible Phase III trial¶
Burke and colleagues analyze nine real Phase II thrombolysis trials and model a possible new Phase III study with intracranial-hemorrhage outcome. Under an assumed control risk of 0.01, success means the new finite trial's 95% Bayesian credible interval for its odds ratio lies wholly below 1. The authors report about 0.4 modeled success probability at 2,000 patients per arm and about 0.6 at 4,000 per arm. Mapped back: the possible new Phase III study and readout are the study and horizon; the interval criterion is the observed event; prior trial data, risk assumption and enrollment are the information and design; between-trial and future-sample variation form the future-observation model; and 0.4/0.6 are its event probabilities. The 0.824 true-effect probability is a different quantity. This is a retrospective model illustration using real prior trials, not a logged forecast or the result of that future trial.[^ref-badf0b583d11]
Relationships to Other Abstractions¶
Current abstraction Probability of Success Domain-specific
Parents (1) — more general patterns this builds on
-
Probability of Success is a kind of Probability Prime
A study-endpoint probability of success is a probability of a declared future observed study event.
Hierarchy paths (2) — routes to 2 parentless roots
- Probability of Success → Probability → Measure → Aggregation → Micro Macro Linkage
- Probability of Success → Probability → Measure → Set and Membership
Neighborhood in Abstraction Space¶
Probability of Success sits in a sparse region of the domain-specific corpus (74th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Causal Inference & Regression Modeling (15 abstractions)
Nearest neighbors
- Kaplan–Meier estimator — 0.84
- Continuous Individualized Risk Index — 0.84
- Clinical Endpoint — 0.84
- Clinical Trial — 0.83
- Event Sampling Methodology — 0.83
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Current posterior effect probability: uncertainty about a parameter, distinct from the future observed study decision.[ref-7d3216c17f6b][ref-badf0b583d11]
- Statistical Power: a detection probability under a specified effect and test setup; a related variant, not the whole admitted PoS class.[^ref-7d3216c17f6b]
- Predictive probability in one monitoring design: a method for obtaining study PoS, not a required beta prior or interim look for every member.[^ref-678eb869f67a]
- Development-program PoS: Hampson's broader approval and product-profile event, not this one-study endpoint measure.[^ref-a0ccdb300ae2]
- A realized success indicator: the eventual pass/fail observation, not the probability assigned before it is observed.
References¶
[^ref-678eb869f67a]: Benjamin R. Saville, Jason T. Connor, Gregory D. Ayers and JoAnn Alvarez, “The utility of Bayesian predictive probabilities for interim monitoring of clinical trials”, Clinical Trials 11 (2014), 485–493, DOI 10.1177/1740774514531352. Original abstract inspected; full original was not directly accessible for this review. [^ref-7d3216c17f6b]: Benjamin R. Saville, “The Utility of Bayesian Predictive Probabilities for Interim Monitoring of Clinical Trials”, DIA KOL Lecture Series, 2015, PDF pp. 5–6, 8–11 and 21–23. Full author lecture inspected; the numerical study is explicitly illustrative and this is not the journal article. [^ref-badf0b583d11]: Danielle L. Burke, Lucinda J. Billingham, Alan J. Girling and Richard D. Riley, “Meta-analysis of randomized phase II trials to inform subsequent phase III decisions”, Trials 15 (2014), 346, DOI 10.1186/1745-6215-15-346. Full original publisher article inspected; the future Phase III case is a model illustration using real earlier data. [^ref-a0ccdb300ae2]: Lisa V. Hampson et al., “Improving the assessment of the probability of success in late stage drug development”, Pharmaceutical Statistics 21 (2022), 439–459, DOI 10.1002/pst.2179. Publisher abstract and author preprint inspected for the broader program-level boundary.