Skip to content

Probability of Success

A model-based probability that a specified planned or ongoing study will meet its declared success criterion at a future assessment.

Version
v1 · 2026-10-07 · History
Domain-specific #
13990
Domain group
Applied Sciences & Engineering
Origin domain
Medicine & Healthcare
Subdomains
Clinical Trials, Biostatistics → Medicine & Healthcare
Aliases
Study Endpoint Probability of Success

Core Idea

A study probability of success (PoS) is the model-based probability that a specified planned or ongoing study will meet a declared success rule at a future assessment. The target is the observed study result under its design: possible final data pass or fail the rule. The calculation needs a model for those future data, and its number is conditional on the stated information and assumptions. It is neither the realized result nor merely a probability that an underlying treatment effect is favorable.[1][2]

The same structure appears at two horizons. Saville's author-worked example predicts whether an ongoing single-arm study will meet its final evidence rule after more responses. Burke and colleagues use data from earlier Phase II trials to predict whether a possible new Phase III trial's finite-sample interval will show a favorable effect. The cases differ in evidence, model and decision rule; neither demonstrates that a treatment worked.[3][2]

Structural Signature

Sig role-phrases:

  • Specified study and future horizon — identifies which planned or ongoing study will be assessed and when. A drug-development program spanning trials and approval is a broader unit.[1][2][4]
  • Declared observed success event — defines a checkable final-data rule before assigning its probability. It may be a final posterior evidence threshold or a trial interval wholly on the favorable side; no one rule is universal.[3][2]
  • Stated information and design — fixes available interim or earlier-study evidence, sample size, endpoint and any effect or prior assumptions relevant to the forecast. Interim data and Bayesian priors are variants, not all-instance conditions.[3][2]
  • Future-observation probability model — distributes probability over possible final study data or statistics, including future sampling variation. Effect uncertainty may be modeled as well; a parameter-only probability without a future observation model fails this role.[1][2]
  • Interpreted event probability — reports the probability of the declared future success event under that model, a number in [0,1], without promoting it to a guarantee or a treatment-effect conclusion.[1][2]

What It Is Not

A posterior probability that a response parameter exceeds a threshold is about the parameter at the current information state. It need not equal the probability that the remaining observations will carry the study over its final success rule. In Saville's illustration the former is 0.81 and the latter is 0.54. Likewise, Burke's 0.824 probability of a favorable true effect for a new trial is a limiting, infinite-sample comparison, while its finite-trial probabilities of showing that effect are lower under the specified designs.[3][2]

It is not all uses of “PoS” in drug development. Hampson and colleagues analyze a broader program-level success event involving late-stage trials, approval and a target product profile. That legitimate neighboring usage is outside this entry's admitted single-study endpoint scope. A historical industry success rate or an already observed pass/fail label also lacks the prospective, study-specific future-data event required here.[4]

Scope of Application

The admitted class includes forecasts made before a study starts and updates while it is ongoing, provided the future assessment and observed rule are explicit. A study can be evaluated under fixed-effect assumptions or by averaging over effect uncertainty; neither Bayesian posterior conditioning nor one frequentist threshold defines every member. The sources here establish clinical-trial design and monitoring cases, so this entry does not claim that a particular numeric procedure works across every research field.[1][3][2]

Within that scope, PoS can support a design conversation or an interim decision. Those are uses of the probability, not membership conditions. A high PoS can still be wrong when the model is misspecified or new data vary; it is not proof of efficacy, safety, approval or eventual enrollment.[1][2][4]

Clarity

To interpret a reported PoS, ask five questions: Which study and future readout? What exact final-data result counts as success? Which information and design were supplied? What model distributes the unobserved data? Which number was actually computed? If the last number concerns only a latent effect or a broad program event, rename it rather than silently treating it as this study-endpoint PoS.[3][2][4]

A threshold must be read with its context. Saville's illustration sets final success as posterior probability of response rate above 0.5 exceeding 0.95, under its uniform beta prior and 100-person final sample. Burke instead asks whether a possible new trial's 95% Bayesian credible interval for its odds ratio lies wholly on the beneficial side of 1. Neither event definition can be imported into the other trial without changing the calculation.[3][2]

Manages Complexity

The single number summarizes possible future observations against a stated end rule. It helps compare a candidate sample size or an interim continuation policy without listing every possible final dataset. Yet compression hides the model, effect assumptions, endpoint, prior evidence and uncertainty sources. Keeping these visible makes clear why two PoS values can differ even when they concern the same treatment.[1][2]

Burke's example shows why the distinction matters: its estimated probability of a favorable true new-trial effect is 0.824, but finite new trials at 2,000 or 4,000 patients per arm have roughly 0.4 or 0.6 modeled probability of demonstrating benefit by the declared interval rule. A probability of benefit is not automatically a probability of showing benefit in a finite study.[2]

Abstract Reasoning

Let Y_future denote possible future study data, I the declared present information and assumptions, and S(Y_future) the final-data success event. Then PoS has the form P(S(Y_future) | I, model, design). The model must normalize probabilities over possible results; the event and horizon must be fixed enough to calculate. For an ongoing study, conditioning can use observed interim data. For a planned study, design assumptions may be declared before any interim observation.[1][3][2]

A counterfactual changes the target while holding the model fixed. In Saville's worked numbers, asking whether the response parameter now exceeds 0.5 produces 0.81; asking whether enough future responses will occur for final study success produces 0.54. In Burke's model, moving from 2,000 to 4,000 participants per arm changes the future sampling distribution and raises the modeled chance of an interval showing benefit from about 0.4 to 0.6. Those differences follow from event and design, not a universal monotonic effect of every design change.[3][2]

Knowledge Transfer

The role audit transfers between within-trial and future-trial prediction: identify the study, observed success event, information/design, future-data model and resulting event probability. Saville supplies an interim beta-binomial continuation; Burke combines evidence from nine real earlier Phase II trials with predicted variation in a possible new Phase III study. What transfers is this audit, not the beta prior, response threshold, odds-ratio interval or numerical success chance.[3][2]

The broader Probability prime supports probabilities of many future events. Study success is narrower because the event and uncertainty are shaped by research design and evaluative rules. A metaphorical “chance of success” for an unmodeled project does not acquire this trial-specific identity by using the phrase.[1][4]

Examples

Canonical: Saville's illustrative interim study

Saville's author lecture considers an illustrative one-arm study of 100 binary responses with a uniform Beta(1,1) prior. Final success means posterior probability that response rate exceeds 0.5 is above 0.95, which under this setup requires at least 59 total responses. At the first interim look, 12 of 20 responses have occurred, so at least 47 of the next 80 are needed. The beta-binomial predictive chance of reaching final success is 0.54, distinct from the current posterior probability 0.81 that the parameter exceeds 0.5. Mapped back: the 100-person study and final look are the study and horizon; the posterior-evidence threshold is the observed success event; 12/20, prior and remaining 80 are the information and design; beta-binomial continuation is the future-observation model; and 0.54 is the event probability. This is a methodological illustration, not a real observed patient cohort or efficacy claim.[3][1]

Applied: Burke's possible Phase III trial

Burke and colleagues analyze nine real Phase II thrombolysis trials and model a possible new Phase III study with intracranial-hemorrhage outcome. Under an assumed control risk of 0.01, success means the new finite trial's 95% Bayesian credible interval for its odds ratio lies wholly below 1. The authors report about 0.4 modeled success probability at 2,000 patients per arm and about 0.6 at 4,000 per arm. Mapped back: the possible new Phase III study and readout are the study and horizon; the interval criterion is the observed event; prior trial data, risk assumption and enrollment are the information and design; between-trial and future-sample variation form the future-observation model; and 0.4/0.6 are its event probabilities. The 0.824 true-effect probability is a different quantity. This is a retrospective model illustration using real prior trials, not a logged forecast or the result of that future trial.[2]

Structural Tensions

A case-specific design tension appears in Burke's enrollment comparison. Larger modeled recruitment raises the chance that the future interval will show benefit, while also requiring more participants and resources. Moving from 2,000 to 4,000 per arm roughly raises its computed PoS from 0.4 to 0.6 under the authors' assumptions. Diagnostic: is that additional modeled chance worth the recruitment and burden for this study and endpoint? This tension belongs to a design decision; no all-instance conflict or universal sample-size rule follows from the definition of PoS.[2]

Structural–Framed Character

This entry is framed by research practice around a structural probability calculation. Future study data can vary independently of how investigators name them, while investigators and research institutions choose the study, endpoint, final success rule and design. The probability vocabulary travels literally from interim to planned-trial settings when a coherent event model is present; it travels outside study design only as the broader Probability relation. “Success” has evaluative weight because the rule picks a result valued for that particular study, but the reported probability by itself does not make the result desirable or true. Calling a number PoS recognizes a modeled event probability when the roles are documented; importing the label without an event/model creates no valid forecast. Its character: a model-relative future-event probability whose study and success frame depends on research decisions, with a portable mathematical core carried by Probability.[1][2]

Structural Core vs. Domain Accent

The portable skeleton is a probability of a declared future event under specified information and a coherent model. The domain accent is a planned or ongoing study, its final observed-data success criterion, design and future-data uncertainty. Remove the study and evaluative end rule, and the named study-endpoint measure disappears even though a probability can remain. Prime bar: this entry cannot travel intact to a coin toss or weather event, because the study-design and observed endpoint roles would be lost; the live Probability prime already carries the general event-measure structure. A future, broader success-event Prime would require independent cross-domain evidence rather than an analogy drawn from this trial usage.[1][2][4]

This entry is a kind of Probability.

Probability of success is, in every case, a kind of Probability, its one broader abstraction: possible final data, an event, a normalized measure, conditional information where applicable, model dependence and interpretation fill out everything Probability requires. Statistical Power is close in fixed-effect detection designs but requires its own false-null framework, which is not universal here. Conditional Probability describes some interim updates, but a design-stage assignment need not condition on an observed event; it is not listed as a redundant direct broader abstraction. A probability about a trial is not itself an executed Clinical Trial.[1][2]

Relationships to Other Abstractions

Local relationship map for Probability of SuccessParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Probabilityof SuccessDOMAINPrime abstraction: Probability — is a kind ofProbabilityPRIME

Current abstraction Probability of Success Domain-specific

Parents (1) — more general patterns this builds on

  • Probability of Success is a kind of Probability Prime

    A study-endpoint probability of success is a probability of a declared future observed study event.

Hierarchy paths (2) — routes to 2 parentless roots

Neighborhood in Abstraction Space

Probability of Success sits in a sparse region of the domain-specific corpus (74th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Causal Inference & Regression Modeling (15 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • Current posterior effect probability: uncertainty about a parameter, distinct from the future observed study decision.[3][2]
  • Statistical Power: a detection probability under a specified effect and test setup; a related variant, not the whole admitted PoS class.[3]
  • Predictive probability in one monitoring design: a method for obtaining study PoS, not a required beta prior or interim look for every member.[1]
  • Development-program PoS: Hampson's broader approval and product-profile event, not this one-study endpoint measure.[4]
  • A realized success indicator: the eventual pass/fail observation, not the probability assigned before it is observed.

References

[1] Benjamin R. Saville, Jason T. Connor, Gregory D. Ayers and JoAnn Alvarez, “The utility of Bayesian predictive probabilities for interim monitoring of clinical trials”, Clinical Trials 11 (2014), 485–493, DOI 10.1177/1740774514531352. Original abstract inspected; full original was not directly accessible for this review. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n

[2] Danielle L. Burke, Lucinda J. Billingham, Alan J. Girling and Richard D. Riley, “Meta-analysis of randomized phase II trials to inform subsequent phase III decisions”, Trials 15 (2014), 346, DOI 10.1186/1745-6215-15-346. Full original publisher article inspected; the future Phase III case is a model illustration using real earlier data. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r ↩s ↩t ↩u ↩v ↩w

[3] Benjamin R. Saville, “The Utility of Bayesian Predictive Probabilities for Interim Monitoring of Clinical Trials”, DIA KOL Lecture Series, 2015, PDF pp. 5–6, 8–11 and 21–23. Full author lecture inspected; the numerical study is explicitly illustrative and this is not the journal article. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m

[4] Lisa V. Hampson et al., “Improving the assessment of the probability of success in late stage drug development”, Pharmaceutical Statistics 21 (2022), 439–459, DOI 10.1002/pst.2179. Publisher abstract and author preprint inspected for the broader program-level boundary. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g