Eventual Occurrence Containment Design¶
When a harmful outcome retains nonzero probability across many opportunities, design as though it will occur within the relevant horizon: keep reducing risk, but also cap impact, isolate propagation, detect quickly, and prove recovery.
Overview¶
When a harmful outcome retains nonzero probability across many opportunities, design as though it will occur within the relevant horizon: keep reducing risk, but also cap impact, isolate propagation, detect quickly, and prove recovery.
This archetype addresses a recurring error in risk design: a small number attached to one opportunity is treated as if it described the safety of an entire operating horizon. It does not. A request-level, unit-level, dose-level, cycle-level, trip-level, or season-level probability must be combined with the number and structure of opportunities before a system can claim that occurrence is unlikely in practice.
The response is not fatalism. The archetype keeps prevention, opportunity reduction, and safety margins in place. Its distinctive move is to add a governed posture switch: once repeated exposure and consequence make an occurrence foreseeable enough, the system must be designed so that the event is detected, contained, capped, survived or stopped safely, remedied, and learned from.
Mathematical basis and guardrails¶
For independent opportunities with constant per-opportunity probability \(p\), the probability of no occurrence in \(n\) opportunities is \((1-p)^n\). Therefore:
When \(p>0\), this expression approaches 1 as \(n\) grows without bound. In finite systems, the important output is not a slogan about inevitability but a decision-specific probability range over the actual horizon. Even a very small \(p\) can become material when \(n\) is large.
The formula is a guardrail only when its assumptions are defensible. Repeated events may be correlated, clustered, seasonal, learning, adaptive, or driven by a common cause. Probability may change as safeguards degrade, operators learn, attackers probe, assets accumulate, or the environment shifts. The draft therefore requires a dependence and mixing check and allows simulation, empirical frequency ranges, scenario bounds, or other conservative models in place of the simple independent-trial expression.
A second guardrail is equally important: nonzero probability is not moral or operational permission. Where consequence is catastrophic, irreversible, unlawful, or imposed on nonconsenting parties, the correct response may be to eliminate the activity, reduce exposure, or stop. Containment supplements prevention only when a tolerable consequence cap can actually be demonstrated.
Structural lifecycle¶
1. Define the event and opportunity process¶
Name the adverse outcome precisely. Then define one opportunity for occurrence and the horizon relevant to the decision. A denominator such as “per request” is incomplete without the number of requests, their distribution over time, and whether the opportunities share dependencies.
2. Bound probability and dependence¶
Record the evidence for the per-opportunity probability or probability interval, including detection limits and residual uncertainty. Test whether opportunities are independent, weakly dependent, clustered, or adaptive. A quiet history can narrow a bound; it does not automatically prove impossibility.
3. Compute horizon-level occurrence¶
Use an analytic expression, simulation, empirical model, fault tree, or conservative scenario envelope to express the chance or expected frequency of one or more occurrences across the horizon. Examine sensitivity to volume growth and probability error. The result should be understandable to the accountable owner and affected decision-makers.
4. Make the posture switch explicit¶
Define when prevention-only reasoning is no longer acceptable. The trigger can be a cumulative probability threshold, an expected frequency, consequence severity, irreversibility, a legal duty, or an ethical nonwaivable condition. The threshold creates accountability: it identifies when containment, safe-state, and recovery evidence become required rather than optional.
5. Preserve prevention while reducing exposure¶
Continue hardening controls, removing hazardous paths, reducing the number of opportunities, limiting privileges, shortening exposure windows, and using safety margins. The archetype rejects the claim that eventual occurrence makes prevention pointless.
6. Cap and contain consequence¶
Create failure domains, isolation points, maximum affected populations, time limits, loss limits, and protected critical functions. Containment should include detection and response latency and should be tested against bypasses and common-mode dependencies.
7. Enter a safe or degraded state¶
Define what stops, what continues, and who may override. A safe state must be reachable under the same conditions that create the incident; a runbook that depends on failed infrastructure is not a safe state.
8. Recover, remedy, and control recurrence¶
Restoration is only one part of recovery. Preserve evidence, notify and support affected parties, correct records or losses, analyze the probability model and opportunity process, and decide whether recurrence remains credible before full exposure resumes.
9. Refresh on drift¶
Opportunity volume, asset concentration, dependencies, attack behavior, product versions, environment, and control performance change. Those changes can increase horizon-level risk even if the original per-opportunity number appears unchanged. Model refresh must be triggered by exposure drift and actual events.
Components¶
The required components divide into four linked groups:
- Model inputs: adverse outcome, opportunity unit and horizon, probability bound, dependence check, and cumulative occurrence model.
- Governance: posture-switch threshold, residual prevention risk, accountable owner, and stop authority.
- Consequence control: containment boundary, consequence cap, critical-function plan, detection and escalation, and safe or degraded state.
- Learning and repair: recovery and recurrence plan plus exposure-drift and model-refresh trigger.
Optional components add adversarial attempt modeling, common-mode checks, affected-party remedy, reserve or transfer capacity, independent assurance, and a stop-operation criterion. They should remain components because each is a structural part of the lifecycle, not a complete intervention on its own.
Mechanisms¶
Mechanisms instantiate the archetype in domain-specific ways. A cumulative risk horizon table and opportunity exposure register make the denominator and horizon visible. Repeated-trial calculators, fault trees, and probabilistic safety assessments model occurrence. Failure injection, blast-radius testing, automatic isolation, degraded-mode runbooks, and recovery drills test consequence control. Sentinel event monitoring and post-incident recurrence review close the learning loop. A stop-or-scale-back gate protects against the misuse of containment as a license for unbounded exposure.
No single mechanism is the archetype. A probability calculator without authority and containment is analysis only. A bulkhead without horizon-level reasoning is an isolation mechanism. A backup without a tested restore path is an unverified component. The archetype exists only when repeated-opportunity reasoning, posture switching, prevention, containment, safe response, recovery, and refresh are connected.
Key parameters and invariants¶
Parameters that should be explicit include:
- the event definition and severity class;
- opportunity unit, count, distribution, and forecast horizon;
- per-opportunity probability range and evidence quality;
- dependence, stationarity, seasonality, and adaptive-opponent assumptions;
- cumulative occurrence threshold or other posture-switch rule;
- maximum blast radius, duration, affected population, and loss;
- detection and isolation latency;
- recovery time and recovery point objectives where applicable;
- reserve, remedy, and notification capacity;
- model-refresh cadence and change triggers.
The strongest invariants are behavioral: per-opportunity risk is never silently presented as horizon safety; prevention remains active; the consequence cap is tested; rescue layers do not share an unexamined common-mode dependency; affected parties are inside the system boundary; and exposure growth forces reassessment.
Boundaries and neighbor distinctions¶
Probabilistic Risk Weighting is the nearest analytical neighbor. It helps weight likelihood and consequence, but it does not by itself count opportunities or create a posture switch into containment and recovery.
Failure Mode Anticipation finds ways a design can fail. It can supply fault trees and mitigation candidates, but this archetype adds the repeated-opportunity argument for treating occurrence as a design basis.
Fault-Tolerant Operation preserves service after partial failure. It is narrower on continuity and may be one target outcome; Eventual-Occurrence Containment Design may instead require shutdown, quarantine, or exposure elimination.
Fail-Safe Default supplies the least harmful reachable state. It is a component or supporting archetype inside the broader lifecycle.
Rupture Containment and Bulkhead Isolation govern post-break propagation and compartmentalization. They are essential implementation neighbors but do not decide from cumulative opportunity and residual probability when those structures become mandatory.
Safety Margin Design reduces per-opportunity risk by increasing distance from a boundary. This archetype addresses what remains when that risk cannot be proven zero.
Resilience Capacity Building is broader and less tied to a defined event and opportunity process. Wild-Card Contingency Mapping addresses low-probability surprises without requiring repeated trials. Tail-Risk Preservation protects rare cases from common-case optimization rather than modeling at-least-one occurrence across exposure.
Tradeoffs and decision costs¶
Eventual-occurrence containment deliberately commits capacity before the adverse event occurs. Redundant barriers, consequence caps, safe-state machinery, drills, reserves, and recovery paths impose cost, latency, maintenance burden, and sometimes lower throughput. The central tradeoff is not prevention versus containment: prevention and exposure reduction remain mandatory. The decision is how much scarce capacity to allocate to reducing opportunities and how much to allocate to limiting consequences once residual horizon-level occurrence is no longer negligible.
Overbuilding containment can normalize preventable exposure, create complexity that adds common-mode failure, or consume resources better spent eliminating opportunities. Underbuilding containment leaves the system dependent on a prevention claim that repeated operation may eventually falsify. The design should therefore record the marginal risk reduction, residual harm, coupling, and reversibility of each prevention and containment layer; protect non-negotiable harm limits; and revisit the allocation whenever opportunity count, dependence, consequence stock, or operating regime changes.
Failure modes and misuse¶
The most serious technical error is false independence. Correlation can either increase clustering and common-cause risk or reduce the relevance of a simple trial count. A second error is horizon manipulation: owners can choose a short time window or a narrow opportunity unit to keep a cumulative number below a governance threshold.
The most serious organizational error is containment theater. Diagrams, policies, backups, or runbooks can create confidence without proving that detection, isolation, safe-state entry, remedy, and restore work under realistic dependency failures. The most serious ethical error is externalizing the blast radius—preserving the formal system while transferring losses to customers, workers, communities, or people with less power.
The archetype must also resist fatalism. “Eventually” does not mean “nothing can be prevented.” Good use reduces per-opportunity probability, reduces the number of opportunities, and caps consequence. Where no tolerable cap is possible, the design must include authority to stop.
Worked example¶
Suppose a cloud service estimates a catastrophic parser failure probability of \(10^{-8}\) per request after preventive controls and expects 500 million requests in a month. Under the independent constant-probability approximation:
Calling this merely a “one-in-one-hundred-million event” would be misleading at the monthly horizon. The platform keeps reducing risky inputs and hardening the parser, but it also isolates tenants, limits privileges and corruptible state, detects anomalies, trips affected shards, enters degraded mode, preserves evidence, restores from tested checkpoints, notifies affected customers, and reviews recurrence before normal exposure resumes. A common-mode review tests whether detection, isolation, and restore depend on the same control plane. Volume, code, input mix, and attacker behavior trigger model refresh. If the declared blast-radius cap cannot be demonstrated, deployment is scaled back or stopped.
Non-examples¶
This is not a generic slogan that “anything that can go wrong will go wrong.” It requires an event, opportunity process, horizon, probability or conservative bound, consequence, and an actionable containment or stop decision. It is not a substitute for immediate incident response, a generic risk register, a safety margin, or a broad resilience program. It also does not apply where a monitored invariant genuinely makes the event impossible inside the declared envelope.
Review recommendation¶
Use this as a bounded, merge-sensitive full archetype. During global reconciliation, preserve the distinctive chain:
nonzero residual probability → repeated opportunities → horizon-level occurrence → posture switch → prevention plus opportunity reduction → consequence cap and containment → safe response and recovery → recurrence and model refresh.
Collapse it only if an existing accepted archetype is deliberately broadened to own that entire chain.
Common Mechanisms¶
- Automatic Isolation Trip — The instant a trigger fires, it severs the connections around a failing part — confining damage inside a pre-drawn boundary and dropping the isolated piece into a safe state, with no human in the loop.
- Blast-Radius Test — Deliberately fails one component and measures how far the damage actually reaches — sizing the worst-case impact and exposing the shared dependencies that make the blast bigger than the diagram claims.
- Cumulative Risk Horizon Table — Lays a tiny per-opportunity probability across the real number of opportunities in the horizon, turning 'practically zero' into a cumulative chance — and marking the point where prevention-only must give way to containment.
- Degraded-Mode Runbook — The pre-written procedure for running on reduced capability — which functions to shed, which to keep alive by hand, and the verified path back to full service.
- Failure-Injection Test — Deliberately induces a fault in the real system to confirm that detection, isolation, and failover actually fire as designed — proving the defensive chain before a real event exercises it.
- Fault Tree with Repeated-Opportunity Branch — A top-down failure-logic tree with an added branch for the event recurring across many demands — compounding a small per-demand probability into a horizon-level one and exposing where the 'independent trials' assumption quietly breaks.
- Opportunity Exposure Register — Keeps a living inventory of every place the adverse outcome could occur and how fast opportunities are piling up, so the 'many chances' fact never quietly goes stale.
- Post-Incident Recurrence Review — After an occurrence actually happens, makes affected parties whole and traces the shared root cause so the same event cannot recur the same way.
- Probabilistic Safety Assessment — A whole-system probabilistic model that scopes exactly what counts as the adverse outcome, tests the independence assumptions simpler math takes for granted, and records the residual risk no control removes.
- Recovery Drill and Restore Test — Actually restores the system from a simulated occurrence, end to end and on the clock, to prove rather than assume that recovery works and critical functions return within their targets.
- Repeated-Trial Probability Calculator — Converts a small per-opportunity probability and a large number of opportunities into the near-certainty of at least one occurrence over the whole horizon.
- Sentinel Event Monitoring — Watches continuously for specific pre-defined rare events whose single occurrence signals high consequence or systemic failure and warrants immediate response.
- Stop-or-Scale-Back Gate — A pre-committed rule that halts or throttles operation the moment cumulative risk crosses a set line, so stopping doesn't depend on someone finding the nerve in the moment.
Compression statement¶
The archetype corrects per-opportunity myopia. A small residual probability can look negligible when each trial is viewed alone, yet many trials can make one or more occurrences likely, recurrent, or asymptotically almost sure. The intervention defines the event, opportunity unit, horizon, probability and dependence assumptions; estimates cumulative occurrence; selects a posture-switch threshold; and requires prevention, exposure reduction, containment boundaries, consequence caps, safe response, recovery, and model refresh as one governed system.
Canonical formula: For n independent opportunities with constant per-opportunity probability p, P(at least one occurrence) = 1 - (1 - p)^n, which approaches 1 as n approaches infinity when p > 0. For varying or dependent opportunities, use an explicit dependence model, conservative bound, simulation, or empirical frequency range; do not claim inevitability from the independent-trial formula when its assumptions fail.
Related Abstractions¶
Abstractions this archetype builds on — directly (a source ingredient) or as a related pattern. Links follow the typed catalog namespace.
Built directly on (7)
- Containment: Holding a hazard, process, or agent within a deliberately maintained perimeter to prevent its spread or uncontrolled interaction with the surroundings.
- Eventual Realisation of Possibility: In a system subject to many independent trials, every outcome with non-zero per-trial probability eventually occurs, so design posture shifts from per-trial prevention to containment.
- Probability: Quantifies uncertainty and likelihoods.
- Recurrence: The property by which a state, event, or value reappears across time or iterations because the present state depends on prior states, distinct from mere repetition by its measurable lag structure.
- Risk: Exposure to a known distribution of possible outcomes.
- Statistical Independence: Learning one variable gives no information about another; the joint distribution factors.
- Time: The dimension that orders events from earlier to later with measurable duration and an irreversible direction, providing the foundation for change, rate, and causality.
Also references 23 related abstractions
- Boundary: Defines system limits.
- Conditional Probability: Re-normalize a probability measure to the information context that is taken as given.
- Controllability: Ability to steer system.
- Coupling: Interdependence among subsystems.
- Defense In Depth: Stacking multiple independent protective layers between threat and asset so that only a correlated breach across all layers produces total loss.
- Escape and Leakage: Constrained quantities exit through unintended pathways.
- Fail-Safe: Default to safe state on failure.
- Fault Tolerance: Continue operating under failure.
- Irreversibility: Cannot revert state.
- Normalization of Deviance: An operating standard drifts outward because small departures, each repeatedly observed without immediate catastrophe, are reclassified as normal.
Variants¶
Narrower or domain-specific specializations that share this archetype's core structure. Recognized variants are established; candidate variants are provisional.
Finite-Horizon Cumulative Risk Governance · temporal variant · recognized
Use a bounded opportunity count and decision horizon to determine when a small per-opportunity probability becomes a containment and recovery requirement.
- Distinct from parent: The parent includes unbounded and qualitative recurrence reasoning; this variant requires explicit finite-horizon governance.
- Use when: The relevant horizon is finite but contains enough opportunities for at least one occurrence to become materially likely; The organization needs an auditable threshold rather than an informal appeal to eventuality; Opportunity volume and probability bounds can be estimated or conservatively bracketed.
- Typical domains: high volume software, manufacturing, healthcare operations, finance
- Common mechanisms: cumulative risk horizon table, repeated trial probability calculator, stop or scale back gate
Assume-Breach Containment · risk or failure variant · recognized
Treat repeated or adaptive attack opportunities as sufficient reason to design for eventual control bypass, then isolate privileges, data, and propagation before a breach occurs.
- Distinct from parent: The parent can model stochastic repeated opportunities; this variant assumes adaptive adversarial pressure and therefore distrusts independence and stationarity.
- Use when: Attackers can make many attempts or adapt after observing defenses; A single perimeter control cannot be proven complete over the operating horizon; Compartmentalization, detection, credential rotation, and recovery can materially cap harm.
- Typical domains: cybersecurity, fraud, identity and access, abuse prevention
- Common mechanisms: automatic isolation trip, blast radius test, post incident recurrence review
Eventual Defect-Escape Containment · domain variant · recognized
Treat a nonzero residual defect-escape rate across high production volume as an eventual field event and prebuild traceability, quarantine, remedy, and recurrence control.
- Distinct from parent: The parent is domain-general; this variant adds lot traceability, quarantine, recall, and corrective-action machinery.
- Use when: Inspection or testing has a known or credibly nonzero miss rate; Production volume makes at least one field escape materially likely; Lots, versions, suppliers, or affected units can be traced and bounded after discovery.
- Typical domains: manufacturing, medical devices, food safety, software release engineering
- Common mechanisms: opportunity exposure register, sentinel event monitoring, post incident recurrence review
Near names: Eventual-Occurrence Containment, Cumulative-Exposure Containment, Repeated-Trial Hazard Containment, Nonzero-Probability Event Containment, Long-Horizon Residual-Risk Design, Assume-Eventual-Occurrence Design.