Skip to content

Empirical No-Failure Anchor

Version
v5 · 2026-09-07 · History
Prime #
1355
Origin domain
Toxicology
Subdomain
experimental inference → Toxicology
Related primes
Absence Of Evidence Vs Evidence Of Absence, Dose-Response Relationship, Experimental Design, Projection

Core Idea

An empirical no-failure anchor is the highest tested input, exposure, or load at which a specified failure was not detected under a stated protocol. It is evidence for a lower bound on tolerance, not an estimate of where the system actually breaks. [1]

The number is a coordinate in a test record rather than a property of the thing tested. Three commitments had to be in place before it could exist: a challenge variable ordered so that levels can be ranked, a failure criterion written down in advance of observation, and an apparatus with some real chance of registering that failure had it occurred. Remove any one of them and what remains is merely a level at which nothing was noticed, which is a far weaker object than it looks. [2]

The asymmetry between the two directions is the whole content of the prime. A detected failure at a level identifies a level at which the system fails; a non-detection identifies nothing and only excludes. So the value cannot travel alone: tested levels and their spacing, sample size, observation window, endpoint set, and detection sensitivity each move the reported figure while the tested system stays exactly as it was, which makes them part of the claim rather than metadata hanging off it. [1]

Structural Signature

Ordered challenge ladder → pre-specified failure criterion → characterized detector → highest non-detecting level → one-sided bound carrying a named residual.

The shape is a reduction under an extremal rule: an entire test-response record collapses to one scalar, and the rule that picks the scalar is "largest level that cleared the criterion." [1] That rule has a consequence fitted curves and group means do not share — the result can only ever land on a level someone chose to run, so the anchor's resolution is capped by the spacing of the ladder no matter how many units were tested at each rung. [2]

Recurring features:

  • An ordered challenge variable — dose, static load, temperature, cycles, concurrency, packet rate — along which levels are comparable as higher and lower.
  • A failure criterion fixed before observation, so that "nothing happened" has a definite referent instead of being read off the data afterwards.
  • A detection apparatus whose probability of registering the criterion when present is itself characterized; the bound is worth exactly as much as that probability.
  • An extremal selection rule that keeps the largest passing level and discards every lower passing level as carrying no additional information.
  • Lattice dependence: the value is pinned to the tested grid, quantized at the grid's spacing, and can be raised by adding a rung rather than by improving the system.
  • A named residual — the untested interval above the anchor, plus the gap up to the lowest failing level wherever one exists.
  • One-sidedness: the value bounds the latent onset from below and stays silent about its location above.

What It Is Not

The anchor does not assert that the system is safe at the anchor level. Safety is a judgement combining a consequence, an acceptable probability, and a buffer someone chose; the anchor reports only what a particular protocol did and did not register. [2]

It does not assert that failure begins just above. Nothing in the construction places the latent onset near the top passing rung: the next rung may be a factor of three away, and the true onset may lie anywhere in the gap or far beyond it. The impression of a cliff comes from the layout of the results table, not from the evidence.

It does not certify every level beneath it. That certification requires monotonicity in the challenge variable, an assumption that fails wherever response is banded rather than graded — resonance windows, biphasic dose effects, buffers that overflow at one arrival rate and drain cleanly at a higher one. Where monotonicity is not separately argued, the anchor covers its own rung and no other.

It is not strengthened by being repeated. A value quoted downstream without its sample size, endpoint, and sensitivity has been laundered rather than confirmed, and the confidence readers place in it grows with circulation while its actual support does not. [3]

It is not a measurement of the system. Two competent laboratories testing identical material with different grids, group sizes, and endpoints will honestly publish different anchors, and neither has made an error.

It does not become a threshold estimate by being divided. Applying a safety factor to an under-informative number yields a smaller under-informative number; the operation changes the magnitude and leaves the epistemic status untouched.

Finally, it does not arise from accumulated experience. Uninstrumented service history with no stated criterion and no exposure record produces recollection, not an anchor, however long the record runs.

Broad Use

Toxicology and pharmacology. Graded dose groups form the ladder, a pre-registered adverse endpoint forms the criterion, and histopathology or clinical chemistry acts as the detector; the highest dose group without a significant adverse finding is reported as a no-observed-adverse-effect level, the field's named subtype of this pattern. [1]

Reliability and qualification engineering. Temperature, vibration amplitude, humidity, or accumulated cycles supply the ordering; the criterion is a functional or parametric out-of-specification condition; the detector is the post-stress screen, which is why a qualification result is only ever as strong as the test coverage run afterwards.

Structural and pressure testing. A proof load applied and held is a single-rung ladder. The article either shows permanent set, leakage, or rupture or it does not, and a passed proof establishes a bound while supplying no information about how much further the article would have gone.

Security and capacity testing. Fuzzing campaigns, load ramps, and red-team engagements generate highest-clean-level claims constantly, and are unusually exposed to the detector problem, because a crash, a breach, or a latency violation can be present in the run and never surface in the instrumentation.

Clinical dose escalation. Cohorts ascend a dose ladder until dose-limiting toxicity appears, so the highest cleared dose and the first toxic dose arrive together. This is one of the few settings where the two-sided bracket is produced by design rather than by luck. [4]

Environmental and food-safety regulation. Residue limits, emission ceilings, and exposure guidance are commonly derived from a value established in one species, matrix, and duration, then carried through a chain of extrapolations the originating protocol never anticipated.

What unites these is not shared vocabulary. It is a common economic situation: a decision-relevant boundary is unknown, locating it directly is expensive or unethical, and passing a test at a known level is cheap by comparison — so practitioners buy a bound instead of an estimate, and agree to remember which one they bought.

Clarity

The confusion this prime dissolves is the slide from "we looked and saw nothing" to "there is nothing there," performed on a number that looks far too precise to be doing anything so loose. [5] Once a result is written as 100 mg/kg or 10,000 requests per second it acquires the grammar of a measurement, and readers bring the intuitions they bring to measurements: that it is a fact about the object, that a more careful worker would recover the same value, that the error bars are narrow. Each of those intuitions is wrong here, and none of them announces itself.

Naming the pattern separates three questions that the single figure silently fuses. Where does the system actually fail? Unknown, and this test was not built to answer it. How high did we get without seeing trouble? That is the anchor, exactly and only. How hard were we looking? That is the detection term, and it is the part most often left behind at the first retelling.

The clarification also disarms a recurring rhetorical move. In disputes about hazard, "no effect has been demonstrated at these levels" is offered as though it were the mirror image of "an effect has been demonstrated." Holding the structure in view makes the asymmetry visible without anyone needing to argue the substance: the two statements are outputs of different logical operations, and the strength of the negative one is a function of experimental parameters that can simply be read off the protocol.

Manages Complexity

What the anchor buys is permission to stop carrying the response surface. A dose-response campaign or a load-test programme produces a large object — every unit, every rung, every endpoint, every replicate — and downstream users can neither hold that object, reason with it, nor pass it on. The reduction to one number plus a short provenance card is defensible precisely because the extremal rule is deterministic: anyone given the same record and the same criterion recovers the same value, so the compression itself adds no discretion beyond the choice of criterion, which no general rule supplies. [1]

It also lets you stop tracking mechanism. Nothing in the construction requires knowing why the system fails, which failure mode dominates, or what shape the curve takes between rungs. Those questions can remain open indefinitely while the bound stays usable, which is why the pattern persists in fields where mechanism is genuinely unavailable and in the early phases of fields where it is merely not yet known.

The cost is bookkeeping that must not be dropped. The provenance card is not decoration; it is the remainder term of the reduction, and a figure circulating without it has quietly become a different and unsupported claim. The complexity was never eliminated. It was moved into a short fixed list of conditions that travels with the value — a far better place for it than a dataset nobody will reopen.

Abstract Reasoning

The pattern licenses a short diagnostic that can be run against any claim of the form "it held up to X." Confirm that the challenge variable is genuinely ordered and that X names a level on it. Confirm that the failure criterion existed before the observation window opened. Ask what the apparatus would have caught: at the sample size, window length, and endpoint sensitivity used, what rate of failure would have escaped notice? Locate the lowest level at which failure was detected, if one exists, and report the interval instead of its endpoint. Then state the residual out loud — the region above is untested, not cleared.

The load-bearing step is the third. With zero events observed in n independent units, the one-sided ninety-five-percent upper bound on the underlying rate is approximately three divided by n, so ten clean units remain consistent with a failure rate near thirty percent while a hundred clean units narrow it to roughly three; the anchor level is identical in both cases and the claim differs by an order of magnitude. [6] Any argument that treats those two records as interchangeable has dropped the detection term without noticing.

The same procedure disqualifies the most frequent impostor. "The system has never failed in ordinary use" fails at the first step: ordinary use is not a level on a ladder, the exposures were never recorded, and no criterion was fixed, so there is nothing for a bound to attach itself to. Uncontrolled non-failure is fully compatible with a system that has never once been meaningfully stressed.

Knowledge Transfer

Four things move intact across substrates: the role structure of ordered challenge, fixed criterion, characterized detector, and one-sided bound; the extremal selection rule; the quantization of the result at ladder spacing; and the requirement that provenance move with the value. An engineer reading a toxicology dossier and a toxicologist reading a qualification report can each locate all four without learning the other's units, and that is what makes the pattern portable rather than merely suggestive. [7]

Three things do not move and have to be re-established wherever the pattern lands. Monotonicity is the first, since dose-response biology, structural loading, and thermal cycling each have their own reasons for a response to be graded or banded, and importing the assumption unexamined is the most common transfer error. The detector model is the second: a histopathological screen, a post-stress functional test, and a fuzzing harness miss a present failure for entirely unrelated reasons, so the zero-event arithmetic that turns non-detection into a rate bound needs a fresh independence argument each time. The third is downstream convention — some fields apply large default factors to such a value by regulation, others feed it raw into a separate risk model, and carrying one field's habit into another changes what the number means without changing how it is written.

The practical rule that follows is to transfer the skeleton and re-derive the arithmetic. Recognizing the pattern in an unfamiliar field costs nothing and immediately supplies the right questions; assuming that field's anchors were built with the same sensitivity discipline as your own is where the transfer reliably goes wrong.

Examples

Formal/abstract

A subchronic study runs dose groups at 0, 10, 30, 100, and 300 mg/kg/day, ten animals per group, with a pre-specified histopathological lesion as the failure criterion and a ninety-day observation window. No lesion appears at 100. Four of ten animals show it at 300. The recorded value is 100 mg/kg/day and the bracket is the half-open interval from 100 up to and including 300.

Three properties of that record are easy to lose. The strength of the bound is set by group size and not by the value itself: with zero events in ten animals, an underlying lesion incidence approaching a quarter of the population remains entirely compatible with the observation, so a clean 100 group excludes much less than it appears to. [6] The value is quantized by the grid: had 200 been included and passed, the figure would have doubled with no change whatever in the animals' biology, and the honest reading is that the study resolves the boundary only to within a factor of three. And the residual is an interval rather than a point — the lesion could begin at 105 or at 295, and nothing in the design distinguishes those worlds.

Now vary one element. Keep the same doses and the same animals but replace the endpoint with a more sensitive assay that registers a precursor biochemical change. The 100 group is no longer clean and the value falls to 30. The system did not become more dangerous; the apparatus became better at looking, and the figure moved down because it had always been partly a statement about the apparatus.

Mapped back: each property corresponds to one role in the signature. Group size is the detector term, grid spacing is the ladder, and the interval is the named residual. The endpoint substitution then demonstrates the defining commitment directly: this is a joint statement about a system and a protocol, and holding the system fixed while varying the protocol moves the result — which is exactly why a bound, and never a threshold estimate, is all it can supply.

Applied/industry

A team qualifies a request-serving service before launch. The ladder is offered load at 1k, 2k, 5k, 10k, and 20k requests per second. The criterion is fixed in advance as p99 latency above 500 ms or an error rate above 0.1%, sustained for ten minutes. The detector is the existing metrics pipeline at one-minute resolution. The service is clean at 10k and breaches at 20k, and the team records a capacity figure of 10k rps.

Six weeks later the service degrades in production at roughly 7k rps. Nothing in the qualification was falsified. The synthetic traffic used a uniform key distribution while real traffic was heavily skewed, concentrating work on a few partitions; caches were warm throughout the test and cold after each deploy; the ten-minute window was shorter than the time a connection pool needed to saturate; and one-minute metric resolution averaged away latency spikes that lasted seconds. Request rate had been adopted as the challenge variable, but it was not the axis along which the system was actually being stressed. [8]

Mapped back: the result bounds the latent onset only along the axis that was laddered and only under the conditions that were held. A system whose stress is multidimensional — arrival rate, key skew, cache state, sustained duration — has one such bound per axis tested and none at all for the joint space, and the untested combinations belong to the residual whether or not anyone wrote them down. The incident is not evidence that the qualification was wrong. It is evidence that the residual was larger than a single scalar suggested, which is the condition this prime exists to keep in view.

Structural Tensions

T1 — Ladder resolution versus the cost of a rung. Every additional tested level narrows the interval left open, and every additional level costs animals, articles, hours, or budget. Coarse geometric ladders — factors of three, factors of ten — are standard precisely because they span a wide range cheaply, and they guarantee that the reported bound is quantized at that spacing. Refining the grid near the suspected boundary buys resolution where it matters but forfeits range, and the choice must usually be made before anyone knows where the boundary actually sits.

T2 — A better detector lowers the number. Improving sensitivity through a finer endpoint, a longer window, larger groups, or higher-resolution instrumentation raises the chance of registering a failure that was present all along, so the highest clean level tends to fall. A laboratory that upgrades its methods will report lower values than it did before, and the same material will look worse under better looking. Read naively, that sequence reads as deterioration or as an earlier mistake. Read correctly, the newer and smaller figure is the more informative one.

T3 — Weak protocols produce flattering results. Every dimension that weakens a test raises the reported level: fewer units, shorter windows, coarser endpoints, wider rung spacing, less sensitive instruments. Whoever benefits from a high figure therefore benefits from a poor experiment, and the benefit arrives with no misconduct whatever — no data altered, nothing concealed, the protocol merely undemanding. The provenance card is the only defence, and the reviewer's question is not whether the result is honest but whether the design could have detected what it failed to detect.

T4 — Taking the maximum across repeats. The extremal rule is well behaved on one record and pathological across many. Run the same ladder three times, take the highest clean level seen in any run, and the reported figure drifts upward with the number of runs, driven by sampling noise rather than by tolerance. The same drift appears when several laboratories test the same article and the most favourable outcome is the one cited. Guarding against it means fixing in advance how repeats combine — pooling units, or taking the minimum — before any run is seen.

T5 — A one-sided bound asked to do a threshold's work. Decisions need a number to compare against, and this is the only number available, so it gets used as though it located the boundary. Downstream conventions formalize the substitution: divide by a factor, call the result a limit, and a bound has become a point estimate with its uncertainty converted into a constant. The conversion is often the best move available, but it remains a policy choice, and the interval it replaced does not shrink because a factor was applied to its lower end.

T6 — Provenance strips off in transit. The value is a scalar and travels effortlessly; its conditions are a paragraph and travel badly. Each citation hop drops another qualifier — first the endpoint, then the group size, then the species or the traffic model — until a protocol-relative bound is quoted as a property of the substance or the service. Nobody performs the stripping deliberately; it is simply what summarization does to a number with a long tail of caveats. The countermeasure has to be structural: bind the conditions into the reported name and refuse to publish the bare scalar.

Structural–Framed Character

Empirical No Failure Anchor sits at the midline of the structural–framed spectrum — mixed-framed, aggregate 0.5, with every diagnostic at 0.5. What travels is a compact skeleton: an ordered ladder of tested challenge levels, a pre-specified failure criterion, a characterized detection apparatus, and the highest level at which that criterion went undetected — a value that constrains but does not identify the latent failure threshold. Toxicology, reliability testing, load tests, and dose escalation preserve those challenge–criterion–detection–bound roles exactly.

Human-practice-bound at 0.5 explains the reading best: the anchor exists only relative to a stated protocol. Its provenance lies in the test grid, sample, duration, endpoint set, and sensitivity, so changing the regime moves the anchor without changing the tested system, and the non-detection is worth only what the apparatus would have caught.

Vocabulary travels partially at 0.5 — no-observed-adverse-effect level, censoring, dose escalation carry a toxicological and reliability accent each substrate restates. Evaluative weight is 0.5: the criterion is an adverse or unacceptable response someone specified, though the prime stops short of the decision claim a maximum safe load makes. Institutional origin is 0.5, those test disciplines being its home. Import-versus-recognize is 0.5: where a challenge ladder already runs one recognizes the bound in the record; elsewhere the protocol is imported with it.

The grade means the number never travels alone: report it as a protocol-relative lower bound on tolerance, with sample size, sensitivity, and censoring attached — never as an estimate of where failure begins.

Substrate Independence

Empirical No-Failure Anchor is about as substrate-independent as a prime can be — composite 5 / 5 on the substrate-independence scale. What travels is not a fact about the thing tested but an extremal rule over a test record: rank the challenge levels, fix the failure criterion before observing, characterize the detector, keep the highest non-detecting level, and report it as a one-sided bound with a named untested residual above it. Toxicology, reliability trials, structural load testing, cybersecurity stress testing, and dose escalation reproduce those roles exactly, down to the lattice dependence pinning the value to the rungs actually run. The three sub-scores converge because the prime concerns measurement rather than any measured system: wherever a ladder is climbed until something fails to appear, the same object has been built.

  • Composite substrate independence — 5 / 5
  • Domain breadth — 5 / 5
  • Structural abstraction — 5 / 5
  • Transfer evidence — 4 / 5

Relationships to Other Abstractions

Current abstraction Empirical No-Failure Anchor Prime

Parents (4) — more general patterns this builds on

  • Empirical No-Failure Anchor is a kind of Projection Prime

    The anchor is a projection specialized to reducing the full test-response dataset to the maximal tested level that cleared a stated failure criterion.

  • Empirical No-Failure Anchor is part of Absence Of Evidence Vs Evidence Of Absence Prime

    The anchor contains a null finding whose boundary force exists only through the test apparatus's probability of detecting a present failure.

  • Empirical No-Failure Anchor presupposes Dose-Response Relationship Prime

    A maximum passed input level presupposes an ordered mapping from input intensity to the specified response or failure criterion.

  • Empirical No-Failure Anchor presupposes Experimental Design Prime

    The anchor presupposes a designed investigation that fixes tested levels, sampled units, duration, endpoints, replication, and the detection apparatus.

Children (1) — more specific cases that build on this

  • NOAEL (No Observed Adverse Effect Level) Domain-specific is a decomposition of Empirical No-Failure Anchor

    NOAEL is the toxicological framed form of the empirical maximum-passed-level anchor, adding dose groups, adverse endpoints, LOAEL, and regulatory factors.

Hierarchy paths (14) — routes to 8 parentless roots

Neighborhood in Abstraction Space

Empirical No-Failure Anchor sits in a sparse region of abstraction space (78th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely rather than landing on a neighbor.

Family — Tail Risk & Long-Horizon Forecasting (11 primes)

Nearest neighbors

Computed from structural-signature embeddings · 2026-09-10

Not to Be Confused With

The closest neighbour is Absence of Evidence vs Evidence of Absence, and the relation is containment rather than contrast. The anchor holds a null finding inside it, and everything that null finding is worth derives from the detection-side counterfactual that neighbour names. What separates them is scope and output shape. That prime is fully general: it governs any search, ordered or not, for any claim, and what it returns is a weight on a negative conclusion. The anchor is that logic specialized to a ladder — the search is a designed challenge at ranked levels, the claim is a pre-specified failure, and the output is not a weight but a coordinate on an input axis, below which the boundary has been excluded. Strip the ordering and the anchor dissolves back into its parent: an interpretable silence remains, but there is no axis left for it to constrain.

Dose–Response Relationship is presupposed rather than resembled, and confusing the two costs most of the caution. A dose-response relationship is the whole mapping from input intensity to response across the studied range — its shape, its slope, its inflections, its confidence band. The anchor is one selected coordinate on that mapping plus the conditions that produced it. Someone holding the full relationship can interpolate, extrapolate with stated error, and estimate an effect at an untested level; someone holding only the anchor can do none of those things. Reporting the coordinate is the normal outcome when the study was never powered to characterize the curve, and treating a coordinate as though a curve stood behind it is the error the distinction exists to block.

Threshold names the thing the anchor gestures at and never touches. A threshold is a property of the system: the level at which its behaviour actually changes, which exists whether or not anyone has tested for it. The anchor is a property of a test record, which exists only because someone ran a protocol. The two would coincide only under an infinitely fine grid, an infinitely sensitive detector, and unlimited units — which is another way of saying they never coincide. The clearest symptom of the conflation is language: a threshold can be estimated, refined, and converged upon; an anchor can only be raised by testing higher or lowered by looking harder.

Margin of Safety sits downstream and is a different kind of claim altogether. A margin is a deliberate buffer between an operating point and a boundary, chosen against a tolerated consequence and a tolerated probability. The anchor supplies one input to that choice and contains no buffer of its own. The order matters: an anchor established under a protocol, then divided by default factors, then declared an operating limit has passed through two separate acts of judgement that the surviving number no longer displays. Reading a margin backwards as though it revealed where failure begins reconstructs a threshold from a policy decision.

Extrapolation Beyond Sampled Regime describes the characteristic way the anchor is misused rather than a rival reading of it. That prime concerns an apparatus applied outside the regime in which it was calibrated while still reporting its in-regime confidence. An anchor carried into a different species, a different duration, a different traffic mix, or a different failure mode is exactly that: the value keeps its original precision on the page while the conditions that gave it meaning have been left behind. The anchor is well formed inside its regime; the extrapolation failure is what happens at the boundary of that regime.

Benign Sampling Safety Drift is the anchor's degenerate cousin, and the contrast is the sharpest one available. Drift feeds on uneventful outcomes that were never a test — quiet operation near a hazard boundary, read as proof the buffer was unnecessary. The anchor's non-failure is manufactured deliberately, at a level someone chose, against a criterion someone fixed, through a detector someone characterized. Both produce the sentence "nothing went wrong." Only one of them earns it, and the difference lies entirely in whether the silence was designed.

Two further neighbours mark the edges. Projection is the general operation the anchor instantiates — mapping a richer object onto a lower-dimensional target along a fixed direction with a residual that must stay named — and the anchor simply fixes the source, the target, the direction, and the residual to specific things. Statistical Power quantifies the detection term rather than competing with it: power is a property of a design, while the anchor is a value that design emitted, and quoting the value without the power is what leaves the claim unfalsifiable in practice.

Solution Archetypes

No catalogued solution archetypes reference this prime yet.

Notes

Contested claims

  • [086] Crump (1984) - the paper the entry itself cites throughout for this pattern - states that there are no general guidelines or rules for defining the level, and that determination is particularly uncertain where the lesion occurs naturally in untreated animals. Taken together with the EPA's finding that doses meeting NOAEL criteria carry real response rates of roughly 5-20 percent above control, the criterion is a locus of judgement rather than a fixed input, so two competent assessors can derive different anchors from one record. The claim is defensible only under the stipulation that the criterion is fully specified, which the softened sentence now makes explicit rather than assuming.

References

[1] Crump, Kenny S. "A New Method for Determining Allowable Daily Intakes". Fundamental and Applied Toxicology 4(5), 1984. Shows that the no-observed-effect level has no general defining rule, is particularly uncertain where the lesion also occurs in untreated animals, and rises as experiments get smaller, so it records what a protocol did not register rather than where the system fails. registry ↩a ↩b ↩c ↩d ↩e

[2] U.S. Environmental Protection Agency. Benchmark Dose Technical Guidance. EPA/100/R-12/001, 2012. States that the NOAEL is limited to one of the doses included in the study and so is pinned to dose selection, that it tends to be higher in studies with fewer animals per group because detection power falls with sample size, and that observed response rates at doses meeting NOAEL criteria run about 5-20 percent above control rather than zero. registry ↩a ↩b ↩c

[3] Greenberg, Steven A. "How citation distortions create unfounded authority: analysis of a citation network". BMJ 339, 2009. Traces how repeated citation of a claim detached from its supporting data converts a weakly supported statement into an accepted fact without any new evidence being produced. registry

[4] Storer, Barry E. "Design and analysis of phase I clinical trials". Biometrics 45(3), 1989, pp. 925-937. Describes escalation designs in which cohorts ascend a dose ladder until toxicity is observed, so the stopping dose and the highest cleared dose are produced together by the design rather than by luck. registry

[5] Altman, Douglas G., and J. Martin Bland. "Absence of evidence is not evidence of absence". BMJ 311, 1995. Distinguishes a study having shown no difference from a study having failed to show one, and ties the strength of the negative reading to the power of the design that produced it. registry

[6] Hanley, James A., and Abby Lippman-Hand. "If Nothing Goes Wrong, Is Everything All Right? Interpreting Zero Numerators". JAMA 249(13), 1983. Derives the rule of three, under which zero events in n units leaves rates up to roughly three over n consistent with the observation, and tabulates 0 of 10 as excluding only rates above about 26 percent. registry ↩a ↩b

[7] Dixon, W. J., and A. M. Mood. "A Method for Obtaining and Analyzing Sensitivity Data". Journal of the American Statistical Association 43(241), 1948, pp. 109-126. States a sensitivity-testing design and estimator purely in terms of ordered challenge levels and a binary pass/fail criterion, with no substrate commitment, and remains the basis of up-and-down dose-finding procedures decades later and in unrelated fields. registry

[8] Bronson, Nathan, Abutalib Aghayev, Aleksey Charapko, and Timothy Zhu. "Metastable Failures in Distributed Systems". HotOS '21, 2021. Documents production systems entering sustained overload without any increase in load, driven by cache-state loss and retry amplification rather than by the request rate a capacity test had laddered. registry