Metastable Failures in Distributed Systems.¶
Bronson, N., Aghayev, A., Charapko, A., & Zhu, T. (2021). Metastable Failures in Distributed Systems. Proceedings of the Workshop on Hot Topics in Operating Systems (HotOS '21), 221-227.
Cited by¶
2 citations across 2 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Empirical No-Failure Anchor
- Request rate had been adopted as the challenge variable, but it was not the axis along which the system was actually being stressed.
This sourceDocuments production systems entering sustained overload without any increase in load, driven by cache-state loss and retry amplification rather than by the request rate a capacity test had laddered.
- Request rate had been adopted as the challenge variable, but it was not the axis along which the system was actually being stressed.
- Metastability
- Computing and distributed systems — caches in a consistent-looking but stale state; configurations that survive small failures but lose data under one specific failure sequence; load that accumulates silently before a sudden avalanche.
This sourceNames metastable failures in distributed systems — a system trapped in a degraded state by a sustaining feedback (load amplification, sudden avalanche after silent accumulation) even after the trigger is removed.
- Computing and distributed systems — caches in a consistent-looking but stale state; configurations that survive small failures but lose data under one specific failure sequence; load that accumulates silently before a sudden avalanche.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:902c47cecd4c · see in the full table