Skip to content

Tensions in Practice: Automatic recovery in tension with fault investigation

Software dependency · reset after automatic interruption

A protective breaker has already stopped requests to a failing dependency. A timed reset can restore service after a momentary fault, but can also reconnect into the same persistent failure. Holding the path open until someone inspects it avoids repeated automatic attempts while extending the interruption. The reset rule encodes a belief about what happened during the wait.

Resume after transient failure

Restore useful service without requiring attention for every momentary fault.

Avoid repeated re-exposure

Do not repeatedly reconnect while the underlying failure may still be present.

Why these aims pull against each other

Stopping the flow and restoring it are different decisions. Passage of time may fit a self-clearing fault model but does not establish that a persistent cause has been corrected.

Compare the arrangements

Reset automatically

Use a predeclared timer to reset after the automatic trip. A renewed fault trips the breaker again.

What it protects
A self-cleared transient fault need not wait for a person to restore the connection.
What it costs
Persistent failure can produce repeated reconnection and tripping, consuming work and exposing the dependency again.
When it fits
Fits faults credibly expected to clear with time and repeated attempts that are acceptable within the application’s limits.

Illustration note: This chosen model omits half-open probes, adaptive backoff and attempt caps. No reset interval or production policy is recommended.

Hold for inspection

Keep requests disconnected until inspection and an explicit reset authorize reconnection.

What it protects
The system does not automatically retry the same potentially persistent cause.
What it costs
Even a recovered dependency can remain unused while attention or a reset decision is unavailable.
When it fits
Fits persistent or uncertain fault causes where the cost of repeated attempts exceeds the cost of a longer interruption.

Illustration note: Inspection may be mistaken and is not itself repair. The reset decision must be appropriate to the actual fault and operating context.

What this illustration does—and does not—establish

Circuit Breaker: The reset policy embeds a hidden theory of the fault supplies the hidden fault model in reset policy; Circuit Breaker: Held-open safety and recoverability are in tension bounds the claim that suspension restores the system.

  • Both arrangements retain the same automatic trip rule; only the reset condition changes.
  • The sketch is a software dependency model, not instructions for electrical or other hazardous equipment.
  • Interruption does not diagnose or repair the underlying cause.

Source entries

Circuit Breaker

Prime · Source of the tension

Circuit Breaker: The reset policy embeds a hidden theory of the fault supplies the conflict examined here.

The reset policy embeds a hidden theory of the fault

T3: The reset policy embeds a hidden theory of the fault. An automatic reset assumes faults are transient and self-clearing; a manual reset assumes faults are persistent and require human inspection. Choosing wrongly is costly in opposite ways: auto-resetting into a persistent fault re-energizes the danger repeatedly (and can chatter), while requiring manual reset for a momentary glitch strands a healthy system in a safe-but-useless state and consumes scarce human attention. The reset gate is where the designer's beliefs about failure get encoded, often unexamined.

Read the source section

Held-open safety and recoverability are in tension

T6: Held-open safety and recoverability are in tension. The breaker's protective state is a held-open disconnection, but a system that is safe only because it is switched off has not been restored — it has been suspended. The longer and more reliably the breaker holds open, the more it protects against re-injury, yet the harder it becomes to bring the system back, especially if the held-open state itself causes secondary problems (a tripped reactor still needs cooling; a halted market still has positions to settle). Maximizing protective hold and minimizing time-to-recovery are competing goods that the reset architecture must somehow reconcile.

Read the source section