Skip to content

Dual Modular Redundancy

A two-channel computing arrangement that compares corresponding results to detect faults, without by itself identifying which disagreeing channel is correct.

Version
v2 · 2026-10-03 · History
Domain-specific #
13171
Domain group
Applied Sciences & Engineering
Origin domain
Computer Science & Software Engineering
Subdomain
Fault Tolerant Computing → Computer Science & Software Engineering
Aliases
Double modular redundancy

Core Idea

Dual modular redundancy (DMR) executes the same intended computation through two corresponding modules and compares their results at a chosen observation point. A mismatch reveals that the pair is inconsistent, providing a fault signal. It does not, from the two results alone, establish which result is right. Two votes split one to one. NASA's JPL states this limit directly for two computers and contrasts it with a three-computer voter; the original diverse-DMR paper makes the same contrast for circuit modules.[1][2]

The abstraction is a detection architecture, not a promise that failures disappear. A design may retry, stop safely, invoke a third check, or use some independent fault locator after detection. IBM's z900, for instance, compares two duplicated instruction/execution units after each instruction and then uses checkpoints and retry in its broader recovery design. That recovery machinery is distinguishable from the bare pairwise comparison.[3]

Structural Signature

Sig role-phrases: functionally corresponding module pair — aligned inputs and observation points — result comparator — mismatch signal — unresolved two-way correctness — optional independent recovery.

  • Module pair. Two active channels implement the same intended function, whether physically identical or deliberately diverse. A spare that is idle and never compared is not enough.[2]
  • Alignment. The compared results must correspond to the same input or computational step. IBM compares its two I/E units at each instruction, while an FIR-filter design compares corresponding output samples.[3][4]
  • Comparator. An equality or application-appropriate equivalence check turns divergent outputs into a detectable event. The comparator itself may fail, and agreement is not proof of correctness if channels share a common-mode fault.[2]
  • Detection, not automatic correction. With ordinary two-output comparison, a mismatch does not identify the faulty channel. A third voter, an independent diagnostic, or deliberately distinguishable error patterns can change what follows, but those are additional conditions.[1][2]

What It Is Not

DMR is not triple modular redundancy: three outputs can support a majority vote under assumptions about independent faults, whereas two disagreeing outputs cannot. It is not generic duplication of hardware or data without comparison. Nor is it guaranteed failover: choosing the surviving channel requires independent evidence of which channel failed, a retry result, or another recovery protocol.[1][3]

It is also not identical to diverse DMR. Reviriego, Bleakley and Maestro describe diverse two-module circuits in which purposely different error patterns can, for certain soft faults, identify the affected path. That is a specialized augmentation of the two-channel idea, not a property of every DMR pair. Two chronometers aboard a ship may illustrate the logical tie, but the frozen candidate's historical adage was not verified from a primary source and is not used here as an engineering instance.[2]

Scope of Application

The arrangement applies where a computation can be performed twice and corresponding results compared. The compared unit may be an instruction result in a server processor or a sample from a digital filter; the comparison granularity and acceptable equivalence differ. The sources support these computing instances, not a universal rule that any pair of sensors, clocks or people constitutes DMR.[3][4]

The fault model matters. Independent transient faults often yield divergent outputs; the same wrong result from a shared cause may pass comparison. For numeric circuits with structural or input diversity, harmless rounding differences may need tolerance, and a loose tolerance can hide small errors. The 2013 paper states this explicitly for reduced-precision diverse designs.[2]

Clarity

DMR separates three questions often collapsed into “fault tolerant”: Was a discrepancy observed? Which channel is faulty? How will service recover? Two-way comparison answers only the first on its own. IBM's architecture makes the distinction visible: its comparator checks after every instruction, checkpoint/retry handles an error trigger, and a spare processor may be used after a processor failure. None of these later steps is inferred solely from a one-to-one split in outputs.[3]

The term “redundancy” here means concurrently available computational evidence, not simply an unused backup. Agreement is useful but conditional evidence: if both paths are wrong in the same way, the comparator sees no difference.[2]

Manages Complexity

Two channels plus a comparator are a compact reliability pattern. A system designer can reason about the pair's inputs, outputs, synchronization and mismatch handling without first choosing a whole recovery architecture. This decomposition lets detection be evaluated separately from retry, checkpoint, failover and voting.[1][3]

The compression also exposes hidden assumptions. The comparator needs aligned observations; the fault model must explain why errors diverge; recovery must not assume that a mismatch has named the bad replica. The original diverse-DMR work shows how much extra structure is needed when the aim moves from detecting a disagreement to locating and correcting selected errors.[2]

Abstract Reasoning

Let corresponding outputs be \(y_1=f_1(x)\) and \(y_2=f_2(x)\) for the same intended computation. A comparator reports \(d=1\) when \(y_1\) and \(y_2\) are judged different. The inference \(d=1\Rightarrow\) “at least one channel or comparison path is inconsistent with the other” is sound under the comparison model; \(d=1\Rightarrow\) “channel 1 is wrong” is not. Conversely, \(d=0\) means agreement under that comparator, not independently verified truth.[1][2]

This reasoning remains stable when the implementation changes from duplicated instruction units to two digital-filter realizations. It changes when an additional discriminator is introduced: a third independent vote or known distinctive error pattern can supply information the ordinary two-result relation lacks.[3][4]

Knowledge Transfer

The portable lesson inside computing is to separate cross-checking from fault localization. A designer reading IBM's dual execution can recognize the same pair/comparator skeleton in a signal-processing implementation, yet must ask which recovery assumptions transfer. IBM's instruction retry and checkpoint arrays do not become features of an FIR filter by analogy.[3][4]

The converse transfer is equally useful: structural diversity may expose faults that identical paths share, but Reviriego and colleagues' correction argument depends on engineered error patterns and its soft-error setting. Importing its correction conclusion into arbitrary replicated processors would be invalid.[2]

Examples

IBM z900 instruction execution. The processor unit has two duplicated instruction/execution units. A register-unit comparator compares their outputs after each instruction; agreeing results are checkpointed, while divergent results trigger an error and attempted retry. The spare-processor/application-preservation path is a separate recovery facility.[3] Mapped back: module pair = the two I/E units; aligned input/observation = corresponding instruction result; comparator = R-unit compare circuitry; mismatch signal = retry trigger; unresolved correctness = no two-way vote; optional recovery = checkpoint, retry and separately conditioned processor sparing.

FIR digital filtering. The original Structural DMR paper studies two functionally equivalent FIR-filter structures run in parallel and checks output-sample mismatches. It engineers different error patterns to go beyond ordinary pairwise detection for selected single soft errors; this extra diagnostic structure must not be read into every DMR pair.[4][2] Mapped back: module pair = two FIR realizations; aligned input/observation = common signal and corresponding samples; comparator = output mismatch detector; mismatch signal = discrepant samples; unresolved ordinary pair = the mismatch alone; optional recovery = pattern-based error location in this diverse specialization.

Structural Tensions

Detection coverage versus resource cost. A second active module and comparator can reveal divergent faults that a single path would not notice, but consume area, energy or execution resources. Removing the duplicate reduces those costs while giving up this direct cross-check.[1][2] Diagnostic: Is the expected detection value worth the duplicate-channel and comparison cost under the system's actual fault model?

Easy alignment versus common-mode resistance. Identical channels simplify lockstep comparison, but shared design or supply faults may make both return the same wrong result. Structural diversity can reveal some shared-mode failures or produce discriminating error patterns, yet adds design complexity and sometimes legitimate small output differences that a comparator must tolerate.[2] Diagnostic: Are independent or common-mode faults the main concern, and can a diverse pair be compared with a tolerance that neither falsely alarms nor hides relevant errors?

Structural–Framed Character

DMR lies toward the structural end of the structural–framed spectrum: the module-pair, aligned-output and comparator relation reappears across processor execution and digital filtering. Still, computing fault detection is its constitutive frame; the entry is not a generic philosophical claim that two witnesses disagree.[3][4]

It meets the five transfer tests with limits. (1) Mechanism: two corresponding outputs are checked. (2) Invariants: pair count, alignment and the underdetermined mismatch persist. (3) Parameter freedom: module scale and comparison granularity vary. (4) Cross-setting transfer: instruction and FIR cases instantiate the same detection skeleton. (5) Boundary preservation: recovery, fault probabilities and diversity-specific correction do not transfer without their own premises. Its character: a reusable computing architecture whose basic inference is detection without automatic localization.[1][2]

Structural Core vs. Domain Accent

The core is two functionally corresponding active channels, a comparable output relation and a mismatch signal that cannot alone designate the true output. Processor registers, filter taps, checkpoint arrays and soft-error waveforms are domain accents; none is required for every DMR instance.[3][4]

The live prime Self-Checking supplies a broader genus: a system tests its own result through partially independent paths and comparison. DMR narrows that genus to two computing modules. The still more portable two-witness disagreement skeleton, stripped even of computation and fault signaling, is a possible future-prime question, not silently admitted as this domain-specific node.

This entry is a kind of Redundancy (engineering) and is a kind of Self Checking.

Self-Checking is the proposed strict prime parent because the two physically distinct paths are partially independent for channel-local faults and produce an internal inconsistency signal. They are not independent against every fault: shared design and common-mode failures remain possible. Redundancy names the more general use of duplicate resources; Engineering Redundancy is the proposed domain-specific parent because it covers more fault-tolerance arrangements than two compared computational modules.

The relation is not “part of TMR.” A third module plus voting changes the decision rule, and two modules do not become a guaranteed corrector merely because one could add a third.[1]

Relationships to Other Abstractions

Local relationship map for Dual Modular RedundancyParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Dual ModularRedundancyDOMAINDomain-specific abstraction: Redundancy (engineering) — is a kind ofRedundancy(engineering)DOMAINPrime abstraction: Self Checking — is a kind ofSelf CheckingPRIME

Current abstraction Dual Modular Redundancy Domain-specific

Parents (2) — more general patterns this builds on

  • Dual Modular Redundancy is a kind of Redundancy (engineering) Domain-specific

    Two redundant modules are a restricted engineering-redundancy architecture used for fault detection.

  • Dual Modular Redundancy is a kind of Self Checking Prime

    Two physically distinct paths are partially independent for channel-local faults, and a comparator detects disagreement; common-mode failures remain possible.

Hierarchy paths (13) — routes to 8 parentless roots

Neighborhood in Abstraction Space

Dual Modular Redundancy sits in a sparse region of the domain-specific corpus (86th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (2551 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • Triple modular redundancy: three channels and majority voting under a single-fault model.[1][2]
  • Diverse DMR: a specialized two-channel design with engineered differences that may support fault localization/correction for selected circuit faults.[2]
  • Cold standby or failover: an unused spare or a transfer procedure, neither of which necessarily compares two concurrent results.[3]
  • Agreement as verification: two matching outputs can share the same error.[2]

References

[1] NASA Jet Propulsion Laboratory, ST8 Dependable Multiprocessor, “How does it work?” discussion of two versus three computers, accessed 2026-09-30. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i

[2] P. Reviriego, C. J. Bleakley and J. A. Maestro, “Diverse Double Modular Redundancy: A New Direction for Soft Error Detection and Correction”, IEEE Design & Test 30 (2013), abstract, Introduction, FIR discussion and conclusion. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q

[3] IBM, IBM eServer zSeries 900 Technical Guide, Appendix A, printed pp. 245–246, “Dual execution with compare” and recovery discussion. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l

[4] P. Reviriego, C. J. Bleakley and J. A. Maestro, “Structural DMR: A Technique for Implementation of Soft-Error-Tolerant FIR Filters”, IEEE Transactions on Circuits and Systems II 58 (2011), abstract and introduction. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g