Two-Key High-Harm Engagement¶
Dual control — instantiates Self-Targeting Defense Guardrail
Requires two independent authorities to concur before an irreversible defensive action fires, so no single classifier or operator can unilaterally harm protected self.
The most dangerous shape a defense can take is one switch that both labels a target and unleashes harm. Two-Key High-Harm Engagement breaks that switch in half. For actions above a harm threshold — irreversible, destructive, or high-blast — the mechanism refuses to fire on a single authority: it demands that two independent keys turn, a second concurring judgment from a role, model, or person that did not make the first call. Its defining move is to make unilateral high-harm engagement structurally impossible; the classifier's confidence is necessary but never sufficient, and the second key is not a rubber stamp but a genuine, independent re-examination that can veto. It is a gate on whether harm may begin, and its distinctive requirement is concurrence before the fact — two authorities agreeing up front — rather than any ability to stop the action once underway.
Example¶
A regional electric utility runs a protection scheme that can island a substation — cutting it from the grid — when relays detect a fault signature suggesting a compromised or malfunctioning feeder. Islanding a healthy substation blacks out a hospital district, so the scheme classes it as high-harm. Under Two-Key High-Harm Engagement, the automated relay logic can propose the island and arm it, but the breaker will not open until a second, independent key turns: a control-room operator who reviews the fault evidence on a separate console fed by independent telemetry and confirms it is not a sensor glitch or a legitimate load transient. Only when both the relay's automated key and the operator's human key are turned does the island execute.
Setup to outcome: a spiking fault score, on its own, opens nothing. On the night a miscalibrated current transformer produced a false fault signature, the second key caught the discrepancy against independent telemetry and declined to turn — the hospital district stayed powered. The distinction that matters is that the veto came from an authority the relay could not overrule.
How it works¶
- Threshold the requirement to harm level. Only actions above a defined harm-and-irreversibility line invoke dual control; reversible, low-harm responses fire on one key so the friction is spent where it counts.
- Enforce key independence. The two keys must draw on different evidence and answer to different authority — a human reviewing separate telemetry, or a second model with disjoint inputs — so a single compromised or mistaken source cannot turn both.
- Both keys are veto-capable. Concurrence, not majority: either key withheld blocks the action. The second key is a real re-examination with standing to override, not a notification.
- Arm-then-confirm sequencing. The first authority arms and proposes; the action stays inert until the second confirms, making the default state "not fired" whenever the two disagree or one is absent.
Tuning parameters¶
- Harm threshold for two-key — where the line sits between one-key and two-key actions. Lower it and more actions gain protection but throughput drops; raise it and rare catastrophic actions may slip through on a single key.
- Independence strength — how disjoint the two keys' evidence and authority are. Stronger independence resists correlated failure and collusion but is costlier to staff and slower.
- Concurrence window — how long the armed action waits for the second key before timing out. Longer windows tolerate slow reviewers but hold urgent defenses in limbo; shorter windows risk auto-abandoning valid engagements.
- Break-glass exception — whether a single key may proceed under extreme time pressure. Allowing it preserves speed against fast threats but reopens the unilateral-harm hole the mechanism exists to close.
When it helps, and when it misleads¶
Its strength is that it structurally prevents the archetype's signature failure — classifier authority collapse, where one model score both labels and destroys. It is the defensive form of the two-person rule[n1] long used for irreversible, catastrophic actions: two independent judgments must agree before the harm can occur, so a single error, spoof, or bad actor is not enough.
Its failure mode is that the second key degrades into ceremony. If the two authorities share evidence, defer to each other, or are pressured to move fast, dual control becomes single control wearing two badges — the second key turns automatically and the independence is fictional. The classic misuse is the "break-glass" override quietly becoming the default under operational pressure, so the high-harm action routinely fires on one key after all. The guarding discipline is to audit second-key veto rates and evidence independence, not merely whether two signatures exist: a second key that has never once declined is not a control, it is a formality.
How it implements the components¶
response_authorization_gate— it is the gate: high-harm action is unauthorized until two independent keys concur, decoupling suspicion from permission to harm.independent_review_or_override_path— the second key is an independent authority with standing to veto, drawing on separate evidence from the first.
It does not halt an action already underway — the always-reachable emergency stop that drops a firing response into a safe state is fail_safe_and_stop_authority, which belongs to Engagement Kill Switch; this mechanism guards the front door before harm begins, that one is the brake after it has started.
Related¶
- Instantiates: Self-Targeting Defense Guardrail — supplies the pre-action authorization-and-independence layer for irreversible responses.
- Consumes: Graduated Response Matrix — that matrix classifies which actions cross the harm threshold and therefore require two keys.
- Sibling mechanisms: Appeal and Rapid Restoration Workflow · Engagement Kill Switch · False-Positive Harm Budget Dashboard · Graduated Response Matrix · Post-Incident Autoimmune Review · Protected-Self Allowlist with Expiry · Quarantine-Before-Destroy Rule · Self-Status Cross-Check · Shadow Mode and Canary Enforcement
Editorial Notes¶
Form Classification¶
Form family: Decision, Gate & Allocation
Rationale: Two-Key High-Harm Engagement operates as a case-specific gate, selection, routing, prioritization, or resource disposition because it requires two independent authorities to concur before an irreversible defensive action fires, so no single classifier or operator can unilaterally harm protected self.
Independent corroboration: The frozen evidence defines Two-Key High-Harm Engagement as 'Requires two independent authorities to concur before an irreversible defensive action fires, so no single classifier or operator can unilaterally harm protected self', so its operative form is Decision, Gate & Allocation.
Nearest alternative: Rule, Policy & Commitment — Two-Key High-Harm Engagement includes features of a standing rule, threshold, contractual commitment, or policy constraint governing future conduct, but its defining operation is a case-specific gate, selection, routing, prioritization, or resource disposition.
Review outcome: Independent reviewer agreement; medium confidence.
Origin Attribution¶
Primary origin: Security Studies & Intelligence Analysis
Origin pattern: Single lineage
Present-day reach: Multi-domain
Rationale: Requiring two independent authorities before a high-harm action is a direct two-person security control. NIST and DOE define dual authorized surveillance or control specifically to detect unauthorized or incorrect procedures in high-consequence contexts; governance supplies authorization scope.
Related originating lineages:
- Computer Science & Software Engineering — Computer science and software-engineering practice supplies a parallel or contributing lineage for the mechanism's defining operation: requires two independent authorities to concur before an irreversible defensive action fires, so no single classifier or operator can unilaterally harm protected self.
- Law & Governance — Legal doctrine, regulatory governance, and procedural accountability supplies a parallel or contributing lineage for the mechanism's defining operation: requires two independent authorities to concur before an irreversible defensive action fires, so no single classifier or operator can unilaterally harm protected self.
- Military & Strategic Studies — Military planning, readiness, and strategic operations supplies a parallel or contributing lineage for the mechanism's defining operation: requires two independent authorities to concur before an irreversible defensive action fires, so no single classifier or operator can unilaterally harm protected self.
- Organizational & Management Science — organizational_management contributes organizational design, management, and operational governance to this mechanism's defining operation—Requires two independent authorities to concur before an irreversible defensive action fires, so no single classifier or operator can unilaterally harm protected self—without displacing the selected primary historical lineage.
- Systems Thinking & Cybernetics — Feedback, system boundaries, stocks, flows, and regulation supplies a distinct formative lineage for the mechanism's two key high harm engagement logic.
Review resolution: The blind reviewers disagree on primary lineage (organizational_management versus security_intelligence). Authoritative or primary research supports security_intelligence as the best historical origin: Requiring two independent authorities before a high-harm action is a direct two-person security control. NIST and DOE define dual authorized surveillance or control specifically to detect unauthorized or incorrect procedures in high-consequence contexts; governance supplies authorization scope. The cited NIST CSRC Glossary, Two-Person Control; U.S. Department of Energy, Nuclear Materials Control and Accountability directly supports the mechanism's defining operation. All independently supported contributing domains are retained without an arbitrary cap. origin_mode=single_lineage records lineage, while domain_reach=multi_domain records later applicability separately from provenance.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Researched adjudication after independent review; high confidence.
Sources consulted:
- NIST CSRC Glossary, Two-Person Control
- U.S. Department of Energy, Nuclear Materials Control and Accountability
Notes¶
[n1] The two-person rule (or two-man rule) requires that two authorized individuals both act to carry out a critical, irreversible operation, each able to veto — a control originally formalized for handling nuclear weapons. Its safety comes not from redundancy alone but from the independence of the two judgments. ↩