Value-Destruction Red Team¶
Test / assessment — instantiates Endogenous-Pie Payoff Design
Stress-tests whether the design invites sabotage, hold-up, gaming, retaliation, or externalized loss.
A Value-Destruction Red Team is an adversarial review whose job is to attack the design before reality does. A team is deputized to think like every party who might exploit the proposed cooperative arrangement, and to answer one question: given these rules, what is the most damaging thing a self-interested actor could rationally do — and does the design invite it? Its defining stance is negative and hostile: it does not look for ways to create value or improve the deal; it looks for how the deal, as written, rewards sabotage, hold-up, gaming a metric, retaliation, or quietly shoving costs onto outsiders. This is what separates it from every constructive mechanism in the family. It builds nothing and allocates nothing. It is the pre-mortem that surfaces the negative-sum strategies a hopeful "win-win" design tends to leave unexamined, so they can be closed off before the arrangement goes live.
Example¶
A video platform is about to launch a new creator revenue-share program: a pooled fund distributed to creators in proportion to watch-time, meant to grow the whole ecosystem — more creators, more viewers, more value for everyone. Before launch, a value-destruction red team is set loose on the rules with a single mandate: break them. They quickly find the openings a celebratory design review had missed. Watch-time is a metric, so it will be gamed — engagement-bait, artificially lengthened videos, view-farming rings that inflate hours without creating anything anyone wants. The pool is fixed, so it is zero-sum among creators — which turns "grow the ecosystem" into a scramble where creators have reason to mass-report rivals to suppress their share (retaliation and sabotage). The rules count only on-platform watch-time, so creators who bring in external audiences generate value the formula ignores while low-quality high-retention content is over-rewarded — an externality the boundary was drawn to miss. The team writes each of these up as a concrete exploit with the incentive that drives it. None of it is hypothetical hand-wringing; it is a list of moves a rational creator will make on day one. The program launches with quality gates, anti-farming detection, and an appeals process bolted on — because the red team named the hazards while they were still cheap to fix.
How it works¶
- Assign the adversary's mindset. Reviewers are explicitly tasked to want the design to fail, and to reason from each party's incentives rather than its stated good intentions.
- Enumerate exploits, not worries. Each finding is a concrete move — who does it, why it pays them, and what value it destroys — not a vague risk.
- Trace incentives to metrics and boundaries. Look hardest where a number is rewarded (it will be gamed) and where the design's boundary is drawn (costs will be pushed just outside it).
- Rank by damage and ease. Sort exploits by how much value they destroy and how cheaply an actor can run them, so the design team fixes the rational, high-damage attacks first.
Tuning parameters¶
- Adversary scope — which parties the team is allowed to role-play (insiders, counterparties, outsiders, future entrants). Wider scope catches more exploits but costs more and can drift into implausible threats.
- Independence — whether the red team is separate from the designers. Genuine independence surfaces uncomfortable findings the authors are blind to; embedding it is cheaper but risks pulled punches.
- Depth — a half-day challenge session versus a funded simulation of adversarial behavior. Deeper testing finds subtler exploits but delays launch and spends real resources.
- Externality reach — how far outside the direct parties the team looks for pushed-off costs. A wide reach protects third parties and the design's legitimacy; a narrow one is faster but lets externalized harm through.
- Fix threshold — how severe an exploit must be before it blocks launch versus being logged as a monitored risk. A high bar ships faster; a low one is safer but can stall a design under a pile of edge cases.
When it helps, and when it misleads¶
Its strength is that it counteracts the optimism baked into cooperative design: a "win-win" arrangement is authored by people who want it to work, and that very hope blinds them to the exploits a self-interested party will find immediately. A structured adversarial review is how latent negative-sum strategies — the hold-up problem, where one party, once another has sunk an irreversible investment, exploits that dependence to extract value — get named before they can be sprung.[n1]
Its failure mode is that a red team can cry wolf: presented with a wall of clever but far-fetched exploits, designers either over-armor the arrangement into something too rigid and distrustful to function, or tune out and miss the one attack that matters. An adversarial review can also become theater — a checkbox that lends false assurance without real independence or teeth. The guarding discipline is to rank findings by rational damage and ease of execution rather than mere ingenuity, to keep the team independent enough to deliver unwelcome news, and to feed its findings into concrete guardrails and monitors rather than treating the report as the safeguard itself.
How it implements the components¶
A Value-Destruction Red Team fills the adversarial-diagnosis subset — it hunts for how the design can be exploited; it builds, allocates, and commits nothing:
value_destruction_hazard_set— its core product: the enumerated, incentive-grounded catalog of sabotage, hold-up, gaming, and retaliation the design invites.opportunism_and_defection_monitor— by naming the concrete exploits, it specifies exactly what an ongoing monitor should watch for once the arrangement is live.externality_boundary— it probes where the design's boundary is drawn and finds the costs a party can shove just outside it onto third parties.
It does not implement value_creation_lever_set or surplus_allocation_rule — inventing value-creating moves belongs to the Mutual-Gains Negotiation Protocol and dividing the gain to the Gainsharing Contract; the red team only attacks the design, never builds or splits its surplus.
Related¶
- Instantiates: Endogenous-Pie Payoff Design — the adversarial stress test that keeps a cooperative design from being naive positive-sum rhetoric.
- Consumes: Joint Payoff Matrix Workshop — the payoff map gives the team the incentives it reasons from when hunting exploits.
- Sibling mechanisms: Gainsharing Contract · Mutual-Gains Negotiation Protocol · No-Harm Standstill Agreement · Public-Goods Contribution Rule · Shared Savings Pool · Side-Payment Compensation Package · Staged Reciprocal Commitment · Joint Payoff Matrix Workshop · Shared Success Dashboard
Editorial Notes¶
Form Classification¶
Form family: Experiment, Test & Rehearsal
Rationale: Value-Destruction Red Team operates as an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation because it stress-tests whether the design invites sabotage, hold-up, gaming, retaliation, or externalized loss.
Independent corroboration: The frozen evidence defines Value-Destruction Red Team as 'Stress-tests whether the design invites sabotage, hold-up, gaming, retaliation, or externalized loss', so its operative form is Experiment, Test & Rehearsal.
Nearest alternative: Assessment, Review & Assurance — Value-Destruction Red Team includes features of a bounded evaluation of existing evidence or work that produces a finding or disposition, but its defining operation is an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation.
Review outcome: Independent reviewer agreement; medium confidence.
Origin Attribution¶
Primary origin: Security Studies & Intelligence Analysis
Origin pattern: Cross-disciplinary synthesis
Present-day reach: Universal
Rationale: NIST AI RMF 1.0 documents that risk management uses adversarial challenge and impact analysis to surface how a system can destroy stakeholder value. This is direct, mechanism-specific evidence for security intelligence as the best-evidenced historical home of the operation—Stress-tests whether the design invites sabotage, hold-up, gaming, retaliation, or externalized loss.—rather than evidence merely that the operation is useful there. The retained alternates record genuine adjacent lineages; later portability is represented separately by domain_reach=universal.
Related originating lineages:
- Accounting & Auditing — Accounting and audit's variance, evidence, ledger, and assurance tradition contributes a separate formative lineage to the mechanism's value destruction red team logic.
- Computer Science & Software Engineering — Computer science and software-engineering practice supplies a parallel or contributing lineage for the mechanism's defining operation: stress-tests whether the design invites sabotage, hold-up, gaming, retaliation, or externalized loss.
- Economics & Finance — Economics Finance supplies a historically relevant adjacent lineage or formative practice for the operation—Stress-tests whether the design invites sabotage, hold-up, gaming, retaliation, or externalized loss.—but the adjudicated evidence more directly locates the defining lineage in security intelligence.
- Organizational & Management Science — Organizational design, management, and operational governance supplies a parallel or contributing lineage for the mechanism's defining operation: stress-tests whether the design invites sabotage, hold-up, gaming, retaliation, or externalized loss.
Review resolution: The blind reviewers disagree on primary lineage (economics_finance versus security_intelligence). The defining operation is: Stress-tests whether the design invites sabotage, hold-up, gaming, retaliation, or externalized loss. The researched NIST AI RMF 1.0 establishes that risk management uses adversarial challenge and impact analysis to surface how a system can destroy stakeholder value. That source therefore supports security intelligence as the historical origin. economics finance remains in the uncapped alternates where it contributes a formative practice, but application or governance is not itself proof of origin. origin_mode=cross_disciplinary_synthesis records lineage construction; domain_reach=universal separately records later applicability.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Researched adjudication after independent review; high confidence.
Sources consulted:
Notes¶
[n1] The hold-up problem (analyzed by Oliver Williamson and by Klein, Crawford, and Alchian) arises when one party makes a relationship-specific, irreversible investment, and the other party then exploits that sunk dependence to renegotiate terms in its favor. The fear of hold-up deters the value-creating investment in the first place — exactly the kind of latent value destruction a red team is meant to surface. ↩