Stress-Test Scorecard¶
Scoring artifact — instantiates Assumption Stress Testing
A one-page verdict sheet that consumes the results of the stress tests and gives each key assumption a confidence grade, a reversibility flag, and a disposition — safeguarded, monitored, or knowingly accepted.
A Stress-Test Scorecard is not a test — it is the verdict that comes after the tests. It takes the outputs of whatever stressing was done — a scenario run, a sensitivity sweep, a red-team case, a tabletop's findings — and reduces each key assumption to a single, legible row: how much confidence we now hold in it and on what evidence, how reversible the commitment riding on it is, and what we have therefore decided to do about it. Its one defining move is disposition: for every load-bearing premise it forces a choice among a fixed menu — safeguarded, monitored, or explicitly accepted as a known risk — and records that choice with the reasoning behind it. The scorecard runs no scenario and breaks no premise; it grades premises already stressed and turns a pile of test results into a decision-ready one-pager a sponsor can sign.
Example¶
A pharmaceutical company is at the go/no-go gate for launching a new therapy, and a stack of stress-test outputs has piled up: forecasts under different reimbursement scenarios, a sensitivity sweep on manufacturing yield, a red-team memo on the approval timeline. Nobody can act on the pile as-is. The scorecard collapses it into one page, a row per key assumption. Regulatory approval lands by Q3 — confidence: medium, evidence: the agency's written guidance but no formal decision; reversibility: low, because the launch supply build commits capital months ahead; disposition: safeguarded, with a delayed-approval fallback plan. Payers reimburse at the target tier — confidence: low, evidence: two advisory-board opinions; reversibility: high, since pricing can be revisited; disposition: monitored, with a payer-signal trigger. Manufacturing scales to full yield in six months — confidence: high, evidence: a completed pilot run; reversibility: moderate; disposition: accepted risk, logged with the reasoning that the residual exposure is tolerable. The one page is what the launch committee actually signs — a graded, dispositioned summary, not another test.
How it works¶
- One row per key assumption, drawn from the inventory the stress tests examined — the scorecard is a summary layer, so it inherits its list rather than generating one.
- Grade confidence against evidence. Each premise gets a confidence level paired with the evidence that justifies it, so a high-confidence row backed by two opinions cannot pass as settled fact.
- Flag reversibility. Mark how reversible the commitment resting on each premise is; low-reversibility rows must clear a higher bar before any disposition but "safeguard" is allowed.
- Assign a disposition. Force each row to a fixed menu — safeguard, monitor, or accept-as-known-risk — and record the reasoning, especially for accepted risks.
Tuning parameters¶
- Confidence scale — a three-band traffic light versus a calibrated probability. Bands are legible to a signing sponsor; probabilities discriminate more but imply precision the evidence rarely supports.
- Reversibility gate strength — how hard a low-reversibility premise is barred from an "accept" disposition. A strict gate protects against irreversible bets on thin evidence; a lax one lets convenience through.
- Disposition menu — how many options a premise can be routed to, and whether "accept" requires sign-off. Requiring a signature on accepted risks makes the acceptance conscious rather than default.
- Evidence-citation rigor — a one-word source tag versus a linked basis. More rigor stops confident-sounding rows from hiding thin backing, at the cost of upkeep.
When it helps, and when it misleads¶
Its strength is that it makes the residual picture legible and forces the honest, uncomfortable act this archetype exists for: naming which premises are being knowingly accepted on thin evidence rather than leaving that acceptance implicit.[n1] It converts scattered test outputs into a single accountable page a decision-maker can actually sign.
Its failure mode is that a scorecard summarizes, and a summary can launder: a green cell or a confident band can paper over a test that never really bit, giving a plan the look of having been stressed without the substance. Reversibility flags get rubber-stamped, and "accepted risk" becomes a quiet dumping ground for premises no one wanted to safeguard. The classic misuse is building the scorecard after the decision to make it look diligent. The guarding discipline is that the scorecard is only as honest as the tests feeding it and the sign-off behind each accepted risk — the grade must trace to a real test, and an accepted-risk row must carry a named owner who put their name to accepting it.
How it implements the components¶
confidence_and_evidence_rating— its core column: each assumption gets a confidence level bound to the specific evidence that warrants it, so support and confidence are graded together rather than asserted.reversibility_guardrail— the reversibility flag raises the required bar for any premise whose underlying commitment is hard to undo, refusing casual acceptance of irreversible bets.accepted_assumption_risk_note— the "accept" disposition records, with reasoning and an owner, the premises the plan is knowingly proceeding on despite residual risk.
It runs no test of its own: it does not construct a scenario or break a premise (stress_scenario, assumption_break_test) — that is Scenario Stress Test, its nearest twin, which generates the evidence this scorecard grades. It also does not monitor those premises over time after the verdict is signed (monitoring_trigger, assumption_owner) — that is Trigger Dashboard.
Related¶
- Instantiates: Assumption Stress Testing — the scorecard is where the archetype's testing converges into a graded, dispositioned decision.
- Consumes: Scenario Stress Test and Sensitivity Analysis Workshop supply the test results the scorecard grades.
- Sibling mechanisms: Scenario Stress Test · Sensitivity Analysis Workshop · Trigger Dashboard · Failure Mode and Effects Table · Premortem · Red-Team Future Challenge · Resilience Tabletop Exercise · Assumption Register
Editorial Notes¶
Form Classification¶
Form family: Assessment, Review & Assurance
Rationale: Stress Test Scorecard is defined in the frozen evidence as: A one-page verdict sheet that consumes the results of the stress tests and gives each key assumption a confidence grade, a reversibility flag, and a disposition — safeguarded, monitored, or knowingly accepted. Its operative deployed or enacted form is therefore Assessment, Review & Assurance.
Nearest alternative: Representation, Specification & Plan — Representation, Specification & Plan can support this mechanism, but the evidence centers the concrete operation described above rather than the alternative family's defining operation.
Review outcome: Adjudicated after independent review; high confidence.
Origin Attribution¶
Primary origin: Organizational & Management Science
Origin pattern: Cross-disciplinary synthesis
Present-day reach: Universal
Rationale: A scorecard that compares scenario assumptions, failure modes, capacity margins, and corrective actions turns stress testing into governed enterprise review. GAO enterprise-risk guidance emphasizes scenario analysis, risk assessment, response, and monitoring; quantitative domains supply the modeled shocks.
Related originating lineages:
- Data Science & Analytics — Data science, analytics, and operational monitoring supplies a parallel or contributing lineage for the mechanism's defining operation: a one-page verdict sheet that consumes the results of the stress tests and gives each key assumption a confidence grade, a reversibility flag, and a disposition — safeguarded,….
- Disaster Management & Risk Reduction — disaster_management contributes continuity, incident recovery, readiness, and resource coordination to this mechanism's defining operation—A one-page verdict sheet that consumes the results of the stress tests and gives each key assumption a confidence grade, a reversibility flag, and a disposition — safeguarded, monitored, or knowingly accepted—without displacing the selected primary historical lineage.
- Economics & Finance — economics_finance contributes economics, finance, and mechanism-design practice to this mechanism's defining operation—A one-page verdict sheet that consumes the results of the stress tests and gives each key assumption a confidence grade, a reversibility flag, and a disposition — safeguarded, monitored, or knowingly accepted—without displacing the selected primary historical lineage.
- Engineering & Design — Reversibility and safeguard status matter.
- Mathematics — Mathematical modeling, proof, and abstract-structure practice supplies a parallel or contributing lineage for the mechanism's defining operation: a one-page verdict sheet that consumes the results of the stress tests and gives each key assumption a confidence grade, a reversibility flag, and a disposition — safeguarded,….
- Operations Research — Scorecards support decisions.
- Statistics & Experimental Design — statistics_experimental_design contributes statistics, experimental design, and measurement theory to this mechanism's defining operation—A one-page verdict sheet that consumes the results of the stress tests and gives each key assumption a confidence grade, a reversibility flag, and a disposition — safeguarded, monitored, or knowingly accepted—without displacing the selected primary historical lineage.
Review resolution: The blind reviewers disagree on primary lineage (organizational_management versus statistics_experimental_design). Authoritative or primary research supports organizational_management as the best historical origin: A scorecard that compares scenario assumptions, failure modes, capacity margins, and corrective actions turns stress testing into governed enterprise review. GAO enterprise-risk guidance emphasizes scenario analysis, risk assessment, response, and monitoring; quantitative domains supply the modeled shocks. The cited U.S. GAO, Enterprise Risk Management: Selected Agencies' Experiences Illustrate Good Practices; U.S. GAO, Disaster Resilience Framework directly supports the mechanism's defining operation. All independently supported contributing domains are retained without an arbitrary cap. origin_mode=cross_disciplinary_synthesis records lineage, while domain_reach=universal records later applicability separately from provenance.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Researched adjudication after independent review; high confidence.
Sources consulted:
- U.S. GAO, Enterprise Risk Management: Selected Agencies' Experiences Illustrate Good Practices
- U.S. GAO, Disaster Resilience Framework
Notes¶
[n1] The IPCC's calibrated-language framework — pairing a confidence level with the type, amount, and agreement of evidence behind each finding — is the mature form of what a scorecard's confidence column does at the row level: it refuses to state a confidence without disclosing the evidence basis that earns it. ↩