{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"bounded_rivalry_governance__futurism_foresight:P1:v0","cell_id":"bounded_rivalry_governance__futurism_foresight","search_queries":["site:fema.gov HSEEP doctrine exercise program priorities scenario design evaluation after action report PDF","site:cisa.gov tabletop exercise infrastructure scenario planning guidance stakeholder exercise package","site:nerc.com GridEx lessons learned scenario exercise report 2023 PDF","scenario planning cognitive bias selection scenarios decision usefulness empirical research","site:gov.uk futures toolkit scenarios quality challenge assumptions decision making PDF Government Office for Science","site:oecd.org strategic foresight scenarios toolkit decision making public sector PDF","site:iso.org ISO 22398 exercise guidelines scenario objectives evaluation","site:nerc.com GridEx VII lessons learned report scenario objectives stakeholder 2024","NERC GridEx VII 2023 lessons learned scenario development objectives report","site:nerc.com GridEx VI lessons learned 2022 exercise design scenario development objectives","site:energy.gov infrastructure exercise scenario evaluation preparedness exercise utility","site:gao.gov critical infrastructure exercises scenario lessons learned preparedness","competition select scenarios for tabletop exercises scenario challenge foresight contest","foresight scenario competition submissions judged scenarios challenge prize","preparedness exercise scenario selection criteria portfolio complementarity","forecasting tournament scenario planning competition prior art","open access empirical study scenario planning bias decision quality full text","site:pmc.ncbi.nlm.nih.gov scenario planning experiment decision making bias","site:researchgate.net \"Cognitive benefits of scenario planning\"","doi 10.1016/j.techfore.2012.09.011 full text"],"sources":[{"source_id":"S1","title":"The Futures Toolkit HTML","publisher":"UK Government Office for Science","url":"https://www.gov.uk/government/publications/futures-toolkit-for-policy-makers-and-analysts/the-futures-toolkit-html","source_class":"OFFICIAL_GUIDANCE","publication_date":"2024-08-29","accessed_at":"2026-08-03","claims_supported":["Scenario work is used to rehearse future decisions and trade-offs.","Official guidance explicitly states that more scenarios exist than can reasonably be included and that choosing sufficiently different and interesting scenarios is partly judgmental.","Scenario work should include diverse stakeholders, distinguish scenarios, and connect scenarios to present decisions.","A scenario-building workshop alone requires at least four to five hours, illustrating nontrivial resource use."]},{"source_id":"S2","title":"Exercises and Training","publisher":"U.S. Department of Energy, Office of Cybersecurity, Energy Security, and Emergency Response","url":"https://www.energy.gov/ceser/exercises-and-training","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["DOE has an identifiable mandate and operational role in energy-sector preparedness exercises.","DOE describes exercises as threat-informed, objectives-based activities used to measure resilience, identify gaps, and generate actionable executive outputs.","DOE conducts after-action improvement planning and carries validated improvements into plans and future exercises."]},{"source_id":"S3","title":"HSEEP Policy and Guidance","publisher":"Federal Emergency Management Agency","url":"https://preptoolkit.fema.gov/web/hseep-resources/policy-and-guidance","source_class":"OFFICIAL_GUIDANCE","publication_date":"2020","accessed_at":"2026-08-03","claims_supported":["HSEEP already supplies a common lifecycle for exercise-program management, design, conduct, evaluation, and improvement planning.","Templates and established doctrine make a retrospective, no-stakes exercise-process comparison technically feasible.","The proposal would operate inside a mature exercise-governance practice rather than inventing exercise governance from scratch."]},{"source_id":"S4","title":"ISO 22398:2013 — Societal security — Guidelines for exercises","publisher":"International Organization for Standardization","url":"https://www.iso.org/standard/50294.html","source_class":"STANDARD","publication_date":"2013-09-13","accessed_at":"2026-08-03","claims_supported":["An international standard already covers planning, conducting, and improving organizational exercise projects and programs.","The guidance is adaptable to organizational objectives, resources, and constraints across public and private organizations."]},{"source_id":"S5","title":"GridEx VII Lessons Learned Report","publisher":"North American Electric Reliability Corporation, Electricity Information Sharing and Analysis Center","url":"https://www.nerc.com/globalassets/programs/electricity-isac/gridex/gridex-vii-report.pdf","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2024-04-01","accessed_at":"2026-08-03","claims_supported":["Critical-infrastructure operators and government leaders participate in large, consequential, scenario-based exercises.","GridEx VII involved 252 registered organizations and an estimated 15,000-plus players, demonstrating that exercise design can matter at operational scale.","Planners expressed differing needs for complexity, customization, earlier materials, and support; NERC committed to improving future scenario materials and exercise structures.","After-action evidence feeds changes to later exercise design, closely paralleling the proposal's recalibration loop."]},{"source_id":"S6","title":"Planning and conducting crisis management exercises for decision-making: the do’s and don’ts","publisher":"Springer Nature, EURO Journal on Decision Processes","url":"https://link.springer.com/article/10.1007/s40070-017-0065-0","source_class":"PRIMARY_RESEARCH","publication_date":"2017-08-21","accessed_at":"2026-08-03","claims_supported":["Analysis drawing on 12 tabletop and functional exercises found that design and conduct affect the relevance of learning.","Scenario selection and development must be tied to exercise goals and client requirements.","Excess detail and poorly paced injects can create overload, while documented feedback and post-exercise learning improve later practice.","The authors note that empirical evidence about which exercise-design choices work remains limited."]},{"source_id":"S7","title":"Challenges in Navigating Scenario Planning","publisher":"California Management Review, UC Berkeley Haas School of Business","url":"https://cmr.berkeley.edu/2026/04/68-3-challenges-in-navigating-scenario-planning/","source_class":"PRIMARY_RESEARCH","publication_date":"2026-04-30","accessed_at":"2026-08-03","claims_supported":["An autoethnographic NASA case reports that leadership bias and internal organizational politics initially limited acceptance of scenarios developed outside the senior team.","The case makes agenda power and leadership resistance plausible problems in organizational scenario processes.","It does not establish the prevalence or effect size of those problems across infrastructure operators."]},{"source_id":"S8","title":"Scenarios in the strategy process: a framework of affordances and constraints","publisher":"Springer Nature, European Journal of Futures Research","url":"https://link.springer.com/article/10.1186/s40309-019-0160-5","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2019-12-30","accessed_at":"2026-08-03","claims_supported":["Scenario processes are constrained by bounded rationality, selective attention, participant intentions, and organizational politics.","Scenario tools can be used to support predetermined decisions or personal interests, and participant selection can be an exercise of power.","The paper recommends broad participation and collaborative development to reduce political misuse, creating a credible non-rivalrous comparator.","It also emphasizes that not every conceivable scenario can be considered and that causal links from scenario use to organizational performance remain unclear."]}],"problem_evidence":{"support":"MODERATE","rationale":"The general problem is visible: official foresight guidance explicitly recognizes that more scenarios exist than can reasonably be exercised and that selection for difference is partly art; exercise research links scenario design to learning; and organizational studies document bounded attention, leadership bias, and political influence. Infrastructure exercises demonstrably matter at substantial scale. However, no external source verifies this candidate's specific three-slot constraint, prevalence of lobbying or presentation spending, repeated-winner capture, or resulting loss of preparedness learning at a particular regional operator.","source_ids":["S1","S5","S6","S7","S8"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"DOE is an identifiable exercise authority with an explicit need for objectives-based exercises that produce actionable executive learning. NERC represents a credible infrastructure-operator setting and reports planner demand for better-tailored, more usable, and more sophisticated scenario materials and structures. This establishes a plausible adopter class and expressed need for exercise improvement, but not expressed demand for a competitive scenario league, masked judging, resource caps, appeals, or anti-collusion screening.","source_ids":["S2","S5"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"HSEEP and ISO 22398 exercise-program lifecycle","similarity":"Both establish objectives, structured design, conduct, evaluation, improvement planning, and adaptation to organizational constraints.","remaining_difference":"They govern exercises themselves but do not prescribe a recurring competition among scenario-producing teams, portfolio-level allocation of scarce rehearsal slots, entrant resource caps, masked standings, or anti-collusion screens.","source_ids":["S3","S4"]},{"name":"DOE CESER exercise program and NERC GridEx","similarity":"These are live critical-infrastructure exercise programs using threat-informed scenarios, planners, evaluation, after-action learning, stakeholder coordination, and iterative redesign.","remaining_difference":"Scenario materials are centrally developed or locally tailored; the sources do not describe rival teams competing under a frozen rubric for a complementary portfolio of rehearsal slots.","source_ids":["S2","S5"]},{"name":"UK Government Office for Science scenario method","similarity":"It explicitly addresses selecting a limited, sufficiently different set of scenarios, stakeholder inclusion, decision rehearsal, and policy stress-testing.","remaining_difference":"It is participatory and collaborative, with selection acknowledged as partly art; it lacks comparative entrant scoring, audits, appeals, effort caps, fouls, challenger access, and portfolio awards.","source_ids":["S1"]},{"name":"Evidence-based crisis-exercise planning practice","similarity":"It ties scenario choice to exercise goals, uses client feedback, manages realism and participant overload, and documents lessons for later designs.","remaining_difference":"It treats planner-client collaboration as the preferred mechanism rather than governed rivalry among scenario suppliers.","source_ids":["S6","S8"]}],"distinctive_claim_remaining":"On the same bounded set of candidate scenarios, a frozen and identity-masked rivalry with auditable scoring, hidden perturbations, and portfolio-level selection will produce a three-scenario set with measurably greater distinct decision coverage and lower causal/signal redundancy than both executive deliberation and masked score-only ranking, while not materially reducing scenario robustness or imposing unacceptable identity leakage, stakeholder harm, appeal burden, or evaluation cost.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"The intervention uses technically ordinary components: secure submission, masking, structured rubrics, separated roles, audits, appeals, exercise evaluation, and after-action review. Existing government programs and standards demonstrate that structured exercise lifecycles and large multi-organization scenario exercises are implementable. A retrospective no-stakes test is therefore feasible. Important gaps remain: masking may fail through writing style or topic; portfolio complementarity can reintroduce discretion; preparation caps are hard to audit; overlap screens cannot establish collusion; sensitive infrastructure data require operator authorization and access controls; and no source demonstrates that the full bundle improves selection.","source_ids":["S1","S2","S3","S4","S5","S6"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Selecting less redundant, more decision-relevant rehearsals could improve preparedness learning where exercises are consequential and resource-intensive, but the proposal has no evidence connecting its selection mechanism to operational outcomes.","source_ids":["S1","S2","S5","S6"]},"stakeholder_pull":{"score":2,"rationale":"Infrastructure exercise authorities express demand for better exercises and scenario materials, but no identified operator has requested or committed to this rivalry mechanism.","source_ids":["S2","S5"]},"incremental_advantage":{"score":2,"rationale":"Most governance components are established in exercise doctrine, standards, and participatory foresight practice. The incremental benefit of combining them as a competitive portfolio-selection process is plausible but wholly untested.","source_ids":["S1","S3","S4","S5","S6"]},"distinctiveness_plausibility":{"score":3,"rationale":"No direct match for a governed scenario-rehearsal league was found in the bounded search. The combination is distinguishable from collaborative scenario workshops and centrally planned exercises, although it is assembled from familiar contest, audit, and exercise-governance elements.","source_ids":["S1","S2","S3","S5","S6"]},"technical_implementability":{"score":4,"rationale":"A no-stakes version needs no novel technology and can reuse mature exercise-management, evaluation, and improvement workflows. Masking, secure data handling, and portfolio optimization require careful implementation but are tractable.","source_ids":["S3","S4","S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"DOE and infrastructure-sector organizations clearly possess exercise authority, and a regional operator could plausibly authorize an internal retrospective test. Exact authority over employee evaluation, consultant rules, sensitive records, appeals, and stakeholder participation is unverified.","source_ids":["S2","S5"]},"evidence_readiness":{"score":2,"rationale":"The proposal has measurable mechanisms and falsifiers, but no verified baseline records, adopter commitment, scoring reliability data, identity-leakage data, or comparative results.","source_ids":["S6","S7","S8"]},"safety_net_benefit":{"score":4,"rationale":"The proposed first step is retrospective, no-stakes, reversible, and can use synthetic packages; standings can be voided without changing live rehearsal allocations. Residual confidentiality and reputational risks require strict de-identification and access controls.","source_ids":["S3","S4","S5"]},"scalability":{"score":3,"rationale":"Digital submission and common rubrics can scale, but auditing, appeals, stakeholder review, sensitive-data controls, and verification of resource caps grow with entrants. GridEx also shows that participant sophistication and resource levels vary substantially.","source_ids":["S5","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"One preregistered, no-stakes retrospective comparison using 24-40 archived or synthetic packages, three independent selection panels, secure de-identification, rubric calibration, limited source audits, stakeholder review, analysis, and a go/no-go report.","confidence":"LOW","assumptions":["Packages already exist or can be synthesized without extensive new research.","No live exercise or employment decision is affected.","Costs are resource-equivalent and include internal staff time, independent evaluators, security review, and analysis.","No source supplied operator-specific wage rates or data-remediation costs."],"source_ids":["S1","S3","S5","S6"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Design and approval of a production rulebook, secure submission and scoring workflow, masking and conflict controls, audit and appeal procedures, labor/privacy/security review, evaluator training, and one dry run before live allocation.","confidence":"LOW","assumptions":["Existing enterprise workflow tools can be configured rather than custom-built.","Sensitive-infrastructure controls and identity masking require professional review.","The operator has an existing exercise-program office and does not need to create one from zero."],"source_ids":["S3","S4","S5"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"First live annual cycle, including competition administration, judges, audits, appeals, stakeholder safeguards, post-contest review, and delivery of three regional cross-departmental rehearsals.","confidence":"LOW","assumptions":["Three rehearsals are regional rather than GridEx-scale national exercises.","Rehearsal delivery, staff participation, consultant support, security, and opportunity cost are included.","If exercise delivery is budgeted separately, the league-only launch could fall in the 250K_TO_1M band."],"source_ids":["S2","S3","S4","S5","S6"]},"annual_recurring":{"band_2026_usd":"1M_TO_5M","scope":"One competition and three rehearsals per year, including recurring administration, secure systems, evaluator and audit effort, appeals, participant protections, after-action analysis, and annual rule revision.","confidence":"LOW","assumptions":["Entrant volume remains in the tens, not hundreds.","Two to four staff-equivalents administer the cycle, supplemented by judges, auditors, facilitators, and exercise participants.","Materially larger exercises or extensive external consulting could exceed this band."],"source_ids":["S2","S3","S5","S6"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"External official guidance directly identifies scenario abundance and difficult diversity selection, while research documents exercise-design consequences and political or cognitive distortions.","source_ids":["S1","S6","S7","S8"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"DOE CESER is an explicit energy-sector exercise authority, and NERC documents extensive participation and improvement needs among infrastructure operators. This verifies a credible adopter class, not commitment to the proposed mechanism.","source_ids":["S2","S5"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal can be compared on identical packages against executive deliberation and masked score-only ranking using portfolio redundancy, decision coverage, hidden-perturbation performance, leakage, burden, and harm outcomes.","source_ids":["S1","S3","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A preregistered retrospective using archived or synthetic packages can test feasibility and mechanisms without allocating live funds or exercises.","source_ids":["S3","S4","S5"]},"no_unresolved_safety_or_authority_stop":{"status":"UNCERTAIN","reason":"The proposed rollback and no-stakes boundary reduce risk, but authority to use proprietary submissions, mask employee identities, review consultant work, expose sensitive vulnerabilities to judges, and involve worker or community representatives has not been verified for an actual operator.","source_ids":["S3","S4","S5"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The process scope and dominant resource drivers are identifiable, but no operator-specific staffing, security, exercise-delivery, or opportunity-cost data were found. The bands are broad resource-equivalent estimates rather than externally validated budgets.","source_ids":["S1","S3","S5","S6"]}},"next_evidence_step":"With one consenting infrastructure operator, preregister a no-stakes retrospective using 24-40 archived or synthetic scenario packages and three separated panels: A reproduces executive pitch/deliberation, B uses masked fixed-rubric individual ranking, and C uses the full masked, audited portfolio-selection arena. A fourth blinded assessment team should score the resulting three-scenario portfolios for distinct consequential decisions covered, pairwise signal/causal-chain overlap, stakeholder-harm coverage, and performance under undisclosed perturbations. Record judge reliability, identity guesses, preparation and evaluation hours, audit changes, appeals, conflicts, and data incidents. Falsify or halt the proposed mechanism if inter-rater reliability is below ICC 0.60; identity inference exceeds chance by more than 10 percentage points; C fails to reduce overlap by at least 15% or improve distinct decision coverage by at least 10% against both comparators; median hidden-perturbation performance falls by more than 10%; audit and appeal work exceeds 25% of evaluation hours; total evaluation burden exceeds twice B without corresponding portfolio improvement; stakeholder-harm flags increase; confidential data escape; or reasonable complementarity-weight variations reverse the selected portfolio. Make no live rehearsal, funding, personnel, misconduct, or reputational decision from the results.","blocking_evidence":["Three to five years of operator records verifying actual rehearsal scarcity, candidate volume, selection criteria, scenario redundancy, preparation spending, lobbying access, repeat-winner concentration, and post-rehearsal learning are unavailable.","No operator has committed data, staff time, security approval, or authority for the retrospective comparison.","Reliability and predictive validity of the proposed scoring rubric and hidden perturbations are unknown.","The feasibility of masking scenario authors and auditing preparation-resource caps has not been demonstrated.","No comparative evidence shows that portfolio scoring improves consequential decision coverage rather than adding discretionary judgment.","Legal, labor, contractual, privacy, confidentiality, and sensitive-infrastructure-data authority remain operator-specific and unverified.","Operator-specific implementation and recurring cost data are unavailable."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This bounded search found mature exercise standards, participatory scenario-selection practice, infrastructure exercise programs, and research on political and cognitive constraints, but no direct published match for the complete governed rehearsal-league bundle. That absence is not a world-novelty finding. Patentability, freedom to operate, market size, realized impact, and comprehensive world novelty remain unmeasured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain written participation and data-governance approval from one infrastructure operator and its relevant security, legal, labor, procurement, and stakeholder authorities.","Audit three to five prior selection cycles to verify scarcity, candidate volume, redundancy, resource disparities, access effects, and repeat-winner concentration.","Finalize and preregister the three-arm retrospective protocol, scoring rubric, perturbation tests, complementarity sensitivity analysis, and quantitative falsifiers.","Run the no-stakes retrospective with independent panels and publish a de-identified methods-and-results report.","Produce an operator-specific cost model separating league administration from the cost of delivering three rehearsals.","Advance only if the full arena outperforms both comparators without breaching reliability, leakage, burden, sensitivity, safety, or authority thresholds."],"reason":"Bounded web research establishes a real general selection problem, a credible adopter class, substantial adjacent practice, and a testable contrast, but cannot determine whether the proposed league improves portfolio selection. The decisive evidence requires proprietary historical submissions, operator records, participant involvement, security review, and live human evaluation. Under the controller rule, this requires STOP_EMPIRICAL_RESEARCH_NEEDED rather than further web research."},"proposal_index":1}