{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"bounded_rivalry_governance__human_computer_interaction:P3:v0","cell_id":"bounded_rivalry_governance__human_computer_interaction","search_queries":["site:nist.gov incident response incident command roles evidence rollback change management guide","site:sre.google incident response incident commander operations lead communication lead hypothesis","software incident response competing hypotheses cognitive bias war room study","incident management tool command center hypothesis testing runbook change lock production","site:pagerduty.com incident response roles incident commander changes only responder operations lead postmortem","site:incident.io incident response decision log actions permissions war room hypotheses","\"Failures and Fixes\" software system incident response study PDF","site:bls.gov software developers occupational employment wage May 2025","site:help.incident.io incident actions decision log roles incident commander workflows","site:firehydrant.com documentation incident command center change tracking runbook roles","site:rootly.com incident response hypothesis decision log roles commands","software incident response multiple hypotheses war room cognitive bias empirical study"],"sources":[{"source_id":"S1","title":"Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile (NIST SP 800-61r3)","publisher":"National Institute of Standards and Technology","url":"https://nvlpubs.nist.gov/nistpubs/specialpublications/nist.sp.800-61r3.pdf","source_class":"STANDARD","publication_date":"2025-04","accessed_at":"2026-08-03","claims_supported":["Incident response is consequential, organization-wide risk management work.","Leadership oversees incident response, allocates funding, and may authorize high-impact actions.","Policies, authorization boundaries, coordinated roles, information sharing, exercises, after-action review, and continuous improvement are recognized practices.","Recovery actions should be selected using predefined criteria and available resources, then changed when needs or resources are reassessed."]},{"source_id":"S2","title":"Incident Management: Key to Restore Operations","publisher":"Google Site Reliability Engineering","url":"https://sre.google/sre-book/managing-incidents/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2016","accessed_at":"2026-08-03","claims_supported":["An unmanaged incident can deteriorate when responders make uncoordinated production changes.","Clear command, role separation, communication, and control are established incident-management practices.","Google recommends that only the operations group modify the system during an incident."]},{"source_id":"S3","title":"Incident Commander: Roles, Responsibilities and Best Practices","publisher":"Rootly","url":"https://rootly.com/incident-response/incident-commander","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"2026-07-23","accessed_at":"2026-08-03","claims_supported":["Current incident-command guidance already distinguishes hypotheses from confirmed facts, uses time-boxed operational periods, and records decisions, owners, expected results, and verification timers.","Current guidance calls for framing options, effects, risks, rollback paths, safety constraints, permissions, evidence capture, and command transfer.","Incident commanders and incident-program owners are identifiable potential authorizers."]},{"source_id":"S4","title":"The Command Center","publisher":"FireHydrant","url":"https://docs.firehydrant.com/docs/the-command-center","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026 (updated)","accessed_at":"2026-08-03","claims_supported":["Commercial incident-management software already aggregates responders, roles, runbooks, messages, actions, change events, and an incident timeline.","Existing command-center infrastructure could host proposal records and audit events, but the documented product does not implement competitive scoring or command-lease allocation."]},{"source_id":"S5","title":"Failures and Fixes: A Study of Software System Incident Response","publisher":"IEEE Computer Society","url":"https://conferences.computer.org/icsme/pdfs/ICSME2020-1oOutvkGTwF4GyVvNtr3Mm/561900a185/561900a185.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2020","accessed_at":"2026-08-03","claims_supported":["A qualitative study analyzed 30 production incidents using interviews and public reports.","Production incidents involve uncertainty, cascading failures, opportunistic investigation, changes, restarts, failovers, and rollbacks under time pressure.","The study supports the importance of incident investigation and mitigation but does not establish the prevalence of strategic rivalry for command access."]},{"source_id":"S6","title":"The Analysis of Competing Hypotheses in Intelligence Analysis","publisher":"ERIC, U.S. Department of Education","url":"https://eric.ed.gov/?id=EJ1262011","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2019","accessed_at":"2026-08-03","claims_supported":["A randomized study of 50 intelligence analysts found mixed evidence for Analysis of Competing Hypotheses.","Participants did not reliably follow all ACH steps, and the method could increase inconsistency and error.","Structuring competing hypotheses should not be presumed beneficial without testing in the target workflow."]},{"source_id":"S7","title":"Effects of Task Structure and Confirmation Bias in Alternative Hypotheses Evaluation","publisher":"Springer Nature","url":"https://link.springer.com/article/10.1186/s41235-024-00560-y","source_class":"PRIMARY_RESEARCH","publication_date":"2024-06-13","accessed_at":"2026-08-03","claims_supported":["In Study 1 with 161 participants, interface orientation affected confirmation-bias outcomes.","An ACH-style matrix did not reduce confirmation bias or improve sensitivity to evidence credibility, while a different row-oriented presentation did reduce bias.","The exact proposal layout and scoring interface are consequential design variables requiring empirical evaluation."]},{"source_id":"S8","title":"Occupational Employment and Wages: May 2025 National Data","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.toc.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-05-15","accessed_at":"2026-08-03","claims_supported":["The May 2025 Occupational Employment and Wage Statistics release provides a public basis for valuing U.S. technical and managerial labor used in the resource-equivalent cost estimates.","The source supports labor-cost framing, not vendor pricing or organization-specific implementation costs."]}],"problem_evidence":{"support":"MODERATE","rationale":"The broad problem visibly exists: Google documents an unmanaged incident worsened by an uncoordinated production change, and the IEEE study shows production investigation and mitigation occurring amid uncertainty, cascading effects, and potentially consequential changes. NIST treats incident response as critical organizational risk management. However, none of the eight sources measures the proposal's narrower prevalence claim that multiple responder teams strategically consume diagnostic resources, conceal evidence, or compete for production command. That component remains a plausible but unverified extrapolation.","source_ids":["S1","S2","S5"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"NIST identifies organizational leadership as funder and high-impact-action authority, while Google and Rootly identify the incident commander, operations lead, safety roles, and incident-program owners as operational authorizers. These publishers express demand for controlled changes, clear authority, evidence capture, rollback, time-boxed decisions, and review. No named organization was found requesting the full rivalry arena, equal-budget rehearsal, masked scoring, or competitive command leases.","source_ids":["S1","S2","S3"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Google incident command and single-operations-writer practice","similarity":"Uses a hierarchical incident commander, separated roles, one operations group modifying production, shared incident documentation, and explicit control to prevent freelancing.","remaining_difference":"It does not create a scored competition among independently registered hypotheses, equal-budget finalist rehearsals, fouls, or prediction-triggered lease reopening.","source_ids":["S2"]},{"name":"Rootly time-boxed incident action planning","similarity":"Already marks hypotheses versus facts, frames options with expected effects and rollback, selects a safe action, records owner and verification timer, uses operational periods, and supports documented command transfer.","remaining_difference":"It relies on incident-command judgment rather than masked portfolio selection, equal diagnostic budgets, audited competitive scoring, and cross-incident anti-collusion screening.","source_ids":["S3"]},{"name":"FireHydrant Command Center","similarity":"Provides roles, runbooks, tasks, change events, messages, and an automatically accumulated incident timeline in a shared interface.","remaining_difference":"The documented product does not allocate exclusive command leases through hypothesis competition or enforce diagnostic-resource caps and scoring.","source_ids":["S4"]},{"name":"Analysis of Competing Hypotheses","similarity":"Structures multiple hypotheses and evidence for comparison instead of evaluating only one favored explanation.","remaining_difference":"It does not govern production authority, safety gates, sandbox budgets, rollback, or incident-command transfer; controlled evidence of cognitive benefit is mixed and layout-sensitive.","source_ids":["S6","S7"]},{"name":"NIST criteria-based recovery selection and reassessment","similarity":"Requires authorized, scoped, prioritized recovery actions selected against predefined criteria and available resources, with reassessment and after-action learning.","remaining_difference":"It is governance guidance rather than a rivalrous selection mechanism and does not prescribe independent submissions, finalist rehearsals, leaderboards, or expiring leases.","source_ids":["S1"]}],"distinctive_claim_remaining":"In simulated software incidents with at least two credible diagnoses and one safely serializable remediation opportunity, adding independent hypothesis registration, equal-budget rehearsal of nonredundant finalists, common audited scoring, and a prediction-triggered expiring command lease to an otherwise conventional incident-command workflow will improve blinded selection quality, causal attribution, and rollback completeness without materially increasing time to safe mitigation, delaying critical evidence, or misrouting protected emergency actions.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"The workflow is technically plausible because commercial command centers already manage roles, runbooks, change events, timelines, and structured records, while current guidance already uses permissions, rollback paths, decision timers, hypothesis labels, and command transfer. A synthetic tabletop can be run safely under NIST-recognized exercise practice. Production implementation is less certain: automatic leases require reliable privileged-access enforcement, atomic revocation and handoff, tamper-resistant logs, service-specific rollback semantics, latency limits, privacy controls, and explicit authority from leadership, asset owners, security, legal, and incident command. The proposal's masking may fail through authorship cues, and experimental evidence warns that structured hypothesis interfaces can worsen inconsistency or depend on layout.","source_ids":["S1","S3","S4","S6","S7"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Poorly coordinated incident changes can worsen outages, and incident response affects critical services, but the frequency and attributable harm of diagnostic rivalry are unmeasured.","source_ids":["S1","S2","S5"]},"stakeholder_pull":{"score":3,"rationale":"Incident leaders visibly seek command clarity, safe changes, evidence capture, rollback, and review; expressed demand for a competitive arena itself was not found.","source_ids":["S1","S2","S3"]},"incremental_advantage":{"score":2,"rationale":"Single-writer control, time-boxed plans, hypothesis labels, expected-result logging, rollback, reassessment, and command transfer already cover much of the proposed value. Added rehearsal and rivalry controls may help, but could also add delay and ceremony.","source_ids":["S1","S2","S3","S4"]},"distinctiveness_plausibility":{"score":3,"rationale":"The integrated combination of equal-budget hypothesis rehearsal, audited scoring, fouls, and an expiring command lease was not found, although its constituent practices substantially collide with incident command, command-center products, and ACH.","source_ids":["S1","S2","S3","S4","S6","S7"]},"technical_implementability":{"score":4,"rationale":"A nonproduction prototype can be assembled from established roles, workflows, structured records, timers, runbooks, and audit timelines. Production-grade access revocation and safe handoff remain significant engineering work.","source_ids":["S1","S3","S4"]},"adoption_authority_feasibility":{"score":3,"rationale":"Leadership and incident commanders are identifiable authorizers, and a training lead can authorize a simulator exercise. Production adoption would require service-owner, security, legal, and privileged-access approval.","source_ids":["S1","S2","S3"]},"evidence_readiness":{"score":3,"rationale":"The proposal has a bounded simulator experiment and observable metrics, but no comparative incident-response evidence yet establishes benefit or safety.","source_ids":["S1","S6","S7"]},"safety_net_benefit":{"score":2,"rationale":"Safer incident response can protect all downstream users, but the proposal is not specifically targeted to safety-net institutions or populations, and distributional benefit is unmeasured.","source_ids":["S1","S5"]},"scalability":{"score":3,"rationale":"The software workflow and audit components can scale across services, but rehearsal capacity, expert scoring, severity-specific tuning, and incident-time latency may not scale uniformly.","source_ids":["S1","S3","S4"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Design, preregister, facilitate, instrument, and analyze three isolated tabletop scenarios with trained responders and blinded outcome review.","confidence":"MODERATE","assumptions":["Approximately 200-400 total hours across scenario design, facilitation, participants, safety review, instrumentation, and analysis.","Existing chat, observability-demo, and simulator infrastructure is reused.","No production credentials, customer data, procurement, or custom privileged-access integration is included.","Labor is valued as 2026 resource-equivalent U.S. technical and managerial time using the May 2025 OEWS release as a public baseline."],"source_ids":["S1","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Build a nonproduction console prototype with proposal templates, scoring, timers, synthetic telemetry, immutable exercise logs, role controls, and facilitator dashboards.","confidence":"MODERATE","assumptions":["A small product/engineering team works for roughly two to four months.","Existing incident-management and identity systems are integrated through available APIs.","Security review is limited to the isolated training environment.","No automatic production command enforcement is enabled."],"source_ids":["S1","S3","S4","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Launch a limited production pilot for selected noncritical services, including privileged-access lease enforcement, revocation, audit retention, rollback integration, training, policy approval, and on-call support.","confidence":"LOW","assumptions":["Three to six engineers plus incident-program, security, legal, service-owner, and training effort over six to twelve months.","The organization already operates mature incident command, observability, identity, and change-control systems.","The pilot excludes life-safety and regulatorily mandated emergency actions.","External penetration testing or independent assurance may be required.","No first-party implementation quote or adopter-specific architecture was available."],"source_ids":["S1","S3","S4","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Maintain integrations and rules, run recurring exercises, review incidents and access logs, recalibrate scoring, train responders, and provide limited operational support.","confidence":"LOW","assumptions":["One partial-to-full-time program owner plus fractional engineering, security, and training support.","Existing incident-management licensing is not materially increased; incremental vendor fees are unknown.","Rules and integrations are reviewed after material incidents and at least annually.","This estimate excludes outage losses, compensation changes, and organization-wide rollout."],"source_ids":["S1","S3","S4","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Authoritative guidance and a real-practice research study establish consequential incident uncertainty and the danger of unmanaged or uncoordinated production changes, although strategic rivalry prevalence remains unmeasured.","source_ids":["S1","S2","S5"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Organizational leadership, incident-program owners, incident commanders, operations leads, safety/security authorities, and service owners are externally documented roles with funding, selection, execution, or veto authority.","source_ids":["S1","S2","S3"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim contrasts the arena with conventional incident command and specifies measurable gains and noninferiority constraints for time, evidence sharing, rollback, and protected actions.","source_ids":["S2","S3","S6","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A counterbalanced, synthetic tabletop comparison can test protocol coherence and safety without production access, consistent with NIST-recognized exercise practice.","source_ids":["S1"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The next step is limited to inert credentials, synthetic telemetry, predefined bypass actions, independent facilitation, and immediate simulator shutdown. Production authority remains explicitly out of scope.","source_ids":["S1","S3"]},"credible_cost_scope_and_range":{"status":"YES","reason":"All four estimates state bounded scopes, staffing assumptions, exclusions, confidence, and public labor-cost framing. Operational bands remain low-confidence because no architecture-specific quote exists.","source_ids":["S1","S4","S8"]}},"next_evidence_step":"Run a preregistered, counterbalanced crossover tabletop with 8-12 trained response teams in an isolated simulator. Each team handles three synthetic scenarios: two plausible diagnoses, four proposals competing for a scarce diagnostic service, and a protected containment trigger. Compare (A) conventional incident command with a single operations writer, decision log, rollback gate, and ordinary collaboration against (B) the same workflow plus independent registration, two equal-budget rehearsals, audited scoring, one expiring lease, and prediction-triggered reopening. Blind adjudicators to condition and score selected-action evidential support, prediction-to-telemetry correspondence, causal attribution, and rollback completeness; also measure time to safe mitigation, critical-evidence sharing delay, scorer agreement, resource exceptions, lease violations, bypass accuracy, participant workload, and retained alternatives. Red-team duplicate authorship, log alteration, status cues, flooding, reciprocal endorsements, and sunk-cost resistance. Falsify the incremental claim if the arena does not improve the prespecified composite selection-quality score, increases median time to safe mitigation by more than 10%, causes any protected containment action to await scoring, increases critical-evidence delay, produces lower rollback completeness, or is routinely bypassed.","blocking_evidence":["No field estimate establishes how often competing responder teams strategically affect diagnosis selection, diagnostic resources, or command access.","No controlled comparison shows that the integrated arena improves incident decisions over modern incident-command practice.","No evidence establishes a safe latency bound by severity or shows that independent registration will not delay critical evidence.","No production test establishes reliable lease revocation, command handoff, rollback ownership, or failure behavior during control-plane degradation.","No committed adopter has validated the proposed scoring rubric, resource caps, masking, foul rules, or workflow burden.","No organization-specific legal, privacy, labor, privileged-access, or regulatory review has been completed.","No architecture-specific implementation quote, vendor price, or observed recurring operating cost is available.","World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The eight-source bounded search found substantial collision with incident command, single-writer operations, time-boxed decision and rollback practice, commercial command centers, and structured competing-hypothesis analysis. It did not find the exact integrated equal-budget rehearsal and expiring-command-lease protocol. This is not an exhaustive prior-art, patent, product, or global-practice search; world novelty, patentability, freedom to operate, market size, and realized impact are expressly unmeasured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Complete the preregistered tabletop comparison against a modern incident-command baseline.","Demonstrate improved blinded selection quality and rollback completeness without more than a 10% median mitigation-time penalty.","Show zero delayed or misrouted protected containment actions and no material increase in critical-evidence sharing delay.","Measure scorer reliability, masking failure, resource-cap exceptions, lease bypass, participant workload, and reopening compliance.","Obtain a named incident-program owner, service owner, and security authority willing to review results and define conditions for any further pilot.","Produce a production threat model and failure-mode analysis for lease issuance, atomic revocation, handoff, logging, and degraded-control-plane operation."],"reason":"Bounded web research establishes the broad problem, credible authorities, and substantial adjacent practice, but cannot establish the remaining incremental efficacy and safety claim. The decisive evidence requires human-participant tabletop fieldwork and later live-system engineering validation, so an empirical-research stop is required."},"proposal_index":3}