{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"predictive_residual_processing__futurism_foresight:P4:v0","cell_id":"predictive_residual_processing__futurism_foresight","search_queries":["site:fema.gov HSEEP doctrine exercise evaluation guides observations after action report 2020 PDF","site:gao.gov emergency preparedness exercises evaluation after action reports weaknesses report","tabletop exercise evaluation workload observers data after action review study","hierarchical exercise evaluation inject response matrix tabletop after action residual analysis","site:emilms.fema.gov/is_0120c evaluation collect data exercise performance against targets capability critical tasks root cause analysis","site:emilms.fema.gov/is_0130a exercise evaluator observation data collection expected actions time after action report","site:nist.gov tabletop exercise after action evaluation guide injects objectives observers","site:iso.org ISO 22398 exercises evaluation observation records safety participants","tabletop exercise software after action report timeline inject tracking evaluation observations official product","Crisis simulation exercise platform inject observer evaluation after action report product","tabletop exercise platform automated evaluation participant actions injects first party","strategic foresight tabletop exercise after action review evaluation methodology","\"The use of trained observers as an evaluation tool\" DOI 10.1017/S1049023X00002387","site:bls.gov Occupational Outlook Handbook emergency management directors median pay 2025","site:hhs.gov OHRP quality improvement activities human subjects research exercise evaluation identifiable data","tabletop exercise evaluation observer bias data analysis study after action report"],"sources":[{"source_id":"S1","title":"Exercise Evaluation Guides (EEGs)","publisher":"Federal Emergency Management Agency, Preparedness Toolkit","url":"https://preptoolkit.fema.gov/web/hseep-resources/eegs","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["HSEEP EEGs are explicitly intended to streamline exercise-data collection, assess capability targets, support after-action reports, and map results to objectives, capabilities, targets, and critical tasks.","FEMA supplies evaluation templates spanning operational coordination, planning, communications, situational assessment, and other cross-organizational capabilities.","This is established prior art for structured, objective-relative exercise evaluation and an identifiable institutional adoption channel."]},{"source_id":"S2","title":"IS-0130.a: How To Be An Exercise Evaluator","publisher":"Federal Emergency Management Agency, Emergency Management Institute","url":"https://emilms.fema.gov/is_0130a/content.html","source_class":"OFFICIAL_GUIDANCE","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["FEMA's evaluator curriculum already combines Situation Manuals, EEGs, Master Scenario Events Lists, observation and data collection, issue identification, root-cause analysis, after-action reporting, corrective actions, and improvement tracking.","The official workflow identifies evaluation teams and leads as credible adopters and authorizers.","The proposal substantially overlaps a mature exercise-evaluation workflow, although FEMA's published course outline does not describe residual-only hierarchical reconstruction or independent raw-sample auditing."]},{"source_id":"S3","title":"Biodefense: After-Action Findings and COVID-19 Response Revealed Opportunities to Strengthen Preparedness","publisher":"U.S. Government Accountability Office","url":"https://www.gao.gov/products/gao-21-513","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2021-08-04","accessed_at":"2026-08-03","claims_supported":["GAO identified 74 interagency biological-incident exercises conducted from 2009 through 2019.","Agencies did not routinely monitor exercise results together to identify patterns and root causes of systemic challenges.","GAO recommended consistent capability reporting and routine interagency monitoring, demonstrating consequential institutional need, while later governance changes caused several named recommendations to be closed as no longer valid."]},{"source_id":"S4","title":"Research and Practice of Delivering Tabletop Exercises","publisher":"Association for Computing Machinery authors via arXiv","url":"https://arxiv.org/abs/2404.10206","source_class":"PRIMARY_RESEARCH","publication_date":"2024-04-16","accessed_at":"2026-08-03","claims_supported":["A systematic review screened 140 papers and examined 14 in detail.","The reviewed literature was dominated by linear exercises and generally did not systematically collect learning data.","Only three reviewed papers addressed assessment; evaluation was described as underdeveloped, with little dedicated software and considerable reliance on manual preparation and unstructured observer or facilitator assessment.","The review's computing-education scope limits generalization to strategic-foresight and operational preparedness exercises."]},{"source_id":"S5","title":"From Paper to Platform: Evolution of a Novel Learning Environment for Tabletop Exercises","publisher":"Association for Computing Machinery authors via arXiv","url":"https://arxiv.org/abs/2404.10988","source_class":"PRIMARY_RESEARCH","publication_date":"2024-04-17","accessed_at":"2026-08-03","claims_supported":["The INJECT Exercise Platform encodes structured scenarios, injects, milestones, participant actions, and automated analytics.","Its three-run study involved 91 computing students and reported detailed comparisons of team behavior.","The paper states that manual tabletop assessment can be highly time-consuming and delay feedback by days or weeks.","INJECT demonstrates technical feasibility for structured event capture, conditional injects, repeatable scenario definitions, and automated analysis, but not the proposed hierarchical residual reconstruction and independent raw-audit design."]},{"source_id":"S6","title":"Crisis Simulation After Action Report (AAR)","publisher":"Immersive Labs","url":"https://support.immersivelabs.com/hc/en-us/articles/47423867668881-Crisis-Simulation-After-Action-Report-AAR","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"2026-06-05","accessed_at":"2026-08-03","claims_supported":["A current commercial product updates after-action results from real-time participant data, tags injects to crisis phases, and rolls scores up by phase, inject, participant, confidence, and team alignment.","The product surfaces strengths, improvement areas, blind spots, metadata, and downloadable reports shortly after exercise completion.","This creates substantial product collision with phase-based, confidence-aware, data-driven evaluation, although the documentation does not claim reconstructive expected-plus-residual coding, independent raw sampling, or hierarchical explaining-away with decompression triggers."]},{"source_id":"S7","title":"Quality Improvement Activities FAQs","publisher":"U.S. Department of Health and Human Services, Office for Human Research Protections","url":"https://www.hhs.gov/ohrp/regulations-and-policy/guidance/faq/quality-improvement-activities/index.html","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"n.d.; guidance predates the 2018 Common Rule revisions","accessed_at":"2026-08-03","claims_supported":["Internal quality-improvement activity limited to practical or administrative purposes may fall outside the HHS definition of research.","A systematic investigation designed to produce generalizable knowledge may instead constitute research, and identifiable private information can make it human-subjects research.","An institution-specific determination remains necessary; other privacy, employment, records, consent, contractual, or sectoral requirements may apply independently."]},{"source_id":"S8","title":"Emergency Management Directors: Occupational Outlook Handbook","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/ooh/management/emergency-management-directors.htm","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2025-08-28","accessed_at":"2026-08-03","claims_supported":["The May 2024 median wage for emergency-management directors was $86,130 annually or $41.41 hourly.","Emergency-management directors organize exercises, review and revise emergency plans, coordinate organizations, and assess outcomes, supporting their identification as plausible adopters or authorizers.","The wage benchmark supports labor-based order-of-magnitude estimates but does not provide software, overhead, legal-review, or organization-specific implementation prices."]}],"problem_evidence":{"support":"STRONG","rationale":"The problem is visible in both research and official oversight. Studies report manual, delayed, and underdeveloped tabletop assessment, while GAO found failures to synthesize exercise results into cross-agency patterns and root causes. FEMA's extensive evaluation machinery also indicates that structured collection, reconstruction, analysis, and improvement planning consume real organizational effort. Evidence does not establish the prevalence of evaluator overload specifically in strategic-foresight exercises or quantify how much routine behavior consumes the review budget.","source_ids":["S1","S2","S3","S4","S5"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Exercise directors, evaluation leads, emergency-management directors, FEMA-aligned program managers, and interagency preparedness owners are identifiable adopters or authorizers. FEMA explicitly seeks streamlined collection and consistent assessment, and GAO documented demand for cross-exercise pattern and root-cause monitoring. No source expresses demand for the proposal's exact residual hierarchy, and several GAO recommendations were later closed because the named governance bodies changed.","source_ids":["S1","S2","S3","S8"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"HSEEP objective-based evaluation, EEGs, MSELs, and AAR/IP workflow","similarity":"Already defines objectives and capability targets, organizes expected exercise events, captures observations, analyzes issues and root causes, produces after-action reports, and tracks improvements.","remaining_difference":"Published materials reviewed do not make expected-plus-residual reconstruction the primary representation, propagate only unresolved errors through explicit team/coordination/strategic model layers, or require random independent raw-record audits and decompression thresholds.","source_ids":["S1","S2"]},{"name":"INJECT Exercise Platform","similarity":"Uses machine-readable scenario definitions, injects, milestones, captured interactions, automated analysis, repeatable delivery, and team-level behavioral comparisons to reduce manual evaluation burden.","remaining_difference":"The demonstrated educational platform does not report synchronized uncertain expected-response models, hierarchical residual explaining-away, complete structured-timeline reconstruction from residuals, protected-signal bypasses, or independent raw sampling.","source_ids":["S4","S5"]},{"name":"Immersive Labs Crisis Simulation AAR","similarity":"Provides real-time, inject- and phase-aligned scoring with confidence adjustments, team-alignment measures, blind-spot views, metadata, and rapid after-action reporting.","remaining_difference":"Its documented approach scores responses against ranked options or best practices rather than reconstructing an open-ended chronology from versioned expectations plus structured residuals, and it does not document raw-audit sampling or residual-triggered decompression.","source_ids":["S6"]},{"name":"GAO-recommended interagency monitoring of exercise findings","similarity":"Calls for consistent reporting, cross-exercise pattern detection, root-cause analysis, responsible agencies, and improvement recommendations across organizational boundaries.","remaining_difference":"This is governance and synthesis practice rather than a predictive residual data architecture for evaluating one exercise within a bounded attention budget.","source_ids":["S3"]}],"distinctive_claim_remaining":"For a bounded completed exercise, a version-synchronized hierarchy of uncertain inject-response expectations plus structured residuals will reconstruct the decision-relevant chronology with predeclared fidelity, preserve every protected or safety-relevant signal under independent raw audit, surface more independently confirmed cross-team coordination failures or surface them sooner, and use at least 20% fewer evaluator hours than ordinary full-chronology HSEEP-style review after all modeling, coding, audit, reconciliation, and fallback work is counted.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"The principal components—structured inject schemas, event capture, objective-relative evaluation, conditional scenario logic, confidence-aware scoring, rapid reporting, root-cause analysis, and role-based exercise workflows—are already implemented in official practice, research prototypes, and commercial software. The unverified work is semantic residual coding of open-ended behavior, synchronization across hierarchical expectations, reconstruction fidelity, unbiased raw sampling, protected-signal bypass, and net workload reduction. Record access, privacy, labor-relations, retention, and human-subjects classification are organization-specific and unresolved.","source_ids":["S1","S2","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Better identification of systemic coordination failures could materially improve preparedness decisions, but no source measures downstream preparedness gains from this intervention.","source_ids":["S3","S4"]},"stakeholder_pull":{"score":3,"rationale":"Official programs seek streamlined, consistent evaluation and cross-agency synthesis, but there is no expressed demand for residual reconstruction specifically.","source_ids":["S1","S2","S3"]},"incremental_advantage":{"score":2,"rationale":"Current official, research, and commercial approaches already cover most workflow functions; superiority on time, cross-layer findings, or signal preservation is wholly untested.","source_ids":["S1","S2","S5","S6"]},"distinctiveness_plausibility":{"score":3,"rationale":"The combined reconstructive hierarchy, raw-audit control, and decompression rule remains meaningfully contrastive, although its ingredients are familiar and the search cannot establish world novelty.","source_ids":["S1","S5","S6"]},"technical_implementability":{"score":4,"rationale":"Existing platforms demonstrate structured scenarios, event capture, conditional injects, analytics, and fast reporting; the bounded shadow prototype is technically straightforward.","source_ids":["S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"Exercise directors and evaluation leads can authorize a read-only shadow study, but record-owner approval, participant protections, and any research or labor review must be resolved locally.","source_ids":["S2","S7","S8"]},"evidence_readiness":{"score":3,"rationale":"A comparative retrospective protocol and measurable outcomes can be preregistered, but no authorized exercise record, baseline workload measurement, or validated residual taxonomy is currently available.","source_ids":["S4","S5"]},"safety_net_benefit":{"score":4,"rationale":"Independent raw review, protected-signal bypass, participant contestability, and rollback to ordinary full-record review plausibly reduce harm versus ungoverned automated scoring, but these controls have not been tested.","source_ids":["S2","S7"]},"scalability":{"score":3,"rationale":"Reusable schemas and software can scale across repeated exercises, but scenario preparation, expert expectation setting, audits, and local governance may remain labor intensive.","source_ids":["S5","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"One preregistered retrospective shadow study on one completed low-stakes exercise, limited to approximately 6–10 injects and two evaluation arms, including authorization, expectation design, dual coding, independent audit, analysis, and reporting.","confidence":"MODERATE","assumptions":["Approximately 250–550 blended staff hours across an exercise lead, subject-matter experts, coders, an independent auditor, analyst, and privacy or research reviewer.","Uses existing recordings, notes, and ordinary office or open-source tooling; excludes new live exercise delivery and major software development.","Loaded labor is assumed above the BLS $41.41 median hourly wage to cover benefits, specialist time, and overhead."],"source_ids":["S5","S7","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Build a reusable residual taxonomy, model-versioning workflow, secure data pipeline, audit sampler, reconstruction tests, evaluator training materials, governance documentation, and connectors to an existing exercise platform.","confidence":"LOW","assumptions":["Roughly 0.5–1.5 staff-years distributed across exercise design, software or data engineering, evaluation science, security, privacy, and project management.","Reuses existing exercise-record and identity systems rather than building an end-to-end simulation platform.","No commercial product pricing or internal integration estimate was publicly verified."],"source_ids":["S5","S6","S8"]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"Shadow launch across approximately three to six exercises in one organization, with evaluator training, help desk, independent audit, threshold calibration, fallback drills, and governance review before any operational reliance.","confidence":"LOW","assumptions":["Includes 750–2,000 labor hours plus modest hosting, training, and security-review costs.","Original HSEEP-style evaluation continues in parallel during launch, temporarily duplicating effort.","Excludes participant-performance use and automatic operational-policy changes."],"source_ids":["S1","S2","S5","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Maintain schemas and software, support several exercises annually, conduct raw audits and reconciliations, review model drift and contested findings, retrain evaluators, and retain secure records.","confidence":"LOW","assumptions":["Approximately 0.5–1.5 full-time-equivalent labor plus hosting, security, and periodic independent review.","Exercise volume, record sensitivity, retention requirements, and commercial licensing are unknown.","The band does not include the underlying cost of designing and running the tabletop exercises themselves."],"source_ids":["S5","S6","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent research and official oversight document manual evaluation burden, delayed feedback, underdeveloped assessment, and failure to identify cross-exercise systemic patterns.","source_ids":["S3","S4","S5"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Exercise directors, FEMA-aligned evaluation leads, preparedness agencies, and emergency-management directors already own exercise evaluation and improvement workflows and can authorize a retrospective shadow study.","source_ids":["S1","S2","S3","S8"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim specifies reconstruction fidelity, protected-signal recall, independently confirmed cross-team findings, timeliness, and fully loaded evaluator hours against ordinary full-record review.","source_ids":["S1","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"One completed low-stakes exercise can support a preregistered, read-only, parallel comparison without changing participant assessment or operational decisions.","source_ids":["S2","S5","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"UNCERTAIN","reason":"The design supplies strong procedural safeguards, but the actual record owner, participant notice or consent terms, privacy and labor rules, protected-disclosure handling, and research-versus-internal-QI determination are unknown and must be resolved before data access.","source_ids":["S2","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four bands have bounded scopes and explicit labor assumptions anchored to an official wage benchmark, but confidence is low beyond the first study because integration, licensing, record sensitivity, and exercise volume are unknown.","source_ids":["S5","S6","S8"]}},"next_evidence_step":"Obtain record-owner and privacy or research authorization for one completed low-stakes exercise; preregister 6–10 injects, team/coordination/strategic interfaces, the residual taxonomy, and immutable expected-response vectors using only the scenario, playbook, and inject materials. Randomly assign two evaluator teams to (A) ordinary full-chronology HSEEP-style review and (B) expected-plus-residual hierarchical review, while a blinded third team establishes a full-record reference chronology and audits a random sample plus risk-stratified intervals. Compare macro-F1 for decision-relevant event reconstruction, timestamp error, protected and safety-signal recall, interrater agreement, independently confirmed omission/handoff/assumption findings, cross-layer findings, time to first usable debrief question, false escalations, fallback frequency, and fully loaded hours. Falsify the intervention if protected-signal recall is below 100%, reconstruction macro-F1 is below 0.90, median timestamp error exceeds the preregistered tolerance, the residual arm omits any consequential reference finding not omitted by the comparator, yields no additional or earlier independently confirmed cross-team finding, or fails to reduce total effort by at least 20%. Immediately revert contested scopes to full-record review.","blocking_evidence":["No authorized completed exercise record or confirmed record owner was identified.","No organization-specific baseline shows what fraction of evaluation time is spent on routine reconstruction or whether full review exceeds the available budget.","No validated taxonomy, semantic reconstruction tolerance, interrater-reliability estimate, or protected-signal test set exists for open-ended tabletop behavior.","No evidence shows that hierarchical residual review improves finding quality, timeliness, or total workload relative to HSEEP-style review or current commercial platforms.","No local determination addresses privacy, labor and personnel rules, protected disclosures, retention, participant contestability, or human-subjects-research status.","No organization-specific software, integration, licensing, security-review, or recurring-support estimate was obtained."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This bounded search found substantial adjacent and colliding practice but did not measure world novelty, patentability, freedom to operate, market size, realized impact, or the completeness of unpublished, classified, proprietary, military, or non-English prior art. The only claim preserved is a field-testable performance contrast, not novelty ownership.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure written record-owner authorization and an organization-specific privacy, labor, protected-disclosure, and research/QI determination.","Measure ordinary-review workload and confirm that routine reconstruction is a binding problem in the selected exercise.","Freeze and validate the expected-response schema and residual taxonomy without inspecting exercise outcomes; establish reproducible coding and reconstruction thresholds.","Run the preregistered parallel shadow comparison against full-chronology HSEEP-style review and a blinded full-record reference.","Demonstrate 100% protected-signal recall, acceptable reconstruction fidelity, no additional consequential omissions, at least one additional or earlier independently confirmed cross-team finding, and at least 20% lower fully loaded evaluator effort.","Obtain implementation quotes or internal estimates for integration, security, training, audit, fallback, and annual support before adoption inquiry."],"reason":"Web evidence verifies the problem, credible authorities, implementable components, and substantial prior-art collision, but it cannot establish the proposal's remaining advantage. Reconstruction fidelity, hierarchy-suppression risk, protected-signal preservation, cross-team finding yield, and net evaluator workload require access to proprietary exercise records and a comparative field study. Under the controller rule, that requires STOP_EMPIRICAL_RESEARCH_NEEDED rather than further bounded web research."},"proposal_index":4}