{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","cell_id":"predictive_residual_processing__futurism_foresight","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_id":"prp-foresight-residual-tabletop-evaluation-004","proposal_index":4,"version":0,"title":"Hierarchical Residual Evaluation for Futures Tabletop Exercises","problem":"A multi-team futures tabletop exercise produces a dense chronological record of participant decisions, communications, handoffs, and delays. Evaluators must review routine playbook-conforming behavior alongside omissions, conflicting assumptions, unexpected adaptations, and cross-team coordination failures. A compliance checklist can identify deviations from prescribed actions, but it does not maintain an uncertain forward model, preserve reconstructive context, or distinguish an informative mismatch from harmless variation.","actors":["Exercise designers who define the future scenario and injects","Playbook and scenario stewards who specify expected responses","Participants who make decisions within the tabletop","Observers who record actions, communications, omissions, and timing","Evaluation leads who reconstruct the exercise and conduct the after-action review","Safety and participant-welfare officers with independent halt authority","Operational or strategic owners who may later authorize preparedness changes","People and groups represented or affected by the simulated decisions"],"observable_state":"For a selected exercise, the full timestamped record can be coded into expected actions, observed actions, decision latency, handoffs, information requests, escalation paths, and unresolved issues. Review logs can show how much evaluator effort is spent documenting expected behavior, whether deviations are discovered only after the relevant debrief window, and whether local observations combine into cross-team or system-level failures.","consequence":"If evaluator attention is consumed by reconstructing predictable actions, consequential omissions and interactions may be identified late or reduced to isolated checklist failures. The after-action review can then reward conformity to the existing playbook without revealing where the playbook, scenario model, organizational boundary, or inter-team assumptions failed.","affected_objective":"Detect and explain consequential coordination failures, assumption breaks, and adaptive responses during futures exercises within a bounded evaluation budget, while preserving the complete exercise record, participant rights, dissent, and independent safety oversight.","intervention":"Before the exercise, designers create a versioned expected-response model for each inject and scenario phase. It predicts a structured decision-relevant event vector: expected actions, acceptable timing ranges, information requirements, responsible roles, handoffs, escalation paths, and uncertainty. Observers retain the complete exercise record but compare actual events with the applicable prediction and encode structured residuals such as omitted action, unexpected delay, sequence inversion, ownership conflict, failed handoff, novel workaround, unmodeled dependency, contradictory assumption, or unobserved response. Residuals are weighted by observation reliability, prediction uncertainty, consequence, reversibility, affected-group significance, cross-team reach, and evaluator cost. At the team layer, deviations explainable by a local model remain local; unresolved residuals propagate to coordination and strategic layers, where they are compared with higher-level expectations. Every residual carries the inject, prediction version, timestamp, observer provenance, uncertainty, and raw-record pointer. The evaluation lead reconstructs the structured exercise timeline from expected events plus residuals. Validated residuals open defined replay or debrief questions and enter a separate prediction-error review that may recommend changing the playbook, scenario, training, role boundary, or model. Random and risk-stratified raw intervals, scheduled full-timeline reconciliations, and cross-observer comparisons test what the residual hierarchy suppressed. Real emergencies, participant-welfare concerns, rights issues, protected disclosures, model mismatch, missing observation, drift, or reconstruction failure bypass residual processing and restore full-record review.","structural_mapping":[{"archetype_element":"Prediction target and observation boundary","domain_realization":"The prediction target is the structured decision-relevant response to a named exercise inject over a defined phase, not every word spoken and not the real future represented by the scenario."},{"archetype_element":"Generative model state","domain_realization":"A versioned inject-response model represents expected actions, timing, roles, handoffs, information flows, escalation paths, and uncertainty."},{"archetype_element":"Model scope and horizon","domain_realization":"Each prediction is bounded to specified injects, exercise phases, teams, roles, and timescales and expires after the applicable decision window."},{"archetype_element":"Predictive feedforward model","domain_realization":"Expected responses are frozen before an inject is released, permitting comparison with behavior rather than retrospective construction of an ideal response."},{"archetype_element":"Expected and actual behavior","domain_realization":"The model supplies the expected event vector; observers record the actual structured events with timestamps, provenance, missingness, and links to the full record."},{"archetype_element":"Prediction comparator","domain_realization":"A declared comparator preserves omission, direction, sequence, timing, ownership, handoff, dependency, and assumption differences."},{"archetype_element":"Prediction-error signal","domain_realization":"Each residual carries the unexplained difference plus sufficient prediction identity and context to reconstruct the structured exercise timeline."},{"archetype_element":"Precision weighting and residual budget","domain_realization":"Residual priority combines observer reliability, model uncertainty, consequence, reversibility, cross-team reach, affected-group significance, and evaluator capacity; suppressed residuals remain auditable."},{"archetype_element":"Residual propagation channel","domain_realization":"Local evaluators route unresolved residuals to coordination and strategic review layers rather than retransmitting every expected event upward."},{"archetype_element":"Hierarchical prediction stack","domain_realization":"Team-level models explain local behavior, coordination-level models predict inter-team handoffs, and strategic-level models predict system posture; only unexplained, precision-weighted error rises between layers."},{"archetype_element":"Confidence and uncertainty state","domain_realization":"Uncertainty in the expected response, observer record, residual classification, and reconstruction remains separate and visible."},{"archetype_element":"Update rule","domain_realization":"Reviewed residuals may revise the response model, playbook hypothesis, scenario boundary, training need, or role contract, but no exercise observation automatically changes operational policy."},{"archetype_element":"Model synchronization and provenance","domain_realization":"Designers, observers, and evaluators use the same inject-response checksum, and every residual carries model, inject, team, observer, timestamp, and transformation identifiers."},{"archetype_element":"Freshness and drift","domain_realization":"Predictions expire with their inject window; correlated or repeated residuals across phases indicate that the expected-response model or scenario assumptions no longer fit."},{"archetype_element":"Reconstruction requirement","domain_realization":"The expected event vectors and residuals must reproduce the complete structured timeline within a declared semantic and temporal tolerance, while full recordings and notes remain retrievable."},{"archetype_element":"Raw audit and resynchronization","domain_realization":"Independent reviewers inspect random and risk-stratified raw intervals, and the evaluation team periodically reconciles the residual reconstruction with the complete chronology."},{"archetype_element":"Fallback and decompression","domain_realization":"Version mismatch, missing observers, high uncertainty, unexpected scenario transition, audit disagreement, or reconstruction error restores complete chronological review for the affected scope."},{"archetype_element":"Safety-critical bypass","domain_realization":"Real-world emergencies, participant distress, discriminatory conduct, rights concerns, protected disclosures, and security incidents bypass exercise scoring and reach independent authorities in full."},{"archetype_element":"Attention and bandwidth budget","domain_realization":"The exercise owner declares evaluator hours and counts model construction, observation, residual review, replay, audit, reconciliation, and fallback against the same budget."}],"mechanism_mapping":[{"mechanism_slug":"hierarchical_prediction_error_loop","role":"Compares behavior with expectations at team, coordination, and strategic layers so unresolved errors propagate upward without duplicating routine lower-level events.","counterfactual_removal":"The proposal would be a flat exception log and could miss failures that emerge only from interactions among individually reasonable team actions."},{"mechanism_slug":"predictive_codec","role":"Uses matched inject-response models to reconstruct the structured timeline as expected events plus encoded corrections.","counterfactual_removal":"Residuals would become context-free anomaly labels rather than a reconstructive representation of exercise behavior."},{"mechanism_slug":"innovation_residual_filter","role":"Weights observed deviations according to uncertainty in the response model and observer evidence before correcting the reconstructed timeline.","counterfactual_removal":"Noisy observer disagreements and reliable small failures would be treated as equivalent based only on apparent magnitude."},{"mechanism_slug":"precision_weighted_error_gate","role":"Allocates review capacity using consequence, reliability, reversibility, cross-team reach, affected-group significance, and model uncertainty.","counterfactual_removal":"The residual channel could be saturated by visible but harmless deviations while subtle coordination failures remained buried."},{"mechanism_slug":"event_triggered_residual_reporting","role":"Routes validated omissions, delays, handoff failures, assumption conflicts, and novel adaptations to the appropriate evaluation layer.","counterfactual_removal":"Observers would still transmit complete chronological logs upward, leaving the attention constraint unchanged."},{"mechanism_slug":"model_version_checksum_handshake","role":"Confirms that observers and evaluators interpret behavior against the same inject, response schema, timing rule, and expected-response version.","counterfactual_removal":"A valid residual could be applied to an obsolete inject or playbook expectation and reconstruct a misleading event."},{"mechanism_slug":"prediction_error_replay_buffer","role":"Stores selected residuals with surrounding raw context for debrief replay, causal analysis, calibration, and regression testing of revised exercises.","counterfactual_removal":"Material mismatches would be reduced to summary findings without sufficient context for challenge or later testing."},{"mechanism_slug":"residual_comparison_test","role":"Tests residual structure against a simple checklist model, alternative response model, cross-observer coding, and raw intervals.","counterfactual_removal":"Evaluators could classify patterned model failure as participant error without testing whether the expected response was misspecified."},{"mechanism_slug":"shadow_raw_channel_sampling","role":"Routes random and risk-stratified full exercise intervals to an independent reviewer who does not rely on the production residual hierarchy.","counterfactual_removal":"Low-level dissent, novel behavior, or contextual evidence explained away by the hierarchy could remain invisible."},{"mechanism_slug":"periodic_full_state_resynchronization","role":"Reconciles the complete exercise chronology with the residual reconstruction at phase boundaries and after drift triggers.","counterfactual_removal":"Timing, ordering, and attribution errors could accumulate across injects and distort the final after-action narrative."},{"mechanism_slug":"raw_signal_fallback_switch","role":"Restores full-record review when safety, compatibility, observation, uncertainty, novelty, or reconstruction conditions fail.","counterfactual_removal":"The exercise would remain compressed precisely when the scenario departed from the model or observer coverage failed."},{"mechanism_slug":"prediction_error_review","role":"Determines whether a residual reflects participant action, playbook failure, scenario design, missing information, role ambiguity, observation error, or model-boundary failure.","counterfactual_removal":"After-action findings could blame participants for deviations produced by a wrong scenario model or organizational contract."},{"mechanism_slug":"surprise_to_action_bridge","role":"Connects a validated residual to a named replay, debrief question, evidence request, model revision proposal, or preparedness owner.","counterfactual_removal":"Important surprises could remain dashboard markers without acknowledgement, explanation, or bounded follow-up."}],"causal_chain":["Exercise designers freeze uncertain, versioned response predictions before each inject is released.","Observers capture actual structured actions, omissions, timing, communications, and handoffs while retaining the complete record.","Comparators calculate signed, categorical, temporal, and relational residuals against the applicable prediction.","Precision weighting separates consequential reliable mismatches from benign variation and observation noise.","Team layers absorb locally explained differences, while unresolved coordination errors propagate to higher evaluation layers.","Evaluators reconstruct the structured timeline from synchronized predictions and residuals and request raw context when needed.","Validated residuals trigger named replays, debrief questions, or evidence requests rather than automatic judgments about participants.","Prediction-error review determines whether the behavior, playbook, scenario, role boundary, observation process, or model should change.","Raw interval sampling, cross-observer comparison, and full-timeline reconciliation test for hierarchical suppression and cumulative error.","Safety events, protected concerns, missing observation, drift, incompatibility, or reconstruction failure restore full-record handling."],"baseline":"A conventional tabletop evaluation retains a complete chronology, observer notes, recordings, inject-response matrices, and checklist scores. Evaluators review the full record after the exercise and produce an after-action report, but expected behavior is not maintained as a synchronized uncertain predictor whose residuals control representation, escalation, learning, auditing, and fallback.","nearest_rivals":["A playbook-compliance checklist, which identifies prescribed actions completed or missed but does not model uncertainty, reconstruct the event timeline, or surface interactions among layers.","A standard after-action review based on the full chronology, which preserves context but spends evaluation effort on expected and unexpected behavior alike.","Observer anomaly flags, which can mark unusual events but lack a frozen generative baseline, model-version compatibility, residual reconstruction, and systematic raw audit.","An automated exercise-summary system, which compresses the record by relevance or language patterns rather than by prediction-relative error with synchronized fallback.","An inject-response matrix, which links prompts to observed actions but does not propagate unresolved error through team, coordination, and strategic models.","A live exercise-control dashboard, which supports facilitation but may track events and statuses without using residuals as both the reconstructive message and the governed teaching signal."],"remaining_contrastive_claim":"The proposal should be preferred only if synchronized hierarchical expectations plus structured residuals preserve a decision-relevant reconstruction of the exercise while helping evaluators identify consequential cross-team mismatches and adaptive responses within the declared budget, and only if independent raw review shows that dissent, context, safety issues, and low-level signals are not being explained away.","authority_safety":{"decision_authority":"The exercise director may authorize the evaluation model, scope, and shadow pilot. Scenario and playbook stewards may propose expectations but cannot adjudicate their own model failures alone. The evaluation lead controls analytical findings, the safety officer has independent halt and bypass authority, participants may contest reconstructions, and operational leaders retain authority over real preparedness or policy changes.","authorized_first_step":"Perform a read-only retrospective shadow recoding of one already-authorized, completed low-stakes exercise. Freeze the expected-response model before inspecting the scored outcomes, preserve the original evaluation, and prevent residual findings from affecting participant assessment or operational decisions.","excluded_actions":["Using residual scores for discipline, promotion, individual performance ranking, or blame attribution","Replacing real emergency communications or safety reporting with residual-only messages","Deleting, truncating, or withholding the complete exercise record","Automatically changing operational plans, policy, resource allocation, or training requirements","Allowing scenario authors to classify contradictory evidence without independent review","Suppressing participant dissent, rights concerns, welfare issues, protected disclosures, or discriminatory conduct","Changing injects or expected-response rules after observing behavior without a new version and trace","Expanding participant surveillance beyond the exercise’s existing consent and authorization","Treating model-conforming behavior as proof that the playbook is correct"],"halt_rollback":"Halt residual evaluation for the affected scope if a real safety event occurs, observer coverage is missing, a protected concern is compressed, model versions diverge, participants cannot contest reconstruction, the raw audit reveals a material omission, or reconstruction exceeds its predeclared tolerance. Roll back to the complete chronology and standard after-action process, label affected residual findings provisional, freeze the model and thresholds, preserve the audit trail, and require independent review before reuse."},"negative_tests":{"strongest_counterevidence":"In retrospective shadow coding, the residual hierarchy explains away low-level dissent or context that full-record reviewers identify as consequential, reconstructs event order or responsibility incorrectly, produces fewer useful cross-team findings than the standard review, increases confirmation of the playbook, or costs as much evaluator effort after model construction, coding, audit, replay, and fallback are counted.","problem_falsifier":"The problem is falsified for the selected exercise if the complete record already fits within the declared evaluation budget, routine actions do not materially consume review effort, consequential deviations and cross-team interactions are consistently identified within the debrief window, or the exercise is intentionally so novel that defensible response prediction is infeasible.","intervention_falsifier":"The intervention is falsified if expected events plus residuals cannot reconstruct the structured chronology; independent reviewers cannot reproduce residual classifications; protected, dissenting, or low-level observations are omitted more often than under full review; residuals retain systematic unexplained structure after bounded model revision; hierarchy adds no cross-team findings; or total model and audit cost is not lower at the required fidelity.","risks":["Strong playbook expectations may turn legitimate adaptation into apparent error.","Low-level observers may be overruled when higher layers explain away their residuals.","A shared wrong model may make a quiet strategic layer appear reassuring.","Observers may code behavior to fit expected categories and miss novel action.","Residual prioritization may discount consequences borne by groups poorly represented in the scenario model.","Participant anonymity or protected disclosures may be weakened by detailed residual provenance.","After-action reviewers may use residuals to assign blame despite uncertainty and model misspecification.","Scenario designers may revise expectations retrospectively to make the exercise appear successful.","Frequent fallback may eliminate the attention benefit and create pressure to weaken safety thresholds.","Saved evaluator capacity may be filled with additional injects, reducing rather than improving reflective depth."]},"next_evidence_step":"Pre-register a bounded retrospective shadow study using one completed low-stakes tabletop with an authorized full record. Before examining its scored outcomes, freeze expected-response vectors for a limited set of injects and define team, coordination, and strategic interfaces. Have one group reconstruct and evaluate the exercise using synchronized predictions plus residuals and another use the ordinary chronology, with independent reviewers auditing random and risk-stratified raw intervals. Compare temporal and semantic reconstruction agreement, consequential omission and handoff findings, cross-layer findings, protected-signal preservation, observer agreement, time to prepare debrief questions, false escalation, fallback frequency, and total modeling and review effort. Define fidelity, protected-signal, hierarchy-suppression, workload, and halt thresholds before scoring.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Proposal 1 governs external horizon-scanning throughput by predicting scenario-driver assessments and routing deviations from distributed scanning desks. Proposal 4 instead evaluates enacted decisions inside a bounded futures exercise: it predicts responses to scenario injects and propagates unresolved behavioral and coordination errors through team, coordination, and strategic layers. Its principal outcome is an auditable after-action diagnosis rather than prioritization of environmental evidence, and it can be adopted without a distributed scanning network. Proposal 2 governs repeated Delphi elicitation by predicting each expert’s next-round judgment and reconstructing questionnaires from edits. Proposal 4 does not compress stated expert forecasts; it compares observed simulated actions, omissions, timing, and handoffs with an expected-response model. Its primary hazard is playbook conformity suppressing adaptive behavior, rather than displayed answers anchoring private judgment. Proposal 3 separates environmental signals caused by the focal organization’s real anticipatory actions from plausibly independent evidence using an efference-copy echo model. Proposal 4 makes no causal attribution about public trend signals; it operates within a designed simulation and uses hierarchical residuals to reveal coordination and model failures. It is independently adoptable by an exercise-evaluation team without horizon scanning, Delphi elicitation, or strategic-echo filtering.","revision_record":{"parent_version":null,"progress_targets_addressed":["Created a fourth independently adoptable candidate with a different problem, intervention, actors, evidence unit, and causal path from proposals 1, 2, and 3.","Preserved forward prediction, structured comparison, residual propagation, reconstruction, model updating, synchronization, raw auditing, and full-signal fallback."],"conceptual_changes":["Made simulated organizational behavior the observation target and expected inject response the maintained prediction.","Introduced hierarchical explaining-away across team, coordination, and strategic evaluation layers.","Separated diagnostic residual review from participant assessment and operational decision authority."],"operational_changes":["Specified frozen inject-response vectors, structured behavioral residuals, cross-layer propagation, raw interval audits, full-timeline reconciliation, participant contestability, safety bypasses, and standard-review rollback.","Restricted the first evidence step to retrospective shadow recoding of a completed low-stakes exercise."],"evidence_changes":["Added direct tests of structured timeline reconstruction, cross-layer finding quality, observer agreement, hierarchy suppression, debrief preparation time, protected-signal preservation, fallback use, and total evaluation cost.","Specified problem and intervention falsifiers for exercise-evaluation overload and playbook-model failure."],"claim_changes":["Made no claim of novelty, prevalence, demand, effect size, or superior exercise outcomes.","Conditioned adoption on reconstruction fidelity, consequential residual detection, independent raw-audit performance, participant safeguards, and net evaluation cost."]}}