{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"predictive_residual_processing__futurism_foresight:P1:v0","cell_id":"predictive_residual_processing__futurism_foresight","search_queries":["site:oecd.org strategic foresight horizon scanning capacity challenges weak signals government 2024","site:gov.uk horizon scanning Futures Toolkit scanning resource analyst time","site:knowledge4policy.ec.europa.eu horizon scanning practitioner's guide weak signals 2023","distributed horizon scanning platform signals collaborative foresight software","\"Enhancing horizon scanning\" pre-developed scenarios process improvement weak signals paper","horizon scanning scenario assumptions monitoring signposts indicators prior art strategic foresight","strategic foresight monitoring scenario indicators signposts weak signals scenario planning","site:oecd.org \"Building capacity in technology horizon scanning\" 129 exercises challenges interpreting early signals","\"Web-based horizon scanning: concepts and practice\" information overload","Palomino web-based horizon scanning concepts practice 2012 full text","horizon scanning analyst information overload weak signals study","strategic foresight horizon scanning \"time-consuming\" signals analysts","site:bls.gov 2025 employer costs employee compensation professional occupations hourly December 2025","site:bls.gov/ooh data scientists median pay 2025 software developers management analysts","FIBRES pricing foresight platform 2026","Futures Platform pricing subscription foresight software","\"Scenarios and early warnings as dynamic capabilities\" abstract case study","Ramirez Osterman Gronquist 2013 scenarios early warnings dynamic capabilities managerial attention","scenario early warning monitoring indicators strategic foresight academic case study","OECD Building Anticipatory Capacity with Strategic Foresight in Government publication date 2025","Anticipating surprise case early warning system Rijkswaterstaat publication date 2018 Policy and Society","FIBRES pricing page publication date 2026","Integrating scenario planning indicator-based early detection 2008 publication date"],"sources":[{"source_id":"S1","title":"Building capacity in technology horizon scanning: A guide for policymakers","publisher":"OECD Publishing","url":"https://www.oecd.org/en/publications/building-capacity-in-technology-horizon-scanning_b4f0d383-en.html","source_class":"OFFICIAL_GUIDANCE","publication_date":"2026-04-14","accessed_at":"2026-08-03","claims_supported":["Horizon scanning is used by government units and international initiatives.","A review of 129 international exercises from 2020–2025 found diverse practices and methodological advances.","Interpreting early signals remains challenging, and robust standards and collaboration are needed."]},{"source_id":"S2","title":"Building Anticipatory Capacity with Strategic Foresight in Government: Lessons from Lithuania, Italy, and Malta","publisher":"OECD Publishing","url":"https://www.oecd.org/content/dam/oecd/en/publications/reports/2025/05/building-anticipatory-capacity-with-strategic-foresight-in-government_ed581d05/d7eb0bb6-en.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-05-09","accessed_at":"2026-08-03","claims_supported":["Government foresight work can be time-consuming and benefits from distributed capabilities across agencies.","Singapore's Centre for Strategic Futures and its deputy-secretary-level Strategic Foresight Network are identifiable adopters and authorizers.","Leadership commitment, cross-sector coordination, authorizing environments, pilots, and communities of practice are implementation requirements.","Foresight outputs are often not acted on systematically when they are insufficiently integrated into governance structures."]},{"source_id":"S3","title":"The Futures Toolkit","publisher":"UK Government Office for Science","url":"https://www.gov.uk/government/publications/futures-toolkit-for-policy-makers-and-analysts/the-futures-toolkit-html","source_class":"OFFICIAL_GUIDANCE","publication_date":"2024-08-29","accessed_at":"2026-08-03","claims_supported":["Horizon scanning can be continuous and distributed among contributors from diverse backgrounds.","The documented workflow can generate at least 60 weekly scans from ten contributors over six weeks, creating a substantive organization and analysis task.","Structured scan records, databases, collaborative scanning networks, workshops, and downstream scenario or policy-stress-testing are established practice.","Narrow scanning risks missing important change, while sensitive consequences require careful handling."]},{"source_id":"S4","title":"Enhancing horizon scanning by utilizing pre-developed scenarios: Analysis of current practice and specification of a process improvement to aid the identification of important weak signals","publisher":"University of Strathclyde repository; Technological Forecasting and Social Change","url":"https://strathprints.strath.ac.uk/61436/","source_class":"PRIMARY_RESEARCH","publication_date":"2017-12-01","accessed_at":"2026-08-03","claims_supported":["Using pre-developed scenario storylines to guide weak-signal identification is published prior art.","The paper explicitly integrates scenario planning and horizon scanning to improve organizational preparedness."]},{"source_id":"S5","title":"Anticipating surprise: The case of the early warning system of Rijkswaterstaat in the Netherlands","publisher":"Oxford University Press, Policy and Society","url":"https://academic.oup.com/policyandsociety/article/37/4/473/6402477","source_class":"PRIMARY_RESEARCH","publication_date":"2018-09-30","accessed_at":"2026-08-03","claims_supported":["A government agency implemented an early-warning team with a distributed network of internal and external correspondents.","Signals were filtered into a short list for board attention and connected to follow-up actions.","Institutionalization made signals more routine and less surprising, demonstrating attention and model-lock-in risks.","Effective early warning depended on authority, organizational practice, interpretation, and willingness to challenge existing assumptions."]},{"source_id":"S6","title":"Integrating scenario planning and indicator-based early detection for scenario transfer","publisher":"Fraunhofer-Gesellschaft repository; International Journal of Technology Intelligence and Planning","url":"https://publica.fraunhofer.de/entities/publication/04f1c755-5519-4d3e-9282-c924f57b2c81","source_class":"PRIMARY_RESEARCH","publication_date":"2008","accessed_at":"2026-08-03","claims_supported":["Systematic scenario-based early detection and continuous verification of developments are established research concepts.","Prior work identifies limitations of fixed indicators and uses scenarios to provide a more holistic view of weak signals."]},{"source_id":"S7","title":"FIBRES Pricing","publisher":"FIBRES Online Ltd","url":"https://www.fibresonline.com/pricing","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"No publication date shown; live pricing accessed 2026-08-03","accessed_at":"2026-08-03","claims_supported":["Commercial collaborative foresight software is already available.","Published 2026-accessed base prices are $8,100 per year for ten users and $23,200 per year for an enterprise plan with 100 users.","Existing products include collaboration, visualization, customization, integrations, support, and optional data sources."]},{"source_id":"S8","title":"Employer Costs for Employee Compensation — December 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/archives/ecec_03202026.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-03-20","accessed_at":"2026-08-03","claims_supported":["Average fully loaded civilian labor cost was $48.78 per hour in December 2025.","Average state and local government labor cost was $65.68 per hour, providing a transparent resource-equivalent basis for 2026 cost estimates."]}],"problem_evidence":{"support":"MODERATE","rationale":"Independent official and primary sources show that distributed horizon scanning produces many records, is time-consuming, faces early-signal interpretation and attention problems, and can lose strategic effect as reporting becomes routine. However, no source verifies the proposal's setting-specific assertion that complete weekly driver packets mostly repeat expected trajectories, exceed a declared review budget, or delay material contradictions. Those prevalence and latency claims require local workflow data.","source_ids":["S1","S2","S3","S5"]},"stakeholder_evidence":{"support":"STRONG","rationale":"Identifiable institutions already commission, authorize, and operate closely related work: Singapore's Centre for Strategic Futures and Strategic Foresight Network, the UK Government Office for Science, and Rijkswaterstaat. OECD evidence explicitly describes needs for distributed capacity, authorizing environments, integration, and pilots. This establishes credible adopters and expressed need for better foresight operations, though none has requested this exact residual-exchange design.","source_ids":["S2","S3","S5"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Scenario-enhanced horizon scanning","similarity":"Uses pre-developed scenarios to focus weak-signal identification and challenge scenario assumptions.","remaining_difference":"Does not establish reconstructive residual cards, synchronized baseline versions, independent raw-packet audits, residual-error budgets, or automatic full-packet fallback.","source_ids":["S4"]},{"name":"Scenario-based indicator early-detection system","similarity":"Continuously monitors developments against scenarios and uses indicators to detect unfolding change.","remaining_difference":"Uses monitoring indicators rather than a distributed semantic residual protocol that reconstructs each complete structured driver assessment.","source_ids":["S6"]},{"name":"Rijkswaterstaat Early Warning System","similarity":"Uses a distributed correspondent network, central filtering, scarce board attention, strategic follow-up, and a mechanism intended to surface surprise.","remaining_difference":"Correspondents submit signals directly; there is no published shared expected-state vector, model-version handshake, reconstruction test, random raw audit, or decompression rule.","source_ids":["S5"]},{"name":"FIBRES collaborative foresight platform","similarity":"Provides multi-user foresight collaboration, structured content, visualization, customization, and integrations at organizational scale.","remaining_difference":"The published offering does not claim scenario-relative reconstructive residual encoding, protected-class bypass, independent suppression auditing, or error-triggered restoration of full packets.","source_ids":["S7"]}],"distinctive_claim_remaining":"Against (a) complete-packet review, (b) ordinary tag-and-priority triage, and (c) scenario-signpost or early-warning filtering, a synchronized expected-state-plus-residual workflow will reduce fully loaded review effort while remaining non-inferior on blinded detection of assumption changes, cross-driver interactions, novel actors, and protected-source evidence; matching versions, random raw audits, periodic reconciliation, and scoped fallback will keep consequential omissions within a preregistered tolerance. The claim fails if reconstruction fidelity or protected-signal recall is inferior, silence remains ambiguous, or preparation, audit, and fallback effort eliminates the review savings.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Structured scanning records, distributed contributor networks, scenario-conditioned monitoring, central signal triage, commercial collaboration software, and management follow-up are all demonstrated. Implementing forms, version identifiers, heartbeats, queues, frozen baselines, audit sampling, and fallback is technically conventional. Feasibility remains unverified for semantic reconstruction reliability, cross-analyst consistency, protected-signal classification, independent audit staffing, integration with records systems, and total workload. A shadow retrospective can stay within existing data authority; live suppression would require privacy, records-retention, access-control, labor/workflow, and strategic-decision authorization reviews.","source_ids":["S2","S3","S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Earlier attention to assumption-breaking evidence could improve preparedness, and sources show that early-warning signals can affect policy action. The candidate's actual backlog, missed-signal burden, and effect size are not measured.","source_ids":["S1","S3","S5"]},"stakeholder_pull":{"score":3,"rationale":"Government foresight units express needs for capacity, coordination, timely interpretation, and pilots, but no adopter has requested or committed to this exact residual workflow.","source_ids":["S1","S2","S3"]},"incremental_advantage":{"score":2,"rationale":"The design adds explicit reconstruction, version synchronization, raw audits, protected bypasses, and fallback, but scenario-conditioned scanning, early-warning filtering, collaborative platforms, and scarce-attention triage already exist. Net advantage requires comparative testing.","source_ids":["S4","S5","S6","S7"]},"distinctiveness_plausibility":{"score":3,"rationale":"The exact governed combination was not found in the eight-source search, but its central scenario-monitoring and attention-filtering functions are adjacent to established research and practice. World novelty and patentability remain unmeasured.","source_ids":["S4","S5","S6","S7"]},"technical_implementability":{"score":4,"rationale":"A shadow implementation can use structured records, database versioning, logged queues, sampling, and access-controlled links without novel infrastructure. Semantic comparator reliability and model synchronization across human teams remain testing risks.","source_ids":["S3","S5","S7"]},"adoption_authority_feasibility":{"score":3,"rationale":"Government foresight steering groups, senior networks, and agency boards are credible authorizers for a shadow pilot. Wider adoption depends on leadership commitment, integration with governance routines, and resistance to workflow change.","source_ids":["S2","S5"]},"evidence_readiness":{"score":2,"rationale":"The proposal specifies observables and falsifiers, but no authorized corpus, packet-level workload logs, reliability labels, protected-signal ground truth, or committed evaluation team is available in the external record.","source_ids":["S1","S2","S5"]},"safety_net_benefit":{"score":4,"rationale":"Independent raw sampling, full-state reconciliation, protected bypasses, and fallback directly address confirmation bias and missed-signal risks documented in scanning practice. Their effectiveness and independence are not yet demonstrated.","source_ids":["S3","S5"]},"scalability":{"score":3,"rationale":"Distributed networks and commercial platforms already operate at multi-user scale, including a published 100-user product tier. Shared taxonomies, audit load, source diversity, and semantic disagreement may scale less favorably than storage or routing.","source_ids":["S2","S3","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Preregister and run a retrospective three-arm shadow evaluation on 200–500 historical packets from one authorized scanning program, including corpus preparation, blinded review, adjudication, audit sampling, and analysis.","confidence":"MODERATE","assumptions":["Approximately 250–650 labor hours at a $48.78–$65.68 fully loaded hourly benchmark.","Existing repository and access controls can be reused.","No custom production integration or live packet suppression.","Open-source or existing collaboration tools are used; purchasing a $8,100 annual platform could move the upper estimate toward the band ceiling."],"source_ids":["S7","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Build a bounded shadow prototype for one scanning network: structured driver schema, frozen expectations, residual forms, versioning, heartbeats, audit sampling, dashboards, fallback logging, governance rules, and evaluator training.","confidence":"MODERATE","assumptions":["Roughly 1,000–3,000 mixed analyst, engineering, governance, and training hours.","One commercial foresight subscription or equivalent existing infrastructure.","No automated strategic decisions and no replacement of the baseline workflow.","Existing identity, records, and document-storage systems expose usable integration points."],"source_ids":["S3","S7","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Launch across a multi-unit distributed network with dual-running, security and privacy review, integrations, training, protected-signal governance, independent audit staffing, reliability testing, and production support.","confidence":"LOW","assumptions":["Approximately three to eight full-time-equivalent years of combined implementation and change-management effort.","Deployment spans multiple organizational units but not a whole national government.","Costs include temporary dual operation and contingency for failed integrations.","Procurement, records, and legal review vary substantially by adopter."],"source_ids":["S2","S3","S7","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Operate one bounded network, including scenario-steward maintenance, audit review, adjudication, training refresh, software, support, periodic reconciliation, and evaluation reporting.","confidence":"MODERATE","assumptions":["Approximately 0.5–2.0 FTE of recurring stewardship and audit effort.","Published software base fees of $8,100–$23,200 annually, with integrations or data sources potentially extra.","Fallback review uses existing analyst capacity unless failure frequency becomes high.","Major model redesigns and organization-wide expansion are excluded."],"source_ids":["S2","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"External evidence verifies time, interpretation, information-selection, and attention difficulties in distributed foresight, but not the candidate's specific prevalence claim that routine complete packets are substantially redundant or cause delayed material discoveries in the intended setting.","source_ids":["S1","S2","S3","S5"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Singapore's Centre for Strategic Futures and senior Strategic Foresight Network, the UK Government Office for Science, and agency boards such as Rijkswaterstaat are identifiable operators or authorizers of closely related workflows.","source_ids":["S2","S3","S5"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The candidate can be tested against full-packet review, ordinary triage, and established scenario early-warning methods on effort, reconstruction, signal recall, protected-source coverage, and all-in cost.","source_ids":["S4","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A single historical corpus, sequential windows, frozen expectations, blinded reviewers, three comparators, preregistered metrics, and explicit falsifiers bound the next study.","source_ids":["S2","S5"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"A retrospective shadow study can retain baseline delivery and existing access permissions, make no strategic decisions, and halt on protected-signal suppression or audit failure. Live suppression remains outside the authorized first step.","source_ids":["S2","S3","S5"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The bands state deployment scope and labor assumptions and are anchored to current BLS fully loaded hourly costs and published foresight-platform prices. Integration and procurement uncertainty remains explicit.","source_ids":["S7","S8"]}},"next_evidence_step":"Secure one authorized historical corpus containing 200–500 sequential driver-assessment packets and packet-level timestamps. Before inspection of each window, freeze scenario-conditioned expectations using only prior material. Randomly assign blinded reviewers to: A) complete packets, B) complete packets with ordinary tags and priorities, or C) synchronized expectation-plus-residual cards with logged full-packet fallback. Preregister a primary success rule such as at least 20% lower fully loaded review time for C, at least 0.90 agreement on reconstructed structured driver state, no more than a five-percentage-point loss versus A in recall of adjudicated assumption changes and cross-driver findings, no loss in protected-source recall, and lower total hours after expectation preparation, auditing, reconciliation, and fallback. Include random and risk-stratified raw audits, version-mismatch and missing-heartbeat injections, and independent adjudication. Falsify the intervention if any safety-class item is suppressed, semantic reconstruction breaches tolerance, protected or underrepresented sources perform worse than A, residuals retain systematic unmodeled structure, fallback cannot restore context, or total effort is not lower. This evidence requires proprietary workflow data and participant testing, not additional bounded web search.","blocking_evidence":["Packet-level evidence that repeated expected content materially consumes the intended adopter's review budget.","Evidence that intake saturation, rather than absent evidence or downstream decision blockage, delays material findings.","Inter-rater reliability for the proposed structured driver-state and residual taxonomy.","Ground-truth adjudication of assumption changes, cross-driver interactions, novel actors, and protected-source evidence.","Comparative review-time, recall, reconstruction, fallback, and fully loaded maintenance-cost results.","Adopter confirmation that an independent audit path, protected bypass, and full-packet fallback can be authorized and staffed.","Privacy, records-retention, access-control, and labor/workflow review for any subsequent live deployment."],"research_disposition":"PROBLEM_PREVALENCE_STUDY","world_novelty_boundary":"This evaluation used a bounded eight-source open-web search. It found substantial adjacent research, government practice, and commercial tooling, but no exact published match for the complete reconstructive, versioned residual-plus-independent-audit architecture. That absence is not evidence of world novelty. Patentability, freedom to operate, patent and non-web literature, market size, and realized impact were not measured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":false,"progress_targets":["Obtain an authorized historical corpus and quantify redundancy, review demand, latency, and missed material signals before building a live workflow.","Preregister and execute the three-arm blinded comparison with explicit non-inferiority, safety, and net-effort thresholds.","Demonstrate reliable semantic reconstruction and no protected-source or underrepresented-perspective degradation.","Measure preparation, synchronization, auditing, reconciliation, fallback, and adjudication effort rather than review time alone.","Secure written authority for audit independence, protected bypasses, retention, access controls, and rollback before any live suppression."],"reason":"Web evidence establishes a credible domain problem, adopters, and close prior art, but cannot verify the candidate's setting-specific backlog or its incremental performance. The decisive evidence requires proprietary packet logs, human reviewer comparisons, semantic adjudication, and workflow testing. Under the controller rule, this must stop for empirical research rather than continue with more bounded web search; STOP recommendations are non-repairable."},"proposal_index":1}