{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"bounded_rivalry_governance__environmental_climate:P2:v0","cell_id":"bounded_rivalry_governance__environmental_climate","search_queries":["site:noaa.gov Subseasonal Climate Forecast Rodeo prize competition official","site:weatherbench2.readthedocs.io evaluation benchmark held out data reproducibility compute WeatherBench 2","WMO seasonal forecast verification standard multi-model ensemble operational seasonal prediction official","seasonal climate forecast model selection public leaderboard overfitting hindcast test data leakage research","NOAA Subseasonal Climate Forecast Rodeo official $ prize 2017","Subseasonal Climate Forecast Rodeo official challenge rules NOAA","Forecast Rodeo prize climate.gov competition improve forecasts official","Subseasonal Climate Forecast Rodeo paper results contest forecasting","WeatherBench benchmark overfitting test dataset repeated evaluation live leaderboard paper","WeatherBench 2 limitations benchmark evaluation period overfitting leaderboard paper","climate prediction benchmark data leakage rolling origin evaluation paper climate 2026","weather forecasting benchmark reproducibility compute cost leaderboard limitations primary research","seasonal forecast multi-model ensemble complementary models error correlation skill primary research","operational seasonal forecasting multi model ensemble weighting model diversity correlated errors study","WMO S2S multi model ensemble forecast combination diversity model dependence","seasonal climate forecast multi-model ensemble model independence performance weighting","site:challenge.gov America COMPETES Act prize authority federal agencies official toolkit","site:noaa.gov prize competition authority America COMPETES NOAA challenge","site:gao.gov federal prize competitions costs administration challenge authority","site:nist.gov AI Risk Management Framework testing evaluation verification validation official","\"Setting the Standard\" \"Data Preprocessing\" climate prediction authors 2026","\"Outcomes of the WMO Prize Challenge\" full text","doi BAMS-D-22-0046.1 WMO Prize Challenge full text","doi BAMS-D-24-0292.1 data preprocessing climate prediction","Challenge.gov federal agency toolkit prize competition legal authority 2025","site:challenge.gov/toolkit legal authority prize competitions America COMPETES","site:challenge.gov Federal agency challenge toolkit budget costs prize competition"],"sources":[{"source_id":"S1","title":"Setting the Standard: Recommended Practices for Data Preprocessing in Data-Driven Climate Prediction","publisher":"American Meteorological Society; record hosted by NSF NCAR/UCAR","url":"https://impacts.ucar.edu/en/publications/setting-the-standard-recommended-practices-for-data-preprocessing/","source_class":"PRIMARY_RESEARCH","publication_date":"2026-06","accessed_at":"2026-08-03","claims_supported":["Climate-prediction results can change materially with preprocessing choices, nonstationarity, spatial-temporal correlation, and treatment of extremes.","Transparent, reproducible preprocessing protocols are needed for trustworthy AI/ML climate prediction.","Climate data present small-sample and evolving-state conditions relevant to leakage-aware evaluation."]},{"source_id":"S2","title":"Sub-Seasonal Climate Forecast Rodeo Winners Announced","publisher":"National Integrated Drought Information System / NOAA","url":"https://www.drought.gov/news/sub-seasonal-climate-forecast-rodeo-winners-announced","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2019-03-21","accessed_at":"2026-08-03","claims_supported":["The Bureau of Reclamation and NOAA ran a year-long real-time forecasting competition against official NOAA benchmarks.","The competition had an $800,000 total prize pool and required real-time performance plus hindcast and documentation criteria.","The agencies stated that improved forecasts would help water managers prepare for drought and wet extremes.","At least five teams outperformed benchmarks, demonstrating feasible external participation and comparative testing."]},{"source_id":"S3","title":"Outcomes of the WMO Prize Challenge to Improve Subseasonal to Seasonal Predictions Using Artificial Intelligence","publisher":"American Meteorological Society; record hosted by Max Planck Society","url":"https://pure.mpg.de/pubman/faces/ViewItemOverviewPage.jsp?itemId=item_3488918","source_class":"PRIMARY_RESEARCH","publication_date":"2022-12","accessed_at":"2026-08-03","claims_supported":["WMO coordinated a 2021 prize challenge because of high demand and expectations for subseasonal-to-seasonal prediction.","The challenge prospectively scored weeks 3–4 and 5–6 temperature and precipitation forecasts.","Top submissions significantly outperformed a bias-corrected ECMWF operational reference, including through multimodel combination.","The challenge is a close precedent for a climate-forecast tournament."]},{"source_id":"S4","title":"WIPPS Seasonal Prediction","publisher":"World Meteorological Organization","url":"https://public.wmo.int/activities/wmo-integrated-processing-and-prediction-system-wipps/wipps-seasonal-prediction","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-11","accessed_at":"2026-08-03","claims_supported":["WMO already operates sub-seasonal and seasonal multi-model ensemble infrastructure through designated and lead centres.","Centres submit forecasts on fixed schedules and lead centres verify hindcast accuracy using standardized WMO metrics.","Fifteen centres contribute seasonal forecasts, showing that operational multi-model selection and verification are mature practices.","WMO Members, regional centres, and national meteorological services are identifiable adopters and users."]},{"source_id":"S5","title":"WeatherBench 2: A Benchmark for the Next Generation of Data-Driven Global Weather Models","publisher":"WeatherBench 2 research consortium / arXiv","url":"https://arxiv.org/abs/2308.15560","source_class":"PRIMARY_RESEARCH","publication_date":"2023-08-29","accessed_at":"2026-08-03","claims_supported":["An open-source weather-model evaluation framework with public data, baselines, operational metrics, and multiple headline scores is technically feasible.","The authors caution that a benchmark should not be reduced to a single leaderboard and identify operational-validity limitations, including ERA5 initialization and ground-truth caveats.","Training-resource requirements vary substantially among models, making resource accounting relevant but difficult.","Probabilistic evaluation, extremes, calibration, operational inputs, and multiple metrics are established evaluation concerns."]},{"source_id":"S6","title":"ESD Reviews: Model Dependence in Multi-Model Climate Ensembles: Weighting, Sub-selection and Out-of-Sample Testing","publisher":"Copernicus Publications / Earth System Dynamics","url":"https://esd.copernicus.org/articles/10/91/2019/","source_class":"PRIMARY_RESEARCH","publication_date":"2019-02-13","accessed_at":"2026-08-03","claims_supported":["Nominally different climate models can share code, data, parameterizations, and biases, weakening independence assumptions.","Performance-only sub-selection can perform worse than random selection when dependence is ignored.","There is no universally accepted measure of model independence.","Any weighting or sub-selection method must be tested out of sample in a way that emulates its intended application."]},{"source_id":"S7","title":"Prize Competitions","publisher":"U.S. General Services Administration","url":"https://www.gsa.gov/technology/government-it-initiatives/prize-competitions","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-03-26","accessed_at":"2026-08-03","claims_supported":["Federal agencies have statutory authority to run prize competitions under the America COMPETES framework, subject to agency-specific legal review.","Prize competitions can use contractors or facilitators for design, administration, marketing, and judging.","Prize challenges differ legally and procedurally from grants and procurement contracts.","Operational model appointment or acquisition may require authority beyond prize authority."]},{"source_id":"S8","title":"Artificial Intelligence Risk Management Framework (AI RMF 1.0)","publisher":"National Institute of Standards and Technology","url":"https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10","source_class":"OFFICIAL_GUIDANCE","publication_date":"2023-01-26","accessed_at":"2026-08-03","claims_supported":["AI risk management should cover design, development, deployment, use, and evaluation.","Trustworthiness requires attention to validity, reliability, safety, security, resilience, transparency, privacy, and harmful bias.","A shadow evaluation can apply recognized testing and risk-governance principles, but the framework does not itself authorize operational warning use."]}],"problem_evidence":{"support":"MODERATE","rationale":"Climate-prediction preprocessing sensitivity, nonstationarity, benchmark-to-operation gaps, unequal compute requirements, and model dependence are directly documented. Official competitions and operational ensembles show that comparative model selection matters to water and climate services. However, no located source measures how often an operational seasonal service actually chooses a brittle model because of public-leaderboard tuning, leakage, compute escalation, or incumbent lock-in; those candidate-specific prevalence claims remain inferred rather than verified.","source_ids":["S1","S2","S3","S5","S6"]},"stakeholder_evidence":{"support":"STRONG","rationale":"NOAA and the Bureau of Reclamation funded and operated an $800,000 real-time forecast competition, WMO operated another S2S challenge in response to expressed demand, and WMO currently coordinates operational multi-model services for member forecasting organizations. These identify credible adopters, authorizers, funders, competitors, and downstream water-management users.","source_ids":["S2","S3","S4","S7"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"NOAA–Bureau of Reclamation Sub-Seasonal Climate Forecast Rodeo","similarity":"A year-long real-time competition scored temperature and precipitation forecasts against official operational benchmarks, with hindcast and documentation requirements and a substantial prize purse.","remaining_difference":"The located account does not establish a common compute cap, reproducibility execution environment, complementary operational portfolio prize, independent appeal process, or time-limited operational appointment.","source_ids":["S2"]},{"name":"WMO AI Prize Challenge for S2S Prediction","similarity":"A prospective international contest compared AI forecasts for weeks 3–4 and 5–6 against an operational reference and rewarded methods including multimodel combinations.","remaining_difference":"It was an innovation challenge rather than a recurring governance process for allocating scarce operational ensemble membership, and no resource cap or challenger-rights package was located.","source_ids":["S3"]},{"name":"WMO WIPPS Multi-Model Ensemble Operations","similarity":"An established operational system aggregates models from multiple producing centres, uses fixed schedules and formats, and publishes standardized hindcast verification.","remaining_difference":"It is a coordinated operational standard and ensemble service, not a capped tournament with a scarce prize, fouls, appeals, portfolio ablations, or time-limited winner authority.","source_ids":["S4"]},{"name":"WeatherBench 2","similarity":"Provides open-source reproducible evaluation, public baselines, operationally grounded metrics, multiple headline scores, probabilistic evaluation, and explicit caveats.","remaining_difference":"It covers medium-range weather models and expressly avoids being a single-leaderboard challenge; it does not allocate operational authority, cap development resources, or govern recurring challengers.","source_ids":["S5"]},{"name":"Climate-ensemble independence weighting and sub-selection","similarity":"Existing research explicitly addresses correlated models, performance/independence weighting, sub-selection, and out-of-sample validation.","remaining_difference":"It supplies statistical methods and cautions rather than a complete institutional tournament with eligibility, sanctions, audit, appeals, and operational-role governance.","source_ids":["S6"]}],"distinctive_claim_remaining":"For a forecasting service with a genuinely scarce operational ensemble allocation, the combined addition of sealed rolling evaluation, audited executable reproducibility, a predefined resource boundary, complementary-error portfolio selection, time-limited authority, independent process review, and realized-season recalibration will select a more stable and decision-relevant model set than either a persistent public single-score winner or simpler established benchmark and equal-weight ensemble practices. This incremental claim is falsifiable but untested.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Open-source multi-metric evaluation, prospective competitions, operational multi-model exchanges, standardized verification, federal prize authority, contractor-supported administration, and recognized AI risk-governance practices all exist. A non-operational replay is therefore technically and procedurally credible. Feasibility is not yet verified for reconstructing issue-time data, comparing pretrained and proprietary resources, protecting code and data, obtaining participant consent, resolving common-control questions, setting user-specific loss weights, or integrating results into an operational appointment. Prize authority also does not automatically confer warning-issuance or procurement authority.","source_ids":["S2","S3","S4","S5","S7","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Subseasonal forecasts affect drought and water-management preparation, and official organizations describe strong demand. Better robustness could matter, although realized decision impact is unmeasured.","source_ids":["S2","S3","S4"]},"stakeholder_pull":{"score":5,"rationale":"Multiple public agencies and WMO have already sponsored closely related competitions and operate multi-model forecast services.","source_ids":["S2","S3","S4","S7"]},"incremental_advantage":{"score":2,"rationale":"The package adds coherent governance, but prospective contests, reproducible multi-metric benchmarks, operational multi-model ensembles, and dependence-aware selection already cover much of the mechanism. No comparative effect estimate exists.","source_ids":["S2","S3","S4","S5","S6"]},"distinctiveness_plausibility":{"score":2,"rationale":"The full conjunction may be uncommon, but its major elements have substantial prior-art collision; distinctiveness rests on the integrated resource-capped, appealable, time-limited operational allocation and must be demonstrated contrastively.","source_ids":["S2","S3","S4","S5","S6"]},"technical_implementability":{"score":4,"rationale":"Comparable competitions, open evaluation code, standardized verification, and multi-model operations exist. The main technical uncertainties are issue-time reconstruction, portable execution, resource accounting, and stable portfolio optimization.","source_ids":["S2","S3","S4","S5"]},"adoption_authority_feasibility":{"score":4,"rationale":"Federal services can authorize research challenges and shadow evaluation, and established WMO centres can run verification. Operational appointment, procurement, data-rights, and warning authority would require separate sponsor-specific review.","source_ids":["S2","S4","S7","S8"]},"evidence_readiness":{"score":4,"rationale":"A preregistered archived replay with explicit comparators can measure reproducibility, ranking stability, correlation, compute association, and burden without changing forecasts. It still requires participating pipelines or proprietary service data.","source_ids":["S1","S2","S5","S6"]},"safety_net_benefit":{"score":4,"rationale":"A shadow-only replay leaves current forecasts unchanged and can be voided after leakage, conflict, or execution failures. Data, IP, confidentiality, and cybersecurity exposure still require controls.","source_ids":["S7","S8"]},"scalability":{"score":3,"rationale":"Shared standards and common evaluation infrastructure support repetition, but secure custody, audits, appeals, compute accounting, and specialized model execution impose continuing administrative and technical costs.","source_ids":["S4","S5","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"One preregistered archived shadow replay for one variable and lead-time pair, using 4–8 existing submissions, one sealed period, common evaluation code, limited independent custody, and no prize or operational integration.","confidence":"LOW","assumptions":["Existing archived forecasts, observations, and executable pipelines are available without new licensing fees.","Approximately 2–4 staff-equivalent months cover protocol, custody, execution, scoring, and reporting.","Compute is bounded to a modest shared cloud or agency environment.","Participant model-development costs and foregone commercial value are excluded."],"source_ids":["S2","S5"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Build reusable secure submission and execution infrastructure, rulebook, data-custody process, audit trail, resource accounting, conflict review, appeal workflow, and one multi-team pre-operational tournament.","confidence":"LOW","assumptions":["The service reuses open evaluation components and existing observations.","Costs include legal, security, scientific, and administrative design but not a major prize purse.","Specialized proprietary hardware and model redevelopment are excluded.","The $800,000 historical NOAA competition provides an order-of-magnitude anchor, not a direct cost estimate."],"source_ids":["S2","S5","S7","S8"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Launch a recurring externally accessible tournament with a prize or participation support, independent test custody, reproducibility audits, security review, adjudication, ensemble-integration testing, and operational change-control preparation.","confidence":"MODERATE","assumptions":["A prize or participation pool is roughly $0.5–1.5 million, comparable in scale to the historical $800,000 Rodeo after allowing for 2026 scope.","A multidisciplinary sponsor team and external administrator operate for one annual cycle.","Operational forecast production by participating teams remains outside the sponsor cost.","No new supercomputing facility is purchased."],"source_ids":["S2","S4","S7","S8"]},"annual_recurring":{"band_2026_usd":"1M_TO_5M","scope":"Annual or seasonal rounds, secure test refresh, participant support, common compute, finalist audits, appeals, post-season review, ensemble maintenance, and challenger access.","confidence":"LOW","assumptions":["The service retains reusable platform and standards from startup.","Each cycle includes several serious finalists and independent scientific and integrity review.","Costs exclude the full underlying research budgets of model developers.","Major incident response, litigation, or procurement transition is excluded."],"source_ids":["S2","S4","S7"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"External research and official practice support the material risks of preprocessing sensitivity, non-operational benchmark conditions, compute disparities, and correlated models, although the prevalence of strategic gaming in any particular service remains unknown.","source_ids":["S1","S2","S5","S6"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"NOAA, the Bureau of Reclamation, WMO, designated climate centres, and national meteorological services are identifiable sponsors or adopters of closely related forecast competitions and multi-model verification.","source_ids":["S2","S3","S4","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The integrated package can be compared prospectively with a public single-score winner, equal-weight eligible ensemble, and simpler sealed benchmark using preregistered stability, calibration, correlation, compute, and burden outcomes.","source_ids":["S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A one-variable, one-lead, archived shadow replay with frozen rules, explicit comparators, and no operational consequence is bounded and reversible.","source_ids":["S1","S2","S5","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"For the shadow replay only, existing competition and research authority is credible and no warning or allocation change is required. Operational adoption remains outside this gate and would need data-rights, cybersecurity, procurement, and warning-authority approval.","source_ids":["S7","S8"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The bands state sponsor-side scope and exclusions and are anchored by an $800,000 prior federal forecast competition plus documented availability of reusable evaluation infrastructure and challenge-administration support. Precision remains low because staffing, compute, prizes, and proprietary access are service-specific.","source_ids":["S2","S5","S7"]}},"next_evidence_step":"With one forecasting-service partner, preregister a non-operational replay for one variable such as two-metre temperature, one lead such as weeks 3–4, one public development window, and at least two later rolling sealed windows. Admit 4–8 de-identified existing models under a documented issue-time data rule, executable-reproduction test, common inference envelope, submission cap, and disclosed treatment of pretraining and proprietary inputs. Compare: (A) public-development single-score winner, (B) sealed single-score winner, (C) all eligible models with equal weights, and (D) the proposed minimum-floor complementary portfolio. Include ablations for sealed evaluation plus audit without a resource cap and for portfolio selection without independence terms. Before opening sealed outcomes, set thresholds for score reproduction, ranking stability, calibration, worst-region and rare-regime error, finalist error correlation, compute–rank association, user interpretability, evaluator disagreement, exclusions, and staff hours. Falsify the incremental claim if the proposed portfolio does not improve prespecified sealed-period stability or decision-weighted loss over both B and C, if gains disappear under reasonable score weights, if the compute boundary is materially circumvented, if reproducibility excludes a prespecified unacceptable share of otherwise credible models, or if administrative delay and burden exceed the predefined value threshold. Do not change an operational forecast or represent the replay as future effectiveness.","blocking_evidence":["No service-level measurement yet shows the prevalence or magnitude of leaderboard tuning, leakage, compute-driven rank, or incumbent lock-in.","No controlled comparison shows that the full governance package outperforms simpler sealed testing, audit, certification, or equal-weight ensembles.","Participant willingness and ability to provide reproducible executable pipelines, provenance, and resource disclosures are unknown.","Issue-time data availability, observation revisions, sealed-test custody, and comparable execution have not been demonstrated for a named service.","The stability and user value of complementary portfolio weights across regions, regimes, metrics, and asymmetric-loss assumptions are unknown.","Sponsor-specific intellectual-property, confidentiality, FOIA or records, cybersecurity, procurement, conflict, and operational-warning authorities remain unverified.","The four cost bands lack adopter quotations, workload measurements, and service-specific compute inventories."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This bounded search found substantial collision with real-time NOAA and WMO forecast challenges, WeatherBench-style reproducible multi-metric evaluation, established WMO operational multi-model ensembles, and climate-model dependence weighting. It did not establish an exact match for the complete resource-capped, appealable, time-limited operational-allocation package. World novelty, patentability, freedom to operate, market size, and realized impact were not measured, and absence from these eight sources is not evidence of novelty.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":false,"progress_targets":["Secure a forecasting-service partner and named decision owner for the shadow replay.","Execute the preregistered replay against single-score, equal-weight, and simpler sealed-audit comparators.","Report component ablations so any incremental benefit can be attributed to the resource cap, reproducibility audit, or complementary portfolio rule.","Measure ranking stability, calibration, rare-regime and regional loss, error dependence, compute association, exclusions, user interpretation, and administrative burden.","Complete sponsor-specific data-rights, IP, confidentiality, cybersecurity, procurement, conflict, appeal, and operational-authority review.","Replace broad resource bands with service workload estimates, compute inventories, and at least two administrator or vendor quotations."],"reason":"Bounded web research establishes a meaningful forecasting need, credible public adopters, feasible evaluation infrastructure, legal pathways for a shadow challenge, and substantial prior art. It cannot determine whether the integrated package improves selection over simpler established practices. That question requires participating model pipelines, controlled sealed execution, service data, user judgments, and operational workflow measurements rather than additional bounded web search."},"proposal_index":2}