{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","cell_id":"bounded_rivalry_governance__environmental_climate","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_id":"brg-climate-forecast-arena-002","proposal_index":2,"version":0,"title":"Governed Tournament for Operational Climate-Forecast Models","problem":"Climate-model teams compete for the scarce opportunity to supply an operational seasonal drought or heat-risk forecast. If selection rests on a public hindcast leaderboard and one aggregate accuracy score, teams can improve rank by tuning to known evaluation years, exploiting data leakage, concentrating compute on the benchmark, or converging on the same high-scoring architecture. The contest can consequently select a brittle or correlated winner rather than a reproducible set of models that provides useful forecasts under future conditions.","actors":["Public climate or environmental forecasting service","University, public-sector, nonprofit, and commercial forecast-model teams","Independent benchmark administrator","Climate-data custodians","Conflict-screened scientific evaluation panel","Operational forecasters and emergency-planning users","Communities and sectors affected by drought or heat advisories","Independent appeal and integrity reviewer"],"observable_state":"Teams seek a limited operational forecast slot, ensemble weight, computing allocation, or validation contract. Public records from the contest show leaderboard rank, score changes across evaluation windows, computing use, model lineage, submission timing, reproducibility results, error correlations, and audit findings. Warning signs include repeated tuning to disclosed test periods, sharp performance loss on newly sealed periods, rising compute without stable out-of-sample gains, unverifiable pipelines, highly correlated finalist errors, or an incumbent retaining the operational slot despite reproducible challenger performance.","consequence":"Operational selection can become coupled to benchmark exploitation, compute resources, and incumbent access rather than forecast robustness. A single winning pipeline can create correlated blind spots, reduce methodological diversity, impede later challengers, and expose forecast users to poorly characterized failure under climate regimes not represented by the public benchmark.","affected_objective":"Select and continually reassess a reproducible, resource-bounded portfolio of climate-forecast models whose combined performance is robust across locations, lead times, climate regimes, and decision-relevant failure modes.","intervention":"Create a recurring, audited forecast tournament for a fixed operational slot or ensemble allocation. Publish and freeze a rulebook covering eligibility, permitted data, forecast issue dates, compute accounting, reproducibility, conflicts, scoring families, fouls, appeals, and the duration of any selection. Separate a public development set from sealed rolling test periods and undisclosed stress cases. Require containerized or otherwise reproducible submissions and audit finalists before selection. Cap contest compute and standardized tuning opportunities so rank cannot be purchased through unbounded search. Score calibration, discrimination, spatial and temporal robustness, performance under rare regimes, operational timeliness, and predefined asymmetric errors rather than one pooled accuracy number. Select a complementary ensemble whose members add marginal skill or reduce correlated failure, subject to a minimum individual-performance floor. Prohibit test-data leakage, post-deadline revision, identity concealment, fabricated provenance, interference with rivals, and undisclosed common control. Award only a time-limited operational role, retain common interfaces and evaluation infrastructure outside any winner's control, reopen the arena on a stated cadence, and use realized-season review to revise or retire the tournament.","structural_mapping":[{"archetype_element":"Rivalry purpose statement","domain_realization":"Use model rivalry to reveal reproducible forecast skill and complementary failure patterns for operational climate decisions."},{"archetype_element":"Scarce prize or selection constraint","domain_realization":"A limited operational forecast slot, ensemble weight, validation contract, and associated computing allocation for a defined service period."},{"archetype_element":"Competitor eligibility boundary","domain_realization":"Teams that disclose control relationships, document data rights and model lineage, satisfy a minimum reproducibility check, and accept common issue-time and compute-accounting rules."},{"archetype_element":"Contest arena boundary","domain_realization":"Permitted modeling, feature design, and tuning use only declared development data and capped resources; leakage, post-deadline changes, fabricated provenance, rival interference, and concealed coordinated entries are fouls."},{"archetype_element":"Performance metric and scoring basis","domain_realization":"A preregistered scorecard covering calibration, discrimination, robustness by region and regime, timeliness, asymmetric decision errors, reproducibility, and marginal ensemble contribution on sealed tests."},{"archetype_element":"Fair process and due process layer","domain_realization":"Rules and test custody are separated from competitors; finalists receive auditable score explanations and may challenge calculation, eligibility, or process errors before an independent reviewer."},{"archetype_element":"Anti-sabotage and anti-collusion guardrail","domain_realization":"Submission provenance, access logs, common-control disclosures, code hashes, and a graduated foul schedule protect sealed tests and distinguish independent entries from coordinated duplicates."},{"archetype_element":"Externality and spillover boundary","domain_realization":"The scorecard includes costly false reassurance, excessive false alarms, regional performance disparities, operational latency, and failure to communicate forecast uncertainty to downstream users."},{"archetype_element":"Escalation and arms-race damper","domain_realization":"A measured compute ceiling, fixed submission count, common execution environment, and audit of external preprocessing constrain mutually escalating search expenditure."},{"archetype_element":"Winner power and lock-in review","domain_realization":"Selection is time-limited; evaluation data custody, interfaces, and operational archives remain with the forecasting service; qualified challengers receive recurring access."},{"archetype_element":"Learning and recalibration loop","domain_realization":"After each forecast season, predicted probabilities are compared with observed outcomes and user-relevant failures, informing the next round's metrics, stress cases, caps, ensemble size, or suspension."}],"mechanism_mapping":[{"mechanism_slug":"contest_rulebook","role":"Freezes eligibility, permitted data, resource accounting, scoring, tie-breaks, foul definitions, selection duration, and appeals before sealed evaluation begins.","counterfactual_removal":"Without the rulebook, test boundaries or scoring weights could shift after organizers see competitors, and teams could not reliably distinguish legitimate modeling from prohibited benchmark exploitation."},{"mechanism_slug":"ranked_leaderboard_with_audit","role":"Provides comparable development standings while withholding part of the evaluation and requiring reproducibility and data-provenance audits before finalists receive operational status.","counterfactual_removal":"Without finalist audit and held-out evaluation, the visible ranking could primarily reward leakage, fragile tuning, or irreproducible submissions."},{"mechanism_slug":"spending_cap_or_resource_cap","role":"Caps measured compute, submission count, and standardized tuning opportunities within a defined accounting boundary.","counterfactual_removal":"Without a cap, teams could enter a compute escalation in which bankroll and search volume dominate methodological quality while every rival incurs higher costs to maintain rank."},{"mechanism_slug":"multiple_award_or_portfolio_selection","role":"Allocates operational ensemble weight to several models according to complementary skill and error structure rather than automatically selecting the top isolated score.","counterfactual_removal":"Without portfolio selection, a marginally higher-ranked but highly correlated model could displace a different model that protects the operational system against a distinct failure mode."},{"mechanism_slug":"sabotage_or_foul_penalty_schedule","role":"Predefines graduated consequences for test leakage, provenance fabrication, post-deadline alteration, compute-accounting evasion, interference, and concealed common control.","counterfactual_removal":"Without known and consistently applied consequences, prohibited conduct may remain a viable route to rank whenever its detection risk is uncertain."},{"mechanism_slug":"challenger_access_window","role":"Limits the operational appointment and schedules recurring opportunities for qualified challengers to contest ensemble membership or weight.","counterfactual_removal":"Without reopening, an early winner could turn operational access, archived feedback, and interface control into a durable advantage unrelated to continuing forecast performance."},{"mechanism_slug":"post_contest_impact_review","role":"Compares issued probabilities with realized environmental conditions, examines region-specific and decision-relevant failures, reviews concentration, and mandates changes before the next round.","counterfactual_removal":"Without ex-post review, a metric that selects poorly under new climate regimes could remain in use because its weakness was invisible on the original benchmark."}],"causal_chain":["A forecasting service has fewer operational slots and computing resources than eligible model teams, creating a genuine scarce-prize rivalry.","A public benchmark gives teams a common target, but repeated exposure permits tuning to its particular years, regions, variables, and scoring function.","Unbounded compute and submissions reward larger searches, while a single-score ranking encourages methodological convergence and hides correlated failure.","A frozen rulebook, sealed rolling tests, reproducibility audits, and enforceable fouls restrict winning strategies to declared data and verifiable forecasting performance.","Compute and submission caps reduce the value of escalating search expenditure and make comparison occur within a common resource boundary.","Multi-dimensional scoring tests whether performance persists across regimes and decision-relevant errors rather than only in the pooled headline metric.","Portfolio selection rewards a model for marginal contribution and error complementarity, preserving useful rivalry while avoiding automatic winner-take-all dependence.","A time-limited appointment and recurring challenger window prevent one victory from granting permanent control of evaluation infrastructure or operational access.","Review against realized seasons tests whether the tournament selected operational value; observed failures trigger rule revision, reweighting, suspension, or retirement."],"baseline":"A forecasting service compares teams on a published historical hindcast dataset, ranks them using one aggregate score, and grants the leading model or incumbent pipeline an operational role. Compute use, repeated tuning, model dependence, provenance, reproducibility, regional failure, selection duration, and post-season recalibration receive limited or separate treatment.","nearest_rivals":["Expert appointment without a contest: a scientific panel chooses an operational model from professional judgment and published evidence, avoiding leaderboard gaming but sacrificing a common, bounded comparative test.","Equal-weight open ensemble: every model meeting a certification floor receives the same weight, preserving diversity but providing little rivalry-driven discipline over marginal contribution, redundancy, or resource use.","Threshold certification: models either satisfy fixed reliability requirements or fail, which is appropriate when the operational capacity is not scarce but does not resolve selection when only a limited ensemble can be maintained.","Ordinary benchmark competition: ranks reproducible models on held-out data, but lacks the combined resource cap, complementary portfolio prize, time-limited operational authority, downstream harm criteria, and recurring challenger path.","Conventional forecasting-service procurement: evaluates vendors through an RFP and contract terms, but may select proposal quality or a single delivery package without maintaining an adaptive scientific tournament across realized forecast seasons."],"remaining_contrastive_claim":"The proposal's limited contrastive claim is that, when an operational forecasting service wants rivalry but can support only a bounded model set, combining sealed evaluation, audited reproducibility, resource caps, complementary portfolio selection, time-limited authority, due process, and realized-season recalibration keeps the route to selection more closely aligned with robust forecast contribution than a persistent public single-score race. Its practical performance is untested.","authority_safety":{"decision_authority":"The forecasting-service director may authorize a non-operational shadow tournament using data the service is already permitted to evaluate. Existing scientific-governance, data-custody, procurement, warning-issuance, and appeal authorities retain their current responsibilities.","authorized_first_step":"Run a preregistered shadow replay using archived issue-date data: release an earlier development window, keep later periods sealed with an independent custodian, execute de-identified reproducible submissions inside a common compute envelope, and compare single-winner and complementary-portfolio selections without changing any operational forecast.","excluded_actions":["Changing an operational drought, heat, or climate advisory","Granting a model direct authority to issue public warnings or trigger resource allocation","Publishing team identities or rankings from the shadow tournament without prior authorization and disclosure rules","Using confidential, proprietary, personal, tribal, or security-sensitive data beyond existing permission","Treating a foul-screen anomaly as proof or imposing a sanction without investigation and appeal","Allowing competitors or evaluators to access the sealed test period before submission closure","Automatically excluding an incumbent or appointing a challenger from the shadow result","Representing retrospective hindcast performance as demonstrated future operational effectiveness"],"halt_rollback":"Stop the shadow tournament if sealed data are exposed, issue-time data availability cannot be reconstructed, observations are too inconsistent for the preregistered score, common execution materially changes model behavior, compute cannot be comparably accounted, conflicts cannot be managed, or portfolio selection is unstable across reasonable specifications. Void the affected ranking, preserve an access-controlled incident and methods record, release no operational conclusion, and leave the existing forecast process unchanged."},"negative_tests":{"strongest_counterevidence":"The strongest counterevidence would show that the baseline single-score selection is stable across genuinely unseen periods and climate regimes, finalists are reproducible, compute and submission volume do not materially determine rank, model errors are not problematically correlated, challengers face no durable access disadvantage, and an ensemble selected for complementarity supplies no additional decision-relevant information.","problem_falsifier":"The inferred problem is falsified in the tested service if there is no scarce operational slot or resource, teams do not respond strategically to the benchmark, test leakage and repeated tuning are absent, resource escalation does not occur, rankings generalize across sealed regimes, and winner control creates no barrier to later entry.","intervention_falsifier":"The intervention is falsified as a useful selection design if blinded reruns cannot reproduce scores; compute boundaries are routinely circumvented; sealed results are no more stable than public-benchmark rankings; complementary selection produces no distinguishable protection against correlated failure; operational users cannot interpret the scorecard; or tournament delay, complexity, and exclusion exceed the decision value produced.","risks":["The sealed test may still be too similar to the development data to reveal regime-specific brittleness.","A compute cap may favor incumbents that possess pretrained models, proprietary data, or reusable infrastructure outside the accounting boundary.","Strict reproducibility requirements may exclude scientifically valuable models dependent on specialized systems or licensed inputs.","Multi-metric scoring can conceal subjective value judgments in weights and thresholds.","Portfolio optimization can appear opaque and weaken participant acceptance.","Selecting for error diversity may preserve a weak model whose errors are merely unusual rather than useful.","Repeated challenger rounds may discourage long-term operational investment or create disruptive model churn.","Audit access to code and data may create intellectual-property, confidentiality, or cybersecurity risks.","Observed environmental outcomes may be revised or spatially incomplete, complicating post-season evaluation.","False-alarm and missed-event costs differ among users, so one downstream harm weighting may not represent all affected groups.","Teams may coordinate through shared code, data, or personnel without intending collusion, generating misleading integrity flags.","A tournament may shift scientific effort toward scored forecast variables and away from unscored but important climate-system understanding."]},"next_evidence_step":"Before viewing sealed outcomes, specify one forecast variable, lead time, operational issue schedule, public development period, sealed archival period, compute boundary, submission limit, reproducibility protocol, scorecard, portfolio rule, minimum performance floor, and sensitivity analysis. Run the non-operational replay with de-identified submissions. Report score reproducibility, ranking stability across periods and regions, association between compute and rank, finalist error correlations, differences between single-winner and portfolio selection, evaluator disagreement, detected leakage routes, operational interpretability, and administrative burden. Do not alter forecasts or infer future effect from the replay.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Proposal 1 governed competition among basin jurisdictions for scarce physical adaptation capital, addressing locally optimized projects that could transfer flood risk and using hydrologically coupled portfolio funding, implementation liability, and grant-process safeguards. This proposal instead governs competition among scientific model teams for scarce operational forecast status. Its problem is benchmark overfitting, compute escalation, correlated model failure, and evaluator lock-in; its intervention is a sealed, reproducible, resource-capped forecast tournament with complementary ensemble selection and recurring model challenges. It is independently adoptable by a forecasting service without changing adaptation-grant allocation or constructing basin projects, and its causal path runs from benchmark visibility and computational rivalry through bounded evaluation to operational model diversity rather than from jurisdictional project rivalry through spillover-aware capital allocation.","revision_record":{"parent_version":null,"progress_targets_addressed":["Second independently adoptable candidate for the same archetype-domain cell","Material separation from basin adaptation-funding governance","Complete mapping of rivalry, prize, arena, scoring, safeguards, authority, falsifiers, and bounded evidence"],"conceptual_changes":["Initial version; no parent proposal","Applied bounded rivalry to operational climate-forecast model selection rather than physical adaptation-project funding","Made correlated forecast failure and benchmark adaptation central to the problem structure"],"operational_changes":["Specified sealed rolling tests, reproducible execution, compute and submission caps, complementary ensemble selection, time-limited appointments, and challenger rounds","Restricted the first step to a non-operational archived-data replay"],"evidence_changes":["Prior art remains unsearched","Defined a preregistered shadow comparison using archived issue-date data and sealed later outcomes"],"claim_changes":["Made no claim of novelty, prevalence, demand, effect size, or demonstrated forecast improvement","Limited the claim to a testable governance contrast with public single-score selection"]}}