{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","cell_id":"bounded_rivalry_governance__aviation_aeronautics","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_id":"brg-aa-02-envelope-robustness-downselect","proposal_index":2,"version":0,"title":"Envelope-Robustness Downselect Arena for Flight-Control Candidates","problem":"An aircraft development program must choose which competing flight-control candidates receive scarce hardware-in-the-loop integration and flight-test resources. When teams know the scored simulation cases, they can improve rank by tuning narrowly to those cases, exploiting simulator artifacts, withholding reproducibility information, or consuming escalating compute and test-labor resources. The contest can therefore select the team best at optimizing the visible evaluation environment rather than the controller that behaves acceptably across relevant but unseen conditions.","actors":["Aircraft program sponsor and system integrator","Competing internal or supplier flight-control teams","Independent simulation and hardware-in-the-loop laboratory","Flight-test engineers and test pilots","Software assurance and configuration-management personnel","Safety review board","Human-factors specialists","Aircraft maintainers and downstream operators","Independent judges and appeal reviewer","Applicable certification or operational-approval authority"],"observable_state":"For each candidate and test round: public-case score; hidden-case score; rank movement between public and hidden cases; safety-envelope violations; recovery success by disturbance class; pilot-intervention demand; control activity and actuator saturation; compute and engineering resources consumed; unreproduced results; configuration changes after deadlines; simulator-specific dependencies; audit findings; common code or result anomalies; appeals; and subsequent performance in independently constructed validation cases. The diagnostic signature is a contest in which visible-case rank, resource expenditure, or harness-specific behavior predicts advancement more strongly than consistent behavior across independently generated conditions.","consequence":"A benchmark-optimized winner can consume scarce integration resources before hidden brittleness is discovered, while capable but less benchmark-specialized alternatives are eliminated. Public scores can also induce a resource arms race, narrow design diversity, and shift safety, maintainability, pilot-workload, or validation costs beyond the scored arena.","affected_objective":"Use rivalry to discover and compare robust flight-control approaches while keeping advancement coupled to safe envelope behavior, reproducibility, maintainability, human interaction, and performance across conditions not available for contestant tuning.","intervention":"Establish a staged envelope-robustness downselect arena whose scarce prizes are two independently funded hardware-in-the-loop integration slots followed by at most one bounded flight-test recommendation. Before submissions, freeze a rulebook defining eligible configurations, allowable training and simulation resources, configuration-control deadlines, safety floors, scored objectives, audit evidence, fouls, appeals, and the authority of the independent safety board. Teams may iterate on a disclosed development suite, but final ranking uses an independently held scenario generator with undisclosed seeds and disturbance combinations. Every candidate must first pass noncompensable safety-envelope gates; among passers, scoring combines hidden-case robustness, pilot-intervention demand, control smoothness, reproducibility, and declared integration burden. A resource cap limits compute, simulator hours, and sponsor-provided engineering support. The first downselect selects a complementary two-candidate portfolio rather than a single leaderboard winner, preserving materially different control approaches for hardware-in-the-loop validation. Leaders are independently reproduced and audited before advancement. Submission similarities and cross-round behavior are screened uniformly for possible coordination or shared prohibited artifacts, but any flag only triggers inquiry. Hardware-in-the-loop outcomes, maintainer findings, and pilot assessments then determine whether either candidate merits a flight-test recommendation; no contest result itself authorizes flight. A later challenger window permits a qualified revised or external candidate to contest the incumbent before architectural lock-in.","structural_mapping":[{"archetype_element":"Rivalry purpose statement","domain_realization":"Use controlled competition to reveal which flight-control approaches remain acceptable across uncertain conditions and deserve scarce integration effort."},{"archetype_element":"Scarce prize or selection constraint","domain_realization":"Two funded hardware-in-the-loop integration slots and, after independent validation, at most one recommendation for a bounded flight-test program."},{"archetype_element":"Competitor eligibility boundary","domain_realization":"Only configuration-controlled candidates satisfying interface, documentation, provenance, cybersecurity, and preliminary safety-evidence requirements may enter scored evaluation."},{"archetype_element":"Contest arena boundary","domain_realization":"Permitted moves include controller design, training on the disclosed suite, documented parameter selection, and timely submission. Test leakage, post-deadline modification, simulator tampering, unreported external compute, interference with rivals, fabricated evidence, and coordinated scoring strategies are prohibited."},{"archetype_element":"Performance metric and scoring basis","domain_realization":"Noncompensable safety gates precede a composite score over hidden-condition robustness, intervention demand, control activity, reproducibility, and integration burden; public-suite performance alone cannot win."},{"archetype_element":"Fair process and due process layer","domain_realization":"Eligibility, caps, scenario-generation custody, scoring, audit sampling, tie-breaks, judge conflicts, debriefs, and a time-boxed independent appeal procedure are fixed before final submissions."},{"archetype_element":"Anti-sabotage and anti-collusion guardrail","domain_realization":"Configuration custody, uniform anomaly screens, access logs, separated test administration, and graduated penalties address interference, leakage, fabricated evidence, and prohibited coordination without treating statistical similarity as a verdict."},{"archetype_element":"Externality and spillover boundary","domain_realization":"Safety, pilot workload, actuator use, maintainability, reproducibility, and integration labor are brought into gates or scores rather than left for flight-test crews and operators to absorb later."},{"archetype_element":"Escalation and arms-race damper","domain_realization":"Audited ceilings on compute, simulator hours, sponsor engineering support, and final submissions prevent rank from becoming primarily a contest of resource expenditure."},{"archetype_element":"Prize decomposition or multiple-winner design","domain_realization":"Two complementary candidates advance to hardware-in-the-loop testing so an uncertain simulation ranking does not prematurely collapse the design space."},{"archetype_element":"Winner power and lock-in review","domain_realization":"Advancement grants validation access rather than permanent architectural control; common interfaces, retained test artifacts, and a scheduled challenger window preserve future contestability."},{"archetype_element":"Learning and recalibration loop","domain_realization":"The post-contest review compares simulation judgments with hardware-in-the-loop and human evaluations, then revises scenario families, weights, caps, gates, or the use of rivalry itself."}],"mechanism_mapping":[{"mechanism_slug":"contest_rulebook","role":"Binds entrants, laboratories, and judges to advance eligibility, legal moves, gates, scoring, evidence, configuration-control, and appeal terms.","counterfactual_removal":"Without a frozen rulebook, test administrators could alter cases or weights after seeing candidate performance, and teams could not distinguish legitimate iteration from prohibited adaptation."},{"mechanism_slug":"ranked_leaderboard_with_audit","role":"Provides comparable standings on disclosed development cases while requiring independent reproduction and hidden-case audit before any leader advances.","counterfactual_removal":"Without auditing and held-out evaluation, the visible leaderboard would reward simulator-specific optimization or irreproducible claims without testing whether the rank survives independent scrutiny."},{"mechanism_slug":"spending_cap_or_resource_cap","role":"Caps compute, simulator time, sponsor engineering support, and final submissions within a defined evaluation period.","counterfactual_removal":"Without resource caps, teams could improve rank through escalating expenditure that reveals organizational bankroll more than controller quality and pressures rivals to imitate."},{"mechanism_slug":"multiple_award_or_portfolio_selection","role":"Awards the first-stage integration resource to two complementary control approaches rather than automatically advancing the top two versions of the same approach.","counterfactual_removal":"Without portfolio selection, small score differences could eliminate design diversity before higher-fidelity testing reveals which assumptions matter."},{"mechanism_slug":"sabotage_or_foul_penalty_schedule","role":"Defines graduated consequences for test leakage, configuration substitution, fabricated evidence, undisclosed resources, interference, and repeated procedural violations.","counterfactual_removal":"Without precommitted penalties, prohibited tactics could remain useful winning strategies whenever their expected competitive benefit exceeds the uncertain consequence."},{"mechanism_slug":"anti_collusion_monitoring","role":"Applies fixed cross-submission and cross-round screens for suspiciously shared prohibited artifacts, coordinated abstention, or result patterns and refers anomalies to independent inquiry.","counterfactual_removal":"Without monitoring, nominally separate teams could coordinate submissions or share restricted test knowledge while preserving the appearance of independent rivalry."},{"mechanism_slug":"challenger_access_window","role":"Allows a qualified revised or external controller to contest the selected architecture at a scheduled pre-lock-in review point.","counterfactual_removal":"Without a credible reopening, the first simulation winner could become the permanent default because later switching costs, interfaces, and accumulated data favor it."},{"mechanism_slug":"post_contest_impact_review","role":"Tests whether simulation advancement predicted hardware-in-the-loop behavior, pilot interaction, integration burden, and maintainability, and feeds discrepancies into the next rule set.","counterfactual_removal":"Without this review, the program could repeatedly use scenario families or weights that clear contests cleanly but select poorly for downstream validation."}],"causal_chain":["Scarce integration and flight-test capacity forces a development program to downselect among rival flight-control candidates.","A fully visible benchmark gives every team an incentive to specialize to scored cases, exploit harness artifacts, conceal reproducibility weaknesses, and escalate resources.","A frozen arena separates disclosed development cases from independently held final cases and defines prohibited routes to advantage.","Noncompensable safety gates prevent a high aggregate score from offsetting unacceptable envelope behavior.","Resource caps preserve useful design pressure while limiting compute and test-labor escalation.","Audited hidden-case scoring makes cross-condition behavior, reproducibility, human interaction, and integration burden more reliable routes to advancement than visible-case tuning.","Portfolio selection preserves two complementary approaches through higher-fidelity testing, reducing premature concentration around a noisy simulation winner.","Independent hardware-in-the-loop, pilot, and maintainer evidence tests whether the contest selected what its purpose required.","A challenger window prevents initial advancement from becoming permanent architectural ownership.","Post-contest comparison recalibrates or retires the arena if strategic adaptation again separates winning from robust downstream contribution."],"baseline":"A prespecified conventional downselect in which all teams receive the complete scored simulation suite, submit final results to a review panel, and the highest aggregate scorer satisfying a basic safety threshold receives the sole integration slot. The comparison should use the same candidate builds, computing envelope, scenario families, and safety constraints.","nearest_rivals":["A conventional tender or paper design review that selects from proposed architecture, experience, cost, and safety plans before comparative executable testing; it reduces contest infrastructure but relies heavily on forecast claims and panel judgment.","A fully public audited leaderboard; it improves reproducibility but leaves all scored conditions available for iterative tuning and may intensify rank-focused resource expenditure.","Independent pass-fail assurance for every candidate without ranking; it protects safety floors but does not decide which candidate receives scarce integration resources when several pass.","A centralized expert panel comparing unranked simulation dossiers; it can integrate qualitative evidence but does not use bounded rivalry to expose performance differences under common conditions.","A single-winner hidden-test challenge; it limits direct overfitting but concentrates the outcome on one potentially noisy evaluation and can create early architectural lock-in.","Parallel funding of all eligible candidates through hardware-in-the-loop testing; it preserves diversity but does not address the binding scarcity that motivates the downselect."],"remaining_contrastive_claim":"The testable contrast is whether a safety-gated, resource-capped, hidden-condition portfolio contest produces advancement decisions that remain more stable under independent hardware-in-the-loop and human evaluation than a public single-winner benchmark, without requiring the program to fund every candidate. This is a design hypothesis, not an effect claim.","authority_safety":{"decision_authority":"The aircraft program's formally designated engineering and safety authorities control development downselection. The independent safety board controls test-gate clearance, and the applicable certification or operational authority retains all approval powers. Airworthiness, crew protection, range safety, and flight authorization are noncontestable and cannot be conferred by score or prize.","authorized_first_step":"Run a non-operational shadow contest using configuration-frozen controller builds in offline simulation only. An independent test custodian may generate held-out cases and calculate hypothetical rankings, but those rankings do not alter contracts, employment decisions, integration access, certification evidence, or flight plans.","excluded_actions":["No command of an aircraft, actuator, flight simulator occupied by operational crews, or flight-test asset","No waiver or aggregation-away of a safety-envelope violation","No flight authorization, certification credit, procurement award, or supplier sanction based on the shadow result","No hidden test drawn from conditions outside the approved simulation validity envelope","No use of anomaly screens as proof of collusion or misconduct","No disclosure of one entrant's protected code, parameters, or results to another entrant","No post-deadline rule or score-weight change based on entrant identity or performance"],"halt_rollback":"Halt if scenario custody is compromised, candidate configurations cannot be frozen, simulation validity is disputed for a score-driving condition, a safety constraint has been encoded as compensable, or protected technical data cannot be separated. Quarantine affected results, disclose the invalidating condition to all entrants, restore the unchanged configuration archive, and either rerun under predeclared correction rules or void the shadow contest."},"negative_tests":{"strongest_counterevidence":"Existing downselects already use independently generated inaccessible cases, configuration-controlled reproduction, noncompensable safety gates, comparable resource limits, and later architectural reopening, while visible-case rank remains consistent through hardware-in-the-loop and human evaluation.","problem_falsifier":"Across a predeclared sample of candidate controllers, public-case rankings remain stable under independently generated conditions; simulator-specific dependencies and unreproduced results are absent; and additional compute or test effort does not materially alter relative rank after basic competence is reached.","intervention_falsifier":"The shadow arena's portfolio fails the same safety or robustness checks as the conventional winner, hidden-case rank is unstable under innocuous scenario-seed changes, complementary selection advances inferior redundancy rather than useful diversity, or the resulting ranking predicts hardware-in-the-loop and human evaluation no better than the conventional baseline.","risks":["Hidden cases may still share assumptions or defects with the disclosed simulator.","Scenario secrecy may reduce entrants' ability to diagnose legitimate modeling errors.","Composite scoring can conceal consequential tradeoffs unless safety floors remain noncompensable.","Resource accounting may favor incumbents that possess pretrained models, reusable tools, or unpriced infrastructure.","A cap may prevent a challenger from performing legitimate integration work needed to become comparable.","Portfolio complementarity can become a discretionary rationale for advancing a favored lower scorer.","Test custodians or judges may have conflicts of interest or inadvertently leak scenario information.","Teams may optimize to inferred scenario-generator structure rather than operational robustness.","Statistical similarity screens can confuse common engineering choices with coordination.","Interface requirements intended to preserve portability may privilege the incumbent architecture.","A challenger window set after major integration commitments may be ceremonial rather than contestable.","Simulation success may not transfer to hardware, pilot interaction, maintenance, or certification review."]},"next_evidence_step":"Pre-register one offline shadow downselect using three to six configuration-frozen, non-operational controller builds and a bounded scenario matrix covering approved nominal states, disturbances, sensor faults, actuator constraints, and recovery tasks. Divide the matrix into disclosed development cases, a custodian-held generator, and a second independently seeded validation set. Apply identical resource accounting, safety gates, scoring, audit, and portfolio rules; compare the resulting choices with the conventional public-suite winner and an unranked expert-panel choice. Report rank stability, gate failures, reproducibility, control activity, intervention demand, resource use, portfolio diversity, audit disputes, and sensitivity to lawful score weights and seeds. Do not proceed to hardware or flight; use the result only to decide whether a separately authorized hardware-in-the-loop study is warranted.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Proposal 1 governs airline rivalry for scarce departure priority during capacity-constrained airport recovery using verified gate readiness, capped operational credits, and repeated release nominations. This proposal instead governs engineering-team rivalry for scarce flight-control integration resources, addressing benchmark overfitting, simulator exploitation, resource escalation, and architectural lock-in through hidden-condition safety-gated testing and complementary portfolio downselection. Its actors, contested prize, observable state, intervention, harms, evidence environment, and causal path are distinct. It can be adopted by an aircraft development program without changing airport surface operations or departure sequencing, while Proposal 1 can be adopted without conducting a controller-development contest.","revision_record":{"parent_version":null,"progress_targets_addressed":["Second independently adoptable proposal","Material problem and intervention separation from proposal 1","Concrete engineering actors and observables","Complete archetype-to-domain mapping","Mechanism removal counterfactuals","Authority-preserving bounded evidence","Explicit diversity from every earlier proposal"],"conceptual_changes":[],"operational_changes":[],"evidence_changes":[],"claim_changes":[]}}