{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"bounded_rivalry_governance__aviation_aeronautics:P2:v0","cell_id":"bounded_rivalry_governance__aviation_aeronautics","search_queries":["site:nasa.gov flight control challenge benchmark hidden scenarios hardware in the loop controller competition robustness","site:faa.gov flight control system simulation validation hardware in the loop AC 25.1309","flight control benchmark competition robust controller hidden test scenarios research","NASA flight control system hardware-in-the-loop verification validation robust controller development","NASA control challenge problem flight control benchmark teams controller comparison hidden evaluation","\"flight control\" \"challenge problem\" benchmark robust NASA","AIAA flight control benchmark challenge controller competition","aerospace control design challenge benchmark downselect hardware in loop","site:nasa.gov multiple teams flight control algorithms competition hardware in the loop selected","site:darpa.mil aircraft flight control competition teams simulation challenge","site:afresearchlab.com flight control challenge simulation teams","NASA \"flight control\" \"downselect\" controller","Springer Robust Flight Control A Design Challenge GARTEUR Research Civil Aircraft Model 1997","\"Robust Flight Control\" \"Design Challenge\" GARTEUR official","doi Robust Flight Control A Design Challenge Magni Bennani Terlouw","\"A Benchmark Comparison of Learned Control Policies for Agile Quadrotor Flight\" DOI","site:ieeexplore.ieee.org \"A Benchmark Comparison of Learned Control Policies\"","site:roboticsproceedings.org quadrotor benchmark sim-to-real control policies 45 km/h","primary research benchmark overfitting public leaderboard hidden test set competition leaderboard shakeup","Kaggle leaderboard overfitting public private leaderboard research paper","benchmark overfitting adaptive holdout test set competition paper"],"sources":[{"source_id":"S1","title":"AC 25.1309-1B—System Design and Analysis","publisher":"Federal Aviation Administration","url":"https://www.faa.gov/regulations_policies/advisory_circulars/index.cfm/go/document.information/documentID/1043037","source_class":"OFFICIAL_GUIDANCE","publication_date":"2024-08-30","accessed_at":"2026-08-03","claims_supported":["The active advisory circular supplies acceptable, nonexclusive means for showing compliance with 14 CFR 25.1309.","Engineering and operational judgment and the FAA compliance process remain authoritative; a contest score cannot itself confer airworthiness approval."]},{"source_id":"S2","title":"AC 25.1329-1B—Approval of Flight Guidance Systems","publisher":"Federal Aviation Administration","url":"https://www.faa.gov/sites/faa.gov/files/2022-11/AC25.1329-1B.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2006-07-17","accessed_at":"2026-08-03","claims_supported":["Flight-guidance-system validation typically combines analysis, laboratory tests, simulation, and flight tests.","Representative operational conditions, environmental conditions, failure scenarios, pilot-in-the-loop evaluation, human factors, and workload can be material to validation.","Applicants should coordinate validation and certification methods with the FAA; laboratory or simulation evidence does not independently authorize flight."]},{"source_id":"S3","title":"NASA Systems Engineering Handbook, Section 5.0—Product Realization","publisher":"National Aeronautics and Space Administration","url":"https://www.nasa.gov/reference/5-0-product-realization/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2024-08-12","accessed_at":"2026-08-03","claims_supported":["NASA describes iterative verification and validation progressing from models and simulations to integrated hardware and software.","NASA guidance calls for baselined validation plans, objective evidence, end-to-end testing, operational scenarios, anomaly resolution, and testing in representative environments.","NASA explicitly identifies hardware-in-the-loop testing and integration of operators and maintainers as relevant verification and validation activities."]},{"source_id":"S4","title":"Flight and Ground Experimental Test Capabilities—Verification and Validation","publisher":"National Aeronautics and Space Administration, Armstrong Flight Research Center","url":"https://www.nasa.gov/centers-and-facilities/armstrong/flight-and-ground-experimental-test-capabilities/","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026-06-22","accessed_at":"2026-08-03","claims_supported":["NASA Armstrong operates a reconfigurable simulation test bench for verification and validation of F/A-18 software and hardware.","The bench is intended to support digital flight-control systems, faster data collection, more detailed modeling, and reduced dependence on expensive flight tests.","NASA Armstrong is an identifiable technically capable potential adopter or evaluation partner, although this source does not request the proposed downselect governance."]},{"source_id":"S5","title":"Robust Flight Control Design Challenge: Problem Formulation and Manual—The Research Civil Aircraft Model","publisher":"Group for Aeronautical Research and Technology in Europe","url":"https://garteur.org/wp-content/reports/FM/FM_AG-08_TP-088-3.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"1995-06-15","accessed_at":"2026-08-03","claims_supported":["GARTEUR established an open robust-flight-control design challenge using a common civil-aircraft model, specifications, disturbances, constraints, and evaluation procedures.","Comparative flight-control benchmark challenges involving research establishments and industry are established prior art rather than a novel contest form.","The source does not document hidden final scenarios, resource caps, complementary portfolio selection, or downstream prediction against independent hardware-in-the-loop and human evaluations."]},{"source_id":"S6","title":"The Ladder: A Reliable Leaderboard for Machine Learning Competitions","publisher":"Proceedings of Machine Learning Research","url":"https://proceedings.mlr.press/v37/blum15.html","source_class":"PRIMARY_RESEARCH","publication_date":"2015-06-01","accessed_at":"2026-08-03","claims_supported":["Repeated, adaptive feedback from a public leaderboard can permit participants to overfit the holdout data underlying the leaderboard.","Reliable-leaderboard methods can limit disclosed information while retaining useful comparative feedback.","This is general competition research, not evidence that aircraft flight-control teams currently exhibit the proposed pathology."]},{"source_id":"S7","title":"A Meta-Analysis of Overfitting in Machine Learning","publisher":"Neural Information Processing Systems Foundation","url":"https://papers.neurips.cc/paper_files/paper/2019/hash/ee39e503b6bedf0c98c388b7e8589aca-Abstract.html","source_class":"PRIMARY_RESEARCH","publication_date":"2019","accessed_at":"2026-08-03","claims_supported":["A study of more than 100 Kaggle competitions compared repeatedly viewed public rankings with separate final rankings.","The study found little evidence of substantial overfitting, providing important counterevidence against assuming that visible leaderboards routinely produce severe rank distortion.","The result is not flight-control-specific and does not exclude overfitting where scenario counts are small, simulators are exploitable, or engineering teams receive richer feedback."]},{"source_id":"S8","title":"DARPA Lift Challenge—Competitors, Rules, Scoring, Safety, and Protest Procedure","publisher":"Defense Advanced Research Projects Agency","url":"https://www.darpa.mil/research/challenges/lift/competitors","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-07-09","accessed_at":"2026-08-03","claims_supported":["A current aviation challenge uses eligibility rules, objective scoring, preflight checks, safety officials with stop authority, disqualification, and a formal protest process.","DARPA separates competition scoring from compliance with applicable FAA rules and requires safety evidence before participation.","This is strong prior art for bounded aviation rivalry, but it evaluates complete VTOL aircraft in physical flight rather than selecting flight-control candidates through hidden simulation and portfolio downselection."]}],"problem_evidence":{"support":"MODERATE","rationale":"The stakes are visible: FAA and NASA guidance require flight-control performance, safety, human interaction, failure behavior, configuration-consistent evidence, and higher-fidelity validation beyond a narrow simulation score. Adaptive-leaderboard research establishes a mechanism by which repeated score feedback can induce overfitting. However, no opened source documents the stated pathology in an actual aircraft-development downselect, and a large competition meta-analysis found little substantial public-to-private overfitting. The exact prevalence, magnitude, and resource-arms-race component therefore remain unverified.","source_ids":["S1","S2","S3","S6","S7"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"NASA Armstrong is an identifiable organization with a relevant flight-control V&V bench, while NASA program engineering authorities and FAA certification personnel are credible authorizers for progressively higher-fidelity evidence. Their published materials express a need for representative scenarios, integrated testing, human evaluation, anomaly resolution, and economical reduction of flight-test dependence. None expressly requests a hidden-condition, capped, portfolio downselect, so adopter pull for this governance package is inferred rather than demonstrated.","source_ids":["S1","S2","S3","S4"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"GARTEUR Robust Flight Control Design Challenge","similarity":"A common aircraft model, disturbances, requirements, constraints, evaluation procedures, and multiple independent controller-design teams directly match the comparative flight-control arena.","remaining_difference":"The opened manual does not show custodian-held final scenarios, public-versus-hidden rank auditing, resource caps, two-candidate portfolio selection, a challenger window, or validation of selection quality against later HIL and human evidence.","source_ids":["S5"]},{"name":"DARPA Lift Challenge","similarity":"A current government aviation competition combines eligibility, scoring, safety gates, compliance obligations, official stop authority, penalties, inspections, and protests.","remaining_difference":"It is a physical-aircraft prize challenge with a simple public metric, not an internal engineering-resource downselect for flight-control software using hidden cases and complementary selection.","source_ids":["S8"]},{"name":"Reliable leaderboard and separate final evaluation practices","similarity":"The Ladder directly addresses adaptive public-score feedback and protects the validity of comparative rankings; the NeurIPS meta-analysis evaluates public-versus-final ranking behavior at scale.","remaining_difference":"These methods concern machine-learning competitions, not safety-gated flight-control simulation, and neither incorporates aircraft safety authority, integration burden, pilot workload, portfolio diversity, or HIL transfer.","source_ids":["S6","S7"]},{"name":"NASA/FAA staged flight-control verification and validation","similarity":"Established practice already uses analysis, laboratory tests, simulation, operational scenarios, pilot involvement, configuration evidence, and progressively higher-fidelity testing before approval.","remaining_difference":"Assurance practice validates candidates but does not itself govern rivalry for scarce integration slots or test whether a bounded contest predicts downstream validation better than a conventional downselect.","source_ids":["S1","S2","S3","S4"]}],"distinctive_claim_remaining":"Holding candidate builds, approved scenario families, safety constraints, and resource accounting constant, a safety-gated and hidden-condition contest that advances two complementary candidates will produce a selection with greater agreement and rank stability under independently seeded validation and later authorized HIL/human evaluation than a public-suite single-winner benchmark or an unranked expert-panel choice, without funding HIL work for every candidate. This comparative effect remains unmeasured.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"The constituent pieces are implementable: GARTEUR demonstrates comparative flight-control benchmarks; reliable-leaderboard research supplies methods for limiting adaptive feedback; DARPA demonstrates eligibility, safety, scoring, penalties, and protest governance; and NASA/FAA materials establish simulation, HIL, configuration evidence, and human evaluation workflows. Integration is not turnkey. A program would still need lawful access to proprietary controller builds, an accepted simulation-validity envelope, independent scenario custody, normalized accounting for legacy tools and compute, protected technical-data handling, conflict management, and safety-authority approval before any HIL or crew-involved work. The proposed offline shadow step avoids operational control and certification claims.","source_ids":["S1","S2","S3","S4","S5","S6","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Selecting a brittle flight-control candidate for scarce integration could waste substantial engineering capacity and delay discovery of safety, workload, and integration problems. Realized impact is unmeasured because the frequency and magnitude of bad downselects are unknown.","source_ids":["S1","S2","S3","S4"]},"stakeholder_pull":{"score":3,"rationale":"NASA and FAA materials show strong demand for robust, representative, economical flight-control V&V, and NASA Armstrong has relevant infrastructure. No source expresses demand for this particular contest design.","source_ids":["S1","S2","S3","S4"]},"incremental_advantage":{"score":3,"rationale":"Hidden final cases, noncompensable gates, auditing, resource normalization, and two-candidate advancement plausibly improve on a visible single-winner score. Whether they outperform ordinary expert review or conventional assurance is an empirical question.","source_ids":["S5","S6","S7"]},"distinctiveness_plausibility":{"score":3,"rationale":"Flight-control challenges, governed aviation competitions, protected leaderboards, and staged V&V are all established. The remaining plausible distinction is their combination plus the downstream decision-stability claim, not any individual mechanism.","source_ids":["S3","S5","S6","S8"]},"technical_implementability":{"score":4,"rationale":"Offline simulations, configuration freezing, held scenario seeds, audit logs, safety gates, and comparative scoring use mature techniques. Fair resource accounting and generator validity are difficult but not fundamental technical blockers.","source_ids":["S3","S4","S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"A program sponsor can authorize an internal shadow comparison, while designated safety authorities and the FAA retain authority over HIL credit, crew involvement, flight, and certification. Proprietary-data, supplier-contract, conflict-of-interest, and export-control arrangements remain program-specific gaps.","source_ids":["S1","S2","S3","S4"]},"evidence_readiness":{"score":2,"rationale":"A protocol can be preregistered now, but decisive evidence requires proprietary frozen builds, scenario and resource data, an independent custodian, and eventually separately authorized HIL/human results. Web research cannot establish the effect.","source_ids":["S4","S5","S7"]},"safety_net_benefit":{"score":4,"rationale":"Advancing two materially different candidates and treating simulation rank as access to validation rather than flight authorization preserves fallback capacity and limits premature architectural lock-in.","source_ids":["S2","S3","S8"]},"scalability":{"score":3,"rationale":"The governance template can recur across programs, but each aircraft needs a validated model, approved envelopes, interfaces, safety constraints, scenario families, and protected data handling. Those local adaptations constrain low-cost scaling.","source_ids":["S1","S2","S3","S4","S5"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"One preregistered, offline shadow downselect for three to six existing configuration-frozen controllers, including scenario-custodian labor, resource metering, reproducibility checks, independent scoring, sensitivity analysis, debriefs, and reporting; excludes new controller development, HIL, occupied simulators, flight assets, and certification credit.","confidence":"LOW","assumptions":["Existing controller builds and a suitable approved simulation harness are contributed in kind.","Approximately 4-10 person-months of engineering, assurance, statistics, and governance effort are required.","No protected-data remediation or export-control restructuring is needed.","This is a resource-equivalent estimate; no direct 2026 facility or labor quotations were located."],"source_ids":["S3","S4","S5","S6"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Program-specific productionization of the arena: independent scenario-generation and custody pipeline, configuration archive, reproducibility environment, audit and access logging, resource-accounting rules, safety-gate implementation, scoring validation, legal/data-rights controls, conflict review, appeals, and dry runs.","confidence":"LOW","assumptions":["The program already has an aircraft model, simulation infrastructure, and candidate interfaces.","Startup does not purchase a new full-scale HIL laboratory.","One aircraft program and one controller family are in scope.","Protected technical data can remain inside an approved program environment."],"source_ids":["S1","S2","S3","S4","S5","S8"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"First consequential program cycle through two-candidate HIL integration and non-operational pilot/maintainer evaluation, including independent test operations, anomaly resolution, configuration assurance, facility time, governance, and post-contest review; excludes flight testing, aircraft modification, certification campaigns, and entrant controller-development costs.","confidence":"LOW","assumptions":["An existing HIL or representative simulation facility is available rather than constructed from scratch.","Two candidates receive bounded integration support.","Pilot participation occurs only in an appropriately authorized simulator.","Major model redevelopment or hardware procurement would move the estimate upward."],"source_ids":["S2","S3","S4"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"One recurring program-scale downselect or challenger cycle per year, including scenario refresh, custodian and judge labor, security and configuration audits, limited HIL regression time, appeals, and outcome review; excludes ongoing controller R&D and flight-test operations.","confidence":"LOW","assumptions":["Core infrastructure is reused after startup.","One major cycle and limited challenger activity occur annually.","Three to six entrants use stable standardized interfaces.","Material simulator revalidation or extensive HIL integration would exceed this range."],"source_ids":["S3","S4","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"External sources establish that robust, representative and higher-fidelity validation matters and that adaptive leaderboard overfitting is possible, but they do not demonstrate the stated pathology in aircraft-development downselects. The large NeurIPS meta-analysis supplies counterevidence to assuming substantial leaderboard overfitting is common.","source_ids":["S1","S2","S3","S6","S7"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"NASA Armstrong is an identifiable technically capable potential adopter with flight-control V&V infrastructure; NASA engineering authorities and the FAA are credible authorizers for progressively higher-fidelity validation and approval. Expressed pull for the exact governance package remains absent.","source_ids":["S1","S2","S3","S4"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal specifies named comparators and a measurable contrast: decision agreement and rank stability under independent validation and later HIL/human evidence, subject to common builds, scenarios, constraints, and resource accounting.","source_ids":["S5","S6","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A three-to-six-candidate offline shadow contest can be preregistered, run without operational control or allocation consequences, and compared with a public-suite winner and expert-panel selection using rank stability, gate failures, reproducibility, intervention demand, resource use, and seed/weight sensitivity.","source_ids":["S3","S5","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"For the bounded offline step, no contest output commands hardware, authorizes flight, supplies certification credit, or overrides a safety violation. FAA and program safety authority remain intact. HIL, occupied simulation, contracting consequences, and flight would require separate authorization.","source_ids":["S1","S2","S3","S8"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The four estimates have explicit scopes, exclusions, and resource assumptions, but no direct program labor rates, simulation-custody bids, proprietary-data remediation estimates, or HIL facility quotations were found. NASA only supports the relative proposition that simulation facilities can reduce expensive flight testing.","source_ids":["S3","S4"]}},"next_evidence_step":"With an aircraft-program partner, preregister one non-operational offline shadow study using three to six existing configuration-frozen controllers. Keep candidate builds, approved scenario families, safety constraints, and nominal compute envelopes identical across three selection methods: (A) the proposed noncompensable-gate, custodian-held, resource-capped two-candidate portfolio arena; (B) a fully disclosed public-suite single-winner benchmark; and (C) an unranked expert-panel choice. Reserve a second independently seeded validation set that is inaccessible to teams and the first custodian. Primary endpoints are selection agreement and rank stability on the second set; secondary endpoints are safety-gate failures, reproducibility, actuator saturation/control activity, simulated pilot-intervention demand, declared integration burden, resource consumption, portfolio diversity, disputes, and sensitivity to lawful seeds and score weights. Falsify the intervention if method A is no more stable or predictive than B or C, if innocuous seeds reverse its choices, if portfolio selection merely advances an inferior redundant candidate, or if resource accounting systematically favors incumbents. Void the study upon scenario leakage, unverifiable builds, compensable safety violations, or disputed score-driving simulator validity. Use results only to decide whether to seek separate authority and funding for an HIL study.","blocking_evidence":["No opened source measures how often public-case tuning, simulator-artifact exploitation, unreproduced results, or resource escalation changes real flight-control downselects.","The comparative effect cannot be estimated without proprietary configuration-frozen controller builds and program-specific scenario data.","It is unknown whether the proposed selection predicts independent HIL, pilot, maintainer, or integration outcomes better than the conventional winner or expert panel.","The fairness and enforceability of compute, simulator-hour, legacy-tool, and sponsor-support accounting have not been tested with actual suppliers or internal teams.","No program partner has expressed willingness to adopt the arena, provide protected artifacts, or accept an independent scenario custodian.","No direct 2026 labor, security, data-rights, simulation, or HIL facility quotations support the cost bands.","Composite-score construct validity, portfolio complementarity criteria, scenario-generator validity, and minimum detectable effect remain to be preregistered."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search found established robust-flight-control design challenges, aviation competitions with safety and protest governance, adaptive-leaderboard protections, and staged NASA/FAA verification practices. It did not establish whether the exact combined governance package exists elsewhere. World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain a named aircraft-program partner, three to six frozen controller builds, lawful technical-data access, and written authority for a consequence-free shadow study.","Preregister the simulation-validity envelope, independent scenario custody, second validation set, noncompensable safety gates, metrics, effect thresholds, comparator procedures, and falsifiers.","Demonstrate auditable and entrant-neutral accounting for compute, reusable tools, simulator hours, external data, and sponsor engineering support.","Run the shadow comparison and quantify public-to-hidden rank movement, selection stability, reproducibility, safety failures, resource effects, seed/weight sensitivity, disputes, and portfolio diversity.","If the offline result survives its falsifiers, secure separate safety, facility, human-subject/workforce, contracting, and data-governance approvals plus direct cost quotations for an HIL study."],"reason":"Bounded web research established high safety relevance, credible authorities, technical feasibility, and substantial adjacent prior art, but not the candidate's central empirical premise or incremental effect. Resolving the remaining questions requires proprietary builds, program data, partner participation, and live comparative simulation followed potentially by authorized HIL/human testing; additional web search cannot decide them."},"proposal_index":2}