{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp09_archetype_breadth150_20260804","research_id":"eoa_inverse_innovation_exp09_light_prior_art_20260804","cell_id":"bounded_rivalry_governance__computer_science","search_lanes":{"direct_problem_and_intervention":{"queries":["query scheduler benchmark overfitting production promotion sealed workloads canary","\"query scheduler\" competition benchmark production default governance"],"source_ids":["SRC1","SRC2","SRC3"],"no_result_note":null},"synonyms_and_historical_terms":{"queries":["database workload management scheduler benchmark generalization workload drift tail latency fairness","query optimizer benchmark specialization workload change production robustness"],"source_ids":["SRC1","SRC2"],"no_result_note":null},"products_practices_and_standards":{"queries":["TPC benchmark rules full disclosure special optimization audit","MLPerf inference rules benchmark detection audit appeal conflict of interest","Kubernetes deployment canary rollback official documentation"],"source_ids":["SRC2","SRC3","SRC4"],"no_result_note":null},"component_combination":{"queries":["sealed tests resource budget audited leaderboard appeals canary deployment","database scheduler tail latency fairness resource consumption benchmark","benchmark governance anti-gaming reproducibility adjudication challenger window"],"source_ids":["SRC1","SRC2","SRC3","SRC4"],"no_result_note":null}},"sources":[{"source_id":"SRC1","title":"Bao: Learning to Steer Query Optimizers","publisher":"arXiv","url":"https://arxiv.org/abs/2004.03814","source_type":"PRIMARY_RESEARCH","claims_supported":["Query-optimization approaches can suffer from impractical training overhead, inability to adapt to workload changes, and poor tail performance.","Bao was evaluated for adaptation to changing workloads, data, and schemas and for end-to-end and tail-latency performance, supporting the relevance of production-transfer testing."]},{"source_id":"SRC2","title":"TPC Benchmark DS Standard Specification, Revision 4.0.0","publisher":"Transaction Processing Performance Council","url":"https://tpc.org/TPC_Documents_Current_Versions/pdf/TPC-DS_v4.0.0.pdf","source_type":"OFFICIAL_STANDARD","claims_supported":["A database benchmark standard already prescribes execution, validation, disclosure, reproducibility, and audit requirements.","TPC-DS limits query-specific hints, regulates and requires disclosure of profile-directed optimization, and requires disclosure of tunable optimizer, recovery, operating-system, and compiler settings."]},{"source_id":"SRC3","title":"MLPerf Inference Rules","publisher":"MLCommons","url":"https://github.com/mlcommons/inference_policies/blob/master/inference_rules.adoc","source_type":"OFFICIAL_STANDARD","claims_supported":["A mature competitive benchmark already prohibits benchmark detection and input-specific optimization and requires replicability.","Its governance includes submission review, randomized or committee-selected audits, conflict-free auditors, evidence requirements, sanctions for material audit failure, and an appeal route."]},{"source_id":"SRC4","title":"Deployments","publisher":"Kubernetes Authors","url":"https://kubernetes.io/docs/concepts/workloads/controllers/deployment/","source_type":"OFFICIAL_GUIDANCE","claims_supported":["Controlled software rollout, rollout-status monitoring, pausing, failure detection, and rollback are established production practices.","Kubernetes documents canary deployment to a subset of users or servers, supporting the feasibility of reversible provisional promotion."]}],"problem_evidence":{"status":"PARTLY_SUPPORTED","finding":"The sources make the technical hazards visible: query optimizers can have adaptation and tail-performance problems, while benchmark standards explicitly control special-casing, input-specific optimization, reproducibility, disclosure, and auditing. The bounded search did not directly document multiple database teams escalating tuning expenditure for one shared default-scheduler slot or an observed instance in which that exact contest entrenched a fragile winner.","source_ids":["SRC1","SRC2","SRC3"]},"closest_prior_art":[{"name":"TPC-DS governed database benchmarking","source_ids":["SRC2"],"overlap":"Detailed database benchmark rules, fixed execution and validation procedures, configuration and optimization disclosure, restrictions on query-specific behavior, reproducibility, and audit.","remaining_difference":"TPC-DS governs publication of system benchmark results, not an internal contest for one revocable production scheduler slot; it lacks sealed production traces, an equal tuning-compute cap, organizational appeals, shadow traffic, and recurring incumbent challenges."},{"name":"MLPerf Inference submission and audit regime","source_ids":["SRC3"],"overlap":"A competitive benchmark with formal rules, standardized tests, anti-special-casing provisions, replicability, independent auditing, review, appeals, and penalties.","remaining_difference":"It validates and publishes benchmark submissions rather than delegating a single live production default. It does not govern rollback authority, winner incumbency, time-limited control, or recurring access to one production slot."},{"name":"Kubernetes controlled rollout, canary, and rollback practices","source_ids":["SRC4"],"overlap":"Controlled progression, status checks, deployment to a subset, pausing, failure handling, and rollback overlap the proposed provisional promotion stage.","remaining_difference":"Deployment control does not govern competition among candidate teams, sealed evaluation, tuning-resource escalation, conflicts, appeals, or post-win lock-in."},{"name":"Bao adaptive query optimization","source_ids":["SRC1"],"overlap":"Workload-sensitive optimizer evaluation, adaptation to changes, operational practicality, training cost, and tail-latency performance.","remaining_difference":"Bao is an optimizer approach rather than a neutral arena governing multiple teams' eligibility, optimization budgets, evaluation conduct, adjudication, and time-limited control of a default."}],"prior_art_disposition":"ADJACENT_PRIOR_ART","contrastive_claim_remaining":"For two or more conforming query schedulers competing for one production-default slot, the combined regime of preregistered multidimensional scoring, untouched traces, equal tuning-resource limits, independent adjudication, reversible shadow or canary validation, and a time-limited default with recurring challenger windows will yield rankings that transfer more reliably to untouched workload periods, without increasing safety-floor violations, than selection by a visible throughput-only benchmark. The retained sources establish many components separately but not this end-to-end governance of the scarce production slot.","contrastive_claim_falsifier":"Falsify the claim if a preregistered retrospective shows no improvement in untouched-period rank stability or production proxies, any increase in correctness, reliability, privacy, or tenant-isolation violations, or that eligibility and compute constraints determine winners independently of scheduler behavior. Distinctness is also falsified by finding an existing database-scheduler program that already combines the specified rulebook, untouched tests, resource cap, independent adjudication, reversible production validation, time-limited appointment, and recurring challenge mechanism.","gates":{"adequate_source_search":{"status":"PASS","rationale":"The bounded search covered direct formulations, synonyms and older terminology, standards and operational practices, and combinations of component subproblems. Exactly four opened sources from four publishers include primary research, two official standards, and official guidance.","source_ids":["SRC1","SRC2","SRC3","SRC4"]},"supported_problem":{"status":"PASS","rationale":"The exact organizational rivalry was not observed, but workload-transfer and tail-performance risks plus explicit standards against benchmark-specific behavior and irreproducible results support PARTLY_SUPPORTED problem evidence.","source_ids":["SRC1","SRC2","SRC3"]},"distinct_testable_claim":{"status":"PASS","rationale":"The remaining claim identifies a specific combined intervention, comparator, population, and measurable outcomes: untouched-period rank stability and safety-floor violations. No retained source implements the full combination for a production query-scheduler slot.","source_ids":["SRC1","SRC2","SRC3","SRC4"]},"bounded_next_test":{"status":"PASS","rationale":"An offline retrospective can rebuild existing schedulers under fixed resources, use one public and one authorized untouched trace set, and compare rankings, resource use, reproducibility, and safety violations without changing production.","source_ids":["SRC1","SRC2","SRC3"]},"no_obvious_safety_or_authority_stop":{"status":"PASS","rationale":"The proposed first test is offline, preserves the current production default, requires privacy-approved traces and reproducible builds, and has explicit halt conditions. Canary and rollback practices support a reversible later stage only if separately authorized.","source_ids":["SRC2","SRC4"]}},"screen_survival":true,"world_novelty_boundary":"This bounded four-source screen supports only coarse researchability and an adjacent-prior-art disposition. It cannot establish world novelty, patentability, market size, expert acceptance, realized value, or the absence of unpublished or unindexed scheduler-selection practices."}