{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp09_archetype_breadth150_20260804","cell_id":"bounded_rivalry_governance__chemistry_materials","arm":"BREADTH_PROBE_ONE_SHOT","candidate_id":"bounded_rivalry_governance__chemistry_materials__P1","proposal_index":1,"version":0,"title":"Governed Qualification Challenge for Scarce Materials Scale-Up Slots","problem":"Research teams competing for a shared facility's limited pilot-scale synthesis slots are compared through self-selected battery-material results produced under different test protocols. Because teams can improve their selection chances by reporting unusually favorable cells, omitting failed replicates, choosing resource-intensive formulations, or deferring safety and waste questions, the competition can select the submission best optimized for presentation rather than the material most ready for reproducible and responsible scale-up.","actors":["Materials-synthesis research teams","Electrochemical testing staff","Shared pilot-scale facility director","Independent technical reviewers","Environmental health and safety staff","Laboratory technicians handling synthesis and waste","Downstream prototype engineers"],"observable_state":"Candidate dossiers contain non-comparable cycling conditions, inconsistent replicate counts, selectively reported cells, incomplete failure data, and uneven accounting of precursor scarcity, hazardous operations, and waste. Rankings depend materially on which submitted result or protocol is accepted, while access to the next scale-up cycle remains scarce.","consequence":"A fragile, unsafe, or impractical candidate can receive scarce scale-up capacity, exposing staff and equipment to avoidable hazards, consuming materials and facility time, and displacing candidates whose performance is more reproducible under common conditions.","affected_objective":"Select a small portfolio of battery-material candidates that can reproduce relevant performance under common tests and meet explicit safety, resource, and scale-up constraints.","intervention":"Replace the informal pitch competition with a bounded qualification challenge for two provisional scale-up slots. Publish eligibility and allowed-evidence rules; require a fixed number of randomly identified samples and disclosure of all runs in a defined window; have a neutral facility team conduct blinded common-protocol tests; score reproducibility, retained performance, synthesis yield, hazardous-operation burden, critical-material intensity, and waste alongside peak performance; impose a common sample and testing-resource cap; apply non-waivable safety floors; disclose reviewer conflicts; provide a documented correction and appeal window; prohibit sample interference, undisclosed protocol changes, and coordination over submitted results; award two staged slots rather than permanent priority; and reopen qualification after post-scale-up review.","structural_mapping":[{"archetype_element":"Rivalry Purpose Statement","domain_realization":"Use rivalry to reveal which battery-material candidates are sufficiently reproducible and scale-ready, not which team can present the highest isolated measurement."},{"archetype_element":"Scarce Prize or Selection Constraint","domain_realization":"Two provisional pilot-scale synthesis slots in the next facility cycle."},{"archetype_element":"Competitor Eligibility Boundary","domain_realization":"Teams must have a documented composition, bench-scale synthesis procedure, minimum replicate set, complete run inventory for the submission window, and preliminary hazard review."},{"archetype_element":"Contest Arena Boundary","domain_realization":"Teams may improve their material and documented process before submission but may not choose which submitted units are tested, alter blinded samples, suppress in-window runs, or interfere with another team's samples or equipment."},{"archetype_element":"Performance Metric and Scoring Basis","domain_realization":"A predeclared composite score combines blinded common-protocol performance, between-sample dispersion, synthesis yield, critical-material intensity, hazardous-operation burden, and waste; safety floors operate as gates rather than compensable score components."},{"archetype_element":"Information Disclosure and Observability Rule","domain_realization":"The facility publishes protocols, scoring weights, coded test outputs, deviations, reviewer conflicts, and reasons for qualification decisions while protecting legitimate composition secrets from rival teams."},{"archetype_element":"Fair Process and Due Process Layer","domain_realization":"Teams can correct clerical or protocol-attribution errors and appeal documented departures from the rulebook to reviewers who did not make the original decision."},{"archetype_element":"Anti-Sabotage and Anti-Collusion Guardrail","domain_realization":"Chain-of-custody records, coded samples, access logs, and a penalty schedule address sample interference, result suppression, and coordinated misreporting."},{"archetype_element":"Externality and Spillover Boundary","domain_realization":"Worker hazard, scarce-precursor use, solvent burden, and waste are evaluated before scarce scale-up access is granted."},{"archetype_element":"Escalation and Arms-Race Damper","domain_realization":"A common cap on submitted samples, instrument hours, and reimbursable characterization prevents teams from winning mainly through greater testing expenditure."},{"archetype_element":"Prize Decomposition or Multiple-Winner Design","domain_realization":"Two staged, conditional scale-up slots preserve portfolio diversity and prevent one favorable round from creating permanent facility priority."},{"archetype_element":"Learning and Recalibration Loop","domain_realization":"After the scale-up runs, reviewers compare predicted and observed reproducibility, hazards, yield, and waste, then revise or retire scoring rules before the next round."}],"mechanism_mapping":[{"mechanism_slug":"contest_rulebook","role":"Defines eligibility, evidence windows, sample selection, common tests, scoring, safety floors, prohibited conduct, and appeals before submissions are judged.","counterfactual_removal":"Without the rulebook, teams and judges can reinterpret acceptable evidence after seeing results, restoring the informal and strategically manipulable arena."},{"mechanism_slug":"ranked_leaderboard_with_audit","role":"Produces an inspectable provisional ranking from blinded tests and disclosed calculations, with an audit trail for samples, deviations, and score corrections.","counterfactual_removal":"Without auditability, ranking errors, selective exclusions, and protocol deviations cannot be distinguished from genuine performance differences."},{"mechanism_slug":"spending_cap_or_resource_cap","role":"Limits submitted samples and shared characterization time per team so selection is not dominated by the ability to run more trials and cherry-pick extremes.","counterfactual_removal":"Without a cap, teams can escalate testing volume, increasing selection advantage through search intensity rather than material readiness."},{"mechanism_slug":"multiple_award_or_portfolio_selection","role":"Allocates two provisional slots and preserves more than one technical pathway through initial scale-up validation.","counterfactual_removal":"With a single durable winner, measurement noise can prematurely eliminate alternatives and give one team disproportionate influence over later qualification."},{"mechanism_slug":"post_contest_impact_review","role":"Checks whether qualification scores predicted pilot-scale reproducibility, yield, hazards, and waste and triggers rule recalibration.","counterfactual_removal":"Without review, strategically adapted or invalid metrics persist even when scale-up outcomes contradict the ranking."}],"causal_chain":["Pilot-scale capacity is scarce, so research teams compete for access.","Informal comparison lets each team choose protocols, samples, and disclosures that favor its candidate.","Selection pressure rewards peak-result presentation, extensive search, and postponed accounting of hazards or resource burdens.","A predeclared arena constrains eligibility, evidence, testing resources, allowed conduct, and noncontestable safety floors.","Random sample identification and blinded common-protocol testing reduce teams' control over which outcomes enter comparison.","Audited multi-criteria scoring makes reproducibility and scale-up burdens part of the route to qualification.","Resource caps dampen testing escalation, while chain-of-custody, conflict disclosure, and appeals constrain abuse and judging error.","Two provisional awards preserve competing pathways until pilot-scale evidence is available.","Post-contest review compares qualification predictions with observed scale-up behavior and recalibrates the next contest."],"baseline":"Facility leaders solicit slide decks and investigator-selected data, discuss candidates in committee, and grant the next scale-up slot to one team based on scientific promise, peak reported performance, and reviewer judgment. Safety review occurs after technical selection, testing effort is uncapped, negative runs need not be disclosed consistently, and unsuccessful teams lack a defined correction or appeal path.","nearest_rivals":["A common electrochemical test protocol without governed eligibility, resource caps, appeals, anti-abuse controls, or post-selection review.","A conventional scientific peer-review panel that evaluates dossiers but leaves sample selection and evidence disclosure to each team.","An environmental health and safety gate applied after technical ranking, which can reject hazards but does not govern metric gaming or testing escalation.","A first-come, first-served or rotating scale-up queue that is procedurally simple but does not use rivalry to compare readiness.","A single-winner prize challenge based mainly on peak performance, without portfolio selection or spillover accounting."],"remaining_contrastive_claim":"The candidate's distinctive contribution is to govern the entire rivalry for scale-up access—entry, permitted evidence, resource expenditure, common testing, safety floors, abuse controls, appeals, provisional prize structure, and recalibration—so that the winning strategy is coupled to reproducible and responsible scale readiness. Standardized testing or safety screening alone addresses only one part of that causal structure.","authority_safety":{"decision_authority":"The shared-facility program director may establish a pilot qualification procedure for facility-controlled slots, but environmental health and safety staff retain independent authority over hazard floors, an independent technical panel scores submissions, and the relevant institutional bodies retain authority over research conduct, procurement, and personnel matters.","authorized_first_step":"The program director may authorize a retrospective, nonbinding tabletop re-score using already-generated, de-identified candidate dossiers and existing test records; this step changes no access decision and requires no new synthesis or hazardous handling.","excluded_actions":["Starting new chemical synthesis or cell fabrication for the evidence step","Relaxing institutional safety, waste, ethics, procurement, or research-integrity requirements","Publicly ranking named teams during the pilot","Disclosing proprietary formulations to competing teams","Penalizing personnel or adjudicating research misconduct through the contest process","Changing already-awarded scale-up access based on the retrospective exercise","Treating a composite score as permission to compensate for failure of a safety floor"],"halt_rollback":"Halt the pilot if de-identification fails, required records are too incomplete for comparable scoring, the scoring procedure rewards a known unsafe route, reviewer conflicts cannot be separated, or participants identify credible retaliation or confidentiality risk. Roll back by discarding the provisional rankings, retaining the existing allocation process, documenting the failure, and allowing only aggregate methodological findings to proceed."},"negative_tests":{"strongest_counterevidence":"Under blinded common-protocol testing, rankings are stable across reasonable scoring weights, align with the existing committee's choices, and show no association with selective reporting, testing volume, safety burden, resource intensity, or reviewer identity.","problem_falsifier":"The inferred problem is falsified if scale-up access is not meaningfully scarce or rivalrous, submissions already include complete comparable run inventories and independent common-protocol tests, safety and spillovers are integrated before selection, and teams cannot improve their prospects through strategic effects on rivals, evidence, or judging.","intervention_falsifier":"The intervention is falsified if the governed re-score is no more stable or predictive of held-out scale-up-readiness indicators than the baseline judgment, or if teams can still improve rank through omitted failures, greater testing volume, protocol manipulation, or uncompensated hazardous and resource-intensive choices.","risks":["Composite scoring can create new gaming targets or conceal value judgments in weights.","Common protocols can disadvantage materials that require legitimately different operating conditions.","Disclosure requirements can expose proprietary information or encourage premature convergence around familiar chemistries.","Resource caps can burden small teams if fixed compliance costs are high.","Multiple provisional winners can divide scale-up capacity below a scientifically useful minimum.","Chain-of-custody and audit requirements can add technician workload and delay facility operation.","Incumbents may shape eligibility or metrics to preserve structural advantage.","A formal ranking can be misused as a general judgment of team quality rather than a bounded scale-up-readiness decision."]},"next_evidence_step":"Using no more than six de-identified historical candidate dossiers and their existing raw run inventories, have one facility scientist and one independent reviewer apply a draft rulebook and calculate rankings under three predeclared weight sets. Compare rank sensitivity, missing-data incidence, testing-volume effects, safety-floor conflicts, and agreement with any existing pilot-scale outcomes. Conduct the exercise within two reviewer-days, make no operational allocation change, and decide only whether a prospective protocol is sufficiently coherent to warrant separate approval.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"No other proposals or experiment cells were inspected. This candidate was derived solely from the supplied bounded-rivalry archetype and chemistry/materials domain card, so no cross-proposal diversity claim is made.","revision_record":{"parent_version":null,"progress_targets_addressed":["Initial one-shot construction from the supplied archetype and domain record"],"conceptual_changes":[],"operational_changes":[],"evidence_changes":[],"claim_changes":[]}}