{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp12_substrate_denial72_20260805","cell_id":"incentive_compatible_rule_design__chemistry_materials","arm":"ORDINARY_MAX","candidate_id":"incentive_compatible_rule_design__chemistry_materials__ORDINARY_MAX","proposal_index":1,"version":0,"title":"Random-Draw Readiness Gates for Reproducible Materials Batches","problem":"In a stage-gated materials-development program, a research team privately observes results from many synthesis batches but can submit its strongest coupon to clear a property threshold. When funding, status, or scale-up access depends on that submitted maximum, selecting an exceptional coupon and omitting weak sibling batches can be a rational strategy even though the program needs a fabrication process that performs reproducibly across batches.","actors":["Materials R&D team that chooses synthesis runs and observes their outcomes","Program stage-gate owner who grants readiness status and downstream resources","Independent characterization core that codes, selects, and measures specimens","Data steward who controls authorized access to batch and instrument records","Downstream scale-up team that relies on the readiness decision","Independent appeal reviewer for exclusions, deviations, and disputed measurements"],"observable_state":"At a gate based on team-selected coupons, the submitted coupon clears the target while recorded sibling batches from the same recipe are absent from the gate dossier, independently selected coupons show a different or wider property distribution, or the pass decision changes materially depending on who selects the specimen. These observations must be distinguished from instrument drift, sample aging, and legitimate process changes.","consequence":"A formulation can advance because of an exceptional specimen rather than a reproducible synthesis window, causing downstream teams to spend time and material on scale-up candidates that may not reliably satisfy the required property envelope.","affected_objective":"Select material formulations whose specified synthesis and processing window reproducibly yields the required multi-property distribution, rather than merely producing one exceptional coupon.","intervention":"Replace the best-coupon readiness rule with two explicit tracks. The exploratory track permits flexible iteration and retains ordinary research support but cannot confer readiness. A team seeking readiness enters a qualification track and, before characterization results are known, commits the formulation identity, process window, fixed batch plan, stopping rule, valid exclusion criteria, and specimen identifiers. An independent core then uses a concealed randomization seed to select coded specimens across committed batches. Gate credit is determined by a predeclared distribution score—such as the fraction of batches inside the property envelope together with a variability constraint—not by the maximum result. The team receives fixed reimbursement for completing a valid qualification run whether it passes or fails, reducing the cost of honestly discovering a weak process. Deliberate omission, relabeling, or post-result substitution can remove qualification credit only after a proportionate record audit and an independent appeal. A failed qualification returns the work to exploratory status rather than triggering employment, publication, or misconduct sanctions.","structural_mapping":[{"archetype_element":"Strategic participants and private information","domain_realization":"The synthesis team sees all attempts, deviations, and preliminary measurements, while the gate owner initially sees only the submitted coupon and dossier."},{"archetype_element":"Exploitable rule","domain_realization":"Readiness is awarded for exceeding a threshold with a team-selected maximum, so the formal success condition can be met without demonstrating cross-batch reproducibility."},{"archetype_element":"Desired outcome specification","domain_realization":"The real target is a bounded process window that repeatedly produces material within a multi-property envelope."},{"archetype_element":"Action and choice set","domain_realization":"The team can remain exploratory, enter qualification, disclose all committed batches, cherry-pick a coupon, stop selectively, relabel specimens, or create undeclared qualification runs."},{"archetype_element":"Incentive payoff map","domain_realization":"Readiness brings resources and status; valid completion is reimbursed even after failure; only distribution-qualified work advances; evidenced intentional substitution risks loss of gate credit."},{"archetype_element":"Information structure map","domain_realization":"Pre-result commitments, coded specimens, concealed random selection, and limited reconciliation with existing batch records reduce the team's ability to condition inclusion on observed performance."},{"archetype_element":"Truthfulness condition","domain_realization":"Complete execution of the committed plan becomes preferable when honest failure retains reimbursement and exploratory access, while omission offers no sample-selection advantage and creates a bounded risk of losing qualification credit."},{"archetype_element":"Verification rule","domain_realization":"The independent core checks specimen identity, randomization, cross-batch measurements, declared exclusions, and a proportionate sample of already-authorized batch or instrument records."},{"archetype_element":"Participation and fairness constraints","domain_realization":"Teams may remain exploratory without penalty; qualification batch counts and reimbursements are fixed; legitimate deviations, measurement faults, and safety interruptions have documented exception and appeal paths."},{"archetype_element":"Strategic response test and gaming monitor","domain_realization":"Shadow analyses test undeclared batches, strategic process-window definitions, relabeling, selective stopping, metric substitution, collusion, and disproportionate burden before the rule affects decisions."}],"mechanism_mapping":[{"mechanism_slug":"anti_gaming_scoring_rule","role":"Replaces the reward for a selected maximum with a preregistered cross-batch distribution score tied to the actual reproducibility objective.","counterfactual_removal":"If the maximum-result score remains, producing or finding one exceptional coupon can still dominate stabilizing the synthesis process."},{"mechanism_slug":"blind_or_randomized_review_rule","role":"Transfers specimen selection to an independent core after batch commitment and conceals the selection seed until identifiers are locked.","counterfactual_removal":"If the team retains post-result specimen choice, it can satisfy the new reporting language while continuing to submit unusually strong coupons."},{"mechanism_slug":"self_selection_menu","role":"Separates flexible exploration from a readiness-bearing qualification track, allowing uncertainty to be disclosed through track choice without ending legitimate research participation.","counterfactual_removal":"If every exploratory run bears qualification burdens, early-stage teams may exit, delay documentation, or create shadow work; if no qualification track exists, readiness remains unsupported."},{"mechanism_slug":"audit_and_penalty_system","role":"Uses limited record reconciliation and loss of qualification credit, subject to evidence and appeal, to make deliberate omission or substitution unattractive without continuous surveillance.","counterfactual_removal":"Random selection among declared batches alone leaves a profitable path through undeclared runs, relabeling, or selective stopping."}],"causal_chain":["The existing gate ties advancement to a team-selected maximum property value.","The team privately observes a larger set of batch outcomes and controls which coupon enters the dossier.","Because weak disclosed outcomes can threaten advancement while an exceptional coupon can secure it, selective inclusion can become the team's best response.","The two-track menu lets uncertain work remain exploratory while requiring an affirmative commitment before making a readiness claim.","Pre-result batch commitments and independent random selection remove most post-result control over the qualification sample.","Distribution-based scoring redirects the readiness payoff from exceptional specimens toward repeatable process performance.","Completion reimbursement reduces the private penalty for an honestly reported failed qualification.","Proportionate audits, bounded loss of gate credit, and appeal change the expected payoff of intentional omission while protecting measurement errors and legitimate deviations.","If these assumptions hold, concentrating effort on a stable process and submitting the complete committed run becomes more attractive than searching for a gate-clearing hero coupon.","The resulting gate record can then provide a more direct test of cross-batch reproducibility for downstream scale-up decisions."],"baseline":"The program specifies a property threshold, permits the team to choose the tested or submitted coupon, and advances the material when that coupon passes. Replicate counts, failed-run reporting, selection timing, and exclusion rules are informal or reviewed only after a discrepancy appears.","nearest_rivals":["Require a larger replicate count and electronic capture of all batch and instrument records while retaining the existing threshold; this could solve the problem if capture is complete and specimen selection no longer matters, but recordkeeping alone does not change the advancement payoff.","Commission fully independent resynthesis before every gate; this directly tests transferability but may require substantially more material and time and does not by itself align the originating team's disclosure incentives.","Standardize synthesis procedures, calibrate instruments, and conduct measurement-system analysis; these are stronger explanations and remedies when variability comes from capability or metrology rather than strategic sample selection.","Use random specimen selection without changing the maximum-based score or failure consequences; this removes some selection control but can produce a noisy single-draw gate and preserves incentives to hide runs upstream.","Increase audits and punish omitted data under the existing gate; this may deter concealment but relies on enforcement, can intensify adversarial behavior, and offers no low-cost path for honest qualification failure."],"remaining_contrastive_claim":"The claim is conditional: where discretionary post-result specimen selection and a maximum-based advancement reward cause the reporting gap, coupling readiness to precommitted, independently sampled cross-batch performance should change the team's best response in a way that extra team-selected replicates, metrology improvements, or audit-only enforcement do not. No advantage is claimed if sample selection is not causal, complete capture already removes discretion, or the program's legitimate objective is discovery of any attainable extreme rather than reproducible fabrication.","authority_safety":{"decision_authority":"Only the designated materials-program stage-gate owner may authorize a study, with approval from the data steward and laboratory-safety authority; the independent characterization lead controls coding and randomization, and an uninvolved reviewer controls appeals.","authorized_first_step":"Authorize only a de-identified retrospective shadow analysis of already-completed gates and already-authorized records. Shadow scores must not alter any past or current decision.","excluded_actions":["Using shadow scores for funding, employment, authorship, publication, intellectual-property, or misconduct decisions","Imposing financial bonds, fines, withheld wages, or personal penalties","Compelling new synthesis, destructive testing, or handling outside approved safety procedures","Inspecting notebooks, instrument records, or communications outside existing consent and access authority","Publicly ranking, blacklisting, or accusing teams or individuals","Changing material release, safety acceptance, or scale-up status from the shadow analysis","Deploying covert or continuous surveillance","Removing qualification credit without documented evidence, notice, and independent appeal"],"halt_rollback":"Stop if records cannot be de-identified, batch inclusion cannot be reconstructed, participation or confidentiality terms are breached, or analysis is used for an adverse decision. Quarantine the shadow results, remove person-level linkages, notify the responsible data steward, preserve correction and appeal rights, and leave the existing gate status unchanged."},"negative_tests":{"strongest_counterevidence":"The diagnosis is undermined if team-selected and independently selected coupons have comparable distributions after controlling for instrument, operator, batch age, and declared process changes; complete batch capture is already routine; and the selected maximum predicts downstream resynthesis at least as well as the proposed distribution score.","problem_falsifier":"The problem is absent if participants cannot influence specimen inclusion after observing results, omissions cannot improve advancement, or the legitimate scientific objective is to demonstrate that an exceptional material state is physically attainable rather than that a process is reproducible.","intervention_falsifier":"The intervention fails if a prospective non-decisional pilot later shows that precommitment and random selection do not reduce selector-dependent gate outcomes or correspondence with independent resynthesis, or if they instead increase undeclared work, strategic relabeling, invalid exclusions, withdrawal by legitimate teams, or unresolvable measurement disputes.","risks":["Gaming may move upstream to formulation boundaries, batch declarations, stopping rules, or qualification timing.","Teams may optimize the declared property score while neglecting durability, toxicity, manufacturability, or other unscored properties.","Resource-constrained teams and materials with intrinsically stochastic early-stage synthesis may face disproportionate qualification burdens.","Independent measurement error, sample aging, contamination, or chain-of-custody failures could falsely imply irreproducibility or conceal it.","Qualification records may expose confidential formulations or process know-how.","The rule could chill exploratory science if exploratory status is stigmatized despite the formal participation protection.","Audits could become punitive surveillance or generate unsupported allegations.","Core-facility staff could become a bottleneck or collude in coding, selection, or exclusions.","Randomly selected specimens may require additional safety review when composition or handling risk varies across batches."]},"next_evidence_step":"For one material family, reconstruct at most 12 completed gate events using only existing authorized records. Before examining downstream outcomes, specify how to identify committed-equivalent batches, simulate independent random specimen draws, and compute the proposed distribution score. Compare the existing selected-coupon decision with the shadow decisions, documented missingness, selector dependence, and any already-scheduled downstream resynthesis result; separately record reconstruction uncertainty and staff burden. If batch capture is too incomplete, the selector gap is not observable, or the score merely tracks metrology artifacts, stop. This bounded step can justify—but cannot replace—a later prospective, non-decisional pilot.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Not assessed against prior proposals under the required runtime isolation. Within this record, the distinctive domain realization is a materials-R&D readiness gate in which synthesis batches and coupons are strategically selectable units and advancement is reassigned from an exceptional maximum to independently sampled cross-batch performance.","revision_record":{"parent_version":null,"progress_targets_addressed":["Initial complete candidate with a concrete materials-science problem, strategic actors, observable diagnostics, causal mapping, operational mechanisms, serious rivals, authority limits, falsifiers, risks, and a bounded evidence step."],"conceptual_changes":["Initial formulation; no parent version.","Separated the legitimate objective of reproducible fabrication from the distinct objective of demonstrating an attainable extreme material state."],"operational_changes":["Specified exploratory and qualification tracks, pre-result batch commitments, concealed random selection, distribution-based scoring, completion reimbursement, proportionate audit, and independent appeal."],"evidence_changes":["Limited the first evidence step to a de-identified retrospective shadow analysis of at most 12 completed gates from one material family, with no decision use."],"claim_changes":["Restricted the contrastive claim to a conditional best-response prediction and made no novelty, prevalence, demand, or effect-size claim."]}}