{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp09_archetype_breadth150_20260804","cell_id":"prediction_error_learning_calibration__chemistry_materials","arm":"BREADTH_PROBE_ONE_SHOT","candidate_id":"prediction_error_learning_calibration__chemistry_materials__P1","proposal_index":1,"version":0,"title":"Prediction-Error-Gated Learning for Solid-State Electrolyte Screening","problem":"In a recurring solid-state electrolyte formulation campaign, candidate compositions and processing recipes are promoted or abandoned according to raw measured room-temperature ionic conductivity. The team and its predictive model estimate performance before synthesis, but those estimates are not frozen and compared with outcomes. Consequently, a high-conductivity result can drive a large recipe update even when it was expected, while a modest result that strongly contradicts the composition model may receive little investigation.","actors":["Materials-science principal investigator who owns the screening strategy","Formulation scientist selecting compositions and processing variables","Laboratory automation operator executing synthesis and pellet preparation","Electrochemical metrology scientist measuring ionic conductivity","Data scientist maintaining the composition–process performance model","Laboratory safety officer enforcing material and process constraints"],"observable_state":"For each batch, the laboratory has a composition, milling and heat-treatment settings, pellet density, environmental metadata, and a conductivity result, but no immutable pre-synthesis prediction record. Queue changes correlate with raw conductivity rank; replicate disagreements, reference-pellet drift, and deviations from the model's expectation are discussed inconsistently rather than converted into an explicit learning signal.","consequence":"The screening trajectory can chase lucky measurements, repeatedly reinforce already-predicted successes, discard informative negative surprises, and attribute a deviation to composition when pellet preparation, humidity, instrument drift, or another contextual factor is responsible. Experimental capacity is then spent on updates that do not reliably correct the campaign's predictive model.","affected_objective":"Reliable selection of composition and processing changes that improve out-of-sample prediction and decision quality for room-temperature ionic conductivity, subject to fixed electrochemical-stability, handling, and laboratory-safety constraints.","intervention":"Insert a prediction-error calibration layer between conductivity measurement and recipe promotion. Before each approved synthesis, freeze the predicted log conductivity and uncertainty interval under a specified temperature, pellet-density, and measurement protocol. After measurement, record the received value and compute the signed prediction–outcome delta. Use replicate agreement, reference-pellet drift, and process metadata to classify the delta as likely model error, process variation, metrology error, context change, or unresolved noise. Assign any update to a named model feature or process assumption, then scale the learning gain: reliable positive or negative surprises receive review and bounded model updates, while expected or poorly attributable outcomes mainly maintain the current model. Keep safety and stability constraints as non-tradeable gates rather than allowing conductivity surprise to override them.","structural_mapping":[{"archetype_element":"Prior Prediction Record","domain_realization":"An immutable pre-synthesis record of predicted log ionic conductivity, uncertainty, composition, process settings, and measurement conditions for each electrolyte batch."},{"archetype_element":"Value Reference Frame","domain_realization":"Room-temperature log ionic conductivity measured under a fixed protocol; electrochemical stability, chemical compatibility, and handling safety remain separate eligibility constraints."},{"archetype_element":"Received Outcome Record","domain_realization":"Measured conductivity with replicate values, pellet density, impedance-fit diagnostics, reference-pellet result, instrument identifier, and environmental metadata."},{"archetype_element":"Signed Error Signal","domain_realization":"Measured minus predicted log conductivity, preserving whether the material performed better or worse than expected."},{"archetype_element":"Credit Assignment Window","domain_realization":"A batch-level review linking the deviation to composition descriptors, heat treatment, milling, pellet preparation, atmosphere, measurement execution, or unresolved variation."},{"archetype_element":"Noise and Volatility Filter","domain_realization":"Replicate consistency, reference-pellet drift, fit quality, and recent environmental or equipment changes determine whether the deviation is eligible to teach."},{"archetype_element":"Learning Gain Rule","domain_realization":"Model or queue changes are larger only for reproducible, well-attributed deviations and are bounded or withheld for isolated, ambiguous, or drift-contaminated results."},{"archetype_element":"Update Target","domain_realization":"A named composition–property relation, process-effect estimate, model parameter, confidence interval, or experimental priority—not an undifferentiated judgment about the material."},{"archetype_element":"Shortcut Learning Guard","domain_realization":"Matched holdout batches break incidental correlations between composition families and press, furnace, operator, dry-room slot, or pellet-density conditions."},{"archetype_element":"Feedback Timing Alignment","domain_realization":"Predictions are frozen immediately before synthesis and deltas are reviewed after validated measurement but before the next queue revision."}],"mechanism_mapping":[{"mechanism_slug":"prediction_outcome_delta_log","role":"Stores the frozen prediction, received conductivity, signed residual, context, attribution, and resulting update decision for every batch.","counterfactual_removal":"Without the log, expectations can be rewritten after measurement and raw conductivity again becomes indistinguishable from instructional surprise."},{"mechanism_slug":"credit_assignment_trace","role":"Documents which composition feature, processing step, measurement condition, or contextual change is considered responsible for an eligible residual and why.","counterfactual_removal":"Without the trace, a genuine surprise can update the nearest visible variable rather than its probable cause."},{"mechanism_slug":"learning_rate_schedule","role":"Bounds update strength according to replicate support, measurement quality, attribution confidence, context stability, and safety sensitivity.","counterfactual_removal":"Without gain control, single lucky or unlucky pellets can produce unstable swings in the model and experimental queue."},{"mechanism_slug":"calibration_curve_review","role":"Periodically checks whether predicted conductivity distributions are systematically high, low, or overconfident across composition and process strata.","counterfactual_removal":"Without calibration review, individual deltas may be processed while persistent bias or misestimated uncertainty remains invisible."},{"mechanism_slug":"shortcut_probe_holdout_set","role":"Tests whether apparent learning survives batches in which incidental laboratory cues are decoupled from composition families.","counterfactual_removal":"Without the holdout, the model may appear to improve by learning furnace, operator, density, or scheduling proxies that fail under changed laboratory conditions."}],"causal_chain":["The campaign repeatedly chooses compositions and processing recipes whose measured conductivity influences future choices.","Pre-synthesis performance expectations exist but are not frozen as decision inputs.","Raw high and low conductivity values therefore act as teaching signals regardless of how expected they were.","Batch noise, pellet preparation, instrument drift, and incidental laboratory correlations can then receive incorrect credit.","The intervention records the prior prediction and computes a signed prediction–outcome delta under a fixed value frame.","Replicates, controls, contextual metadata, and a credit-assignment trace filter unreliable or misattributed deltas.","A bounded learning-gain rule updates only the relevant model component or queue priority in proportion to reliable surprise.","Subsequent calibration and shortcut-holdout checks test whether the update reduced repeated prediction error without learning incidental cues or crossing safety gates."],"baseline":"The current comparison condition is raw-outcome screening: rank candidates by validated measured conductivity, use ordinary replicate discussion and scientist judgment to resolve anomalies, retrain the model on accumulated outcomes, and revise the next synthesis queue without a mandatory frozen prediction, signed residual, attribution record, or residual-dependent learning gain.","nearest_rivals":["Bayesian optimization that selects experiments from posterior mean, uncertainty, and an acquisition function but does not require a batch-level signed residual and attribution gate before learning from each outcome.","Factorial design or response-surface modeling that estimates composition and process effects from planned contrasts but does not make surprise relative to each frozen prediction the update trigger.","Statistical process control using reference materials and control charts to detect laboratory drift but not to decide which scientific model component should learn from a signed conductivity deviation.","Replicate-and-review practice that repeats anomalous measurements but leaves the magnitude, sign, attribution, and gain of the subsequent model update implicit."],"remaining_contrastive_claim":"The candidate's distinguishing proposition is that an electrolyte result should change the scientific model and screening queue according to its reliable signed deviation from a frozen expectation, after explicit credit and noise checks—not according to raw conductivity rank, anomaly status, or posterior ingestion alone. Expected successes are maintenance evidence; reproducible positive and negative surprises are the primary candidates for targeted learning.","authority_safety":{"decision_authority":"The materials-science principal investigator may authorize changes to the screening model and experimental queue; the metrology lead validates measurements, and the laboratory safety officer retains independent veto authority over compositions and processing conditions.","authorized_first_step":"Run the layer in shadow mode on the next 24 already approved, non-hazardous screening batches: freeze predictions, calculate deltas, complete attribution and filter records, and simulate update decisions without changing the live recipe queue.","excluded_actions":["Autonomous changes to synthesis recipes, furnace programs, or experimental ordering during the shadow test","Introduction of compositions or reagents outside existing environmental, health, and safety approvals","Relaxation of electrochemical-stability, compatibility, waste-handling, or exposure constraints in response to positive conductivity surprise","Use of the records to score, reward, punish, or rank individual laboratory personnel","Backfilling predictions after conductivity results are visible","Suppressing failed, ambiguous, or constraint-violating measurements from the delta log"],"halt_rollback":"Stop the trial and revert to the existing review process if predictions are repeatedly entered after outcomes become inferable, metrology cannot produce comparable measurements, attribution discussions become personnel-blaming, safety gates are treated as reward tradeoffs, or shadow recommendations cannot be reconstructed from the recorded rule. Preserve the audit log but apply none of its simulated queue changes."},"negative_tests":{"strongest_counterevidence":"An audit could show that the existing optimizer and review process already freeze pre-run predictive distributions, standardize residuals against measurement uncertainty, use reference and replicate filters, trace deviations to named composition or process features, and bound updates accordingly. That would make the proposed layer a relabeling rather than a distinct intervention.","problem_falsifier":"The problem is falsified if blinded reconstruction of recent queue decisions shows that raw conductivity rank did not govern promotion, decisions already followed reproducible expectation-adjusted deviations, and apparent oscillations are explained by deliberate exploration or changed safety and stability constraints rather than noisy learning.","intervention_falsifier":"The intervention is falsified for this setting if, during the bounded shadow test, residual-gated simulated updates yield no better one-step-ahead calibration than the baseline retraining rule, flagged surprises do not reproduce when repeated, or attribution and filtering systematically direct updates away from the variables supported by matched controls.","risks":["Scientists may sandbag predictions to manufacture positive surprises or avoid negative-error review.","A narrow conductivity value frame may accelerate learning toward materials that fail stability, compatibility, manufacturability, or safety requirements.","Small sample strata may make signed residuals appear systematic when they reflect sparse noise.","Credit assignment may favor readily measured process variables and miss latent chemical mechanisms.","Frequent negative-error reviews may create blame or discourage reporting of anomalous results.","Holdout batches and extra replicates consume synthesis and metrology capacity.","Instrument or protocol changes may be mistaken for material-model error.","Overly conservative gain rules may suppress useful adaptation; aggressive rules may amplify noise."]},"next_evidence_step":"For the next 24 already scheduled and safety-approved batches, freeze model and scientist predictions before synthesis and collect the complete delta log in shadow mode. Predeclare comparison against the existing retraining workflow using one-step-ahead calibration, reproducibility of flagged positive and negative surprises, attribution agreement with matched controls, and constraint-violation counts. At the end of the 24 batches, permit consideration of a limited live pilot only if the records are prospective and complete, the flagged errors reproduce, and simulated residual-gated updates improve calibration without weakening any safety or stability gate.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Not assessed; this one-shot candidate was generated solely from the supplied archetype record and domain card without inspecting other proposals or experiments.","revision_record":{"parent_version":null,"progress_targets_addressed":[],"conceptual_changes":[],"operational_changes":[],"evidence_changes":[],"claim_changes":[]}}