{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","cell_id":"predictive_residual_processing__chemistry_materials","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_id":"governed_residual_reference_potential","proposal_index":2,"version":0,"title":"Governed Residual Correction for Reference-Level Materials Calculations","problem":"A materials simulation may need energies, atomic forces, and stresses at many configurations while its designated reference calculation is available for only a bounded set of queries. Using a cheaper physical model everywhere can carry systematic configuration-dependent error, while learning the complete reference mapping from scratch can discard useful baseline structure and still become confidently wrong outside its sampled configurations. The resulting uncertainty can limit which simulated transformations or property estimates are defensible.","actors":["Computational materials scientist","Electronic-structure method owner","Baseline-model owner","Residual-model owner","High-performance computing scheduler","Downstream simulation owner","Independent numerical-validation reviewer"],"observable_state":"For each atomic configuration, the system records species, geometry, cell, charge and electronic-state assumptions, provenance, baseline-model energy, forces and stress, baseline uncertainty, predicted residual correction, corrected output, residual-model version and scope status. At selected reference queries it also records the complete converged reference output, convergence diagnostics, exact signed reference-minus-baseline residual, reconstruction check, precision weight, update decision, audit status and any fallback or halt condition.","consequence":"Reference calculations may be spent repeatedly measuring discrepancies the correction model already represents, or a cheaper evaluator may be trusted where its unobserved error changes a scientifically relevant energy ordering, force direction, stress response or trajectory. Without independent reference audits, the correction model can appear accurate because it determines where reference evidence is requested.","affected_objective":"Support bounded use of a corrected energy-force-stress evaluator within a declared reference-calculation budget while preserving reference reconstruction, uncertainty calibration, scientific invariants, independent audits and immediate return to the complete reference method when predictive assumptions fail.","intervention":"Treat a versioned lower-cost physical model as the expected reference output and learn only its signed discrepancy from a fixed high-fidelity reference method. At each authorized reference configuration, freeze the baseline energy, force and stress prediction before the reference result is revealed; compute the complete reference result; then form component-wise reference-minus-baseline residuals with convergence provenance. A precision-weighted Bayesian residual model learns how that discrepancy varies across the explicitly bounded configuration domain. Within validated scope, the downstream evaluator returns baseline plus predicted residual and exposes uncertainty and model identity. High uncertainty, cross-model disagreement, invariant failure, unfamiliar local environments or protected configuration classes trigger an actual reference query or prohibit use. Random reference queries and risk-stratified full-reference challenges proceed independently of the model's uncertainty, and their complete outputs are compared with reconstructed values. Version incompatibility, residual drift, audit failure or cumulative error suspends corrected evaluation and restores the full reference method or halts the affected calculation. Residual-model updates occur in reviewed batches and cannot silently alter an active simulation.","structural_mapping":[{"archetype_element":"Explicit prediction target and observation boundary","domain_realization":"The target is the energy, atomic-force vector and cell stress that one fixed reference method would return for a configuration under declared convergence, charge, spin and boundary assumptions; a completed reference calculation is the actual observation."},{"archetype_element":"Maintained generative model with bounded scope","domain_realization":"The expected output is generated by a versioned lower-cost physical model, while a separate uncertainty-tagged residual model represents its reference-method discrepancy only within approved elements, phases, local environments and thermodynamic ranges."},{"archetype_element":"Prediction before observation","domain_realization":"Baseline outputs and the current predicted correction are frozen before the corresponding reference result is revealed."},{"archetype_element":"Signed structured comparator","domain_realization":"The comparator preserves signed energy, per-atom force and stress differences, units, convergence diagnostics and configuration provenance rather than reducing disagreement to one accuracy score."},{"archetype_element":"Precision-weighted prediction error","domain_realization":"Residual labels are weighted by reference convergence, numerical noise, configuration reliability and the consequence of error in the relevant energy, force or stress component."},{"archetype_element":"Residual propagation channel","domain_realization":"Reference workers return discrepancy records and provenance to the correction learner; the unchanged baseline output need not become the learned target because the receiver can regenerate it from the versioned baseline model."},{"archetype_element":"Reconstruction of the reference output","domain_realization":"For an audited reference query, the receiver reconstructs the complete reference result as the compatible baseline output plus the transmitted residual and verifies it against the full reference record."},{"archetype_element":"Residual-driven update rule","domain_realization":"Accepted residuals revise the posterior correction model and its uncertainty during a slower reviewed update cycle; failed or intervention-contaminated calculations do not enter automatically."},{"archetype_element":"Synchronization and provenance","domain_realization":"Every residual names the baseline implementation, parameters, reference method, convergence recipe, preprocessing convention, configuration identity and residual-model version."},{"archetype_element":"Freshness and drift controls","domain_realization":"Residual calibration, local-environment coverage, model age and realized audit errors determine whether corrected evaluation remains valid."},{"archetype_element":"Independent raw audit","domain_realization":"Random configurations and predeclared high-consequence configuration classes receive complete reference calculations irrespective of predicted residual magnitude or uncertainty."},{"archetype_element":"Fallback and protected bypass","domain_realization":"Out-of-scope chemistry, uncertain electronic state, nonconvergence, invariant violations, strong model disagreement or audit failure bypasses the correction path and invokes the complete reference method or halts use."},{"archetype_element":"Explicit capacity and residual-error budget","domain_realization":"Acceptance counts reference queries, baseline and residual-model computation, audits, fallbacks, review work, component-wise reconstruction error and cumulative downstream error rather than relying on average test error alone."}],"mechanism_mapping":[{"mechanism_slug":"bayesian_model_update","role":"Represents uncertainty over the reference-minus-baseline correction and revises that belief from accepted residual observations.","counterfactual_removal":"Without uncertainty-aware residual updating, the correction would become a static offset or an uncalibrated regression rather than an adaptive prediction-error model."},{"mechanism_slug":"precision_weighted_error_gate","role":"Weights each residual by reference convergence, numerical reliability, component consequence, configuration coverage and update cost.","counterfactual_removal":"Without precision weighting, noisy or poorly converged reference results could dominate reliable small corrections in consequential force or energy components."},{"mechanism_slug":"confidence_threshold_table","role":"Maps uncertainty, disagreement, scope and consequence classes to corrected use, human review, reference query or prohibited-use actions.","counterfactual_removal":"Without the table, reference-query and fallback decisions could be tuned to computational convenience rather than an explicit error budget."},{"mechanism_slug":"forecast_backtesting","role":"Uses held-out configurations and regime slices to define where the residual evaluator may substitute for a reference query.","counterfactual_removal":"Without strictly separated testing, fit to queried configurations could be mistaken for permission to use the correction elsewhere."},{"mechanism_slug":"residual_comparison_test","role":"Tests residuals for systematic dependence on composition, local environment, strain, force magnitude and configuration source, and compares them with a direct-reference surrogate.","counterfactual_removal":"Without residual-structure comparison, the workflow could assume the discrepancy is simpler or more stable than the full target without evidence."},{"mechanism_slug":"anomaly_detection_model","role":"Identifies configurations or residual patterns outside the validated correction distribution and routes them to reference calculation or review.","counterfactual_removal":"Without an out-of-envelope detector, a numerically ordinary uncertainty score could permit use on unfamiliar chemistry or geometry."},{"mechanism_slug":"model_version_checksum_handshake","role":"Ensures reference workers, the learner and downstream evaluator use compatible baseline parameters, reference definitions and residual conventions.","counterfactual_removal":"Without version gating, a residual computed against one baseline could be applied to another and produce a well-formed but incorrect corrected output."},{"mechanism_slug":"prediction_error_replay_buffer","role":"Stores residuals with complete configuration and convergence context for regression testing, calibration and model-revision review.","counterfactual_removal":"Without replay, later correction versions could not be challenged consistently on earlier failures and boundary cases."},{"mechanism_slug":"shadow_raw_channel_sampling","role":"Selects random and risk-stratified configurations for complete reference evaluation independently of correction confidence.","counterfactual_removal":"Without independent reference sampling, the residual model would largely choose the evidence used to validate its own omissions."},{"mechanism_slug":"model_drift_monitoring","role":"Tracks changes in queried configuration distributions, residual calibration, local-environment coverage and realized reference error.","counterfactual_removal":"Without drift monitoring, a corrected simulation could move into a new regime while retaining confidence inherited from its starting configurations."},{"mechanism_slug":"raw_signal_fallback_switch","role":"Disables corrected evaluation and invokes complete reference calculation or halt when validity, synchronization, convergence or invariant conditions fail.","counterfactual_removal":"Without fallback, the approximate evaluator could remain active precisely where no defensible residual prediction exists."},{"mechanism_slug":"prediction_error_review","role":"Requires experts to classify material residuals as baseline-model deficiency, reference-convergence fault, scope failure, configuration error or learnable correction before updating.","counterfactual_removal":"Without review, failed reference calculations or malformed configurations could be learned as physical discrepancy."}],"causal_chain":["A versioned lower-cost physical model produces the expected reference energy, forces and stress for a configuration.","Within authorized scope, a residual model predicts the signed discrepancy and uncertainty needed to correct that baseline.","Selected reference calculations reveal the actual complete target, allowing an exact, provenance-tagged reference-minus-baseline residual to be computed.","Precision and convergence weighting determine whether that residual may teach the correction model, requires review or invalidates the configuration.","Validated residuals update the correction posterior, concentrating representation on what the baseline physics does not explain.","For later in-scope configurations, baseline plus predicted residual supplies the candidate corrected output and its uncertainty to the downstream calculation.","Random reference queries, protected challenge classes and rival-model comparisons reveal confident omissions and measure reconstruction error independently.","Uncertainty, disagreement, drift, invariant failure, version mismatch or audit error restores complete reference evaluation or halts the affected use before correction resumes."],"baseline":"Use the designated reference method for every configuration requiring energy, force or stress, with no residual surrogate; if the reference-query budget is exhausted, either stop or use the lower-cost baseline without learned correction. The comparison must count all reference and baseline computations, model training, uncertainty estimation, audits, fallback queries, failed convergence, expert review and downstream error checks.","nearest_rivals":["A direct surrogate trained to predict complete reference energies, forces and stresses, which does not preserve a versioned physical baseline as the expected component or make its signed discrepancy the learned signal.","An uncorrected lower-cost physical model, which supplies an expectation but has no residual learning, uncertainty-governed reference queries or independent audits.","A fixed empirical offset or linear calibration, which corrects a predetermined bias without a scoped adaptive residual distribution and decompression path.","Ordinary active learning based only on surrogate uncertainty, which selects reference configurations but need not reconstruct results from baseline plus discrepancy or test independent raw-reference samples.","A composite calculation with fixed hand-designed corrections, which does not update from observed prediction errors under versioned residual governance.","Running the complete reference method on every configuration, which preserves the target directly but does not use predictable baseline structure to govern the bounded reference-query path."],"remaining_contrastive_claim":"The proposal is a governed residual physics evaluator: a declared physical baseline supplies the expected reference output, a learned signed discrepancy supplies the correction and teaching signal, and complete reference calculations remain an independent oracle, audit source and fallback. Its load-bearing causal path is numerical baseline-plus-residual correction under uncertainty, not scientific-stream compression, campaign attention triage, command-conditioned sensory cancellation or an ungoverned surrogate.","authority_safety":{"decision_authority":"The reference-method owner defines the authoritative calculation and convergence rules. The computational materials scientist approves chemical scope and scientific use. The residual-model owner may propose updates but cannot activate them unilaterally. The downstream simulation owner decides whether corrected outputs may influence a calculation. The system may request reference queries or halt use but may not redefine the reference method, expand scope or certify a result.","authorized_first_step":"Conduct an offline shadow evaluation on one bounded material system and configuration domain. Compute and retain the complete reference result for every pilot configuration, keep those results authoritative, and prohibit corrected outputs from controlling production simulations or supporting certification claims.","excluded_actions":["Replacing authoritative reference calculations in safety, release, certification or publication-critical decisions during the first evidence step","Changing the reference method, convergence recipe, baseline model or residual target during scoring","Allowing the residual model to expand automatically to new elements, charge states, phases, bonding regimes or thermodynamic ranges","Using corrected forces after an invariant, scope, uncertainty or model-disagreement trigger","Learning from unconverged, provenance-incomplete or intervention-contaminated reference calculations","Deleting complete pilot reference outputs after residual construction","Activating a model update without regression tests and recorded method-owner approval","Interpreting absence of a reference query as confirmation that the corrected output is accurate"],"halt_rollback":"Immediately disable the corrected evaluator for the affected scope and return to complete reference calculation or halt if versions mismatch, reference convergence is invalid, uncertainty or disagreement crosses its tabled bound, energy-force consistency or another protected invariant fails, the configuration leaves scope, residual calibration drifts, or an independent audit exceeds a component-wise or cumulative error budget. Preserve the configuration, baseline output, predicted correction, full reference result when available and all model versions; roll back to the last approved correction model only after method-owner and simulation-owner review."},"negative_tests":{"strongest_counterevidence":"Independent random reference calculations reveal configurations where the corrected evaluator is confidently wrong in an energy ordering, force direction or stress response that changes the downstream scientific interpretation, while all scope and validity checks remain nominal.","problem_falsifier":"The complete reference method fits within the declared computational and review budget for every required configuration, or the uncorrected baseline already satisfies all predeclared downstream tolerances; a residual correction layer would then solve no binding problem.","intervention_falsifier":"At an equal bounded set of reference observations, the residual formulation fails its predeclared component-wise, calibration, invariant and downstream-decision criteria; requires fallback so often that its governed cost exceeds the accepted budget; or provides no defensible advantage over a direct-reference surrogate or the uncorrected baseline.","risks":["The baseline and correction model can share a blind spot that makes their combined output confidently wrong.","The reference-minus-baseline discrepancy may be less regular than the complete reference mapping in a boundary regime.","Reference convergence noise can be learned as a physical correction.","Separate energy and force corrections can violate gradient consistency and yield unstable dynamics.","Average error cancellation can conceal incorrect energy ordering or force direction in a small consequential region.","A corrected simulation generates its own future configurations, creating distribution shift that was absent from the training set.","Uncertainty estimates can be overconfident outside sampled local environments.","Random audits may miss rare reaction pathways, while risk-stratified audits may encode only anticipated failures.","A compatible checksum proves model agreement, not agreement with the reference physics.","Reference-query thresholds may be relaxed to meet a compute target rather than an error target.","Residual records may omit electronic-state or convergence context needed to reproduce the reference comparison."]},"next_evidence_step":"Pre-register an offline comparison for one material composition, one bounded structural regime, one baseline method, one reference method and convergence recipe, and a finite configuration set spanning approved equilibrium perturbations, strain states, local coordination changes and scope-boundary challenges. Compute and retain complete reference outputs for every configuration, but simulate a fixed reference-query budget for training and auditing. Freeze data partitions, residual definitions, invariants, uncertainty method, confidence table and acceptance rules before fitting. Compare the uncorrected baseline, the residual-corrected evaluator and a direct-reference surrogate trained on the same queried configurations. On untouched configurations, measure component-wise energy, force and stress reconstruction, uncertainty calibration, energy ordering, force-direction agreement, invariant violations, residual structure, fallback decisions and total governed computation. Include forced checksum, convergence, scope and missing-data failures. Reject downstream use if any protected challenge bypasses full reference evaluation, any confident error changes a predeclared scientific decision, energy-force consistency fails, independent audits exceed budget or the correction path does not fit the bounded reference-query objective after all overhead is counted.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Proposal 1 uses matched endpoint predictors to compress and reconstruct a live Raman stream for remote reaction monitoring. This replacement does not encode a chronological experimental signal or optimize communication; it composes a lower-cost physical calculation with a learned reference-method discrepancy to govern when a complete numerical oracle is required. The superseded proposal 2 encoded atomistic trajectory frames for storage, whereas this replacement does not predict or archive successive trajectory states: it predicts the difference between two energy-force-stress evaluators and changes the numerical evaluator available to any compatible calculation. Proposal 3 routes unexpected candidate characterization packages into human review and follow-up planning; this replacement produces a corrected computational physics output under a reference-query budget and has no discovery-review queue as its primary intervention. Proposal 4 subtracts the sensory consequences of an outgoing recoater command to reveal external powder interactions; this replacement has no motor command, live machinery or self-signal cancellation. It is independently adoptable as a computational-method layer without a Raman monitor, trajectory codec, discovery campaign workflow or powder recoater.","revision_record":{"parent_version":null,"progress_targets_addressed":["Replace proposal 2 while preserving proposal_index 2 and version 0","Remove the auditor-identified overlap with proposal 1's scientific-stream predictive codec","Create an independently adoptable numerical residual-correction opportunity","Preserve prediction, signed residual learning, uncertainty, provenance, independent full-reference audits and fallback","Explain material diversity from proposals 1, 3 and 4 and from the superseded proposal 2"],"conceptual_changes":["Replaced trajectory-stream archival with reference-minus-baseline correction of energies, forces and stresses","Changed the scarce resource from storage and stream representation to authoritative reference calculations","Changed the downstream result from reconstructed trajectory frames to a governed corrected physics evaluator"],"operational_changes":["Made complete reference calculation the independent oracle and fallback","Added scope, convergence, invariant and energy-force-consistency gates","Separated reviewed residual-model updates from active downstream simulations"],"evidence_changes":["Replaced trajectory reconstruction testing with an equal-query offline comparison against uncorrected and direct-surrogate rivals","Required random full-reference audits and protected scope-boundary challenges","Added downstream scientific-decision and invariant falsifiers"],"claim_changes":["Removed the claim that changing from spectra to atomistic frames constitutes an independent intervention","Limited the remaining claim to governed baseline-plus-residual numerical correction under a reference-query budget"]}}