{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"predictive_residual_processing__chemistry_materials:P2:v0","cell_id":"predictive_residual_processing__chemistry_materials","search_queries":["delta machine learning baseline correction quantum chemistry reference method energies forces materials","machine learning interatomic potentials reference calculations computational cost uncertainty active learning materials","official documentation machine learning force fields VASP on the fly reference calculations uncertainty","materials project machine learned interatomic potentials official documentation","site:nature.com delta machine learning quantum chemistry 2015 Ramakrishnan baseline correction","delta learning interatomic potential baseline potential forces energy materials primary research","physics-informed delta learning interatomic potential baseline correction reference energy forces","machine learning interatomic potentials uncertainty extrapolation failure primary research energy force consistency","Ramakrishnan Dral Rupp von Lilienfeld Big Data Meets Quantum Chemistry Approximations delta machine learning journal","\"Δ-machine learning\" reference baseline energies forces molecular dynamics","\"delta-learning\" \"interatomic potentials\" active 2025 journal","site:openkim.org reliability reproducibility interatomic potentials verification checks","\"Active delta-learning for fast construction\" IOPscience 2025","\"Kohn-Sham accuracy from orbital-free\" journal delta machine learning 2024","site:openkim.org \"verification checks\" interatomic models","AWS EC2 on demand pricing official 2026","site:bls.gov/ooh physical scientists materials scientists median pay 2025","site:bls.gov/oes materials scientists annual mean wage 2025","site:bls.gov software developers median pay 2025","site:openkim.org verification checks models OpenKIM documentation","site:openkim.org \"Testing Framework\" interatomic potentials verification","site:docs.openkim.org model verification checks"],"sources":[{"source_id":"S1","title":"Big Data meets Quantum Chemistry Approximations: The Δ-Machine Learning Approach","publisher":"arXiv","url":"https://arxiv.org/abs/1503.04987","source_class":"PRIMARY_RESEARCH","publication_date":"2015-03-17","accessed_at":"2026-08-03","claims_supported":["Expensive target-level quantum-chemistry properties can be approximated by adding a learned correction to an inexpensive baseline.","The original Δ-ML work reported out-of-sample gains over direct target prediction and showed that baseline quality affects transferability.","High-level reference calculations such as post-Hartree–Fock methods carry steep computational scaling."]},{"source_id":"S2","title":"Kohn-Sham accuracy from orbital-free density functional theory via Δ-machine learning","publisher":"arXiv; subsequently Journal of Chemical Physics","url":"https://arxiv.org/abs/2310.06598","source_class":"PRIMARY_RESEARCH","publication_date":"2023-10-10","accessed_at":"2026-08-03","claims_supported":["A materials-specific Δ-ML force field has already learned Kohn-Sham-minus-orbital-free energies and forces for on-the-fly molecular dynamics.","The reported method improved representative energy and force accuracy by more than two orders of magnitude and was tested on molten Al-Si.","Baseline-plus-learned-residual evaluation is technically implementable for atomistic materials simulations."]},{"source_id":"S3","title":"Active delta-learning for fast construction of interatomic potentials and stable molecular dynamics simulations","publisher":"IOP Publishing; record hosted by Cambridge Open Engage","url":"https://www.cambridge.org/engage/chemrxiv/article-details/673a046f5a82cea2fa6c52a5","source_class":"PRIMARY_RESEARCH","publication_date":"2025-07-14","accessed_at":"2026-08-03","claims_supported":["Active Δ-learning already combines baseline residual learning with iterative selection of expensive reference calculations.","The authors reported approximately tenfold fewer sampled points and iterations than non-Δ active learning at similar accuracy on their test reactions.","The study reported improved simulation stability relative to pure learned potentials, making it the closest collision with the proposed causal core."]},{"source_id":"S4","title":"Machine learning force field calculations: Basics","publisher":"VASP Software GmbH","url":"https://vasp.at/wiki/Machine_learning_force_field_calculations","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-03-20","accessed_at":"2026-08-03","claims_supported":["VASP officially supports on-the-fly force-field training in which Bayesian error thresholds determine whether an ab-initio calculation is performed.","VASP documentation says reference data must be consistent and well converged, warns that misconfiguration can severely damage force-field quality, and recommends benchmarking against ab-initio quantities and conservation principles.","VASP users and computational-method owners are identifiable potential adopters with an existing workflow and authority surface for reference-query governance."]},{"source_id":"S5","title":"Active learning of uniformly accurate interatomic potentials for materials simulation","publisher":"American Physical Society","url":"https://journals.aps.org/prmaterials/abstract/10.1103/PhysRevMaterials.3.023804","source_class":"PRIMARY_RESEARCH","publication_date":"2019-02-25","accessed_at":"2026-08-03","claims_supported":["DP-GEN established an exploration, reference-data generation, and training loop for interatomic potentials.","Applications to Al, Mg, and Al-Mg demonstrated that active learning can reduce required reference data.","Uncertainty- or model-driven reference acquisition is established prior art even when the learned target is not a residual."]},{"source_id":"S6","title":"Training data selection for accuracy and transferability of interatomic potentials","publisher":"Nature Portfolio","url":"https://www.nature.com/articles/s41524-022-00872-x","source_class":"PRIMARY_RESEARCH","publication_date":"2022-09-01","accessed_at":"2026-08-03","claims_supported":["Machine-learned interatomic potentials can have large out-of-sample errors and struggle with transferability.","Expert-curated configuration classes can produce misleading validation, while diverse coverage of descriptor space reduces extrapolation.","The paper identifies LAMMPS as an actively used deployment path and reports DOE-supported development, supporting identifiable institutional users and funders."]},{"source_id":"S7","title":"OpenKIM verification-check dashboard for Tersoff-style iron potential MO_137964310702_000","publisher":"Open Knowledgebase of Interatomic Models","url":"https://openkim.org/cite/MO_137964310702_000","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2015","accessed_at":"2026-08-03","claims_supported":["OpenKIM operationalizes versioned interatomic-model records and automated verification checks.","Existing checks include species support, periodicity, permutation symmetry, energy-force derivative consistency, continuity, objectivity, inversion symmetry, memory behavior, and deterministic threaded execution.","Invariant and version checks proposed by the candidate are feasible but are not novel as general interatomic-model governance practices."]},{"source_id":"S8","title":"Chemists and Materials Scientists","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/ooh/life-physical-and-social-science/chemists-and-materials-scientists.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2025-08-28","accessed_at":"2026-08-03","claims_supported":["The May 2024 median annual wage for materials scientists was $104,160, with research-and-development industry wages reported at $118,010.","Materials scientists use modeling and simulation and commonly work in teams, supporting labor-equivalent cost assumptions for scientific design, validation, and review.","The cost bands are resource-equivalent estimates rather than vendor quotes or measured project budgets."]}],"problem_evidence":{"support":"STRONG","rationale":"Primary studies and official VASP guidance visibly establish the binding tradeoff: authoritative electronic-structure evaluations are expensive, inexpensive potentials can be systematically wrong, reference-data acquisition is a bottleneck, and out-of-sample transferability and convergence quality are material scientific risks. The evidence supports the general problem strongly, although it does not quantify prevalence for the candidate's unspecified material system.","source_ids":["S1","S2","S3","S4","S5","S6"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"VASP exposes an official on-the-fly MLFF workflow to computational materials scientists, and DOE-supported LAMMPS work identifies institutional method developers and a large user community. These are credible adopters, authorizers, and funders. No source documents a commitment to adopt this exact independent-audit, batch-approval governance package, so demonstrated pull is weaker than demonstrated workflow relevance.","source_ids":["S4","S6","S7"]},"prior_art":{"proximity":"ESTABLISHED_PRACTICE","closest_analogues":[{"name":"Active delta-learning for interatomic potentials","similarity":"It learns target-minus-baseline energy and force differences, uses active selection of expensive reference points, and reports efficiency and simulation-stability gains—the candidate's principal numerical causal path.","remaining_difference":"The candidate adds predeclared random and risk-stratified oracle audits independent of model uncertainty, explicit version/scope authorization, reviewed batch activation, protected scientific-decision falsifiers, and mandatory fallback or halt.","source_ids":["S3"]},{"name":"Orbital-free-to-Kohn-Sham Δ-ML force field","similarity":"It uses a lower-cost physical method as the baseline and learns the difference to Kohn-Sham energies and forces for materials molecular dynamics.","remaining_difference":"The published demonstration does not establish the candidate's full governance system of independent audits, authority separation, challenge classes, or frozen production versions.","source_ids":["S2"]},{"name":"VASP on-the-fly MLFF and DP-GEN active learning","similarity":"Both alternate learned force evaluation with uncertainty-triggered ab-initio reference calculations and maintain training/application workflows under a reference-compute constraint.","remaining_difference":"These approaches generally learn the complete potential rather than necessarily learning a signed correction to a separately versioned physical baseline; the reviewed sources do not show model-independent random reference audits.","source_ids":["S4","S5"]},{"name":"OpenKIM verification and versioned model records","similarity":"OpenKIM already applies invariant, consistency, compatibility, and provenance checks to deployable interatomic potentials.","remaining_difference":"It does not itself implement the proposed baseline-plus-residual active reference-query loop or demonstrate that the combined governance improves downstream decisions.","source_ids":["S7"]}],"distinctive_claim_remaining":"For one frozen material domain and equal reference-query budget, adding model-independent random and protected-class reference audits, explicit baseline/reference/residual version gating, reviewed batch activation, and automatic halt/fallback to an active Δ-learning evaluator will reduce confidently wrong downstream energy-ordering, force-direction, or stress decisions relative to ordinary active Δ-learning and a direct-reference surrogate, without eliminating the governed compute advantage. This is contrastive and falsifiable, but currently unsupported by comparative evidence.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"The numerical primitives—baseline-plus-residual energy/force learning, Bayesian or ensemble uncertainty, active reference acquisition, versioned model records, invariant checks, and fallback to ab-initio evaluation—exist separately in research or production tooling. An offline shadow implementation faces no obvious legal barrier beyond software licenses and institutional compute/data policy. However, no reviewed source validates the full integrated workflow, calibrated tail-risk thresholds, energy-force-stress consistency, independent audit sampling, or authority handoffs on the proposed material domain.","source_ids":["S2","S3","S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":4,"rationale":"If reference calculations are genuinely binding, a reliable corrected evaluator could expand accessible simulation length or system size while protecting consequential scientific decisions.","source_ids":["S1","S2","S3"]},"stakeholder_pull":{"score":3,"rationale":"There is clear workflow demand and official tooling for reference-efficient force fields, but no named organization has requested or committed to the exact governance package.","source_ids":["S4","S6"]},"incremental_advantage":{"score":3,"rationale":"Δ-learning and active acquisition already report substantial benefits; the incremental benefit of independent audits, version gates, and fallback over competent existing active-learning practice remains unmeasured.","source_ids":["S2","S3","S4"]},"distinctiveness_plausibility":{"score":2,"rationale":"The numerical core is established practice, and verification, provenance, invariants, uncertainty triggers, and reference fallback are known practices. Only their strict combined governance and comparator-defined tail-safety claim remains distinctive.","source_ids":["S2","S3","S4","S5","S7"]},"technical_implementability":{"score":4,"rationale":"All major computational components have credible implementations or close demonstrations; integration, calibration, and scientific validation rather than basic technical invention are the main work.","source_ids":["S2","S3","S4","S7"]},"adoption_authority_feasibility":{"score":4,"rationale":"An offline shadow pilot can be authorized by a reference-method owner and simulation owner without allowing the correction to control production or publication-critical results.","source_ids":["S4","S7"]},"evidence_readiness":{"score":4,"rationale":"A bounded computational experiment can retain complete reference outputs and evaluate all comparators offline, although it requires proprietary or institutionally controlled code, compute, and new reference calculations.","source_ids":["S2","S3","S4"]},"safety_net_benefit":{"score":4,"rationale":"Independent oracle audits and immediate fallback directly address extrapolation, convergence, invariant, and self-selection failures, but their rare-event coverage must be measured rather than assumed.","source_ids":["S4","S6","S7"]},"scalability":{"score":3,"rationale":"Software patterns scale, but each new composition, phase, electronic state, baseline/reference pairing, and thermodynamic regime needs renewed reference coverage and validation.","source_ids":["S2","S4","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Pre-registration, one bounded material/configuration set, complete reference outputs, simulated query budgets, three comparators, uncertainty calibration, invariant tests, and an independent numerical review.","confidence":"MODERATE","assumptions":["Approximately 0.4-1.0 labor-year split among a computational materials scientist, ML engineer, and reviewer.","Existing baseline/reference software licenses, workflow code, and institutional HPC access are available.","Reference calculations are bounded to hundreds or low thousands of tractable configurations rather than a production-scale reactive domain.","The BLS materials-scientist wage is used only as a labor-equivalent anchor; institutional overhead and scarce HPC can materially change cost."],"source_ids":["S2","S3","S8"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Integrate version gating, provenance records, scheduler-controlled oracle queries, uncertainty and disagreement checks, audit sampler, immutable model releases, dashboards, and rollback into one institution's simulation workflow.","confidence":"LOW","assumptions":["Roughly 1-3 multidisciplinary labor-years plus HPC and software integration.","Deployment remains limited to one baseline/reference pairing and a small number of approved material regimes.","No new electronic-structure code or general-purpose MLIP architecture is invented.","Security, procurement, and software-license changes are modest."],"source_ids":["S4","S7","S8"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Validated institutional service across several material regimes with independent review, production observability, regression and challenge suites, user training, incident response, and reserved reference-compute capacity.","confidence":"LOW","assumptions":["Several material domains require separate reference campaigns and downstream scientific-decision tests.","A multidisciplinary team maintains method, platform, and independent-validation functions.","The service supports consequential research workflows but not regulated certification or safety-critical control.","Reference-compute demand and fallback frequency remain uncertain until the first experiment."],"source_ids":["S4","S6","S7","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Model monitoring, periodic independent reference audits, reviewed retraining, regression testing, HPC/storage, support, and governance for a bounded institutional deployment.","confidence":"LOW","assumptions":["Approximately 1-3 continuing full-time-equivalent roles plus compute and storage.","Material scope changes are controlled; major expansion would be a new launch cost.","Audit and fallback rates are low enough that complete reference calculations do not erase the compute advantage.","No cost evidence from an actual adopter was found."],"source_ids":["S4","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Multiple primary studies and official product guidance independently document expensive reference computation, reference-data bottlenecks, transferability failures, and the need for convergence and quality control.","source_ids":["S1","S3","S4","S6"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Computational materials scientists and reference-method owners using VASP or LAMMPS are identifiable adopters; VASP's official MLFF workflow gives them a concrete implementation and authorization surface, although adoption commitment to this proposal is absent.","source_ids":["S4","S6"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim compares governed active Δ-learning against ordinary active Δ-learning, a direct-reference surrogate, and an uncorrected baseline at equal reference budget, with tail scientific-decision errors and total governed cost as outcomes.","source_ids":["S2","S3","S4"]},"bounded_next_evidence_step":{"status":"YES","reason":"A single-composition offline shadow benchmark with frozen methods, finite configurations, full reference retention, equal-query comparators, forced failures, and explicit rejection rules is bounded and executable.","source_ids":["S2","S3","S4","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step leaves every complete reference result authoritative, prohibits production control, and can halt on scope, convergence, invariant, version, uncertainty, disagreement, or audit failures. Remaining risks are scientific-validity questions measurable offline rather than an authority stop.","source_ids":["S4","S6","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The ranges explicitly separate experiment, integration, operational launch, and recurrence, anchor labor to official wage evidence, and state reference-compute and scope assumptions. Precision remains low because no adopter-specific configuration count or HPC benchmark was supplied.","source_ids":["S2","S3","S8"]}},"next_evidence_step":"Pre-register a shadow-only experiment on one composition and bounded structural regime. Freeze the baseline implementation, reference method and convergence recipe, residual definition, energy-force-stress consistency rule, uncertainty method, query budget, random-audit fraction, protected challenge classes, and model versions before revealing test references. Compute and retain complete reference outputs for every configuration, while simulating a fixed subset available to training and audit. At the same reference-query count, compare: (1) uncorrected baseline, (2) ordinary uncertainty-driven direct-reference surrogate, (3) ordinary active Δ-learning, and (4) the proposed governed active Δ-learning. Use untouched equilibrium perturbations, strains, coordination changes, trajectory-generated configurations, and scope-boundary challenges. Measure component-wise errors, calibration and coverage, energy-ordering and force-direction decisions, stress response, invariant failures, reference-query and fallback rates, wall-clock/resource-equivalent cost, and errors discovered only by independent audits. Inject checksum mismatch, missing provenance, unconverged reference, out-of-scope chemistry, and missing-data faults. Falsify the incremental claim if governed Δ-learning fails any protected challenge without oracle fallback, remains confidently wrong on a predeclared scientific decision, violates energy-force consistency, has no statistically or practically meaningful reduction in tail decision errors versus ordinary active Δ-learning, or loses its compute advantage after audit and governance overhead.","blocking_evidence":["No empirical comparison of the full governed package against ordinary active Δ-learning and a direct-reference surrogate at equal reference-query budget.","No adopter-specific material composition, phase space, electronic-state assumptions, reference method, baseline method, convergence recipe, or downstream tolerances.","No measured calibration or coverage of the proposed uncertainty and disagreement triggers under configuration shift.","No measured random-audit rate sufficient to bound rare consequential misses.","No evidence that corrected energy, force, and stress outputs remain mutually consistent and dynamically stable across the declared scope.","No measured end-to-end cost including failed reference convergence, audits, reviews, fallbacks, versioning, and scheduler overhead.","No named adopter commitment, data-access agreement, HPC allocation, or reference-method-owner approval."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This eight-source evaluation establishes that Δ-learning, active reference acquisition, on-the-fly uncertainty gating, model verification, and invariant checking are established or adjacent practices. It does not measure world novelty, patentability, freedom to operate, market size, or realized impact. The only boundary retained for testing is the combined independent-audit, version-and-authority-gated, tail-decision-safety claim within one frozen materials domain.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain a named computational-materials partner, method-owner approval, and a bounded reference-compute allocation.","Freeze one material domain, baseline/reference pairing, convergence recipe, configuration generator, scientific-decision tolerances, and protected challenge classes.","Run the equal-reference-budget four-arm offline comparison with complete oracle retention and forced fault injections.","Demonstrate calibrated uncertainty coverage and quantify confidently wrong energy-ordering, force-direction, and stress decisions, including errors found only by random audits.","Show that energy-force-stress invariants and simulation stability pass predeclared tests.","Measure total governed resource cost and prove that audits, fallbacks, reviews, and model maintenance do not erase the reference-compute advantage."],"reason":"Bounded web research resolved the problem, adopter plausibility, implementation feasibility, and prior-art landscape. It also found that the proposal's numerical core is established practice and that only the integrated governance and tail-safety advantage remains. That advantage depends on new reference calculations and live computational testing against comparators, not further web search; under the required controller rule this is an empirical-research stop and therefore repairable is false."},"proposal_index":2}