{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"predictive_residual_processing__chemistry_materials:P3:v0","cell_id":"predictive_residual_processing__chemistry_materials","search_queries":["site:nist.gov autonomous materials laboratories characterization data bottleneck campaign scientists review","materials discovery active learning characterization phase mapping closed loop CAMEO paper","A-Lab autonomous materials discovery XRD characterization Nature 2023","materials characterization anomaly detection residual prediction scientific discovery workflow","A-Lab Nature correction 2026 crystallography automated characterization correction","NIST autonomous experimentation materials community needs human in loop characterization official report","DOE autonomous materials science laboratory data characterization human review bottleneck report","materials characterization data deluge scientist attention review bottleneck high throughput official","NeXus data format official standard provenance materials characterization NXentry instrument sample data","FAIR data materials characterization standard raw data provenance official","NIST materials data provenance characterization standard raw data"],"sources":[{"source_id":"S1","title":"Achieving AI-Driven Autonomous Laboratories","publisher":"U.S. Department of Energy","url":"https://www.energy.gov/undersecretaryforscience/genesis-mission/achieving-ai-driven-autonomous-laboratories","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["DOE identifies traditional human-driven experimentation, combinatorial design spaces, and non-deterministic AI control as bottlenecks slowing discovery and wasting critical national assets.","DOE proposes integrated AI, real-time analysis, intelligent feedback, and data curation in National Laboratory infrastructure, identifying a credible funder and authorizing environment."]},{"source_id":"S2","title":"Autonomous Methods","publisher":"National Institute of Standards and Technology","url":"https://www.nist.gov/mml/mmsd/data-and-ai-driven-materials-science-group/autonomous-methods","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2023-05-05; updated 2025-03-24","accessed_at":"2026-08-03","claims_supported":["NIST says materials search spaces can be too large or expensive to measure comprehensively and that beamline experiments must be chosen carefully.","NIST operates stakeholder-facing autonomous methods that analyze prior data, target informative measurements, quantify uncertainty, and provide interpretable human-in-the-loop interfaces.","NIST reports CAMEO and ANDiE as deployed analogues, including a reported fivefold characterization acceleration for ANDiE."]},{"source_id":"S3","title":"Workshop Report on Autonomous Methodologies for Accelerating X-ray Measurements","publisher":"National Institute of Standards and Technology and International Centre for Diffraction Data","url":"https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.1500-25.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2024-11","accessed_at":"2026-08-03","claims_supported":["A predominantly industry workshop explicitly prioritized challenges and solutions for automated X-ray structural analysis and future public-private cooperation.","The workshop recorded demand for AI phase identification, prediction of additional phases in multiphase mixtures, ML data organization, uncertainty quantification, coordinated multi-technique analysis, and storage permitting reinterpretation.","Participants identified missing standardized diffraction formats and metadata as a costly distraction and called for industry adoption, robust tools, standards, and workforce development.","The report documents unresolved data-quality, trust, interoperability, intellectual-property, and economic issues relevant to adoption."]},{"source_id":"S4","title":"An autonomous laboratory for the accelerated synthesis of inorganic materials","publisher":"Nature","url":"https://www.nature.com/articles/s41586-023-06734-w","source_class":"PRIMARY_RESEARCH","publication_date":"2023-11-29; corrected 2026-01-19","accessed_at":"2026-08-03","claims_supported":["A-Lab integrated model-derived targets, robotic synthesis, automated XRD characterization, probabilistic phase analysis, active learning, and feedback across 353 experiments.","The workflow already treats failed or discrepant outcomes as inputs to subsequent recipe selection and reports that discrepancies are crucial signals for improving materials-synthesis understanding.","The study identifies model, kinetics, volatility, amorphization, and measurement-interpretation limitations, and states that multiphase characterization and automated interpretation still need improvement."]},{"source_id":"S5","title":"Author Correction: An autonomous laboratory for the accelerated synthesis of inorganic materials","publisher":"Nature","url":"https://www.nature.com/articles/s41586-025-09992-y.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2026-01-19","accessed_at":"2026-08-03","claims_supported":["Post-publication concerns required manual re-analysis of diffraction patterns.","Four of forty reported successes were found inconclusive from XRD alone, one compound had mistakenly appeared in training data, and novelty language required correction.","The episode directly supports independent review, provenance checks, multimodal fallback, and separation of model-relative success from scientific novelty."]},{"source_id":"S6","title":"On-the-fly closed-loop materials discovery via Bayesian active learning","publisher":"Nature Communications","url":"https://www.nature.com/articles/s41467-020-19597-w","source_class":"PRIMARY_RESEARCH","publication_date":"2020-11-24","accessed_at":"2026-08-03","claims_supported":["CAMEO combined probabilistic phase-map and property predictions, uncertainty, physics-informed active learning, human input, and real-time XRD control.","The system prioritized informative measurements and reported a tenfold reduction in experiments, demonstrating the feasibility and value of model-governed scientific attention.","CAMEO is a close analogue but chooses measurements and objectives rather than presenting reconstructive multimodal residuals through an audited review queue."]},{"source_id":"S7","title":"Human-in-the-loop for Bayesian autonomous materials phase mapping","publisher":"National Institute of Standards and Technology / Matter","url":"https://www.nist.gov/publications/human-loop-bayesian-autonomous-materials-phase-mapping","source_class":"PRIMARY_RESEARCH","publication_date":"2024-02-07","accessed_at":"2026-08-03","claims_supported":["The study probabilistically integrates expert indications of phase boundaries, phase regions, uncertainty, and regions of interest into autonomous phase mapping.","It reports improved phase-mapping performance given appropriate human input, supporting governed expert intervention rather than fully automated interpretation.","It is adjacent prior art for attributable model updates but does not test residual-first package triage with random full-package audits."]},{"source_id":"S8","title":"NCAL: Data Management","publisher":"National Institute of Standards and Technology","url":"https://www.nist.gov/programs-projects/ncal-data-management","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2025-02-04; updated 2025-11-25","accessed_at":"2026-08-03","claims_supported":["NIST reports no widely accepted materials-test-data management solution and describes requirements beyond ordinary commercial archives.","Its deployed system preserves material, test-history, machine-configuration, analysis, documentation, and end-to-end provenance across more than 10,000 records and nearly 20 TB.","The implementation shows that raw-data retention, versioned instrument context, searchable metadata, and restoration are technically feasible, although laboratory integration is substantial."]}],"problem_evidence":{"support":"STRONG","rationale":"DOE and NIST independently describe combinatorial materials spaces, costly measurements, human-driven bottlenecks, and demand for informative experiment targeting. The X-ray workshop documents industry-facing needs in phase identification, data organization, uncertainty, and scalable analysis. A-Lab generated 353 experiments in 17 days and later required manual diffraction re-analysis, visibly demonstrating both characterization volume and the consequence of trusting automated interpretation. The exact prevalence of undifferentiated full-package review queues, however, was not measured.","source_ids":["S1","S2","S3","S4","S5","S8"]},"stakeholder_evidence":{"support":"STRONG","rationale":"DOE identifies National Laboratories as infrastructure and funding nuclei for AI-driven laboratories. NIST develops autonomous materials methods for stakeholders and its own systems, while the NIST–ICDD workshop explicitly solicited predominantly industry input and called for public-private action and adoption. These are credible adopters, authorizers, and funders, though none expressly requested this exact residual-first queue.","source_ids":["S1","S2","S3"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"CAMEO closed-loop Bayesian materials discovery","similarity":"Uses physics-informed predictions, uncertainty, human-in-the-loop input, and information-directed XRD measurements to allocate scarce experimental capacity.","remaining_difference":"It selects the next measurement or experiment; it does not make a frozen prediction-plus-reconstructive-residual package the primary scientist-review message, nor report random full-package audits and version-triggered decompression as the governing comparison.","source_ids":["S2","S6"]},{"name":"A-Lab autonomous inorganic synthesis and characterization","similarity":"Integrates predicted targets, automated XRD interpretation, active learning, observed discrepancies, and iterative synthesis decisions across a materials campaign.","remaining_difference":"It automates synthesis success assessment and recipe selection rather than shadow-testing an attention queue against complete multimodal package review. Its 2026 correction underscores the proposed safeguards but does not establish their effectiveness.","source_ids":["S4","S5"]},{"name":"NIST human-in-the-loop Bayesian phase mapping","similarity":"Combines probabilistic phase predictions, uncertainty, expert knowledge, and governed updates in an autonomous exploration campaign.","remaining_difference":"Human input modifies priors and regions of interest; the method does not test residual-only presentations, action concordance, raw audit sampling, or total reviewer burden.","source_ids":["S7"]},{"name":"NIST autonomous-methods and X-ray workflow program","similarity":"Targets informative measurements, physically meaningful outputs, quantified uncertainty, interpretable interfaces, standards, and coordinated characterization workflows.","remaining_difference":"This is a programmatic family of methods and requirements, not a demonstrated residual-first candidate-review protocol with a frozen baseline and explicit fallback rules.","source_ids":["S2","S3"]}],"distinctive_claim_remaining":"On a preregistered finite campaign within one material family, a queue presenting frozen, versioned multimodal predictions plus signed, categorical, missingness, and unpredicted-feature residuals—while forcing complete review for protected cases and a random/risk-stratified audit sample—will reduce total reviewer time by at least 30% relative to blinded complete-package review while retaining at least 95% overall action concordance, 100% capture of protected challenge cases, zero acceptance of missing or version-incompatible observations, and no increase in follow-up-characterization demand after audit, maintenance, and fallback work are counted.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"CAMEO, A-Lab, NIST human-in-the-loop methods, and NCAL collectively demonstrate probabilistic prediction, uncertainty, automated characterization, expert interfaces, active updates, provenance, versioned machine context, large raw-data retention, and feedback loops. The proposed system can remain decision support in shadow mode, avoiding autonomous laboratory actions. No source demonstrates the complete multimodal residual representation, safe thresholds, audit sampling rate, blinded workflow equivalence, or integration across the proposed actors. Legal exposure appears primarily contractual, cybersecurity, export-control, trade-secret, and research-governance dependent; no general legal prohibition was found, but institution-specific review remains necessary.","source_ids":["S2","S3","S4","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"If the review constraint is real, redirecting expert attention while preserving full evidence could materially accelerate campaigns and catch model-breaking results; prior systems report fivefold to tenfold measurement efficiencies, but impact from this exact queue is untested.","source_ids":["S2","S6"]},"stakeholder_pull":{"score":4,"rationale":"DOE, NIST, and predominantly industry workshop participants express strong demand for autonomous, trustworthy, information-efficient materials workflows, although not for this exact product specification.","source_ids":["S1","S2","S3"]},"incremental_advantage":{"score":3,"rationale":"The auditable residual-presentation layer could add value beyond active learning and anomaly detection, but the claimed workload and decision-quality advantage requires a direct comparator study.","source_ids":["S4","S5","S6","S7"]},"distinctiveness_plausibility":{"score":3,"rationale":"No exact match was found among the eight sources for reconstructive residual-first review plus independent full-package audits, yet most constituent mechanisms are already established adjacent practice.","source_ids":["S2","S3","S4","S6","S7"]},"technical_implementability":{"score":4,"rationale":"Existing autonomous laboratories and deployed data-management systems demonstrate most required components. Multimodal standardization, calibration, and audit independence remain nontrivial integration work.","source_ids":["S2","S4","S6","S8"]},"adoption_authority_feasibility":{"score":3,"rationale":"A laboratory director, campaign committee, characterization lead, data steward, safety owner, and model owner can authorize a shadow deployment without surrendering scientific authority, but operational adoption requires multi-role agreement and institution-specific data/IP review.","source_ids":["S1","S2","S3"]},"evidence_readiness":{"score":2,"rationale":"The problem and technical ingredients are well supported, but no externally reported evaluation measures the exact intervention against complete-package review, and target-campaign queue prevalence is unknown.","source_ids":["S2","S3","S4","S5","S6","S7"]},"safety_net_benefit":{"score":4,"rationale":"Manual re-analysis changing four A-Lab success classifications provides direct motivation for protected bypasses, raw audits, provenance, and full-review fallback. Whether the proposed sampling catches rare blind spots remains empirical.","source_ids":["S5","S8"]},"scalability":{"score":3,"rationale":"Digital routing, prediction, and provenance can scale, and NIST reports a 20-TB characterization archive, but assay-specific standardization, expert review capacity, and fallback frequency may limit cross-laboratory scaling.","source_ids":["S3","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"One shadow comparison on an existing finite candidate matrix, using existing instruments and complete-review workflow; includes model freezing, residual generation, challenge-package preparation, blinded review, workload capture, statistical analysis, and governance documentation.","confidence":"LOW","assumptions":["Existing raw data, instruments, model outputs, and laboratory safety processes are available.","Approximately 0.3–1.0 FTE-year is distributed among a data scientist, materials scientist, characterization specialist, and study coordinator.","No autonomous synthesis or new major instrument purchase occurs.","The range is a resource-equivalent estimate, not a vendor quote."],"source_ids":["S2","S3","S4","S8"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Production-quality integration for one laboratory and one material family: data connectors, standardized outputs, model registry/checksums, queue interface, audit sampler, access controls, replay store, validation, and staff training.","confidence":"LOW","assumptions":["Existing electronic laboratory and characterization systems expose usable exports or APIs.","No new XRD, microscopy, spectroscopy, or robotic hardware is included.","Two to five technical and scientific contributors work part-time for 6–12 months.","Cybersecurity and IP controls use existing institutional infrastructure."],"source_ids":["S3","S8"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Validated launch across multiple assays or campaigns, including assay-specific comparators, uncertainty calibration, multimodal provenance, failure drills, independent audit operations, workflow change management, security review, and dedicated support.","confidence":"LOW","assumptions":["The launch spans several characterization modalities or organizational teams.","Existing laboratory instruments remain in place; major robotics or beamline construction is excluded.","Costs include scientific validation and change management, not only software development.","Fallback review capacity is funded during rollout."],"source_ids":["S1","S3","S4","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Annual operation for a deployed laboratory: model and threshold governance, data engineering, calibration, raw-package audits, fallback reviews, security maintenance, incident investigation, storage, and user support.","confidence":"LOW","assumptions":["One to three resource-equivalent FTEs cover scientific ownership, data/model operations, and characterization QA.","Instrument operating and synthesis costs that would exist without the queue are excluded.","Storage growth is comparable to an established characterization program, while complete raw records remain retained.","Fallback frequency is moderate; frequent decompression could exceed the band."],"source_ids":["S3","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent DOE, NIST, industry-workshop, and primary-research evidence supports combinatorial scale, expensive characterization, analysis/data burdens, and consequential automated-interpretation failures.","source_ids":["S1","S2","S3","S4","S5","S8"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"DOE National Laboratories, NIST autonomous-methods programs, user facilities, and industry participants are identifiable funders, adopters, or convening authorizers for this class of workflow.","source_ids":["S1","S2","S3"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal can be distinguished from active learning by testing whether residual-first presentation reduces total review work without degrading blinded actions or protected-case capture.","source_ids":["S4","S5","S6","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A shadow study on one frozen candidate matrix can compare residual and complete-package paths without changing experiments or scientific authority and has explicit success and rejection criteria.","source_ids":["S2","S5","S6","S8"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step is non-interventional: existing laboratory, safety, characterization, and planning authorities remain controlling; all raw data are preserved; missingness, incompatibility, out-of-scope cases, and safety flags force complete review. Institution-specific IP, cybersecurity, and data approvals are still prerequisites but are not an intrinsic stop.","source_ids":["S3","S5","S8"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The four ranges are bounded resource-equivalent estimates with explicit exclusions, but the opened sources provide no priced bill of materials, labor quote, or observed cost for this exact queue. Integration difficulty and fallback workload could move deployment by more than one band.","source_ids":["S3","S8"]}},"next_evidence_step":"Partner with one laboratory to preregister and execute a shadow, paired, blinded study on 60–150 already-authorized candidates from one material family and fixed assay suite. Freeze the campaign model, uncertainty calibration, preprocessing, checksum, thresholds, protected classes, audit draw, review-capacity budget, and prohibition on model updates before revealing results. Preserve the existing complete-package review as the authoritative comparator. Independently generate residual presentations and include concealed challenges covering small reliable property shifts, categorical phase conflicts, unpredicted peaks/features, corrupted or low-quality assays, missing observations, version mismatch, and out-of-scope chemistry. Randomize presentation order and reviewers where feasible. Primary comparators are complete-package review, uncertainty-only prioritization, and a fixed anomaly/feature threshold. Measure total reviewer minutes including audits and fallback, action concordance, protected-case recall, false deferrals, reconstruction error, unpredicted-feature recall, follow-up requests, audit disagreement, and fallback correctness. Falsify operational adoption if protected-case recall is below 100%, overall action concordance is below 95%, any missing/version-incompatible package passes, residual review systematically misses a discrepancy class, or total workload falls by less than 30% after all overhead is counted.","blocking_evidence":["No field evidence shows that the target laboratory's complete-package review queue exceeds its declared capacity or delays model-relevant discrepancies.","No paired blinded trial establishes noninferior scientific actions or lower total workload for residual presentations.","Safe consequence weights, residual thresholds, model-expiry rules, and random/risk audit rates are unknown.","Multimodal comparators may fail to preserve unfamiliar peaks, morphology, minority phases, or contextual patterns not represented in standardized outputs.","The independence and statistical power of raw-package audits, especially for rare events, have not been demonstrated.","Institution-specific ownership, trade-secret, export-control, cybersecurity, retention, and publication requirements have not been reviewed.","No vendor quotes or observed labor accounting validate the 2026 cost bands.","World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This eight-source search found established active learning, autonomous XRD control, human-in-the-loop Bayesian phase mapping, automated synthesis/characterization, provenance systems, and explicit demand for trustworthy AI-ready workflows. It did not find an exact documented implementation combining frozen multimodal predicted packages, reconstructive residual-first scientist review, protected bypass classes, independent random and risk-stratified complete-package audits, checksum-triggered decompression, and paired workload/action equivalence testing. This is only a contrast within the opened sources; world novelty, patentability, freedom to operate, market size, and realized impact were not assessed.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain campaign-level evidence that complete-package review is a binding constraint and quantify current delay, reviewer time, and discrepancy outcomes.","Run the preregistered paired shadow comparison against complete review, uncertainty-only prioritization, and fixed anomaly screening.","Demonstrate 100% protected-case capture, at least 95% action concordance, zero missing/version-incompatible passes, and at least 30% net reviewer-time reduction.","Estimate rare-event audit power and validate random plus risk-stratified audit rates on concealed challenge cases.","Complete institution-specific scientific-authority, safety, IP, export-control, cybersecurity, and retention review.","Replace resource-equivalent cost estimates with observed labor, storage, integration, and fallback accounting from the shadow study.","Differentiate the residual queue operationally from CAMEO, A-Lab, human-in-the-loop phase mapping, and ordinary anomaly dashboards using the paired results."],"reason":"Bounded web research supports the problem, adopter class, technical ingredients, adjacent prior art, and a testable incremental claim, but cannot establish the proposal's decisive assertions: workload reduction, action equivalence, rare-discrepancy safety, audit effectiveness, or real integration cost. Those require proprietary campaign data, expert fieldwork, and live or retrospective paired testing. Under the required controller rule, this empirical-research stop is terminal and therefore marked non-repairable."},"proposal_index":3}