{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","source_assessment_id":"predictive_residual_processing__chemistry_materials:P3:v0","cell_id":"predictive_residual_processing__chemistry_materials","proposal_index":3,"qualification":"EMPIRICAL_PARTNER_CANDIDATE","criteria":{"specific_differentiated_claim":{"status":"YES","reason":"The proposal retains a specific, falsifiable contrast with active learning and anomaly screening: frozen multimodal predictions plus reconstructive residual-first review, protected bypasses, independent full-package audits, and decompression should reduce total reviewer time by at least 30% while preserving at least 95% action concordance, 100% protected-case capture, and zero missing or version-incompatible passes."},"credible_problem_signal":{"status":"YES","reason":"Independent government, standards, workshop, and primary-research evidence supports costly high-volume characterization, analysis bottlenecks, and consequential failures of automated interpretation, although prevalence of the exact queue constraint remains to be measured at a partner laboratory."},"identifiable_partner_or_adopter":{"status":"YES","reason":"A materials laboratory running a model-guided campaign—such as a DOE or NIST laboratory, user facility, autonomous-materials program, or industrial characterization group—is a concrete partner class with laboratory leadership, characterization, model, safety, and data-governance authorities."},"partner_access_is_necessary":{"status":"YES","reason":"The decisive uncertainties require nonpublic campaign packages, frozen model outputs, real reviewer behavior, workflow timing, protected-case handling, audit outcomes, and institution-specific integration costs; ordinary public research cannot establish these effects."},"safe_authorized_first_step":{"status":"YES","reason":"The proposed shadow study preserves the existing complete-package workflow as authoritative, uses already-authorized candidates and assays, retains all raw data, prohibits model updates and operational influence, includes explicit fallback rules, and proceeds only under the partner's existing safety, data, IP, and scientific authorities."},"bounded_decisive_empirical_design":{"status":"YES","reason":"A preregistered paired, blinded study of 60–150 candidates from one material family compares residual-first review with complete-package review, uncertainty-only prioritization, and fixed anomaly screening using explicit workload, concordance, protected-case, reconstruction, audit, and fallback endpoints with hard rejection thresholds."},"no_material_negative_gate":{"status":"YES","reason":"All substantive pipeline gates are YES; cost scope is uncertain but not a safety, authority, problem, adopter, differentiation, or test-design failure, and observed partner-specific labor and integration accounting can be collected within the same bounded study."},"not_merely_more_research":{"status":"YES","reason":"The next step specifies a partner, sample range, fixed campaign scope, comparators, concealed challenges, frozen controls, measured outcomes, and quantitative success and falsification rules rather than requesting open-ended investigation."}},"uncertainty_types":["PROBLEM_PREVALENCE","ADOPTER_PULL","INCREMENTAL_EFFECT","WORKFLOW_FIT","DATA_ACCESS","COST_SCOPE"],"partner_profile":"One materials-discovery laboratory with an active model-guided campaign in a single material family, a fixed characterization-assay suite, retained raw and standardized packages, identifiable complete-package reviewers, and laboratory, characterization, model, safety, and data-governance owners able to authorize a non-interventional shadow study.","required_access":"Access to 60–150 already-authorized candidate packages, raw and standardized characterization data, prerevelation campaign-model predictions and uncertainties, provenance and version records, current complete-review decisions, reviewer participation and timing, follow-up requests, and partner approval for blinded challenge insertion and institution-specific governance review.","bounded_empirical_test":"Preregister and run a paired, blinded shadow comparison on 60–150 candidates from one material family. Freeze the model, calibration, preprocessing, checksum, thresholds, protected classes, audit draw, review budget, and update prohibition. Preserve complete-package review as authoritative; independently generate residual presentations; insert concealed protected and failure-mode challenges; compare against complete review, uncertainty-only prioritization, and fixed anomaly screening; and count reviewer, audit, fallback, maintenance, and follow-up workload.","success_condition":"After all audit, fallback, maintenance, and follow-up overhead is counted, residual-first review reduces total reviewer time by at least 30%, achieves at least 95% overall action concordance with blinded complete-package review, captures 100% of protected challenge cases, permits zero missing or version-incompatible packages to pass, preserves unfamiliar-feature detection, and does not increase follow-up-characterization demand.","falsification_condition":"Reject operational triage if workload falls by less than 30%, action concordance is below 95%, any protected case is missed, any missing or incompatible observation passes, any discrepancy class is systematically suppressed, reconstruction or audit disagreement exceeds preregistered bounds, or total follow-up demand increases after all overhead is included.","rationale":"This is an appropriate empirical-partner lane because adjacent prior art establishes the components and problem context but does not resolve the differentiated claim. The remaining questions are partner-specific and experimentally decidable, access to real campaign data and reviewers is essential, and the proposed shadow comparison is safe, bounded, comparator-based, and capable of either supporting or falsifying adoption."}