{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"predictive_residual_processing__environmental_climate:P2:v0","cell_id":"predictive_residual_processing__environmental_climate","search_queries":["site:docs.esmvaltool.org climate model evaluation diagnostics provenance recipe official","ILAMB climate model benchmarking hierarchical scoring regions observations paper","DOE climate model evaluation diagnostics expressed need E3SM official","climate model evaluation limited metrics process oriented diagnostics observations structured residuals paper","site:gmd.copernicus.org ESMValTool v2 provenance recipe diagnostics climate model evaluation 2020","CMEC framework climate model evaluation benchmark official DOE GMD 2023","site:wcrp-cmip.org model evaluation benchmarking observations working group official need","BLS data scientists hourly median pay 2025 official","site:e3sm.org model evaluation diagnostics package land atmosphere observations official E3SM","site:pcmdi.llnl.gov PCMDI metrics package climate model evaluation official","site:wcrp-cmip.org tools model benchmarking evaluation CMIP7 official","site:obs4mips.org official climate model evaluation observations uncertainty provenance","site:ipcc.ch/report/ar6/wg1 Chapter 10 model evaluation process based diagnostics metrics limitations final PDF","site:ipcc.ch/report/ar6/wg1 climate model evaluation observations uncertainty metrics process based diagnostics","climate model aggregate metrics mask regional errors process oriented evaluation primary paper","\"The International Land Model Benchmarking (ILAMB) System\" full text PDF","site:osti.gov 2018 ILAMB Collier Hoffman Lawrence benchmarking system"],"sources":[{"source_id":"S1","title":"The International Land Model Benchmarking (ILAMB) System: Design, Theory, and Implementation","publisher":"Journal of Advances in Modeling Earth Systems; University of California eScholarship record","url":"https://escholarship.org/uc/item/0d91b7bj","source_class":"PRIMARY_RESEARCH","publication_date":"2018","accessed_at":"2026-08-03","claims_supported":["Increasing Earth-system-model complexity creates demand for rigorous, multifaceted model-observation evaluation.","ILAMB already evaluates many land variables using in-situ, remote-sensing, and reanalysis data and presents hierarchical statistical pages and figures.","ILAMB is used by modeling teams and centers during development and intercomparison, establishing both prior art and an adopter class."]},{"source_id":"S2","title":"Earth System Model Evaluation Tool (ESMValTool) v2.0 – technical overview","publisher":"Geoscientific Model Development, Copernicus Publications","url":"https://gmd.copernicus.org/articles/13/1179/2020/index.html","source_class":"PRIMARY_RESEARCH","publication_date":"2020-03-16","accessed_at":"2026-08-03","claims_supported":["Growing model resolution, process complexity, experiments, and data volume create significant evaluation challenges.","Existing open-source infrastructure supports configurable recipes, model-observation preprocessing, task execution, provenance logging, standards-based data, and parallel or out-of-core operation.","ESMValTool provides technically mature components for versioned, reconstructible evaluation workflows."]},{"source_id":"S3","title":"Advanced climate model evaluation with ESMValTool v2.11.0 using parallel, out-of-core, and distributed computing","publisher":"Geoscientific Model Development, Copernicus Publications","url":"https://gmd.copernicus.org/articles/18/4009/2025/index.html","source_class":"PRIMARY_RESEARCH","publication_date":"2025-07-02","accessed_at":"2026-08-03","claims_supported":["Published CMIP6 output reached approximately 20 PB, making timely and efficient evaluation a visible problem.","ESMValTool is established and widely used, supports custom recipe-defined workflows, and records inputs, processing, diagnostics, and software versions for traceability.","Demonstrated parallel, distributed, and out-of-core operation makes the proposed bounded retrospective computation technically credible."]},{"source_id":"S4","title":"CMIP 2026: Launch of the Rapid Evaluation Framework","publisher":"World Climate Research Programme Coupled Model Intercomparison Project","url":"https://wcrp-cmip.org/cmip-2026-launch-of-the-rapid-evaluation-framework/","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026-06-19","accessed_at":"2026-08-03","claims_supported":["The WCRP Model Benchmarking and Evaluation Task Team created REF to streamline and standardize multi-model evaluation.","WCRP explicitly recognizes researcher time spent processing CMIP data and provides an identifiable authorizer, coordinator, and potential adopter for improved evaluation workflows.","REF supplies a current comparator with dashboards, observational comparisons, downloadable data, local customization, and planned legacy records."]},{"source_id":"S5","title":"Coordinated Model Evaluation Capabilities","publisher":"Lawrence Livermore National Laboratory","url":"https://cmec.llnl.gov/","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["DOE-supported CMEC reports that proliferating tools and absent standards can make even one evaluation package require extensive intervention.","CMEC coordinates PMP, ILAMB, IOMB, and other packages through lightweight standards, installation, execution, and shared data-product mechanisms.","Existing PMP and ILAMB capabilities already provide high-level scores, detailed diagnostics, observational comparisons, and interactive exploration, narrowing the candidate's incremental claim."]},{"source_id":"S6","title":"E3SM Diagnostics Package v3 documentation","publisher":"U.S. Department of Energy Energy Exascale Earth System Model Project","url":"https://docs.e3sm.org/e3sm_diags/_build/html/v3.0.1/index.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2020; version 3 documentation accessed 2026-08-03","accessed_at":"2026-08-03","claims_supported":["DOE's E3SM project is an identifiable modeling organization with an established diagnostics workflow and authority to authorize a shadow evaluation.","The package uses remote-sensing, reanalysis, and in-situ observations and connects atmosphere, land, and coupled-model focus groups.","Existing maps, tables, time series, regional configuration, custom diagnostics, testing, and output viewing reduce implementation risk."]},{"source_id":"S7","title":"Climate Change 2021: The Physical Science Basis, Chapter 1: Framing, Context and Methods","publisher":"Intergovernmental Panel on Climate Change","url":"https://www.ipcc.ch/report/ar6/wg1/chapter/chapter-1/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2021","accessed_at":"2026-08-03","claims_supported":["Climate-model evaluation commonly compares climatologies or time series with observations while considering observational uncertainty.","Process-oriented diagnostics and comprehensive tools are established, but dedicated fitness-for-purpose evaluation remains necessary for particular regions and uses.","Climate simulations involve forcing, response, internal-variability, and natural-variability uncertainties, limiting causal interpretation of coherent residuals."]},{"source_id":"S8","title":"Occupational Employment and Wages — May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-05-15","accessed_at":"2026-08-03","claims_supported":["May 2025 mean annual wages were $126,800 for data scientists, $148,100 for software developers, $106,110 for atmospheric and space scientists, and $97,110 for environmental scientists and geoscientists.","These wage levels provide a public labor-cost anchor for 2026 resource-equivalent estimates, before benefits, overhead, computing, and institutional indirect costs."]}],"problem_evidence":{"support":"MODERATE","rationale":"The problem is visible at the general level: model complexity and data volume have expanded sharply; current evaluation tools still require intervention; and WCRP explicitly seeks faster, standardized evaluation. However, the eight sources do not measure how often complete-field review exhausts expert attention or how frequently aggregate products conceal consequential land-atmosphere process failures. The candidate's specific prevalence and harm magnitude therefore remain unquantified.","source_ids":["S1","S2","S3","S4","S5","S7"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"DOE E3SM, DOE-supported CMEC, and WCRP's Model Benchmarking and Evaluation Task Team are identifiable adopters, authorizers, or funders with active diagnostics programs and expressed need for comprehensive, streamlined evaluation. No source records a commitment to adopt this exact hierarchical residual queue, independent raw-panel audit, or decompression protocol.","source_ids":["S4","S5","S6"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"International Land Model Benchmarking (ILAMB)","similarity":"ILAMB already performs observation-based evaluation across land variables, weights evidence by certainty, scale, and process importance, aggregates metrics, and exposes hierarchical pages and detailed figures to help scientists locate strengths and weaknesses.","remaining_difference":"The candidate would make unresolved signed residuals the governed review message across explicit site-to-region-to-process contracts, with a fixed attention budget, random and risk-stratified non-escalated panels, exact reconstruction tests, and mandatory scope decompression. Those combined workflow controls were not found in the opened ILAMB source.","source_ids":["S1","S5"]},{"name":"ESMValTool","similarity":"ESMValTool already supplies frozen recipe-like configurations, model-observation comparison, arbitrary regions and variables, provenance, version traceability, scalable computation, and numerous process-oriented diagnostics.","remaining_difference":"The opened sources do not describe an expert-review queue whose exclusions are independently sampled and whose audit disagreements automatically restore complete-field review.","source_ids":["S2","S3","S7"]},{"name":"CMIP Rapid Evaluation Framework","similarity":"REF centralizes standardized model diagnostics and observation comparisons, accelerates evaluation, supports interactive filtering and local customization, and plans legacy records.","remaining_difference":"REF is an access and execution framework, not an evidenced hierarchical signed-residual propagation and human error-classification protocol with independently sampled suppressed cases.","source_ids":["S4"]},{"name":"Coordinated Model Evaluation Capabilities and PMP","similarity":"CMEC standardizes execution and outputs across evaluation packages; PMP produces quick high-level observation-based scores, while CMEC integrates detailed packages such as ILAMB.","remaining_difference":"The candidate's incremental unit is a provenance-bearing, reconstructible residual case with queue and fallback governance, rather than package interoperability or a high-level score.","source_ids":["S5"]},{"name":"E3SM Diagnostics Package","similarity":"E3SM already uses comprehensive observational diagnostics, regional and custom configuration, maps, tables, time series, and multiple process-focused diagnostic sets during model development.","remaining_difference":"The documented package does not establish random full-field auditing of omitted cases, explicit suppressed-error budgets, or hierarchical escalation non-inferiority against complete expert review.","source_ids":["S6"]}],"distinctive_claim_remaining":"On a frozen run and held-out observations, compared with both complete-field review and an ILAMB/ESMValTool-style dashboard without the proposed queue-and-audit controls, hierarchical reconstructible residual triage will reduce median expert-review time by at least 25% while detecting every preregistered planted process failure and losing no more than five percentage points of recall for independently adjudicated material natural mismatches; every sampled comparison will reconstruct exactly, and any material suppressed mismatch found by audit will trigger scope decompression.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Technical and data feasibility are strong: open tools already align model and observational data, compute regional diagnostics, scale to large archives, and retain provenance. Workflow authority is plausible inside E3SM or another modeling center, and a retrospective shadow study avoids changing operational forecasts or published assessments. The unverified elements are the scientific validity of layer-specific 'explaining away,' threshold calibration under multiple testing and sparse coverage, true independence of audit observations, reviewer behavior, and whether the added audit/governance burden is lower than complete-field review. No material legal barrier is apparent for authorized internal data, but dataset-specific licenses and redistribution terms must be checked by the observation steward.","source_ids":["S1","S2","S3","S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"On a 1–5 scale, improved detection of model process failures could strengthen model-development evidence and avoid wasted review, but downstream climate impact is indirect and no effect size has been measured.","source_ids":["S3","S7"]},"stakeholder_pull":{"score":4,"rationale":"WCRP, DOE CMEC, and E3SM visibly invest in comprehensive, faster, standardized model evaluation; pull for this exact residual-audit variant is not yet demonstrated.","source_ids":["S4","S5","S6"]},"incremental_advantage":{"score":2,"rationale":"ILAMB, ESMValTool, REF, CMEC, and E3SM Diags cover most computational, hierarchical-display, weighting, provenance, and workflow infrastructure. Advantage rests on the narrower audited-attention protocol and requires direct comparison.","source_ids":["S1","S2","S3","S4","S5","S6"]},"distinctiveness_plausibility":{"score":2,"rationale":"A contrastive and falsifiable workflow claim remains, but the core intervention substantially collides with established benchmarking practice; absence of the exact combination in eight sources is not world-novelty evidence.","source_ids":["S1","S2","S4","S5","S6"]},"technical_implementability":{"score":4,"rationale":"Mature open tooling supports the necessary data alignment, diagnostics, provenance, regional slicing, customization, parallelism, and output viewing. Human queue, audit, and fallback modules still require integration and validation.","source_ids":["S2","S3","S5","S6"]},"adoption_authority_feasibility":{"score":4,"rationale":"A modeling center's evaluation lead can authorize a retrospective shadow workflow without altering the authoritative model or observations; E3SM and WCRP provide concrete institutional analogues. Dataset permissions remain local checks.","source_ids":["S4","S6"]},"evidence_readiness":{"score":3,"rationale":"A frozen run, held-out observations, existing diagnostics, and complete-field baseline permit a bounded experiment, but reviewer ground truth, materiality rules, thresholds, and audit independence must be preregistered.","source_ids":["S2","S3","S6","S7"]},"safety_net_benefit":{"score":4,"rationale":"Random full-field panels, provenance, reconstruction, missingness states, and decompression directly address silent suppression and invalid inference. Their sensitivity and operational reliability remain untested.","source_ids":["S2","S3","S7"]},"scalability":{"score":3,"rationale":"The computation can scale through established parallel and distributed tooling, but scientific-review labor, observation stewardship, region-specific hierarchy design, and audit rates may scale poorly.","source_ids":["S3","S5","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"One retrospective shadow test on one frozen run, one variable family, two regions, and one held-out seasonal cycle, using existing infrastructure; approximately 4–8 combined scientist/engineer person-weeks plus reviewer sessions and modest compute.","confidence":"MODERATE","assumptions":["Model output and observation products are already available and authorized.","Existing ESMValTool, ILAMB, or E3SM Diags components are reused.","The estimate values internal labor and compute as resource equivalents rather than requiring new hardware.","BLS mean wages are uplifted for benefits and institutional overhead."],"source_ids":["S2","S3","S6","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Build and validate one center-specific prototype: residual schema, hierarchy configuration, provenance binding, queue interface, audit sampler, reconstruction tests, replay storage, and fallback controls.","confidence":"MODERATE","assumptions":["Approximately 0.5–1.3 loaded FTE-years across research software engineering, climate science, observation stewardship, and review design.","Open-source evaluation infrastructure is extended rather than replaced.","No new model simulation campaign or observation acquisition is included."],"source_ids":["S2","S5","S6","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Production launch across several variable families and regions, including hardened pipelines, access controls, regression and fallback testing, documentation, reviewer training, observation-release management, and an initial evaluation cycle.","confidence":"LOW","assumptions":["Approximately 1.5–4 loaded FTE-years plus institutional computing and storage.","Existing center HPC, archives, and identity systems are available.","Launch remains an internal evaluation aid and does not include certification of public climate products."],"source_ids":["S3","S5","S6","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Annual operation at one modeling center: observation and model-version updates, diagnostic maintenance, compute and storage, audit-panel review, threshold governance, incident handling, and scientific error-review meetings.","confidence":"LOW","assumptions":["Approximately 1–3 loaded FTE equivalents plus reviewer allocations and computing.","Run frequency, archive volume, variable coverage, and audit sampling rate dominate variance.","No new observing system, model-development program, or major HPC acquisition is included."],"source_ids":["S3","S6","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent sources document rapidly increasing model complexity and data volume, extensive intervention in current evaluation workflows, and institutional demand for faster comprehensive assessment. The exact prevalence of attention-related missed failures remains an evidence gap but does not negate the visible general problem.","source_ids":["S2","S3","S4","S5"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"DOE E3SM is a concrete model-development organization with an established diagnostics package, while WCRP's Model Benchmarking and Evaluation Task Team and DOE-supported CMEC are credible coordinating or funding authorities.","source_ids":["S4","S5","S6"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim specifies two comparators, a review-time improvement threshold, non-inferiority tolerance, planted-failure recall, exact reconstruction, and audit-triggered decompression; it is narrower than the established tooling and can fail.","source_ids":["S1","S2","S4","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A single frozen run, variable family, two regions, held-out season, fixed reviewer pool, scripted perturbations, and unchanged complete-field package bound the experiment without changing authoritative outputs.","source_ids":["S2","S3","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"A retrospective internal shadow evaluation can remain subordinate to complete-field review, prohibit automatic model changes or publication claims, and stop on provenance, reconstruction, coverage, or audit failure. Observation licenses and local data authority still require routine confirmation.","source_ids":["S6","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four bands distinguish experiment, prototype, production launch, and recurring operation; public wage data anchor labor while mature open infrastructure constrains technical scope. Confidence is lower for production because local HPC, archive, and reviewer arrangements are unknown.","source_ids":["S2","S3","S6","S8"]}},"next_evidence_step":"Obtain one modeling-center partner and preregister a retrospective crossover study on one frozen run, one land-atmosphere variable family, two contrasting regions, and one held-out seasonal cycle. Randomly assign blinded reviewers to: A) the unchanged complete-field package; B) an ILAMB/ESMValTool-style hierarchical diagnostics package without residual queue, shadow sampling, or decompression; and C) the candidate residual-audit workflow. Seed known mean-bias, phase-shift, variance, cross-variable, missingness, sparse-region, and version-mismatch failures without revealing locations. Independently adjudicate natural material mismatches from complete fields. Measure reviewer time, inspected items, planted-failure recall, natural-mismatch recall, false escalations, classification agreement, exact reconstruction, audit discoveries, sparse-region coverage, fallback correctness, compute, and maintenance effort. Advance only if C reduces median review time by at least 25% versus A, detects every planted material failure, loses no more than five percentage points of adjudicated natural-mismatch recall versus A, reconstructs every sampled comparison exactly, and correctly decompresses every scripted validity failure. Falsify the intervention if any planted material process failure is suppressed, missingness is treated as agreement, a sampled comparison cannot be reconstructed, a material audit mismatch fails to decompress, sparse/high-uncertainty regimes are systematically omitted, or total labor and compute exceed A at the required fidelity.","blocking_evidence":["No comparative human-review study establishes that the proposed queue reduces time while preserving or improving discovery of material process failures.","No partner has committed a frozen run, observation releases, reviewers, or authority to execute the shadow test.","Layer-specific explanatory contracts, consequence weights, multiple-testing controls, materiality thresholds, and suppressed-error budgets are not empirically calibrated.","Independence and licensing of the observation and raw-audit paths have not been verified for a concrete deployment.","The eight-source search cannot establish absence of closer implementations, world novelty, patentability, freedom to operate, market size, or realized impact.","Production and recurring costs remain sensitive to local archive size, HPC accounting, reviewer seniority, variable coverage, and audit rate."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"World novelty is unmeasured. The search establishes substantial collision with ILAMB, ESMValTool, CMEC/PMP, CMIP REF, and E3SM Diags, and only a narrow unresolved contrast around governed hierarchical residual queues, independent sampling of suppressed full fields, exact reconstruction, and automatic scope decompression. It does not measure patentability, freedom to operate, market size, or realized impact.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure written participation from a climate-model evaluation lead and observation-data steward for one bounded retrospective shadow study.","Preregister the three-arm comparison, hierarchy, uncertainty treatment, materiality rules, planted perturbations, error budget, audit sampling, multiple-testing controls, and analysis plan before reviewers see results.","Demonstrate at least 25% lower median reviewer time than complete-field review, 100% recall of preregistered material planted failures, and no more than five percentage points lower recall of independently adjudicated natural material mismatches.","Demonstrate exact reconstruction for all sampled cases, correct handling of missingness and version mismatch, representation of sparse regions, and successful decompression for every scripted or audit-detected validity failure.","Produce observed labor, compute, storage, maintenance, false-escalation, and fallback costs and compare them with both the complete-field baseline and the established-tool dashboard comparator."],"reason":"Bounded web research confirms a real evaluation burden, credible institutional adopters, mature implementation components, and substantial prior-art collision. The only decision-relevant remaining advantage—saving expert attention without suppressing consequential structured errors—requires proprietary run data and a live reviewer comparison, not additional web search. Therefore empirical partnered research is required."},"proposal_index":2}