{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","cell_id":"predictive_residual_processing__environmental_climate","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_id":"prp-hierarchical-climate-model-residual-audit-002","proposal_index":2,"version":0,"title":"Hierarchical Residual Audit for Climate-Model Process Failures","problem":"A climate-model development team must compare simulated land-atmosphere fields with observations across many locations, variables, seasons, and spatial scales. Complete-field plot collections consume scarce scientific review time, while aggregate error scores can collapse geographically or seasonally coherent mismatches into acceptable averages. The team can therefore complete an evaluation without identifying which unresolved residual patterns indicate measurement limitations, local representativeness errors, or a misspecified modeled process.","actors":["Climate-model evaluation lead","Land-atmosphere process scientist","Observation-data steward","Regional environmental scientist","Independent evaluation reviewer"],"observable_state":"For each frozen model run and evaluation window, the system retains the versioned predicted field, provenance-bearing observations, signed point-level residuals, observation and model uncertainty, hierarchical assignments from site to catchment or ecoregion to region, residuals explained or forwarded at each layer, analyst-review dispositions, raw-audit selections, and reconstruction tests. Every quiet layer has explicit coverage and missingness states, so absence of an escalated residual is distinguishable from absent observations or failed processing.","consequence":"Review capacity may be spent repeatedly examining expected spatial and seasonal structure, while coherent errors confined to a transition, sparse region, or coupled variable relationship remain unexamined or are diluted by aggregation. This weakens the evidentiary basis for accepting, revising, or delimiting the model configuration.","affected_objective":"Identify decision-relevant, structured model-observation mismatches within a bounded scientific-review budget while preserving complete reconstructability, uncertainty, regional context, and independent access to observations the evaluation model may explain away.","intervention":"Freeze a candidate climate-model run, its evaluation targets, spatial hierarchy, uncertainty model, residual comparator, and version identifier before observing evaluation results. At the lowest layer, compare predicted and observed variables at eligible sites or grid supports. Remove only preregistered measurement-noise and representativeness components, then pass unresolved signed residuals upward to catchment or ecoregion layers. Those layers test whether residuals share timing, direction, covariance, or spatial structure; only precision- and consequence-weighted unresolved patterns enter the process-scientist review queue. Each escalated item includes the predicted baseline, residual field, uncertainty, model version, observation provenance, suppressed-error summary, and reconstruction instructions. Random full-field panels and risk-stratified panels from sparse regions, seasonal transitions, and coupled-variable regimes are reviewed independently. Structured drift, failed reconstruction, incompatible versions, excessive suppressed-error mass, missing coverage, or disagreement with an audit panel decompresses the affected scope to full-field review. Validated residuals enter a replay buffer and a human prediction-error review; any parameter or structural revision is tested as a new model version rather than silently incorporated into the evaluated run.","structural_mapping":[{"archetype_element":"Prediction target definition","domain_realization":"The target is a preregistered set of observed land-atmosphere variables at declared spatial supports and time intervals, such as surface temperature, soil moisture, snow state, evapotranspiration, and precipitation, with units, evaluation seasons, and fidelity requirements fixed before residual inspection."},{"archetype_element":"Generative model state","domain_realization":"A frozen climate-model configuration and run archive produces the expected fields; its code, parameters, forcing data, initialization, and post-processing identifiers form the versioned predictor state."},{"archetype_element":"Model scope and horizon","domain_realization":"The evaluation license is limited to named variables, regions, observation types, temporal aggregations, and forcing conditions. Residual suppression is not authorized outside that envelope."},{"archetype_element":"Expected and actual behavior","domain_realization":"Model output supplies the committed expectation, while independently stewarded observations supply the actual field with timestamps, coverage, calibration, transformation history, and quality flags."},{"archetype_element":"Prediction comparator and error signal","domain_realization":"The comparator retains signed errors, timing offsets, covariance discrepancies, missingness, and categorical transition mismatches instead of reducing all disagreement to one scalar score."},{"archetype_element":"Hierarchical prediction stack","domain_realization":"Site-level residuals feed catchment or ecoregion tests, and unresolved regional structures feed process-level review. Every layer declares what variation it may explain, what it must forward, and which lower-layer context accompanies escalation."},{"archetype_element":"Precision and consequence weighting","domain_realization":"Residual priority reflects observational uncertainty, sampling density, model spread where available, persistence, spatial coherence, variable coupling, and the scientific consequence of the process being evaluated rather than magnitude alone."},{"archetype_element":"Residual reconstruction and provenance","domain_realization":"The archived prediction plus the retained signed residual reconstructs each evaluated observation-model comparison; every residual names its model run, observation release, comparator, hierarchy, and transformation version."},{"archetype_element":"Raw audit and error budget","domain_realization":"Random full-field panels and panels stratified toward sparse regions, seasonal transitions, extremes, and coupled-variable boundaries are reviewed outside the residual-ranking process and compared with a preregistered suppressed-error budget."},{"archetype_element":"Decompression and bounded updating","domain_realization":"Audit disagreement, structured suppressed residuals, version mismatch, stale observations, reconstruction failure, or inadequate coverage restores complete-field review for the affected scope. Model changes require a separately versioned rerun and cannot rewrite the evaluated prediction retrospectively."}],"mechanism_mapping":[{"mechanism_slug":"hierarchical_prediction_error_loop","role":"Organizes model-observation discrepancies from sites through environmental regions to process-level review, allowing expected local structure to stop at the lowest authorized explanatory layer while unresolved structure rises.","counterfactual_removal":"Without the layered contracts and upward residual path, the proposal becomes a flat anomaly list and cannot distinguish local noise from spatially or temporally coherent process error."},{"mechanism_slug":"forecast_backtesting","role":"Freezes the evaluation protocol and tests the model on data and periods excluded from configuration, defining where residual triage may be used.","counterfactual_removal":"Without walk-forward or held-out evaluation, residual thresholds can be tuned to flatter the same run they are supposed to challenge."},{"mechanism_slug":"precision_weighted_error_gate","role":"Allocates the bounded expert-review queue using uncertainty, coverage, coherence, consequence, and attention cost while retaining a record of suppressed residual mass.","counterfactual_removal":"Without it, noisy dense regions can dominate review while reliable small errors or sparsely observed consequential regimes receive little attention."},{"mechanism_slug":"residual_comparison_test","role":"Tests residuals for trend, autocorrelation, seasonal timing, spatial coherence, changing variance, and differences from a rival model or raw audit panel.","counterfactual_removal":"Without it, residual magnitude alone cannot distinguish unstructured imprecision from systematic misspecification."},{"mechanism_slug":"model_version_checksum_handshake","role":"Binds every residual, reconstruction, and replay item to the exact model run, observation release, and transformation pipeline that produced it.","counterfactual_removal":"Without it, residuals can be interpreted against revised fields or observations and yield plausible but invalid reconstructions."},{"mechanism_slug":"shadow_raw_channel_sampling","role":"Selects random and risk-stratified complete-field panels for independent inspection, including cases the production hierarchy did not escalate.","counterfactual_removal":"Without it, the residual system grades only the mismatches it already recognizes and cannot reveal systematic explaining-away."},{"mechanism_slug":"prediction_error_replay_buffer","role":"Stores escalated residuals, sampled non-escalated cases, predictions, observations, and provenance for later regression testing of revised model versions.","counterfactual_removal":"Without it, scientific review findings cannot be replayed consistently and later configurations may repeat or merely conceal earlier failures."},{"mechanism_slug":"prediction_error_review","role":"Requires scientists to classify material mismatches as observation, representativeness, forcing, parameter, process, boundary, or unresolved problems before authorizing changes.","counterfactual_removal":"Without it, residual escalation produces visual attention but no governed inference or attributable revision decision."},{"mechanism_slug":"bayesian_model_update","role":"Provides an optional bounded method for revising parameter or process-hypothesis uncertainty after residuals survive provenance, comparison, and review checks.","counterfactual_removal":"Without an uncertainty-preserving update rule, accepted residuals may prompt ad hoc parameter changes that erase uncertainty or overreact to one evaluation window."},{"mechanism_slug":"raw_signal_fallback_switch","role":"Restores full-field plots and uncompressed observation-model comparisons whenever audit, coverage, validity, compatibility, or cumulative-error conditions fail.","counterfactual_removal":"Without it, an invalid hierarchy or comparator can continue filtering the evidence needed to diagnose its own failure."}],"causal_chain":["A frozen model run generates explicit expectations for declared environmental variables, places, and times.","Independent observations are aligned to those expectations with provenance and uncertainty preserved.","Signed and structured residuals expose what the model did not predict rather than repeating the expected simulated field.","Preregistered lower layers remove only authorized noise and representativeness components; unresolved residual structure propagates upward.","Precision and consequence weighting directs bounded scientific attention toward residuals that are reliable, coherent, and relevant to a declared process question.","Escalated residuals arrive with reconstructed baseline context, enabling reviewers to distinguish isolated measurement issues from candidate process misspecification.","Random and risk-stratified full-field audits test whether the hierarchy suppressed consequential patterns or marginalized poorly sampled regimes.","Validated errors enter a versioned review and replay process; any accepted model change produces a new predictor whose residuals can be compared with the frozen predecessor.","Failed compatibility, reconstruction, coverage, drift, or audit checks decompress the affected scope to complete-field evaluation."],"baseline":"A fixed evaluation package containing complete model and observation maps, time-series panels, and aggregate bias and error metrics for every declared variable and region. Reviewers inspect the package without residual prioritization, hierarchical propagation, an explicit attention budget, or independent sampling of panels omitted from detailed discussion.","nearest_rivals":["An aggregate climate-model scorecard summarizes bias, correlation, and error by variable or region but does not preserve structured residuals as the review message or require reconstruction and raw-panel audits.","A flat anomaly detector ranks unusual grid cells or time steps but lacks layer-specific explanatory contracts, upward propagation of unresolved error, and governed model revision.","An ensemble-disagreement display identifies where model configurations differ, but agreement does not establish consistency with observations and the display does not audit what all models jointly suppress.","Manual expert browsing preserves full context but provides no explicit residual budget, reproducible prioritization rule, model-version binding, or test of what limited attention failed to inspect.","Automated parameter calibration minimizes a loss function but can absorb structural error into parameters and does not require human classification, independent raw review, or decompression when the evaluation representation fails."],"remaining_contrastive_claim":"This candidate makes hierarchical, reconstructible model-observation residuals the governed unit of scientific attention and model learning: expected simulated structure is available from a frozen predictor, only unresolved error rises across explicit environmental scales, and independent full-field audits retain authority over what the predictor suppresses.","authority_safety":{"decision_authority":"The climate-model evaluation lead may approve the frozen protocol and shadow comparison. The observation-data steward controls observation provenance and exclusions. Process leads may propose revisions, but the project science authority must approve any new model configuration or external validation statement.","authorized_first_step":"Apply the residual workflow retrospectively to one completed model run, one declared variable family, a bounded set of regions, and a held-out observation window while the complete-field evaluation remains authoritative.","excluded_actions":["Replace an operational environmental forecast or assessment product during the test","Delete, overwrite, or down-rank the authoritative model outputs or observations","Automatically retune parameters or alter model structure from residuals","Declare model validity, causal process failure, or observational error from residual ranking alone","Exclude sparse regions or observation types merely because their uncertainty lowers throughput","Publish comparative performance claims from the bounded shadow evaluation","Use review-queue size as the target for threshold tuning"],"halt_rollback":"Stop residual triage and return the affected variable, region, or time window to complete-field review if any comparison cannot be reconstructed, provenance or version checks fail, an independent panel reveals a material suppressed pattern, hierarchy assignments are disputed, missingness is treated as agreement, or fallback tests fail. Preserve the frozen run and logs; discard only the unapproved triage disposition and restart from the preregistered full-field package."},"negative_tests":{"strongest_counterevidence":"The complete-field baseline enables reviewers to identify and classify the same consequential structured mismatches with less total analyst time and lower processing burden, while the hierarchical system either adds no distinct diagnosis or suppresses errors that the baseline exposes.","problem_falsifier":"The bounded evaluation contains few enough model-observation comparisons for complete expert inspection, or aggregate metrics and existing panels preserve all distinctions needed for the model decision, so scientific attention is not meaningfully constrained by repeated expected structure.","intervention_falsifier":"The system cannot reconstruct evaluated comparisons from the frozen prediction and residual archive; preregistered timing, bias, covariance, or spatial-pattern perturbations fail to reach the correct review layer; random full-field panels repeatedly reveal unforwarded structured error; sparse or high-uncertainty regimes are systematically omitted; version and fallback tests do not fail safely; or model, audit, and review costs exceed those of complete-field evaluation at the required fidelity.","risks":["Strong model priors at a low layer may explain away the evidence of a real process failure.","Dense observation regions may dominate precision weights and marginalize sparsely monitored environmental regimes.","A hierarchy based on existing process categories may prevent genuinely cross-boundary residuals from reaching an appropriate reviewer.","Repeated residual testing can manufacture apparent structure through multiple comparisons.","Reviewers may mistake a coherent residual for proof of a particular causal mechanism.","Parameter updates may absorb forcing, observation, or structural errors and make the next residual field look quieter without making the model more defensible.","A shared observation-processing error can affect both production comparisons and nominally independent audits.","Visible queue metrics may induce threshold changes aimed at comfortable workload rather than the declared error budget.","Residual archives may lose baseline spatial context if provenance or reconstruction dependencies are not preserved."]},"next_evidence_step":"Preregister a retrospective shadow evaluation using one frozen model run, one related variable family, two contrasting environmental regions, and one held-out seasonal cycle. Freeze the spatial hierarchy, comparator, uncertainty treatment, error budget, threshold table, and observation release. Include random full-field audit panels, panels stratified toward sparse coverage and seasonal transitions, one unchanged baseline package, one rival-model residual comparison, and scripted test cases for mean bias, seasonal phase shift, changing variance, coupled-variable inconsistency, missing observations, and version mismatch. Have reviewers record which items they inspect, their classifications, time spent, reconstruction failures, audit discoveries, fallback use, and proposed model actions. The test determines whether the representation and safeguards are workable; it does not establish model validity or scientific effect.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Proposal 1 addressed live transmission and energy constraints at a remote peatland-rewetting sensor station. It used matched edge/server predictors to reconstruct high-frequency field measurements and route operational departures to monitoring or safety owners. This proposal addresses a different object, bottleneck, actor set, and decision: retrospective scientific evaluation of climate-model simulations under limited expert attention. Its intervention is a hierarchical model-observation residual audit across spatial and process layers, and its causal endpoint is an attributable decision to retain, delimit, or retest a model configuration—not reduced station traffic, field-event detection, or land management. It can be adopted by a modeling group without deploying remote sensors or changing any environmental monitoring link, while Proposal 1 can operate without a climate-model development workflow.","revision_record":{"parent_version":null,"progress_targets_addressed":["Additional independently adoptable candidate","Materially different environmental problem","Distinct hierarchical causal path","Operational authority and safeguards","Explicit contrast with proposal 1","Bounded falsifiable first evidence"],"conceptual_changes":["Created a scientific model-evaluation opportunity rather than extending the remote peatland telemetry proposal","Made hierarchical unresolved residual propagation, not edge transmission, the central intervention"],"operational_changes":["Defined frozen evaluation versions, layer contracts, independent full-field panels, replay, human error classification, and scope-specific decompression"],"evidence_changes":["Specified a bounded retrospective shadow evaluation with held-out observations, a rival model, random audits, and scripted failure tests"],"claim_changes":["Made no novelty, prevalence, demand, model-validity, or effect-size claim; limited the first test to representational feasibility and safeguard performance"]}}