{"schema_version":1,"assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","source_experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"computability_boundary_mapping__geoengineering_planetary_science","archetype_slug":"computability_boundary_mapping","domain_slug":"geoengineering_planetary_science","title":"Label-preserving verification boundary for geoengineering controllers","opportunity_summary":"Evaluate an offline assurance interface that admits controller/model pairs only to a proved decidable fragment, routes other inputs to bounded fallbacks, and preserves UNKNOWN, TIMEOUT, and OUT_OF_SCOPE labels. The mechanism could prevent bounded analysis from being presented as universal certification, but the packet does not establish that real geoengineering workflows exhibit this failure or that a useful fragment would cover decision-relevant controllers.","adopter_authorizer":"A geoengineering assurance program's analysis or independent-review team could adopt the interface; research sponsors and regulators could authorize evaluation, while environmental deployment authority remains outside the pilot.","scores":{"meaningful_impact":{"score":3,"rationale":"Incorrectly clearing an unsafe controller or overstating planetary-safety evidence could have serious consequences, but the packet provides no evidence that an actual program demands the universal verdict at issue or how many decisions would be affected."},"stakeholder_pull":{"score":2,"rationale":"Modelers, independent reviewers, sponsors, and regulators are plausible stakeholders, but no documented adopter demand, requirements artifact, workflow observation, or willingness to change interfaces is supplied."},"incremental_advantage":{"score":3,"rationale":"A proved verifier, enforced scope boundary, and preserved non-Boolean labels directly improve on the stated Boolean simulation baseline, but comparative coverage, runtime, false alarms, and reviewer usefulness are entirely untested."},"distinctiveness_plausibility":{"score":2,"rationale":"The combination is coherently specified, but prior art is explicitly unsearched and distinctiveness from formal verification, robust-control assurance, and model-governance practice cannot be inferred closed-book."},"technical_implementability":{"score":3,"rationale":"Defining a finite fragment and testing a total verifier on toy or synthetic pairs is plausible, while sound abstraction of executable planetary models, useful controller coverage, and tractable verification remain unresolved."},"adoption_authority_feasibility":{"score":3,"rationale":"The pilot team has authority to define an offline interface and recommend assurance language, with clear deployment exclusions; operational adoption would still require unidentified program owners, reviewers, and regulators to accept the restrictions and labels."},"evidence_readiness":{"score":3,"rationale":"The packet supplies observable states, separate problem and intervention falsifiers, halt rules, and a bounded offline design, but supplies neither workflow evidence nor a representative archived corpus and comparative metrics."},"safety_net_benefit":{"score":4,"rationale":"Explicit UNKNOWN, TIMEOUT, and OUT_OF_SCOPE states, prohibition on treating abstention as SAFE, guarantee withdrawal, and reversion to non-certifying use provide a strong procedural safety net even if the verifier is incomplete."},"scalability":{"score":2,"rationale":"Scaling is threatened by potentially narrow fragment coverage, expensive abstractions, false alarms, and computational complexity; the packet contains no evidence that the mechanism remains useful across models, controllers, or assurance programs."}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"A bounded audit of requirements, interface specifications, and decision records from one accessible assurance setting, including stakeholder interpretation of SAFE, UNKNOWN, TIMEOUT, and scope labels.","confidence":"LOW","assumptions":["One assurance setting grants access to a small predeclared artifact set.","The step is an offline requirements audit rather than verifier construction.","Costs include analyst labor, partner coordination, data handling, and synthesis."]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Construct and independently review one offline prototype: define a finite controller/model fragment, formalize the verifier contract, implement label-preserving routing, and compare it with the stated baseline on toy and archived synthetic pairs.","confidence":"LOW","assumptions":["The fragment and property are deliberately narrow.","No physical intervention or deployment authorization is included.","Formal-methods, domain-modeling, software, proof-review, and evaluation labor are included."]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Integrate a reviewed verifier and fallback interface into one assurance program, validate model adapters and downstream label preservation, train reviewers, and establish governance, audit, and rollback procedures.","confidence":"LOW","assumptions":["Launch is limited to decision support in one program.","Existing model and controller infrastructure can be adapted rather than rebuilt.","Independent scientific, software, compliance, and governance review is required.","Verifier output is not the sole basis for environmental authorization."]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain proofs, model and controller adapters, regression suites, assurance records, reviewer training, audits, and reclassification when languages, properties, or models change.","confidence":"LOW","assumptions":["One operational assurance program is supported.","Material changes trigger renewed proof and validation work.","Costs include specialist labor, compute, software, partner coordination, and independent review."]}},"research_burden":"HIGH","earliest_credible_horizon":"3_TO_12_MONTHS","pipeline_gates":{"recognizable_externally_supportable_problem":{"status":"UNCERTAIN","reason":"The failure mode is recognizable and operationally specified, but the packet contains no external or target-workflow evidence that unrestricted exact verdicts are demanded or that UNKNOWN, TIMEOUT, and SAFE are collapsed."},"identifiable_adopter_or_authorizer":{"status":"YES","reason":"The sealed candidate identifies the pilot analysis team as able to define the interface and recommend assurance language, with independent reviewers, sponsors, and regulators as adoption or authorization roles."},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal can be tested against the Boolean simulation baseline on soundness, admitted coverage, label correctness, abstention, runtime, false alarms, and reviewer usefulness."},"bounded_next_evidence_step":{"status":"YES","reason":"An offline audit of a predeclared assurance setting can test the stated problem falsifier without environmental deployment, followed only conditionally by the specified toy or synthetic comparison."},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"Physical deployment, SAFE treatment of unknown results, silent out-of-fragment admission, and sole reliance on verifier output are excluded, and concrete proof, trace, and label failures trigger withdrawal and rollback."},"implementation_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The packet bounds a toy pilot conceptually but provides no staffing, corpus, proof-system, integration, compliance, or maintenance requirements from which an externally supportable implementation range could be established."}},"blocking_evidence":["Evidence that at least one real assurance workflow demands or implies an unrestricted exact terminating verdict, or collapses UNKNOWN, TIMEOUT, OUT_OF_SCOPE, and SAFE.","Evidence that a certified fragment covers a useful share of decision-relevant controller/model pairs without unacceptable false alarms or reviewer bypass.","Comparative evidence that the mechanism adds useful assurance over conservative simulation and expert review at comparable resources.","A bounded prior-art review establishing whether the proposed mechanism or combination is materially distinct.","Program-specific implementation requirements sufficient to validate resource and integration ranges."],"next_evidence_step":"Audit a predeclared, access-bounded set of requirements, interface specifications, and decision records from one candidate geoengineering assurance setting; compare its actual quantifiers and handling of UNKNOWN, TIMEOUT, and OUT_OF_SCOPE with the proposal's observable state, and falsify the problem if the setting requires only bounded or probabilistic analysis, makes no universal exactness claim, and already preserves all non-safe labels.","research_questions":["Does any identified assurance workflow require a correct, terminating verdict over an open-ended controller/model class?","Where, if anywhere, are UNKNOWN, TIMEOUT, or OUT_OF_SCOPE results converted into Boolean approval recommendations?","What controller, model, property, horizon, and numeric-encoding restrictions define a useful certified fragment?","What proportion of a predeclared representative corpus would the fragment admit?","Does the verifier miss any admitted unsafe trace or assign an exact verdict to an out-of-fragment input?","How do coverage, abstention, runtime, false alarms, and reviewer behavior compare with conservative simulation and expert review?","Can model-semantic assumptions and reclassification triggers remain visible through downstream approval interfaces?","Which elements, if any, are absent from established formal-verification, robust-control, and planetary-model governance practice?"] ,"recommendation":"VALIDATE_PROBLEM_FIRST","uncertainty_constraints":["Problem prevalence and stakeholder demand are unsupported hypotheses.","Prior art and world distinctiveness are unmeasured.","Planetary-model fidelity may dominate computability concerns.","Useful fragment coverage and abstraction soundness are untested.","Operational costs cannot be established exactly from the sealed packet.","No conclusion about market size or realized impact is available closed-book."],"closed_book_prior_art_boundary":"No external sources or prior-art search were used. The candidate marks prior art as UNSEARCHED, so no claim is made about novelty, prevalence, existing practice, market size, realized impact, or exact cost; distinctiveness is assessed only as an unsupported plausibility within the sealed packet."}