{"schema_version":1,"assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","source_experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"computability_boundary_mapping__astronomy_astrophysics","archetype_slug":"computability_boundary_mapping","domain_slug":"astronomy_astrophysics","title":"Guarantee-tiered routing for astrophysical eventual-outcome queries","opportunity_summary":"Audit and formalize one astrophysical model language and event predicate, restrict exact classification to enforceable complete fragments, and route other cases to certified YES, bounded UNKNOWN, out-of-scope, or numerical-failure outputs instead of treating timeout as NO. The opportunity is conditional on verifying that a deployed service actually promises unrestricted, exact, terminating Boolean classification.","adopter_authorizer":"The simulation platform's scientific owner can authorize an audit and interface pilot; any computability verdict requires independent formal-methods review and domain-science review.","scores":{"meaningful_impact":{"score":4,"rationale":"If the stated Boolean interface converts unresolved trajectories into negative outcomes, separating NO from UNKNOWN and numerical failure could materially improve scientific defensibility and prevent effort directed at an impossible total-classifier requirement. The frequency and consequence of such misclassification are not established."},"stakeholder_pull":{"score":2,"rationale":"Researchers, catalog users, reviewers, and platform operators have identifiable reasons to value honest classifications, but the sealed candidate provides no evidence of complaints, demand, adoption commitments, or prevalence of the alleged unrestricted guarantee."},"incremental_advantage":{"score":3,"rationale":"Compared with run-until-event-or-timeout simulation and numerical uncertainty analysis, the proposed formal guarantee audit and tiered abstaining interface address a different failure mode: unsupported total-exact claims. Advantage is conditional because existing workflows may already bound inputs and distinguish timeout from NO."},"distinctiveness_plausibility":{"score":2,"rationale":"The combination of enforceable fragments, certificates, explicit abstention, and versioned evidence is coherent, but prior art is unsearched and the packet supplies no comparison with existing reachability tools, simulation contracts, or abstaining scientific workflows."},"technical_implementability":{"score":3,"rationale":"Adding explicit output labels and replaying a bounded case set appears technically feasible, while formalizing model semantics, enforcing fragment membership, preserving the intended event predicate, and producing valid certificates could require substantial specialized work."},"adoption_authority_feasibility":{"score":4,"rationale":"A platform scientific owner is explicitly identified as able to approve interface labels and the pilot, with formal and domain reviewers assigned authority over stronger computability conclusions. Downstream label preservation remains an implementation dependency."},"evidence_readiness":{"score":3,"rationale":"The candidate provides a bounded audit, a replay design, explicit counterevidence, and separate problem and intervention falsifiers. It supplies no current interface specification, timed-out case set, measured misclassification rate, or domain-specific proof."},"safety_net_benefit":{"score":5,"rationale":"Explicit UNKNOWN, OUT_OF_SCOPE, and NUMERICAL_FAILURE outputs directly reduce false certainty while halt and rollback rules limit harm from semantic mismatch or label collapse. The pilot does not require asserting undecidability or treating model certification as physical truth."},"scalability":{"score":3,"rationale":"The guarantee-tiering pattern could be reused across predicates and model languages, but each extension may need new semantics, fragment proofs, membership enforcement, certificates, and domain review, limiting low-cost scaling."}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Formalize one model language and one event predicate, inspect the current interface contract, define comparison labels, assemble and replay a bounded set of previously timed-out cases, and obtain limited formal and domain review.","confidence":"LOW","assumptions":["Existing specifications, code, and timed-out cases are accessible.","The audit uses existing compute infrastructure.","One formal-methods specialist, platform staff, and domain reviewers participate part-time.","No new theorem-proving platform or large simulation campaign is required."]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Implement enforceable input fencing, tiered output labels, audit logging, certificate-checking hooks, downstream schema changes, documentation, and validation for one restricted production pathway.","confidence":"LOW","assumptions":["The service architecture permits interface and validation changes.","Certificate checking is limited to one defined fragment.","Downstream consumers can accept revised schemas without wholesale replacement.","Security, legal, and compliance requirements are routine research-platform requirements."]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Launch the restricted exact mode and abstaining routes across a working service, including integration testing, user migration, reviewer coordination, monitoring, governance, training, and rollback readiness.","confidence":"LOW","assumptions":["Launch covers one platform rather than a multi-institution standard.","Multiple downstream catalogs or workflows require coordinated changes.","Formal and domain reviewers validate the release boundary.","No astronomy-specific undecidability proof is required for launch."]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Maintain semantic versions, fragment definitions, certificate checkers, regression cases, monitoring, reviewer cycles, user support, and revalidation when model expressiveness or assumptions change.","confidence":"LOW","assumptions":["The initial scope remains one platform with a limited number of supported fragments.","Reviews are periodic rather than continuous.","Compute demand is comparable to existing bounded replay and validation workloads.","Major new model languages are separately funded."]}},"research_burden":"HIGH","earliest_credible_horizon":"3_TO_12_MONTHS","pipeline_gates":{"recognizable_externally_supportable_problem":{"status":"UNCERTAIN","reason":"The candidate specifies an observable and falsifiable interface failure—timeouts or inconclusive runs becoming Boolean negatives—but supplies no external evidence that an operational astrophysical service makes the unrestricted total-exact promise or exhibits this behavior."},"identifiable_adopter_or_authorizer":{"status":"YES","reason":"The simulation platform's scientific owner is explicitly authorized to approve interface labels and the pilot, with independent formal and domain reviewers required for computability conclusions."},"distinct_testable_incremental_claim":{"status":"YES","reason":"For one formalized language and predicate, tiered routing can be compared with the current Boolean treatment on the same timed-out cases to test whether it reduces unsupported negative conclusions while preserving event semantics and useful coverage."},"bounded_next_evidence_step":{"status":"YES","reason":"The authorized audit is limited to one model language, one event predicate, and a bounded replay set, with explicit comparison outputs and problem and intervention falsifiers."},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step is non-decisional and reversible; it prohibits unsupported undecidability claims and timeout-as-NO reporting, requires appropriate reviews, and includes halt and rollback conditions for label collapse, unenforceable membership, or semantic divergence."},"implementation_cost_scope_and_range":{"status":"YES","reason":"The candidate identifies the principal implementation units—formalization, fragment enforcement, certificates, interface labels, replay evaluation, reviews, monitoring, and downstream coordination—supporting broad resource bands despite low cost confidence."}},"blocking_evidence":["The current service specification must show whether model language, representations, horizons, and guarantees are actually open-ended or enforceably finite and bounded.","Current interface behavior must establish whether timeout, UNKNOWN, out-of-scope inputs, and numerical failure are presently collapsed into NO.","A bounded set of representative timed-out or inconclusive cases and their downstream handling must be available for comparison.","Formal and domain reviewers must agree that the encoded event predicate preserves the intended scientific semantics.","The proposed exact fragment must have enforceable membership and a reviewed terminating procedure or certificate-checking rule.","Downstream consumers must demonstrate that the new labels remain distinct rather than being recoded as Boolean negatives."],"next_evidence_step":"Audit one deployed model language and one event predicate, then replay a preregistered bounded sample of previously timed-out cases through both the existing Boolean interface and the proposed YES_WITH_CERTIFICATE, UNKNOWN_BOUND, OUT_OF_SCOPE, and NUMERICAL_FAILURE routing. Falsify the motivating problem if the specification is enforceably finite and bounded, already preserves UNKNOWN separately from NO, and makes no class-wide total-exact claim; reject the intervention if unsupported Boolean conclusions do not decrease or scientifically intended event semantics and useful coverage cannot be preserved.","research_questions":["Does the deployed interface actually promise exact, terminating YES/NO classification for every admitted model and initial condition?","Are model syntax, numerical representation, external evidence, and time horizons enforceably bounded in practice?","How are timeout, numerical failure, out-of-scope inputs, and inconclusive evidence currently represented downstream?","What proportion of audited inconclusive cases are currently interpreted as negative outcomes, without extrapolating beyond the bounded sample?","Can one useful exact fragment be defined with enforceable membership and a reviewed terminating procedure?","Can certified event witnesses be checked independently without implying that the encoded model confirms physical reality?","Do tiered labels preserve the intended astrophysical event semantics and provide enough coverage to justify integration burden?","Will catalog and review workflows preserve UNKNOWN and other non-Boolean labels?","Which elements of guarantee-tiered routing are already established in prior systems, and what incremental feature, if any, is distinct?"],"recommendation":"VALIDATE_PROBLEM_FIRST","uncertainty_constraints":["The central undecidability proposition is only a hypothesis; chaos, continuum mathematics, sensitivity, and long runtime are not proofs.","The operational existence and prevalence of an unrestricted total-exact requirement are unsupported.","Finite-precision states, fixed horizons, and finitely encoded integrators could make the relevant workflow decidable in principle, leaving only complexity or numerical-validity problems.","A proof about an encoded model does not establish the behavior of the physical universe.","Stakeholder demand, realized impact, market size, and adoption willingness are unmeasured.","Prior-art position and distinctiveness are unmeasured.","Cost bands are resource-equivalent planning ranges, not observed vendor or program prices."],"closed_book_prior_art_boundary":"Prior art is unsearched and unverified in this closed-book assessment. No claim is made about novelty, prevalence, existing reachability systems, simulation-service practices, or whether the proposed combination has already been implemented."}