{"schema_version":1,"assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","source_experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"computability_boundary_mapping__astronomy_astrophysics","archetype_slug":"computability_boundary_mapping","domain_slug":"astronomy_astrophysics","title":"Tiered Reachability Guarantees for Astrophysical Simulation Services","opportunity_summary":"Audit whether an astrophysical simulation service actually promises exact terminating eventual-outcome classifications over an unrestricted model class, then replace unsupported Boolean conclusions with enforceable exact fragments, certified event witnesses, bounded approximations, and explicit UNKNOWN, OUT_OF_SCOPE, or NUMERICAL_FAILURE outputs. The opportunity is conditional because the packet does not establish that the unrestricted promise or timeout-as-NO practice occurs in a real workflow.","adopter_authorizer":"The simulation platform's scientific owner can authorize the interface audit and replay pilot; computability conclusions additionally require independent formal-methods review and domain-scientific review.","scores":{"meaningful_impact":{"score":4,"rationale":"If the stated interface exists, preventing unresolved trajectories from being labeled as impossible events would materially improve scientific defensibility and downstream catalog integrity. The packet does not establish how often such errors occur or how consequential they have been."},"stakeholder_pull":{"score":2,"rationale":"Researchers, platform operators, reviewers, and catalog users have plausible reasons to value honest uncertainty labels, but the sealed candidate contains no demonstrated requests, complaints, adoption commitments, or evidence that the described unrestricted service contract is in use."},"incremental_advantage":{"score":3,"rationale":"The proposed routing directly improves on the stated Boolean timeout baseline by separating certificates, bounded results, failures, and unknowns. Its advantage is uncertain because existing workflows may already use bounded horizons and non-Boolean uncertainty handling, and the exact fragment may provide little useful coverage."},"distinctiveness_plausibility":{"score":2,"rationale":"The integrated application of computability-boundary mapping to an astrophysical simulation interface is coherent, but prior-art status is explicitly UNSEARCHED and the packet supplies no basis for distinguishing it from existing abstaining workflows, reachability tools, or simulation contracts."},"technical_implementability":{"score":4,"rationale":"Auditing one model language and event predicate, introducing differentiated labels, and replaying a bounded case set are technically bounded and reversible. Proving fragment properties, enforcing membership, and preserving astrophysical semantics could still require difficult specialist work."},"adoption_authority_feasibility":{"score":4,"rationale":"The packet identifies the platform scientific owner as able to approve interface labels and the pilot, while reserving computability conclusions for formal and domain review. Feasibility is reduced by dependence on downstream consumers preserving the new labels."},"evidence_readiness":{"score":3,"rationale":"The candidate specifies a concrete audit, bounded replay, comparison outputs, problem falsifier, and intervention falsifier. It does not supply an actual service specification, historical case inventory, baseline misclassification rate, formal encoding, or agreed scientific semantics."},"safety_net_benefit":{"score":5,"rationale":"Explicit UNKNOWN, OUT_OF_SCOPE, and NUMERICAL_FAILURE outputs directly prevent timeout from becoming a false negative, while exclusions, review requirements, audit preservation, halt conditions, and rollback to non-decisional outputs provide strong safeguards."},"scalability":{"score":3,"rationale":"The labeling and routing pattern could be reused across predicates and model languages, but each extension may require new formal semantics, enforceable fragment checks, certificate procedures, domain validation, and evidence versioning."}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Audit one model language and one event predicate, formalize the claimed guarantees and encodings, select and replay a bounded set of timed-out cases, and conduct formal and domain review of the resulting labels.","confidence":"LOW","assumptions":["A usable service specification and historical timed-out cases can be accessed.","The audit is limited to one language, one predicate, and a modest replay set.","Specialist formal-methods and astrophysics labor is included.","No new large-scale simulation campaign or formal undecidability proof is required."]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Implement pilot-only output labels, provenance records, fragment-membership checks, basic certificate validation, evaluation instrumentation, and reversible integration with one simulation workflow.","confidence":"LOW","assumptions":["The existing platform can represent non-Boolean outcomes without architectural replacement.","Certificate checking and fragment enforcement remain limited to the audited regime.","Security, compliance, and data-access requirements are ordinary research-platform requirements.","This band excludes organization-wide rollout."]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Productionize the routing interface for the initial supported service, validate downstream label preservation, document guarantees, train users, establish review and incident procedures, and complete launch evaluation.","confidence":"LOW","assumptions":["Launch covers one platform rather than an astronomy-wide standard.","Downstream catalogs and clients require integration work but not complete reconstruction.","Formal and scientific reviewers are available.","Useful exact or bounded coverage is demonstrated before launch."]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Maintain semantics, certificate checkers, fragment rules, evidence versions, user documentation, monitoring, audits, and periodic formal and domain review as models or assumptions change.","confidence":"LOW","assumptions":["The supported model and predicate set grows slowly.","Compute-intensive revalidation is selective rather than exhaustive.","No continuous large-scale proof-development program is required.","Recurring partner coordination and evaluation labor are included."]}},"research_burden":"HIGH","earliest_credible_horizon":"3_TO_12_MONTHS","pipeline_gates":{"recognizable_externally_supportable_problem":{"status":"UNCERTAIN","reason":"The packet clearly describes a falsifiable problem, but it does not establish that a real platform promises universal exact termination, collapses timeout into NO, or produces consequential unsupported classifications."},"identifiable_adopter_or_authorizer":{"status":"YES","reason":"The simulation platform's scientific owner is explicitly authorized to approve the audit, labels, and pilot, with independent formal and domain reviewers identified for stronger verdicts."},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal makes a bounded incremental claim that tiered routing will reduce unsupported Boolean conclusions while preserving intended event semantics, testable against the current Boolean handling on replayed timed-out cases."},"bounded_next_evidence_step":{"status":"YES","reason":"Auditing one language and predicate and replaying a bounded historical case set is reversible, non-live, comparison-based, and includes both problem and intervention falsifiers."},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step stays within platform ownership, forbids an unsupported undecidability claim, requires formal and domain review, and includes halt and rollback conditions for semantic divergence or downstream label collapse."},"implementation_cost_scope_and_range":{"status":"YES","reason":"The pilot and broader implementation can be scoped around one language, predicate, platform, label set, membership checks, review process, and replay evaluation, although the resource bands remain low-confidence until the existing architecture is inspected."}},"blocking_evidence":["Whether an actual service specification makes a reusable total, exact, terminating Boolean guarantee over an open-ended model class.","Whether timeouts, unresolved runs, or numerical failures are currently reported or consumed as NO rather than as distinct inconclusive states.","Whether the model language and event predicate can be formalized without changing the scientifically intended question.","Whether exact-fragment membership and any proposed certificates can be enforced and independently checked.","Whether downstream systems preserve the tiered labels and avoid interpreting UNKNOWN as evidence of stability.","Whether the restricted exact or bounded modes cover enough scientifically relevant cases to justify implementation."],"next_evidence_step":"Obtain the current specification for one simulation model language and one event predicate, then have platform, formal-methods, and domain reviewers audit its quantifiers, representations, horizons, and timeout semantics. Replay a predeclared bounded sample of previously timed-out cases through both the existing interface and the proposed YES_WITH_CERTIFICATE, UNKNOWN_BOUND, OUT_OF_SCOPE, and NUMERICAL_FAILURE routing. Reject the motivating problem if the specification is enforceably finite and bounded, already distinguishes inconclusive outcomes from NO, and makes no universal exact claim; reject the intervention if it does not reduce unsupported Boolean outputs or fails blinded domain review of semantic fidelity and useful coverage.","research_questions":["Does the deployed interface actually claim exact termination for every admitted model and initial condition, or is the candidate responding to a guarantee that is not made?","Which inputs, horizons, numerical representations, evolution laws, and event predicates are enforceably bounded in the current service?","How frequently do current timeouts or unresolved executions become negative classifications in downstream records?","Can a useful exact fragment be defined with decidable membership, a proved terminating procedure, and independently checkable event certificates?","Do the proposed output categories preserve the scientific meaning of collision, escape, instability, or threshold events for the audited models?","Will downstream catalogs, reviewers, and users preserve UNKNOWN and failure labels rather than collapsing them into stability or NO?","What existing guarantee-labeling, reachability, certificate, and abstention practices overlap with the proposed integrated workflow?"],"recommendation":"VALIDATE_PROBLEM_FIRST","uncertainty_constraints":["No domain-specific undecidability result is supplied, so no impossibility conclusion is justified.","Problem prevalence, stakeholder demand, realized harm, and useful exact-fragment coverage are unmeasured.","Prior art and world-level distinctiveness are unmeasured because prior-art status is UNSEARCHED.","Chaos, continuum modeling, sensitivity, and long runtime cannot serve as evidence of undecidability.","A certificate about an encoded mathematical model cannot be treated as confirmation of physical reality.","Cost bands are resource-equivalent planning ranges based only on the stated scope, not inspected architecture or vendor pricing.","If all admitted instances are enforceably finite and bounded, the central issue may be complexity or numerical validity rather than computability."],"closed_book_prior_art_boundary":"No conclusion is made about novelty, prevalence, market size, or existing adoption. The sealed packet only supports assessing the internal coherence and testability of the proposed audit-and-routing pattern; external research would be required to compare it with simulation-service contracts, reachability methods, certificate systems, abstaining classifiers, and timeout-handling practices."}