{"schema_version":1,"assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","source_experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"computability_boundary_mapping__physics","archetype_slug":"computability_boundary_mapping","domain_slug":"physics","title":"Model-Relative Reachability Contracts for Computational Physics","opportunity_summary":"Replace unsupported unrestricted reachability answers with a frozen, model-relative contract that distinguishes bounded or sound verdicts from timeout, out-of-scope, and UNKNOWN. The immediate opportunity is an offline comparison on one model family and 30 archived queries, not a universal undecidability claim or live experiment gate.","adopter_authorizer":"A physics program lead can authorize the shadow pilot; the model owner and an independent methods reviewer must authorize changes to published scientific claims.","scores":{"meaningful_impact":{"score":4,"rationale":"Avoiding unjustified negative predictions could protect experiment selection and scientific claim integrity. The consequence is substantial when the diagnosed behavior occurs, but its prevalence and realized impact are unsupported."},"stakeholder_pull":{"score":2,"rationale":"The candidate names model authors, experimental teams, software users, and program leadership, but provides no evidence that they currently experience the problem, demand a remedy, or would accept additional UNKNOWN outputs."},"incremental_advantage":{"score":4,"rationale":"Compared with budget-limited simulation or simply adding compute, explicit boundedness and separate UNKNOWN routing directly prevent timeout from being represented as NO. Whether this improves practical calibration enough to justify added inconclusive results remains untested."},"distinctiveness_plausibility":{"score":3,"rationale":"The combination of a frozen computation contract, admitted fragment, evidence-relative labels, and reclassification triggers is internally coherent, but prior art is explicitly unsearched and external distinctiveness cannot be established closed-book."},"technical_implementability":{"score":3,"rationale":"A 30-query offline shadow comparison is plausibly implementable, but the model family, encoding, target semantics, precision rules, admission grammar, and sound checking procedure have not been instantiated. A universal impossibility result may be substantially harder or invalid."},"adoption_authority_feasibility":{"score":4,"rationale":"The candidate identifies authority for a shadow pilot and separately identifies the model owner and independent reviewer for claim changes. Multi-party approval and potential resistance to UNKNOWN labels remain adoption frictions."},"evidence_readiness":{"score":3,"rationale":"The candidate supplies archived-query scope, a baseline comparison, falsifiers, stop conditions, and rollback rules. Readiness is limited by the missing operational specification and absence of evidence that the diagnosed production behavior exists."},"safety_net_benefit":{"score":5,"rationale":"The intervention explicitly retains UNKNOWN, prohibits timeout-as-NO and unvalidated experiment gating, limits generalization to checked bounds, and defines halt and rollback conditions. These protections remain useful even if unrestricted undecidability is not proved."},"scalability":{"score":2,"rationale":"Each additional model family may require bespoke semantics, encoding validation, fragment design, proof or abstraction work, and reviewer agreement. The packet provides no basis that these tasks can be standardized across computational physics."}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Specify one pilot model family and evaluate 30 archived reachability queries using exhaustive bounded checks, shadow routing, baseline comparison, and independent methods review.","confidence":"LOW","assumptions":["Archived models, queries, baseline outputs, and relevant software are accessible.","The pilot requires computational-physics, formal-methods, and evaluation labor.","No new experiments or specialized physical equipment are required.","Encoding and target semantics can be resolved without creating a new modeling platform."]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Convert a successful pilot into a supported workflow for the initial model family, including admission checks, label presentation, provenance, versioned guarantees, tests, documentation, and governance integration.","confidence":"LOW","assumptions":["Deployment remains limited to one organization and a small number of workflows.","Existing simulation infrastructure can be integrated rather than replaced.","Independent review and user training are required.","The router does not control live experiments during initial validation."]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Launch a production-quality service across multiple model families or programs with validated encodings, soundness testing, software integration, monitoring, reviewer processes, and change control.","confidence":"LOW","assumptions":["Multiple families require separate admission and semantic-validation work.","Production reliability, auditability, access controls, and support are required.","No universal proof eliminates the need for family-specific engineering.","The scope excludes major new experimental facilities or hardware programs."]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain guarantee versions, revalidate changed model assumptions, operate compute and software, investigate misencodings or counterexamples, support users, and conduct recurring independent review.","confidence":"LOW","assumptions":["The deployed portfolio remains modest rather than field-wide.","Model grammars and simulation stacks change regularly enough to require rechecking.","Specialist methods and domain staff remain involved.","Compute demand is material but does not require exceptional dedicated infrastructure."]}},"research_burden":"HIGH","earliest_credible_horizon":"3_TO_12_MONTHS","pipeline_gates":{"recognizable_externally_supportable_problem":{"status":"YES","reason":"The packet gives an observable problem definition: unrestricted 'all models' and 'ever' requirements combined with finite simulation evidence and collapsed NO, timeout, unknown, and out-of-scope outputs. Actual occurrence and prevalence still require an audit."},"identifiable_adopter_or_authorizer":{"status":"YES","reason":"The physics program lead is identified as shadow-pilot authorizer, while the model owner and independent methods reviewer are identified for changes to published claims."},"distinct_testable_incremental_claim":{"status":"YES","reason":"For a frozen pilot class, the proposal can be tested on whether explicit routing yields fewer false negatives or misleading labels than the budget-expiry baseline without producing predominantly unhelpful UNKNOWN results."},"bounded_next_evidence_step":{"status":"YES","reason":"The candidate bounds the first step to one model family, 30 archived queries, a frozen encoding and horizon, exhaustive bounded checks, and a shadow comparison that cannot alter live experiments."},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The shadow-only scope, excluded actions, named authorities, halt criteria, and rollback plan address the immediate pilot risks. The unresolved undecidability hypothesis does not prevent testing the narrower labeling intervention."},"implementation_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The 30-query pilot supplies a bounded evidence scope, but the model family, encoding difficulty, software architecture, data condition, staffing, and number of eventual production families are unspecified, so implementation resource ranges remain assumption-sensitive."}},"blocking_evidence":["An audit must establish that the target workflow actually makes an unrestricted guarantee, collapses timeout or UNKNOWN into NO, or presents finite simulation as universal evidence.","The pilot model family needs a frozen encoding, target-region semantics, precision rules, horizon, fragment-membership test, and experimentally meaningful label definitions.","The archived-query comparison must show better claim calibration or fewer misleading negatives than the baseline without mostly substituting unusable UNKNOWN results.","Any sound-verdict mechanism must survive bounded counterexample checks and encoding review before its labels influence scientific claims.","A bounded prior-art and practice review is required before asserting distinctiveness or novelty.","A checked reduction and answer-preservation argument are required before making any undecidability claim about an unrestricted physical-model class."],"next_evidence_step":"For one selected model family, first audit the 30 archived queries for the stated overclaim and timeout-as-NO behavior; then freeze encoding, observable semantics, precision, and horizon, and run an offline shadow comparison of the baseline labels against exhaustive bounded checks plus REACHED, NOT-REACHED-WITHIN-BOUND, UNKNOWN, and OUT-OF-SCOPE routing. Falsify progression if the diagnosed behavior is absent, a supposedly sound verdict misses a bounded counterexample, label comprehension fails, or routing does not reduce misleading labels and mainly adds unusable UNKNOWN results.","research_questions":["Does the actual production requirement claim unrestricted 'all models' and 'ever' coverage, and are timeout or unresolved cases represented as negative answers?","Can one experimentally relevant model family be encoded with semantics and precision rules that preserve the reachability question users actually care about?","What enforceable fragment admits a total or sound procedure, and can membership in that fragment be checked reliably?","Relative to the baseline, how often does shadow routing correct misleading negatives, and how often does it merely replace outputs with UNKNOWN?","Can downstream users reliably distinguish claims about an encoded model and finite bound from claims about nature or unlimited time?","Does the production class already have finite enforceable bounds and a total procedure, making this ordinary complexity or implementation work?","What relevant reachability, hybrid-system verification, computable-analysis, abstraction, and scientific-model-validation practices already exist?","What changes in model assumptions, grammar, software, or observables must trigger guarantee reclassification?"] ,"recommendation":"VALIDATE_PROBLEM_FIRST","uncertainty_constraints":["Closed-book assessment cannot establish problem prevalence, stakeholder demand, realized impact, market size, or prior-art distinctiveness.","Unrestricted undecidability is only a hypothesis; no valid physical-model encoding, reduction, or answer-preservation proof is supplied.","Continuous quantities, measurement noise, approximation, and laboratory interaction may break the proposed computational correspondence.","The useful decidable fragment and the proportion of real queries it would admit are unknown.","Cost bands are resource-equivalent planning bands based on assumed staffing and scope, not observed prices or point estimates.","The candidate supplies no empirical evidence that users will understand or accept the proposed label lattice."],"closed_book_prior_art_boundary":"Prior art is explicitly UNSEARCHED. This assessment makes no claim that the reachability contract, admitted-fragment approach, routing scheme, or broader computability-boundary transfer is novel, uncommon, or superior to existing methods."}