{"schema_version":1,"assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","source_experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"computability_boundary_mapping__performing_arts_theatre","archetype_slug":"computability_boundary_mapping","domain_slug":"performing_arts_theatre","title":"Scoped verification and labeled fallbacks for interactive stage scores","opportunity_summary":"The proposal would replace unsupported Boolean conclusions about arbitrary programmable, sensor-responsive stage scores with enforceable finite-fragment checking and explicit UNKNOWN, TIMEOUT, and OUT_OF_SCOPE results. Its value depends on confirming that the theatre company actually makes or operationally relies on universal checker claims; the packet provides no prevalence, pilot-result, or prior-art evidence.","adopter_authorizer":"Production safety lead and stage-management authority for operational reliance; artists retain authority over altering, simplifying, or withdrawing scores.","scores":{"meaningful_impact":{"score":4,"rationale":"A false SAFE result could expose performers, crew, audiences, machinery, and venue operations to harm, while false rejection could suppress viable artistic work. Impact is conditional because the packet supplies no evidence about how often unsupported Boolean conclusions occur."},"stakeholder_pull":{"score":2,"rationale":"The proposal addresses recognizable safety, auditability, and artistic-exclusion concerns, but the sealed candidate contains no requests, commitments, observed incidents, budget ownership, or evidence that theatre stakeholders want formal scope routing."},"incremental_advantage":{"score":4,"rationale":"Compared with more rehearsal and simulation coverage, enforceable fragment membership and preserved UNKNOWN, TIMEOUT, and OUT_OF_SCOPE states directly target unsupported Boolean conclusions that additional testing alone cannot justify. Advantage disappears if existing staff already use equivalent distinctions or if labels do not affect decisions."},"distinctiveness_plausibility":{"score":2,"rationale":"The theatre-specific combination may be useful, but prior-art status is explicitly UNSEARCHED and the packet does not establish separation from existing show-control verification, automation assurance, or scoped formal-analysis practice."},"technical_implementability":{"score":3,"rationale":"A ten-score, non-public pilot with seeded unsafe variants is bounded and technically conceivable, but explicit semantics, enforceable fragment membership, faithful environmental abstractions, and comparison against physical traces require substantial multidisciplinary work."},"adoption_authority_feasibility":{"score":4,"rationale":"Operational and artistic authorities are explicitly identified, live machinery autonomy is excluded, and rollback returns control to existing interlocks and manual stage management. Adoption may still fail if the restricted fragment excludes routine artistic practices or UNKNOWN results are informally overridden."},"evidence_readiness":{"score":3,"rationale":"The packet supplies a specific rehearsal setting, ten representative scores, seeded variants, comparison modes, expert review, falsifiers, and halt conditions. It lacks baseline observations, results, and a prespecified acceptance rule for classification errors, scope mismatches, overrides, and decision changes."},"safety_net_benefit":{"score":5,"rationale":"The intervention explicitly preserves unresolved states, prohibits SAFE labels for inconclusive cases, excludes autonomous live-machinery control, and mandates immediate reversion to established rehearsal interlocks after false SAFE or model mismatch."},"scalability":{"score":2,"rationale":"The routing pattern could recur across productions, but each score language, sensor environment, machinery model, artistic practice, and guarantee may require bespoke semantics, abstraction validation, and authority review. The packet gives no evidence of reuse across venues or systems."}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Audit one planned production's current claims and decision practice, then—only if the problem survives—specify and evaluate ten representative scores plus seeded unsafe variants on a non-public rehearsal system against witnessed traces and expert review.","confidence":"LOW","assumptions":["Requires coordinated time from stage management, safety staff, artists, automation engineers, and a formal-methods specialist.","Existing rehearsal hardware, interlocks, logs, and score artifacts are accessible without major procurement.","The comparison includes baseline review, exact fragment checking, finite abstraction, bounded search, and prespecified error and mismatch measures.","No live machinery is autonomously controlled by pilot outputs."]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Production-grade semantics, fragment-membership enforcement, checker and fallback integration, evidence versioning, operator interfaces, validation, training, and safety review for one venue or theatre company.","confidence":"LOW","assumptions":["The existing score language and cue infrastructure can expose stable machine-readable interfaces.","Safety and stage-management authorities require documented validation and workflow integration.","The implementation covers one principal technical stack rather than many incompatible venue systems.","Physical interlocks and manual authority remain separate and funded."]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"Launch on an initial production cohort, including score onboarding, model review, rehearsal comparison, staff training, incident procedures, and monitored decision use.","confidence":"LOW","assumptions":["Core tooling already exists from startup work.","Launch remains advisory and does not authorize autonomous machinery control.","Several productions require individual score and environment modeling.","Artists must consent to any score alteration rather than having changes imposed by the system."]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Ongoing specialist review, model and semantics maintenance, per-production onboarding, software support, evidence retention, training, and periodic safety reassessment for one theatre company.","confidence":"LOW","assumptions":["A limited number of productions use the system each year.","Changes to sensors, machinery, score syntax, or guarantees trigger renewed review.","Recurring costs exclude major replacement of theatre machinery and expansion to unrelated venue stacks.","Manual rehearsal, expert judgment, and established interlocks continue alongside the verification process."]}},"research_burden":"HIGH","earliest_credible_horizon":"3_TO_12_MONTHS","pipeline_gates":{"recognizable_externally_supportable_problem":{"status":"UNCERTAIN","reason":"The failure mode is concrete and falsifiable, but the proposal depends on an actual universal or total checker claim and operational reliance. The packet provides no observation showing that current staff coerce timeouts, unknowns, or out-of-scope inputs into Boolean conclusions."},"identifiable_adopter_or_authorizer":{"status":"YES","reason":"The production safety lead and stage-management authority can approve operational reliance, while artists control score alteration or withdrawal."},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal claims that enforceable routing and explicit unresolved labels will reduce false Boolean conclusions or reveal scope and assumption mismatches relative to baseline rehearsal review; the stated intervention falsifier directly tests this."},"bounded_next_evidence_step":{"status":"YES","reason":"A one-production audit followed conditionally by a non-public ten-score comparison is bounded, safe, and decision-relevant, and can be falsified if existing practice already distinguishes scoped results or the new routing produces no improvement."},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"Live machinery autonomy is excluded, inconclusive outputs cannot receive SAFE labels, artistic changes require artist control, and false SAFE, unclassified input, abstraction omission, or semantic mismatch triggers rollback to existing safeguards."},"implementation_cost_scope_and_range":{"status":"YES","reason":"The candidate specifies a bounded pilot and enough components to state broad resource-equivalent ranges, although exact costs remain uncertain because score-language complexity, integration conditions, and modeling effort are unknown."}},"blocking_evidence":["Evidence that the company actually asserts or operationally relies on a universal, terminating safe/unsafe checker rather than already preserving bounded and unresolved outcomes.","A specification of the accepted score language, execution semantics, environmental inputs, quantifiers, and safety or termination guarantee, including whether the system is already finite and covered by a proven total checker.","Prespecified pilot acceptance criteria covering false Boolean classifications, missed concrete behaviors, scope mismatches, overrides, and changes to production decisions.","Pilot evidence comparing the proposed routing modes with baseline expert review and witnessed traces, including seeded unsafe cases.","Evidence that the enforceable fragment includes enough routine artistic practice to be useful and that membership cannot be bypassed or silently rewritten.","External prior-art evidence distinguishing this proposal from theatre automation, show-control verification, and safety-assurance practice."],"next_evidence_step":"On one planned production, conduct a non-public document-and-staff audit comparing current checker outputs, documentation, and readiness decisions across completed cases, timeouts, abstraction failures, changed scores, and out-of-scope inputs. Falsify the problem if current practice already preserves bounded results, UNKNOWN, TIMEOUT, OUT_OF_SCOPE, and manual authority with scope-linked evidence; otherwise use the audit to prespecify error measures and an acceptance rule for the ten-score rehearsal comparison.","research_questions":["Does the company make or operationally rely on a universal total-checker claim, and are inconclusive cases currently reported as safe, unsafe, or impossible?","Is the accepted score language and modeled environment finite and fully specified, or does it include unrestricted computation, unbounded inputs, dynamic code, or semantic changes?","What exact property is checked: termination, reachability of forbidden encoded states, real-world safety, or a combination, and what quantifiers apply?","Can fragment membership be mechanically enforced before results are used, including after score, sensor, machinery, or environment changes?","How often do baseline review and the proposed modes disagree, and which disagreements are supported by witnessed traces or seeded cases?","What false-SAFE, false-UNSAFE, unresolved-rate, scope-mismatch, override, and decision-change thresholds would justify continuation?","Does the finite abstraction preserve relevant performer and machinery behavior, and how are missing behaviors detected?","Which routine artistic practices fall outside the fragment, and do artists regard the resulting restrictions or UNKNOWN rates as acceptable?","What existing theatre automation, show-control verification, or safety-assurance approaches provide equivalent scope contracts and labeled unresolved states?","What staffing, integration, validation, and recurring per-production modeling effort is required? "],"recommendation":"VALIDATE_PROBLEM_FIRST","uncertainty_constraints":["Closed-book assessment: no claims are made about prior art, prevalence, market size, realized impact, or exact cost.","The candidate is conditional on an expressive programmable score system and a genuine universal terminating verification claim.","Formal conclusions apply only to the encoded score-and-environment model, not to theatrical meaning, improvisation, or complete real-world safety.","Unrestricted computability remains unresolved because the packet supplies no checked domain-faithful proof or reduction.","Cost bands are resource-equivalent planning ranges based only on the stated scope and may shift materially with language complexity, system access, compliance demands, and venue heterogeneity."],"closed_book_prior_art_boundary":"Prior-art status is explicitly UNSEARCHED. This assessment cannot determine novelty, prevalence, or whether theatre automation, show-control verification, formal methods, and safety-assurance systems already implement equivalent scoped fragments, unresolved-state labels, or authority controls."}