{"schema_version":1,"assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","source_experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"computability_boundary_mapping__human_computer_interaction","archetype_slug":"computability_boundary_mapping","domain_slug":"human_computer_interaction","title":"Proof-scoped preflight validation with explicit unknown routing","opportunity_summary":"Evaluate whether automation-authoring interfaces misrepresent inconclusive analysis as universal binary verdicts, and whether enforceable scope restrictions, guarantee-labelled routing, and distinct UNKNOWN, TIMEOUT, and OUT-OF-SCOPE states improve user understanding and safer deployment choices. The candidate is structurally testable, but the problem's prevalence, stakeholder demand, distinctiveness, and practical fragment coverage remain unestablished.","adopter_authorizer":"The product owner can authorize a read-only shadow prototype; security, accessibility, and formal-methods reviewers must approve production guarantees or enforcement changes.","scores":{"meaningful_impact":{"score":4,"rationale":"If the specified binary-verdict failure exists, false reassurance can contribute to unsafe workflow deployment and timeout-derived rejection can waste author effort; however, prevalence and realized impact are unsupported."},"stakeholder_pull":{"score":2,"rationale":"The candidate identifies affected authors, operators, reviewers, and maintainers, but provides no evidence that any adopter currently experiences, prioritizes, or seeks a remedy for this hypothesized problem."},"incremental_advantage":{"score":4,"rationale":"Compared with mapping failures to pass/fail or merely adding heuristic confidence, enforceable scope classification and explicit abstention states directly target misleading definiteness and provide a clear safety-oriented increment."},"distinctiveness_plausibility":{"score":3,"rationale":"The composition of formal scope, enforceable routing, guarantee-labelled outcomes, and user-comprehension testing is coherent, but prior art is explicitly unsearched and no distinctiveness claim can be established closed-book."},"technical_implementability":{"score":3,"rationale":"A read-only prototype and explicit UI states appear feasible, while sound fragment definitions, enforceable membership, correct routing, and maintenance across arbitrary code, services, and plug-ins may be difficult."},"adoption_authority_feasibility":{"score":4,"rationale":"The candidate names a product owner for shadow testing and identifies required security, accessibility, and formal-methods reviewers, with production authority explicitly bounded."},"evidence_readiness":{"score":4,"rationale":"A safe 40-workflow comparison, problem and intervention falsifiers, excluded actions, and halt criteria are specified; readiness is reduced by missing decision thresholds and an independent classification-ground-truth procedure."},"safety_net_benefit":{"score":4,"rationale":"UNKNOWN, TIMEOUT, and OUT-OF-SCOPE routing could prevent inconclusive computation from being presented as assurance, and the proposal includes rollback triggers; flawed formalization or automation bias could still create new false confidence."},"scalability":{"score":3,"rationale":"The interface pattern could transfer across workflow systems, but each language, integration, plug-in change, and computation model may require renewed boundary proofs, routing validation, and evidence updates."}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Audit the product claim and accepted language, construct or consent 40 representative workflows, independently classify scope, build a read-only interface comparison, conduct the user study, and complete formal, security, and accessibility review of the protocol.","confidence":"LOW","assumptions":["An existing product or realistic authoring prototype is available for inspection.","No external actions are executed and no production integration is required.","The study requires specialized formal-methods and HCI labor plus participant coordination.","Ground-truth adjudication can be completed without developing a production-grade verifier."]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Engineer enforceable fragment membership, routing logic, guarantee-labelled UI states, logging, versioned evidence records, recheck triggers, and production-quality tests for one bounded automation product.","confidence":"LOW","assumptions":["The initial scope is one product and a limited set of workflow modes.","Existing authoring, analyzer, and policy infrastructure can be extended.","The exact fragment is technically definable and covers enough workflows to justify implementation.","This excludes a broad rewrite of the underlying workflow language."]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Complete production assurance, independent formal review, security and accessibility validation, staged rollout, monitoring, documentation, operator training, and rollback preparation.","confidence":"LOW","assumptions":["Startup engineering produces a sound and enforceable implementation.","Launch is staged within one organization rather than across a large multi-product portfolio.","No new regulated certification regime is triggered.","User testing shows that labels improve discrimination without materially increasing unsafe choices."]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Maintain classifications and evidence records, review language and plug-in changes, rerun comprehension and accessibility checks, investigate routing errors, and update guarantees.","confidence":"LOW","assumptions":["The product changes at a moderate cadence.","Most rechecks reuse established proofs and test infrastructure.","External-service integrations do not require continuous full re-verification.","One product team retains formal-methods review access."]}},"research_burden":"HIGH","earliest_credible_horizon":"3_TO_12_MONTHS","pipeline_gates":{"recognizable_externally_supportable_problem":{"status":"UNCERTAIN","reason":"The candidate specifies an observable and consequential interface failure, but whether any target product actually makes a universal exact binary claim or collapses inconclusive states requires external audit."},"identifiable_adopter_or_authorizer":{"status":"YES","reason":"A product owner is identified as the shadow-prototype authorizer, with security, accessibility, and formal-methods reviewers identified for later production decisions."},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal claims that enforceable proof-scoped routing with explicit inconclusive states will improve discrimination of FALSE versus UNKNOWN and reduce unsafe choices relative to binary failure mapping."},"bounded_next_evidence_step":{"status":"YES","reason":"The sealed candidate specifies a read-only, sandboxed comparison on 40 consented or synthetic workflows without executing external actions, together with an intervention falsifier."},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step is non-deploying, authority is bounded, prohibited interpretations are listed, and hazardous SAFE labels, unsafe UNKNOWN interpretation, unenforceable membership, or failed accessibility review trigger halt and rollback."},"implementation_cost_scope_and_range":{"status":"YES","reason":"A bounded one-product implementation scope can be described and assigned broad resource bands, although estimates remain low-confidence because the language, integrations, and proof obligations are unknown."}},"blocking_evidence":["Whether a real target interface promises or communicates a universal exact binary verdict for unrestricted workflows.","Whether the accepted workflow language is actually unrestricted or is already enforceably bounded and decidable for the represented property.","Whether authors currently confuse timeout or analyzer failure with definitive safe, unsafe, valid, or invalid outcomes.","An independently reviewed ground truth for fragment membership and routed outcomes in the comparison set.","Whether the exact fragment covers enough representative workflows and UNKNOWN rates remain usable.","Whether guarantee-labelled routing improves FALSE-versus-UNKNOWN discrimination and reduces unsafe choices relative to the baseline at comparable task time.","Whether prospective adopters and required reviewers regard the problem as sufficiently important to fund implementation.","The relationship of the proposed composition to existing formal-verification interfaces, abstaining systems, and end-user programming research."],"next_evidence_step":"First audit one candidate product's accepted syntax, computation model, user-visible claims, and timeout handling; proceed only if the diagnosed universal-binary condition is present. Then run the authorized read-only 40-workflow comparison against the current binary baseline using independently reviewed classifications and predeclared comprehension, unsafe-choice, task-time, and UNKNOWN-rate thresholds. Falsify the intervention if labelled routing fails to improve FALSE-versus-UNKNOWN discrimination or reduce unsafe choices despite correct routing and comparable task time.","research_questions":["Does the target interface actually make or imply a total exact binary guarantee across all accepted workflows?","Are workflow inputs mechanically restricted enough that the relevant property is already decidable?","How often do users interpret timeout, analyzer failure, confidence, or UNKNOWN as a safety assurance?","Which formal properties are being judged, under which encoding, quantifiers, and computation model?","Can fragment membership be mechanically enforced without silent syntax narrowing or unsafe escape hatches?","What proportion of representative workflows falls into exact, one-sided, heuristic, timeout, unknown, and out-of-scope modes?","Do users perceive and correctly act on guarantee labels across relevant accessibility conditions?","What thresholds for comprehension, unsafe choices, task time, and UNKNOWN frequency would justify advancement?","How sensitive are classifications and proofs to plug-in, language, and external-service changes?","What existing work most closely matches the proposed mechanism composition and evaluation design?","Will product, security, accessibility, and formal-methods stakeholders authorize and resource a staged implementation if the test succeeds?"],"recommendation":"VALIDATE_PROBLEM_FIRST","uncertainty_constraints":["Problem prevalence is a hypothesis and cannot be inferred from the candidate's detailed specification.","No stakeholder-demand, adoption, market-size, or realized-impact evidence is supplied.","Prior art is unsearched, so novelty and relative distinctiveness are unmeasured.","The actual workflow language, property, encoding, quantifiers, and computation model are unspecified.","Cost bands are resource-equivalent planning ranges, not estimates based on an inspected product architecture.","Technical feasibility depends on sound formalization, enforceable fragment membership, and representative workflow coverage.","The proposed study lacks sealed decision thresholds and a complete independent-adjudication procedure.","A computability analogy is insufficient without a valid reviewed reduction that preserves the deployed HCI task."],"closed_book_prior_art_boundary":"No claim is made about novelty, prevalence, market size, existing implementations, or realized effectiveness. The assessment is limited to the sealed candidate's internal structure, stated baseline and rival, falsifiers, authority boundaries, and proposed evidence step; external prior-art research is required before asserting distinctiveness."}