{"schema_version":1,"assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","source_experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"computability_boundary_mapping__human_computer_interaction","archetype_slug":"computability_boundary_mapping","domain_slug":"human_computer_interaction","title":"Proof-scoped preflight validation with explicit unknown routing","opportunity_summary":"Evaluate whether an automation-authoring interface that promises universal binary preflight verdicts for unrestricted workflows misrepresents inconclusive analysis, and whether enforceable scope restrictions, guarantee-labelled routing, and distinct UNKNOWN, TIMEOUT, and OUT-OF-SCOPE states improve user understanding and deployment choices. The existence and prevalence of the diagnosed interface condition remain hypotheses.","adopter_authorizer":"The product owner can authorize a shadow prototype; security, accessibility, and formal-methods reviewers must approve any production guarantee or enforcement change.","scores":{"meaningful_impact":{"score":4,"rationale":"If the hypothesized interface condition exists, false reassurance or timeout-derived rejection could materially affect deployment safety, valid workflow retention, user understanding, and development effort. The sealed candidate does not establish prevalence or realized harm."},"stakeholder_pull":{"score":2,"rationale":"The proposal identifies affected authors, operators, reviewers, maintainers, and people acted upon, but supplies no evidence that any adopter currently experiences the problem, requests the intervention, or prioritizes it."},"incremental_advantage":{"score":4,"rationale":"Relative to mapping analyzer failure to pass/fail or using confidence-scored heuristics without formal scope, enforceable fragment membership and explicit inconclusive states directly address the stated causal mechanism. Whether they improve decisions is still an empirical hypothesis."},"distinctiveness_plausibility":{"score":2,"rationale":"The composition is coherent, but prior art is explicitly unsearched and relationships to end-user programming, abstaining classifiers, explainable verification, and formal-methods interface research are unknown."},"technical_implementability":{"score":3,"rationale":"A read-only labelled prototype and workflow classification exercise appear bounded, but sound formalization, independently checked guarantees, enforceable fragment membership, routing accuracy, and versioned rechecks may be difficult for workflows containing arbitrary code and external interaction."},"adoption_authority_feasibility":{"score":4,"rationale":"The candidate identifies a product owner for shadow-prototype approval and names security, accessibility, and formal-methods reviewers for production changes. Multi-reviewer coordination remains necessary."},"evidence_readiness":{"score":4,"rationale":"The candidate specifies a safe 40-workflow comparison, baseline, observable comprehension target, intervention falsifier, and halt conditions. It lacks predeclared thresholds and a complete independent classification-validation procedure."},"safety_net_benefit":{"score":4,"rationale":"Distinct UNKNOWN, TIMEOUT, and OUT-OF-SCOPE states, non-assertive fallback messaging, hazardous-label halt criteria, and recheck triggers could reduce false certainty even when exact verification is unavailable. Guarantee labels could themselves induce automation bias."},"scalability":{"score":3,"rationale":"Routing and versioned evidence could extend across workflows, but scalability depends on how many real workflows fit enforceable exact fragments and on maintaining classifications as syntax, plug-ins, services, and interaction models change."}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Design and run the read-only sandboxed comparison on 40 consented or synthetic workflows, including prototype UI, workflow classification review, participant sessions, accessibility checks, and analysis.","confidence":"LOW","assumptions":["No external actions are executed.","Existing prototyping and sandbox infrastructure is available.","The study uses a modest participant sample and synthetic or already-consented workflows.","Formal-methods review is limited to study classifications rather than a production proof system."]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Formalize supported workflow classes and guarantees, build enforceable membership checks and routing, integrate user-visible states, and complete security, accessibility, formal-methods, and product validation before limited production use.","confidence":"LOW","assumptions":["The current authoring system can expose syntax and analyzer boundaries for enforcement.","Only selected exact fragments receive strong guarantees.","No complete redesign of the automation language or execution platform is required.","Production approval requires several specialist reviews."]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Launch the scoped validation modes, migrate preflight messaging, train operators and support staff, instrument user outcomes, validate rollback, and conduct a controlled evaluation.","confidence":"LOW","assumptions":["Launch is limited to one product or workflow platform.","Existing observability, experimentation, and release systems can be reused.","Rollout is staged and does not require executing previously prohibited external actions during evaluation.","Major customer-specific integrations are excluded."]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain proofs, abstractions, routing rules, accessibility behavior, boundary records, regression tests, reviewer processes, and reclassification triggers as the workflow language and plug-ins evolve.","confidence":"LOW","assumptions":["Language and integration changes occur regularly.","Specialist engineering and periodic formal-methods, security, and accessibility review remain necessary.","The system serves one principal platform rather than many independently evolving products.","No exact staffing levels or vendor prices are assumed."]}},"research_burden":"HIGH","earliest_credible_horizon":"3_TO_12_MONTHS","pipeline_gates":{"recognizable_externally_supportable_problem":{"status":"UNCERTAIN","reason":"The failure mode is recognizable and observable, but the sealed candidate labels the target problem and prevalence as hypotheses and supplies no external evidence that a relevant interface actually makes the universal binary claim."},"identifiable_adopter_or_authorizer":{"status":"YES","reason":"The product owner is identified as the shadow-prototype authorizer, with security, accessibility, and formal-methods reviewers identified for production authority."},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal claims that enforceable proof-scoped routing with explicit inconclusive states will improve discrimination of FALSE versus UNKNOWN and reduce unsafe deployment choices relative to binary timeout-to-pass/fail behavior."},"bounded_next_evidence_step":{"status":"YES","reason":"A read-only, sandboxed comparison on 40 consented or synthetic workflows is specified, with no external actions and an explicit intervention falsifier."},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step excludes production decisions and external actions, prohibits unsafe label interpretations, names required authority, and defines halt and rollback conditions for hazardous SAFE labels, misunderstood UNKNOWN states, unenforceable membership, or accessibility failure."},"implementation_cost_scope_and_range":{"status":"YES","reason":"The candidate defines separable evidence, production integration, launch, and maintenance work; broad resource-equivalent ranges can be bounded despite low confidence about platform complexity."}},"blocking_evidence":["Whether any target product actually accepts unrestricted workflows while promising universal exact binary preflight verdicts.","Whether accepted inputs, analyzer scope, timeout behavior, and user-visible guarantees are already enforceably bounded and accurately understood.","Whether workflow classifications and exact-fragment membership can be independently validated and mechanically enforced.","Whether guarantee-labelled routing improves FALSE-versus-UNKNOWN comprehension and reduces unsafe choices relative to the stated baseline without materially increasing task time.","Whether UNKNOWN frequency leaves the intervention useful across representative workflows.","Whether status distinctions pass accessibility review and avoid transferring automation bias to guarantee labels.","Whether close prior work already contains the proposed composition, interface contract, or evaluation design."],"next_evidence_step":"Run the authorized read-only, sandboxed study on 40 consented or synthetic workflows. Independently establish classification ground truth, compare the existing binary timeout-to-pass/fail presentation with enforceable guarantee-labelled TRUE/FALSE/UNKNOWN/TIMEOUT/OUT-OF-SCOPE routing, and measure comprehension, unsafe deployment choices, and task time without executing external actions. Falsify the intervention if correctly classified routing does not improve FALSE-versus-UNKNOWN discrimination or reduce unsafe choices at comparable task time; halt on any hazardous exact SAFE label, systematic interpretation of UNKNOWN as safe, unenforceable membership, or accessibility failure.","research_questions":["Does an auditable target interface actually make a universal exact binary claim for workflows containing loops, arbitrary code, external services, or unbounded interaction?","What formal class, property, quantifiers, computation model, and user-visible guarantee accurately describe the deployed HCI task?","Which useful workflow fragments admit independently checked exact verification, and can membership be mechanically enforced?","What proportion of representative workflows receives UNKNOWN, TIMEOUT, or OUT-OF-SCOPE outcomes?","Do guarantee-labelled outcomes improve users' discrimination of FALSE from UNKNOWN and reduce unsafe deployment choices relative to the baseline?","Do users transfer automation bias from binary verdicts to official-looking guarantee labels?","Can the distinctions be perceived and understood across relevant accessibility needs?","What decision thresholds, exclusion rules, reviewer-disagreement process, and ground-truth procedure should govern the study?","How often must boundary records and classifications be rechecked after language, plug-in, or service changes?","What existing research or products are the closest precedents for the mechanism composition and evaluation design?"],"recommendation":"VALIDATE_PROBLEM_FIRST","uncertainty_constraints":["Problem existence and prevalence are unsupported hypotheses.","Realized impact, stakeholder demand, market size, and adoption willingness are unmeasured.","Prior art and distinctiveness are unverified.","The deployed workflow language, formal property, encoding, and computation model are unspecified.","The proportion of real workflows inside useful exact fragments is unknown.","User comprehension gains, unsafe-choice reduction, task-time effects, and automation-bias risks are untested.","Cost bands are resource-equivalent planning ranges, not quotations or point estimates."],"closed_book_prior_art_boundary":"Prior art is unsearched and unverified. This closed-book assessment makes no claim about novelty, prevalence, market position, or absence from existing end-user programming, verification, abstention, or HCI research."}