{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"computability_boundary_mapping__human_computer_interaction:P2:v0","cell_id":"computability_boundary_mapping__human_computer_interaction","search_queries":["site:learn.microsoft.com Power Automate Copilot troubleshoot flow errors repair action user need","workflow reachability undecidable executable processes halting problem paper","automated planning incomplete search timeout unknown plan existence PDDL","human computer interaction uncertainty status assistant timeout impossible user trust study","research contextual help task completion plan current state cross application assistant workflow planning user goal","ACM intelligent user interface task assistant workflow plan explanation execution external human collaboration","Microsoft contextual help Copilot can I complete task current state workflow plan","standard user interface status uncertainty unknown error messages ISO 9241-110","Unified Planning PlanGenerationResultStatus UNSOLVABLE_INCOMPLETELY TIMEOUT documentation","planning API result status timeout unsolvable proven incomplete planner official documentation","PDDL planner status timeout unknown unsolvable proven plan validation standard","AI planning human action external actor workflow planning contingent plans current state assistant","Turing 1936 computable numbers Entscheidungsproblem original paper PDF Proceedings London Mathematical Society","site:academic.oup.com Turing 1936 computable numbers Entscheidungsproblem","site:royalsocietypublishing.org halting problem undecidable original Turing paper","site:plato.stanford.edu halting problem Turing computability","\"Scaling Context-Aware Task Assistants\" PDF","\"Scaling Context-Aware Task Assistants that Learn from Demonstration\" arxiv","site:arxiv.org context-aware task assistants demonstration mixed-initiative dialogue UIST 2025","site:dl.acm.org \"Scaling Context-Aware Task Assistants\""],"sources":[{"source_id":"S1","title":"Troubleshoot cloud flow errors","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/power-automate/troubleshoot-flow-errors","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-04-10","accessed_at":"2026-08-03","claims_supported":["Production workflows visibly encounter trigger failures, permission loss, moved resources, stale reads, action failures, wrong outputs, and timeouts.","Microsoft directs users to inspect run history and distinguish error causes rather than treating every failure as one state.","Power Automate is a concrete application environment and Microsoft/Power Platform product owners are identifiable potential authorizers."]},{"source_id":"S2","title":"FAQ for Copilot in cloud flows","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/power-automate/faqs-copilot","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-01-16","accessed_at":"2026-08-03","claims_supported":["Power Automate already embeds an assistant that answers questions about the current flow and helps build, edit, and run automations.","Microsoft documents limited connector coverage, incomplete error-fixing support, review and undo requirements, and administrator control.","The documented capability and limitations identify Microsoft Power Platform product owners and administrators as plausible adopters or authorizers, but do not establish willingness to adopt this broker."]},{"source_id":"S3","title":"GUIDE: A Benchmark for User Context Understanding and Assistance in GUI Workflow Videos","publisher":"Google Research","url":"https://research.google/pubs/guide-a-benchmark-for-user-context-understanding-and-assistance-in-gui-workflow-videos/","source_class":"PRIMARY_RESEARCH","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["Context-aware GUI help is an active research problem spanning complex software applications.","Across 120 novice demonstrations and ten applications, evaluated models struggled with behavior-state and help prediction.","Structured behavioral state and intent materially improved help prediction, supporting the value of explicit state and goal representations while showing that natural-language grounding remains difficult."]},{"source_id":"S4","title":"unified_planning.engines.results — Unified-Planning 1.3.0 documentation","publisher":"Unified Planning / AIPlan4EU maintainers","url":"https://unified-planning.readthedocs.io/en/latest/_modules/unified_planning/engines/results.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["Existing planning infrastructure already distinguishes solved, proven unsolvable, incompletely unsolved, timeout, memory exhaustion, internal error, unsupported problem, and intermediate results.","Existing validation infrastructure already exposes VALID, INVALID, and UNKNOWN and enforces consistency between positive statuses and attached plans.","This is the closest collision with the proposal's central status-honesty and evidence-carrying result pattern."]},{"source_id":"S5","title":"ISO 9241-110: Ergonomics of human-system interaction — Part 110: Interaction principles","publisher":"International Organization for Standardization","url":"https://www.iso.org/obp/ui?_escaped_fragment_=iso%3Astd%3Aiso%3A9241%3A-110%3Adis%3Aed-2%3Av1%3Aen","source_class":"STANDARD","publication_date":"2020","accessed_at":"2026-08-03","claims_supported":["Established interaction principles include suitability for users' tasks, self-descriptiveness, controllability, and use-error robustness.","The standard identifies misleading or insufficient information, unexpected responses, and inefficient error recovery as usability problems.","It defines interactive systems broadly enough to include software, services, people, online help, and human support, supporting explicit treatment of external actors."]},{"source_id":"S6","title":"Turing Machines","publisher":"Stanford Encyclopedia of Philosophy","url":"https://plato.stanford.edu/entries/turing-machine/index.html","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2025-05-21","accessed_at":"2026-08-03","claims_supported":["The halting problem has no total decider for arbitrary machines and inputs.","A valid undecidability transfer requires a reduction translating every source instance into a target instance so a target solver would solve the known-uncomputable source.","Computability in principle is distinct from practical computability, and external choices must be stated as part of the computation model."]},{"source_id":"S7","title":"SAP Speaks PDDL: Exploiting a Software-Engineering Model for Planning in Business Process Management","publisher":"arXiv / SAP Research","url":"https://arxiv.org/abs/1401.5858","source_class":"PRIMARY_RESEARCH","publication_date":"2014-01-23","accessed_at":"2026-08-03","claims_supported":["Planning over business-process states and actions is established prior art rather than a new application category.","SAP demonstrated compilation of an existing status-and-action model covering more than 400 business-object types into PDDL for process planning.","The paper identifies suitable model construction as potentially prohibitively complicated or costly, directly constraining scalability of the proposed broker."]},{"source_id":"S8","title":"Scaling Context-Aware Task Assistants that Learn from Demonstration and Adapt through Mixed-Initiative Dialogue","publisher":"Carnegie Mellon University Smart Sensing for Humans Lab / ACM UIST","url":"https://smashlab.io/publications/prism_agent/","source_class":"PRIMARY_RESEARCH","publication_date":"2025","accessed_at":"2026-08-03","claims_supported":["Context-aware procedural assistants using explicit step representations, demonstration, and dialogue are established adjacent research.","The PrISM system uses dialogue to refine uncertain state estimates and reduce inappropriate responses.","Task-assistant implementation and user evaluation are feasible, but imperfect sensing and unpredictable behavior remain material limitations."]}],"problem_evidence":{"support":"MODERATE","rationale":"Official Power Automate documentation visibly establishes stale state, permissions, connector, action, logic, and timeout failures in a relevant workflow environment, while GUI-assistance research shows substantial state- and intent-understanding errors. However, no source measures how often deployed help assistants specifically convert timeout or incomplete search into an 'impossible' verdict, so the proposal's exact failure pattern and prevalence remain unverified.","source_ids":["S1","S2","S3","S8"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Microsoft already operates contextual flow assistance, documents limitations and failure modes, and gives Power Platform administrators control, making its product owners and administrators credible adopters or authorizers. The documentation expresses a need to answer current-flow questions while avoiding faulty states, but there is no adoption commitment, budget, procurement signal, or explicit request for computability-boundary labels.","source_ids":["S1","S2"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Unified Planning result and validation statuses","similarity":"Already separates witnessed solutions, proven unsolvability, incomplete failure to find a plan, timeout, memory exhaustion, internal errors, unsupported inputs, and UNKNOWN validation; it also associates positive results with plan artifacts.","remaining_difference":"It is a planner-facing library, not a cross-application user-help broker with mechanically enforced finite-world promises, snapshot-staleness handling, external-actor capability contracts, and user-tested labels.","source_ids":["S4"]},{"name":"SAP Status and Action Management compiled to PDDL","similarity":"Uses enterprise business-object status variables and actions to generate processes toward desired state changes, closely matching state-relative workflow planning.","remaining_difference":"The reported work focuses on model reuse and process creation, not open executable workflows, semi-decision behavior, evidence-labelled help responses, or external-actor disclosure.","source_ids":["S7"]},{"name":"GUIDE contextual GUI assistance benchmark","similarity":"Represents behavior state and intent to decide when and how to assist users across complex GUI applications.","remaining_difference":"It benchmarks perception and help prediction from videos rather than formally deciding or semi-deciding task reachability and attaching proof-status semantics.","source_ids":["S3"]},{"name":"PrISM context-aware task assistant","similarity":"Tracks procedural state, answers task questions, and uses clarification dialogue to recover from uncertainty.","remaining_difference":"It supports learned procedural tracking rather than exact finite-state reachability, certified negative answers, or explicit computability-mode routing.","source_ids":["S8"]},{"name":"ISO 9241-110 interaction principles","similarity":"Already calls for task suitability, self-descriptiveness, controllability, and robust error recovery while treating people and support as parts of an interactive system.","remaining_difference":"It supplies general interaction requirements rather than the proposed computational status lattice, certificates, or routing algorithm.","source_ids":["S5"]}],"distinctive_claim_remaining":"For cross-application help, mechanically gate exact negative reachability claims on a checked finite, terminating, snapshot-stable, closed-world promise; otherwise expose replayable positive witnesses, incomplete search, staleness, tool failure, and named external-capability dependence as distinct user-facing results. The falsifiable increment over Unified Planning is whether the additional promise, staleness, and external-capability layer prevents materially more misleading help outcomes than a status-preserving planner wrapper alone.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Finite-state planning, explicit result enums, enterprise status/action models, and context-aware task assistants are already implemented, so an offline non-executing broker is technically credible. Major unresolved work is faithful goal formalization, enforceable completeness of cross-application action models, state freshness, permission and privacy handling, external-actor semantics, user comprehension of labels, and the modeling cost highlighted by SAP. The unrestricted-workflow reduction is theoretically plausible but has not been independently checked against a concrete workflow language.","source_ids":["S2","S3","S4","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Misleading feasibility guidance can waste user and support effort, and the relevant systems exhibit many distinct failure modes. Prevalence and realized harm are not measured.","source_ids":["S1","S3"]},"stakeholder_pull":{"score":3,"rationale":"A credible product family and administrative authority exist, and current-flow assistance is already an intended feature. No buyer, budget holder, or product owner has requested this intervention.","source_ids":["S1","S2"]},"incremental_advantage":{"score":2,"rationale":"The principal distinction between incomplete search, timeout, failure, and proven unsolvability is already implemented in Unified Planning. Added staleness and external-capability contracts may help, but their incremental effect is untested.","source_ids":["S4","S5"]},"distinctiveness_plausibility":{"score":2,"rationale":"The exact cross-application packaging is distinguishable, but its components substantially overlap planning-result APIs, enterprise process planning, interaction standards, and context-aware assistants.","source_ids":["S3","S4","S5","S7","S8"]},"technical_implementability":{"score":4,"rationale":"A frozen offline fixture broker can reuse established planning models, status types, traces, and finite-state algorithms. Live cross-application completeness is much harder and is not yet evidenced.","source_ids":["S4","S7","S8"]},"adoption_authority_feasibility":{"score":3,"rationale":"A non-executing pilot fits ordinary Power Platform product and help-operations authority, and administrators already control related assistant functionality. Data access and cross-application model ownership remain unresolved.","source_ids":["S1","S2"]},"evidence_readiness":{"score":3,"rationale":"The proposed fixture tests are bounded and technically executable, but no concrete workflow language, reduction proof, fixtures, archived cases, partner, or outcome baseline has been supplied.","source_ids":["S4","S6","S7"]},"safety_net_benefit":{"score":4,"rationale":"Preserving unknown, stale, unsupported, and failure states and withholding execution directly supports self-descriptiveness, controllability, and error robustness. Whether users interpret the labels correctly requires testing.","source_ids":["S4","S5"]},"scalability":{"score":2,"rationale":"Cross-application deployment requires reliable state, action, permission, and external-capability models. Prior enterprise planning work explicitly identifies modeling as potentially prohibitively complicated or costly.","source_ids":["S3","S7","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Four-to-eight-week offline study for one workflow family: formal model, approximately 60 fixed fixtures or archived de-identified cases, three comparator implementations, independent reduction/algorithm review, and audit report.","confidence":"LOW","assumptions":["0.5-1.5 full-time-equivalent engineering/research effort plus limited formal-methods review","Existing planning libraries and a frozen workflow model can be reused","No production connectors, live user actions, or new data collection"],"source_ids":["S4","S6","S7"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Shadow integration for one application family, including model adapter, mechanically checked promise gate, evidence trace store, status UI, access-control review, privacy review, accessibility review, and monitoring.","confidence":"LOW","assumptions":["Three-to-six multidisciplinary staff for six to nine months","Application owner supplies action semantics and test snapshots","Deployment remains advisory and non-executing"],"source_ids":["S1","S2","S4","S5","S7"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Production launch across several applications or workflow families with authenticated connectors, snapshot freshness, model-version governance, external-capability registry, reliability engineering, user research, support training, and incident response.","confidence":"LOW","assumptions":["Eight-to-fifteen multidisciplinary staff for roughly one year","Existing enterprise identity, telemetry, and planning infrastructure are reusable","Excludes licensing costs imposed by third-party applications and large-scale model creation"],"source_ids":["S1","S2","S3","S7","S8"]},"annual_recurring":{"band_2026_usd":"1M_TO_5M","scope":"Ongoing maintenance of application models and connectors, guarantee reclassification, security and privacy review, reliability operations, user-support escalation, audits, and regression fixtures for a multi-application service.","confidence":"LOW","assumptions":["Six-to-twelve continuing engineering, product, assurance, and support staff","Multiple applications change APIs, permissions, and workflow semantics each year","Compute cost is secondary to modeling and integration labor"],"source_ids":["S1","S2","S7"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official workflow documentation and primary GUI-assistance research establish heterogeneous failures, stale state, permissions, timeouts, and context-understanding errors in relevant environments, though the precise mislabeling prevalence is unknown.","source_ids":["S1","S3","S8"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Microsoft Power Platform product owners and administrators are identifiable because they already operate and control a current-flow assistant and associated failure-handling experience. Willingness to pilot is not established.","source_ids":["S1","S2"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The full broker can be compared against both a timeout-based planner and a Unified Planning-style status wrapper on promise enforcement, stale-state handling, external-dependency disclosure, and misleading-negative rates.","source_ids":["S4"]},"bounded_next_evidence_step":{"status":"YES","reason":"A fixed one-family, approximately 60-case, non-executing shadow evaluation with frozen models, predefined comparators, metrics, and falsifiers is bounded.","source_ids":["S4","S6","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first study can remain offline, frozen, de-identified, advisory, and non-executing; no collaborator contact, permission change, or support denial is required. Privacy review is still required for archived snapshots.","source_ids":["S2","S5"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"Scopes and labor-equivalent bands are explicit, but none of the eight sources supplies 2026 labor pricing, vendor quotes, integration estimates, or measured maintenance effort. Modeling is known to be costly, but the dollar ranges remain assumption-driven.","source_ids":["S7"]}},"next_evidence_step":"With a willing Power Platform or comparable workflow owner, pre-register a non-executing shadow study on one workflow family using 60 frozen cases: 20 finite reachable/unreachable cases, 15 open-search cases including seeded nontermination before a feasible path, 10 stale or permission-mismatch cases, 10 external-actor/service cases, and five malformed-goal or tool-failure cases. Compare (A) the current timeout-based help baseline, (B) a Unified Planning-style status-preserving wrapper, and (C) the full broker with mechanically checked finite-world promises, snapshot versions, evidence traces, and external-capability contracts. Primary measures are false exact-negative rate, witness replay rate, status-classification accuracy, omitted-dependency rate, model-authoring hours, and blinded user interpretation of each label. Falsify the intervention if any exact negative escapes the checked finite class, any accepted witness fails replay, fair search starves the seeded feasible path, any external dependency is omitted, or the status-only wrapper matches the full broker within a pre-registered non-inferiority margin while requiring materially less modeling effort. Also stop if archived-case review finds no timeout/incomplete-search collapse in the intended workflow family.","blocking_evidence":["No field evidence measures the prevalence of timeout or incomplete search being presented as task impossibility.","No product owner, workflow owner, help-operations manager, or budget holder has expressed willingness to pilot or fund the broker.","The unrestricted-workflow reduction has not been checked against a specified executable workflow language and semantics.","No real workflow corpus establishes that finite-world promises, goal predicates, and external dependencies can be authored and mechanically enforced at acceptable cost.","No comparative test shows that promise, staleness, and external-capability contracts outperform the already-established status distinctions in Unified Planning.","User comprehension, accessibility, privacy, and behavioral effects of the proposed labels require live or archived-case empirical study.","The four cost bands lack vendor quotes, wage benchmarks, or measured integration and maintenance effort."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This evaluation found substantial adjacent and colliding practice but did not perform an exhaustive literature, product, standards, patent, or code search. World novelty, patentability, freedom to operate, market size, realized impact, and commercial demand remain unmeasured. The only surviving testable distinction is the cross-application enforcement and user-facing effect of finite-world promises, staleness controls, evidence traces, and external-capability disclosure beyond existing planner result statuses.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure a named product/workflow owner and data-governance approval for one workflow family.","Measure the prevalence and consequences of collapsed timeout, stale-state, permission, dependency, and impossibility outcomes in archived or live help cases.","Independently verify the unrestricted reachability reduction and the finite-state algorithm against a concrete workflow language.","Run the pre-registered three-arm shadow comparison and test whether the full broker adds value over Unified Planning-style result statuses.","Measure goal/action-model authoring and maintenance effort, label comprehension, accessibility, privacy exposure, and support-workflow effects.","Replace assumption-driven cost bands with observed staff time, infrastructure consumption, and partner or vendor estimates."],"reason":"Bounded web research establishes a relevant problem setting, credible authorizers, technical feasibility, and substantial prior-art collision. The remaining questions—actual failure prevalence, stakeholder pull, modelability, user comprehension, comparative advantage, and operating cost—require proprietary cases, partner participation, or live/shadow testing and cannot be resolved by further bounded web search."},"proposal_index":2}