{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"computability_boundary_mapping__human_computer_interaction:P5:v0","cell_id":"computability_boundary_mapping__human_computer_interaction","search_queries":["site:learn.microsoft.com Power Automate limits run duration retry policy timeout official","site:docs.temporal.io durable execution retries idempotency workflow official","site:docs.camunda.io job retries incidents user task workflow state official","workflow termination analysis sound decidable fragment research workflow nets termination","site:docs.temporal.io workflow durable execution event history retries activities idempotency","site:docs.aws.amazon.com step functions execution status redrive retry idempotent official","site:docs.camunda.io docs process instance state incidents retries wait states official","site:help.zapier.com loops limits timeout workflow run statuses official","CHI end user automation failures trigger action programming errors study IFTTT","end user programming automation debugging workflow failures study no code HCI","Microsoft Power Automate user research run failure status confusion study","site:dl.acm.org automation authoring debugging end users trigger action programming","Turing 1936 computable numbers decision problem PDF original paper proceedings London Mathematical Society","Rice theorem semantic properties programs original paper 1953 PDF","ISO 9241-210 status feedback user control error official standard human system interaction","Nielsen visibility system status user control error prevention official heuristics","site:bls.gov/ooh software developers quality assurance analysts testers median pay 2025","site:bls.gov/oes software developers annual mean wage May 2025","site:aws.amazon.com/step-functions/pricing pricing state transitions official","site:learn.microsoft.com power automate pricing official 2026"],"sources":[{"source_id":"S1","title":"On Computable Numbers, with an Application to the Entscheidungsproblem","publisher":"Proceedings of the London Mathematical Society; authorized web copy hosted by Abelard","url":"https://www.abelard.org/turpap2/turpap2.htm","source_class":"PRIMARY_RESEARCH","publication_date":"1937","accessed_at":"2026-08-03","claims_supported":["A uniform mechanical procedure cannot decide every well-formed question in an unrestricted computational model.","The proposal's impossibility premise is credible only after the workflow language is shown capable of the relevant simulation; the source does not itself establish that product-specific reduction."]},{"source_id":"S2","title":"Soundness of workflow nets: classification, decidability, and analysis","publisher":"Springer Nature, Formal Aspects of Computing","url":"https://link.springer.com/article/10.1007/s00165-010-0161-4","source_class":"PRIMARY_RESEARCH","publication_date":"2010-08-03","accessed_at":"2026-08-03","claims_supported":["Workflow nets are an established formal method for analyzing workflow soundness, including deadlocks and livelocks.","Eight soundness notions are decidable for the studied workflow-net class, while most examined expressive extensions make them undecidable.","Restricting expressiveness to a formally analyzable workflow class is established prior art rather than a novel mechanism."]},{"source_id":"S3","title":"Supporting mental model accuracy in trigger-action programming","publisher":"University of Washington Human-Centered Robotics Lab; published in ACM UbiComp","url":"https://hcrlab.cs.washington.edu/publications/huang2015ubicomp/","source_class":"PRIMARY_RESEARCH","publication_date":"2015","accessed_at":"2026-08-03","claims_supported":["Trigger-action automation was already used at substantial scale.","Two user studies found inconsistent interpretations of automation behavior and errors when users created programs for desired behavior.","Interface-level representations of automation semantics can materially affect end-user understanding."]},{"source_id":"S4","title":"Limits of automated, scheduled, and instant flows","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/he-il/power-automate/limits-and-config","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-03-24","accessed_at":"2026-08-03","claims_supported":["A major no-code workflow platform enforces finite limits on actions, nesting, loop iterations, runtime, retries, concurrency, and external calls.","Power Automate permits runs containing pending approvals for up to 30 days and distinguishes flow suspension reasons in its management surfaces.","A test view may display a timeout after ten minutes while the flow continues in the background, directly demonstrating that a UI timeout need not mean execution termination.","Retries, failed actions, connector limits, throttling, pending work, cancellation, and suspension are operationally distinct conditions."]},{"source_id":"S5","title":"Employ robust error handling","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/power-automate/guidance/coding-guidelines/error-handling","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-07-11","accessed_at":"2026-08-03","claims_supported":["Microsoft expressly recommends separate handling for failed, timed-out, skipped, and successful actions.","Microsoft recommends run metadata, logging, notifications, bounded retry policies, and explicit termination status and messages.","This establishes an identifiable platform authorizer and expressed operational need for more informative workflow outcomes, though not demand for the proposal's exact completion-contract design."]},{"source_id":"S6","title":"Restarting state machine executions with redrive in Step Functions","publisher":"Amazon Web Services","url":"https://docs.aws.amazon.com/step-functions/latest/dg/redrive-executions.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["AWS already supports resuming unsuccessful, aborted, or timed-out workflows from an unsuccessful step while preserving successful-step results and history.","Redrive is permission-controlled, version-linked, status-aware, and constrained by retention and event-history limits.","Checkpoint-like continuation, preservation of completed work, explicit resume authority, and differentiated unsuccessful statuses substantially overlap the proposal."]},{"source_id":"S7","title":"Choosing workflow type in Step Functions","publisher":"Amazon Web Services","url":"https://docs.aws.amazon.com/step-functions/latest/dg/choosing-workflow-type.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["AWS exposes different maximum durations, persistence behavior, execution histories, and exactly-once, at-least-once, and at-most-once guarantees for different workflow types.","AWS explicitly connects non-idempotent effects to exactly-once execution and idempotent effects to at-least-once execution.","Durable execution, explicit guarantee labels, bounded workflow modes, persisted state, audit history, and idempotency-aware routing are established commercial practices."]},{"source_id":"S8","title":"Software Developers, Quality Assurance Analysts, and Testers","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/ooh/computer-and-information-technology/software-developers.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2025-08-28","accessed_at":"2026-08-03","claims_supported":["The May 2024 median annual wage was $133,080 for software developers and $102,610 for software QA analysts and testers.","Software development and QA normally involve requirements, security, maintenance, test plans, risk assessment, usability, and stakeholder feedback, supporting a multidisciplinary labor-cost model.","The cost bands are resource-equivalent planning estimates, not vendor quotations, and require loaded compensation and 2026 escalation assumptions."]}],"problem_evidence":{"support":"STRONG","rationale":"The problem is visible in both research and deployed systems. End-user automation studies found interpretation and authoring errors, while current Power Automate documentation explicitly says a test UI can time out while execution continues and separately documents pending approvals, retries, throttling, suspension, connector failures, and cancellation. These sources support the need to avoid treating elapsed time as a semantic nontermination verdict. They do not measure how often duplicate effects or misleading completion labels occur across platforms.","source_ids":["S3","S4","S5"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Microsoft Power Automate and AWS Step Functions are identifiable platform owners with authority over workflow status, retry, suspension, execution guarantees, and resume permissions. Their official documentation expresses needs for differentiated outcomes, resilient error handling, histories, and controlled redrive. No source shows either organization requesting, funding, or agreeing to pilot the proposal's exact proven-fragment-plus-completion-contract package.","source_ids":["S4","S5","S6","S7"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Formal workflow-net soundness analysis","similarity":"Uses an enforceable formal workflow class to obtain decidable guarantees and identifies expressive extensions that cross into undecidability.","remaining_difference":"The research concerns formal soundness classes, not the proposed end-user status alphabet, resource suspension, checkpoints, effect ledger, or resume/cancel interface.","source_ids":["S2"]},{"name":"Microsoft Power Automate limits and error-handling model","similarity":"Already bounds loops and runtime, distinguishes several action outcomes and suspension reasons, supports pending approvals, retries, logging, notifications, and explicit termination status.","remaining_difference":"The reviewed documentation does not advertise a mechanically proved guaranteed-completion fragment or distinguish SUSPENDED_BUDGET from a semantic nontermination finding under the proposed contract.","source_ids":["S4","S5"]},{"name":"AWS Step Functions Standard workflows and redrive","similarity":"Provides durable state, execution history, differentiated failed/aborted/timed-out outcomes, permissioned redrive, preservation of successful work, version-linked continuation, and explicit execution/idempotency guarantees.","remaining_difference":"AWS documents operational execution guarantees rather than a proof-carrying end-user language fragment or the proposal's exact completion labels and author-facing boundary record.","source_ids":["S6","S7"]}],"distinctive_claim_remaining":"For an end-user automation builder that currently exposes unrestricted executable behavior, the combined contract—mechanically admitting only proved-terminating workflows to GUARANTEED_COMPLETE while routing all others to checkpointed budget suspension with non-Boolean status labels and effect-ledger-backed authorized resumption—will reduce false completion judgments and duplicate mock effects relative to both a timeout/failure baseline and a durable-retry system without the proof/status contract. This is contrastive and falsifiable, but its integrated advantage is untested; the constituent mechanisms are substantially established.","confidence":"HIGH"},"implementation_evidence":{"support":"STRONG","rationale":"Formal decidable workflow subclasses are demonstrated in research, and major commercial orchestration systems already implement bounded modes, durable histories, explicit execution semantics, redrive permissions, retry policies, and idempotency-aware guarantees. A side-effect-free prototype is therefore technically credible. Remaining uncertainties are product-specific: whether the actual grammar can be safely restricted, whether fragment checks match runtime semantics, how external calls are modeled, checkpoint confidentiality and compatibility, and whether every effect can satisfy an idempotency or compensation contract.","source_ids":["S1","S2","S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Truthful status and controlled resumption could prevent unsafe retries, duplicated effects, and wasted effort on impossible universal prediction. The realized frequency and severity of these failures remain unmeasured.","source_ids":["S3","S4","S5","S6","S7"]},"stakeholder_pull":{"score":3,"rationale":"Major platform owners visibly invest in limits, differentiated error handling, histories, retries, and redrive, but no external source expresses demand for the exact integrated intervention.","source_ids":["S4","S5","S6","S7"]},"incremental_advantage":{"score":3,"rationale":"The integration could improve semantic honesty over timeout/failure interfaces and add a provable guarantee absent from ordinary durable execution. Its advantage over combining existing formal verification and orchestration practices has not been measured.","source_ids":["S2","S4","S6","S7"]},"distinctiveness_plausibility":{"score":2,"rationale":"Decidable workflow fragments, explicit runtime limits, differentiated statuses, durable histories, permissioned resumption, and idempotency-aware guarantees are all established. Distinctiveness is confined to the exact contract and UX integration.","source_ids":["S2","S4","S5","S6","S7"]},"technical_implementability":{"score":4,"rationale":"Each major technical building block has a credible analogue. Correct product-specific formalization, proof/runtime correspondence, external-call modeling, secure checkpoints, and effect-ledger coverage remain material engineering obligations.","source_ids":["S2","S4","S6","S7"]},"adoption_authority_feasibility":{"score":4,"rationale":"A workflow-platform release owner can control grammar admission, statuses, budgets, retry permissions, and production-effect access. External-service owners must separately authorize called effects; an offline pilot avoids that dependency.","source_ids":["S4","S5","S6","S7"]},"evidence_readiness":{"score":3,"rationale":"A frozen fixture suite and three-arm prototype comparison are bounded and measurable, but proof review, product-specific runtime instrumentation, and user comprehension testing cannot be completed by web research.","source_ids":["S2","S3","S4","S6"]},"safety_net_benefit":{"score":5,"rationale":"Explicit suspension, permissioned resumption, preserved execution history, and idempotency-aware effect handling directly reduce the risk that uncertainty is converted into a false failure or destructive duplicate retry.","source_ids":["S4","S5","S6","S7"]},"scalability":{"score":4,"rationale":"Static admission checks, status routing, histories, and redrive already operate in large commercial workflow systems. Proof-rule maintenance, checkpoint storage, history limits, and complex external effects may constrain scale.","source_ids":["S2","S4","S6","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Freeze one workflow grammar and mocked primitive set; independently review one unrestricted-model reduction and the terminating-fragment argument; build a side-effect-free three-arm runner; execute the fixture matrix; and run a small moderated comprehension study.","confidence":"MODERATE","assumptions":["Roughly 0.8-1.5 FTE-years spread across a formal-methods reviewer, workflow engineer, QA engineer, and HCI researcher.","Loaded 2026 labor is estimated above BLS cash wages to cover benefits, overhead, contracting premiums, and wage escalation.","No production connectors, migrations, regulated data, or external effects are included."],"source_ids":["S8"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Implement grammar admission, proof-rule versioning, checkpoint format, status APIs and UI, mock effect ledger, observability, security review, and limited sandbox integrations for one platform.","confidence":"MODERATE","assumptions":["Approximately 3-6 multidisciplinary FTEs for six to nine months.","Reuses an existing workflow runtime and identity system.","Excludes broad connector remediation and production migration."],"source_ids":["S4","S5","S6","S7","S8"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Production hardening, connector classification, secure checkpoint storage, authorization, recovery tooling, effect-ledger integration, proof/runtime conformance testing, documentation, migration controls, and staged release.","confidence":"LOW","assumptions":["Approximately 8-15 multidisciplinary FTEs for nine to eighteen months plus infrastructure, security, legal, support, and contingency.","Existing workflow execution infrastructure is retained.","Cost rises materially if many existing connectors lack usable idempotency, compensation, or bounded-wait contracts."],"source_ids":["S4","S5","S6","S7","S8"]},"annual_recurring":{"band_2026_usd":"1M_TO_5M","scope":"Operate checkpoint and ledger services; monitor status integrity and duplicate effects; maintain proof rules and connector contracts; conduct security reviews, incident response, support, and reclassification after language changes.","confidence":"LOW","assumptions":["Approximately 5-12 continuing engineering, QA, formal-review, security, and support FTEs.","Infrastructure cost depends on run volume, history retention, checkpoint size, and suspended-run duration, none of which is externally established for the candidate.","No market-size or revenue assumptions are included."],"source_ids":["S4","S6","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Research documents end-user automation interpretation and authoring errors, and current product documentation demonstrates that UI timeout, pending work, suspension, retries, failures, and cancellation are distinct.","source_ids":["S3","S4","S5"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Microsoft Power Automate and AWS Step Functions product organizations are identifiable authorities already governing workflow limits, statuses, execution guarantees, histories, retry, and redrive. Exact pilot interest is not established.","source_ids":["S4","S5","S6","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The integrated contract can be compared against a timeout/failure baseline and a durable-retry comparator on status accuracy, unauthorized retries, duplicate mock effects, task completion, and user interpretation.","source_ids":["S2","S4","S6","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A frozen grammar, finite synthetic fixture suite, independently reviewed proofs, three side-effect-free prototype conditions, and prespecified pass/fail metrics bound the next step.","source_ids":["S2","S3","S4","S6","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"For the next step only, production effects and live external services can be prohibited, capabilities mocked, datasets synthetic, and all resumption participant-authorized. Production adoption still requires checkpoint-security and connector-authority review.","source_ids":["S5","S6","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four bands state staffing, duration, reuse, integration, and infrastructure assumptions and are anchored to government wage data. Production and recurring bands remain low-confidence because connector inventory and run volume are unknown.","source_ids":["S8","S4","S6","S7"]}},"next_evidence_step":"Conduct a preregistered, side-effect-free evaluation on a frozen grammar and synthetic workflow suite. First, have an independent reviewer check the unrestricted-model reduction and the guaranteed-fragment evaluator's termination argument. Then compare three prototypes: A, syntactic linting plus wall-clock timeout mapped to FAILED; B, bounded durable execution with checkpoints and retry but ordinary success/failure/timed-out labels; and C, the full completion contract with enforced guaranteed fragment, COMPLETED_OBSERVED, SUSPENDED_BUDGET, WAITING_EXTERNAL, USER_ACTION_REQUIRED, CANCELLED, TOOL_FAILURE, permissioned resume, and a mock-effect ledger. Use fixtures covering acyclic flows, maximum finite loops, valid and invalid decreasing measures, unchecked code, finite slow runs, infinite loops, external waits, human pauses, cancellation, crashes, stale checkpoints, and suspension after a committed mock effect. With 24-30 representative automation authors, measure semantic-label accuracy, correct resume/cancel choice, duplicate mock effects, unauthorized retries, task completion, time, and confidence calibration. Falsify the intervention if any admitted guaranteed fixture violates its proof, any rejected fixture obtains the guarantee, any budget exhaustion is labeled nonterminating or failed, any committed mock effect duplicates on authorized resume, condition C does not improve semantic-label accuracy over both comparators, or its added complexity materially increases unsafe actions.","blocking_evidence":["No measured prevalence of false timeout/nontermination interpretations, duplicate effects, or unsafe retries in a defined target platform.","No adopter commitment, product-owner interview, funding signal, or authorization for a platform-specific pilot.","No product-specific proof that the accepted workflow language can simulate the required unrestricted computation under its actual semantics.","No independently checked termination proof or runtime-conformance evidence for a concrete guaranteed fragment.","No comparative user evidence that the expanded status alphabet improves understanding without increasing confusion or abandonment.","No security and privacy assessment for checkpointed workflow state, credentials, histories, or effect-ledger data.","No connector inventory showing which external effects are idempotent, compensable, bounded, or unsuitable for resumption.","World novelty, patentability, freedom to operate, market size, and realized impact are unmeasured."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This assessment measures neither world novelty nor patentability. It found substantial collisions for formal decidable workflow subclasses, bounded execution, differentiated outcomes, durable histories, permissioned redrive, and idempotency-aware execution. The only surviving distinctiveness candidate is the exact integration and user-facing completion contract, whose novelty, freedom to operate, market size, and realized impact remain unmeasured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain a named workflow-platform product owner or funder willing to authorize the frozen, side-effect-free pilot.","Produce an independently checked product-specific reduction and guaranteed-fragment termination/conformance artifact.","Complete the three-arm fixture and user study with prespecified comparators, safety metrics, and falsifiers.","Quantify target-platform prevalence of ambiguous timeouts, unsafe retries, and duplicate effects from proprietary run telemetry or a representative field sample.","Complete checkpoint-security, retention, authorization, and connector idempotency/compensation assessments before any production-effect trial."],"reason":"Bounded web research verified the problem, credible authorizers, technical feasibility, and substantial prior-art collision, but it cannot establish the proposal's remaining incremental claim. That claim requires a live prototype comparison, independent proof review, user testing, and preferably proprietary operational telemetry. Under the controller rule, the need for empirical testing requires STOP_EMPIRICAL_RESEARCH_NEEDED, and every STOP is non-repairable in this evaluation cycle."},"proposal_index":5}