{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"catalytic_pathway_enablement__human_computer_interaction:P4:v0","cell_id":"catalytic_pathway_enablement__human_computer_interaction","search_queries":["site:pair.withgoogle.com guidebook Wizard of Oz prototyping AI","site:microsoft.com HAX toolkit Wizard of Oz prototyping human AI","Wizard of Oz prototyping intelligent agent CHI 1993 ACM","Wizard of Oz studies why and how 1993 Dahlback Jonsson Ahrenberg","site:dl.acm.org Wizard of Oz prototyping platform reusable operator console HCI","Wizard of Oz prototyping platform operator interface behavior script logging HCI study","site:protopie.io official conditional interactions variables formulas prototyping user testing","site:figma.com prototyping conditional logic variables official","site:hhs.gov OHRP informed consent deception research debriefing official","site:acm.org code ethics human subjects research privacy informed consent","site:bls.gov OEWS web and digital interface designers May 2025 wage","site:bls.gov occupational outlook handbook web developers digital designers pay 2025","\"Mapping the Wizards' Path\" PDF","\"Mapping the Wizards’ Path\" arxiv","\"Wizard of Oz studies — why and how\" PDF"],"sources":[{"source_id":"S1","title":"HAX Playbook","publisher":"Microsoft","url":"https://www.microsoft.com/en-us/haxtoolkit/playbook/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Teams developing user-facing AI have difficulty anticipating interaction failures before a system is built.","Microsoft explicitly offers inexpensive simulation and early user testing before fully functional implementation.","Product and design teams are an identifiable adopter class expressing demand for early behavioral prototyping."]},{"source_id":"S2","title":"Mapping the Wizards' Path: A Systematic Review of Wizard-of-Oz in HCI","publisher":"Association for Computing Machinery","url":"https://doi.org/10.1145/3772318.3791174","source_class":"PRIMARY_RESEARCH","publication_date":"2026-04-13","accessed_at":"2026-08-03","claims_supported":["Wizard-of-Oz is a longstanding core HCI prototyping technique spanning numerous application domains.","A review of 194 SIGCHI papers found recurring concerns including operator variability and incomplete reporting of behavior, rules, and logging.","The literature already includes specialized tools and automation intended to reduce wizard cognitive load and inconsistency.","The review recommends explicit methodological justification, standardized protocols, operator decision rules, training, and consistency checks."]},{"source_id":"S3","title":"Wizard of Oz Studies—Why and How","publisher":"Elsevier, author-hosted copy at Linköping University","url":"https://www.ida.liu.se/~arnjo82/papers/kbs.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"1993-12","accessed_at":"2026-08-03","claims_supported":["Wizard-of-Oz studies obtain empirical interaction evidence before the target system exists.","The reported ARNE environment already combined participant and wizard interfaces, canned responses, templates, automation, timestamped logs, and application-specific customization.","Timing, consistency, scenario design, wizard guidance, and post-study debriefing materially affect validity.","Customization could require 20–40 pilot studies, and human operators had difficulty maintaining consistent answers."]},{"source_id":"S4","title":"WebWOZ: A Wizard of Oz Prototyping Framework","publisher":"Association for Computing Machinery / Trinity College Dublin","url":"https://www.scss.tcd.ie/Gavin.Doherty/papers/EICS-WebWOZ.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2010-06-19","accessed_at":"2026-08-03","claims_supported":["One-off participant and wizard interfaces are a documented recurring setup burden.","WebWOZ proposed a reusable framework integrating human operation and automated language-technology components.","Wizard interfaces support response consistency, dialogue state, history, options, and the operator's cognitively demanding task.","Early-stage implementation and tuning can be time- and cost-intensive, motivating reusable mixed-fidelity prototyping."]},{"source_id":"S5","title":"A Web-Based Wizard-of-Oz Platform for Collaborative and Reproducible Human-Robot Interaction Research","publisher":"Bucknell University","url":"https://digitalcommons.bucknell.edu/honors_theses/747/","source_class":"PRIMARY_RESEARCH","publication_date":"2026-05-06","accessed_at":"2026-08-03","claims_supported":["HRIStudio is an open-source platform with a visual experiment designer, guided wizard interface, hierarchical specification, event-driven execution, timestamped logging, deviation tracking, plugins, and role-based access control.","A six-session pilot reported directionally better fidelity, execution reliability, and usability than a robot-programming comparator, while being too small for inferential claims.","This implementation closely overlaps the proposed shell, behavior-contract, operator-console, automation, logging, and access-control elements."]},{"source_id":"S6","title":"Use Expressions in Prototypes","publisher":"Figma","url":"https://help.figma.com/hc/en-us/articles/15253194385943-Use-expressions-in-prototypes","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["A widely used commercial prototyping product already supports realistic dynamic behavior using variables, expressions, conditional transitions, and viewer access.","Existing general-purpose tools can implement fixed branching without a hidden operator, creating a credible low-cost comparator.","The documentation does not establish governed human operation, consent, debriefing, operator workload control, or cross-session reset."]},{"source_id":"S7","title":"Informed Consent FAQs","publisher":"U.S. Department of Health and Human Services, Office for Human Research Protections","url":"https://www.hhs.gov/ohrp/regulations-and-policy/guidance/faq/informed-consent/index.html","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["For covered human-subjects research, legally effective prospective consent is generally required unless an IRB approves a permitted waiver or alteration.","Consent is an ongoing process requiring disclosure, understanding, voluntariness, confidentiality information, and withdrawal rights.","An IRB may authorize altered disclosure only when regulatory criteria are met; local applicability and authorization cannot be inferred from a product team's approval.","This supports independent ethics-review authority and creates a deployment dependency rather than a technical impossibility."]},{"source_id":"S8","title":"National Employment and Wage Data by Occupation, May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.t01.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-05-15","accessed_at":"2026-08-03","claims_supported":["May 2025 mean wages were $148,100 for software developers, $117,490 for web and digital interface designers, and $98,770 for web developers.","These wage anchors support order-of-magnitude labor-equivalent cost estimates, but not vendor quotes or fully loaded organizational costs."]}],"problem_evidence":{"support":"STRONG","rationale":"Independent research and first-party guidance repeatedly document the underlying problem: teams need to study adaptive behavior before full implementation, while static prototypes omit it and one-off wizard interfaces are costly, cognitively demanding, and inconsistent. The evidence establishes the class-level problem, not its prevalence or cost inside any particular prospective adopter.","source_ids":["S1","S2","S3","S4"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Microsoft explicitly addresses product teams building user-facing AI and offers early behavioral simulation because those teams face the stated need. Bucknell's HRIStudio demonstrates an academic research adopter building closely related infrastructure, while OHRP identifies the IRB as an authorizer for covered research. No named organization has expressed intent, budget, or demand for this particular cross-team foundry.","source_ids":["S1","S5","S7"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"HRIStudio","similarity":"Very high technical overlap: reusable web platform, hierarchical behavior specification, guided wizard execution, event-driven behavior, logging, deviation tracking, plugins, role-based access, and an empirical comparator.","remaining_difference":"HRIStudio is scoped to human-robot interaction and its reported pilot does not establish the proposal's broader research-operations service, consent/debrief templates, operator-welfare limits, privacy reset, turnover accounting, or independent shutdown governance.","source_ids":["S5"]},{"name":"WebWOZ","similarity":"High overlap with a reusable framework combining participant interface, wizard interface, automated components, dialogue state, consistent responses, and repeated early-stage experiments.","remaining_difference":"It is language-technology-focused and does not supply the proposal's full institutional admission, evidence-release, privacy, welfare, regeneration, and claim-boundary regime.","source_ids":["S4"]},{"name":"ARNE Wizard-of-Oz simulation environment","similarity":"Early established implementation already used separate operator controls, response templates, automation, timing control, timestamped logs, repeated scenarios, and debriefing.","remaining_difference":"It required application-specific customization and lacked the proposed reusable service governance, explicit behavioral contracts, capacity metrics, privacy reset, and matched turnover assay.","source_ids":["S3"]},{"name":"Microsoft HAX Playbook","similarity":"First-party toolkit for cheaply simulating human-AI behavior and failure scenarios before building a functional system.","remaining_difference":"It is principally planning and methodological guidance, not an operated shared facility with controlled sessions, reusable specialist capacity, fidelity auditing, reset, and evidence release.","source_ids":["S1"]},{"name":"Figma dynamic prototypes","similarity":"Commercially established reusable shells can express variables, conditional transitions, and dynamic behavior without custom backend implementation.","remaining_difference":"Figma does not itself provide governed hidden operation, operator consistency controls, consent/debriefing, study authorization, fidelity-deviation records, or cross-session regeneration.","source_ids":["S6"]}],"distinctive_claim_remaining":"For low-risk behavioral hypotheses that exceed ordinary conditional prototypes, a shared service combining a versioned shell, constrained human assistance, standardized contracts, independent study authorization, deviation-linked evidence release, workload limits, and verified reset will reduce per-hypothesis preparation effort by at least 30% across three successive variants versus thin coded and controlled ad-hoc Wizard-of-Oz comparators, while remaining non-inferior on preregistered participant-visible fidelity and producing no consent, privacy, or prohibited-claim breach.","confidence":"HIGH"},"implementation_evidence":{"support":"STRONG","rationale":"The technical core is demonstrably implementable: historical and current systems provide reusable participant shells, guided operator consoles, scripted or event-driven behavior, automation, logs, deviation tracking, plugins, and role-based access. Conditional commercial prototypes cover simpler cases. The workflow is feasible if eligibility, rehearsal, fallbacks, evidence labeling, reset, and operator limits are enforced. Legal authority is contextual: covered studies require appropriate institutional/IRB determination and consent or an approved alteration. Safety is manageable for the proposed fictional, non-consequential probe, but no source verifies its privacy reset, operator-welfare regime, general-HCI portability, or claimed labor advantage in live operation.","source_ids":["S2","S3","S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"The foundry could prevent premature backend investment and improve experiential evidence, but the prevalence and current cost of repeated setup have not been quantified and ordinary prototyping tools cover some cases.","source_ids":["S1","S4","S6"]},"stakeholder_pull":{"score":3,"rationale":"Product teams and HCI/HRI researchers visibly use early simulation methods, but no adopter has committed to this governed service or stated a budget.","source_ids":["S1","S5"]},"incremental_advantage":{"score":2,"rationale":"Most technical elements already appear in ARNE, WebWOZ, HRIStudio, and commercial dynamic prototyping. The remaining advantage is an untested governance-and-operations bundle.","source_ids":["S3","S4","S5","S6"]},"distinctiveness_plausibility":{"score":2,"rationale":"A broader governed service with welfare, privacy reset, evidence-release constraints, and turnover measurement is contrastive, but it may be an institutional packaging of established Wizard-of-Oz methodology rather than a distinct intervention.","source_ids":["S2","S5","S7"]},"technical_implementability":{"score":5,"rationale":"Multiple implementations demonstrate the shell, operator console, automation, logging, behavior specification, and access-control components.","source_ids":["S3","S4","S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"UX research and research-operations leaders can sponsor a low-risk service, while an IRB or equivalent authority can approve covered studies. Applicability, data policy, and any altered disclosure remain institution-specific.","source_ids":["S1","S7"]},"evidence_readiness":{"score":4,"rationale":"The proposal defines a small, reversible test with concrete comparators, measurements, and stop rules. Live-session evidence is still absent.","source_ids":["S2","S5","S7"]},"safety_net_benefit":{"score":4,"rationale":"Explicit fallbacks, withdrawal, debriefing, deviation disclosure, reset, and reversion to an ordinary prototype pathway provide meaningful containment if actually enforced.","source_ids":["S3","S7"]},"scalability":{"score":3,"rationale":"Reusable platforms and templates can scale across sessions, but wizard cognitive load, scenario customization, analysis backlog, and authorization create real throughput limits.","source_ids":["S2","S3","S4","S5"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Prepare three contracts and comparators; complete ethics/privacy review, rehearsals, recruitment, 12 sessions, fidelity coding, residue checks, operator-workload review, and a short analysis.","confidence":"MODERATE","assumptions":["Approximately 150–350 staff hours across UX research, prototype design/development, operator training, privacy/ethics review, and analysis.","Uses an existing prototyping platform and internal research infrastructure rather than building a new production service.","Participant incentives, software subscriptions, and overhead remain modest.","BLS wages are converted to loaded resource-equivalent rates; this is not a procurement quote."],"source_ids":["S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Build or adapt a reusable shell and operator console, contract templates, deterministic transitions, secure study-scoped logging, reset tooling, deviation reports, access controls, and initial operator training.","confidence":"MODERATE","assumptions":["Roughly 0.5–1.5 software-developer years plus part-time design, research-operations, privacy, security, and training effort.","Reuses established web, identity, storage, and prototyping infrastructure.","Excludes production-system integration and high-consequence use cases."],"source_ids":["S3","S4","S5","S6","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Harden the facility for several internal teams with documented eligibility, role-based access, audit and incident procedures, operator coverage, scenario library, service metrics, privacy testing, and independent ethics governance.","confidence":"LOW","assumptions":["Requires approximately 2–5 loaded staff-years across engineering, UX research operations, operator coverage, security/privacy, and governance.","Includes organizational integration and initial service capacity but no production actions or regulated high-risk studies.","Actual cost depends heavily on security controls, study volume, and whether existing research operations can absorb authorization duties."],"source_ids":["S5","S7","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Operate a small multi-team foundry: intake, rehearsals, sessions, operator rotation and training, module maintenance, privacy and fidelity audits, incident handling, platform upkeep, and evidence support.","confidence":"LOW","assumptions":["Approximately 2–4 loaded full-time-equivalent roles plus infrastructure, participant, and periodic review costs.","Throughput is capped by operator and downstream analysis capacity rather than software concurrency alone.","Does not include product engineering that follows successful research."],"source_ids":["S2","S3","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent research and first-party guidance directly document costly one-off interfaces, inadequate low-fidelity substitutes, early-testing need, wizard cognitive load, and response inconsistency.","source_ids":["S1","S2","S3","S4"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Product teams building user-facing AI and HCI/HRI researchers are externally visible users of this method; Bucknell implemented a close platform, and an IRB or equivalent institutional process is a credible authorizer for covered studies.","source_ids":["S1","S5","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The residual governance bundle can be tested against thin coded, dynamic-tool, and controlled ad-hoc Wizard-of-Oz pathways using preparation effort, fidelity, deviation, latency, safety, workload, reset, and reuse measures.","source_ids":["S2","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A maximum of 12 fictional, non-sensitive sessions across three behavior variants is reversible, comparator-based, measurable, and equipped with precommitted stop rules.","source_ids":["S5","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"No inherent stop was found for a fictional, non-consequential study if local ethics/privacy determination, prospective consent or approved alteration, withdrawal, debriefing, data minimization, and immediate suspension rules are satisfied. Approval cannot be presumed for later expansion.","source_ids":["S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four scopes are bounded and labor-equivalent ranges are anchored to current federal occupational wage data, although loaded rates, utilization, security, and organizational overhead remain uncertain.","source_ids":["S8"]}},"next_evidence_step":"With a named research-operations partner and required ethics/privacy authorization, preregister three fictional task-planning behavior contracts and run no more than 12 consented sessions, balanced as four foundry, four thin-coded, and four controlled ad-hoc Wizard-of-Oz sessions while holding task, visible surface, disclosure, staffing standard, and analysis rubric constant. Measure preparation and rehearsal hours, participant-visible state/response fidelity, latency, operator interventions and deviations, fallback use, participant comprehension of the simulated capability, discomfort or withdrawal, prohibited-data events, cross-session residue, operator workload and recovery, reset time, and analysis backlog. Advance only if the foundry shows at least 30% lower median preparation effort than the thin-coded pathway across successive variants, no greater effort than controlled ad-hoc operation, fidelity within the preregistered non-inferiority margin, and zero consent, privacy, real-action, or prohibited-claim breach. Falsify the intervention if custom work or operator labor scales roughly linearly with each hypothesis, either comparator matches cost and fidelity, simulation artifacts materially change conclusions, reset cannot be verified, or any safety stop occurs. Treat the small study as feasibility evidence, not an impact estimate.","blocking_evidence":["No live evidence shows that the complete foundry reduces preparation effort relative to modern conditional prototyping or controlled ad-hoc Wizard-of-Oz work.","No named organization has committed adoption authority, staff, budget, or a recurring study pipeline.","HRIStudio's six-session result is domain-specific and too small for inferential comparison.","The incremental value of the governance bundle over ordinary research operations, IRB review, and existing Wizard-of-Oz protocols is unmeasured.","Cross-domain behavior-module reuse, reset reliability, operator-welfare limits, and downstream analysis capacity require field observation.","Cost estimates are labor-equivalent ranges, not vendor quotes or an institution-specific staffing model.","Regulatory and ethics-review applicability varies by institution, funding, jurisdiction, participant population, and study design."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"World novelty, patentability, freedom to operate, market size, and realized impact were not measured. This bounded search found substantial collisions with established Wizard-of-Oz practice, ARNE, WebWOZ, HRIStudio, Microsoft HAX, and dynamic commercial prototyping. It supports only a potentially distinctive governance-and-service claim, not a world-novelty conclusion.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure a named research-operations adopter and documented ethics/privacy authorization for the bounded probe.","Run the preregistered 12-session, three-comparator feasibility study with blinded fidelity coding where practicable.","Demonstrate a repeated-variant preparation-effort advantage without proportional custom engineering or hidden operator labor.","Show non-inferior participant-visible fidelity and preserve deviation records and claim limitations.","Verify zero cross-session residue and successful reset after every cycle.","Measure operator workload, recovery, succession coverage, and analysis backlog under real scheduling.","Differentiate the service empirically from HRIStudio, WebWOZ, Microsoft HAX, and ordinary research-operations governance.","Replace labor-equivalent estimates with an institution-specific staffing, infrastructure, authorization, and recurring-volume cost model."],"reason":"Web evidence establishes the problem and technical feasibility, but also reveals substantial prior-art collision. The only material remaining claim is that the integrated governance, welfare, privacy-reset, evidence-release, and turnover regime produces a reusable cross-team advantage. That claim requires proprietary workflow data and live human-subject testing; bounded web research cannot resolve it."},"proposal_index":4}