{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"representation_independent_interface_contract__human_computer_interaction:P4:v0","cell_id":"representation_independent_interface_contract__human_computer_interaction","search_queries":["Wizard of Oz prototyping systematic review HCI reproducibility wizard behavior automation handoff","site:dl.acm.org Wizard of Oz prototyping operator behavior specification conversational agents study","Wizard of Oz study ethics deception IRB official guidance human subjects research","Google Calendar API event ID duplicate create idempotency official","site:dl.acm.org \"Managing Consistency in Wizard of Oz Studies\"","WebWOZ Wizard of Oz framework reproducible operator behavior paper","Suede Wizard of Oz prototyping tool speech user interfaces ACM","Wizard of Oz protocol constrained wizard behavior decision rules reproducibility HCI","site:hhs.gov OHRP deception incomplete disclosure informed consent IRB official guidance","site:ecfr.gov 45 CFR 46.116 deception waiver informed consent minimal risk","Wizard of Oz studies reporting guidelines Riek full text PDF","Managing consistency wizard performance varies significantly wizard of oz study full text","Laurel Riek Wizard of Oz Studies in HRI systematic review reporting guidelines PDF 2012","\"Wizard of Oz Studies in HRI\" \"pdf\"","site:human-robot-interaction.org Riek Wizard Oz guidelines","site:dl.acm.org wizard of oz conversational calendar system scheduling assistant prototype","site:bls.gov/ooh software developers occupational outlook handbook 2025 median pay","site:bls.gov/ooh web developers digital designers occupational outlook 2025 median pay","\"Mapping the Wizards' Path\" PDF","\"Mapping the Wizards’ Path\" authors PDF CHI 2026","3772318.3791174 pdf","\"Wizards in the Middle\" comparing humans and robots paper 2022","human wizard autonomous system shared interface comparison behavioral equivalence Wizard of Oz","Wizard of Oz transition to automation compare wizard autonomous implementation HCI"],"sources":[{"source_id":"S1","title":"Mapping the Wizards' Path: A Systematic Review of Wizard-of-Oz in HCI","publisher":"Association for Computing Machinery","url":"https://doi.org/10.1145/3772318.3791174","source_class":"PRIMARY_RESEARCH","publication_date":"2026-04-13","accessed_at":"2026-08-03","claims_supported":["A systematic review of 194 SIGCHI papers found Wizard-of-Oz concerns involving wizard variability and bias, ecological validity, deception and ethics, and transparency and reproducibility.","The review reports that operator actions, decision rules, and logging are sometimes underreported and recommends standardized, documented wizard behavior, explicit decision protocols, training, and consistency checks.","The review describes constrained interfaces, logs, hybrid wizard-automation systems, and deliberate decisions about what to simulate versus automate, establishing close prior art."]},{"source_id":"S2","title":"Managing Consistency in Wizard of Oz Studies: A Challenge of Prototyping Natural Language Interactions","publisher":"University of Edinburgh Research Explorer","url":"https://www.research.ed.ac.uk/en/publications/managing-consistency-in-wizard-of-oz-studies-a-challenge-of-proto/","source_class":"PRIMARY_RESEARCH","publication_date":"2013-08-01","accessed_at":"2026-08-03","claims_supported":["A meta-analysis of three Wizard-of-Oz studies found significant within- and between-study variation in wizard timing that could affect participant experience.","The authors identify wizard consistency as a methodological problem and call for additional operator support."]},{"source_id":"S3","title":"SUEDE: A Wizard of Oz Prototyping Tool for Speech User Interfaces","publisher":"Association for Computing Machinery","url":"https://hci.stanford.edu/publications/2000/suede/suede-uist2000.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2000","accessed_at":"2026-08-03","claims_supported":["SUEDE represents dialogue as prompt/response state transitions, restricts the wizard to valid controls and links for the current state, records transcripts, and can inject recognition errors mechanically.","Its design was based on interviews with six speech-interface designers from SRI, Sun Microsystems Laboratories, Nuance Communications, and Sun Microsystems; those practitioners wanted better support for analyzing test data and failures.","SUEDE demonstrates the technical feasibility and prior-art status of a constrained operator console, executable dialogue graph, logging, and error simulation."]},{"source_id":"S4","title":"Wizards in the Middle: An Approach to Comparing Humans and Robots","publisher":"IEEE","url":"https://doi.org/10.1109/RO-MAN53752.2022.9900655","source_class":"PRIMARY_RESEARCH","publication_date":"2022-08","accessed_at":"2026-08-03","claims_supported":["The WITM system directed both human and robot interviewers through a common collaborative wizarding system to reduce behavioral variability and improve protocol consistency.","The system was used across three studies totaling 217 child interactions, demonstrating workflow feasibility beyond a conceptual prototype.","This is close prior art for separating the visible agent from behavior-producing resources, although its robots remained wizard-mediated rather than independently automated implementations passing a shared conformance oracle."]},{"source_id":"S5","title":"Informed Consent FAQs","publisher":"U.S. Department of Health and Human Services, Office for Human Research Protections","url":"https://www.hhs.gov/ohrp/regulations-and-policy/guidance/faq/informed-consent/index.html","source_class":"OFFICIAL_GUIDANCE","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Covered non-exempt human-subject research generally requires legally effective prospective informed consent unless an IRB documents an applicable waiver or alteration.","An IRB must approve an altered consent process; disclosure, comprehension, and voluntariness remain core requirements.","Participant-facing concealed wizarding therefore requires institution-specific ethics and consent review, while the proposed synthetic, nonparticipant test avoids this authority dependency."]},{"source_id":"S6","title":"Create events — Google Calendar API","publisher":"Google for Developers","url":"https://developers.google.com/workspace/calendar/api/guides/create-events","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["Calendar events can be created with a caller-generated event ID.","Google documents that caller-generated event IDs can prevent duplicate creation after a request fails following successful backend execution, supporting an idempotent adapter design.","Calendar writes and attendee notifications are real external side effects, supporting the proposal's insistence on a sandbox ledger before network integration."]},{"source_id":"S7","title":"Software Developers, Quality Assurance Analysts, and Testers — Occupational Outlook Handbook","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/ooh/computer-and-information-technology/software-developers.htm","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2025-08-28","accessed_at":"2026-08-03","claims_supported":["May 2024 median annual wages were $133,080 for software developers and $102,610 for software quality-assurance analysts and testers.","These wage benchmarks support resource-equivalent estimates for contract adapters, test harnesses, mutation testing, integration, and maintenance."]},{"source_id":"S8","title":"Web Developers and Digital Designers — Occupational Outlook Handbook","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/ooh/computer-and-information-technology/web-developers.htm","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2025-08-28","accessed_at":"2026-08-03","claims_supported":["May 2024 median annual wages were $98,090 for web and digital-interface designers and $90,930 for web developers.","These wage benchmarks support resource-equivalent estimates for interaction design, the operator console, and participant-facing workflow work."]}],"problem_evidence":{"support":"STRONG","rationale":"The problem is visible beyond the candidate. A 2026 review of 194 SIGCHI papers identifies wizard variability, bias, transparency, reproducibility, and ethics as recurring concerns and recommends standardized decision protocols and logging. A separate three-study meta-analysis measured significant wizard-performance variation capable of changing participant experience. These sources establish the general methodological problem, though they do not estimate its prevalence specifically in conversational scheduling or quantify downstream wizard-to-automation failures.","source_ids":["S1","S2"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Identifiable HCI stakeholders exist: SUEDE was designed from interviews with six practitioners at SRI, Sun, and Nuance who requested better prototype-analysis and failure support, and current CHI guidance calls on researchers to standardize wizard protocols. Research leads and prototype owners can adopt an offline contract, while an IRB is the recognizable authorizer for participant-facing concealment or altered consent. No organization has expressed demand for this exact scheduling-session contract, committed a budget, or agreed to run the proposed study.","source_ids":["S1","S3","S5"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"SUEDE constrained Wizard-of-Oz dialogue prototyping","similarity":"Provides an executable prompt/response state graph, exposes only valid current-state controls to the wizard, records sessions, and injects errors mechanically—most of the proposed operator-boundary and logging logic.","remaining_difference":"It does not establish a representation-independent semantic contract shared with independently implemented automation, formalize scheduling side-effect invariants, or use seeded nonconforming implementations as an acceptance test.","source_ids":["S3"]},{"name":"Wizards in the Middle shared control layer","similarity":"Separates visible human or robot agents from common behavior-producing resources and routes both through one protocol-oriented system to improve behavioral consistency.","remaining_difference":"Both agent conditions remained wizard-mediated; the work did not test an autonomous replacement against a black-box scheduling-session oracle or confirmation/idempotency side effects.","source_ids":["S4"]},{"name":"2026 standardized Wizard-of-Oz protocol guidance","similarity":"Recommends constraint-based simulations, explicit decision rules, standardized wizard behavior, training, consistency checks, and logging while directly considering the boundary between wizarding and automation.","remaining_difference":"It is methodological guidance rather than an executable, domain-specific abstract state contract with a substitution gate and mutation-tested calendar ledger.","source_ids":["S1"]},{"name":"Wizard-consistency support","similarity":"Empirically identifies operator variability as capable of altering the user experience and calls for added support to improve consistency.","remaining_difference":"It does not define a common abstraction mapping, automated comparator, or side-effect safety oracle.","source_ids":["S2"]}],"distinctive_claim_remaining":"Relative to constrained WoZ tools and standardized-protocol guidance, the remaining falsifiable claim is that one predeclared semantic state-and-side-effect contract can govern a console-constrained human wizard and an independently built automated scheduling implementation, reject specified unsafe mutants, and still permit noncontractual variation in wording, timing, and internal reasoning.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Constrained wizard consoles, state-transition graphs, shared human/robot control systems, logging, and injected failures have all been implemented and used in studies. Calendar APIs support stable caller-generated IDs useful for retry idempotency. Thus an offline synthetic implementation is technically credible. Evidence does not yet show that the proposed semantic response-act vocabulary is complete, that independently authored implementations will agree, that wording and timing can safely remain noncontractual, or that a participant-ready system will pass privacy, security, and ethics review.","source_ids":["S1","S3","S4","S6","S5"]},"scores":{"meaningful_impact":{"score":4,"rationale":"If successful, the intervention would prevent prototype findings from being attributed to capabilities unavailable to automation and would detect premature or duplicate bookings. Realized impact and prevalence remain unmeasured.","source_ids":["S1","S2","S6"]},"stakeholder_pull":{"score":3,"rationale":"HCI practitioners and reviewers visibly want better consistency, analysis, and standardized wizard protocols, but no named adopter has requested or funded the exact scheduling contract.","source_ids":["S1","S3"]},"incremental_advantage":{"score":2,"rationale":"The proposal improves on transcript-only handoff by adding executable invariants and negative controls, but constrained consoles, state machines, logging, common control layers, and protocol standardization already cover much of the intervention.","source_ids":["S1","S3","S4"]},"distinctiveness_plausibility":{"score":2,"rationale":"The formal cross-implementation contract plus seeded calendar-side-effect mutants is plausibly narrower than the located analogues, but the conceptual core substantially collides with established WoZ tooling and protocol practice.","source_ids":["S1","S3","S4"]},"technical_implementability":{"score":4,"rationale":"Existing tools demonstrate constrained state-based wizarding and shared agent-control workflows, while calendar APIs provide identifiers useful for idempotency. Semantic completeness and independent conformance remain untested.","source_ids":["S3","S4","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"A research lead can authorize a synthetic offline study. Participant-facing deception or incomplete disclosure requires IRB review, and production calendar integration requires service-owner, security, and privacy authority not yet identified.","source_ids":["S5","S6"]},"evidence_readiness":{"score":4,"rationale":"The proposal supplies a bounded trace set, explicit observables, five mutants, comparators, and halt conditions. The missing evidence is executable empirical output rather than further conceptual specification.","source_ids":["S1","S3","S4"]},"safety_net_benefit":{"score":4,"rationale":"Preconfirmation-write and duplicate-retry mutants directly test consequential side effects, and a non-networked ledger keeps first evidence reversible. This does not establish production safety.","source_ids":["S6"]},"scalability":{"score":3,"rationale":"A reusable contract harness can be parameterized across implementations, but semantic-act taxonomies, adapters, leakage reviews, and stewardship remain domain-specific and labor-intensive.","source_ids":["S1","S3","S4"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Two-to-six-week offline mutation study: specify 20 traces, implement a synthetic ledger, deterministic model, constrained wizard adapter, independent automated adapter, five mutants, and an audit report.","confidence":"MODERATE","assumptions":["Approximately 6-12 loaded person-weeks split among an HCI researcher/designer, software developer, and QA/test contributor.","May 2024 wage medians are escalated modestly to 2026 and loaded by roughly 35%-60% for benefits, facilities, management, and tools.","No participant recruitment, network calendar, legal engagement, or production security work is included."],"source_ids":["S7","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Reusable prototype infrastructure: contract schema, operator console, adapters, conformance and property tests, leakage instrumentation, documentation, access controls, and sandbox deployment.","confidence":"MODERATE","assumptions":["Roughly 3-10 loaded staff-months across interaction design, software engineering, QA, and research stewardship.","Existing identity, hosting, logging, and test infrastructure can be reused.","Only synthetic calendars and nonparticipant rehearsals are included."],"source_ids":["S7","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Participant-ready and integration-ready launch: hardened system, privacy/security review, ethics protocol, consent or approved alteration, recruitment and pilot operations, calendar-service integration, monitoring, rollback, and incident procedures.","confidence":"LOW","assumptions":["A multidisciplinary team works for roughly 9-24 loaded staff-months.","The organization already has an IRB or equivalent review pathway and enterprise identity/calendar infrastructure.","The band excludes large-scale product development, model training, indemnification, and organization-wide deployment."],"source_ids":["S5","S6","S7","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Contract stewardship, adapter and test maintenance, regression runs, audit-log review, platform updates, incident exercises, and limited study operations.","confidence":"LOW","assumptions":["Approximately 0.4-1.3 loaded full-time-equivalent staff plus modest hosting and review costs.","Usage remains research-scale or limited internal deployment.","Major model replacement, new domains, participant recruitment waves, and security incidents would require separate funding."],"source_ids":["S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"A large current systematic review and an independent three-study meta-analysis document wizard variability, reproducibility, and consistency problems that can affect participant experience.","source_ids":["S1","S2"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"HCI research and interaction-design teams are credible adopters: named practitioners at SRI, Sun, and Nuance expressed adjacent tooling needs, and current CHI guidance addresses this research community. An IRB is the externally established authorizer for participant-facing altered consent. Exact-candidate adoption commitment remains absent.","source_ids":["S1","S3","S5"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim can be tested by running human-wizard and independently authored automated implementations against the same predeclared semantic-state and ledger oracle while injecting specified mutants.","source_ids":["S3","S4","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"Twenty synthetic traces, three valid implementations, five named mutants, contract-level comparators, and explicit continuation and falsification conditions bound the next study.","source_ids":["S3","S4","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"For the next step only, synthetic availability and a non-networked ledger avoid participant, privacy, and external side-effect authority. Any participant study or real calendar connection remains stopped pending IRB, privacy/security, and service-owner approval.","source_ids":["S5","S6"]},"credible_cost_scope_and_range":{"status":"YES","reason":"All four bands identify included work and exclusions and are anchored to official software, QA, web, and interface-design wage benchmarks, with explicit loading and staffing assumptions.","source_ids":["S7","S8"]}},"next_evidence_step":"Pre-register a 20-trace, synthetic-calendar mutation study lasting no more than six weeks. Before coding, freeze the abstract state, semantic response-act vocabulary, error rules, confirmation and idempotency invariants, and which wording/timing differences are explicitly noncontractual. Have separate contributors implement (1) a deliberately simple deterministic model, (2) the console-constrained human wizard, and (3) a rules-based automated adapter without sharing implementation code. Compare semantic acts, abstract-state transitions, error classes, and ledger diffs after every operation; separately record wording and latency to detect leakage without making them automatic failures. Seed five mutants: undeclared constraint use, missing clarification, preconfirmation booking, duplicate retry booking, and expired-proposal selection. Continue only if all five mutants are rejected, all three valid implementations agree on every declared obligation, at least two permissible wording variants pass, and the wizard never requires undeclared information. Falsify the intervention if any mutant passes, any declared divergence survives triage, acceptable wording variation is rejected because the oracle overfits, or operator-only information is necessary.","blocking_evidence":["No executed mutation-study results exist; the candidate's central incremental claim therefore remains empirically unverified.","No evidence establishes that the proposed semantic response-act vocabulary is sufficiently discriminating without overconstraining natural conversation.","No participant evidence shows that wording, timing, social inference, or recovery style can safely remain outside the behavioral contract.","No prevalence estimate exists for unconstrained wizard-to-automation handoff failures in conversational scheduling.","No named organization has committed to adopt, fund, or steward this exact contract.","Participant-facing ethics approval, production privacy/security review, and real calendar-service authorization have not been obtained.","World novelty, patentability, freedom to operate, market size, and realized impact are unmeasured."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"World novelty is unmeasured. This bounded search establishes substantial collision with constrained Wizard-of-Oz consoles, standardized protocol guidance, logging/error tooling, and shared human/robot control layers; it does not constitute an exhaustive literature, product, standards, or patent search. The only preserved novelty boundary is the narrow, unverified combination of a representation-independent scheduling contract, independently implemented human and automated adapters, and mutation-tested calendar side-effect semantics.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Pre-register the semantic contract, noncontractual observables, 20 traces, five mutants, and pass/fail thresholds before implementations are run.","Demonstrate that every seeded unsafe or nonconforming implementation is rejected and report mutation-specific results.","Demonstrate contract-level agreement among a deterministic model, console-constrained wizard, and independently authored automated adapter without shared implementation logic.","Show that at least two semantically equivalent wording variants pass while declared clarification, confirmation, cancellation, retry, and side-effect divergences fail.","Complete a leakage audit showing that success requires no operator-only information and document every observable timing or wording dependency.","Differentiate the resulting evidence explicitly from SUEDE-style constrained state graphs, standardized WoZ protocols, and the Wizards-in-the-Middle shared control layer.","Before any participant or networked phase, obtain the applicable IRB/ethics determination, consent or alteration approval, privacy/security review, and calendar-service-owner authorization."],"reason":"Web research has resolved the general problem and prior-art questions sufficiently to show both strong need and substantial collision. The remaining incremental claim cannot be resolved by further bounded web search: it requires executing independent implementations, mutants, and the shared oracle. Under the controller rule, live testing requires STOP_EMPIRICAL_RESEARCH_NEEDED; repairable is false because advancement depends on new empirical evidence rather than another proposal revision."},"proposal_index":4}