{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"ritualized_meaning_and_commitment_enactment__human_computer_interaction:P4:v0","cell_id":"ritualized_meaning_and_commitment_enactment__human_computer_interaction","search_queries":["site:nist.gov AI RMF human oversight roles responsibilities change management generative AI profile","site:eur-lex.europa.eu AI Act Article 14 human oversight automation bias","human AI collaboration mental models delegation accountability research paper","AI governance tabletop exercise human oversight scenario official guide","EUR-Lex Regulation EU 2024 1689 Article 14 human oversight automation bias official","site:ai.gov.au AI scenario tabletop exercise governance official National AI Centre","site:microsoft.com responsible AI impact assessment guide human oversight accountability PDF","site:pair.withgoogle.com guidebook mental models AI user expectations","site:bls.gov employer costs employee compensation March 2026 private industry professional occupations hourly official","site:bls.gov occupational employment wages management analysts May 2025 official","Buçinca To Trust or to Think cognitive forcing functions AI overreliance CHI 2021 PDF","human oversight AI empirical automation bias approval review ceremonial primary study"],"sources":[{"source_id":"S1","title":"AI Risk Management Framework Core","publisher":"U.S. National Institute of Standards and Technology","url":"https://airc.nist.gov/airmf-resources/airmf/5-sec-core/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2023-01-26","accessed_at":"2026-08-03","claims_supported":["AI risk management should be continuous and include multidisciplinary and external perspectives.","Organizations should differentiate human-AI roles, document human oversight, assign executive responsibility, and periodically review risk management.","Post-deployment plans should cover user input, appeal, override, decommissioning, incident response, recovery, and change management.","AI actors responsible for one lifecycle segment may lack visibility or control over other segments."]},{"source_id":"S2","title":"Regulation (EU) 2024/1689 (Artificial Intelligence Act)","publisher":"Official Journal of the European Union / EUR-Lex","url":"https://eur-lex.europa.eu/eli/reg/2024/1689/oj?locale=en","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2024-07-12","accessed_at":"2026-08-03","claims_supported":["For covered high-risk systems, human oversight must enable understanding of capabilities and limitations, monitoring, resistance to automation bias, interpretation, override, reversal, intervention, and safe stopping.","Deployers must assign oversight to people with necessary competence, training, authority, and support.","Relevant impact assessments should be updated when material factors change and may involve affected-group representatives and independent experts.","The Act supports real technical and organizational oversight controls, not symbolic authorization; its applicability to a generic customer-operations deployment depends on the system's regulated use and risk classification."]},{"source_id":"S3","title":"People + AI Guidebook: Mental Models","publisher":"Google People + AI Research","url":"https://pair.withgoogle.com/guidebook-v2/chapter/mental-models/","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"2019","accessed_at":"2026-08-03","claims_supported":["Users' mental models can differ from actual AI behavior and from one another.","Mismatched expectations can produce confusion, misuse, broken trust, and overestimation of system capability.","Mental models evolve through onboarding and continued interaction, so expectation-setting is an ongoing design task."]},{"source_id":"S4","title":"Test your scenario","publisher":"Australian Government National AI Centre","url":"https://www.ai.gov.au/practical-guides-and-learning/planning-tools-and-templates/test-your-scenario","source_class":"OFFICIAL_GUIDANCE","publication_date":"2026-05-05","accessed_at":"2026-08-03","claims_supported":["Organizations using or planning to use AI are an expressly identified adopter group for team-based governance exercises.","Guided tabletop exercises with mixed roles and realistic AI failures are recommended to identify gaps, clarify responsibilities, capture actions, and agree what should change.","This is close prior art for the proposal's scenario-enactment core."]},{"source_id":"S5","title":"Microsoft Responsible AI Impact Assessment Template","publisher":"Microsoft","url":"https://blogs.microsoft.com/wp-content/uploads/prod/sites/5/2022/06/Microsoft-RAI-Impact-Assessment-Template.pdf","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"2022-06","accessed_at":"2026-08-03","claims_supported":["Established impact-assessment practice already records reviewers, lifecycle stage, intended uses, stakeholders, harms, deployment modes, and responsible oversight roles.","The template treats human oversight and control as applicable to all AI systems and asks who operates, troubleshoots, oversees, and controls a system.","Microsoft directs review at least annually, when intended uses change, and before a new release stage."]},{"source_id":"S6","title":"To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making","publisher":"Proceedings of the ACM on Human-Computer Interaction / Harvard University authors","url":"https://www.eecs.harvard.edu/~kgajos/papers/2021/bucinca2021trust.shtml","source_class":"PRIMARY_RESEARCH","publication_date":"2021-04","accessed_at":"2026-08-03","claims_supported":["In an experiment with 199 participants, cognitive-forcing designs reduced overreliance relative to simpler explainable-AI designs.","Interventions that reduced overreliance received worse subjective ratings, establishing a usability and acceptance tradeoff.","Benefits varied with participants' need for cognition, cautioning against assuming uniform benefit from an effortful enactment."]},{"source_id":"S7","title":"The Flaws of Policies Requiring Human Oversight of Government Algorithms","publisher":"Colorado Technology Law Journal / arXiv","url":"https://arxiv.org/abs/2109.05067","source_class":"PRIMARY_RESEARCH","publication_date":"2021-09-10","accessed_at":"2026-08-03","claims_supported":["A survey of human-oversight policies argues that nominal human review can legitimize faulty systems and permit responsibility avoidance.","Human-oversight arrangements require empirical justification rather than being presumed effective.","Institutional accountability is an important comparator to individual human-in-the-loop controls."]},{"source_id":"S8","title":"Average hourly employer costs for employee compensation, March 2026","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/charts/employer-costs-for-employee-compensation/costs-per-hour.htm","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026-06-12","accessed_at":"2026-08-03","claims_supported":["March 2026 total employer compensation averaged $46.60 per hour for private-industry workers, $49.32 for civilian workers, and $66.41 for state and local government workers.","These figures provide a 2026-USD labor-resource anchor; they are not vendor quotes or a measured project budget."]}],"problem_evidence":{"support":"STRONG","rationale":"External evidence supports the general problem: NIST identifies fragmented lifecycle visibility and requires differentiated human-AI roles and continuing review; Google documents changing and mismatched AI mental models; experimental and policy research documents overreliance and nominal-oversight risk. Evidence does not quantify how often the specific cross-role delegation disagreement occurs in customer operations.","source_ids":["S1","S3","S6","S7"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"The Australian National AI Centre explicitly addresses organizations using or planning to use AI and offers mixed-role scenario exercises to clarify responsibilities. NIST assigns responsibility to organizational leadership, and the EU AI Act creates authority and oversight duties for covered high-risk deployments. These demonstrate institutional pull for the underlying governance job, but no named customer-operations employer has requested the proposal's ritualized variant or committed resources.","source_ids":["S1","S2","S4","S5"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Australian National AI Centre guided AI tabletop exercises","similarity":"Uses realistic AI failure scenarios, mixed organizational roles, group discussion, gap identification, responsibility clarification, and action capture—the proposal's main operational sequence.","remaining_difference":"Does not require a symbolic authorization key, marked observer-only threshold, witnessed re-consent at each delegation transition, ritual stewardship, or ritual retirement.","source_ids":["S4"]},{"name":"Microsoft Responsible AI Impact Assessment","similarity":"Maps intended uses, affected stakeholders, harms, lifecycle stages, reviewers, oversight roles, monitoring, and updates after release or use changes.","remaining_difference":"It is a documented assessment rather than an embodied, recurring scenario enactment and does not test whether cross-role mental models become mutually visible.","source_ids":["S5"]},{"name":"NIST AI RMF Core","similarity":"Already requires differentiated human-AI roles, documented oversight, multidisciplinary participation, periodic review, change management, override, appeal, recovery, and decommissioning.","remaining_difference":"Specifies governance outcomes and processes but not a symbolic, witnessed method for renewing delegation boundaries within a workgroup.","source_ids":["S1"]},{"name":"Cognitive-forcing interfaces and procedures","similarity":"Deliberately interrupt automatic acceptance and require human analytical engagement with AI recommendations.","remaining_difference":"The tested intervention operates at individual decision time, not as a collective lifecycle ritual connecting shared interpretations to permissions and remedies.","source_ids":["S6"]}],"distinctive_claim_remaining":"Compared with an otherwise identical conventional guided tabletop, adding a recurring observer-only threshold, witnessed symbolic transfer of the authorization turn, and explicit boundary re-consent will produce greater cross-role agreement and better detection and correction of delegation drift after a simulated capability change, without increasing perceived coercion or confusing symbolism with operational authority.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"A synthetic, nonproduction implementation is technically straightforward because official guidance already supplies tabletop and impact-assessment patterns, while NIST and the EU AI Act specify the permission, logging, override, change-management, competence, authority, and remedy outputs. Production closure remains organization-specific and requires authorized administrators, workflow integration, privacy review, labor-sensitive consent design, and proof that observer-only mode is real. No external source validates the symbolic key, ritual cadence, or safe nonretaliatory workplace dissent mechanism.","source_ids":["S1","S2","S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Correctly aligning delegation, authority, accountability, and correction could prevent consequential operational and customer harm; realized impact and prevalence are unmeasured.","source_ids":["S1","S2","S7"]},"stakeholder_pull":{"score":4,"rationale":"Regulators and official governance bodies expressly demand or recommend effective oversight, clear responsibilities, mixed-role exercises, and lifecycle updating, although pull for ritualization itself is absent.","source_ids":["S1","S2","S4"]},"incremental_advantage":{"score":2,"rationale":"Most practical functions are already covered by table-top exercises, impact assessments, access controls, and periodic governance review. Incremental benefit from symbolism, witnessing, and re-consent is only a hypothesis.","source_ids":["S1","S4","S5","S6"]},"distinctiveness_plausibility":{"score":3,"rationale":"The ritualized combination is distinguishable within the bounded source set, but it is a narrow process-layer difference over close established analogues and world novelty was not tested.","source_ids":["S4","S5"]},"technical_implementability":{"score":4,"rationale":"Synthetic scenarios, mock logs, disabled actions, and proposed control mappings are readily buildable; production permission and workflow changes are feasible but system-specific.","source_ids":["S1","S2","S4","S5"]},"adoption_authority_feasibility":{"score":3,"rationale":"Business owners and administrators can authorize a sandbox and implement controls, but legal, compliance, security, privacy, labor, and affected-person authority cannot be transferred through the enactment.","source_ids":["S1","S2","S5"]},"evidence_readiness":{"score":4,"rationale":"The contrast with a conventional tabletop is testable using synthetic cases, pre/post decisions, a simulated capability change, verified mock controls, and private safety measures.","source_ids":["S4","S6"]},"safety_net_benefit":{"score":3,"rationale":"The practice may expose role ambiguity and drift before production harm, but it could also create false assurance, coerced affirmation, or ceremonial accountability unless independently audited.","source_ids":["S1","S6","S7"]},"scalability":{"score":3,"rationale":"Templates and synthetic scenarios are replicable, but meaningful enactment requires contextual capability mapping, mixed-role attendance, facilitation, technical verification, and repeated updates.","source_ids":["S1","S4","S5"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Design and run a preregistered, synthetic two-cycle comparator study with 24–36 volunteers, three conditions, private surveys, mock permissions/logs, facilitation, independent safety review, and analysis.","confidence":"MODERATE","assumptions":["Approximately 200–500 compensated staff and specialist hours.","No production data, customer contact, new enterprise software, or production integration.","BLS March 2026 compensation averages are used only as labor-resource anchors; specialist contracting may cost more."],"source_ids":["S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Adapt scenarios and records, map actual capabilities and authority, configure a safe rehearsal environment, design consent and privacy controls, and build verification links to permissions, approvals, logs, escalation, and remedies for one deployment.","confidence":"LOW","assumptions":["Approximately 1,000–3,000 cross-functional hours across operations, security, privacy/legal, engineering, design, facilitation, and affected-party review.","Existing identity, logging, sandbox, and workflow systems can be reused.","Excludes major platform replacement, litigation, regulatory certification, and remediation of pre-existing control failures."],"source_ids":["S1","S2","S5","S8"]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"Launch for one customer-operations workgroup, including production-safe capability mapping, participant preparation, two enactments, authorized control changes, validation, documentation, and escalation readiness.","confidence":"LOW","assumptions":["Approximately 800–2,500 internal and specialist hours plus modest tooling and accessibility costs.","Production changes use existing administrative APIs and approval systems.","Material security engineering or customer remediation would be separately funded."],"source_ids":["S1","S2","S4","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Quarterly and change-triggered enactments, scenario maintenance, stewardship rotation, independent harm audits, permission/log verification, accessibility support, and annual repair or retirement review for one operational program.","confidence":"LOW","assumptions":["Approximately 600–2,000 hours annually, depending on change frequency and number of integrated tools.","Four scheduled cycles plus material-change reviews.","Excludes incident losses, customer compensation, and large engineering remediations."],"source_ids":["S1","S5","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent official, first-party, and primary-research sources support changing or mismatched mental models, fragmented oversight responsibility, automation bias, and nominal human-review risk.","source_ids":["S1","S3","S6","S7"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"The Australian National AI Centre identifies organizations using or planning AI as adopters of mixed-role scenario exercises; NIST identifies executive and organizational authorizers, and EU law assigns deployer oversight authority for covered systems.","source_ids":["S1","S2","S4"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim isolates the ritual additions against an otherwise identical conventional tabletop and names measurable benefits and safety constraints.","source_ids":["S4","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A synthetic, nonproduction two-cycle comparator can measure agreement, drift detection, control correctness, retention, and coercion without changing production authority.","source_ids":["S4","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The next step can be limited to volunteers, synthetic cases, disabled external actions, private dissent, independent stopping authority, and mock controls. Production deployment would require fresh security, privacy, legal, labor, and administrative authorization.","source_ids":["S1","S2","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four ranges have bounded scopes and labor assumptions anchored to March 2026 employer-compensation data, though no vendor quote or site-specific engineering estimate was obtained.","source_ids":["S8"]}},"next_evidence_step":"Preregister and run a synthetic two-cycle study with 24–36 volunteers assigned to: (A) written policy, approval screens, and control review; (B) an official-style guided tabletop using identical scenarios and action capture; or (C) the same tabletop plus the observer-only threshold, circulating symbolic key, witnessed boundary decisions, and explicit re-consent. Privately elicit role-specific draft/recommend/modify/communicate/commit/correct decisions before the exercise, immediately afterward, and after a simulated model-or-tool change. Score cross-role agreement, correctness against a preauthorized boundary matrix, detection of permission-policy mismatches, assignment of accountable owners, completeness and simulated executability of remediation controls, retention after change, time burden, comprehension, and optional private reports of pressure. An independent reviewer must halt for coercion, confidential-data exposure, authority confusion, or uncontrolled action. Falsify incremental advantage if condition C does not materially outperform B on prespecified agreement and drift-detection thresholds, if technical closure remains unverifiable, or if C increases pressure, authority confusion, or responsibility diffusion. Do not proceed to production on attendance, affect, or subjective enthusiasm alone.","blocking_evidence":["No direct prevalence estimate for cross-role delegation-boundary disagreement in customer-operations workgroups.","No comparative trial showing that symbolic transfer, witnessing, and re-consent outperform a conventional guided tabletop.","No evidence that the symbolic key and visible choices remain noncoercive under real workplace power asymmetries.","No production demonstration that enacted decisions reliably become authenticated permissions, approval gates, logs, escalation ownership, and customer remedies.","No named employer has committed to adopt or fund the ritualized variant.","Site-specific privacy, labor, legal, compliance, security, accessibility, and affected-customer approvals remain unavailable."],"research_disposition":"PILOT_OR_ADOPTION_INQUIRY","world_novelty_boundary":"The bounded eight-source search found close prior art for every principal governance function and no exact match for the combined symbolic-key, witnessed re-consent protocol. This is not a world-novelty conclusion. World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Demonstrate prespecified incremental improvement over an otherwise identical conventional tabletop.","Show retained boundary agreement and faster detection of drift after a simulated capability change.","Verify that every enacted closure item maps to an executable permission, approval, log, escalation, or remedy control.","Show no material increase in coercion, retaliation concern, authority confusion, or responsibility diffusion.","Obtain a named adopter's conditional authorization and a site-specific implementation estimate before production use."],"reason":"Web evidence verifies the underlying oversight problem, institutional need, feasibility of a safe sandbox, and substantial collision with existing tabletop and impact-assessment practice. Whether the ritual additions create incremental benefit or new coercion and false-assurance risks requires participant testing, proprietary workflow information, and live control verification; bounded web research cannot resolve those questions."},"proposal_index":4}