{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"bounded_rivalry_governance__human_computer_interaction:P1:v0","cell_id":"bounded_rivalry_governance__human_computer_interaction","search_queries":["site:microsoft.com research Guidelines for Human-AI Interaction CHI 2019 PDF Amershi","site:nist.gov AI RMF human oversight roles responsibilities documentation test evaluation","automation bias AI assisted decision making human corrections primary research interface overreliance 2021","online controlled experiments A/B testing organizational teams metrics guardrails Microsoft experimentation platform primary research","\"To Trust or to Think\" cognitive forcing functions CHI 2021 ACM","site:microsoft.com HAX Toolkit human AI interaction workbook official","site:w3.org/TR/WCAG22 keyboard screen reader accessibility conformance official","site:bls.gov occupational employment wage user experience researcher software developer 2025","site:hhs.gov OHRP quality improvement activities human subjects research generalizable knowledge official","site:ecfr.gov 45 CFR 46 human subject research definition official","site:bls.gov OEWS web and digital interface designers software developers May 2025 wages","human AI decision support interface competition default A/B test accessibility held-out evaluation research","doi 10.1145/3449287 ACM To Trust or to Think full paper"],"sources":[{"source_id":"S1","title":"To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making","publisher":"Proceedings of the ACM on Human-Computer Interaction; author manuscript hosted by arXiv","url":"https://arxiv.org/abs/2102.09692","source_class":"PRIMARY_RESEARCH","publication_date":"2021-02-19","accessed_at":"2026-08-03","claims_supported":["People can accept incorrect AI suggestions; immediate suggestions and explanations do not reliably prevent overreliance.","In the reported experiment, cognitive-forcing interfaces improved correctness and reduced overreliance on incorrect suggestions relative to simple explainable-AI interfaces on specified measures.","Preference and perceived simplicity can diverge from objective performance, supporting decision quality rather than acceptance or speed as the primary endpoint.","Interface effects varied with individual cognitive motivation, supporting heterogeneous-participant evaluation."]},{"source_id":"S2","title":"Guidelines for Human-AI Interaction","publisher":"Microsoft Research; CHI 2019 organized by ACM","url":"https://www.microsoft.com/en-us/research/publication/guidelines-for-human-ai-interaction/","source_class":"PRIMARY_RESEARCH","publication_date":"2019-05","accessed_at":"2026-08-03","claims_supported":["Eighteen human-AI interaction guidelines were evaluated in multiple rounds, including a study in which 49 design practitioners assessed 20 AI-infused products.","The guidance covers interaction when the system is wrong and supports correction, dismissal, and appropriate reliance as legitimate interface objectives.","The authors report remaining knowledge gaps and a need for further study rather than presenting the guidelines as a sufficient evaluation method."]},{"source_id":"S3","title":"Experimentation Platform (ExP)","publisher":"Microsoft Research","url":"https://www.microsoft.com/en-us/research/group/experimentation-platform-exp/","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"undated","accessed_at":"2026-08-03","claims_supported":["Microsoft operates a foundational experimentation platform whose stated mission is trusted, low-friction experimentation for AI builders.","The platform embeds experimentation into product-development workflows and supports engineers, scientists, and product teams in hypothesis validation, impact measurement, and safer iteration.","Microsoft is an identifiable analogous adopter with the technical and organizational capacity to authorize governed interface comparisons, although this source does not express demand for the proposal's rivalry-specific rules."]},{"source_id":"S4","title":"Guidelines for Human-AI Interaction — Microsoft HAX Toolkit","publisher":"Microsoft","url":"https://www.microsoft.com/en-us/haxtoolkit/ai-guidelines/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"undated","accessed_at":"2026-08-03","claims_supported":["The HAX Toolkit targets teams planning AI applications and provides a workbook for jointly prioritizing human-AI guidelines.","Microsoft reports that practitioner demand arose from perceived ambiguity and a lack of rules and tools for AI design.","The toolkit addresses design planning and failure handling but does not provide the proposed scarce-slot contest, resource caps, appeals, or challenger-access mechanism."]},{"source_id":"S5","title":"Artificial Intelligence Risk Management Framework Core","publisher":"U.S. National Institute of Standards and Technology","url":"https://airc.nist.gov/airmf-resources/airmf/5-sec-core/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2023-01-26","accessed_at":"2026-08-03","claims_supported":["NIST calls for documented accountability, differentiated human-AI roles, executive responsibility, multidisciplinary perspectives, accessibility, and ongoing review.","NIST calls for objective, repeatable or scalable test, evaluation, verification, and validation processes with documented metrics and methods.","NIST supports feedback, monitoring, incident handling, and decommissioning, making a reversible governed trial technically and organizationally plausible.","The framework assigns responsibility to organizational leadership but does not itself authorize a specific workplace experiment."]},{"source_id":"S6","title":"Web Content Accessibility Guidelines (WCAG) 2.2","publisher":"World Wide Web Consortium","url":"https://www.w3.org/TR/WCAG22/","source_class":"STANDARD","publication_date":"2024-12-12","accessed_at":"2026-08-03","claims_supported":["WCAG 2.2 is a W3C Recommendation for web accessibility and includes keyboard operation, error identification, predictable interaction, compatibility, and non-interference requirements.","Accessibility can be operationalized as threshold criteria rather than offset by improvements in engagement or speed.","WCAG conformance supplies a standard comparator but does not by itself establish usability with screen readers in the proposal's document-review context."]},{"source_id":"S7","title":"Quality Improvement Activities FAQs","publisher":"U.S. Department of Health and Human Services, Office for Human Research Protections","url":"https://www.hhs.gov/ohrp/regulations-and-policy/guidance/faq/quality-improvement-activities/index.html","source_class":"OFFICIAL_GUIDANCE","publication_date":"undated","accessed_at":"2026-08-03","claims_supported":["An operational quality-improvement activity is not automatically research under the Common Rule, but an activity designed to develop or contribute to generalizable knowledge may be research.","Applicability depends on whether the activity is research, involves human subjects, qualifies for an exemption, and is covered by HHS support or an applicable assurance.","Some covered non-exempt human-subjects research requires IRB review, while minimal-risk research may be eligible for expedited review.","Consent alone does not resolve the classification or authorization question; institutional determination remains necessary."]},{"source_id":"S8","title":"Occupational Employment and Wages in San Jose-Sunnyvale-Santa Clara — May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/regions/west/news-release/occupationalemploymentandwages_sanjose.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-07-09","accessed_at":"2026-08-03","claims_supported":["May 2025 mean annual wages were $221,710 for software developers, $200,670 for web and digital interface designers, $160,040 for software quality-assurance analysts and testers, and $212,760 for data scientists in the San Jose area.","These wage data provide a defensible high-cost labor anchor for broad 2026 resource-equivalent estimates.","The wage data exclude benefits, overhead, participant compensation, and opportunity cost, which must be added as assumptions."]}],"problem_evidence":{"support":"MODERATE","rationale":"Primary HCI evidence shows that interface presentation can change correction of bad AI advice and that preference or simplicity can diverge from correctness. Microsoft and NIST also visibly treat trustworthy human-AI experimentation as an organizational need. No direct source verifies the candidate's more specific prevalence claim that internal interface teams manipulate cohorts, escalate participant use, collude, interfere, or capture integration access while competing for one default slot.","source_ids":["S1","S2","S3","S4","S5"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Microsoft is an identifiable analogous adopter: it operates an experimentation platform for AI builders and a HAX workflow built in response to practitioner demand for clearer human-AI design tools. NIST identifies executive leadership and cross-functional risk roles as appropriate authorizers. None of the eight sources records a commitment by Microsoft or another named organization to adopt the complete rivalry-governance package or confirms a currently contested default-interface slot.","source_ids":["S3","S4","S5"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Microsoft Experimentation Platform (ExP)","similarity":"Shared experimentation infrastructure, hypothesis testing, impact measurement, trustworthy workflows, and safer product iteration substantially overlap the common-arena and audited-comparison components.","remaining_difference":"The opened source does not describe one scarce default slot governed through equal participant and development caps, competitor due process, anti-interference rules, or a challenger window.","source_ids":["S3"]},{"name":"Microsoft Guidelines for Human-AI Interaction and HAX Toolkit","similarity":"Provides evidence-based human-AI design criteria, practitioner workflow support, failure planning, correction, and recovery concepts directly relevant to scoring the interfaces.","remaining_difference":"It guides collaborative product design rather than governing strategically dependent teams competing for a scarce deployment advantage.","source_ids":["S2","S4"]},{"name":"NIST AI RMF plus WCAG 2.2","similarity":"Supplies governance roles, repeatable TEVV, monitoring, accessibility, feedback, review, and threshold-like requirements that cover much of the proposed rulebook and safety layer.","remaining_difference":"Neither source specifies a rival-team tournament, equal resource allocation, collusion screens, appeals over comparative scoring, or reopening an incumbent's default slot.","source_ids":["S5","S6"]},{"name":"Cognitive-forcing evaluation of AI-assisted decisions","similarity":"Directly tests interface variants using correctness and overreliance outcomes and demonstrates that friction can improve detection of unsuitable AI advice despite preference tradeoffs.","remaining_difference":"It evaluates interface treatments, not the organizational governance of teams selecting cohorts, resources, rules, or future access.","source_ids":["S1"]}],"distinctive_claim_remaining":"For multiple teams competing for one reversible default-interface slot, adding one frozen arena that jointly controls cohorts, tasks, development and participant resources, prohibited cross-team conduct, contestability, and the winner's future access will produce a ranking more consistent with blinded correctness-and-harm review than ordinary governed A/B testing that uses the same held-out tasks, accessibility thresholds, and expert review but lacks those rivalry controls. The claim is falsified if rankings, manipulation opportunities, participant burden, or challenger access do not materially differ, or if the failures persist with a single noncompeting team.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"Common sandbox builds, randomized evaluation, audit logs, reproducible scoring, accessibility checks, human-AI failure tasks, organizational TEVV, and reversible review are all supported by established research, standards, and first-party practice. The proposal still lacks empirical evidence that masked judging will remain masked, multidimensional scoring will be reliable, resource caps will be fair to incumbents and challengers, collusion screens will be useful with only three teams, and appeals will justify their overhead. Human-subjects or workplace-research classification must be determined locally.","source_ids":["S1","S2","S3","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Incorrectly accepting AI advice and overlooking accessibility can affect decision quality and affected document subjects; the intervention could redirect optimization toward correction and recovery. Realized impact and local prevalence remain unmeasured.","source_ids":["S1","S2","S6"]},"stakeholder_pull":{"score":3,"rationale":"Microsoft demonstrates organizational investment in trustworthy experimentation and reports practitioner demand for HAX guidance, but there is no expressed demand or adoption commitment for rivalry-specific governance.","source_ids":["S3","S4"]},"incremental_advantage":{"score":3,"rationale":"Resource equality, cross-team conduct rules, appeals, shared infrastructure, and reopening address strategic dependencies omitted from the opened A/B, HAX, NIST, and WCAG sources. Whether they improve selection enough to justify overhead requires a direct comparator.","source_ids":["S3","S4","S5","S6"]},"distinctiveness_plausibility":{"score":3,"rationale":"No complete match was found in the bounded search, while nearly every technical, HCI, accessibility, and governance component has close prior art. Distinctiveness rests on their integration around a scarce slot, not on a new interface-testing method.","source_ids":["S1","S2","S3","S4","S5","S6"]},"technical_implementability":{"score":4,"rationale":"A synthetic-data sandbox, reproducible builds, held-out tasks, accessibility testing, randomized participant allocation, and audit logging use mature methods. Reliable composite scoring, masking, and low-sample collusion analysis remain design risks.","source_ids":["S1","S3","S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"The proposed product owner, research-governance lead, privacy and accessibility reviewers, and independent chair align with NIST's role-based accountability. Actual organizational authority and the human-subjects determination are not externally verified.","source_ids":["S5","S7"]},"evidence_readiness":{"score":4,"rationale":"Three prototypes, synthetic tasks, a frozen protocol, hidden variants, blinded outcome review, and an ordinary-metrics comparator form a bounded test. Sample size, power, scoring weights, reliability thresholds, and preregistered falsifiers still need specification.","source_ids":["S1","S2","S3"]},"safety_net_benefit":{"score":4,"rationale":"Synthetic documents, consented participants, non-production execution, noncompensable accessibility and severe-error thresholds, independent review, and rollback materially limit first-step harm. Workplace privacy and research classification remain unresolved locally.","source_ids":["S5","S6","S7"]},"scalability":{"score":3,"rationale":"Shared experimentation infrastructure and standards can be reused, but masked panels, accessibility sessions, appeals, audits, and task-specific ground truth impose recurring expert labor and may be disproportionate for low-stakes interface choices.","source_ids":["S3","S6","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Preregister and run one four-week, non-production shadow comparison of three existing prototypes using synthetic documents and consented internal participants; includes protocol design, task authoring, participant compensation or backfill, accessibility sessions, blinded scoring, audit, analysis, and one appeal rehearsal.","confidence":"MODERATE","assumptions":["Existing prototypes, sandbox APIs, and internal research tooling are reusable.","Approximately 12-30 loaded specialist person-weeks are required across UX research, interface engineering, accessibility, data science, privacy, and evaluation.","Loaded labor is estimated above cash wages to include benefits, overhead, and opportunity cost; BLS San Jose wages provide a conservative high-cost anchor.","No production integration, customer recruitment, or new model training is included."],"source_ids":["S3","S8"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Build or harden reusable submission packaging, hidden-task management, participant allocation, audit logging, scoring dashboards, accessibility harnesses, conflict controls, evidence publication, and appeal workflow before a formal six-week trial.","confidence":"LOW","assumptions":["Requires roughly 1-3 loaded FTE-years across engineering, research operations, data science, accessibility, security/privacy, and governance.","Existing experimentation infrastructure reduces greenfield platform work.","External legal counsel, procurement, major data licensing, and production model changes are excluded."],"source_ids":["S3","S5","S6","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Operate one six-week formal trial and one eight-week reversible default pilot, including participant operations, panel time, accessibility and privacy review, build audits, incident handling, an appeal window, pilot telemetry, and post-pilot review.","confidence":"LOW","assumptions":["Three teams enter with functional builds.","Production integration uses common existing APIs and does not require core workflow replacement.","The band counts participant burden and competing-team time as resource-equivalent cost.","No severe incident, litigation, or full rerun occurs."],"source_ids":["S3","S5","S6","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain the common evaluation infrastructure and run one to three challenger or reselection rounds per year, with task refresh, accessibility regression testing, privacy review, audit retention, appeals, and post-contest recalibration.","confidence":"LOW","assumptions":["A small standing platform and governance function is retained, with specialists contributing part time.","Task leakage requires periodic hidden-set renewal.","Participant volume remains bounded and synthetic or approved internal data remain adequate.","The estimate excludes the ordinary product-development budgets of competing teams beyond contest caps."],"source_ids":["S3","S5","S6","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Primary research verifies the underlying HCI problem that attractive or simple AI interfaces can induce inappropriate reliance and that objective correction performance can diverge from preference. The exact prevalence of strategic team behavior is not verified, but the problem needed to motivate a bounded test is externally supported.","source_ids":["S1","S2","S4"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Microsoft is an identifiable analogous adopter operating both an AI experimentation platform and a human-AI design toolkit; NIST identifies executive and cross-functional governance roles capable of authorizing such testing. This is credibility evidence, not an adoption commitment.","source_ids":["S3","S4","S5"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim contrasts the full rivalry-governed arena with ordinary governed A/B testing holding tasks, accessibility thresholds, and expert review constant, and specifies ranking, manipulation, burden, and access outcomes that can falsify it.","source_ids":["S1","S3","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A four-week synthetic-data shadow exercise with three existing prototypes, fixed quotas, hidden tasks, blinded scoring, build audits, a red-team, and a no-deployment rule is time-, data-, and authority-bounded.","source_ids":["S1","S3","S5","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"UNCERTAIN","reason":"The non-production and synthetic-data limits substantially reduce risk, but neither consent nor an internal quality-improvement label determines whether local workplace-research, privacy, accessibility, labor, or Common Rule review applies. A documented institutional determination and named authorizer are still required before participant enrollment.","source_ids":["S5","S6","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"All four estimates state included work and assumptions, use broad bands, and are anchored to current government wage data for the relevant high-cost technical labor market. Infrastructure reuse and participant volume remain uncertain.","source_ids":["S3","S8"]}},"next_evidence_step":"After a documented institutional research/privacy determination, preregister a four-week non-production study of three existing prototypes on one synthetic document-review task family. Randomly allocate a fixed participant quota; include keyboard and screen-reader users; freeze eligibility, permitted changes, resource accounting, scoring, and appeals before submission; retain hidden task variants and deliberately seed unsuitable AI suggestions. Primary comparator A is ordinary governed A/B testing using acceptance rate, completion time, held-out tasks, accessibility thresholds, and expert review but no cross-team resource, conduct, appeal, or future-access rules. Comparator B is the full bounded-rivalry protocol. Blinded reviewers must independently score correctness, correction of unsuitable suggestions, comprehension, recovery, accessibility, and time; report inter-rater reliability and disaggregated outcomes. Audit builds and resource use, red-team cohort selection, leakage, deceptive defaults, evaluator contact, off-books effort, and incumbent-asset advantages, and rehearse one appeal. Falsify the incremental claim if protocol B does not improve agreement with blinded correctness-and-harm review, does not close prohibited strategies, does not reduce unequal participant/resource use, materially disadvantages challengers through incumbent assets, or adds governance burden exceeding the value of the decision. Do not select or deploy a production winner.","blocking_evidence":["No direct evidence establishes how often internal interface teams actually manipulate cohorts, consume unequal participant resources, interfere, collude, or capture future integration access when competing for a default slot.","No controlled comparison shows that rivalry-specific controls improve rank validity beyond an ordinary A/B protocol already using held-out tasks, accessibility thresholds, guardrail metrics, and expert review.","No named organization has committed prototypes, participants, staff time, or decision authority to the shadow exercise.","The protocol lacks a power analysis, minimum sample, scoring weights, inter-rater reliability threshold, minimum meaningful rank improvement, and maximum acceptable governance burden.","The feasibility of masking team identity, detecting cap circumvention, and interpreting collusion screens with only three internal teams is untested.","Institutional classification under workplace research, privacy, labor, and any applicable human-subjects rules has not been documented.","Synthetic and internal-participant performance may not generalize to production documents, affected populations, assistive-technology users, repeated use, or high-stakes automation bias."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This evaluation measured neither world novelty, patentability, freedom to operate, market size, nor realized impact. The eight-source search found established prior art for human-AI correction-oriented interfaces, HAX design guidance, enterprise A/B infrastructure, organizational AI risk governance, accessibility thresholds, and human-subjects review boundaries. It did not find the complete combination of a scarce internal default slot, equal team resources, anti-interference and collusion referral, comparative due process, common infrastructure, and a challenger reopening. Absence from this bounded search is not evidence of worldwide novelty; the defensible boundary is only an apparently uncommon integration of established components whose incremental benefit remains untested.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain a written institutional determination covering research status, participant consent, workplace privacy, data retention, accessibility review, and the named official authorized to halt or approve the exercise.","Preregister the ordinary-governed-A/B comparator and full-rivalry comparator, sample and power rationale, scoring weights, noncompensable thresholds, inter-rater reliability target, minimum meaningful rank-agreement improvement, and maximum acceptable governance burden.","Run the bounded shadow exercise and demonstrate that the full protocol improves agreement with blinded correctness-and-harm review while reducing unequal resource use or feasible manipulation; stop if it does not.","Demonstrate through audit and red-team evidence that cohort selection, task leakage, deceptive defaults, evaluator contact, off-books effort, and incumbent-asset circumvention are detectable without disproportionate workplace surveillance.","Document whether at least one qualified challenger can realistically satisfy entry and unseat rules using shared interfaces, and whether one rehearsed appeal can be resolved consistently within the time box.","Do not make a production, personnel, compensation, or disciplinary decision from the shadow ranking; production generalization requires a separately authorized pilot."],"reason":"Bounded web research verifies the underlying human-AI reliance problem, credible experimentation and governance adopters, mature technical components, adjacent prior art, and a testable remaining contrast. It cannot establish the organization-specific strategic problem, comparative rank improvement, cap fairness, masking, audit effectiveness, participant acceptability, or legal classification. Those questions require proprietary workflow access, human participants, and live protocol execution, so the next valid decision point is empirical rather than further general web research."},"proposal_index":1}