{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","cell_id":"bounded_rivalry_governance__human_computer_interaction","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_id":"brg_hci_governed_default_ui_trial_001","proposal_index":1,"version":0,"title":"Governed Default-Interface Trial for AI-Assisted Review","problem":"Several internal interface teams are competing for one default deployment slot in an AI-assisted document-review workflow. Under separate A/B tests, each team can improve its apparent standing by choosing favorable user cohorts, optimizing acceptance clicks instead of decision quality, consuming more research participants, or omitting friction that helps users detect and correct bad suggestions. The eventual winner also gains integration access and interaction data that make later challenges harder, so winning can become control of the future comparison arena rather than evidence of the best human-computer interaction.","actors":["Interface teams competing for the default slot","Document reviewers who use the interfaces","People whose documents are reviewed","UX research operations staff","Accessibility and privacy reviewers","Independent evaluation panel","Workflow product owner","Design-system and integration maintainers","Qualified future challenger teams"],"observable_state":"For the same review workflow, teams run differently timed tests on different cohorts and report noncomparable success measures; acceptance rate and completion time dominate decision-correction, comprehension, accessibility, and recovery evidence; participant contacts and experimentation effort are not jointly budgeted; evaluation changes remain possible after teams see interim results; and the incumbent interface controls integration hooks or telemetry needed by challengers.","consequence":"The organization can select an interface that produces attractive local metrics while weakening error detection, informed override, accessibility, or maintainability. Repeated experiments can burden users, teams can divert effort into metric and cohort strategy, and the selected team can acquire an unrelated advantage in later selection rounds.","affected_objective":"Select a usable, accessible interface that supports accurate and appropriately skeptical human review while retaining useful discovery pressure among alternative designs and limiting burdens on users and maintainers.","intervention":"Create a six-week bounded trial for the single prize of an eight-week, reversible default-interface pilot. Before submissions, an independent rule-setting group freezes eligibility, shared tasks, permitted interface changes, prohibited conduct, resource limits, evaluation criteria, tie-breaks, and appeals. Eligible teams receive the same sandbox APIs, held-out task families, participant allocation, and design-time budget. Winning is based on audited task correctness, detection and correction of unsuitable AI suggestions, user comprehension, keyboard and screen-reader completion, and recovery from mistakes; completion time is secondary, and privacy, accessibility, or severe-error thresholds cannot be offset by gains elsewhere. Teams may alter interaction design but may not alter task labels, recruit outside the common participant pool, preselect favorable cases, use deceptive defaults, interfere with another submission, coordinate scores, or contact evaluators. A masked panel evaluates reproducible builds, publishes evidence after scoring, hears time-boxed appeals, and refers suspicious cross-team patterns for investigation rather than treating them as proof. Production ownership remains with the workflow owner, common evaluation infrastructure remains available to challengers, a qualified challenger window reopens the slot after the pilot, and a post-pilot review determines whether the scoring and arena boundaries should be revised or the contest retired.","structural_mapping":[{"archetype_element":"Rivalry purpose statement","domain_realization":"Use competition to discover which interaction design best supports reliable human review, not to maximize team status or raw AI-acceptance activity."},{"archetype_element":"Scarce prize or selection constraint","domain_realization":"One interface receives an eight-week reversible default-deployment pilot because the workflow cannot expose reviewers to several simultaneous defaults."},{"archetype_element":"Competitor eligibility boundary","domain_realization":"An internal team may enter only with a reproducible sandbox build, required accessibility checks, approved data handling, disclosed dependencies, and no evaluator conflict."},{"archetype_element":"Contest arena boundary","domain_realization":"Teams may redesign the interface within common APIs but may not choose cohorts, change ground-truth tasks, manipulate defaults deceptively, contact evaluators, obstruct rivals, or acquire extra participant exposure."},{"archetype_element":"Performance metric and scoring basis","domain_realization":"Held-out evaluation combines task correctness, appropriate correction of bad suggestions, comprehension, accessibility, and error recovery, with speed treated as secondary and specified harm thresholds treated as noncompensable."},{"archetype_element":"Fair process and due process layer","domain_realization":"Rules are frozen before entry, builds are scored by a masked panel, evidence and conflicts are documented, and teams receive a time-boxed correction and appeal route."},{"archetype_element":"Anti-sabotage and anti-collusion guardrail","domain_realization":"Shared audit logs record participant allocation, build changes, evaluator contacts, and cross-round scoring patterns; anomalies trigger an independent inquiry, not an automatic penalty."},{"archetype_element":"Externality and spillover boundary","domain_realization":"Participant burden, privacy exposure, accessibility exclusion, severe downstream review errors, and future maintenance load are counted as constraints on a submission rather than exported outside its score."},{"archetype_element":"Escalation and arms-race damper","domain_realization":"Each team receives the same participant quota, design-time ceiling, compute allowance, sandbox interfaces, and submission count."},{"archetype_element":"Winner power and lock-in review","domain_realization":"The winning team does not own common telemetry or integration interfaces, deployment is reversible, and a qualified challenger can contest the slot after the pilot."},{"archetype_element":"Learning and recalibration loop","domain_realization":"The post-pilot review compares realized reviewer behavior and spillovers with the trial score, then revises or retires the next round's metrics, eligibility, and boundaries."}],"mechanism_mapping":[{"mechanism_slug":"contest_rulebook","role":"Freezes eligibility, legal design moves, scoring, tie-breaks, evidence requirements, conflicts, and appeals before teams know the outcome.","counterfactual_removal":"Without the rulebook, organizers or teams could change cohorts, criteria, or permissible interface behavior after seeing interim standings, making the result neither comparable nor contestable."},{"mechanism_slug":"ranked_leaderboard_with_audit","role":"Produces a comparable ranking from common held-out tasks while requiring reproducible builds and rechecking leading submissions before selection.","counterfactual_removal":"Without targeted audit, the ranking could reward favorable logging, task leakage, or irreproducible demonstrations rather than observed interaction performance."},{"mechanism_slug":"spending_cap_or_resource_cap","role":"Caps participant exposure, design time, compute, and submission count so teams compete through interface choices rather than research-resource escalation.","counterfactual_removal":"Without the cap, a better-resourced incumbent could buy more iterations or consume more participants, and escalating effort could burden users without improving the comparison's meaning."},{"mechanism_slug":"sabotage_or_foul_penalty_schedule","role":"Publishes graduated responses to cohort manipulation, deceptive defaults, evaluator contact, task leakage, interference, and repeated violations.","counterfactual_removal":"Without defined fouls and proportionate consequences, off-arena manipulation could remain a rational route to the deployment slot or enforcement could become selective and improvised."},{"mechanism_slug":"anti_collusion_monitoring","role":"Applies the same cross-round screens to all teams for suspicious participant swapping, coordinated score patterns, or reciprocal noncompetition, and refers anomalies to independent inquiry.","counterfactual_removal":"Without cross-round monitoring, teams could make the field appear competitive while informally dividing task families or avoiding strong head-to-head comparison; a single-round review would not reveal the pattern."},{"mechanism_slug":"challenger_access_window","role":"Reopens the default slot after the reversible pilot to qualified teams using common interfaces and a predeclared unseat rule.","counterfactual_removal":"Without a credible reopening, the first winner's integration and data advantages could turn a temporary selection into durable control of the arena."},{"mechanism_slug":"post_contest_impact_review","role":"Checks whether trial rank predicted realized review quality, user harms, maintenance burden, and winner lock-in, then changes or retires the arena.","counterfactual_removal":"Without ex-post review, a gamed or incomplete scoring basis could be repeated even when production observations show that winning did not serve the stated purpose."}],"causal_chain":["One default deployment slot makes interface teams strategically dependent rivals rather than independent improvers.","A frozen common arena removes team control over cohorts, tasks, participant volume, and judging conditions.","Held-out multidimensional scoring with noncompensable harm thresholds makes decision support, accessibility, and recoverability more reliable routes to winning than acceptance-click optimization.","Resource caps reduce the advantage from escalating experimentation and limit participant burden.","Audit, defined fouls, and appeals make manipulation detectable and enforcement reviewable without treating anomalies as verdicts.","A reversible award and challenger window prevent the initial winner from converting integration access into permanent ownership of comparison infrastructure.","Post-pilot comparison of trial scores with realized outcomes reveals whether the arena kept winning coupled to the HCI objective and supplies grounds to revise or retire it."],"baseline":"Teams independently run A/B tests or usability studies, report their preferred engagement and speed measures, and seek executive approval for the default slot. The product owner reconciles noncomparable evidence through discretionary design review; participant use, iteration resources, appeals, interference rules, and future challenger access are not governed as one shared contest.","nearest_rivals":["Centralized expert design review: can assess interaction quality and accessibility but replaces rivalry with panel judgment and does not govern strategic behavior among teams pursuing the same slot.","Standard A/B experimentation governance: can require sample plans and metric definitions but usually treats experiments as independent and does not jointly bound rival access to cohorts, resources, or future integration control.","Multi-armed-bandit allocation: can shift traffic toward better-observed variants but relies on a reward signal and does not by itself supply eligibility, noncompensable harm boundaries, appeals, anti-interference rules, or lock-in review.","Participatory co-design: can center reviewer needs and surface harms but is primarily collaborative; it does not specify how rival teams may legitimately compete for a scarce default deployment."],"remaining_contrastive_claim":"Even if ordinary A/B governance adds held-out tasks, accessibility thresholds, and expert review, the remaining distinction is governance of the teams' strategic dependence around one scarce default slot: the same frozen arena bounds entry, actions, participant and development resources, interference, due process, and the winner's future control. If teams do not affect one another's chances or no scarce deployment advantage exists, this distinction collapses into standard experimentation governance.","authority_safety":{"decision_authority":"The workflow product owner may authorize the reversible pilot only after approval from the UX research governance lead, accessibility reviewer, privacy reviewer, and an evaluation chair who is independent of competing teams. Employment, compensation, and disciplinary decisions remain outside the panel's authority.","authorized_first_step":"The evaluation chair may run a non-production shadow exercise using three already-approved prototype builds, a fixed synthetic task set, and consented internal usability participants under existing research limits; the exercise may compare rankings but may not designate a production winner.","excluded_actions":["Sending live customer or case data to prototype interfaces","Changing the production default or routing production traffic","Using contest rank in employee evaluation, compensation, or discipline","Collecting participant data beyond existing consent and retention rules","Treating an anti-collusion anomaly as proof or imposing a penalty without inquiry and appeal","Giving the winning team exclusive ownership of shared telemetry, evaluation tasks, or integration APIs","Waiving accessibility, privacy, or severe-error thresholds to improve an aggregate score"],"halt_rollback":"Halt the exercise if consent scope is exceeded, protected data appears, an evaluator conflict is undisclosed, task leakage is detected, accessibility or severe-error thresholds are breached, or participant burden exceeds the approved quota. Preserve audit records, suspend scoring, withdraw affected sessions, restore the prior sandbox configuration, notify reviewers and teams, and resume only after independent review approves a corrected protocol."},"negative_tests":{"strongest_counterevidence":"Comparable held-out evaluation shows that cohort choice, resource escalation, cross-team interference, and incumbent control do not change rankings or access, while the same acceptance-bias and accessibility failures occur when a single noncompeting team designs the interface. That would indicate an HCI measurement or design-quality problem rather than a rivalry-governance problem.","problem_falsifier":"There is no genuinely scarce default slot or contested downstream advantage, teams cannot strategically affect one another's outcomes, or parallel interfaces can be evaluated and deployed independently without shared participant, attention, integration, or data constraints.","intervention_falsifier":"Under the bounded shadow exercise, arena rank is no more consistent with blinded task-quality and harm review than the baseline ranking, prohibited strategies remain feasible or undetectable, resource caps mainly preserve incumbent assets, or independent reviewers cannot apply the rules reproducibly.","risks":["The composite score becomes a new gaming target and conceals disagreements among correctness, comprehension, accessibility, and recovery.","Uniform resource caps favor an incumbent that already owns reusable components, data, or integration knowledge.","Contest process and appeals consume more research and governance effort than the decision warrants.","Masked evaluation may fail because interface styles or technical dependencies reveal team identity.","Monitoring can create privacy or workplace-surveillance concerns if it records activity beyond contest integrity needs.","A public ranking can damage collaboration or be repurposed for personnel decisions despite the exclusion.","A challenger cadence set too frequently can discourage maintenance investment; one set too rarely can make reopening ceremonial.","Synthetic or internal-user tasks may not reproduce production context, stakes, assistive-technology use, or long-term automation bias."]},"next_evidence_step":"Run one preregistered, four-week non-production shadow comparison involving three eligible prototypes, one document-review task family, a fixed participant quota, synthetic documents, and consented internal participants. Freeze the rulebook before submissions; retain hidden task variants; have blinded reviewers independently score correctness, unsuitable-suggestion correction, comprehension, accessibility, recovery, and time; audit builds and resource use; red-team the listed gaming routes; and compare the resulting rank with the rank that acceptance rate and completion time alone would have produced. Record disagreements, participant burden, rule ambiguities, cap circumvention, and appeal usability. Do not deploy a winner or infer production effects from this step.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"This is proposal 1, version 0; there is no earlier proposal in this cell against which to assess within-cell diversity.","revision_record":{"parent_version":null,"progress_targets_addressed":[],"conceptual_changes":[],"operational_changes":[],"evidence_changes":[],"claim_changes":[]}}