{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"bounded_rivalry_governance__human_computer_interaction:P4:v0","cell_id":"bounded_rivalry_governance__human_computer_interaction","search_queries":["accessibility bug bounty program digital accessibility crowdtesting external testers","site:w3.org accessibility evaluation automated testing cannot identify all issues","bug bounty duplicate reports race first report triage research","government accessibility user testing disabled participants digital services official guidance","\"accessibility bug bounty\"","crowdsourced accessibility testing platform disabled testers first party","accessibility testing challenge competition bounty reports rewards","digital accessibility crowdsourcing research accessibility bugs","site:hackerone.com duplicate reports first reporter bounty program policy safe harbor official","site:bugcrowd.com accessibility testing crowdsourced official accessibility","bug bounty competition duplicates researchers race empirical study primary research","site:bls.gov web digital interface designers median pay 2025","site:cisa.gov vulnerability disclosure policy authorization safe harbor testing systems official","site:justice.gov framework vulnerability disclosure program authorization safe harbor official"],"sources":[{"source_id":"S1","title":"Guidance on Web Accessibility and the ADA","publisher":"U.S. Department of Justice, Civil Rights Division","url":"https://www.ada.gov/resources/web-guidance/","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2022-03-18","accessed_at":"2026-08-03","claims_supported":["Inaccessible digital services can deny people with disabilities access to essential public and commercial services.","Common barriers include inaccessible forms, missing alternatives, missing captions, and mouse-only navigation.","Public entities and businesses have identifiable accessibility responsibilities and enforcement exposure."]},{"source_id":"S2","title":"Selecting Web Accessibility Evaluation Tools","publisher":"W3C Web Accessibility Initiative","url":"https://www.w3.org/WAI/test-evaluate/tools/selecting/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2024-05-13","accessed_at":"2026-08-03","claims_supported":["Automated tools cannot check every accessibility aspect, may return misleading results, and require human judgment.","Manual evaluation and simulation of real user experience remain necessary complements to automated scanning."]},{"source_id":"S3","title":"Tips for Usability Testing with People with Disabilities","publisher":"Section508.gov, U.S. General Services Administration","url":"https://www.section508.gov/test/usability-testing-with-people-with-disabilities/","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2023-09","accessed_at":"2026-08-03","claims_supported":["Federal guidance recommends usability testing with representative people with diverse disabilities and assistive technologies.","Testing requires recruitment, task scenarios, observation, accommodation, careful interpretation, and integration with standards-based evaluation.","Product owners and federal accessibility programs are identifiable adopters of structured accessibility testing."]},{"source_id":"S4","title":"WCAG Evaluation Methodology (WCAG-EM) 2.0","publisher":"World Wide Web Consortium","url":"https://www.w3.org/TR/wcag-em-2/","source_class":"STANDARD","publication_date":"2026-07-23","accessed_at":"2026-08-03","claims_supported":["Accessibility evaluation already has an established methodology for defining scope, selecting representative samples and complete processes, evaluating, and reporting findings.","Combined evaluator expertise and involvement of people with disabilities can reveal barriers not found through expert evaluation alone.","Sample-based evaluation can leave unidentified errors and generally cannot support a whole-product conformance claim by itself."]},{"source_id":"S5","title":"Detailed Platform Standards","publisher":"HackerOne","url":"https://docs.hackerone.com/en/articles/8369826-detailed-platform-standards","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-01-20","accessed_at":"2026-08-03","claims_supported":["A mature bounty platform already distinguishes unique remedial value from duplicate or diminishing-value reports.","Existing bounty practice considers report sequence, impact, comprehensiveness, systemic scope, discretionary consolidation, and equitable distribution among contributors.","Prompt disclosure and coordinated handling are necessary when findings may expose serious risks."]},{"source_id":"S6","title":"Fable: Accessibility Research, Powered by People with Disabilities","publisher":"Fable Tech Labs","url":"https://makeitfable.com/","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["A commercial accessibility-research platform already connects product teams with trained and compensated testers with disabilities.","Fable reports more than 26,000 completed sessions and more than 40 assistive-technology or accommodation configurations.","Paid, disability-led testing across the product lifecycle is established practice, but the page does not describe scarce competitive bounties or batch portfolio selection."]},{"source_id":"S7","title":"Crowdsourced Security Vulnerability Discovery: Modeling and Organizing Bug-Bounty Programs","publisher":"Math-HCOMP Workshop; Pennsylvania State University and University of California, Berkeley authors","url":"https://chienjuho.com/workshops/mathematical-foundations-of-human-computation/papers/Math-HCOMP16-ZLMG.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2016","accessed_at":"2026-08-03","claims_supported":["Bug-bounty participants partially compete, duplicate discoveries consume processing resources, and only the first duplicate was commonly rewarded in the studied model.","The paper reports an industry estimate that 30–40% of submissions were duplicates and models report-processing cost.","Its preliminary model predicts that adding participants can eventually reduce organizer and participant utility, motivating control and diversification of the field."]},{"source_id":"S8","title":"Vulnerability Disclosure Policy","publisher":"U.S. Department of Justice, Office of the Chief Information Officer","url":"https://www.justice.gov/jmd/vulnerability-disclosure-policy","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2024-04-03","accessed_at":"2026-08-03","claims_supported":["External testing authorization depends on compliance with explicit scope and conduct rules.","Official practice prohibits privacy violations, service disruption, destructive actions, unnecessary exploitation, and high-volume low-quality reports.","Reports should include impact, environment or configuration, reproduction steps, proof, and remediation information, with immediate stopping and notification when sensitive data appears."]}],"problem_evidence":{"support":"MODERATE","rationale":"Accessibility barriers visibly matter, automated checks are incomplete, and official methodology anticipates undiscovered barriers and human evaluation. Competitive-bounty research directly documents duplicate effort, processing cost, and first-reporter incentives in security programs. However, no direct source established the prevalence of flooding, fragmentation, withholding, or first-to-file distortion in accessibility-specific bounty programs; that transfer remains an analogy requiring measurement.","source_ids":["S1","S2","S3","S4","S5","S7","S8"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Federal product owners, Section 508 program managers, accessibility evaluators, and commercial product teams are identifiable adopters of structured accessibility testing. Section508.gov expressly recommends testing with people with disabilities, and Fable demonstrates organizational purchasing of compensated external accessibility research. No source expressed demand for this proposal's competitive, sealed-batch portfolio mechanism specifically, and no named organization has committed to authorize or fund a pilot.","source_ids":["S1","S3","S4","S6"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Fable accessibility research platform","similarity":"Uses trained, compensated people with disabilities to identify barriers across assistive-technology configurations and product-development stages.","remaining_difference":"The public description is commissioned accessibility research, not a scarce-prize contest with sealed batches, resource caps, portfolio awards, affiliate aggregation, cross-round abuse screens, or appeals.","source_ids":["S6"]},{"name":"HackerOne bounty standards for systemic and duplicate reports","similarity":"Governs competitive external discovery, duplicates, systemic findings, report sequence, comprehensiveness, remediation value, and contributor compensation.","remaining_difference":"It addresses security vulnerabilities and still gives sequence an explicit role; it does not describe accessibility-task coverage, equal testing ceilings, sealed batch submission, or optimization of an award portfolio for complementary interaction modes.","source_ids":["S5"]},{"name":"WCAG-EM 2.0","similarity":"Provides established scope, representative sampling, complete-process evaluation, expertise, user involvement, documentation, and repeat-evaluation practices.","remaining_difference":"It is an evaluation methodology rather than a rivalry or reward-allocation mechanism and does not govern strategic behavior among competing external reporters.","source_ids":["S4"]},{"name":"Bug-bounty competition-control research","similarity":"Models scarce rewards, duplicate competitive effort, report-processing cost, participant diversity, and the need to control competition.","remaining_difference":"The work is security-focused and primarily analytical; it does not test sealed accessibility batches or complementarity-based portfolio scoring.","source_ids":["S7"]}],"distinctive_claim_remaining":"Holding testing scope, participant expertise, safety rules, and total reward resources constant, sealed batch submission plus complementarity-based portfolio awards and equal resource ceilings will select more distinct, independently reproducible, repair-usable accessibility barriers per reviewer hour than either first-valid-report allocation or a noncompetitive paid panel, without increasing unsafe actions, tester burden, exclusion of assistive-technology-intensive work, or delayed escalation of urgent findings.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"The constituent workflow is implementable with established practices: scoped evaluation, representative tasks, human accessibility expertise, compensated external testers, detailed reproducibility records, bounty duplicate rules, authorization boundaries, stop conditions, and validation. A synthetic-data mock avoids production authority and most privacy risk. Unverified elements are scoring reliability, complementarity optimization, identity/affiliate detection, enforceable equality of resources, accessible submission tooling, urgent-disclosure bypasses, and whether competition improves coverage rather than merely adding administration.","source_ids":["S2","S3","S4","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"If effective, the intervention could redirect scarce review and reward resources from duplicates toward task-blocking, repair-usable barriers affecting access to important services. Realized impact is unmeasured.","source_ids":["S1","S3","S4"]},"stakeholder_pull":{"score":3,"rationale":"There is clear institutional demand for accessibility evaluation and demonstrated purchasing of external disabled-user testing, but no verified demand for a competitive portfolio challenge.","source_ids":["S3","S4","S6"]},"incremental_advantage":{"score":3,"rationale":"Batch sealing and portfolio-level complementarity plausibly improve on sequence-weighted bounty allocation, but comparative performance against a well-run paid panel or ordinary audit is unknown.","source_ids":["S4","S5","S7"]},"distinctiveness_plausibility":{"score":3,"rationale":"No opened source described the full accessibility-specific combination. Most components are established separately, so distinctiveness rests on their joint allocation rule and measurable outcome rather than on novel components.","source_ids":["S4","S5","S6","S7"]},"technical_implementability":{"score":4,"rationale":"An isolated prototype, scoped accounts, sealed submissions, logs, blinded scoring, reproduction checks, and portfolio selection can be built using conventional testing and workflow infrastructure. Reliable identity and collusion inference should remain investigatory rather than dispositive.","source_ids":["S4","S5","S8"]},"adoption_authority_feasibility":{"score":3,"rationale":"A product owner and accessibility-program owner can commission a mock round, while security, privacy, and environment owners can bound authorization. A paid external program would still require organization-specific legal, procurement, tax, labor, privacy, and disclosure review.","source_ids":["S3","S4","S8"]},"evidence_readiness":{"score":4,"rationale":"A preregistered synthetic-environment experiment can directly compare allocation rules, measure coverage and reviewer burden, and red-team the governance without touching production.","source_ids":["S4","S5","S8"]},"safety_net_benefit":{"score":4,"rationale":"Synthetic data, scoped accounts, explicit prohibited actions, immediate stop conditions, and independent verification substantially limit downside and provide rollback. They do not eliminate participant exploitation, sensitive evidence leakage, or urgent-report delay risks.","source_ids":["S3","S8"]},"scalability":{"score":3,"rationale":"External tester communities and repeatable platforms exist, but manual reproduction, appeals, complementarity scoring, accommodations, and abuse investigations may grow faster than report volume and require skilled reviewers.","source_ids":["S3","S4","S6","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"One preregistered two-round synthetic mock with approximately 12 compensated testers, an accessibility lead, independent scorers, a security/privacy review, prototype fixtures, data analysis, and a short findings report.","confidence":"LOW","assumptions":["The prototype already exists or can be instrumented without major product engineering.","Participants receive guaranteed participation compensation independent of competitive points.","Internal loaded labor plus participant compensation dominates cost; no production integration is included.","No opened source published directly comparable prices."],"source_ids":["S3","S4","S6","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Design and assurance for a reusable challenge service: accessible intake, identity and affiliate declarations, synthetic environment, immutable logs, scoring and appeal workflow, policy drafting, privacy/security review, and staff training.","confidence":"LOW","assumptions":["Existing identity, ticketing, storage, and test-environment systems can be adapted.","External counsel or procurement work is bounded rather than organization-wide.","Custom anti-collusion modeling is limited to referral screens, not automated adjudication."],"source_ids":["S4","S5","S6","S8"]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"First paid production-like but nonproduction round, including reward pool, tester guarantees, recruitment and accommodation, verification panel, appeal reviewer, monitoring, remediation handoff, and post-round review.","confidence":"LOW","assumptions":["The first launch remains on a segregated pre-release environment with synthetic data.","A modest fixed reward pool and fewer than roughly 25 entrants are used.","Engineering remediation itself is excluded because it is required under every comparator."],"source_ids":["S3","S4","S5","S6","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Several challenge rounds per year across multiple services, with program management, tester compensation and awards, environments, verification, appeals, security/privacy operations, maintenance, and evaluation.","confidence":"LOW","assumptions":["Three to eight bounded rounds are conducted annually.","Manual accessibility expertise and participant compensation remain material recurring inputs.","Costs exclude broad accessibility remediation and enterprise-wide platform replacement.","No direct market quote was verified, so the band is a resource-equivalent planning estimate."],"source_ids":["S3","S4","S5","S6","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official and standards sources establish consequential digital barriers and limits of automation; bounty research and platform rules establish duplicate and low-quality-report problems. Accessibility-bounty-specific prevalence remains an evidence gap but does not erase the supported underlying problem.","source_ids":["S1","S2","S4","S5","S7","S8"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Product owners, accessibility program managers, Section 508 officials, security/privacy leads, and test-environment owners are identifiable authorizers; government guidance and commercial practice show these organizations conduct or procure external accessibility testing. No adoption commitment is verified.","source_ids":["S3","S4","S6","S8"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal specifies a contrastive comparison against first-valid-report allocation and a noncompetitive paid panel, with measurable distinctness, reproducibility, coverage, repair utility, reviewer effort, burden, and safety outcomes.","source_ids":["S4","S5","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A small preregistered experiment in a synthetic environment can compare mechanisms and exercise stop, appeal, and abuse procedures without production access.","source_ids":["S3","S4","S8"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"For the first evidence step, product, accessibility, security, privacy, and environment owners can authorize an isolated mock; synthetic data, guaranteed compensation, scoped accounts, stop conditions, and no production access bound material risks. A paid operational launch would require additional organization-specific review.","source_ids":["S3","S8"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"Scopes and labor-bearing assumptions are explicit, but no direct vendor quote, tester-compensation schedule, internal loaded-rate evidence, or comparable program budget was verified. The ranges are resource-equivalent planning estimates, not validated budgets.","source_ids":["S3","S4","S5","S6","S8"]}},"next_evidence_step":"Run a preregistered two-round crossover mock on one isolated web-service prototype with synthetic accounts, four essential user journeys, hidden seeded barriers, unseeded states, and approximately 12 compensated qualified testers stratified across assistive-technology and input-mode expertise. Randomize teams initially to (A) the proposed sealed competitive portfolio condition or (B) a noncompetitive flat-fee commissioned-panel condition, then switch conditions on a reset fixture; independently rescore the combined reports under (C) first-valid-report allocation. Guarantee equal base compensation in all conditions and use fictitious competitive points, not contingent cash. Blind initial scorers to condition and seed status. Measure distinct independently reproduced barriers, hidden-seed recall, journey/mode/failure-class coverage, remediation-handoff ratings by maintainers, scorer agreement, reviewer minutes, duplicate and fragmentation rates, tester burden, cap exceptions, appeals, urgent-disclosure latency, and safety events. Red-team identity splitting, automated flooding, reciprocal attribution, fixture alteration, fabricated traces, and rubric-targeted fragmentation. Advance only if the portfolio condition improves distinct reproducible repair-usable findings per reviewer hour by a preregistered material margin, such as 20%, over both comparators while preserving inter-rater reliability and producing no excess safety events, urgent-report delay, or disproportionate loss of manual and assistive-technology-intensive findings. Falsify the incremental claim if the flat-fee panel matches or exceeds it, first-valid allocation performs equivalently, caps suppress valuable manual work, complementarity scoring is unreliable, or competition adds burden or unsafe behavior without coverage gain.","blocking_evidence":["No direct prevalence estimate for strategic flooding, fragmentation, withholding, or first-to-file distortion in accessibility-specific bounty programs.","No field comparison showing that competitive accessibility discovery adds coverage beyond a well-run compensated panel or commissioned WCAG evaluation.","No evidence that the proposed complementarity rubric is reliable across scorers or predicts remediation utility.","No validated method for affiliate aggregation or collusion screening that avoids unfairly merging independent testers or treating anomaly signals as guilt.","No evidence that equal request, time, and submission caps avoid disadvantaging assistive-technology-intensive or manual investigation.","No organization has committed adoption authority, staff, environment access, or funding.","No direct pricing or internal labor data validates the four cost bands.","Legal, procurement, tax, worker-classification, accessibility-accommodation, privacy, and disclosure requirements remain organization- and jurisdiction-specific.","World novelty, patentability, freedom to operate, market size, and realized impact are unmeasured."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The bounded search found established accessibility-evaluation methodologies, compensated external disabled-user testing, sequence- and value-aware bounty rules, and research on controlling duplicate competitive effort. It did not find an opened source describing the full combination of an accessibility barrier bounty with sealed batches, equal resource ceilings, noncompensable harm gates, complementarity-based portfolio awards, affiliate aggregation, appeals, recurring challenger access, and post-round recalibration. This is only a contrast within eight opened sources, not a world-novelty, patentability, freedom-to-operate, or exhaustive prior-art determination.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure one product owner, accessibility lead, security/privacy approvers, and an independent scoring partner for the synthetic mock.","Preregister the three comparators, randomization or crossover design, seed register, primary metric, material-effect threshold, safety outcomes, and falsifiers.","Create and validate an accessible submission format, guaranteed participant-compensation plan, urgent-disclosure bypass, stop procedure, and appeal protocol.","Measure scorer agreement and maintainer-rated remediation utility before interpreting portfolio scores as meaningful.","Audit whether caps differentially suppress manual or assistive-technology-intensive findings.","Obtain bottom-up labor, participant, platform, legal-review, and reward estimates to replace the low-confidence cost bands."],"reason":"Web evidence supports the underlying accessibility problem, credible authorizers, adjacent practices, and a safe bounded mock, but it cannot establish the proposal's remaining causal claim. Whether governed competition improves distinct repair-usable coverage over first-to-file allocation and a compensated noncompetitive panel requires live testing with participants and proprietary workflow data. Under the controller rule, that evidence need requires an empirical-research stop and repairable=false."},"proposal_index":4}