{"schema_version":1,"research_id":"eoa_inverse_innovation_exp04_external_evaluation_20260802","source_assessment_id":"negative_space_design__behavioral_economics:SENTINEL_MATCHED:v0","cell_id":"negative_space_design__behavioral_economics","search_queries":["information cascades experiment private signal observe previous decisions primary research","independent judgment before group discussion conformity accuracy primary study","platform hide community prediction until user predicts first product feature","Delphi method anonymous independent judgments aggregation official guidance","IDEA protocol Investigate Discuss Estimate Aggregate independent estimates before discussion primary paper","estimate talk estimate method independent judgment before social information research","Delphi method first round independent anonymous estimates revise after feedback RAND official","Metaculus hide community prediction until forecast first","Lorenz Rauhut Schweitzer social influence undermines wisdom of crowd PNAS 2011 full text","social influence bias estimation diversity crowd accuracy primary experiment PNAS","independent first group decision making intervention initial judgment before discussion experiment","popularity information online choice herding experiment ratings primary research","site:w3.org WAI status messages WCAG loading errors content hidden accessibility official","site:hhs.gov OHRP Common Rule minimal risk informed consent research official","site:ftc.gov dark patterns delayed disclosure material information online choice official","site:bls.gov occupational employment wage software developers web digital interface designers 2025","Europe PMC How social influence can undermine the wisdom of crowd effect Lorenz abstract","Crossref 10.1073/pnas.1008636108 abstract"],"sources":[{"source_id":"S1","title":"Information Cascades in the Laboratory","publisher":"American Economic Association; indexed by RePEc","url":"https://ideas.repec.org/a/aea/aecrev/v87y1997i5p847-62.html","source_class":"PRIMARY_RESEARCH","publication_date":"1997-12","accessed_at":"2026-08-02","claims_supported":["In incentivized sequential decisions, early public predictions can cause later decision-makers to follow the established pattern regardless of their private signals.","Rational cascades formed in most experimental periods in which the necessary imbalance of prior decisions occurred."]},{"source_id":"S2","title":"How social influence can undermine the wisdom of crowd effect","publisher":"Proceedings of the National Academy of Sciences","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC3107299/","source_class":"PRIMARY_RESEARCH","publication_date":"2011-05-16","accessed_at":"2026-08-02","claims_supported":["In an experiment with 144 participants, even mild exposure to others' estimates reduced diversity without improving collective error.","Social information could increase confidence despite no corresponding accuracy improvement.","The effect was demonstrated in estimation tasks and does not establish that social information is harmful in every decision context."]},{"source_id":"S3","title":"Investigate Discuss Estimate Aggregate for structured expert judgement","publisher":"International Journal of Forecasting; Monash University research portal","url":"https://research.monash.edu/en/publications/investigate-discuss-estimate-aggregate-for-structured-expert-judg/","source_class":"PRIMARY_RESEARCH","publication_date":"2017-01","accessed_at":"2026-08-02","claims_supported":["The IDEA protocol has participants investigate and predict before seeing and discussing others' thinking, then provide a second private judgment for aggregation.","The protocol performed well relative to an equally weighted linear pool and a prediction market, although some results were not statistically significant.","The method is reported as relatively simple to implement."]},{"source_id":"S4","title":"Predicting reliability through structured expert elicitation with the repliCATS process","publisher":"PLOS ONE","url":"https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0274429","source_class":"PRIMARY_RESEARCH","publication_date":"2023-01-26","accessed_at":"2026-08-02","claims_supported":["repliCATS operationalized an independent initial judgment and reasoning, disclosure and discussion of group judgments, and a second private judgment.","A cloud-based platform supported synchronous and asynchronous use, assessed 3,000 claims over 18 months, and enrolled hundreds of participants.","A small validation experiment reported 84% classification accuracy and AUC 0.94, but the authors caution that participants may have known some outcomes and that a more precise accuracy estimate requires further work.","The work was funded by DARPA's SCORE program, and participants consented under institutional ethics approval."]},{"source_id":"S5","title":"Metaculus FAQ","publisher":"Metaculus","url":"https://www.metaculus.com/faq/","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"n.d.","accessed_at":"2026-08-02","claims_supported":["Metaculus users can hide the Community Prediction.","Metaculus initially hides the Community Prediction on newly opened questions to avoid giving disproportionate weight to early predictions that might ground or bias later forecasts.","Metaculus restores and displays an aggregate based on recent individual forecasts, demonstrating technical and operational feasibility for delayed social information."]},{"source_id":"S6","title":"Understanding Success Criterion 4.1.3: Status Messages","publisher":"World Wide Web Consortium Web Accessibility Initiative","url":"https://www.w3.org/WAI/WCAG22/Understanding/status-messages","source_class":"STANDARD","publication_date":"2026-05-11","accessed_at":"2026-08-02","claims_supported":["State changes and exposed or hidden content must remain programmatically understandable through applicable name, role, value, and status-message mechanisms.","User testing is advised to calibrate feedback and avoid an interface that is either silent or excessively chatty for screen-reader users."]},{"source_id":"S7","title":"Informed Consent FAQs","publisher":"U.S. Department of Health and Human Services, Office for Human Research Protections","url":"https://www.hhs.gov/ohrp/regulations-and-policy/guidance/faq/informed-consent/index.html","source_class":"OFFICIAL_GUIDANCE","publication_date":"n.d.","accessed_at":"2026-08-02","claims_supported":["Covered human-subject research generally requires prospective, legally effective informed consent unless an IRB approves an applicable waiver or alteration.","Consent should disclose procedures, foreseeable risks, confidentiality, voluntariness, and withdrawal rights.","An IRB may approve altered consent for qualifying minimal-risk research, but the relevant findings must be documented."]},{"source_id":"S8","title":"National employment and wage data by occupation, May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.t01.htm","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026-05","accessed_at":"2026-08-02","claims_supported":["May 2025 mean annual wages were approximately $148,100 for software developers and $117,490 for web and digital interface designers.","These wage benchmarks support order-of-magnitude labor estimates but do not include benefits, overhead, participant incentives, legal review, or vendor margin."]}],"problem_evidence":{"support":"STRONG","rationale":"The diagnosed mechanism is visible in controlled evidence: sequential public decisions can override private signals, and exposure to others' estimates can reduce diversity without improving collective error. The consequence matters where independent signals are valuable for aggregation. External validity is bounded: social information can be genuinely diagnostic, and these studies do not show that every popularity cue or recommendation degrades decisions.","source_ids":["S1","S2"]},"stakeholder_evidence":{"support":"STRONG","rationale":"Metaculus is an identifiable platform operator that already hides community predictions to avoid grounding or bias, directly expressing the need and demonstrating authorization. DARPA funded a large repliCATS implementation of a closely matching staged-elicitation workflow. These establish adopter and funder credibility, though no party has committed to testing this exact three-arm ordinary-choice variant.","source_ids":["S4","S5"]},"prior_art":{"proximity":"ESTABLISHED_PRACTICE","closest_analogues":[{"name":"IDEA protocol","similarity":"Participants form and record an initial judgment before exposure to others' judgments and reasoning, then make a second private judgment for aggregation.","remaining_difference":"IDEA is structured expert elicitation with discussion and mathematical aggregation, not a general-purpose per-user interface intervention for ordinary sequential choices with a debiasing-prompt comparator.","source_ids":["S3","S4"]},{"name":"repliCATS online elicitation platform","similarity":"A deployed cloud workflow collected initial quantitative judgments and reasons, revealed group judgments and reasoning, and collected revised private judgments.","remaining_difference":"Its validated application was forecasting research replicability among structured groups; it did not test popularity counts, rankings, or recommendations across ordinary consumer or platform choices.","source_ids":["S4"]},{"name":"Metaculus hidden Community Prediction","similarity":"A live prediction platform hides aggregate social information to prevent early grounding or bias and later makes the aggregate available.","remaining_difference":"The documented feature is a platform-wide initial hiding period or user setting, not necessarily a mandatory per-choice provisional judgment and rationale followed by automatic reveal and explicit revision measurement.","source_ids":["S5"]}],"distinctive_claim_remaining":"In a specified low-stakes sequential-choice setting, requiring each chooser to record a provisional choice and brief rationale before complete, clearly framed restoration of nonmaterial social cues will increase private-signal sensitivity and preserve independent information relative to both simultaneous display and simultaneous display plus a debiasing prompt, while being noninferior on post-reveal accuracy, comprehension, accessibility, and chooser trust. This is contrastive and falsifiable, but it is a context-specific incremental-effect claim rather than a novel mechanism.","confidence":"HIGH"},"implementation_evidence":{"support":"STRONG","rationale":"Both Metaculus and repliCATS demonstrate that social information can be withheld, later displayed, and combined with revisable judgments in web platforms. The required state machine, event logging, rationale capture, reveal, revision, and rollback are routine web functionality. Feasibility remains conditional on accessible state announcements, reliable restoration, research review, privacy controls for rationale text, and exclusion of prices, risks, eligibility, conflicts, required disclosures, and other material information.","source_ids":["S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Loss of independent information can create wrong cascades and unjustified confidence, although harm magnitude is setting-dependent.","source_ids":["S1","S2"]},"stakeholder_pull":{"score":4,"rationale":"A live platform explicitly reports hiding aggregates to avoid grounding, and a government-funded program deployed a close staged-judgment workflow.","source_ids":["S4","S5"]},"incremental_advantage":{"score":2,"rationale":"The proposed head-to-head comparison against simultaneous display and a prompt rival is useful, but no direct evidence yet shows that mandatory provisional rationale adds benefit beyond established staged elicitation or optional hiding.","source_ids":["S3","S4","S5"]},"distinctiveness_plausibility":{"score":1,"rationale":"The core sequence substantially overlaps IDEA, repliCATS, and Metaculus. Only a narrower cross-context implementation and comparator-specific effect claim remains.","source_ids":["S3","S4","S5"]},"technical_implementability":{"score":5,"rationale":"Close workflows have already been implemented at scale; remaining work is conventional interface, logging, accessibility, and experiment infrastructure.","source_ids":["S4","S5","S6"]},"adoption_authority_feasibility":{"score":4,"rationale":"A platform owner can authorize a reversible low-stakes experiment, subject to research, privacy, accessibility, and applicable legal review. Consent or an approved waiver is needed where human-subject rules apply.","source_ids":["S5","S7"]},"evidence_readiness":{"score":4,"rationale":"The mechanism, comparators, outcomes, and falsifiers are measurable in a preregistered simulation, but the incremental claim requires live participant evidence.","source_ids":["S1","S2","S4"]},"safety_net_benefit":{"score":3,"rationale":"Explicit restoration, reversible provisional choices, rollback, comprehension checks, and accessible state signaling bound risk, but novices may lose useful social guidance and rationale collection introduces privacy burden.","source_ids":["S6","S7"]},"scalability":{"score":4,"rationale":"Cloud implementations have supported hundreds of participants and thousands of assessments, though added friction and context-specific cue classification limit universal deployment.","source_ids":["S4","S5"]}},"score_confidence":"HIGH","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Preregistered low-stakes online simulation with three interface arms, power analysis, approximately 600-900 adult participants, participant compensation, instrument development, accessibility checks, analysis, and research-review preparation.","confidence":"MODERATE","assumptions":["An existing survey or experiment platform can host the prototype.","One research lead, part-time engineering and design support, and limited statistical support are sufficient.","No production-system integration or proprietary-data purchase is included.","Loaded labor is higher than published wages because benefits and overhead are added."],"source_ids":["S4","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"One reversible production prototype on an existing low-stakes choice platform, including event instrumentation, rationale storage, reveal and rollback logic, privacy review, and accessibility testing.","confidence":"MODERATE","assumptions":["Existing authentication, experimentation, analytics, and feature-flag infrastructure are available.","Roughly two to six person-months of engineering, design, research, QA, and governance effort are required.","No regulated or high-stakes decision surface is included."],"source_ids":["S4","S5","S6","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Multi-surface launch with production hardening, subgroup and accessibility validation, monitoring dashboards, support procedures, security and privacy review, legal review, localization, and rollback operations.","confidence":"LOW","assumptions":["Launch spans several choice contexts but remains outside medical, financial, employment, benefits, and emergency decisions.","The operator already has platform and experimentation teams.","The range excludes major platform rearchitecture and litigation or regulatory-response costs."],"source_ids":["S4","S6","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Ongoing experiment monitoring, accessibility regression testing, privacy retention controls, model and metric review, support, incident response, and periodic revalidation of cue classifications.","confidence":"LOW","assumptions":["Approximately 0.5-1.5 combined full-time-equivalent staff plus infrastructure and periodic participant testing are needed.","Rationale text is retained only as long as necessary and does not require intensive moderation.","No continuous regulated-domain compliance program is included."],"source_ids":["S6","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Primary experiments directly show private-signal suppression, cascade formation, reduced diversity, or misplaced confidence after exposure to others' choices or estimates.","source_ids":["S1","S2"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Metaculus is an identifiable operator already using delayed aggregate visibility to avoid grounding, while DARPA funded a deployed close analogue through SCORE/repliCATS.","source_ids":["S4","S5"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim specifies two comparators, measurable private-signal sensitivity and aggregation outcomes, and noninferiority guardrails; it is narrower than the established mechanism.","source_ids":["S3","S4","S5"]},"bounded_next_evidence_step":{"status":"YES","reason":"A preregistered, consented, low-stakes three-arm simulation can test the claim without production deployment or irreversible choices.","source_ids":["S1","S4","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The bounded test is feasible if an appropriate research authority reviews it, participation is voluntary, all delayed cues are restored, choices remain reversible, accessibility is tested, and material or safety-critical information is never delayed. These conditions do not authorize a high-stakes or commercial launch.","source_ids":["S6","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The ranges are explicitly scoped and bottom-up assumptions are anchored to demonstrated online implementations and current occupational wage data, although no vendor bids or target-platform estimates were obtained.","source_ids":["S4","S5","S8"]}},"next_evidence_step":"Preregister a low-stakes simulated sequential-choice experiment with a short usability pilot followed by a power-calculated confirmatory sample capped at 900 adults. Randomize participants among: (A) simultaneous attributes plus social cues, (B) the same screen plus a debiasing prompt, and (C) attributes first, mandatory provisional choice and brief rationale, then automatic full restoration of social cues and an explicit revision opportunity. Generate private-signal quality and public-cue correctness experimentally. Primary tests are private-signal sensitivity, initial-choice correlation/diversity, and aggregate error. Guardrails are post-reveal final accuracy, cue comprehension, completion time, perceived manipulation, novice and accessibility subgroup outcomes, reliable cue restoration, and screen-reader task success. Falsify the intervention if arm C does not outperform both comparators on preregistered private-information outcomes, or if any gain is offset by worse final accuracy, comprehension, accessibility, trust, or disproportionate novice harm. Stop immediately on restoration failures or participant inability to distinguish deliberate delay from missing or failed data.","blocking_evidence":["No direct three-arm test establishes incremental advantage over a debiasing prompt.","External validity from forecasting and estimation protocols to ordinary digital choices is unknown.","The contexts in which social information is beneficial, especially for novices or weak private signals, are not yet delimited.","No target platform has supplied proprietary logs, integration estimates, or a commitment to this exact experiment.","Jurisdiction-specific privacy, consumer-protection, and disclosure review would be required before any live commercial deployment.","Cost ranges lack vendor bids and target-platform engineering estimates."],"research_disposition":"KNOWN_PRACTICE_DIFFUSION","world_novelty_boundary":"World novelty, patentability, freedom to operate, market size, and realized impact were not measured. The broad mechanism is not novel on the reviewed evidence: staged independent judgment followed by social-information exposure and revision is established in IDEA/repliCATS, and delayed aggregate visibility is deployed by Metaculus. Only the context-specific mandatory-rationale workflow and its comparator-defined incremental effect remain unmeasured.","arm":"SENTINEL","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain preregistered evidence that the protected interval improves private-signal sensitivity and aggregate information retention relative to both simultaneous display and the prompt rival.","Demonstrate noninferiority on final post-reveal accuracy, comprehension, accessibility, trust, and completion burden.","Show that all delayed cues are reliably restored and that participants distinguish deliberate sequencing from error, concealment, or missing data.","Estimate heterogeneous effects for novices, low-confidence participants, weak private signals, and assistive-technology users.","Secure target-platform product, research, privacy, accessibility, and legal authorization plus a platform-specific cost estimate before production testing."],"reason":"Bounded web research establishes the problem, credible adopters, implementability, and substantial prior-art collision, but the remaining incremental claim can only be resolved with live participant testing. Because the required evidence is empirical rather than obtainable through further web search, the evaluation stops for empirical research; as a stop recommendation, it is not marked repairable."}}