{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"bounded_rivalry_governance__futurism_foresight:P4:v0","cell_id":"bounded_rivalry_governance__futurism_foresight","search_queries":["site:iarpa.gov ACE forecasting tournament geopolitical events official","Good Judgment Project superforecasters top 2 percent elite teams forecasting tournament primary research","forecasting tournaments question selection participation bias proper scoring research","Metaculus tournament forecasting official scoring rules team selection council forecasters","site:iarpa.gov research programs ACE aggregative contingent estimation forecasting official","site:iarpa.gov forecasting tournament ACE official 2011 2015","site:goodjudgment.com superforecasters selection top 2% teams official","site:gov.uk forecasting tournament policy superforecasting government official","site:contractsfinder.service.gov.uk superforecasting DEFRA ForgeFront Good Judgment contract 2024","Defra superforecasting foresight contract Good Judgment 2024 £ procurement","government superforecasting challenge contract forecasting strategic foresight cost","Mellers Psychological Strategies for Winning a Geopolitical Forecasting Tournament PMC 2014","Identifying and Cultivating Superforecasters PMC Mellers 2015","A neglected dimension of good forecasting judgment questions we choose full text 2017","forecasting tournament question difficulty participation selection primary study Brier score","site:eeoc.gov employee selection procedures validation Uniform Guidelines official","EEOC Uniform Guidelines employee selection procedures validity job performance official","42 USC employment tests selection procedures adverse impact EEOC official","\"A neglected dimension of good forecasting judgment\" PDF","\"Effects of Choice Restriction on Accuracy\" full text PDF","\"Incentive-Compatible Forecasting Competitions\" Microsoft Research PDF"],"sources":[{"source_id":"S1","title":"ACE — Aggregative Contingent Estimation","publisher":"Intelligence Advanced Research Projects Activity","url":"https://www.iarpa.gov/research-programs/ace","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"Undated; accessed 2026-08-03","accessed_at":"2026-08-03","claims_supported":["IARPA sought more accurate, precise, and timely intelligence forecasts.","The program funded empirical testing of elicitation, aggregation, performance weighting, and forecast representation against real events.","A government funder and authorizer visibly exists for operationally relevant forecasting research."]},{"source_id":"S2","title":"Identifying and cultivating superforecasters as a method of improving probabilistic predictions","publisher":"Perspectives on Psychological Science / PubMed","url":"https://pubmed.ncbi.nlm.nih.gov/25987508/","source_class":"PRIMARY_RESEARCH","publication_date":"2015-05","accessed_at":"2026-08-03","claims_supported":["The Good Judgment Project annually selected top performers into elite superforecaster teams.","Selected forecasters maintained high accuracy across hundreds of questions for two subsequent years.","Performance reflected discovery, task skills, motivation, and enriched environments, making simple attribution to innate individual skill inappropriate."]},{"source_id":"S3","title":"A neglected dimension of good forecasting judgment: The questions we choose also matter","publisher":"International Journal of Forecasting / Elsevier","url":"https://www.sciencedirect.com/science/article/pii/S0169207017300377","source_class":"PRIMARY_RESEARCH","publication_date":"2017","accessed_at":"2026-08-03","claims_supported":["Question selection contains information omitted by conventional score-only evaluation.","Forecaster ability estimates can change when question-selection and nonresponse behavior are modeled.","Good forecasters in the studied tournament tended to select more questions, so exposure is not ignorable."]},{"source_id":"S4","title":"Effects of Choice Restriction on Accuracy and User Experience in an Internet-Based Geopolitical Forecasting Task","publisher":"Frontiers in Psychology","url":"https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2021.662279/pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2021-07-01","accessed_at":"2026-08-03","claims_supported":["Forecasting tournaments face a real load-distribution and question-coverage problem.","Two online novice-forecaster studies found no significant accuracy loss from restricted question choice, including hard assignment.","The experiments were short and used novices; they do not establish effects for trained employees, compulsory balanced blocks, or longitudinal promotion."]},{"source_id":"S5","title":"Incentive-Compatible Forecasting Competitions","publisher":"Management Science / INFORMS","url":"https://pubsonline.informs.org/doi/10.1287/mnsc.2022.4410","source_class":"PRIMARY_RESEARCH","publication_date":"2022","accessed_at":"2026-08-03","claims_supported":["Selecting the highest proper-score contestant for a scarce prize need not elicit truthful reports; contestants can benefit from strategic extremization.","No deterministic forecasting competition is strictly incentive compatible under the paper's formal conditions.","Randomized ELF and I-ELF mechanisms address this incentive problem, including ranking and forecaster-selection applications.","The proposal's deterministic promotion rule therefore retains a theoretically material incentive defect despite using a proper scoring rule."]},{"source_id":"S6","title":"Scores FAQ","publisher":"Metaculus","url":"https://www.metaculus.com/help/scores-faq/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"Undated; scoring documentation current as accessed 2026-08-03","accessed_at":"2026-08-03","claims_supported":["A deployed forecasting platform uses proper log-based scores, coverage incentives, weighted questions, rankings, and tournament prizes.","Metaculus assigns zero for unforecast questions in current tournament scoring and uses hidden periods because visible community predictions can be copied.","Many proposed components—proper scoring, coverage rules, question weighting, hidden forecasts, and audited tournament administration—are established product practice."]},{"source_id":"S7","title":"Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures","publisher":"U.S. Equal Employment Opportunity Commission, jointly with DOJ, OPM, DOL, and Treasury","url":"https://www.eeoc.gov/es/node/130157","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"1979-03-02","accessed_at":"2026-08-03","claims_supported":["Where U.S. employment law applies, a selection procedure with adverse impact requires evidence connecting it to successful job performance, not merely a rational justification.","Unsupported assertions that a selection procedure is valid are insufficient.","Using league results for funded employment opportunities would require jurisdiction-specific personnel review, accommodations, adverse-impact monitoring, and potentially validation."]},{"source_id":"S8","title":"Defra Contract C25475 Conditions of Contract and Specification","publisher":"UK Department for Environment, Food & Rural Affairs / Contracts Finder","url":"https://www.contractsfinder.service.gov.uk/Notice/Attachment/b1c5ef11-66cc-4ed7-bd45-7eb6ec393b65","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2024-09-02","accessed_at":"2026-08-03","claims_supported":["Defra explicitly reported fragmented future-biothreat assessment, no mechanism for systematizing judgments, and limited capacity for anticipatory action based on forecasts and foresight.","Defra commissioned a six-month prototype integrating superforecasting and foresight, identifying a concrete adopter, project manager, and governance workflow.","Disclosed milestones included £1,500 for expert/superforecaster engagement and legal question checks, £18,500 for three superforecasting rounds, and £9,500 for evaluation and workshops, providing a partial cost analogue.","The contract excluded supplier processing of personal data, underscoring that an employee-ranking league would add privacy and personnel-governance scope absent from this analogue."]}],"problem_evidence":{"support":"STRONG","rationale":"The general problem is visible: question exposure is informative and non-ignorable, deployed tournaments hide community forecasts to reduce copying, and formal research shows that a scarce-prize ranking based on proper scores can induce strategic misreporting. However, no external source establishes the claimed prevalence of copying, collusion, resource-cap evasion, resolver influence, or incumbent entrenchment in the candidate's hypothetical forty-person office.","source_ids":["S3","S5","S6"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"IARPA is an identifiable funder seeking empirically validated forecasting improvements, while Defra is an identifiable adopter that documented a need to systematize anticipatory judgments and purchased a superforecasting/foresight prototype. Neither source expresses demand for a promotion-and-relegation personnel league or for allocating council seats by tournament rank.","source_ids":["S1","S8"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Good Judgment Project annual superforecaster selection and elite teams","similarity":"Very close on repeated outcome-resolved forecasting, annual selection of top performers, scarce elite status, team assignment, and demonstrated persistence across subsequent questions.","remaining_difference":"The candidate adds compulsory balanced blocks, sealed rival estimates, equal research credits, fixed update counts, audits, staggered internal council seats, relegation, appeals, and employment-style safeguards.","source_ids":["S1","S2","S5"]},{"name":"Metaculus scored forecasting tournaments","similarity":"Close on proper scoring, rankings, question weights, coverage incentives, hidden community forecasts, and allocation of scarce prizes.","remaining_difference":"Metaculus does not implement compulsory balanced blocks, equal research-resource accounting, employer-controlled council promotion, staggered seats, or relegation from advisory authority.","source_ids":["S6"]},{"name":"ELF and I-ELF incentive-compatible forecaster selection","similarity":"Directly addresses strategic reporting when accurate forecasters compete for a scarce prize and explicitly generalizes to rankings and hiring.","remaining_difference":"The mechanisms randomize selection and do not supply the candidate's seasonal council workflow, assignment balancing, confidentiality controls, audits, or employment governance.","source_ids":["S5"]},{"name":"Choice-limiting question triage","similarity":"Direct empirical test of restricting or assigning forecasting questions to distribute workload and coverage.","remaining_difference":"The studies used short tasks with online novices and did not test balanced strata, mandatory longitudinal completion, high-stakes rank, council selection, or out-of-sample advisory performance.","source_ids":["S4"]}],"distinctive_claim_remaining":"Against both (a) an open, self-selected leaderboard and (b) manager appointment, selecting a temporary internal council through balanced compulsory question blocks, sealed estimates, equal auditable research resources, fixed updates, rolling two-season performance, and staggered promotion/relegation will produce a council with better held-out accuracy and incremental aggregate information without unacceptable burden, confidentiality failure, score instability, or adverse impact. This combination is contrastive and falsifiable, but its individual mechanisms and the core tournament-to-elite-selection pathway are already prior art.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Proper scoring, sealed/hidden forecasts, question weighting, tournament administration, repeated selection, load assignment, and simple randomized incentive-compatible alternatives are technically feasible. Material unresolved issues are whether balanced strata and normalization are reliable with only forty forecasters; whether enough questions resolve promptly; whether equal research credits and independence can be audited; whether sealing removes useful information exchange; how similarity-screen false positives are handled; and whether promotion is a legally valid, accessible, job-related employment-selection procedure. These require local workflow and field evidence.","source_ids":["S2","S4","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Accurate, plural probabilistic advice could improve decisions, but no evidence connects this league's selected council to realized decision quality or establishes the scale of current selection error.","source_ids":["S1","S8"]},"stakeholder_pull":{"score":4,"rationale":"Government funders and foresight users have explicitly sought forecasting improvements and purchased superforecasting prototypes, though not this personnel league.","source_ids":["S1","S8"]},"incremental_advantage":{"score":2,"rationale":"Balanced assignments and council-governance safeguards add plausible value, but annual elite selection, proper-score tournaments, hidden forecasts, coverage incentives, and assignment restriction already exist.","source_ids":["S2","S4","S5","S6"]},"distinctiveness_plausibility":{"score":2,"rationale":"The integrated employer-council bundle is distinguishable, but the central causal pathway—tournament performance selecting an elite forecasting group—substantially collides with Good Judgment prior art.","source_ids":["S2","S5"]},"technical_implementability":{"score":4,"rationale":"The software and scoring components are straightforward and deployed analogues exist; question normalization, resolution latency, auditability, and integrity enforcement remain nontrivial.","source_ids":["S4","S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"A foresight director can authorize a no-stakes pilot, but live seat allocation needs personnel, ethics, privacy, accessibility, and possibly employment-validation authority.","source_ids":["S7","S8"]},"evidence_readiness":{"score":3,"rationale":"Most process measures are observable, but outcome-resolved accuracy, ranking reliability, and council utility require prospective data and potentially long resolution horizons.","source_ids":["S1","S3","S4"]},"safety_net_benefit":{"score":4,"rationale":"A no-stakes, non-sensitive microseason with synthetic identifiers, no personnel use, aggregate-only outputs, preregistration, and void-on-leakage rules is reversible and informative.","source_ids":["S7","S8"]},"scalability":{"score":3,"rationale":"Digital forecasting tournaments scale, but compulsory balanced coverage, audits, appeals, question resolution, and resource accounting grow operationally with participants and questions.","source_ids":["S4","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Design and run one no-stakes microseason for approximately forty existing forecasters on non-sensitive questions, including preregistration, basic platform configuration, question/legal review, scoring analysis, participant survey, and a short report.","confidence":"MODERATE","assumptions":["Existing staff participate as part of paid work and no new platform is built.","The exercise uses 20-30 short-horizon questions and synthetic participant identifiers.","The Defra £1,500 legal/engagement and £18,500 three-round superforecasting milestones are treated as partial 2024 analogues and converted only at broad-band resolution to 2026 USD."],"source_ids":["S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Production rulebook, privacy and security design, HR/legal validation plan, accessible workflow, question bank, resolution committee, audit logging, appeals process, and configured forecasting platform.","confidence":"LOW","assumptions":["No bespoke platform is built from scratch.","Internal legal, HR, security, and data-protection staff contribute material time.","Employment-validation studies and accommodations may push the estimate toward the upper bound."],"source_ids":["S7","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"First full live season for forty contestants, including compensated forecasting time, independent administration, question writing and resolution, integrity audits, appeals, training, analysis, and transition support for eight council seats.","confidence":"LOW","assumptions":["Resource-equivalent staff time is counted even where no cash changes hands.","At least two independent governance functions are staffed throughout the season.","The funded council-seat time allocation is included; salaries and classified-system requirements are unknown."],"source_ids":["S7","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Recurring seasonal question production and resolution, contestant and council time, platform operation, audits, appeals, privacy/security review, post-season evaluation, and staggered promotion/relegation.","confidence":"LOW","assumptions":["Eight council seats consume a meaningful fraction of existing employees' time.","One or two seasons run annually.","No major custom-software redevelopment or classified infrastructure is required."],"source_ids":["S6","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Primary research, formal mechanism analysis, and first-party tournament documentation independently support exposure bias, copying risk, and strategic incentives in ranked forecasting competitions.","source_ids":["S3","S5","S6"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"IARPA is an established forecasting-research funder, and Defra documented an operational foresight need and commissioned a superforecasting prototype. Demand for the exact council-selection mechanism remains unshown.","source_ids":["S1","S8"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The integrated balanced-assignment, sealed, equal-resource, rolling and staggered council-selection bundle can be compared prospectively with open leaderboard and manager-selection baselines on held-out outcomes and safeguards.","source_ids":["S2","S4","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A no-stakes microseason can test workflow integrity, ranking stability, burden, auditability, and preliminary held-out performance without appointments or operational decisions.","source_ids":["S4","S8"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The bounded pilot can proceed without personnel consequences or sensitive publication, provided HR/ethics approval, synthetic identifiers, accessibility accommodations, aggregate-only outputs, and predefined halt rules are used. Live promotion remains unauthorized pending validation.","source_ids":["S7","S8"]},"credible_cost_scope_and_range":{"status":"YES","reason":"An official Defra prototype contract supplies partial superforecasting, legal-review, evaluation, and workshop cost anchors; broad bands explicitly add staff opportunity cost, governance, audit, and council-seat assumptions.","source_ids":["S8"]}},"next_evidence_step":"Preregister and run a no-stakes two-stage microseason with the forty forecasters. During the selection stage, randomly assign forecasters to (A) the candidate's balanced compulsory blocks with sealed estimates, equal research credits, and fixed updates or (B) an open self-selected leaderboard with the same question pool; separately ask managers to nominate eight forecasters without seeing tournament scores. Do not appoint anyone. In the held-out stage, have the eight highest candidate-rule scorers, eight open-leaderboard scorers, and eight manager nominees forecast the same new sealed questions under equal resources. Compare mean proper score, calibration, coverage, rank split-half reliability, aggregate accuracy and incremental value over the all-forecaster aggregate. Record seal breaches, recusals, resource exceptions, resolution agreement, audit corrections, appeals, burden, attrition, confidentiality events, and exploratory selection-rate differences. Falsify or materially undermine the candidate if fewer than 80% of questions resolve as preregistered; candidate ranks have split-half reliability below 0.5 or change materially under reasonable normalization choices; more than 10% of records have unresolvable seal/resource/audit exceptions; participant attrition or burden materially exceeds the open baseline; or the candidate-selected group fails to improve held-out proper score by at least 5% relative to the open-leaderboard group and fails to outperform or add information beyond manager nominees. Treat similarity screens only as investigation referrals. Use no result for employment, compensation, reputation, or operational decisions.","blocking_evidence":["No externally verified prevalence estimate exists for the alleged copying, collusion, resource inequality, outcome influence, or incumbent lock-in in the specific office.","No head-to-head evidence shows that the complete candidate bundle selects a council with better held-out accuracy or incremental decision information than an open leaderboard or manager appointment.","The proposed deterministic proper-score promotion rule has an unresolved incentive-compatibility defect identified in formal prior art.","No local evidence establishes that question strata can be balanced, normalized, and resolved with enough events to produce reliable two-season ranks.","Resource-cap auditability, seal integrity, collusion-screen false-positive rates, and appeal workload are untested.","If funded seats constitute an employment opportunity, job-relatedness, accommodations, adverse-impact monitoring, privacy, and jurisdiction-specific personnel authority remain unvalidated.","No evidence links league-selected council forecasts to improved downstream decisions or guards adequately against decision teams over-relying on the aggregate."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This evaluation measured only visible prior-art proximity using eight web sources. It did not measure world novelty, patentability, freedom to operate, market size, or realized impact. The search found substantial collision in tournament-based elite forecaster selection and in most component mechanisms; it did not establish whether the exact integrated staggered-council bundle has previously been implemented.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":false,"progress_targets":["Obtain preregistered held-out comparison data against both open-leaderboard and manager-selection baselines.","Test whether a randomized incentive-compatible selection rule such as ELF/I-ELF should replace deterministic top-score promotion.","Demonstrate reliable question balancing, resolution, normalization, and rank stability across enough independent events.","Quantify seal failures, resource-cap exceptions, audit corrections, appeals, burden, attrition, and collusion-screen false positives.","Complete jurisdiction-specific HR, privacy, accessibility, job-relatedness, and adverse-impact review before any live seat allocation.","Show that the selected council adds predictive information and decision usefulness beyond the all-forecaster aggregate."],"reason":"Bounded web research establishes a real problem, credible adopters, feasible components, and substantial prior-art collision. The remaining claim is an empirical claim about comparative selection quality, incentive behavior, workflow integrity, and personnel safety. Those questions require proprietary participant data and a live no-stakes trial, not additional bounded web search."},"proposal_index":4}