{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"bounded_rivalry_governance__mathematics:P3:v0","cell_id":"bounded_rivalry_governance__mathematics","search_queries":["site:imo-official.org regulations problems shortlist submitted problem proposals confidentiality marking schemes","mathematics competition problem proposal submission guidelines confidentiality problem selection committee","math olympiad problem selection ambiguity grading dispute research","competition mathematics problem setting guidelines test solving marking scheme","site:maa.org AMC problem submission writer proposal competition problem author","site:ukmt.org.uk \"problem proposals\" mathematics competition","site:imo-official.org \"problem proposals\" \"confidential\" shortlist regulations","mathematical olympiad problem selection committee test solve problems guidelines official","official Standards educational psychological testing pilot test items validity fairness open access","site:testingstandards.net test development item review pilot testing standards","site:bls.gov mathematicians median pay 2025 occupational employment wages","peer reviewed mathematics olympiad test validity item selection ambiguity","mathematics olympiad problem withdrawn ambiguity official","math competition problem error scoring adjustment official ambiguous question","olympiad problem invalid statement competition official correction","math contest question flawed ambiguity grading official announcement"],"sources":[{"source_id":"S1","title":"Regulations — International Mathematical Olympiad","publisher":"International Mathematical Olympiad","url":"https://www.imo-official.org/regulations/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["The host organization has overall responsibility for fair-play arrangements.","Each participating country may securely submit up to six proposed problems with solutions.","Proposals should cover multiple fields and difficulty levels, be new, and not have appeared in another competition.","The Problem Selection Committee creates a shortlist and the Jury selects six problems and approves marking schemes.","Proposals, shortlists, contest problems, and solutions are subject to confidentiality controls.","Leaders must disclose known shortlisted problems; marking disputes and alleged rule violations have escalation procedures."]},{"source_id":"S2","title":"What is IMO?","publisher":"International Mathematical Olympiad","url":"https://www.imo-official.org/about/?language=en","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"undated; current in 2026","accessed_at":"2026-08-03","claims_supported":["The IMO uses a six-problem proof-based examination involving more than 100 national teams.","Participating countries submit proposed problems, a host-appointed committee prepares a shortlist, and the Jury selects the final six.","The final set is intentionally arranged across mathematical areas and a planned difficulty progression.","Proofs receive partial-credit scores and disputed marks can escalate through coordinators to the Jury."]},{"source_id":"S3","title":"28th Junior Balkan Mathematical Olympiad — Regulations","publisher":"Scientific and Technological Research Council of Türkiye (TÜBİTAK)","url":"https://jbmo2024.tubitak.gov.tr/regulations","source_class":"OFFICIAL_GUIDANCE","publication_date":"2024","accessed_at":"2026-08-03","claims_supported":["Countries may propose up to five problems with solutions and subject classifications.","A Problem Selection Committee shortlists proposals and the Jury chooses the final four-problem set.","Unselected proposals remain confidential, and decision-makers are isolated from external communications before the contest.","Contestants may request clarification, scores require coordination agreement, and disputed scores have an appeal path."]},{"source_id":"S4","title":"NCES Statistical Standard 2-6: Educational Testing","publisher":"National Center for Education Statistics, U.S. Department of Education","url":"https://nces.ed.gov/statprog/2002/std2_6.asp","source_class":"STANDARD","publication_date":"2002","accessed_at":"2026-08-03","claims_supported":["Assessment instruments should have explicit, reproducible specifications covering purpose, content, participant characteristics, timing, psychometric properties, administration, and scoring.","Items should be reviewed before and after pilot or field testing using participants similar to the intended population.","Validity, scorer reliability, fairness, documented judge selection, fixed scoring criteria, and monitoring of scorer consistency are recognized requirements.","The standard supports pilot testing and rubric auditing but does not prescribe rivalry governance among problem authors."]},{"source_id":"S5","title":"2025-26 Part B Examination in Mathematics: D. Form of Questions","publisher":"Mathematical Institute, University of Oxford","url":"https://courses.maths.ox.ac.uk/mod/book/view.php?chapterid=1028&id=65690","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-2026","accessed_at":"2026-08-03","claims_supported":["Question setters are instructed to provide complete model solutions and comprehensive marking schemes.","Question papers and marking schemes undergo whole-board and external-examiner review before sign-off.","Review explicitly considers comparative difficulty, similarity to recent questions, clarity, and assessment structure.","These procedures are a close analogue for validation and grading-rubric audit, although not for an author competition."]},{"source_id":"S6","title":"Table 1. National employment and wage data by occupation, May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.t01.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["The May 2025 mean wage for mathematicians was $62.15 per hour and the median was $60.92 per hour.","These wages provide a transparent near-2026 labor-value proxy for expert mathematical validation and review, subject to overhead and geographic uncertainty."]},{"source_id":"S7","title":"Contest Guidelines","publisher":"Philippine Mathematical Olympiad","url":"https://home.pmo.ph/contest-guidelines/","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"undated","accessed_at":"2026-08-03","claims_supported":["The PMO states that preparation, analysis, and evaluation of competition questions continues throughout the year.","Open-ended proof questions are assessed by university mathematics professors, with conflict restrictions and partial-credit judgment.","Some papers are double-checked, illustrating organizational demand for accuracy and independent review.","The page identifies a technical committee and boards of judges as plausible adopters or authorizers."]},{"source_id":"S8","title":"Investigating the treatment of missing data in an Olympiad-type test — the case of the selection validity in the South African Mathematics Olympiad","publisher":"Pythagoras / AOSIS","url":"https://pythagoras.org.za/index.php/pythagoras/article/view/333","source_class":"PRIMARY_RESEARCH","publication_date":"2016-10-31","accessed_at":"2026-08-03","claims_supported":["Olympiad scoring and weighting choices can change which contestants are selected.","The study's microanalyses identified selection-validity concerns and recommended changes to scoring, weighting, and finalist counts.","This supports the materiality of assessment-design choices, but it does not estimate the prevalence of lobbying, leakage, redundant proof problems, or submission-volume gaming."]}],"problem_evidence":{"support":"MODERATE","rationale":"The problem matters because a six-item proof assessment determines consequential rankings, and official systems devote substantial machinery to novelty, field and difficulty coverage, confidentiality, marking schemes, clarification, and dispute resolution. NCES and Oxford guidance independently establish that item specifications, piloting, complete solutions, rubric review, and scorer reliability are necessary controls. Primary research demonstrates that Olympiad scoring choices can alter selection. However, no searched source measures the candidate's specific alleged author-side behaviors—lobbying, concealed overlap, excessive submission volume, or tailoring to familiar contestants—or their prevalence or effect size.","source_ids":["S1","S2","S3","S4","S5","S8"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Identifiable authorizers include an IMO host organization and Jury, a JBMO organizing country and Jury, and the PMO Technical Committee and boards of judges. Their published procedures express operational need for secure problem sourcing, balanced sets, validation, conflict management, grading consistency, and appeals. They do not express demand for this particular integrated author-rivalry arena, anonymized evaluation, two-submission cap, standardized revision round, or portfolio optimizer.","source_ids":["S1","S2","S3","S7"]},"prior_art":{"proximity":"ESTABLISHED_PRACTICE","closest_analogues":[{"name":"International Mathematical Olympiad problem-proposal, shortlist, Jury-selection, confidentiality, and marking-coordination process","similarity":"Very close: it solicits rival proposals with solutions under a submission cap, requires novelty and coverage across fields and difficulty, securely shortlists them, selects exactly six as a set, protects confidentiality, approves marking schemes, handles familiarity disclosures, and resolves scoring disputes.","remaining_difference":"The candidate adds author-level coded submissions, a two-item cap, one common revision round, explicit noncompensable validation gates, ineligible pilot solvers, independent rubric dry runs, formal portfolio optimization and sensitivity analysis, author-facing sanctions and appeals, and a post-contest causal review.","source_ids":["S1","S2"]},{"name":"Junior Balkan Mathematical Olympiad problem selection","similarity":"Close: capped proposals include solutions and fields, a selection committee creates a shortlist, a Jury chooses a multi-problem paper, confidentiality is enforced, clarification is available, and scoring has coordination and appeal.","remaining_difference":"The published rules do not establish anonymized authors, pilot solving, rubric reliability testing, explicit portfolio complementarity metrics, equal revision opportunities, or comparative post-contest evaluation.","source_ids":["S3"]},{"name":"NCES test-development and pilot-testing standard combined with Oxford mathematics examination review","similarity":"Close on explicit specifications, expert review, pilot or field testing, complete model solutions, marking schemes, scorer reliability, external review, and final sign-off.","remaining_difference":"These are assessment-quality systems rather than governed rivalry among credited authors competing for six paid slots.","source_ids":["S4","S5"]},{"name":"Philippine Mathematical Olympiad test-development and conflict-screened grading practice","similarity":"Adjacent: it performs continuing preparation, analysis, and evaluation of problems and uses conflict-restricted expert graders and double-checking.","remaining_difference":"The public guidance does not describe a confidential author competition, portfolio-level selection, submission caps, pilot solvers, or a proposer appeal process.","source_ids":["S7"]}],"distinctive_claim_remaining":"Under a fixed reviewer-hour ceiling, adding author-level caps and coded review, noncompensable mathematical gates, ineligible pilot solving, independent rubric dry runs, and preregistered six-item portfolio optimization to an otherwise competent holistic selection process will reduce undetected defects, redundant problem pairings, and grader disagreement without increasing confidentiality exceptions or selection instability. This comparative claim is falsifiable but untested; it is much narrower than claiming novelty of confidential Olympiad problem selection.","confidence":"HIGH"},"implementation_evidence":{"support":"STRONG","rationale":"All central workflow elements are technically ordinary: secure submissions, confidentiality, capped proposals, independent committees, complete solutions, item review, pilot testing, external examination, marking schemes, conflict screening, score coordination, and appeal are documented practices. A shadow study using retired problems avoids authority over live slots and materially reduces contestant risk. Remaining feasibility uncertainties concern whether mathematical style defeats anonymization, whether pilot access can remain secure, whether suitable ineligible solvers can be recruited, whether portfolio weights are stable, local privacy and contract terms, and the qualified-reviewer burden.","source_ids":["S1","S3","S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"A defective or poorly balanced six-problem set can affect hundreds of contestants, grading labor, and institutional legitimacy, but the direct population and annual frequency are bounded and no effect size for the proposed controls is known.","source_ids":["S2","S4","S8"]},"stakeholder_pull":{"score":2,"rationale":"Multiple organizers visibly invest in related safeguards, but none requests the candidate's incremental integrated arena or commits resources to test it.","source_ids":["S1","S3","S7"]},"incremental_advantage":{"score":2,"rationale":"Pilot solving, rubric dry runs, and explicit portfolio sensitivity could improve an informal process, but the strongest analogue already implements most of the selection, confidentiality, coverage, marking, and dispute architecture; comparative advantage remains unmeasured.","source_ids":["S1","S2","S4","S5"]},"distinctiveness_plausibility":{"score":2,"rationale":"The precise bundled comparison is distinguishable, but the component practices and much of their integration are established. World novelty and patentability were not assessed.","source_ids":["S1","S3","S4","S5"]},"technical_implementability":{"score":4,"rationale":"The intervention mainly combines established administrative and assessment practices. Secure access control, versioning, coded packets, scoring forms, and sensitivity analysis require no speculative technology.","source_ids":["S1","S3","S4","S5","S7"]},"adoption_authority_feasibility":{"score":4,"rationale":"Competition host organizations, technical committees, and juries visibly possess authority over problem intake, review, confidentiality, final selection, marking, and appeals. The proposed shadow study stays within ordinary organizer authority if participation and data handling are consented.","source_ids":["S1","S3","S7"]},"evidence_readiness":{"score":3,"rationale":"Public prior art and standards support a preregistered shadow comparison, but no suitable comparative dataset, adopter commitment, reviewer-time log, or author-behavior prevalence estimate was found.","source_ids":["S1","S4","S5","S8"]},"safety_net_benefit":{"score":4,"rationale":"Retired or purpose-written items, ineligible volunteer solvers, no live-slot consequences, planted defects, access logs, and predefined stopping rules make a useful low-stakes test possible. Confidentiality and allegation-handling protocols remain necessary.","source_ids":["S1","S3","S4"]},"scalability":{"score":3,"rationale":"The workflow can be reused by other proof competitions, but expert validation, pilot solving, and rubric audits scale largely with item count and may be constrained by scarce trusted mathematicians.","source_ids":["S4","S5","S6","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Design and execute one no-stakes 18-item shadow comparison, including preregistration, coded packet preparation, mathematical validation, pilot solving, duplicate grading, portfolio selection, analysis, and a short report.","confidence":"MODERATE","assumptions":["Approximately 250-500 expert hours across validators, solvers, graders, selectors, administration, and analysis.","The May 2025 mathematician median of $60.92 per hour is treated as a near-2026 labor-value proxy.","Some participants may volunteer; loaded coordination and administrative costs offset part of that saving.","Existing secure document and survey tools are used; no custom platform is built."],"source_ids":["S4","S5","S6"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"One-time rulebook, secure workflow, conflict and provenance forms, privacy and license review, portfolio-scoring specification, reserve-item and incident procedures, reviewer recruitment and training, and rehearsal for a real annual competition.","confidence":"LOW","assumptions":["Approximately 600-1,500 expert, legal, administrative, and technical hours.","The organizer adapts existing identity, storage, and access-control systems.","The band excludes the host event's general travel, accommodation, venue, and contestant costs.","Jurisdiction-specific privacy, employment, honorarium, and intellectual-property review may materially change cost."],"source_ids":["S1","S3","S4","S6"]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"Operate the first live sourcing cycle: intake, access control, disclosures, validation, pilot solving, rubric audits, selection meetings, author honoraria, procedural review capacity, monitoring, reserves, and post-contest review.","confidence":"LOW","assumptions":["Approximately 800-2,000 compensated resource-equivalent hours, plus six modest honoraria.","Many committee members may be volunteers, but their time is valued at a professional resource-equivalent rate.","No major security incident, item replacement after exposure, litigation, or custom software build occurs.","Live adoption follows a successful shadow study and separately approved local rules."],"source_ids":["S1","S3","S6","S7"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Repeat one annual six-problem sourcing cycle, maintain secure records, refresh and train reviewers, conduct validation and pilots, administer appeals, pay honoraria, and complete post-contest analysis.","confidence":"LOW","assumptions":["Approximately 650-1,800 annual expert and administrative hours after startup.","A trusted reviewer and ineligible-solver pool can be retained.","Software and legal templates are reused.","The estimate values donated professional time as a resource cost rather than treating it as free."],"source_ids":["S1","S3","S6","S7"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official competition procedures, assessment standards, examination guidance, and primary Olympiad research support the material importance of validity, balanced content and difficulty, confidentiality, grading reliability, and defensible scoring. Specific author-side abuse prevalence remains an evidence gap but does not negate the broader problem.","source_ids":["S1","S2","S3","S4","S5","S8"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"IMO host organizations and Juries, JBMO organizing countries and Juries, and the PMO Technical Committee are identifiable bodies with demonstrated authority over problem development, selection, security, marking, and appeals.","source_ids":["S1","S3","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The candidate can be tested against a competent holistic baseline under equal reviewer-hour constraints on defect detection, redundant pairings, grader disagreement, confidentiality exceptions, and portfolio stability.","source_ids":["S1","S4","S5"]},"bounded_next_evidence_step":{"status":"YES","reason":"An 18-item no-stakes shadow exercise with planted defects, preregistered comparators, a reviewer-hour ceiling, and explicit stop conditions is bounded in time, exposure, authority, and cost.","source_ids":["S4","S5","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"For the shadow step only, retired or purpose-written items, ineligible consenting solvers, no live decisions or public rankings, controlled access, and predefined incident handling avoid material contestant or author consequences. Live adoption would require separate privacy, licensing, security, and procedural authorization.","source_ids":["S1","S3","S4"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four scopes are explicit and broad bands are anchored to enumerated expert-hour assumptions and an official near-2026 mathematician wage benchmark. Confidence is limited because no organizer time logs, local quotes, honorarium policy, or security-platform specification was found.","source_ids":["S6"]}},"next_evidence_step":"With a willing competition organizer, preregister a four-to-six-week, no-stakes shadow study using 18 retired or purpose-written proof problems, including at least two known-invalid statements, three planted ambiguities, three deliberately overlapping solution structures, and sample responses designed to expose rubric disagreement. Randomly assign conflict-screened reviewers to (A) the full governed workflow and (B) an independently staffed holistic expert-selection baseline; add a commissioned-editorial comparator if staffing permits. Give arms equal total expert hours and prevent cross-arm communication. Freeze six-slot coverage criteria, selection weights, sensitivity thresholds, access rules, and stop conditions. Measure known-defect sensitivity and false-positive rate, redundant selected pairs, duplicate-grader agreement, reviewer hours, procedural exceptions, author-identification accuracy from style, and stability of the selected portfolio under preregistered weight perturbations. Falsify advancement if either arm leaks an item; the governed arm admits any known-invalid item; detects fewer planted ambiguities than the holistic baseline; exceeds its reviewer-hour ceiling by more than 20%; produces lower grader agreement; identifies authors above the preregistered anonymity threshold; or changes at least three of six selections under reasonable weight perturbations. Do not let results determine a live slot, payment, sanction, or public author ranking.","blocking_evidence":["No comparative field evidence shows that the full governed workflow outperforms holistic expert selection, commissioning, a collaborative workshop, or stratified lottery under equal expert hours.","No source quantified the prevalence or consequences of author lobbying, submission-volume escalation, concealed overlap, false independence, or live-item tailoring in proof competitions.","No identified organizer expressed demand for the incremental package or committed problems, reviewers, solvers, or funds to a shadow study.","The effectiveness of coded review against recognizable mathematical style is unknown.","The leakage risk introduced by volunteer pilot solvers has not been measured.","Portfolio weights, sensitivity limits, and acceptable redundancy metrics have not been validated with organizers or contestants.","Cost bands lack organizer time logs, vendor quotes, jurisdiction-specific legal review, and a defined honorarium schedule.","World novelty, patentability, freedom to operate, market size, and realized impact are unmeasured."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search establishes extensive prior art for capped problem proposals, secure shortlisting, multi-problem Jury selection, content and difficulty balancing, confidentiality, marking schemes, clarification, conflict controls, pilot testing, external review, scorer reliability, and appeals. It does not establish whether any organizer has previously combined every candidate feature at author level or run the proposed controlled comparison. No conclusion is made about world novelty, patentability, freedom to operate, market size, or realized impact.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain a written partnership and data-governance agreement from a competition organizer for the no-stakes 18-item shadow exercise.","Preregister the two principal arms, optional commissioned comparator, reviewer-hour budget, planted defects, outcomes, sensitivity analysis, and falsifiers before reviewers see items.","Collect blinded item-level validation, pilot-solution, duplicate-grading, portfolio-choice, access-log, anonymity, and reviewer-time data.","Demonstrate fewer undetected defects or redundant pairings and no worse grader agreement than the holistic baseline without exceeding the hour ceiling or triggering a confidentiality stop.","Replace wage-based planning estimates with observed hours, actual honorarium policy, and jurisdiction-specific privacy, licensing, and security costs.","Document whether adopters value the remaining incremental features enough to authorize a separately governed live pilot."],"reason":"Bounded web research establishes that the problem category is consequential, identifies credible authorizers, shows strong implementability, and reveals substantial collision with established Olympiad and assessment practice. The only material remaining claim is comparative incremental performance under a fixed review budget. That claim requires partner access, proprietary workflow data, and a controlled shadow exercise rather than more bounded web search; therefore the required controller action is an empirical-research stop, with repairable set false as required."},"proposal_index":3}