{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","cell_id":"bounded_rivalry_governance__mathematics","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_id":"math_bounded_problem_set_selection_v0","proposal_index":3,"version":0,"title":"Bounded Problem-Set Selection Arena for a Proof Competition","problem":"An organization must select six problems from a larger pool for an annual proof-based mathematics competition. Problem authors receive a scarce combination of inclusion, attribution, and honorarium, so their submissions are rivals. If selectors reward individually impressive problems through informal deliberation, authors can improve their chances through submission volume, obscurity, lobbying, concealed overlap, or tailoring to groups with advance familiarity. Selecting the six highest-rated items independently can also produce a redundant, uneven, ambiguous, or difficult-to-grade set even when each item appears acceptable alone.","actors":["Mathematicians proposing competition problems","Competition organizing committee controlling the six problem slots","Conflict-screened problem selectors","Independent mathematical validators","Volunteer pilot solvers who are ineligible for the live competition","Independent grading-rubric auditors","Live contestants and their coaches","Competition graders","Independent procedural appeal reviewer"],"observable_state":"The organizer can observe versioned problem statements and solutions; author, coauthor, provenance, prior-circulation, coaching, and conflict disclosures; submission counts; selector contacts; mathematical validation results; pilot solution paths and completion records; ambiguity reports; grader agreement on sample responses; similarity among proposed items; portfolio coverage; revision histories; procedural appeals; and post-contest regrading or clarification events. Private communications are not generally monitored and may be examined only when voluntarily submitted in support of a specific alleged process violation.","consequence":"The selected set can overrepresent one topic or solution trick, contain an invalid or ambiguous item, impose avoidable grading disputes, advantage contestants exposed to related material, and consume participant effort on defects created by the selection process. Informal access and submission volume can determine inclusion more readily than the problem set's collective mathematical and assessment value.","affected_objective":"Assemble a confidential six-problem portfolio that is mathematically valid, independently solvable, gradable under a reproducible rubric, complementary across the organizer's declared structural and difficulty dimensions, and selected without rewarding leaks, lobbying, concealed familiarity, or submission-volume escalation.","intervention":"Run a confidential, bounded selection arena for problem proposals. Before submissions, publish an RFP and freeze eligibility, provenance disclosures, allowed assistance, submission limits, validation gates, portfolio criteria, conflicts, sanctions, and appeal procedure. Each eligible author may submit at most two problems and receive one standardized revision round through a common channel. Author identities remain hidden from selectors where the mathematical style does not reveal them. Every item must pass noncompensable checks for a valid statement, a complete reference solution, declared dependencies, and independent solvability. Non-live volunteer solvers then attempt coded items, and separate graders apply draft rubrics to sample solution paths. Passing problems receive individual descriptors, but the award decision optimizes the six-item portfolio for complementary mathematical structures, solution-method diversity, intended progression, statement accessibility, and grading reliability rather than selecting the six highest individual scores. Proposers may disclose and merge overlapping submissions before the freeze. Undisclosed prior circulation, plagiarism, live-item leakage, private selector contact, tampering, false independence, or tailoring with confidential contestant information triggers a published, reviewable response schedule. Winning authors gain only one-cycle inclusion and attribution; they do not acquire access to future live pools or selection authority. A post-contest review compares realized item behavior with the design assumptions and revises or retires the arena.","structural_mapping":[{"archetype_element":"Rivalry purpose statement","domain_realization":"Use rivalry to draw out independently constructed proof problems while making contribution to a coherent, valid, and gradable six-item set the route to selection."},{"archetype_element":"Scarce prize or selection constraint","domain_realization":"Six confidential live-contest slots, each carrying attribution and a fixed honorarium."},{"archetype_element":"Competitor eligibility boundary","domain_realization":"Eligible authors accept confidentiality, disclose provenance, collaborators, coaching relationships, prior circulation, and conflicts, and submit a complete reference solution and grading outline."},{"archetype_element":"Contest arena boundary","domain_realization":"Authors may develop, test privately with declared ineligible reviewers, revise once, withdraw, or merge overlapping proposals. They may not leak live candidates, lobby selectors privately, plagiarize, conceal prior exposure, misstate authorship independence, tamper with rival materials, or use confidential contestant information."},{"archetype_element":"Performance metric and scoring basis","domain_realization":"Validity and independent solvability are noncompensable gates. Selection after those gates scores the six-item portfolio for complementarity, intended progression, solution-method diversity, statement accessibility, and grading reliability."},{"archetype_element":"Fair process and due process layer","domain_realization":"Frozen requirements, coded submissions, conflict-screened validators, identical revision opportunities, recorded rationales, and a separate procedural appeal give proposers a contestable process without disclosing live problems."},{"archetype_element":"Anti-sabotage and anti-collusion guardrail","domain_realization":"Version histories, provenance declarations, controlled item access, similarity screening, selector-contact logs, and reviewable sanctions address leaks, sham independence, coordinated cover submissions, plagiarism, and tampering."},{"archetype_element":"Externality and spillover boundary","domain_realization":"Pilot solving and rubric audits bring ambiguity, familiarity advantage, contestant time loss, and grader burden into selection before those costs reach live participants."},{"archetype_element":"Escalation and arms-race damper","domain_realization":"A two-problem submission limit, common format, one revision round, and one communication channel prevent volume and repeated selector access from becoming dominant competitive resources."},{"archetype_element":"Prize decomposition or multiple-winner design","domain_realization":"Six awards are chosen as a complementary portfolio, so an individually lower-rated problem may be selected when it adds a needed mathematical structure or solution mode."},{"archetype_element":"Winner power and lock-in review","domain_realization":"Inclusion lasts for one contest cycle and grants no future selection role, privileged access, or presumptive slot; later rounds reopen under independently approved rules."},{"archetype_element":"Learning and recalibration loop","domain_realization":"Post-contest evidence on ambiguities, unexpected solution paths, grading disputes, progression, and exposure concerns informs the next rulebook or a decision to stop competitive sourcing."}],"mechanism_mapping":[{"mechanism_slug":"tender_or_rfp_process","role":"Publishes the desired portfolio, eligibility rules, confidentiality duties, validation stages, evaluation criteria, and challenge process before authors submit problems.","counterfactual_removal":"Without the structured solicitation, requirements could be communicated unevenly, selectors could favor familiar authors, and the final rationale could be reconstructed after a preferred set was chosen."},{"mechanism_slug":"contest_rulebook","role":"Freezes allowed assistance, submission and revision limits, mathematical gates, portfolio scoring, conflicts, prohibited conduct, sanctions, tie-breaks, and procedural appeals.","counterfactual_removal":"Without a binding rulebook, selectors could change standards after recognizing authors or seeing which problems benefit from a particular interpretation."},{"mechanism_slug":"multiple_award_or_portfolio_selection","role":"Chooses six problems for their joint coverage and complementarity rather than treating each slot as an independent individual ranking.","counterfactual_removal":"Without portfolio selection, the six highest individual scores could duplicate topics, difficulty levels, or solution techniques and fail the objective at the set level."},{"mechanism_slug":"spending_cap_or_resource_cap","role":"Limits each author to two submissions, one standardized revision, a common document format, and the same selector-contact channel.","counterfactual_removal":"Without resource caps, well-resourced authors could dominate attention through submission volume, repeated polishing, or privileged access rather than through the usefulness of their problems."},{"mechanism_slug":"sabotage_or_foul_penalty_schedule","role":"Assigns proportionate, predeclared responses to negligent nondisclosure, repeated private contact, false provenance, plagiarism, tampering, or live-item leakage, with notice and an opportunity to answer.","counterfactual_removal":"Without a published schedule, serious misconduct might be weakly deterred while minor mistakes could receive arbitrary maximal punishment, and enforcement could vary by author status."},{"mechanism_slug":"post_contest_impact_review","role":"Examines the selected set after grading, including unexpected valid solutions, ambiguity, item interactions, exposure concerns, grader disagreement, and whether selection criteria predicted the properties they were intended to capture.","counterfactual_removal":"Without retrospective review, selectors could defend an ineffective rubric, overlook harms experienced only during the live contest, and repeat the same portfolio or governance errors."}],"causal_chain":["Six live slots, attribution, and honoraria make proposed problems rival claims on a scarce opportunity.","A frozen confidential rulebook defines legitimate competition and removes submission volume, private selector access, and concealed prior exposure from the permitted route to winning.","Independent mathematical validation prevents aesthetic appeal or portfolio usefulness from compensating for an invalid statement or incomplete reference solution.","Pilot solving and grading dry runs expose ambiguity, alternative interpretations, and rubric weaknesses before contestants bear their costs.","Portfolio-level selection changes the target from maximizing an item's isolated impressiveness to adding complementary value to the six-problem set.","Equal submission and revision allowances dampen escalation in authoring volume and selector-facing effort.","Versioning, disclosures, controlled access, and reviewable sanctions make leakage, plagiarism, sham independence, and tampering less useful as contest strategies.","One-cycle awards prevent accepted authors from converting a win into control over future selection.","Post-contest comparison of intended and realized item behavior supplies evidence for recalibrating the next arena or replacing it with a noncompetitive process."],"baseline":"Problem proposals arrive through professional networks and are discussed by an organizing committee without uniform submission limits, coded evaluation, standardized provenance disclosures, pilot-grading requirements, or a formal appeal path. Selectors evaluate attractive items largely one at a time and assemble the final set through informal adjustment.","nearest_rivals":["Commissioned editorial team: a small paid group collaboratively writes the entire set without author rivalry. This may improve coherence and confidentiality but concentrates stylistic and mathematical judgment in one team.","Open call followed by holistic expert selection: selectors consider all submissions and assemble a set without fixed scoring or procedural machinery. This preserves expert discretion but makes access, consistency, and contestability harder to inspect.","Random selection among validated problems within predefined topic strata: a lottery limits favoritism and fine-grained metric gaming, but it does not deliberately optimize interactions among difficulty, solution method, statement form, and grading burden.","Collaborative problem workshop: contributors jointly revise and combine candidate ideas without individual slot attribution. This may be preferable when idea sharing matters more than independent alternatives and confidential competition would inhibit useful synthesis."],"remaining_contrastive_claim":"The candidate's testable structural distinction is to govern problem authors as rivals for a multi-item portfolio while treating mathematical validity as a gate, set-level complementarity as the selection target, contestant and grader burdens as externalities, and leakage and submission escalation as off-arena conduct. This does not assert that competitive sourcing is preferable to commissioning, lottery, expert deliberation, or collaborative authorship.","authority_safety":{"decision_authority":"The organizing committee may govern its own confidential problem pool, select six items, allocate its stated honoraria, and exclude a submission from the current process. Validators may determine mathematical gate results; selectors may assess portfolio fit; an independent reviewer may remedy procedural error. They may not control external publication, academic employment, professional standing, or attribution beyond the submitted contest license.","authorized_first_step":"Authorize only a preregistered shadow selection using retired or purpose-written non-live problems, anonymized author labels, and volunteer solvers who cannot participate in the corresponding live competition. No shadow result may determine a live slot, payment, or public author ranking.","excluded_actions":["Using an unverified shadow result to select or reject a live contest problem","Searching private communications without consent or treating stylistic similarity as proof of collusion","Publishing confidential submissions, pilot responses, or allegations before a reviewable finding","Penalizing an author's employment, publication access, professional membership, or future unrelated work","Allowing selectors to evaluate submissions from undisclosed collaborators, students, supervisors, or coaching clients","Changing validity gates, portfolio weights, submission limits, or sanctions after author identity or performance becomes visible","Trading mathematical validity for portfolio balance or an intended difficulty profile","Giving accepted authors privileged access to future problem pools or authority over future rivals"],"halt_rollback":"Immediately quarantine an item after a credible leak, provenance conflict, mathematical defect, unequal selector contact, or pilot-access breach. Preserve the versioned record, replace conflicted reviewers, and reassess every affected item under the frozen rules. In a live process, replace a compromised item only from a prevalidated reserve before contestants see it; after exposure, halt use of the item and apply the contest's predeclared neutral scoring remedy rather than improvising an author-specific penalty."},"negative_tests":{"strongest_counterevidence":"The strongest counterevidence would be that a commissioned editorial team or collaborative workshop produces a more coherent and reliably gradable set within the same review budget, while the arena suppresses idea sharing, attracts strategic submissions, or requires secrecy and enforcement burdens disproportionate to the six selections.","problem_falsifier":"The inferred problem is falsified if problem authors receive no scarce inclusion, attribution, payment, or strategic advantage; all items are jointly authored without rival claims; there is no bounded live set; or the organizer already uses uniform disclosures, independent validation, portfolio selection, controlled access, equal revision opportunities, and post-contest review without material strategic behavior.","intervention_falsifier":"The intervention is falsified for its intended mechanism if coded review does not meaningfully separate identity from evaluation, pilot solvers cannot expose planted ambiguities, portfolio rankings reverse under reasonable changes to declared weights, submission caps merely shift work to undisclosed coauthors, prohibited exposure cannot be distinguished from innocent mathematical similarity, or the governed process consumes more qualified reviewer capacity than direct commissioning.","risks":["Recognizable mathematical style may defeat author anonymization.","Pilot testing creates an additional path for live-item leakage.","Problems may be optimized for the pilot-solver pool rather than the intended contestants.","Portfolio weights may exclude an excellent problem because selectors overvalue categorical balance.","Submission limits may favor established authors who already possess mature problem inventories.","Confidentiality rules may inhibit legitimate collaboration or prior feedback.","Similarity screening may falsely flag independently discovered constructions.","A published sanction schedule may be too weak for deliberate leakage or too rigid for an innocent disclosure mistake.","Difficulty estimates may remain unstable even after pilot solving.","Post-contest review may blame authors for grading or administration failures outside the problem design." ]},"next_evidence_step":"Pre-register a no-stakes shadow exercise with 18 retired or purpose-written non-live problems, including known invalid statements, ambiguous wording, overlapping solution structures, and uneven grading rubrics. Freeze the six-slot portfolio criteria, submission caps, access rules, reviewer-hour ceiling, and stop conditions before coded packets are assigned. Have one conflict-screened group validate items, a separate group of ineligible volunteers solve them, and independent graders score multiple sample solution paths. Compare a governed portfolio selection with an independently staffed holistic baseline using mathematical defects detected, redundant item pairings, grading disagreements, procedural exceptions, and reviewer time. Stop without a live pilot if any known-invalid item passes the gate, planted ambiguity is not detected, portfolio choice is unstable under the preregistered sensitivity checks, confidentiality cannot be maintained, or governance exceeds its review budget.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Proposal 1 governs teams competing for one formal-verification slot among proof routes to the same theorem; its causal path uses independent proof reproduction, route scoring, and staged formalization to counter premature completeness and concealed dependencies. Proposal 3 instead governs authors competing for six slots in a live problem-set portfolio; its causal path uses confidential validation, pilot solving, grading audits, portfolio complementarity, and submission caps to counter ambiguity, leakage, redundancy, and assessment burdens. Proposal 2 governs foundational interfaces competing to become a shared library default; its causal path centers on semantic translation, secured exit duties, network-effect lock-in, challenger access, and separation of winner maintenance from arena rulemaking. Proposal 3 neither chooses a proof route nor establishes a mathematical foundation or persistent standard: it makes a one-cycle portfolio decision for a time-bounded event, after which every slot reopens. It can be adopted independently of both earlier interventions.","revision_record":{"parent_version":null,"progress_targets_addressed":["Initial complete formulation for proposal index 3","Material separation from sealed proposal indices 1 and 2"],"conceptual_changes":["Instantiated bounded rivalry around authors competing for inclusion in a multi-item mathematical problem set.","Made portfolio complementarity, rather than individual rank, the scarce-prize selection objective.","Located major spillovers in contestant exposure, ambiguity, and grading burden rather than proof formalization or interface migration."],"operational_changes":["Specified coded submissions, provenance and conflict disclosure, validity gates, pilot solving, grading dry runs, portfolio scoring, submission caps, reviewable sanctions, and one-cycle winner limits.","Restricted first evidence to a preregistered shadow exercise using non-live problems."],"evidence_changes":["Defined observable validation, pilot, grading, portfolio, access, and procedural records.","Specified shadow comparison measures and explicit stop conditions.","Kept prior art unsearched and made no novelty, prevalence, demand, or effect-size claim."],"claim_changes":["Restricted the claim to a testable governance distinction among problem-sourcing processes.","Explicitly retained commissioned, collaborative, lottery, and holistic expert selection as potentially preferable alternatives."]}}