{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","cell_id":"bounded_rivalry_governance__mathematics","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_id":"math_bounded_proof_verification_slot_v0","proposal_index":1,"version":0,"title":"Bounded Proof-Verification Slot Challenge","problem":"A mathematics consortium pursuing a specified theorem has enough specialist time to formalize only one of several submitted proof routes. When the slot informally goes to the first apparently complete submission, teams can improve their chances through premature completeness claims, concealed dependencies, excessive presentation work, or withholding discovered flaws in rival routes. The selection can therefore favor contest strategy over the proof route most suitable for rigorous, maintainable verification.","actors":["Mathematical research teams submitting proof routes","Consortium committee controlling the verification slot","Independent mathematical judges","Formal-proof auditors and implementers","Independent appeal reviewer","Future readers and maintainers of the formalized proof"],"observable_state":"For each proof route, the consortium can observe submission and revision timestamps; declared axioms, lemmas, imported results, and unresolved gaps; compliance with a standardized proof-certificate format; independent reproduction outcomes; judge scores and rationales; appeal records; referee and formalizer hours consumed; detected omissions or misrepresentations; and whether the selected route later passes staged formalization checks. Cross-team communications are examined only when voluntarily disclosed or when a specific rule violation is alleged, not through general surveillance.","consequence":"The scarce verification slot may be consumed by a fragile or strategically packaged route, while sound alternatives lose access, referees duplicate work, teams have incentives to hide useful negative findings, and the resulting formal artifact carries avoidable dependency and maintenance burdens.","affected_objective":"Select a proof route for one consortium-funded formal-verification slot such that mathematical correctness is a noncompensable threshold and the chosen route is independently traceable, feasible to formalize, and maintainable within the consortium's bounded review capacity.","intervention":"Run a consent-based, time-bounded proof-route challenge under a frozen rulebook. Eligible teams submit a standardized proof certificate identifying every dependency, unresolved obligation, contributor, and known countercheck. Correctness is a pass/fail gate established by independent reproduction; only passing routes are ranked on predeclared secondary criteria: dependency transparency, estimated formalization burden, modularity, and explanatory coverage. The public status board reports stage and audit disposition rather than exposing hidden test details. Each team receives the same page allowance, number of clarification rounds, and judge-contact channel. Legitimate collaboration and route merging are allowed if disclosed; concealed shared control, sham independent submissions, retaliation, tampering, plagiarism, and knowingly withholding a discovered fatal flaw from the audit process are prohibited. A separate reviewer hears time-boxed procedural appeals. The selected route enters staged formalization, with a predeclared reopening trigger if an audit gate fails. All correctness-passing submissions retain attribution and remain accessible regardless of who receives the scarce slot.","structural_mapping":[{"archetype_element":"Rivalry purpose statement","domain_realization":"Use rivalry to reveal which independently developed proof route is most suitable for rigorous formal verification, not to determine mathematical truth or personal worth."},{"archetype_element":"Scarce prize or selection constraint","domain_realization":"One consortium-funded block of specialist formalization time for the specified theorem."},{"archetype_element":"Competitor eligibility boundary","domain_realization":"Teams consent to the process, disclose contributors and conflicts, submit a minimally complete proof certificate, and accept identical formatting and audit requirements."},{"archetype_element":"Contest arena boundary","domain_realization":"Teams may strengthen, compare, merge, or withdraw routes and answer standardized clarification requests; they may not tamper with submissions, misstate independence or dependencies, plagiarize, retaliate, contact judges privately, or conceal a known fatal defect from an active audit."},{"archetype_element":"Performance metric and scoring basis","domain_realization":"Independent mathematical reproduction is a noncompensable correctness gate; ranking after that gate uses frozen criteria for dependency transparency, modularity, estimated formalization burden, and explanatory coverage."},{"archetype_element":"Fair process and due process layer","domain_realization":"Identical submission constraints, recorded rationales, conflict-screened judges, a common question channel, and a separate time-boxed procedural appeal protect contestability."},{"archetype_element":"Anti-abuse guardrail","domain_realization":"Versioned submissions, contributor disclosures, judge-contact logs, reproducibility audits, and graduated responses distinguish legitimate collaboration from sham rivalry, sabotage, or misrepresentation."},{"archetype_element":"Externality and spillover boundary","domain_realization":"Submission-length and clarification caps bound demands on referees, while dependency and maintainability scoring brings downstream formalizer costs into the selection."},{"archetype_element":"Escalation damper","domain_realization":"A common proof-certificate format, page limit, fixed deadline, and two clarification rounds prevent presentation volume and repeated judge access from becoming the route to winning."},{"archetype_element":"Winner power and recalibration","domain_realization":"Winning controls only the current verification slot, not theorem attribution or future rules; staged audit gates can reopen selection, and a post-contest review can revise or retire the format."}],"mechanism_mapping":[{"mechanism_slug":"contest_rulebook","role":"Freezes eligibility, permitted conduct, the correctness gate, secondary scoring, tie-breaks, disclosures, and appeal procedure before judges see entrant identities.","counterfactual_removal":"Without the rulebook, judges could alter standards after seeing submissions, teams could not distinguish legitimate collaboration from prohibited conduct, and procedural challenges would revert to discretion."},{"mechanism_slug":"ranked_leaderboard_with_audit","role":"Provides an observable stage-and-score record while requiring independent reproduction and dependency checks before any route can be selected.","counterfactual_removal":"Without the audit coupling, a polished or strategically compressed proof certificate could rank highly despite a hidden gap or undeclared dependency."},{"mechanism_slug":"spending_cap_or_resource_cap","role":"Caps the scarce reviewer-facing resources each team may consume through a common page allowance, two clarification rounds, and one shared contact channel.","counterfactual_removal":"Without these caps, better-resourced teams could escalate presentation volume and repeated judge access, increasing referee burden without improving mathematical validity."},{"mechanism_slug":"sabotage_or_foul_penalty_schedule","role":"Defines proportionate responses to misrepresentation, tampering, retaliation, plagiarism, concealed conflicts, and repeated prohibited judge contact, with an opportunity to answer allegations.","counterfactual_removal":"Without a published and reviewable response schedule, misconduct would be either weakly deterred or punished ad hoc, making enforcement vulnerable to favoritism."},{"mechanism_slug":"post_contest_impact_review","role":"Compares the selected route's staged formalization performance with the stated objective and examines burdens imposed on judges, losing teams, and maintainers before another challenge is authorized.","counterfactual_removal":"Without ex-post review, a gamed metric, underestimated dependency burden, or winner lock-in could persist into later rounds even when the selected route performs poorly."}],"causal_chain":["A single formal-verification slot makes otherwise independent proof teams strategically interdependent.","A frozen rulebook separates mathematical contribution from prohibited off-arena tactics and gives every team the same procedural resources.","The correctness gate prevents strengths in presentation, modularity, or estimated cost from compensating for an unreproduced proof step.","Secondary scoring among correctness-passing routes directs rivalry toward transparent dependencies, modular structure, and bounded formalization burden.","Versioning, audits, disclosures, and reviewable sanctions raise the likelihood that concealment or sham independence is detected and cannot determine the award by itself.","Submission and contact caps reduce the advantage obtainable from escalating reviewer-facing effort.","A staged reopening trigger limits the selected route's control of the scarce slot when later formalization reveals a disqualifying defect.","Post-contest comparison of promised and realized properties supplies evidence for revising the next rulebook or abandoning the contest format."],"baseline":"The consortium gives the verification slot to the first route that a small internal panel regards as apparently complete. Review criteria, permitted revisions, dependency disclosures, contact with judges, and reconsideration after selection are handled informally.","nearest_rivals":["Blinded expert triage: a conflict-screened panel selects one route through holistic mathematical judgment without a contestant-facing scoring system; it may preserve nuance with less machinery but offers weaker advance predictability and contestability.","Lottery among routes that pass an independent correctness screen: this resists fine-grained metric gaming and status bias but does not deliberately select for formalization burden or maintainability.","Collaborative synthesis workshop: teams combine routes into one proof plan instead of competing; it may be preferable when contributions are strongly complementary and the scarce slot does not require a discrete choice among incompatible routes.","First independently reproduced proof receives the slot: this is simple and rewards speed to a valid result, but it does not compare multiple valid routes on downstream verification cost."],"remaining_contrastive_claim":"Relative to these baselines, the candidate's testable design distinction is to preserve a scarce, comparative selection while making correctness a noncompensable gate, equalizing reviewer-facing resources, bounding off-arena conduct, allowing procedural challenge, and reopening the slot if staged formalization invalidates the selection. This is a structural description, not a claim of novelty or superior effect.","authority_safety":{"decision_authority":"The consortium committee may govern only its own verification funding, submission process, shared communication channels, and access to consortium-controlled formalizers. Independent judges may determine compliance and scoring within the frozen rules; a separate reviewer may remedy procedural error. None of these actors determines mathematical truth beyond the submitted evidence or controls external publication and attribution.","authorized_first_step":"Authorize only a preregistered shadow exercise using de-identified proof packets for an already-settled theorem, with no funding, priority, employment, publication, or reputational consequence.","excluded_actions":["Claiming ownership of a team's theorem, proof, or authorship because its route wins or loses","Restricting teams from publishing, withdrawing, collaborating outside the challenge, or seeking independent review","Using general surveillance of private communications to search for collusion or misconduct","Treating an anomaly, allegation, or audit flag as a verdict without notice and an opportunity to respond","Changing eligibility, scoring, or sanctions after entrant identities or scores are known","Allowing a secondary score to compensate for failure of mathematical reproduction","Applying contest results to hiring, promotion, discipline, or future eligibility without separate authority and process"],"halt_rollback":"Pause scoring when a material conflict, data leak, rule ambiguity, unequal judge access, or credible due-process failure could affect ranking. Preserve the versioned record, appoint an unconflicted reviewer, and either rescore all affected packets under the frozen rule or void the shadow round. During a live implementation, failure of a staged correctness gate returns the slot to an uncommitted state rather than automatically transferring authorship or penalizing the team."},"negative_tests":{"strongest_counterevidence":"The strongest counterevidence would be that blinded expert triage selects routes that require no more correction or formalizer effort than challenge-selected routes, while the challenge consumes more reviewer time or induces greater strategic packaging. Evidence that teams avoid useful collaboration because of the contest would also weigh strongly against the design.","problem_falsifier":"The inferred problem is falsified in the target consortium if there is no genuine scarcity, proof routes do not compete for the same verification capacity, selection timing and strategic presentation do not affect access, or ordinary independent review already records dependencies, bounds judge contact, supplies appeals, and reopens failed selections.","intervention_falsifier":"The intervention is falsified for its intended mechanism if, under a shadow comparison, correctness-gated scoring does not distinguish known dependency or formalization burdens; judges cannot apply the frozen rubric with acceptable agreement; caps merely shift effort into undisclosed channels; planted rule-gaming survives audit; or the governed process requires more review capacity than the slot it is meant to allocate.","risks":["A visible ranking may convert a proof-development exercise into a status contest.","Any secondary rubric can become a target and favor a recognizable proof style rather than mathematical value.","Submission caps may disadvantage routes whose irreducible explanation is longer.","Requirements to disclose dependencies or known defects could expose unpublished work unless access and retention are tightly limited.","Policing collusion could mistakenly stigmatize legitimate mathematical collaboration.","Judges may underestimate formalization burden or import school-specific preferences.","A reopening trigger set too readily may create churn; one set too narrowly may preserve a failing route.","The process may consume scarce mathematical attention better spent on direct collaboration or verification."]},"next_evidence_step":"Pre-register a no-stakes shadow comparison using eight de-identified proof packets for one already-settled theorem, including packets with known gaps, undeclared dependencies, redundant presentation, and differing modularity. Randomly assign two conflict-screened panels to the frozen challenge procedure or the baseline holistic triage, record reviewer minutes and rationales, then have separate auditors reproduce the selected critical lemma chains without seeing panel assignment. Before inspection, define failure conditions: any incorrect packet passing the correctness gate, planted gaming determining selection, materially inconsistent rubric application, or governance overhead exceeding the preset shadow-review budget. Use the exercise only to decide whether a small consent-based live pilot is justified.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Not applicable to this single-proposal output.","revision_record":{"parent_version":null,"progress_targets_addressed":["Initial complete proposal"],"conceptual_changes":["Instantiated bounded rivalry as governance of competing proof routes for one scarce formal-verification slot.","Separated mathematical correctness from secondary selection criteria by making correctness noncompensable.","Kept theorem attribution and publication outside the contest prize."],"operational_changes":["Specified a frozen proof-certificate rulebook, equal reviewer-facing resource caps, staged audit gates, procedural appeal, and reopening trigger.","Limited the authorized first step to a no-stakes shadow comparison."],"evidence_changes":["Defined observable records and preregistered failure conditions for a bounded shadow exercise.","Marked prior art as unsearched and made no novelty, prevalence, demand, or effect-size claim."],"claim_changes":["Restricted the claim to a testable structural contrast with stated baselines.","Made continued use conditional on correctness, burden, gaming, collaboration, and governance-overhead evidence."]}}