{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"bounded_rivalry_governance__mathematics:P2:v0","cell_id":"bounded_rivalry_governance__mathematics","search_queries":["formal mathematics library interoperability foundations translation semantic preservation MMT Math-in-the-Middle","formal proof assistant libraries interoperability foundation independent interchange standard OpenMath OMDoc","mathlib governance technical committee official","Lean mathlib port maintenance person years paper","site:leanprover-community.github.io mathlib governance committee official","site:github.com/leanprover-community/mathlib4 governance.md committee","OpenDreamKit deliverable interoperability mathematical software expressed need Math in Middle","formal mathematics interoperability benchmark translation proof assistants competition challenge","Logipedia proof library interoperability foundations official project","Dedukti interoperability proof systems common framework library translation paper","universal proof checker foundations interoperability formal mathematics library Dedukti","proof assistant interoperability contest benchmark common corpus semantic equivalence","site:mathlib-initiative.org governance board Mathlib Initiative funding official","Mathlib Initiative board governance fund formal mathematics library official","Mathlib Initiative mission maintenance funding official","Lean FRO mathlib funding maintenance formal mathematics library"],"sources":[{"source_id":"S1","title":"ITPEval: Benchmarking Formal Translation Across Interactive Theorem Provers","publisher":"arXiv","url":"https://arxiv.org/abs/2607.19407","source_class":"PRIMARY_RESEARCH","publication_date":"2026-07-07","accessed_at":"2026-08-03","claims_supported":["Proofs remain siloed across incompatible interactive theorem provers, limiting portability and reusable training data.","A released benchmark compares 1,560 files and 6,848 theorems across Lean 4, Rocq, Isabelle, and HOL Light.","Proof translation achieved at most 10.5% pass@1, while ecosystem-level translation achieved 5.2%, indicating major library-level difficulty.","Native type checking can overstate semantic fidelity: a separate equivalence check confirmed only 54.0% of one set of verified translations."]},{"source_id":"S2","title":"Experiences from Exporting Major Proof Assistant Libraries","publisher":"arXiv","url":"https://arxiv.org/abs/2005.03089","source_class":"PRIMARY_RESEARCH","publication_date":"2020-05-05","accessed_at":"2026-08-03","claims_supported":["Interoperability and integration of proof-assistant libraries is described as highly valued but elusive.","Libraries from Coq, HOL Light, IMPS, Isabelle, Mizar, and PVS were exported into OMDoc/MMT.","The exports encountered theoretical, technical, and social challenges, including unresolved problems.","The paper reports that satisfactory translation and library-access mechanisms remain lacking."]},{"source_id":"S3","title":"Interoperability in the OpenDreamKit Project: The Math-in-the-Middle Approach","publisher":"arXiv","url":"https://arxiv.org/abs/1603.06424","source_class":"PRIMARY_RESEARCH","publication_date":"2016-03-21","accessed_at":"2026-08-03","claims_supported":["The EU-funded OpenDreamKit infrastructure project made interoperability among mathematical software systems an explicit work-package mission.","Its Math-in-the-Middle architecture uses a central ontology and system-function specifications as a common meaning space.","A cooperative neutral-interchange architecture is established prior art and a direct comparator to selecting one default interface."]},{"source_id":"S4","title":"The OpenMath Standard, Version 2.0 Revision 1","publisher":"The OpenMath Society","url":"https://openmath.org/standard/om20-2017-07-22/omstd20.html","source_class":"STANDARD","publication_date":"2017-06","accessed_at":"2026-08-03","claims_supported":["OpenMath is an approved standard for representing and communicating mathematical objects.","The standard encodes mathematical meaning rather than only visual presentation.","It defines abstract objects, machine-readable encodings, content dictionaries, and compliance requirements usable in an interchange layer."]},{"source_id":"S5","title":"The MMT Language","publisher":"UniFormal/MMT Project","url":"https://uniformal.github.io/doc/language/index","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["MMT explicitly targets foundation independence, scalability, and modularity.","MMT supplies foundation- and logic-independent semantics.","Theory morphisms and module operations support translation and combination of formal theories."]},{"source_id":"S6","title":"Project-Team DEDUCTEAM: Overall Objectives","publisher":"Inria","url":"https://radar.inria.fr/rapportsactivite/RA2019/deducteam/uid3.html","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2020","accessed_at":"2026-08-03","claims_supported":["An identifiable public research organization has an explicit program to achieve proof-system interoperability and system-independent proof libraries.","The program develops import, cross-theory translation, export, consistency-analysis, and proof-development tools around Dedukti.","Logipedia provides an implemented multi-theory proof encyclopedia, demonstrating adjacent shared-library practice."]},{"source_id":"S7","title":"Porting wiki","publisher":"Lean prover community / GitHub","url":"https://github.com/leanprover-community/mathlib4/wiki/Porting-wiki","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2023-07-16","accessed_at":"2026-08-03","claims_supported":["Even the Lean 3-to-Lean 4 migration required automated conversion, manual cleanup, builds, linting, documentation changes, review, and many contributors.","The documented porting workflow supports the plausibility of substantial migration and maintenance burdens even within one prover family.","Versioned repositories, automated builds, linting, status tracking, and review are available implementation mechanisms for a shadow trial."]},{"source_id":"S8","title":"The Mathlib Initiative","publisher":"Mathlib Initiative / Renaissance Philanthropy","url":"https://mathlib-initiative.org/","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2025-07-24","accessed_at":"2026-08-03","claims_supported":["The Mathlib Initiative is an identifiable funded program supporting a large formal-mathematics library ecosystem.","Its stated priorities include review capacity, dependency coordination, ecosystem health monitoring, documentation, and AI-assisted contribution tools.","The page identifies Renaissance Philanthropy as program operator and acknowledges external financial support."]}],"problem_evidence":{"support":"MODERATE","rationale":"The general problem is visible: current proofs and libraries are fragmented across incompatible systems; cross-system exports are difficult; verified translations may still fail semantic-equivalence checks; and even an intra-family Lean migration required extensive coordinated labor. However, no source demonstrates the proposal's narrower scenario of a consortium currently forced to award one default foundational interface, nor documents benchmark capture, collusive submissions, or winner control of later compatibility rules in such a consortium.","source_ids":["S1","S2","S7"]},"stakeholder_evidence":{"support":"WEAK","rationale":"OpenDreamKit and Inria DEDUCTEAM expressly pursued interoperability, while the Mathlib Initiative funds ecosystem coordination and maintenance. These establish credible organizations, funders, and technical stakeholders. None expresses demand for a competitive default-interface trial, authority over a multi-foundation shared library, willingness to impose semantic gates, or interest in transferring maintenance funds through a challenger process. The proposed adopter remains hypothetical.","source_ids":["S3","S6","S8"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"ITPEval","similarity":"Uses a common benchmark, multiple proof assistants and foundations, native verification, controlled versus ecosystem tiers, and explicit semantic-fidelity checks.","remaining_difference":"It evaluates automated translations; it does not award default-interface status, govern entrant conduct, retain maintenance funding, provide appeals, or reopen an incumbent's designation.","source_ids":["S1"]},{"name":"OpenDreamKit Math-in-the-Middle","similarity":"Creates a shared ontology and joint vocabulary to mediate interoperability among heterogeneous mathematical systems.","remaining_difference":"It is a cooperative hub architecture rather than a bounded rivalry for one default interface, and it does not include winner-power controls or challenger windows.","source_ids":["S3"]},{"name":"OMDoc/MMT","similarity":"Provides foundation-independent representation, theory morphisms, modular semantics, and demonstrated exports from several major proof libraries.","remaining_difference":"It is an interchange representation and logical framework, not a comparative selection-and-governance protocol; the published export experience also reports unresolved translation difficulties.","source_ids":["S2","S5"]},{"name":"Dedukti and Logipedia","similarity":"Implements import, theory translation, export, independent proof checking, and a multi-theory proof encyclopedia.","remaining_difference":"It pursues system-independent pluralism and translation rather than selecting a scarce default or securing a winner's exit and migration obligations.","source_ids":["S6"]},{"name":"OpenMath","similarity":"Standardizes meaning-oriented exchange of mathematical objects through formal encodings and governed content dictionaries.","remaining_difference":"It does not by itself represent whole proof-library governance or resolve whether a neutral interchange layer privileges one foundation; it supplies a potential cooperative baseline rather than the proposed contest.","source_ids":["S4"]}],"distinctive_claim_remaining":"Against neutral-interchange, permanent-pluralism, and independent architecture-board comparators, adding noncompensable semantic gates, equal trial allowances, secured exit duties, separation of interface maintenance from arena rulemaking, and a credible challenger path will reduce unreconciled semantic discrepancies and downstream handoff/exit effort without causing material corpus-dependent ranking instability or governance cost exceeding the preregistered ceiling.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Common corpora, multi-prover verification infrastructure, semantic-equivalence checks, foundation-independent representations, theory morphisms, library exporters, version control, builds, linting, and review workflows all exist. A copied-corpus shadow trial is therefore technically plausible. Feasibility is not strong because cross-foundation proof translation remains low-performing, ecosystem mismatch dominates results, semantic checking can disagree with native verification, and exporters have encountered unresolved theoretical and social problems. No evidence validates neutral corpus construction, enforceable equal-resource allowances, reliable anonymization, exit-package adequacy, appeal operations, or a live default migration. Legal risk is limited for copied artifacts under compatible licenses, but licenses and contributor permissions would require corpus-specific review. Authority is adequate only for an organization testing artifacts it controls; no source establishes authority for the hypothesized cross-foundation consortium.","source_ids":["S1","S2","S3","S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Interoperability failures can strand reusable formal proofs and impose migration work, so a successful intervention could matter. The scale and incidence of actual forced default-interface decisions are unmeasured.","source_ids":["S1","S2","S7"]},"stakeholder_pull":{"score":2,"rationale":"Credible organizations fund library maintenance and explicitly pursue interoperability, but no organization requests this competitive governance mechanism or commits artifacts, staff, or funds to it.","source_ids":["S3","S6","S8"]},"incremental_advantage":{"score":2,"rationale":"Semantic gates, exit security, rulemaking separation, and reopening add safeguards absent from technical benchmarks, but there is no evidence they outperform cooperative interchange, pluralism, or expert selection after accounting for overhead.","source_ids":["S1","S3","S4","S5","S6"]},"distinctiveness_plausibility":{"score":3,"rationale":"The governance bundle was not found in the searched formal-mathematics prior art, although almost every technical component has close analogues. World novelty remains unmeasured.","source_ids":["S1","S3","S4","S5","S6"]},"technical_implementability":{"score":3,"rationale":"A small shadow trial can reuse extant standards, exporters, verification backends, and repository workflows, but semantic neutrality and cross-foundation reconstruction are difficult and partly unresolved.","source_ids":["S1","S2","S4","S5","S6","S7"]},"adoption_authority_feasibility":{"score":2,"rationale":"Existing library programs can authorize tests on their own copied artifacts, but no identified body controls the proposed shared library or has authority to designate a cross-foundation default.","source_ids":["S3","S6","S8"]},"evidence_readiness":{"score":3,"rationale":"A bounded common-corpus test is immediately specifiable and comparable with existing benchmark methodology, but adopter commitment, neutral cases, labor data, and governance outcomes require new empirical work.","source_ids":["S1","S2","S7"]},"safety_net_benefit":{"score":3,"rationale":"A shadow-only start, semantic stop gate, retained award, export package, rollback, and challenger window plausibly limit lock-in and migration harm, but these safeguards have not been tested in this setting.","source_ids":["S1","S7"]},"scalability":{"score":2,"rationale":"Existing projects have handled multiple systems and large libraries, but ecosystem-level translation performs poorly and major exports report substantial technical and social challenges. Governance and audit labor may grow rapidly with foundations, corpus size, and translations.","source_ids":["S1","S2","S6","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Pre-registration and execution of one nonproduction 24-artifact shadow trial with two or three candidate interfaces, separate translators and semantic auditors, corpus rotation, maintenance handoff, and a short governance-cost report.","confidence":"LOW","assumptions":["Approximately 4-8 person-months of specialist formalization, auditing, coordination, and legal-license review.","Existing prover infrastructure and candidate exporters are reused.","No live migration, cash prize, new prover implementation, or production service is included.","The sources establish labor-intensive workflows but do not publish directly comparable prices."],"source_ids":["S1","S2","S7"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Build a reusable trial harness, neutral-interchange profile, expanded corpus, independent review process, adapter and exit-package requirements, audit logging, and documented governance for a consenting consortium.","confidence":"LOW","assumptions":["Roughly 2-5 specialist full-time-equivalent years distributed across engineering, formal semantics, program management, and review.","Three candidate interfaces and at least two independent verification paths are supported.","Existing OpenMath, MMT, Dedukti, or native-prover infrastructure can be adapted rather than rebuilt.","Live corpus conversion and long-term maintenance allocation are excluded."],"source_ids":["S2","S4","S5","S6","S7"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Run a production-facing selection period, validate full-scale adapters and exit materials, migrate a bounded initial corpus, operate appeals and conflict screening, and fund parallel rollback readiness.","confidence":"LOW","assumptions":["A medium-sized consortium library rather than all existing formal mathematics is in scope.","Production launch requires sustained work by candidate teams, auditors, migration maintainers, governance staff, and contributor-support personnel.","The range excludes compensation for pre-existing candidate-library development and any consortium-wide replacement of proof assistants.","Semantic discrepancies can halt launch and materially increase cost."],"source_ids":["S1","S2","S7","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain adapters and interchange specifications, audit sampled migrations, monitor governance concentration, support contributors, preserve exit readiness, and administer periodic challenger and impact reviews.","confidence":"LOW","assumptions":["Approximately 2-5 ongoing specialist FTE equivalents plus compute, review, and documentation support.","A challenger round is periodic rather than continuous.","The consortium already operates source hosting, CI, and proof-checking infrastructure.","No reliable external cost benchmark for this exact governance model was found."],"source_ids":["S6","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent research directly documents incompatible proof ecosystems, poor translation performance, semantic-fidelity failures, unresolved export challenges, and material migration work.","source_ids":["S1","S2","S7"]},"externally_credible_adopter_or_authorizer":{"status":"NO","reason":"Credible funders and interoperability programs exist, but none is shown to control the hypothesized shared library or to want a competitive default-interface designation with retained funding and challenger governance.","source_ids":["S3","S6","S8"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The governance bundle can be compared with a neutral interchange workflow, permanent pluralism, an architecture-board decision, and the same benchmark without secured exit or reopening; semantic discrepancies, handoff effort, ranking stability, and governance hours are observable.","source_ids":["S1","S3","S4","S5"]},"bounded_next_evidence_step":{"status":"YES","reason":"A 24-artifact copied-corpus shadow trial with frozen rules, independent reconstruction, corpus rotation, explicit comparators, and prespecified stop conditions is bounded and avoids production changes.","source_ids":["S1","S2","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"For the shadow step only, a participating library can authorize use of copied, licensed artifacts and anonymized internal results; the proposal excludes live migration, reputational ranking, surveillance, and compulsory adoption. Live deployment would require new authority and license review.","source_ids":["S7","S8"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"Scopes and staffing assumptions can be bounded, but no direct cost evidence exists for a governed cross-foundation trial, and translation difficulty varies sharply by corpus and ecosystem.","source_ids":["S1","S2","S7"]}},"next_evidence_step":"Secure one consenting library organization and two genuinely different interface teams, then pre-register a six-week shadow study on 24 copied, nonproduction artifacts stratified across definitions, theorem statements, dependencies, notation overloading, and foundation-sensitive constructions. Compare: (A) the proposed governed trial; (B) a jointly governed neutral-interchange workflow with no winner; and (C) blinded independent architecture-board selection using the same evidence. Freeze public and withheld subsets, equivalence judgments, allowed assistance, reviewer-hour ceilings, labor logging, appeal rules, and a maximum $250,000 resource-equivalent budget. Separate translation, native checking, semantic reconstruction, and maintenance-handoff teams. Rotate one-third of withheld artifacts and report semantic discrepancies, unsupported constructions, ranking changes, translator/auditor/maintainer hours, usable round-trip exports, governance hours, and participant willingness to continue. Falsify or stop the proposed mechanism if any selected or passing interface has an unreconciled semantic discrepancy; independent auditors cannot reproduce gate decisions; a reasonable corpus rotation changes the winner or pass set; the retained-duty simulation fails to deliver a usable exit package; the governed arm does not reduce semantic or handoff failures versus both comparators; governance effort exceeds the preregistered ceiling; or no credible organization is willing to authorize a later bounded pilot.","blocking_evidence":["No current consortium with authority over a shared cross-foundation library and maintenance allocation has been identified.","No adopter has expressed demand for competitive default-interface selection rather than cooperative interchange or pluralism.","No empirical evidence shows that the governance bundle improves semantic fidelity, handoff effort, or future contestability over simpler comparators.","Neutrality and independent applicability of the proposed semantic gates and interchange specification are unvalidated.","Corpus sensitivity, equal-resource enforceability, anonymization, exit-package usability, and appeal workload remain untested.","Corpus-specific copyright, contributor-license, attribution, and data-governance review is absent.","No direct cost observations exist for the proposed shadow or production workflow."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"World novelty, patentability, freedom to operate, market size, and realized impact were not measured; the search establishes only that close technical interoperability systems and benchmarks exist while the specific governance bundle was not found in the eight reviewed sources.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain written interest and shadow-trial authority from an organization controlling a relevant formal library and maintenance resources.","Recruit at least two foundationally distinct interface teams and independent semantic auditors.","Run the preregistered three-arm shadow study and publish artifact-level semantic, labor, ranking-stability, and governance results.","Demonstrate independently reproducible semantic gates that do not assume one entrant's foundation.","Show that a usable round-trip export and maintenance exit package can be delivered within the retained-duty and budget ceilings.","Complete corpus-license, attribution, conflict-of-interest, appeal, and data-governance review before any live pilot.","Replace resource-equivalent estimates with observed person-hours and infrastructure costs from the shadow trial."],"reason":"Bounded web research verifies the broad interoperability problem and substantial adjacent prior art, but it cannot establish adopter willingness, neutral semantic comparability, comparative workflow advantage, exit-package usability, or realistic operating cost. Those questions require a consenting partner and live execution of the nonproduction shadow trial; under the controller rule this is an empirical-research stop, not a request for more web search."},"proposal_index":2}