{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"computability_boundary_mapping__economics_finance:P4:v0","cell_id":"computability_boundary_mapping__economics_finance","search_queries":["site:imf.org macroeconomic model governance validation external models uncertainty policy official guidance","site:bankofengland.co.uk model risk management macroeconomic forecasting models validation","site:federalreserve.gov SR 11-7 model risk management validation official","macroeconomic models multiple equilibria solver convergence human selection Dynare documentation","central bank macroeconomic model governance validation policy models uncertainty official report","macroeconomic forecasting model governance central bank independent validation official","policy model transparency documentation central bank model uncertainty external judgment official","IMF macroeconomic model review governance model uncertainty policy analysis","undecidability economic models equilibrium computation Turing complete agent based models paper","computational undecidability general equilibrium dynamic economic models paper","agent-based model undecidable convergence halting problem economics","economic equilibrium computation undecidable arbitrary programs","model cards model reporting intended use limitations versioning primary paper","NIST AI RMF model documentation limitations human oversight traceability official","ISO 42001 AI system impact assessment documentation external dependencies human oversight standard","model passport governance algorithmic systems prior art","site:bls.gov oes software developers economists mathematicians annual mean wage 2025","BLS employer costs employee compensation professional occupations 2025 official","federal salary economist software developer model validation cost pilot labor rates 2026"],"sources":[{"source_id":"S1","title":"Annual Report 2025, Box 7: Macroeconomic modelling in times of uncertainty","publisher":"European Central Bank","url":"https://www.ecb.europa.eu/press/annual-reports-financial-statements/annual/html/ecb.ar2025~b7f898b33d.mt.html","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["The ECB is an identifiable prospective adopter because it operates forecasting and policy models and has an internal macroeconomic-model governance guide.","The ECB expressly seeks robustness, transparency and adaptability and applies documentation, validation, auditability and lifecycle-risk-management guidelines.","Macroeconomic modelling matters to monetary-policy preparation and is being expanded as uncertainty and model complexity increase."]},{"source_id":"S2","title":"The Dynare Reference Manual, version 7.1: The model file","publisher":"Dynare Team","url":"https://www.dynare.org/manual/the-model-file.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["A widely used macroeconomic modelling platform employs iterative equilibrium solvers with configurable iteration bounds.","Steady-state results depend on user-supplied initial guesses, and the manual calls finding good initial values the trickiest part for complicated models.","Dynare documents cases with infinitely many steady states requiring user selection and warns that numerical convergence can be spurious or economically meaningless."]},{"source_id":"S3","title":"No-Regret Learning in Games is Turing Complete","publisher":"arXiv / Cornell University","url":"https://arxiv.org/abs/2202.11871","source_class":"PRIMARY_RESEARCH","publication_date":"2022-02-24","accessed_at":"2026-08-03","claims_supported":["Replicator dynamics on matrix games can simulate arbitrary Turing machines.","Reachability can be undecidable for these dynamics, with equilibrium convergence identified as a special case.","This supports the plausibility of computability boundaries in some expressive multi-agent dynamics but does not itself prove the proposal's unrestricted macro-model claim."]},{"source_id":"S4","title":"Supervisory Guidance on Model Risk Management, SR 26-2 Attachment","publisher":"Board of Governors of the Federal Reserve System, FDIC, and OCC","url":"https://www.federalreserve.gov/supervisionreg/srletters/SR2602a1.pdf","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-04-17","accessed_at":"2026-08-03","claims_supported":["Model risk can cause financial loss, reporting errors and flawed financial or risk-management decisions.","Existing financial-sector practice already includes model inventories, documentation, validation, effective challenge, clear roles, ongoing monitoring and controls over external resources and third-party products.","The revised guidance says validation should identify limitations and errors and that material model risk may remain even after rigorous validation."]},{"source_id":"S5","title":"Model Cards for Model Reporting","publisher":"arXiv / Cornell University","url":"https://arxiv.org/abs/1810.03993","source_class":"PRIMARY_RESEARCH","publication_date":"2019-01-14","accessed_at":"2026-08-03","claims_supported":["Model cards are established prior art for structured, accompanying documentation of intended use, evaluation conditions, limitations and performance.","The model-card concept is presented as applicable beyond its initial computer-vision and language-model examples.","A passport-shaped documentation artifact is therefore not distinctive by itself."]},{"source_id":"S6","title":"Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1","publisher":"National Institute of Standards and Technology","url":"https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf","source_class":"STANDARD","publication_date":"2023-01","accessed_at":"2026-08-03","claims_supported":["Existing governance prior art calls for documenting system knowledge limits, human oversight, third-party software and data, testing, independent review and lifecycle risk tracking.","NIST distinguishes human roles from autonomous system action and warns that mathematical representations can remove necessary context.","The framework supports versioned, traceable governance but does not provide the proposal's oracle-relative computability classification."]},{"source_id":"S7","title":"ECB macroeconometric models for forecasting and policy analysis: Development, current practices and prospective challenges","publisher":"European Central Bank / Publications Office of the European Union","url":"https://op.europa.eu/en/publication-detail/-/publication/6d61e9bb-e5c7-11ee-8b2b-01aa75ed71a1/language-en","source_class":"PRIMARY_RESEARCH","publication_date":"2024-03-18","accessed_at":"2026-08-03","claims_supported":["Central-bank macroeconomic models are used for forecasting, scenario analysis and policy preparation.","The ECB describes these models as stylised representations that can fail to predict or assess major events.","The evidence supports consequential use and acknowledged uncertainty, not the prevalence of hidden oracle-relative convergence claims."]},{"source_id":"S8","title":"Employer Costs for Employee Compensation—March 2026","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/pdf/ecec.pdf","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-06-12","accessed_at":"2026-08-03","claims_supported":["March 2026 employer compensation averaged $96.92 per hour for professional and related occupations in the financial-activities industry.","The compensation data provide a resource-equivalent labor anchor for the broad pilot, startup, launch and recurring-cost estimates.","The source does not estimate the proposal's implementation costs directly."]}],"problem_evidence":{"support":"MODERATE","rationale":"The operational ingredients are visible: central banks use consequential macroeconomic models, acknowledge model uncertainty and governance needs, and Dynare documents initial-value sensitivity, bounded iterations, multiple steady states and potentially spurious numerical convergence. Primary research also proves undecidable equilibrium-convergence phenomena for one expressive game-dynamics class. However, no opened source documents the proposal's specific alleged failure—policy platforms systematically laundering external services or adaptive human branch choices into unconditional universal convergence guarantees—or its prevalence. The general problem matters, but the exact failure mode remains an evidence gap.","source_ids":["S1","S2","S3","S4","S7"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"The ECB's Forecasting and Policy Modelling function is an identifiable adopter, its model-governance function is a plausible authorizer, and the ECB expressly maintains an internal guide emphasizing robustness, transparency, documentation, validation, auditability and lifecycle risk management. Federal regulators independently express a strong need for model inventories, limitations, external-resource oversight and effective challenge. No source expresses demand for oracle-relative computability passports, formal reductions or the proposed status vocabulary specifically.","source_ids":["S1","S4","S7"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"ECB internal guide to governance of macroeconomic models","similarity":"Same institutional setting and model lifecycle; already addresses usability, documentation, validation, auditability, transparency and lifecycle risk management for projection and policy-analysis models.","remaining_difference":"The public description does not state that it classifies answers by base versus external computational capability, enforces solver promises, performs oracle-withdrawal tests, or preserves computability-specific UNKNOWN statuses.","source_ids":["S1"]},{"name":"SR 26-2 model-risk-management framework","similarity":"Established financial-sector practice covering model purpose, limitations, testing, independent challenge, inventories, documentation, external resources, third-party products and ongoing monitoring.","remaining_difference":"It is risk-based governance rather than a formal computability-boundary protocol, and its principal scope is regulated banking organizations rather than central-bank macro-policy models.","source_ids":["S4"]},{"name":"Model Cards","similarity":"A structured companion artifact reports intended use, evaluation context, performance and limitations so downstream users do not overgeneralize model outputs.","remaining_difference":"Model cards do not distinguish base-computable from oracle-relative results or test withdrawal, abstention, promise violations and semi-decision semantics.","source_ids":["S5"]},{"name":"NIST AI RMF 1.0","similarity":"Requires documented knowledge limits, human oversight, third-party components, independent testing, traceability and lifecycle governance.","remaining_difference":"It does not supply model-relative computability proofs, reduction obligations, oracle contracts or the proposal's seven-state result router.","source_ids":["S6"]},{"name":"Dynare solver controls and user-assisted equilibrium selection","similarity":"Current product practice already exposes iteration bounds, initial guesses, convergence criteria and user assistance for multiple or difficult steady states.","remaining_difference":"The documentation exposes solver behavior but does not create a policy-governance passport that propagates capability provenance and non-universal guarantees into decision materials.","source_ids":["S2"]}],"distinctive_claim_remaining":"For executable macroeconomic policy models, adding mechanically preserved BASE_CERTIFIED, ORACLE_RELATIVE, COUNTEREXAMPLE_FOUND, BOUNDED_ONLY, UNKNOWN, OUT_OF_MODEL and SYSTEM_FAILURE labels—derived from named capability contracts, promise checks and withdrawal/abstention tests—will detect and prevent materially more capability-laundering errors than an existing model inventory plus ordinary validation documentation, without changing the intended economic query. This is contrastive and falsifiable, but untested.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"A synthetic offline exercise is technically implementable with rational-valued encodings, stub services, scripted human responses, bounded trace enumeration, signed records and independent review. Existing solver controls, model-documentation practices and governance frameworks supply many components. Data needs are minimal because the proposed first test is synthetic; no confidential feeds or policy instruments are required. A model-governance director could authorize the workflow while the policy committee retains policy authority. Safety risk is bounded if labels cannot enter production materials. The difficult parts remain unvalidated: semantic fidelity of the formal economic query, construction and checking of the proposed undecidability reduction, enforceability of semantic solver promises, downstream status preservation, and the usefulness burden created by UNKNOWN results.","source_ids":["S1","S2","S3","S4","S5","S6"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Incorrectly elevating a conditional numerical result into an unconditional policy-model guarantee could affect consequential analysis, but occurrence and realized harm are not measured.","source_ids":["S1","S4","S7"]},"stakeholder_pull":{"score":3,"rationale":"ECB and financial regulators express strong adjacent governance needs, but no adopter has requested the computability-specific intervention.","source_ids":["S1","S4"]},"incremental_advantage":{"score":3,"rationale":"Capability provenance, withdrawal tests and non-collapsible statuses add a clear function beyond generic documentation; comparative performance against current governance remains untested.","source_ids":["S1","S4","S5","S6"]},"distinctiveness_plausibility":{"score":3,"rationale":"The computability-specific combination appears differentiated from opened analogues, but passports, model inventories and lifecycle governance are established and world novelty was not assessed.","source_ids":["S1","S4","S5","S6"]},"technical_implementability":{"score":4,"rationale":"The offline artifact, stubs, bounded search and routing logic are straightforward; formalization and proof obligations require scarce expertise but do not block a synthetic pilot.","source_ids":["S2","S3","S6"]},"adoption_authority_feasibility":{"score":4,"rationale":"Existing model-governance and independent-validation functions are plausible owners, and an offline pilot does not require policy authority or production access.","source_ids":["S1","S4"]},"evidence_readiness":{"score":3,"rationale":"The problem class and adjacent practices are externally supported, but prevalence, workflow fit and incremental detection advantage require empirical testing.","source_ids":["S1","S2","S3","S4","S7"]},"safety_net_benefit":{"score":4,"rationale":"Explicit UNKNOWN, OUT_OF_MODEL and SYSTEM_FAILURE states plus withdrawal tests directly reduce the risk of treating timeouts, abstentions or human choices as proofs, although downstream effectiveness is untested.","source_ids":["S2","S4","S6"]},"scalability":{"score":3,"rationale":"A common schema and router could scale across models, but each formal query, promise screen and proof classification may require substantial bespoke economic and formal-methods review.","source_ids":["S1","S4","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"One offline six-model exercise: formalize the query, implement seven labels and stub capabilities, run withdrawal/abstention/promise tests, check one witness and review one proposed reduction.","confidence":"MODERATE","assumptions":["Approximately 250-450 total professional hours across a macroeconomist, software/formal-methods engineer, model-risk reviewer and project owner.","Synthetic models and existing internal tooling are used; no procurement, live data or production integration.","The BLS financial-activities professional rate of about $97 per employer-paid hour is a labor anchor; specialist contracting or institutional overhead could push the exercise above this band."],"source_ids":["S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Convert the pilot into a reusable passport schema, validator, capability registry, status-preserving export and review procedure for one modelling team.","confidence":"LOW","assumptions":["Roughly 0.5-1.5 resource-equivalent FTE-years distributed among engineering, macroeconomic modelling, governance and independent review.","Existing repositories, identity controls and model inventory can be extended rather than replaced.","No formal verification of an entire production model language is included."],"source_ids":["S1","S4","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Institutional launch across a bounded portfolio: integrations, migration of priority models, training, independent-review capacity, audit trail, downstream presentation controls and acceptance testing.","confidence":"LOW","assumptions":["Launch covers one institution and a prioritized portfolio, not every historical model or all Eurosystem institutions.","Three to eight resource-equivalent FTE-years plus limited infrastructure and assurance support.","Models with confidential code or semantic promises require extra manual review."],"source_ids":["S1","S4","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Operate the capability registry, review new and changed passports, rerun dependency tests, maintain tooling and audit downstream label preservation.","confidence":"LOW","assumptions":["Two to six resource-equivalent FTEs support a material policy-model portfolio.","Review is risk-tiered and triggered by model, oracle, promise or human-procedure changes.","Compute costs remain modest relative to specialist labor; large exhaustive-search workloads are excluded."],"source_ids":["S1","S4","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Consequential use, model uncertainty, solver sensitivity, multiple equilibria and model-risk consequences are externally supported. The narrower prevalence claim about hidden capability laundering remains unmeasured and must be tested.","source_ids":["S1","S2","S3","S4","S7"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"The ECB operates a forecasting and policy-modelling portfolio and an internal macro-model governance framework; its model-governance or audit function is a credible authorizer for an offline exercise.","source_ids":["S1","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The passport can be compared with an existing model inventory and ordinary validation template on correct dependency classification, preservation of UNKNOWN and prevention of false BASE_CERTIFIED outputs.","source_ids":["S1","S4","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"The six-model synthetic exercise has fixed model cases, rational encoding, ten-step search, explicit comparator, review outputs and falsifiers, with no live policy action.","source_ids":["S2","S3","S4","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"An offline synthetic pilot can be authorized by model governance, leaves policy authority untouched, uses no confidential feed and can be rolled back by withdrawing pilot records. Production or official-forecast use remains excluded.","source_ids":["S1","S4","S6"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four ranges identify concrete scopes and labor assumptions and use current BLS employer-compensation data as a 2026 resource-equivalent anchor. Confidence is low beyond the first exercise because no implementation benchmark or model inventory was available.","source_ids":["S8"]}},"next_evidence_step":"Preregister and run the proposed offline exercise on six synthetic models. Randomize each case between (A) the institution's ordinary model-inventory/validation template plus raw solver logs and (B) the passport workflow. Use blinded reviewers to classify whether the result is internally computed, externally assisted, human-selected, counterexample-bearing, bounded, unknown, out of model or failed. Primary comparators are dependency-classification accuracy, false BASE_CERTIFIED rate, UNKNOWN-preservation rate, detection of promise violations, reviewer time and semantic-fidelity ratings. Require independent checking of the proposed reduction and countertrajectory. Falsify the incremental claim if the passport does not improve classification accuracy, produces any false BASE_CERTIFIED result, loses UNKNOWN downstream, cannot localize dependency under withdrawal, or formalization materially changes the economic question for at least two independent macroeconomists. Halt if an undeclared capability affects an output or any pilot label enters an official forecast or policy comparison.","blocking_evidence":["No direct evidence establishes how often central-bank policy-model outputs currently conceal external-solver or adaptive-human dependence.","No adopter has expressed demand for oracle-relative computability labels or committed staff, authority or funding.","The unrestricted-class reduction has not been constructed or independently checked for the proposed model language and convergence query.","Semantic fidelity between the formal convergence property and the economic question used by policy staff has not been demonstrated.","Comparative evidence against existing ECB-style governance, model inventories and validation templates is absent.","Promise conditions for real external solvers may be semantic or empirical and therefore not mechanically enforceable.","Downstream systems have not been tested for preservation of UNKNOWN and ORACLE_RELATIVE labels.","Implementation hours and portfolio-scale recurring costs have not been observed."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"World novelty, patentability, freedom to operate, market size and realized impact were not measured. The bounded search found established adjacent practices—macroeconomic-model governance, model inventories, model cards, AI risk frameworks and solver diagnostics—but did not establish whether the exact capability-relative passport and withdrawal-test combination exists anywhere.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain primary workflow evidence from at least two central-bank or public-policy model-governance teams on whether external services, human branch choices, timeouts and unresolved branches are currently recorded and preserved downstream.","Run the preregistered six-model comparison and demonstrate higher dependency-classification accuracy than ordinary validation documentation, zero false BASE_CERTIFIED labels and 100% preservation of UNKNOWN.","Produce and independently check the unrestricted-class reduction, including source-to-target direction, total computable mapping, encoding assumptions and exact theorem scope.","Have two independent macroeconomists assess whether each formalized query preserves the intended economic question; treat material disagreement as a falsifier.","Measure actual person-hours, specialist mix and integration effort to replace extrapolated startup and recurring-cost estimates.","Secure a written offline-pilot authorization from a model-governance owner specifying excluded production, forecast and policy uses."],"reason":"Web evidence establishes consequential model use, recognized governance needs, solver ambiguity and adjacent documentation practice, and it leaves a plausible computability-specific distinction. The decisive remaining questions—prevalence, semantic fidelity, comparative workflow advantage, reduction validity and downstream label preservation—require interviews, proprietary workflow observation or live synthetic testing rather than further bounded web search. Under the controller rule, this requires an empirical-research stop, and every STOP is non-repairable."},"proposal_index":4}