{"schema_version":1,"assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","source_experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"computability_boundary_mapping__political_science","archetype_slug":"computability_boundary_mapping","domain_slug":"political_science","title":"Model-Relative Verification with Explicit Abstention for Executable Public Policies","opportunity_summary":"Assess whether an authority's universal, exact, always-terminating compliance requirement is computable for a frozen executable-policy class, then replace unsupported binary verdicts with an enforceable decidable fragment and labeled UNKNOWN or OUT_OF_SCOPE routing. The opportunity is conditional on an operative class-wide guarantee actually existing and on the formal equal-treatment property faithfully representing the intended institutional claim.","adopter_authorizer":"A legally accountable regulator or election authority, with procurement, legal, policy, and independent formal-verification reviewers; technical reviewers classify evidence but do not authorize policy deployment.","scores":{"meaningful_impact":{"score":4,"rationale":"False clearance or rejection could affect votes, benefits, burdens, and public accountability. The consequence is potentially serious, but the packet provides no evidence about how often the specified universal-verification demand occurs."},"stakeholder_pull":{"score":2,"rationale":"The candidate identifies regulators, election authorities, auditors, policy submitters, and affected groups, but supplies no observed requirements, procurement records, demand signals, or evidence that any institution currently forces binary universal verdicts."},"incremental_advantage":{"score":4,"rationale":"Relative to expanded simulation and adversarial testing, computability classification, checked certificates, enforceable fragments, and explicit UNKNOWN routing directly address unjustified extrapolation and false binary verdicts. Advantage remains conditional on the requirement being genuinely universal and open-ended."},"distinctiveness_plausibility":{"score":2,"rationale":"The composition is coherent and proposal-specific, but prior art is explicitly unsearched. The packet cannot establish whether comparable formal-methods, election-verification, assurance, or abstention practices already exist."},"technical_implementability":{"score":3,"rationale":"A frozen-language offline pilot with synthetic inputs, constructive checking, checked reduction attempts, and UNKNOWN routing is technically bounded. Faithful semantics, proof checking, open-ended state, and correspondence between formal and political equality create substantial implementation uncertainty."},"adoption_authority_feasibility":{"score":4,"rationale":"The accountable regulator or election authority is clearly identified, retains certification authority, and can authorize an offline pilot without delegating live decisions. Procurement, legal interpretation, and cross-disciplinary review could still slow adoption."},"evidence_readiness":{"score":2,"rationale":"The candidate is labeled HYPOTHESIS and provides falsifiers and an authorized experiment, but contains no empirical baseline, operative institutional requirement, validated formalization, pilot comparison, or prior-art assessment."},"safety_net_benefit":{"score":5,"rationale":"Preserving UNKNOWN and OUT_OF_SCOPE, prohibiting timeout-to-verdict conversion, preventing live-rights effects, and reverting to accountable human review provide strong safeguards against overclaiming during evaluation."},"scalability":{"score":3,"rationale":"The boundary-mapping and abstention pattern could transfer across executable policy systems, but each policy language, equality property, legal interpretation, proof obligation, and escalation process may require institution-specific formalization and governance."}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Review one operative requirement and interface, freeze one formal policy language and equal-treatment property, construct finite synthetic cases, attempt both a checker and a checked reduction, and compare explicit UNKNOWN routing with the stated baseline offline.","confidence":"LOW","assumptions":["A single institution supplies requirements and staff interviews without purchasing sensitive individual-level data.","The pilot uses one property and a modest synthetic corpus.","The work requires formal-methods, policy or legal, evaluation, and affected-party review labor.","No pilot result affects live certification or individual rights."]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Build and independently review a production-oriented verifier or boundary classifier, enforce the accepted policy fragment, implement versioned evidence labels and escalation routing, and complete governance, security, documentation, and staff preparation.","confidence":"LOW","assumptions":["Deployment is limited to one authority and one policy system.","Existing submission and case-management systems can be integrated rather than replaced.","Independent proof review and semantic-validity governance are required.","No assurance is offered beyond the frozen formal model."]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Launch within one accountable authority with validated integration, independent assurance, legal and policy review, affected-group governance, monitoring of UNKNOWN outcomes, incident procedures, and a parallel human decision pathway.","confidence":"LOW","assumptions":["The system concerns high-consequence public decisions and therefore requires extensive assurance and coordination.","Launch covers multiple submission cases but not a nationwide multi-authority rollout.","Human escalation capacity and evaluation are funded as part of launch.","Material changes to the language or property trigger renewed validation."]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain formal semantics, proof tooling, guarantee versions, independent review, security, audit records, human escalation, outcome monitoring, and periodic revalidation for one institutional deployment.","confidence":"LOW","assumptions":["Policy language and legal interpretations change periodically.","Specialist formal-methods and governance personnel remain involved.","UNKNOWN cases receive substantive human review rather than automatic disposition.","The estimate excludes expansion to unrelated policy domains or authorities."]}},"research_burden":"HIGH","earliest_credible_horizon":"3_TO_12_MONTHS","pipeline_gates":{"recognizable_externally_supportable_problem":{"status":"UNCERTAIN","reason":"The packet describes recognizable requirement language, interface behavior, and harms, but the problem is explicitly hypothetical and includes no operative requirement showing that an authority demands a class-wide exact terminating verdict."},"identifiable_adopter_or_authorizer":{"status":"YES","reason":"The legally accountable regulator or election authority is explicitly assigned certification and pilot authority, while technical reviewers have a bounded advisory role."},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal can test whether enforceable restrictions and explicit UNKNOWN routing reduce unsupported binary verdicts relative to simulation, finite tests, expert review, and timeout-based handling."},"bounded_next_evidence_step":{"status":"YES","reason":"The authorized first step freezes one language and one property, uses finite synthetic inputs, attempts constructive and impossibility evidence in parallel, exercises UNKNOWN routing, and excludes live-rights effects."},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"For the offline evidence step, certification authority remains with the accountable institution, synthetic data are used, proxy overclaiming is prohibited, and explicit halt conditions address unchecked proofs, coerced UNKNOWN states, and effects on live rights. These controls do not authorize deployment."},"implementation_cost_scope_and_range":{"status":"UNCERTAIN","reason":"A one-authority implementation can be scoped by formalization, tooling, integration, assurance, governance, and escalation work, but the policy language, system scale, assurance standard, integration burden, and case volume are unspecified, making the resource bands provisional."}},"blocking_evidence":["An operative requirement or interface demonstrating a class-wide exact, terminating binary compliance demand rather than explicitly bounded analysis or declared human judgment.","A faithful, independently governed mapping between the formal equal-treatment property and the operative legal or policy claim, including explicit limits and affected-group review.","Independent checking of any constructive algorithm, impossibility reduction, certificate, and computation-model assumptions.","A blinded offline comparison showing that restrictions and UNKNOWN routing reduce false or unsupported binary verdicts relative to the stated baseline.","Evidence that UNKNOWN outcomes, language restrictions, and human escalation do not systematically shift delay, exclusion, or discretion onto particular groups.","Prior-art research sufficient to evaluate whether the proposed mechanism composition is distinct from existing formal-verification and public-sector assurance practices."],"next_evidence_step":"Conduct a time-bounded offline requirements-and-formalization pilot for one authority: compare the operative specification and interface against the problem falsifier, freeze one policy language and one equal-treatment property, and have independent reviewers attempt both a total checker and a checked impossibility reduction. On a preregistered synthetic case set, compare forced-binary baseline handling with explicit UNKNOWN and OUT_OF_SCOPE routing; stop if no universal guarantee exists, semantics are not faithful, proof obligations remain unchecked, or routing does not reduce unsupported binary verdicts.","research_questions":["Does any operative authority actually require an exact, always-terminating binary verdict across an open-ended executable policy class?","Is the submitted policy language finite-state or otherwise already within a decidable fragment, which would falsify the need for a computability-boundary intervention?","Can the formal equal-treatment property be justified as a faithful but explicitly limited representation of the operative legal and policy claim?","Can independent reviewers validate either a total checker for the declared class or a source-to-target impossibility reduction under the stated computation model?","Compared with current testing and review, does explicit abstention reduce unsupported binary verdicts without merely increasing delay or discretionary rejection?","Are UNKNOWN and OUT_OF_SCOPE outcomes distributed unevenly across affected groups or policy types?","What existing verification, assurance, abstention, and election or regulatory auditing approaches are relevant to incremental distinctiveness?","What staffing, integration, assurance, and recurring review burdens arise for one institutional deployment?"],"recommendation":"VALIDATE_PROBLEM_FIRST","uncertainty_constraints":["Problem prevalence and stakeholder demand are unsupported in the sealed packet.","The candidate is conditional on an open-ended executable class and a genuine universal exact terminating requirement.","The formal property may be only a politically inadequate proxy for substantive or legal equality.","Prior art, prevalence, market size, realized impact, and comparative performance are unmeasured.","Cost bands are resource-equivalent planning ranges, not estimates, and depend heavily on language complexity, institutional integration, assurance requirements, and escalation volume.","A theorem about the formal model cannot by itself establish justice, legitimacy, legal compliance, or acceptable real-world effects.","The authorized evidence step is offline only and provides no authority for live approval, rejection, or rights-affecting use."],"closed_book_prior_art_boundary":"Prior art is explicitly UNSEARCHED. This assessment makes no claim about novelty, prevalence, existing products or practices, market size, realized impact, or exact cost; distinctiveness remains provisional pending external research."}