{"schema_version":1,"assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","source_experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"computability_boundary_mapping__public_administration_policy","archetype_slug":"computability_boundary_mapping","domain_slug":"public_administration_policy","title":"Enforceable assurance boundaries for executable public policy","opportunity_summary":"Evaluate an offline boundary-and-router design that restricts executable policies to proved-decidable fragments, preserves an explicit UNKNOWN state outside justified guarantees, and compares its labeling and disposition against an unrestricted Boolean analyzer and expert adjudication.","adopter_authorizer":"An agency program executive may authorize the offline pilot, with semantic-scope approval from legal and policy owners and boundary-claim approval from an independent technical reviewer.","scores":{"meaningful_impact":{"score":4,"rationale":"False clearance, false rejection, and opaque non-decisions could affect public entitlements and obligations; however, the frequency and scale of the stated unrestricted-assurance practice are unsupported."},"stakeholder_pull":{"score":3,"rationale":"Agency policy owners, developers, caseworkers, affected parties, and oversight bodies have identifiable interests in reviewable decisions, but the packet contains no demonstrated demand, procurement commitment, or measured operational pain."},"incremental_advantage":{"score":4,"rationale":"Compared with an unrestricted Boolean analyzer or ordinary risk-based review, enforced fragment membership, guarantee-matched routing, and a non-collapsed UNKNOWN result directly target unsupported assurance claims; realized workload and service advantages remain untested."},"distinctiveness_plausibility":{"score":3,"rationale":"The integration of scope-matched proofs, enforced language boundaries, explicit UNKNOWN routing, and administrative rollback is a coherent composition, but prior art is explicitly unsearched and world novelty cannot be inferred."},"technical_implementability":{"score":3,"rationale":"An offline pilot on one policy family is bounded and technically plausible, but faithful semantics, enforceable membership, useful decidable fragments, and deadline-compatible complexity are unresolved."},"adoption_authority_feasibility":{"score":4,"rationale":"The candidate names the program executive, legal and policy owners, and independent technical reviewer, while limiting the first step to non-live analysis; multi-party approval and semantic accountability still add coordination burden."},"evidence_readiness":{"score":3,"rationale":"The packet supplies separate problem and intervention falsifiers, comparison modes, expert adjudication, halt conditions, and rollback, but remains a hypothesis and lacks prespecified measures, denominators, and acceptance thresholds."},"safety_net_benefit":{"score":5,"rationale":"Explicit UNKNOWN and OUT_OF_SCOPE states, prohibitions on consequential action from those states, expert escalation, boundary enforcement, and rollback provide a strong safety net against mislabeled automation."},"scalability":{"score":3,"rationale":"The routing pattern could be reused across policy families, but each language, semantic model, invariant set, legal interpretation, and boundary change may require substantial independent formalization and review."}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Design and execute one offline, non-live policy-family study covering grammar and invariant formalization, a small enforced fragment, bounded-corpus evaluation, expert adjudication, subgroup reporting, and baseline comparison.","confidence":"LOW","assumptions":["A usable bounded corpus and policy documentation are available without major acquisition expense.","The pilot requires a small interdisciplinary team spanning formal methods, rules-engine engineering, policy or legal interpretation, and evaluation.","No production integration or live adjudication is included.","Formalization difficulty and corpus size are not specified in the packet."]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Prepare a limited agency deployment after successful evidence, including production-grade boundary enforcement, analyzer routing, audit logs, reviewer interfaces, validation, security and compliance work, training, and governance procedures.","confidence":"LOW","assumptions":["Deployment remains limited to one policy family and excludes automatic adverse decisions from UNKNOWN, TIMEOUT, or OUT_OF_SCOPE.","Existing rules-engine and review infrastructure can be integrated rather than replaced.","Independent technical and legal review are required.","Procurement, legacy integration, and assurance requirements are unknown."]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Launch an operational service for multiple policy families with hardened infrastructure, integrations, monitoring, appeals and escalation workflows, accessibility, documentation, independent assurance, and change-control processes.","confidence":"LOW","assumptions":["Multiple policy families require separate semantic and invariant work.","Agency-grade reliability, privacy, security, records, and review requirements apply.","Human review capacity must absorb conservative and UNKNOWN outputs.","The number and complexity of programs are unspecified."]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain analyzers, policy semantics, fragment definitions, boundary checks, monitoring, reviewer support, audits, retraining, and revalidation after language, policy, property, or infrastructure changes.","confidence":"LOW","assumptions":["Operation covers a limited portfolio rather than an agency-wide universal service.","Policy and implementation changes trigger recurring independent review.","Human escalation remains material.","Case volume, change frequency, and reviewer workload are unknown."]}},"research_burden":"HIGH","earliest_credible_horizon":"3_TO_12_MONTHS","pipeline_gates":{"recognizable_externally_supportable_problem":{"status":"YES","reason":"The candidate specifies an observable practice—unrestricted assurance claims from finite tests with collapsed timeout and inconclusive states—and identifies concrete risks to entitlements, obligations, and reviewability; prevalence still requires external validation."},"identifiable_adopter_or_authorizer":{"status":"YES","reason":"The agency program executive is identified as pilot authorizer, with legal and policy owners and an independent technical reviewer assigned specific approval roles."},"distinct_testable_incremental_claim":{"status":"YES","reason":"The bounded pilot can test whether enforced boundaries and explicit UNKNOWN routing reduce mislabeled results and unsupported assurance claims relative to the unrestricted Boolean analyzer, while monitoring unresolved consequential cases."},"bounded_next_evidence_step":{"status":"YES","reason":"The sealed candidate authorizes an offline study of one non-live policy family using a bounded corpus, three output modes, expert adjudication, and explicit halt conditions."},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step excludes production denials or sanctions, requires independent review, preserves existing review as rollback, and states concrete halt triggers; possible subgroup exclusion and reviewer overload remain evaluation risks rather than an immediate stop."},"implementation_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The candidate bounds the pilot activities but provides no staffing, corpus size, policy complexity, integration condition, review volume, or procurement basis sufficient to validate even a broad implementation range externally."}},"blocking_evidence":["Whether the agency actually asserts a reusable total-exact guarantee over an extensible policy language rather than a fixed finite class.","Whether the selected policy semantics and invariants faithfully represent legally relevant administrative behavior.","Whether fragment membership can be mechanically enforced without bypass or silent admission of out-of-scope policies.","Whether the restricted fragment retains necessary distinctions, including those important to vulnerable groups.","Whether the design improves labeling and safe disposition without unacceptable reviewer workload, unresolved consequential cases, or service delay.","Whether existing public-sector rules-engine, policy-as-code, formal-verification, or decision-governance approaches already provide the proposed composition."],"next_evidence_step":"Run a preregistered offline study on one non-live policy family: compare the current unrestricted Boolean analyzer with an enforced-fragment router producing exact, conservative, UNKNOWN, and OUT_OF_SCOPE labels against blinded expert adjudication on a fixed corpus; falsify continuation if it does not reduce mislabeled or unsupported assurance outcomes, if semantic disagreement crosses a prespecified halt threshold, or if safer routing materially increases unresolved consequential cases or projected service delay.","research_questions":["Is the claimed service genuinely quantified over an extensible or unbounded program and case class, or is it a fixed finite service with an effective complete enumeration?","What exact grammar, semantics, computation model, quantifiers, legality properties, and service invariants define the assurance claim?","Can a useful decidable fragment be proved and its membership enforced at every admission path?","How often do baseline timeout or inconclusive searches become approvals, failures, or opaque manual exceptions?","What prespecified thresholds should govern mislabeled outputs, semantic disagreement, subgroup effects, reviewer workload, unresolved consequential cases, and service delay?","Does explicit UNKNOWN routing improve safe disposition relative to both the unrestricted analyzer and ordinary risk-based human review?","What prior work already exists across policy-as-code, public-sector rules engines, formal verification, and administrative decision-support governance?","How should language, model, property, or external-capability changes trigger reclassification and independent revalidation?"],"recommendation":"PARTNERED_RESEARCH","uncertainty_constraints":["Closed-book assessment cannot establish problem prevalence, stakeholder demand, market size, realized impact, or distinctiveness.","Prior art is explicitly unsearched, so no novelty claim is supported.","Cost bands are resource-equivalent planning ranges inferred from pilot scope, not observed prices or quotations.","The proposal applies only if the agency seeks reusable total-exact assurance over an extensible class; it does not apply merely because a particular case is difficult.","Undecidability of a formal property would not establish legal invalidity, case-level irresolvability, or correctness of human judgment.","Pilot acceptance thresholds and operational denominators are absent and must be specified before evidence collection."],"closed_book_prior_art_boundary":"No conclusion is made about novelty, prevalence, existing implementations, or superiority to established methods. The sealed packet marks prior art as UNSEARCHED; external research is required to determine whether any element or the overall composition is already established."}