{"schema_version":1,"research_id":"eoa_inverse_innovation_exp04_external_evaluation_20260802","source_assessment_id":"computability_boundary_mapping__public_administration_policy:PROPOSAL_FIRST:v0","cell_id":"computability_boundary_mapping__public_administration_policy","search_queries":["site:gao.gov benefits eligibility systems automated rules errors Medicaid SNAP report","site:gov.uk rules as code public benefits policy automation guidance","site:oecd.org rules as code public administration report","site:nist.gov AI risk management automated eligibility public benefits government","official government rules as code pilot eligibility benefits Better Rules New Zealand","OpenFisca official social benefits rules engine test legislation","Catala language law social benefits formal verification paper","public benefits automated eligibility due process algorithm official guidance US","site:hhs.gov \"public benefits\" \"AI\" guidance benefits administrators 2026","site:cms.gov automated renewal systems glitch children Medicaid 2023 official letter states","site:gao.gov automated eligibility systems benefits due process errors computer system","site:congress.gov automated decision systems public benefits due process","\"Plan for Promoting Responsible Use\" HHS public benefits date","site:hhs.gov \"public-benefits-and-ai.pdf\" 2024","Catala programming language law PLDI 2021 official PDF","formal verification policy rules public benefits model checking","HHS public benefits AI plan April 2024 published date state local tribal territorial","\"Public Benefits and AI\" HHS May 2024","site:hhs.gov/news \"public benefits\" algorithmic systems plan"],"sources":[{"source_id":"S1","title":"CMS Takes Action to Protect Health Care Coverage for Children and Families","publisher":"Centers for Medicare & Medicaid Services","url":"https://www.cms.gov/newsroom/press-releases/cms-takes-action-protect-health-care-coverage-children-families","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2023-08-30","accessed_at":"2026-08-02","claims_supported":["CMS identified incorrectly programmed state eligibility systems that could improperly disenroll eligible people, especially children.","CMS required affected states to evaluate their systems, pause disenrollments, reinstate coverage, mitigate harm, and correct systems and processes.","CMS is an identifiable oversight authority able to require corrective action and potentially withhold enhanced funding."]},{"source_id":"S2","title":"Medicaid Eligibility: Accuracy of Determinations and Efforts to Recoup Federal Funds Due to Errors","publisher":"U.S. Government Accountability Office","url":"https://www.gao.gov/products/gao-20-157","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2020-01-13","accessed_at":"2026-08-02","claims_supported":["GAO found eligibility-accuracy issues across 47 federal and state audits covering 21 states.","Observed issues included incomplete information, untimely redeterminations, unresolved discrepancies, and enrollment under an incorrect eligibility basis.","Medicaid eligibility accuracy materially affects beneficiaries and large federal and state expenditures."]},{"source_id":"S3","title":"Plan for Promoting Responsible Use of Artificial Intelligence in Automated and Algorithmic Systems by State, Local, Tribal, and Territorial Governments in Public Benefit Administration","publisher":"U.S. Department of Health and Human Services","url":"https://www.hhs.gov/sites/default/files/public-benefits-and-ai.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2024-04-29","accessed_at":"2026-08-02","claims_supported":["HHS identifies state, local, Tribal, and territorial public-benefit administrators and their vendors as intended actors.","HHS calls for evaluation of automated systems, detection of unjust denials, human appeal, retained expert discretion, inventories, validation, and disclosure of out-of-scope uses.","HHS recommends controlled sandboxes or limited pilots before expensive or lengthy commitments.","The recommendations are general and nonmandatory, and HHS acknowledges budget, staffing, data, and legacy-technology constraints."]},{"source_id":"S4","title":"Cracking the Code: Rulemaking for Humans and Machines","publisher":"OECD Publishing","url":"https://www.oecd.org/en/publications/cracking-the-code_3afe6ba5-en.html","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2020-10-12","accessed_at":"2026-08-02","claims_supported":["Rules as Code proposes official machine-consumable versions of government rules.","Public-sector teams internationally were already exploring the approach in response to complexity and pressure on rulemaking systems.","The report addresses both potential and limitations, showing that machine-consumable policy is an established research and practice area."]},{"source_id":"S5","title":"Better Rules for Government Discovery Report","publisher":"New Zealand Department of Internal Affairs, Digital Government","url":"https://www.digital.govt.nz/dmsdocument/95-better-rules-for-government-discovery-report","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2018-03","accessed_at":"2026-08-02","claims_supported":["A New Zealand government multidisciplinary team examined human- and machine-consumable rules and recognized that not all rules are suitable for machine consumption.","The government piloted a scoped cross-agency eligibility engine covering roughly 18 benefits and tax credits and used ambiguous as a distinct outcome.","The report emphasizes legal traceability, multidisciplinary sign-off, iterative testing, and preserving legal meaning between statute and code.","The discovery itself took three weeks, providing a rough resource analogue for a bounded preflight exercise."]},{"source_id":"S6","title":"OpenFisca Documentation: Introduction","publisher":"OpenFisca","url":"https://openfisca.readthedocs.io/en/latest/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2018 (continuously updated documentation)","accessed_at":"2026-08-02","claims_supported":["OpenFisca is an existing open platform that transforms legislation into executable tax-and-benefit calculations.","It supports test cases, population simulation, country packages, APIs, and benefit-entitlement products.","It is a close functional analogue for executable benefit rules but does not document a universal termination or guarantee-typed preflight contract on the opened page."]},{"source_id":"S7","title":"Catala: A Programming Language for the Law","publisher":"Merigoux, Chataing, and Protzenko; arXiv preprint, subsequently ACM PACMPL","url":"https://arxiv.org/abs/2103.03198","source_class":"PRIMARY_RESEARCH","publication_date":"2021-03-04","accessed_at":"2026-08-02","claims_supported":["Catala is an implemented domain-specific language for systematically translating computational statutes into executable specifications.","The authors formally verified core compiler transformations and evaluated the language on U.S. tax law and French family benefits.","The work found a discrepancy in an official benefits implementation, demonstrating both feasibility and the importance of formalization.","The paper limits the approach to computational parts of law and identifies human judgment and classification as outside the general automated scope.","Formalization required interdisciplinary effort, including four two-hour lawyer-programmer sessions for about 15 percent of one tax-code section and a 1,500-line family-benefits implementation."]},{"source_id":"S8","title":"NIST IR 8539: Security Property Verification by Transition Model","publisher":"National Institute of Standards and Technology","url":"https://csrc.nist.gov/pubs/ir/8539/final","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-01-31","accessed_at":"2026-08-02","claims_supported":["NIST documents converting policy logic into transition models and using model checking to verify formally specified properties.","The report explains automata, temporal-logic specifications, and available verification tools.","It demonstrates technical feasibility for model-relative policy verification, while its application is access-control security rather than public-benefit eligibility."]}],"problem_evidence":{"support":"MODERATE","rationale":"The consequential broad problem is visible: CMS found incorrectly programmed eligibility systems causing potential unlawful disenrollment, and GAO found recurring eligibility-accuracy issues across many audits. HHS treats benefit-administration automation as rights-impacting and calls for validation, out-of-scope disclosure, human review, and detection of unjust denials. However, no opened source documents an agency promising an exact terminating checker for arbitrary rule packages or routinely translating timeout directly into FAIL. The specific computability-failure mechanism therefore remains unverified.","source_ids":["S1","S2","S3","S5","S7"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"CMS is an identifiable authorizer and funder with demonstrated power to require state eligibility-system evaluation and remediation. HHS explicitly addresses public-benefit administrators and vendors and recommends evaluation, governance, inventories, disclosure, and sandbox trials. New Zealand agencies have conducted a scoped cross-agency eligibility-engine pilot. These establish credible institutional pull for safer rule automation, but none expresses demand for this exact computability-boundary record or guarantee-typed router.","source_ids":["S1","S3","S5"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"HHS public-benefits automated-system risk governance and sandboxing","similarity":"Calls for pre-acquisition evaluation, validation, inventories, out-of-scope disclosure, human oversight, appeal, and controlled pilots for rights-impacting benefits systems.","remaining_difference":"It is general risk guidance, not a computability classification, formal totality proof, decidable-fragment gate, or mandated output alphabet separating timeout from semantic failure.","source_ids":["S3"]},{"name":"New Zealand Better Rules and scoped cross-agency eligibility engine","similarity":"Combines machine-consumable benefit rules, legal traceability, multidisciplinary review, limited scope, testing, and an ambiguous outcome.","remaining_difference":"The report does not establish a per-package decidability proof obligation, independently checked computability boundary, or exact-versus-bounded guarantee router.","source_ids":["S5"]},{"name":"OpenFisca and Catala computational-law toolchains","similarity":"Both encode tax or benefit legislation as executable rules; Catala adds formalized compiler correctness, lawyer-programmer collaboration, and demonstrated defect discovery.","remaining_difference":"Neither opened source documents the candidate's organizational release preflight that classifies every submitted package and binds SAFE, UNKNOWN, OUT_OF_SCOPE, TIMEOUT, and SYSTEM_FAILURE to versioned guarantee records.","source_ids":["S6","S7"]},{"name":"NIST transition-model policy verification","similarity":"Provides an official, technically concrete model-checking approach for converting policies into automata and checking temporal properties.","remaining_difference":"It addresses access-control security properties and does not solve legal-semantic fidelity, unrestricted eligibility-rule computability, benefit-workflow authority, or downstream handling of unknown results.","source_ids":["S8"]}],"distinctive_claim_remaining":"For a frozen executable benefit-rule package, adding an enforceable fragment-membership check plus guarantee-typed routing and a versioned boundary record will, relative to conventional testing/timeouts, standalone model checking, and legal review alone, reduce false Boolean release recommendations and make every result traceable to its formal scope without allowing UNKNOWN or OUT_OF_SCOPE to affect an applicant. This is contrastive and falsifiable, but no comparative outcome evidence was found.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"A synthetic bounded audit is technically credible: New Zealand completed a short multidisciplinary discovery and scoped eligibility prototype; OpenFisca executes benefit rules; Catala demonstrates legal-rule formalization and verified compilation; and NIST documents model checking of transition-policy models. HHS explicitly recommends sandboxes and limited pilots. Production feasibility is unresolved because the actual rule grammar, external-service semantics, data interfaces, legal interpretation, proof obligations, state-space size, integration points, and staff capacity are unavailable. Formal verification can certify the wrong model, while necessary discretionary rules may not be computable without accountable human judgment.","source_ids":["S3","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Eligibility-system defects can improperly terminate health coverage, and eligibility accuracy affects beneficiaries and substantial public expenditure. The candidate could prevent a narrow but serious class of assurance errors; realized impact is unmeasured.","source_ids":["S1","S2"]},"stakeholder_pull":{"score":3,"rationale":"HHS, CMS, and New Zealand public agencies demonstrate demand for safer, testable benefits automation, but no named agency has requested or committed to this exact preflight.","source_ids":["S1","S3","S5"]},"incremental_advantage":{"score":2,"rationale":"The proposed separation of timeout, unknown, and out-of-scope is logically useful, but no comparative evidence shows an advantage over a well-governed Rules-as-Code pipeline combining validation, model checking, and legal review.","source_ids":["S3","S5","S7","S8"]},"distinctiveness_plausibility":{"score":2,"rationale":"Most components already appear across HHS governance, Better Rules, OpenFisca, Catala, and model checking. The remaining distinction is their particular computability-centered integration and release contract, not a new underlying method.","source_ids":["S3","S4","S5","S6","S7","S8"]},"technical_implementability":{"score":3,"rationale":"A small synthetic audit is implementable with existing formal-language and model-checking techniques. Production-wide totality, semantic fidelity, state-space control, and external-service modeling remain uncertain.","source_ids":["S5","S6","S7","S8"]},"adoption_authority_feasibility":{"score":3,"rationale":"CMS and HHS provide credible oversight and funding relationships, and program executives can plausibly authorize a nonproduction audit. Exact state or agency authority, procurement rules, counsel concurrence, and labor obligations must be established locally.","source_ids":["S1","S3"]},"evidence_readiness":{"score":3,"rationale":"The candidate specifies a bounded synthetic exercise and measurable output states, and close technical precedents exist. It lacks an agency package, adjudicated semantic oracle, baseline measurements, and pre-registered acceptance thresholds.","source_ids":["S3","S5","S7","S8"]},"safety_net_benefit":{"score":4,"rationale":"Explicit UNKNOWN, OUT_OF_SCOPE, and human-review routing could prevent unresolved analysis from directly denying or delaying benefits. The benefit depends on downstream systems preserving the labels and on review capacity.","source_ids":["S1","S3","S5"]},"scalability":{"score":2,"rationale":"Reusable grammars and tooling could scale within one stable rule family, but Catala and Better Rules show substantial interdisciplinary formalization effort; legal changes, external dependencies, and state-space growth impose recurring costs.","source_ids":["S5","S7","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Two-week paper-and-sandbox audit of one frozen package using synthetic records, including part-time policy counsel, one formal-methods engineer, one rule-engine engineer, an independent reviewer, and a short decision record.","confidence":"MODERATE","assumptions":["No production integration, personal data, procurement, or live eligibility action.","Existing open-source tools and agency development environments are reused.","Approximately 2-3 full-time-equivalent person-weeks plus limited lawyer and reviewer time.","The three-week Better Rules discovery and documented Catala lawyer-programmer sessions are resource analogues, not price quotations."],"source_ids":["S3","S5","S7"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Create and validate an enforceable fragment and checker for one benefit program, formalize selected procedural invariants, implement typed outputs and audit records, and conduct independent technical and legal review.","confidence":"LOW","assumptions":["A multidisciplinary 4-7 person team works for roughly 3-6 months.","Existing rule-engine interfaces are documented and accessible.","Scope is one program and excludes wholesale replacement of its eligibility platform.","No proprietary tool licensing or major data remediation is required."],"source_ids":["S3","S5","S7","S8"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Production hardening and integration for one agency program, including security and privacy review, workflow changes, staff training, monitoring, rollback, accessibility, procurement, and independent assurance.","confidence":"LOW","assumptions":["Launch integrates with an existing eligibility platform rather than replacing it.","Applicant-facing determinations remain governed by existing program law and appeal processes.","External-service failure modes and data contracts require integration testing.","HHS notes that implementation may require substantial changes amid legacy-system and resource constraints."],"source_ids":["S1","S3","S5"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain rule-language classifications, models, proofs, test suites, dependency contracts, independent reviews, training, monitoring, and version-linked records for one program.","confidence":"LOW","assumptions":["A small permanent technical-policy team supports the service.","Every material rule-language or dependency change triggers reclassification.","Human-review workload remains bounded; large UNKNOWN queues would raise cost substantially.","Infrastructure cost is secondary to specialized personnel and governance effort."],"source_ids":["S3","S5","S7"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official CMS and GAO evidence verifies consequential eligibility-system and determination errors, while HHS verifies that automated benefit administration creates rights and safety risks. The narrower prevalence of computability-specific misclassification remains unknown.","source_ids":["S1","S2","S3"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"CMS is a credible oversight authorizer and funder, HHS explicitly addresses state and local benefit administrators and vendors, and New Zealand agencies have piloted a scoped eligibility engine. Exact local sponsorship is not yet secured.","source_ids":["S1","S3","S5"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The candidate can be compared against testing/timeouts, standalone model checking, and expert review on false PASS/FAIL, correct abstention, traceability, review burden, and time-to-verdict.","source_ids":["S3","S5","S7","S8"]},"bounded_next_evidence_step":{"status":"YES","reason":"A two-week, synthetic, nonproduction audit of one frozen package is bounded and consistent with HHS sandbox guidance and prior short government discovery work.","source_ids":["S3","S5"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The proposed step can avoid personal data and live-case effects, preserve existing eligibility and appeal authority, and stop if labels enter an operational workflow. Agency permission and counsel participation remain prerequisites, not reasons to expose applicants.","source_ids":["S1","S3"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"Scopes and broad labor-equivalent bands are stated, and prior work shows substantial interdisciplinary effort, but no agency staffing rates, vendor quotations, integration inventory, or production architecture were available.","source_ids":["S3","S5","S7"]}},"next_evidence_step":"With a named public-benefit agency partner, preregister a two-week retrospective sandbox study on one frozen, production-representative rule package and synthetic or fully de-identified inputs. Inventory its grammar, recursion or iteration, external calls, history bounds, current output semantics, and downstream consumers. Define one procedural invariant and create 20-30 adjudicated cases or mutations spanning ordinary success, genuine counterexample, nonterminating loop, external-service delay or refusal, malformed input, out-of-fragment construct, and resource exhaustion. Compare: (A) the existing schema/unit/regression/timeout release process, (B) standalone finite-state model checking, (C) lawyer-plus-caseworker review, and (D) the candidate's fragment check, model-relative analysis, guarantee-typed router, and versioned record. Measure false PASS, false FAIL, correct UNKNOWN/OUT_OF_SCOPE classification, semantic mismatches found by independent counsel, percentage of outputs retaining scope metadata downstream, analysis time, reviewer hours, and false-alarm burden. Falsify the intervention if the existing process already has a checked total procedure and distinct governed outputs; fragment membership cannot be enforced; any feasible behavior is omitted by the abstraction; bounded coverage is overstated; UNKNOWN reaches release approval; or option D produces no material reduction in invalid Boolean recommendations or traceability failures versus A-C. No result may affect a live eligibility determination.","blocking_evidence":["No direct source showed that an agency currently promises an exact terminating verifier for arbitrary executable benefit rules or collapses timeout into a semantic FAIL; the candidate's precise problem prevalence is unknown.","No named agency has supplied its rule grammar, frozen package, execution logs, external-service contract, current assurance language, or downstream output handling.","No comparative evidence shows that computability classification and guarantee typing outperform a mature Rules-as-Code, model-checking, and legal-review workflow.","No independent legal-semantic oracle or checked totality/impossibility proof exists for a target benefit-rule language.","No evidence establishes acceptable fragment coverage, UNKNOWN rate, false-alarm burden, reviewer capacity, or production scalability.","Cost estimates lack agency wage data, vendor quotations, architecture discovery, and observed maintenance effort.","Local statutory, procurement, privacy, cybersecurity, accessibility, labor, appeal, and records-management authority must be mapped before production use."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"World novelty, patentability, freedom to operate, market size, and realized impact were not measured. The search establishes substantial collision with Rules as Code, executable tax-and-benefit platforms, computational-law languages, model checking, risk inventories, sandbox pilots, human-review safeguards, and legal traceability. It does not establish whether the exact integrated computability-boundary preflight and guarantee-typed release contract has or has not been implemented anywhere.","arm":"PROPOSAL_FIRST","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":true,"material_progress_observed":true,"progress_targets":["Secure a named agency sponsor, program owner, CIO or delegated technical authority, counsel, and independent reviewer for a synthetic nonproduction study.","Obtain a frozen production-representative rule package, enforceable grammar, dependency inventory, current test suite, timeout behavior, and downstream output contract.","Document whether any unrestricted total-verification claim is actually made and estimate how often timeout, unknown, malformed input, system failure, and semantic false are currently conflated.","Preregister the contrastive claim, comparators, adjudicated cases, metrics, acceptance thresholds, and falsifiers before executing the audit.","Produce an independently reviewed formal statement of the model, quantifiers, property, fragment membership rule, termination argument, abstraction soundness obligations, and limits.","Collect observed staff hours, integration tasks, reviewer burden, UNKNOWN rate, false-alarm rate, and vendor or agency cost inputs to replace broad resource bands.","Demonstrate that guarantee labels remain attached through release dashboards and that UNKNOWN, TIMEOUT, or OUT_OF_SCOPE can never authorize or block a live benefit action."],"reason":"Open-book research verified a consequential broad problem, credible public authorities, implementable technical ingredients, and substantial prior-art collision. It did not verify the candidate's specific prevalence claim or its incremental advantage. Answering those questions requires access to an actual rule package, workflow, staff, logs, and controlled comparative execution rather than additional bounded web search."}}