{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp04_retrieval_first_paired20_20260802","cell_id":"computability_boundary_mapping__public_administration_policy","round_index":0,"assessments":[{"hypothesis_id":"H1","search_queries":["Rules as Code public benefits eligibility OpenFisca official rule engine","Catala language social benefits legislation decidable termination","New Zealand Better Rules rules as code official","public benefits eligibility rule engine timeout unknown decidable grammar"],"sources":[{"source_id":"H1-S1","title":"What Better rules – better outcomes is all about","publisher":"New Zealand Better Rules initiative","url":"https://www.betterrules.govt.nz/about","source_class":"GOVERNMENT_OR_REGULATOR","claims_supported":["Government rules-as-code practice already translates legislation into concept models, decision models, rule statements, and software code.","Eligibility and benefit calculations are explicit use cases, establishing substantial workflow overlap."]},{"source_id":"H1-S2","title":"Write rules as code","publisher":"OpenFisca","url":"https://openfisca.org/en/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","claims_supported":["OpenFisca models tax and benefit legislation as executable rules and exposes eligibility calculations through APIs.","OpenFisca is already used to inform Barcelona residents about social-benefit eligibility and to analyze interactions among reforms."]},{"source_id":"H1-S3","title":"Catala: A Programming Language for the Law","publisher":"arXiv / Catala research team","url":"https://arxiv.org/abs/2103.03198","source_class":"PRIMARY_RESEARCH","claims_supported":["Catala is a domain-specific language for translating statutes into executable implementations.","The compiler's core compilation steps were formally verified, and the language was evaluated on French family benefits and U.S. tax law."]}],"closest_analogue":"Catala-based computational-law authoring, supplemented by government Rules-as-Code workflows and OpenFisca benefit models.","overlap":"All three analogues translate legislation or benefits policy into executable rules; Catala adds formal verification of compiler transformations, while Better Rules integrates policy, legal, and software authoring.","remaining_difference":"The hypothesis specifically requires a mechanically enforceable terminating fragment and a rule-set-level totality-and-correctness witness before deployment. The opened analogues do not establish that every admitted benefit rule set is guaranteed to terminate, nor do they measure the share of historical constructs excluded by such a fragment. That delta is testable by defining the grammar, mechanically checking membership, proving the evaluator total, and measuring retained constructs.","classification":"POSSIBLE_DISTINCTION","disposition":"ADVANCE","rationale":"The rules-as-code portion is established, but the computability-boundary control—enforced total language plus per-rule deployment witness—remains a concrete, falsifiable distinction rather than an inference from search failure."},{"hypothesis_id":"H2","search_queries":["automated building permit review software unresolved manual review official government","digital building permit code compliance automated plan review human review official","automated code compliance checking \"unknown\" \"manual review\" building permit","Solibri checking results undefined unclassified first party"],"sources":[{"source_id":"H2-S1","title":"Understanding Checking","publisher":"Solibri","url":"https://help.solibri.com/hc/en-us/articles/1500005009042-Understanding-Checking","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","claims_supported":["Solibri performs rule-based model and compliance checking.","Its workflow supports accept, reject, undefined, and unhandled decisions, plus a blocked-error state when preconditions fail.","Results are organized for investigation, reporting, and human disposition."]},{"source_id":"H2-S2","title":"Plan2Permit - AI-Powered Permit Review","publisher":"Plan2Permit","url":"https://plan2permit.com/","source_class":"COMMERCIAL_FIRST_PARTY","claims_supported":["Plan2Permit reviews permit applications against local codes and cites relevant provisions.","Findings are assigned Confirmed, Likely, or Verify confidence tiers.","Licensed plans examiners validate every review, providing an explicit human-review route for unresolved findings."]},{"source_id":"H2-S3","title":"2024 ICC Performance Code for Buildings and Facilities, Chapter 1","publisher":"International Code Council","url":"https://codes.iccsafe.org/content/ICCPC2024P1/chapter-1-general-administrative-provisions","source_class":"STANDARD","claims_supported":["The code official remains responsible for reviewing or ensuring compliance before permit approval.","Performance-based designs may require specialized review rather than automatic acceptance."]}],"closest_analogue":"Plan2Permit's confidence-tiered, citation-backed automated permit review with mandatory examiner validation.","overlap":"The analogue already replaces undifferentiated automated findings with confidence tiers, identifies items requiring verification, supplies code-section evidence, and routes all outputs through accountable human review. Solibri independently implements undefined, unhandled, and blocked states.","remaining_difference":"The proposed experiment uniquely labels the unresolved state as a bounded semi-decision UNKNOWN and evaluates timeout-derived misclassification and review-time effects. That is principally a formal interpretation and evaluation of an already deployed triage pattern, not a distinct operational intervention.","classification":"OBVIOUS_COLLISION","disposition":"REJECT","rationale":"The core output-protocol and escalation mechanism is directly present in current first-party permit-review products; changing the theoretical justification or metric set does not create enough separation for advancement in this screen."},{"hypothesis_id":"H3","search_queries":["model checking emergency response plan formal verification interagency","formal verification emergency plans deadlock model checking","Petri net emergency response plan model checking coordination","government emergency management plan simulation formal model verification"],"sources":[{"source_id":"H3-S1","title":"Petri Net Based Modeling and Correctness Verification of Collaborative Emergency Response Processes","publisher":"Cybernetics and Information Technologies / Eindhoven University of Technology repository","url":"https://pure.tue.nl/ws/portalfiles/portal/101860483/_13144081_Cybernetics_and_Information_Technologies_Petri_Net_Based_Modeling_and_Correctness_Verification_of_Collaborative.pdf","source_class":"PRIMARY_RESEARCH","claims_supported":["The paper formalizes collaborative emergency-response processes with finite Petri nets extended for resources and messages.","It defines and verifies correctness through reachability analysis.","Its case includes an emergency command center, fire brigade, and hospital, with message exchange, shared resources, synchronization, and potential conflicts."]},{"source_id":"H3-S2","title":"Integrated modelling of medical emergency response process for improved coordination and decision support","publisher":"Healthcare Technology Letters / PubMed Central","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC5048341/","source_class":"PRIMARY_RESEARCH","claims_supported":["The research models heterogeneous emergency-response organizations, actors, communications, and data flows.","It uses Petri nets to represent event-driven dynamics and coordination points for decision support."]},{"source_id":"H3-S3","title":"Modeling and Simulation for Emergency Response: Workshop Report, Standards and Tools","publisher":"National Institute of Standards and Technology / GovInfo","url":"https://www.govinfo.gov/content/pkg/GOVPUB-C13-bc62daf0582d06254b25f2535794242d/pdf/GOVPUB-C13-bc62daf0582d06254b25f2535794242d.pdf","source_class":"GOVERNMENT_OR_REGULATOR","claims_supported":["NIST documented emergency-response modeling and simulation as a government preparedness concern.","The report identifies validation, verification, accreditation, interoperability standards, metadata, and audit trails as relevant requirements."]}],"closest_analogue":"Liu and Zhang's Petri-net modeling and formal correctness verification of collaborative interagency emergency-response processes.","overlap":"The primary analogue already formalizes a multi-organization emergency plan, models messages, resources, synchronization, concurrency, and conflict, and applies reachability-based correctness verification before real-world execution.","remaining_difference":"The hypothesis emphasizes a deliberately conservative over-approximation, one-directional safety certification, an explicit hazard envelope, and measured false alarms. Those are refinements to abstraction construction and evaluation; the central plan-level formal-verification intervention has already been demonstrated.","classification":"OBVIOUS_COLLISION","disposition":"REJECT","rationale":"A direct primary-research analogue matches the domain, unit of analysis, interagency coordination mechanism, predeployment timing, and formal verification objective closely enough to eliminate the candidate at a shallow screen."},{"hypothesis_id":"H4","search_queries":["exhaustive enumeration household states policy eligibility benefits cliffs simulation","policy microsimulation all combinations household characteristics eligibility cliffs","government benefits calculator policy interactions household types official","combinatorial testing tax benefit policy rules exhaustive enumeration"],"sources":[{"source_id":"H4-S1","title":"How would reforms affect cliffs?","publisher":"PolicyEngine","url":"https://www.policyengine.org/us/research/how-would-reforms-affect-cliffs","source_class":"OFFICIAL_ORGANIZATION_DATA","claims_supported":["PolicyEngine analyzes benefit cliffs and policy reforms at household and population levels.","Its microsimulation uses Current Population Survey records and stochastic assignments for eligibility attributes that are missing from the survey.","The approach therefore evaluates a representative modeled population rather than certifying every state in a declared finite encoding."]},{"source_id":"H4-S2","title":"Resources to Support System and Logic Testing for Unwinding when the PHE Ends","publisher":"Centers for Medicare & Medicaid Services","url":"https://www.medicaid.gov/resources-for-states/downloads/resources-support-sys-logic-test-unwinding.pdf","source_class":"OFFICIAL_GUIDANCE","claims_supported":["CMS recommends testing selected complex beneficiary scenarios, boundary values, negative inputs, and edge cases in eligibility systems.","The guidance also calls for governance, remediation capacity, manual overrides, and independent verification and validation.","It does not characterize the listed scenarios as an exhaustive census of all encoded combinations."]},{"source_id":"H4-S3","title":"Combinatorial Testing Applied","publisher":"National Institute of Standards and Technology","url":"https://tsapps.nist.gov/publication/get_pdf.cfm?pub_id=913806","source_class":"PRIMARY_RESEARCH","claims_supported":["NIST distinguishes exhaustive coverage over finite equivalence-class combinations from full value-level exhaustiveness.","An exhaustive suite of 36,626 combinations found failures, while reduced t-way suites traded coverage for efficiency.","The source demonstrates both the feasibility of bounded exhaustive enumeration and the need to state precisely what is exhaustive."]},{"source_id":"H4-S4","title":"Benefits calculators","publisher":"UK Government","url":"https://www.gov.uk/benefits-calculators?ac=%2F","source_class":"OFFICIAL_GUIDANCE","claims_supported":["Government-endorsed calculators estimate benefits and effects of changing household circumstances.","The guidance lists populations for which the calculators are inaccurate, illustrating explicit but instance-oriented scope limits rather than exhaustive policy-space coverage."]}],"closest_analogue":"PolicyEngine's household-level tax-benefit microsimulation and benefit-cliff analysis.","overlap":"Existing tools encode interacting tax and benefit rules, vary household circumstances, detect cliffs, and support ex ante reform analysis. Government eligibility-system guidance already promotes boundary, negative, and complex-scenario testing.","remaining_difference":"The proposed unit is the complete, explicitly finite cross-product of encoded states for one jurisdiction, package, and horizon, accompanied by a machine-checkable coverage certificate. The closest analogue instead uses representative survey records and stochastic imputations. The distinction is testable by auditing the state-space generator for completeness, rerunning every encoded state, comparing findings with sampled scenarios, and recording the exact bound and resource cost.","classification":"POSSIBLE_DISTINCTION","disposition":"ADVANCE","rationale":"The substantive policy-analysis use case is established, but exhaustive bounded coverage with a scope certificate differs concretely from both microsimulation and selected test scenarios. Advancement is warranted for deeper prior-art checking and feasibility measurement, not because the shallow search failed to locate an identical deployment."},{"hypothesis_id":"H5","search_queries":["government AI procurement guidelines performance claims evidence fallback monitoring re-evaluation official","public procurement algorithmic systems independent review technical claims official standard","NIST AI procurement contract requirements validity evidence monitoring recheck","government algorithm procurement impact assessment vendor claims universal accuracy"],"sources":[{"source_id":"H5-S1","title":"Guidelines for AI procurement","publisher":"UK Government","url":"https://www.gov.uk/government/publications/guidelines-for-ai-procurement/guidelines-for-ai-procurement","source_class":"OFFICIAL_GUIDANCE","claims_supported":["Public buyers are advised to define acceptable performance, test systems under varied conditions, evaluate suppliers with multidisciplinary expertise, and require accountability and reproducibility.","The guidance calls for explicit evaluation, reassessment, continuous-monitoring, and follow-up timeframes."]},{"source_id":"H5-S2","title":"Responsible AI in Recruitment","publisher":"UK Government","url":"https://www.gov.uk/government/publications/responsible-ai-in-recruitment-guide/responsible-ai-in-recruitment","source_class":"OFFICIAL_GUIDANCE","claims_supported":["The guidance explicitly warns that suppliers make performance, efficiency, fairness, and capability claims during procurement.","It recommends evidence such as impact assessments, risk assessments, model cards, performance tests, intended scope, and documented limitations.","It also recommends repeated assurance, monitoring, human review, and alternative workflows when limitations affect users."]},{"source_id":"H5-S3","title":"Artificial Intelligence: Key Practices to Help Ensure Accountability in Federal Use","publisher":"U.S. Government Accountability Office","url":"https://www.gao.gov/products/gao-23-106811","source_class":"GOVERNMENT_OR_REGULATOR","claims_supported":["GAO's framework organizes federal AI accountability around governance, data, performance, and monitoring.","GAO identifies third-party assessment and audit as important accountability practices."]},{"source_id":"H5-S4","title":"AI strategies and compliance plan","publisher":"U.S. General Services Administration","url":"https://www.gsa.gov/artificial-intelligence/resources/ai-strategies-and-compliance-plan","source_class":"OFFICIAL_GUIDANCE","claims_supported":["GSA requires impact statements, real-world test plans, independent evaluation, ongoing monitoring, human review, and fallback or escalation for high-impact AI.","Written waivers are centrally tracked, published, and reassessed annually, showing a mature record-and-recheck analogue."]}],"closest_analogue":"The UK Government Guidelines for AI Procurement, reinforced by GSA's independent-evaluation, monitoring, fallback, and reassessment controls.","overlap":"Existing public-sector guidance already requires scoped requirements, evidence for vendor capability claims, independent or third-party review, documented limitations, fallback, monitoring, and reassessment after deployment or change.","remaining_difference":"The hypothesis adds a model-relative computability classification before award: it would require the input class and universal quantifiers to be explicit, distinguish decidability from benchmark performance, and demand a constructive totality proof or valid impossibility reduction when such claims are made. The opened procurement frameworks do not prescribe that computability-specific evidentiary gate. Its incremental effect is testable by coding solicitations for unsupported universal guarantees and measuring reviewer agreement and cycle time.","classification":"POSSIBLE_DISTINCTION","disposition":"ADVANCE","rationale":"Most governance components have close official analogues, but none of the opened sources substitutes formal computability evidence for ordinary performance assurance. The narrow gate is therefore distinguishable enough for further review, while the broader procurement-record concept is not novel."}],"nominated_ids":["H1","H4","H5"],"replenishment_recommended":false}