{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp04_retrieval_first_paired20_20260802","cell_id":"computability_boundary_mapping__public_administration_policy","arm":"RETRIEVAL_FIRST","round_index":0,"hypotheses":[{"hypothesis_id":"H1","title":"Decidable policy-rule authoring","problem":"Benefits rules encoded in increasingly expressive policy languages may make universal eligibility determination nonterminating or unverifiable.","affected_stakeholder":"Benefits administrators and applicants","workflow_boundary":"From legislative rule translation through production eligibility determination","failure_mode":"An unrestricted rule engine promises exact, terminating decisions for every case and collapses timeout into ineligibility.","unit_of_analysis":"One machine-executable eligibility rule set","causal_lever":"Restrict authors to an enforceable decidable grammar and require a totality-and-correctness witness before deployment.","archetype_mapping":"Maps the unrestricted eligibility language, draws a decidable fragment, and separates exact decisions from rejected out-of-fragment cases.","expected_value":"Potentially fewer erroneous denials and less engineering effort spent repairing fundamentally overbroad automation claims.","falsifiable_claim":"In a pilot across newly authored rule sets, the restricted workflow will eliminate timeout-coded denials while retaining at least 80% of historically used rule constructs.","diversity_rationale":"Targets policy authoring, a rule-set unit, syntactic restriction, and applicant-facing denial errors rather than runtime triage, simulation, or procurement assurance.","mechanism_slugs":["language_fragment_restriction","constructive_algorithm_and_correctness_proof"],"search_questions":["Have public benefits systems adopted decidable rule-language subsets?","Which eligibility constructs fall outside known decidable fragments?","What share of real rule sets requires unrestricted recursion or external calls?"]},{"hypothesis_id":"H2","title":"Explicit unknown in permit review","problem":"Automated permit-compliance searches may find confirming evidence but cannot reliably establish that no violation exists within operational time.","affected_stakeholder":"Permit reviewers, applicants, and affected residents","workflow_boundary":"From submitted application through automated compliance triage to human review","failure_mode":"Search exhaustion or timeout is reported as compliant or noncompliant instead of unresolved.","unit_of_analysis":"One permit application checked against applicable constraints","causal_lever":"Replace Boolean output with witness-backed YES and bounded UNKNOWN, routing unresolved cases to accountable review.","archetype_mapping":"Treats compliance search as potentially one-sided, preserves unknown at a declared bound, and governs escalation rather than inventing a negative verdict.","expected_value":"Potentially reduces false clearance or rejection while concentrating expert attention on unresolved applications.","falsifiable_claim":"Compared with Boolean triage, the protocol will reduce audit-discovered timeout-derived misclassifications by at least 50% without increasing median end-to-end review time by more than 20%.","diversity_rationale":"Changes the output protocol and escalation boundary for individual permit cases; unlike H1 it does not restrict the policy language or claim complete automation.","mechanism_slugs":["semi_decision_with_explicit_unknown","fallback_mode_router"],"search_questions":["Do permitting systems currently distinguish timeout, unknown, and no violation?","Which permit checks admit verifiable positive witnesses but weak negatives?","How do unknown rates affect reviewer workload and applicant delay?"]},{"hypothesis_id":"H3","title":"Sound abstraction for emergency plans","problem":"Agencies cannot exhaustively predict every behavior of adaptive, interagency emergency-response plans.","affected_stakeholder":"Emergency managers and populations exposed to service failure","workflow_boundary":"Predeployment validation of response plans before exercises or activation","failure_mode":"Finite scenarios are presented as proof that no coordination deadlock or coverage gap can occur.","unit_of_analysis":"One formalized interagency response plan and hazard envelope","causal_lever":"Over-approximate plan states in a finite model and certify only properties proven safe on that abstraction.","archetype_mapping":"Replaces an unrestricted behavioral guarantee with sound finite-state validation whose false alarms and uncovered assumptions remain explicit.","expected_value":"Potentially detects coordination hazards earlier while making the limits of scenario testing auditable.","falsifiable_claim":"Against conventional tabletop review, abstraction-based validation will identify at least one independently confirmed coordination hazard in 20% more plans without producing more than two false alarms per confirmed hazard.","diversity_rationale":"Uses a plan-level finite abstraction and one-directional safety guarantee in emergency coordination, structurally distinct from language restriction and case triage.","mechanism_slugs":["abstract_interpretation_or_model_checking","proof_checking"],"search_questions":["Where has model checking been applied to public emergency plans?","Which operational behaviors can be conservatively abstracted without losing real failures?","What false-alarm rates remain usable for emergency managers?"]},{"hypothesis_id":"H4","title":"Bounded policy-interaction census","problem":"Policy analysts may generalize from sampled scenarios when assessing all combinations of a finite local-policy package.","affected_stakeholder":"Local policy analysts, council members, and regulated residents","workflow_boundary":"Ex ante review between drafting a policy package and legislative adoption","failure_mode":"A partial simulation is described as exhaustive, or a bounded result is generalized beyond its encoded jurisdiction and horizon.","unit_of_analysis":"All encoded household or firm states within one jurisdiction, policy package, and fixed horizon","causal_lever":"Define a finite state bound, enumerate every encoded case, and attach a scope-limited certificate to the result.","archetype_mapping":"Converts a vague universal interaction claim into a terminating exhaustive claim over a precisely encoded finite domain, followed by a separate feasibility assessment.","expected_value":"Potentially reveals rare eligibility cliffs or contradictory obligations with a checkable coverage claim.","falsifiable_claim":"For selected finite policy packages, exhaustive enumeration will uncover materially different outcomes missed by the agency's sampled scenarios in at least 10% of evaluations and will complete within the declared resource budget.","diversity_rationale":"Analyzes a complete bounded population of policy-state combinations and focuses on coverage cliffs, not individual decisions, language expressiveness, or behavioral abstractions.","mechanism_slugs":["bounded_domain_exhaustive_search","computational_complexity_analysis"],"search_questions":["Do agencies enumerate finite policy-interaction spaces today?","Which administrative datasets support complete, nonduplicative state encodings?","At what bounds does exhaustive policy analysis become operationally infeasible?"]},{"hypothesis_id":"H5","title":"Computability gate for algorithm procurement","problem":"Public buyers may accept vendor claims of universal, exact, terminating detection across open-ended fraud or misconduct patterns.","affected_stakeholder":"Procurement officers, oversight bodies, case subjects, and taxpayers","workflow_boundary":"From solicitation requirements through technical evaluation, contract award, and material system changes","failure_mode":"An impossibility analogy, benchmark, or vendor assertion substitutes for a model-matched solvability argument; guarantees then drift after award.","unit_of_analysis":"One procured algorithmic capability and its contractual guarantee","causal_lever":"Require an independently reviewed computability-boundary record specifying input class, guarantee, evidence, fallback, and recheck triggers.","archetype_mapping":"Moves boundary classification upstream into procurement, checks reduction direction and assumptions, and binds the shipped claim to a versioned decision record.","expected_value":"Potentially prevents impossible specifications, misleading guarantees, and repeated spending on unattainable universal automation.","falsifiable_claim":"Solicitations using the gate will contain fewer unsupported universal guarantees at award than matched conventional solicitations, with no increase above 10% in procurement cycle time.","diversity_rationale":"Intervenes at organizational acquisition governance, uses the procurement contract as the unit, and targets claim validity and guarantee drift rather than operational classification accuracy.","mechanism_slugs":["computability_boundary_decision_record","reduction_direction_checklist","proof_by_counterexample"],"search_questions":["Do public procurement standards require computability or totality evidence?","Which government algorithm solicitations make unrestricted universal claims?","Can independent reviewers apply a boundary gate reliably and within procurement timelines?"]}]}