{"schema_version":1,"research_id":"eoa_inverse_innovation_exp04_external_evaluation_20260802","source_assessment_id":"computability_boundary_mapping__public_administration_policy:RETRIEVAL_FIRST:v0","cell_id":"computability_boundary_mapping__public_administration_policy","search_queries":["site:gov.uk \"Guidelines for AI procurement\" acceptable performance evaluation reassessment","site:gao.gov GAO-23-106811 Artificial Intelligence Key Practices accountability federal use","site:gsa.gov artificial intelligence resources AI compliance plan independent evaluation ongoing monitoring fallback","site:sei.cmu.edu \"Incorporating Software Requirements into the System RFP\" assurance case","site:sei.cmu.edu/documents/1864/2009_003_001_15017.pdf assurance case RFP","site:omg.org/spec/SACM/2.3 Structured Assurance Case Metamodel","site:canada.ca Guide Scope Directive Automated Decision-Making significant modification fraud probability","\"Public Procurement for Responsible AI\" U.S. Cities Practices Needs pdf","site:opm.gov 2026 general schedule pay tables hourly annual rates Chicago","site:nist.gov assurance case glossary claims evidence assumptions","site:nist.gov AI RMF procurement third party evaluation monitoring performance official"],"sources":[{"source_id":"SRC1","title":"Guidelines for AI procurement","publisher":"UK Government, Office for Artificial Intelligence","url":"https://www.gov.uk/government/publications/guidelines-for-ai-procurement/guidelines-for-ai-procurement","source_class":"OFFICIAL_GUIDANCE","publication_date":"2020-06-08","accessed_at":"2026-08-02","claims_supported":["Public procurement officials are identifiable intended users with authority to define requirements, evaluate suppliers, and manage contracts.","The guidance recommends multidisciplinary evaluation, testing under varied conditions, acceptable-performance definitions, auditability, reassessment, and continuous monitoring.","Proofs of concept and challenge-based procurement are established bounded evaluation approaches."]},{"source_id":"SRC2","title":"Artificial Intelligence: Key Practices to Help Ensure Accountability in Federal Use","publisher":"U.S. Government Accountability Office","url":"https://www.gao.gov/products/gao-23-106811","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2023-05-16","accessed_at":"2026-08-02","claims_supported":["GAO identifies governance, data, performance, and monitoring as complementary federal AI-accountability principles.","Third-party assessment and audit are important, while public agencies face shortages of relevant technical expertise."]},{"source_id":"SRC3","title":"AI strategies and compliance plan","publisher":"U.S. General Services Administration","url":"https://www.gsa.gov/artificial-intelligence/resources/ai-strategies-and-compliance-plan","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-09-30; updated 2026-05-11","accessed_at":"2026-08-02","claims_supported":["GSA assigns review and authorization responsibilities to its CAIO, AI Oversight Committee, EDGE Board, system owners, OCIO, and counsel-related offices.","For high-impact systems, GSA requires impact documentation, real-world testing, independent evaluation, monitoring, human review, and records in a central repository.","GSA provides fallback or escalation options and periodic or event-triggered reassessment, establishing close workflow and safety analogues."]},{"source_id":"SRC4","title":"Incorporating Software Requirements into the System RFP: Survey of RFP Language for Software by Topic, v. 2.0","publisher":"Carnegie Mellon University Software Engineering Institute","url":"https://www.sei.cmu.edu/documents/1864/2009_003_001_15017.pdf","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2009-05","accessed_at":"2026-08-02","claims_supported":["A government-oriented RFP template requires an initial ISO/IEC 15026 software assurance case containing claims, arguments, evidence, explicit assumptions, and an approving authority.","The assurance case can be evaluated before award, incorporated into the contract, used as an acceptance condition, refined during development, and monitored throughout the contract.","This is a close prior-art collision with the proposed record and gate, though it is not computability-specific."]},{"source_id":"SRC5","title":"Structured Assurance Case Metamodel, Version 2.3","publisher":"Object Management Group","url":"https://www.omg.org/spec/SACM/2.3/PDF","source_class":"STANDARD","publication_date":"2023-10","accessed_at":"2026-08-02","claims_supported":["SACM standardizes structured assurance-case packages, argumentation, claims, asserted evidence, context, artifacts, and terminology.","A computability-specific record could reuse an established assurance-case representation rather than require a new information model.","SACM is domain-neutral and supplies no computability classification or totality-proof procurement rule."]},{"source_id":"SRC6","title":"Guide on the Scope of the Directive on Automated Decision-Making","publisher":"Government of Canada","url":"https://www.canada.ca/en/government/system/digital-government/digital-government-innovations/responsible-use-ai/guide-scope-directive-automated-decision-making.html","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"Publication date not stated on opened page","accessed_at":"2026-08-02","claims_supported":["The directive applies to procured automated systems used in federal administrative decisions and explicitly includes estimating fraud probability.","Its stated purposes include transparency, accountability, fairness, quality, recourse, and public reporting.","Significant changes in scope, functionality, population, or use trigger renewed assessment, providing an official recheck analogue.","Research-only testing that does not affect real clients is outside the directive's operational scope, supporting a safe retrospective first step."]},{"source_id":"SRC7","title":"Public Procurement for Responsible AI? Understanding U.S. Cities’ Practices and Needs","publisher":"2nd Workshop on Regulatable ML at NeurIPS 2024 / OpenReview","url":"https://openreview.net/attachment?id=pw6AEHJfvM&name=pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2024","accessed_at":"2026-08-02","claims_supported":["Semi-structured interviews with 18 employees in seven U.S. cities found oversight gaps, limited capacity, vendor opacity, missing performance information, and difficulty interpreting vendor metrics.","Some participants performed independent evaluations because vendor demonstrations or disclosures were inadequate and expressed interest in clearer, standardized or centralized review support.","The study documents a general procurement-assurance problem and stakeholder pull but does not document universal, exact, terminating fraud or misconduct claims."]},{"source_id":"SRC8","title":"Salary Table 2026-CHI","publisher":"U.S. Office of Personnel Management","url":"https://piv.opm.gov/policy-data-oversight/pay-leave/salaries-wages/salary-tables/26Tables/html/CHI.aspx","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026-01","accessed_at":"2026-08-02","claims_supported":["The 2026 Chicago locality table places GS-13 through GS-15 annual salaries at approximately $119,000 to $197,200 before benefits and overhead.","These rates provide a transparent public-sector labor-cost anchor for resource-equivalent estimates, not a quote or observed cost for computability review."]}],"problem_evidence":{"support":"WEAK","rationale":"The problem visibly exists at the broader level: public buyers report vendor opacity, insufficient performance evidence, selective demonstrations, limited technical capacity, and difficulty interpreting metrics, while official bodies require stronger evaluation and monitoring. Fraud-related automated decisions are an identified public-administration use. However, none of the eight opened sources documents the candidate's narrow failure mode—public award of universal, exact, guaranteed-terminating fraud or misconduct detection—or estimates its prevalence or consequences. The specific problem therefore remains an empirically plausible but unverified subset of a supported general assurance problem.","source_ids":["SRC1","SRC2","SRC3","SRC6","SRC7"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Public procurement practitioners, technical officials, CAIO-led oversight bodies, contracting authorities, and city employees are identifiable adopters or authorizers. Municipal interviewees expressed needs for technical interpretation, standardized review, independent evaluation, and stronger vendor accountability; official guidance assigns related responsibilities and requires evaluation. No source expresses demand for computability classification, reductions, or totality proofs specifically, so stakeholder pull is adjacent rather than direct.","source_ids":["SRC1","SRC2","SRC3","SRC7"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Government RFP software-assurance-case submission","similarity":"It already requires pre-award claims, arguments, evidence, assumptions, an approving authority, contract incorporation, acceptance conditions, lifecycle refinement, and review.","remaining_difference":"The template addresses software safety and security generally; it does not mandate an input-class and quantifier audit, decidability-status classification, termination proof, or scope-matched undecidability reduction.","source_ids":["SRC4"]},{"name":"OMG Structured Assurance Case Metamodel","similarity":"It provides a standardized representation for claims, evidence, context, argumentation, terminology, and assurance-case packages.","remaining_difference":"It is domain-neutral and contains no computability-specific evidence rules or public-procurement award gate.","source_ids":["SRC5"]},{"name":"GSA high-impact AI governance and compliance process","similarity":"It assigns authorizers, requires documentation, real-world testing, independent evaluation, monitoring, human review, fallback, centralized records, and reassessment.","remaining_difference":"It tests operational performance and risk rather than proving whether a declared open-ended problem class admits a correct total algorithm.","source_ids":["SRC3"]},{"name":"Government of Canada automated-decision scope and material-change review","similarity":"It covers procured public algorithms, fraud-related assessments, impact controls, and renewed review after significant scope or functionality changes.","remaining_difference":"It governs administrative impacts and quality but does not classify computability or require constructive totality evidence.","source_ids":["SRC6"]}],"distinctive_claim_remaining":"For public procurements that contain a genuinely universal, exact, terminating capability claim over an open-ended encoded class, adding a computability-specific subcase to an existing assurance process will cause independent reviewers to identify and narrow, relabel, or reject more unsupported guarantees than conventional assurance review, with reproducible classifications and no more than a 10% increase in procurement-cycle time. This is falsified if no such claims occur, conventional review already handles every one, reviewer classifications are not reproducible, or cycle time exceeds the threshold.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"Technical implementation is feasible as a structured document and review protocol: assurance-case standards and government RFP language already support claims, evidence, assumptions, approving authorities, contract incorporation, and lifecycle review. Existing public-sector workflows also demonstrate independent evaluation, fallback, monitoring, and material-change reassessment. The document-only first step needs no operational model or personal data. Material gaps are access to complete procurement files, trade-secret restrictions, availability of computability reviewers, jurisdiction-specific authority to impose proof-related conditions, formal-to-operational fidelity, and the risk that officials mistake absence of proof for impossibility. These require legal tailoring, reviewer training, confidentiality controls, and an explicit UNRESOLVED outcome.","source_ids":["SRC1","SRC3","SRC4","SRC5","SRC6","SRC7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Preventing an impossible or materially overstated guarantee could protect public funds and affected case subjects, but the frequency and realized harm of the narrow failure mode are unmeasured.","source_ids":["SRC6","SRC7"]},"stakeholder_pull":{"score":3,"rationale":"Officials express demand for better technical evaluation, standardized support, disclosure, and vendor accountability, but not for computability review specifically.","source_ids":["SRC1","SRC3","SRC7"]},"incremental_advantage":{"score":2,"rationale":"The intervention adds a precise theoretical screen, but most governance, documentation, independent-review, contract, fallback, and recheck components already exist; no comparative outcome evidence shows that the added subcase improves decisions.","source_ids":["SRC3","SRC4","SRC5","SRC6"]},"distinctiveness_plausibility":{"score":3,"rationale":"The opened sources contain no direct computability-specific procurement gate, yet bounded search cannot establish world novelty and substantial assurance-case collision remains.","source_ids":["SRC3","SRC4","SRC5"]},"technical_implementability":{"score":4,"rationale":"A form, decision lattice, checklist, and versioned record can be layered onto established assurance-case and RFP structures without building a new production algorithm.","source_ids":["SRC4","SRC5"]},"adoption_authority_feasibility":{"score":3,"rationale":"Contracting and AI-governance authorities capable of imposing evaluation and documentation requirements are identifiable, but authority and acceptable evidence standards require jurisdiction-specific procurement and legal review.","source_ids":["SRC1","SRC3","SRC4","SRC6"]},"evidence_readiness":{"score":2,"rationale":"General procurement-assurance needs and implementation analogues are documented, but prevalence, baseline handling, reviewer reliability, added time, and comparative effectiveness for universal terminating claims require new empirical data.","source_ids":["SRC4","SRC7"]},"safety_net_benefit":{"score":4,"rationale":"Explicit UNRESOLVED, UNKNOWN, and OUT_OF_SCOPE states plus retrospective testing can reduce overclaiming without influencing live cases; safeguards must prevent missing proof from being treated as proof of impossibility.","source_ids":["SRC3","SRC6"]},"scalability":{"score":3,"rationale":"Reusable templates and assurance standards support diffusion, but scarce formal-methods expertise, heterogeneous procurement authority, trade secrecy, and potentially low claim prevalence constrain scale.","source_ids":["SRC2","SRC4","SRC5","SRC7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Preregister definitions; obtain and code 20 completed procurement records; use two independent reviewers, adjudication, de-identification, reliability analysis, and a review-hours estimate.","confidence":"MODERATE","assumptions":["Approximately 140-240 combined hours for protocol design, retrieval, duplicate coding, adjudication, analysis, and reporting.","Public-sector technical labor is anchored to 2026 GS-13 through GS-15 Chicago salaries, then increased for benefits, overhead, and limited specialist review.","No proprietary-data purchase, live procurement intervention, or formal theorem proving is included."],"source_ids":["SRC7","SRC8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Draft the computability subcase, evidence rubric, reviewer manual, confidentiality process, procurement clauses, training materials, and legal/technical review for one authority.","confidence":"LOW","assumptions":["One jurisdiction and one procurement category.","Uses an existing assurance-case or governance workflow.","Includes part-time procurement, legal, technical, records, and independent formal-methods expertise but no production software platform."],"source_ids":["SRC3","SRC4","SRC5","SRC8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Run a prospective matched pilot across approximately 5-10 relevant procurements, fund independent reviews and adjudication, integrate records into contract workflow, measure delay, and conduct safety and process evaluation.","confidence":"LOW","assumptions":["Several universal-claim procurements can be recruited within the pilot period.","External reviewers are needed where agency expertise is unavailable.","No adverse procurement action is automated; contracting authorities retain decisions.","Secure document handling and legal review are included."],"source_ids":["SRC1","SRC2","SRC3","SRC4","SRC7","SRC8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain a small review function for roughly 10-30 screened procurements annually, including triage, specialist review of in-scope claims, training, record retention, audits, and material-change reassessment.","confidence":"LOW","assumptions":["Most procurements exit after inexpensive quantifier and scope triage; only a minority require formal proof review.","Costs are dominated by specialized labor and governance rather than computing infrastructure.","Volume, case complexity, external-review rates, and jurisdictional overhead are not yet observed."],"source_ids":["SRC2","SRC3","SRC4","SRC6","SRC8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"A general public AI procurement-assurance problem is externally supported, but the defining universal-exact-terminating claim failure has not been observed or measured in the opened evidence.","source_ids":["SRC6","SRC7"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Contracting authorities, procurement officials, CAIO-led oversight bodies, and public-sector technical reviewers already exercise adjacent evaluation, documentation, contract, and reassessment powers.","source_ids":["SRC1","SRC3","SRC4","SRC6"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal isolates a computability-specific addition to established assurance review and specifies observable outcomes: unsupported-claim disposition, reviewer agreement, and no more than 10% cycle-time increase.","source_ids":["SRC3","SRC4","SRC5"]},"bounded_next_evidence_step":{"status":"YES","reason":"A preregistered duplicate-coded sample of 20 completed records is bounded, comparator-based, non-operational, and has explicit prevalence, reliability, utility, and delay falsifiers.","source_ids":["SRC6","SRC7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The authorized first step is retrospective and document-only, leaves all award and case authority with officials, and can use de-identified findings; it must halt if protected information cannot be separated or findings begin affecting live decisions.","source_ids":["SRC3","SRC6"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"Scope and public-sector labor anchors support broad resource-equivalent bands, but no source reports actual computability-review effort, specialist rates, claim volume, or integration cost.","source_ids":["SRC4","SRC8"]}},"next_evidence_step":"Preregister the exact inclusion rule for a universal, exact, guaranteed-terminating claim and sample 20 completed public solicitations or award packages involving fraud or misconduct detection without selecting on outcome. Two independent reviewers, blinded to each other's coding, will classify quantifiers, input-class enforceability, exactness, termination, evidence type, fallback states, and change triggers under (A) conventional assurance criteria and then (B) the draft computability subcase; adjudication must preserve silence as missing information, not impossibility. Report prevalence with uncertainty, inter-rater agreement, changed dispositions, review hours, inaccessible-record rate, and formal-to-operational mismatches. Falsify progression if no in-scope claim is found, every claim is already resolved by baseline review, agreement is below a preregistered acceptable threshold, protected information prevents review, the gate systematically confuses missing evidence with impossibility, or projected cycle-time increase exceeds 10%. Proceed to a prospective matched pilot only if at least one unsupported claim is observed, classifications are reproducible, and the delay bound is met.","blocking_evidence":["No direct evidence that public buyers award universal, exact, guaranteed-terminating fraud or misconduct detection claims.","No prevalence estimate for in-scope claims across solicitations or contracts.","No evidence that conventional assurance review fails on identified in-scope claims.","No comparative evidence that the computability subcase changes award-language quality or reduces unsupported guarantees.","No measured inter-reviewer reliability for the proposed classification.","No observed review-hours or procurement-cycle-time effect.","No jurisdiction-specific legal analysis of proof obligations, trade secrets, bid fairness, protests, records disclosure, or reviewer authority.","No demonstrated supply, accreditation, or conflict-of-interest model for independent computability reviewers.","No validation that a formalized class faithfully represents the operational fraud or misconduct task.","No quantified false-positive risk from treating missing proof as evidence of impossibility."],"research_disposition":"PROBLEM_PREVALENCE_STUDY","world_novelty_boundary":"Assurance cases, structured claim-evidence-assumption records, approving authorities, pre-award review, contract incorporation, acceptance conditions, independent evaluation, fallback, monitoring, and change-triggered reassessment are established. The bounded search found no direct source requiring a model-relative computability classification, input-class and universal-quantifier audit, constructive correctness-and-totality evidence, or scope-matched impossibility reduction as a procurement condition. That absence is not a world-novelty finding. World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured.","arm":"RETRIEVAL_FIRST","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":true,"material_progress_observed":true,"progress_targets":["Obtain a preregistered prevalence estimate for universal, exact, guaranteed-terminating claims in relevant completed procurements.","Measure whether conventional assurance review already identifies and resolves every observed in-scope claim.","Demonstrate reproducible computability classifications with two independent reviewers and adjudication.","Estimate added review hours and projected procurement-cycle-time change, testing the 10% bound.","Document inaccessible or trade-secret-protected evidence rates and establish a lawful confidentiality and reviewer-access pathway.","Conduct jurisdiction-specific procurement and administrative-law review of gate authority, evidence demands, vendor fairness, records treatment, and protest risk.","Validate formal-to-operational fidelity on each candidate claim and quantify errors caused by confusing absent proof, timeout, poor scaling, and undecidability.","Only after those targets pass, run a prospective matched pilot comparing unsupported-guarantee disposition under baseline review versus the computability subcase."],"reason":"Bounded web research established a meaningful general procurement-assurance problem, credible public authorizers, substantial prior-art collision, and a narrow testable distinction. It did not establish that the defining universal-exact-terminating failure occurs or that the proposed gate improves decisions. Resolving those blockers requires document sampling, reviewer coding, access to procurement records, legal analysis, and eventually prospective workflow testing rather than additional bounded web search."}}