{"schema_version":1,"research_id":"eoa_inverse_innovation_exp03_external48_20260801","source_assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","cell_id":"computability_boundary_mapping__gender_studies","selection_stratum":"REJECTION_LOW_BAND_AUDIT","search_queries":["formal verification fairness probabilistic programs FairSquare paper","VeriFair probabilistic fairness verification paper CAV 2019","official algorithmic impact assessment gender equality audit limitations uncertainty guidance","gender equality impact assessment algorithm automated audit fairness policy official guidance","site:nist.gov AI RMF uncertainty limitations test evaluation documentation fairness official","site:canada.ca Algorithmic Impact Assessment peer review gender impact automated decision official","site:eeoc.gov algorithmic fairness audit bias audit limitations gender official","automated bias audit false assurance limitations empirical study algorithmic auditing","Rice theorem semantic properties programs undecidable textbook official university","halting problem policy verification fairness undecidable paper","formal verification fairness undecidability fairness verification paper limitations","assurance case AI fairness scope claims uncertainty official guidance","site:bls.gov/ooh software developers median pay 2025","site:bls.gov/ooh social scientists median pay sociologists 2025","site:bls.gov/ooh compliance officers median pay 2025","site:bls.gov/ooh computer and information research scientists median pay 2025","PRISM model checker official finite state exhaustive reachability verification documentation"],"sources":[{"source_id":"S1","title":"FairSquare: Probabilistic Verification for Program Fairness","publisher":"Microsoft Research / Proceedings of the ACM on Programming Languages","url":"https://www.microsoft.com/en-us/research/publication/fairsquare-probabilistic-verification-program-fairness/","source_class":"PRIMARY_RESEARCH","publication_date":"2017-10","accessed_at":"2026-08-02","claims_supported":["Formal fairness definitions can be encoded as probabilistic program properties.","FairSquare automatically certifies specified fairness properties for a supported class of decision-making programs."]},{"source_id":"S2","title":"Probabilistic Verification of Fairness Properties via Concentration","publisher":"Proceedings of the ACM on Programming Languages","url":"https://arxiv.org/abs/1812.02573","source_class":"PRIMARY_RESEARCH","publication_date":"2019-10","accessed_at":"2026-08-02","claims_supported":["VeriFair verifies formal fairness properties using sampling and explicitly probabilistic guarantees.","Its termination guarantee has technical preconditions and excludes pathological boundary cases.","It demonstrates that formal fairness verification, confidence bounds, and qualified termination contracts are established technical prior art."]},{"source_id":"S3","title":"AI Risk Management Framework Core","publisher":"National Institute of Standards and Technology","url":"https://airc.nist.gov/airmf-resources/airmf/5-sec-core/","source_class":"STANDARD","publication_date":"2023-01-26","accessed_at":"2026-08-02","claims_supported":["AI evaluation should document uncertainty, benchmarks, methods, and results.","Risks or characteristics that cannot be measured should be documented.","Generalization limits, knowledge limits, fairness results, affected-community input, and lifecycle updates should be addressed."]},{"source_id":"S4","title":"Algorithmic Impact Assessment tool","publisher":"Government of Canada","url":"https://www.canada.ca/en/government/system/digital-government/digital-government-innovations/responsible-use-ai/algorithmic-impact-assessment.html","source_class":"OFFICIAL_GUIDANCE","publication_date":"2026-05-28","accessed_at":"2026-08-02","claims_supported":["The federal AIA is a mandatory, multidisciplinary assessment process for covered automated decisions.","Departments must review and update assessments when functionality or scope changes.","The guidance assigns oversight to the Office of the Chief Information Officer and calls for legal, privacy, diversity, inclusion, bias, and intersectionality expertise.","The process includes testing, monitoring, human involvement, recourse, documentation, and public results."]},{"source_id":"S5","title":"Auditing Work: Exploring the New York City Algorithmic Bias Audit Regime","publisher":"arXiv","url":"https://arxiv.org/abs/2402.08101","source_class":"PRIMARY_RESEARCH","publication_date":"2024-02-12","accessed_at":"2026-08-02","claims_supported":["Interviews with 16 experts and practitioners identified unclear scope, auditor definitions, data-access problems, and accountability weaknesses in an operational bias-audit regime.","The study supports concern about underspecified audit claims but does not report universal terminating policy-analysis guarantees."]},{"source_id":"S6","title":"What Is Gender Impact Assessment","publisher":"European Institute for Gender Equality","url":"https://eige.europa.eu/gender-mainstreaming/toolkits/gender-impact-assessment/what-gender-impact-assessment","source_class":"OFFICIAL_GUIDANCE","publication_date":"undated","accessed_at":"2026-08-02","claims_supported":["Gender-impact assessment is an ex ante, structured evaluation of likely policy or program effects on gender equality.","Established gender-impact practice evaluates projected effects rather than purporting to prove a universal computational property."]},{"source_id":"S7","title":"Computability Theory: Rice's Theorem","publisher":"Loyola Marymount University, Department of Computer Science","url":"https://cs.lmu.edu/~ray/notes/computabilitytheory/","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"undated","accessed_at":"2026-08-02","claims_supported":["Rice's theorem rules out a general decider for nontrivial semantic properties of unrestricted programs.","The theorem does not apply merely because a question is difficult, nor does it cover syntactic or trivial properties.","Static analyzers consequently may need approximation or an unknown result for general semantic questions."]},{"source_id":"S8","title":"Annexe 5: Argument-Based Assurance Cases","publisher":"UK Information Commissioner's Office","url":"https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/explaining-decisions-made-with-artificial-intelligence/annexe-5-argument-based-assurance-cases/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2020","accessed_at":"2026-08-02","claims_supported":["Assurance claims should be tied to evidence and qualified by scope, definitions, assumptions, operational context, and risks.","Fairness and bias mitigation are already contemplated as assurance-case goals."] 	} 	 
   	    	
		]  ,"problem_evidence":{"support":"WEAK","rationale":"There is external evidence that real bias-audit regimes suffer from unclear scope and definitions, while official frameworks require uncertainty, generalization limits, and unmeasurable characteristics to be disclosed. However, the bounded search found no institution or product promising an exact, terminating gender-equity verdict for every unrestricted executable policy, and no source documented timeouts being relabeled as compliance. The candidate's exact antecedent therefore remains unverified rather than established.","source_ids":["S3","S4","S5"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Official Canadian requirements demonstrate institutional authority, multidisciplinary review, Gender-based Analysis Plus participation, public assessment records, and scope-change review for automated decisions. EIGE documents an established policy-facing gender-impact-assessment function. These sources establish plausible authorizers and users, but not expressed demand for computability-boundary labels or willingness to adopt this specific mechanism.","source_ids":["S4","S6"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"VeriFair","similarity":"It verifies formally specified fairness properties, exposes a probabilistic error contract, and states explicit conditions under which termination is guaranteed.","remaining_difference":"It targets fairness properties of machine-learning programs and does not provide the proposed institutional admission control for executable policy languages, governance-controlled guarantee records, or a required UNKNOWN/out-of-scope workflow for gender-policy audits.","source_ids":["S2"]},{"name":"FairSquare","similarity":"It encodes formal fairness definitions as program properties and automatically certifies supported decision-making programs.","remaining_difference":"It is a technical verifier rather than a governed gender-impact-audit protocol that prevents unsupported policy models or inconclusive analyses from receiving broader clearance.","source_ids":["S1"]},{"name":"NIST AI RMF and ICO assurance-case guidance","similarity":"They already call for scoped claims, explicit assumptions, documented uncertainty and limitations, evidence, contextual interpretation, affected-party input, and lifecycle updates.","remaining_difference":"They do not prescribe an enforceable finite executable-policy fragment, mechanically prevent claim upgrades, or evaluate whether UNKNOWN routing reduces false definitive gender-equity clearances.","source_ids":["S3","S8"]},{"name":"Government of Canada Algorithmic Impact Assessment","similarity":"It combines automated-decision assessment with multidisciplinary governance, gender and intersectionality expertise, public records, human involvement, testing, monitoring, and reassessment after scope changes.","remaining_difference":"It is a questionnaire-based risk process, not a formal verifier with decidable-fragment admission, exhaustive reference semantics, or explicit computability-boundary result types.","source_ids":["S4"]}],"distinctive_claim_remaining":"The remaining testable distinction is the integrated mechanism: mechanically admit only a governed executable-policy fragment, produce separately typed EXACT_OR_SOUND_RESULT, UNKNOWN, TIMEOUT, and OUT_OF_SCOPE outcomes, and prevent technical staff from representing a fragment-bound result as universal gender-equity clearance. The differentiation is an untested systems-and-governance integration, not an established novelty claim.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Existing fairness verifiers demonstrate that formal fairness specifications, confidence-qualified results, and conditional termination contracts can be implemented. PRISM demonstrates automated exhaustive model checking over precisely specified models, while assurance guidance supplies a template for scoped, evidence-backed claims. A synthetic finite-state prototype is therefore technically credible. Evidence does not establish that real institutional policies can be encoded faithfully, that the proposed gender-disparity predicate is legitimate and stable, or that useful policies will remain inside the admitted fragment.","source_ids":["S1","S2","S7","S8","S13"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Avoiding an unjustified clearance could be consequential, and documented audit-scope failures make the safety objective credible. The exact failure mode's prevalence and realized harm remain unmeasured.","source_ids":["S3","S5"]},"stakeholder_pull":{"score":2,"rationale":"There are mandatory automated-decision assessments and established gender-impact processes, but no located demand for computability-boundary classification or evidence that the diagnosed universal guarantee is currently offered.","source_ids":["S4","S6"]},"incremental_advantage":{"score":3,"rationale":"Mechanical admission control and typed UNKNOWN/out-of-scope outputs are more enforceable than narrative limitation statements, but no comparative evidence yet shows fewer false clearances or better interpretation than existing assurance and mixed-method practices.","source_ids":["S3","S8"]},"distinctiveness_plausibility":{"score":2,"rationale":"Formal fairness verification, qualified guarantees, scoped assurance claims, multidisciplinary impact assessment, and scope-change review all exist. Only their specific integration into executable gender-policy auditing remains differentiable.","source_ids":["S1","S2","S3","S4","S8"]},"technical_implementability":{"score":3,"rationale":"A synthetic finite-state analyzer is feasible using established model-checking and fairness-verification techniques. Semantic fidelity, fragment-membership enforcement, and policy translation are substantial unresolved work.","source_ids":["S1","S2","S13"]},"adoption_authority_feasibility":{"score":3,"rationale":"Canada provides a concrete example of accountable institutional oversight, multidisciplinary review, and gender/intersectionality expertise. Authority exists in an adjacent setting, but access and willingness for this proposal are unknown.","source_ids":["S4"]},"evidence_readiness":{"score":4,"rationale":"Both the problem claim and the intervention admit safe bounded tests: public-artifact coding can test prevalence first, followed by synthetic seeded-policy comparisons if warranted. Neither requires live decisions or identity inference.","source_ids":["S4","S5","S13"]},"safety_net_benefit":{"score":4,"rationale":"Explicit scope, uncertainty, evidence, affected-community input, lifecycle review, and safe-failure practices are supported by official frameworks; mechanical routing could strengthen those controls. It cannot make an unjust disparity definition substantively legitimate.","source_ids":["S3","S4","S8"]},"scalability":{"score":2,"rationale":"The mechanism can reuse model-checking infrastructure, but each policy family still requires precise semantics, construct validation, governance, and updates when scope changes. Those requirements make cross-domain scaling labor intensive.","source_ids":["S3","S4","S13"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"A six-to-eight-week, non-deployed prevalence study: pre-register a coding protocol, sample and double-code 30 public audit or impact-assessment artifacts, interview no affected individuals, compare definitive-clearance language with explicit limitation or UNKNOWN language, and publish an adjudicated evidence table.","confidence":"MODERATE","assumptions":["Public artifacts are accessible without purchasing data.","Work uses approximately 0.15 to 0.30 annual FTE-equivalent across an audit researcher, gender-policy researcher, and formal-methods reviewer.","The band includes labor burden, protocol design, coordination, software, quality review, and reporting.","No sensitive identity data, legal representation, or production integration is involved."],"source_ids":["S4","S5","S9","S10","S11","S12"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"One non-production institutional prototype for one policy family, including formal semantics, fragment admission checks, model checker integration, typed outputs, guarantee records, security and privacy review, gender and intersectionality review, stakeholder workshops, and controlled evaluation.","confidence":"LOW","assumptions":["A three-to-six-person interdisciplinary team works for roughly six to twelve months.","Existing open-source verification infrastructure is reused.","The institution supplies policy documentation and accountable reviewers.","The range includes loaded labor, software, computing, legal and compliance review, coordination, documentation, and evaluation but excludes live consequential use."],"source_ids":["S3","S4","S9","S10","S11","S12","S13"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"A governed advisory launch for one institution and narrow policy class, covering production engineering, integration, access controls, security and privacy, independent verification, reviewer training, affected-party governance, monitoring, incident response, evaluation, documentation, and rollback capability.","confidence":"LOW","assumptions":["Launch occurs only after separate authorization and successful non-deployed evidence.","Formal outputs remain advisory and cannot trigger sanctions or eligibility decisions.","Several full-time technical and governance roles plus external review are required.","The band includes labor, software, data preparation, compliance, coordination, equipment, evaluation, and contingency capacity."],"source_ids":["S3","S4","S9","S10","S11","S12"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintenance of one narrow operational program: policy-model updates, guarantee versioning, scope revalidation, software and security maintenance, human escalation, affected-party oversight, compliance review, monitoring, and periodic comparative evaluation.","confidence":"LOW","assumptions":["Two to five FTE-equivalent roles are shared across technical maintenance, policy and gender analysis, governance, compliance, and evaluation.","Major expansion to new policy families or institutions is excluded.","Recurring costs include loaded labor, software, computing, coordination, independent review, training, and evaluation.","Material semantic or scope changes can require a new startup project rather than routine maintenance."],"source_ids":["S3","S4","S9","S10","S11","S12"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"Adjacent audit-scope and accountability failures are documented, but the defining claim—an exact terminating class-wide gender-equity verdict for unrestricted executable policies, or timeout treated as compliance—was not located.","source_ids":["S3","S5"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"The Government of Canada provides an official example of institutional oversight for automated-decision assessments, including multidisciplinary review, gender and intersectionality expertise, publication, reassessment, and central compliance oversight.","source_ids":["S4"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal can be compared with narrative or mixed-method labels on false definitive clearance, interpretation accuracy, coverage, and UNKNOWN rates. Existing prior art makes the components concrete even though their incremental effect is untested.","source_ids":["S1","S2","S3","S8"]},"bounded_next_evidence_step":{"status":"YES","reason":"A fixed public-artifact corpus can safely test whether the diagnosed guarantee and uncertainty-collapsing behavior occur before any analyzer is built or deployed.","source_ids":["S4","S5"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The proposed next step is document research only, involves no identity inference or consequential decisions, and can stop if the problem antecedent is absent. This does not clear the proposal for live deployment.","source_ids":["S3","S4"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The bands enumerate labor, software, data work, compliance, coordination, security, evaluation, and governance. Official wage data anchor labor magnitude, while low confidence reflects the absence of an institutional design partner or observed implementation history.","source_ids":["S9","S10","S11","S12"]}},"next_evidence_step":"Pre-register a six-to-eight-week document study and double-code 30 public artifacts—15 NYC Local Law 144 bias-audit reports and 15 Canadian federal Algorithmic Impact Assessments—for policy/model scope, guarantee language, timeout or failed-analysis handling, uncertainty labels, and whether 'no disparity found' is represented as class-wide clearance. Compare definitive-clearance frequency and coder interpretation error against artifacts with explicit limitation or UNKNOWN language. Falsify the diagnosed problem and stop advancement if no artifact claims unrestricted exact termination or collapses inconclusive analysis into compliance; do not build or deploy an analyzer unless the antecedent is observed.","blocking_evidence":["Observed examples of exact, terminating, class-wide gender-equity claims for unrestricted executable policies or of timeout, abstraction uncertainty, or out-of-scope input being reported as compliance.","A governed disparity predicate that does not materially erase nonbinary, intersectional, contextual, or changing identities.","Evidence that mechanically typed boundary labels improve interpretation or reduce false definitive clearance relative to scoped mixed-method and assurance-case labels.","Evidence that an enforceable finite fragment covers policy questions stakeholders materially need assessed without bypass-inducing UNKNOWN rates.","An institutional partner with affected-party representation and authority over scope, claims, escalation, versioning, and rollback.","A valid reduction or other proof would be required before describing any specific unrestricted gender-equity policy property as undecidable; Rice's theorem alone does not establish that exact result."],"research_disposition":"PROBLEM_PREVALENCE_STUDY","world_novelty_boundary":"No world-novelty conclusion is made or supported. This bounded search cannot establish world novelty; it found substantial adjacent prior art in fairness verification, model checking, scoped assurance cases, and governed impact assessment, while only the specific integration and comparative effect described here remain untested."}