{"schema_version":1,"research_id":"eoa_inverse_innovation_exp03_external48_20260801","source_assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","cell_id":"computability_boundary_mapping__gender_studies","selection_stratum":"REJECTION_LOW_BAND_AUDIT","search_queries":["site:dl.acm.org FairSquare probabilistic program verification fairness","formal verification fairness probabilistic programs VeriFair PLDI official paper","site:nist.gov AI RMF limitations testing evaluation bias documentation uncertainty","official gender impact assessment automated decision systems equality audit limitations","FairSquare fairness verification tool unknown timeout result","fairness verification abstract interpretation unknown result formal tool","formal methods fairness verification survey limitations undecidable","algorithmic audit uncertainty no disparity found limitations study","gender impact assessment automated tool algorithm policy","automated gender impact assessment policy","gender equality policy simulation automated audit software","gender audit tool automated policy compliance software","nuXmv official documentation finite state model checking manual","TLA+ model checker finite state official documentation","site:bls.gov occupational employment wages software developers mathematicians 2025"],"sources":[{"source_id":"S1","title":"FairSquare: Probabilistic Verification for Program Fairness","publisher":"Microsoft Research / ACM","url":"https://www.microsoft.com/en-us/research/publication/fairsquare-probabilistic-verification-program-fairness/","source_class":"PRIMARY_RESEARCH","publication_date":"2017-10","accessed_at":"2026-08-02","claims_supported":["Fairness definitions can be encoded as probabilistic program properties.","FairSquare automatically certifies whether decision-making programs satisfy specified fairness properties.","Formal fairness verification of programs substantially predates this candidate."]},{"source_id":"S2","title":"Probabilistic Verification of Fairness Properties via Concentration","publisher":"Proceedings of the ACM on Programming Languages","url":"https://arxiv.org/abs/1812.02573","source_class":"PRIMARY_RESEARCH","publication_date":"2019-10","accessed_at":"2026-08-02","claims_supported":["VeriFair verifies specified fairness properties using adaptive sampling and concentration bounds.","Its correctness guarantee is probabilistic rather than absolute.","Its termination guarantee has stated exceptional cases, including boundary/pathological instances.","The paper compares VeriFair with FairSquare on explicit benchmarks."]},{"source_id":"S3","title":"AI Risk Management Framework Core","publisher":"National Institute of Standards and Technology","url":"https://airc.nist.gov/airmf-resources/airmf/5-sec-core/","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2023-01-26","accessed_at":"2026-08-02","claims_supported":["Risks that cannot be measured should be documented.","Testing should report uncertainty, benchmarks, methods, and limitations on generalizability.","Fairness and bias evaluations should be documented.","Affected communities and independent or domain experts should be consulted.","Feedback, monitoring, and fail-safe behavior outside knowledge limits are recommended."]},{"source_id":"S4","title":"Mitigating AI/ML Bias in Context: Establishing Practices for Testing, Evaluation, Verification, and Validation of AI Systems","publisher":"National Institute of Standards and Technology","url":"https://csrc.nist.gov/pubs/pd/2022/11/09/mitigating-ai-ml-bias-in-context/final","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2022-11-09","accessed_at":"2026-08-02","claims_supported":["Bias is context-dependent and requires socio-technical testing and evaluation.","Technical evaluation should be connected to societal values and deployment context.","A proof-of-concept implementation is an accepted preliminary evidence form."]},{"source_id":"S5","title":"About Gender Impact Assessments","publisher":"Commission for Gender Equality in the Public Sector, Victoria","url":"https://www.genderequalitycommission.vic.gov.au/about-gender-impact-assessments","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2025-07-31","accessed_at":"2026-08-02","claims_supported":["Victoria requires roughly 300 defined public-sector entities to conduct gender impact assessments for policies, programs, or services directly and significantly affecting the public.","Official practice includes policy-context research, stakeholder engagement, options analysis, recommendations, and reporting to the Commission.","The guidance explicitly includes women, men, and gender-diverse people."]},{"source_id":"S6","title":"NSW Gender Impact Assessment Resource Hub","publisher":"NSW Government / NSW Treasury","url":"https://www.nsw.gov.au/departments-and-agencies/nsw-treasury/projects-reviews-and-consultation/gender-equality-and-womens-economic-outcomes/gender-impact-assessment","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"n.d.","accessed_at":"2026-08-02","claims_supported":["NSW agencies must complete gender impact assessments for covered budget proposals.","Official assessment practice combines internal data, research, and meaningful stakeholder engagement.","Assessors must identify risks, limitations, knowledge gaps, intersectional data gaps, and monitoring plans.","NSW describes gender impact assessment as qualitative and complementary to distributional analysis, with decision-makers retaining responsibility."]},{"source_id":"S7","title":"nuXmv 2.2.0 User Manual","publisher":"Fondazione Bruno Kessler","url":"https://nuxmv.fbk.eu/downloads/nuxmv-user-manual.pdf","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026","accessed_at":"2026-08-02","claims_supported":["Existing model-checking infrastructure supports invariant and temporal-property analysis of finite-state transition systems.","The tool separates finite-state algorithms from techniques for infinite-state systems.","The manual explicitly describes general incompleteness of its infinite-state approaches, while identifying cases where counterexamples or proofs can be concluded."]},{"source_id":"S8","title":"Occupational Employment and Wages — May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-05-15","accessed_at":"2026-08-02","claims_supported":["May 2025 mean annual wages were approximately $148,100 for software developers, $111,490 for software quality-assurance analysts and testers, and $129,260 for mathematicians.","A multidisciplinary effort involving several partial or full annual labor equivalents can therefore cross the broad research and deployment cost-band boundaries before benefits, procurement, governance, or infrastructure costs."]}],"problem_evidence":{"support":"WEAK","rationale":"The broader epistemic risk is externally supported: NIST requires documentation of uncertainty, unmeasurable risks, generalizability limits, and contextual fairness evaluation, while VeriFair itself states probabilistic and conditional termination guarantees. However, the bounded search found no gender-impact institution, audit product, or published case promising a terminating exact verdict for every unrestricted executable policy, and official gender-impact practice is qualitative, evidence-based, and explicit about knowledge gaps. Thus the proposed failure mode is plausible but its defining prevalence claim remains unverified.","source_ids":["S2","S3","S4","S5","S6"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Credible institutional authorizers exist: Victoria imposes gender-impact-assessment duties on defined public bodies and receives periodic reports, while NSW Treasury requires covered proposals to include gender analysis for decision-makers. These sources establish authority and a recurring assessment workflow, but they do not show demand for computability-boundary labels or willingness to sponsor this mechanism.","source_ids":["S5","S6"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"FairSquare","similarity":"It formalizes fairness as a program property and automatically certifies decision-making programs against a specified probabilistic fairness condition—the central technical move proposed by the candidate.","remaining_difference":"It targets decision-making programs and population models rather than open-ended institutional policy workflows; the reviewed publication does not establish the candidate's particular combination of compiler-enforced fragment membership, explicit UNKNOWN routing, versioned guarantee records, affected-party governance, and comparison with qualitative gender-impact assessment.","source_ids":["S1"]},{"name":"VeriFair","similarity":"It produces formal fairness verdicts with explicit probabilistic correctness and termination qualifications, directly anticipating bounded guarantee language and the need to distinguish a result from an unconditional universal proof.","remaining_difference":"It uses adaptive sampling and high-probability guarantees for specified ML fairness properties, not exhaustive finite-state policy analysis, and it does not supply the proposed institutional escalation and claim-authorization workflow.","source_ids":["S2"]},{"name":"NIST AI RMF measurement and governance controls","similarity":"It already calls for documenting what cannot be measured, uncertainty, test conditions, generalizability limits, fairness results, independent review, affected-community input, monitoring, and fail-safe operation outside knowledge limits.","remaining_difference":"It is general risk-management guidance rather than an executable type system that rejects unsupported policy models and attaches a machine-checkable guarantee contract to each equity result.","source_ids":["S3","S4"]},{"name":"Official mixed-method gender impact assessment","similarity":"Victoria and NSW already require scoped policy assessment, contextual evidence, stakeholder engagement, limitations or knowledge gaps, recommendations, monitoring, and accountable decision-makers.","remaining_difference":"The official processes are qualitative or mixed-method and do not formally model executable policies, exhaust finite state spaces, or classify computability and analyzer outcomes.","source_ids":["S5","S6"]}],"distinctive_claim_remaining":"The remaining testable claim is that adding enforceable policy-fragment membership, machine-readable guarantee versions, and an explicit UNKNOWN state to an accountable gender-impact-assessment workflow will reduce false definitive clearances and interpretation errors without making coverage unusably narrow. This is a potentially distinctive integration claim, not evidence of world novelty.","confidence":"HIGH"},"implementation_evidence":{"support":"STRONG","rationale":"A synthetic finite-state pilot is technically feasible using existing model-checking infrastructure, and prior systems already encode and verify formal fairness properties. nuXmv also illustrates why finite- and infinite-state guarantees must be distinguished. This supports implementation of the bounded laboratory mechanism, not semantic validity for gender justice or safe operational deployment.","source_ids":["S1","S2","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Preventing an inconclusive analysis from becoming an equity clearance could matter substantially, and official guidance treats policy-related gender harm and unmanaged bias as consequential. Impact is capped because no occurrence or prevalence of the candidate's exact failure mode was found.","source_ids":["S3","S4","S5","S6"]},"stakeholder_pull":{"score":2,"rationale":"There are mandated gender-impact workflows and identifiable public-sector decision-makers, but no source showed requests, procurement, adoption, or practitioner demand for computability-boundary classification.","source_ids":["S5","S6"]},"incremental_advantage":{"score":2,"rationale":"Explicit UNKNOWN and enforceable scope are testable improvements over ambiguous outputs, but most communicative and governance controls—uncertainty, limitations, unmeasurable risks, contextual validation, human input, and monitoring—already appear in NIST guidance. Advantage over a well-run mixed-method assessment is therefore unproven.","source_ids":["S3","S4","S6"]},"distinctiveness_plausibility":{"score":2,"rationale":"Formal certification of program fairness, qualified probabilistic guarantees, fairness-specific solvers, and documented uncertainty governance are established prior art. Only the particular integration with executable gender-policy models and guarantee-authority controls remains potentially distinctive.","source_ids":["S1","S2","S3"]},"technical_implementability":{"score":4,"rationale":"Finite-state invariant checking and formal fairness specifications are implemented practices, making the proposed synthetic pilot feasible. Real-policy modeling, state explosion, infinite-state incompleteness, and construct validity prevent a top score.","source_ids":["S1","S2","S7"]},"adoption_authority_feasibility":{"score":4,"rationale":"Official regimes identify organizations that must conduct gender impact assessments, bodies that receive reports, and decision-makers who use the evidence. The unresolved issue is willingness and authority to alter formal audit claims, not the existence of plausible institutional owners.","source_ids":["S5","S6"]},"evidence_readiness":{"score":4,"rationale":"Existing formal tools permit a corpus with exhaustive reference outcomes, while official guidance supplies relevant reporting and interpretation criteria. A synthetic comparative study can proceed without personal data or live decisions.","source_ids":["S3","S6","S7"]},"safety_net_benefit":{"score":4,"rationale":"Separating unsupported, inconclusive, and no-violation-found outcomes accords with official recommendations to document uncertainty and unmeasurable risks and to fail safely outside knowledge limits. It would not protect against an unjust or socially invalid formal predicate.","source_ids":["S3","S4","S6"]},"scalability":{"score":2,"rationale":"Model-checking infrastructure is reusable, but official gender-impact practice requires contextual, intersectional evidence and stakeholder engagement, while infinite-state verification may be incomplete. Policy-specific semantics and governance are likely to dominate expansion costs.","source_ids":["S4","S5","S6","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Preregister and conduct a non-deployed corpus study plus synthetic finite-state benchmark: formalize one narrow policy language and disparity predicate, create seeded cases and exhaustive reference results, dual-code current audit claims, compare baseline and boundary labels, assess reviewer interpretation, and include gender-domain and affected-party review.","confidence":"MODERATE","assumptions":["Three to six months of partial effort from a formal-methods researcher, gender-policy researcher, evaluation researcher, and project coordinator.","Existing model-checking software is reused; no sensitive identity data, production integration, or custom hardware is required.","The band includes labor burden, participant or reviewer compensation, software/cloud use, coordination, governance review, and evaluation.","BLS wage data are only a labor-cost anchor and exclude benefits, contracting premiums, and institutional overhead."],"source_ids":["S3","S4","S7","S8"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Prepare one advisory implementation for one institution and policy family: build the restricted modeling interface and scope checker, integrate labeled outputs and versioned guarantee records, establish security and access controls, complete legal and semantic review, train reviewers, and test rollback and escalation without consequential automation.","confidence":"LOW","assumptions":["Several annual or partial annual labor equivalents across software, formal methods, gender research, governance, legal/compliance, security, and program management.","Institutional policy documentation is available and a responsible authorizer participates.","Existing verification engines can be integrated, but policy semantics and user-facing controls require custom work.","No identity inference, automatic sanctions, or eligibility decisions are included."],"source_ids":["S3","S5","S6","S7","S8"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Launch a governed advisory service for a bounded policy class, including production engineering, independent verification and evaluation, security and compliance, documentation, monitoring, affected-party engagement, reviewer staffing and training, incident response, and tested rollback.","confidence":"LOW","assumptions":["A multidisciplinary team of several full-time-equivalent staff is required through launch.","Formal findings remain advisory and every unsupported or inconclusive result is routed to accountable review.","The estimate includes labor burden, contractors, software/cloud infrastructure, accessibility, coordination, evaluation, and contingency.","Launch is limited to one institutional program rather than multi-jurisdiction deployment."],"source_ids":["S3","S5","S6","S7","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain one narrow operational program through policy-model updates, guarantee versioning, software and security maintenance, semantic and intersectional review, human escalation, affected-party consultation, incident handling, and periodic comparative evaluation.","confidence":"LOW","assumptions":["Two to five combined full-time-equivalent or substantial partial roles across engineering, formal analysis, gender expertise, review, and coordination.","Supported policy scope remains narrow and major expansion is excluded.","Changing policy assumptions, gender concepts, evidence, and system versions require recurring revalidation.","BLS wages anchor labor scale but do not capture all benefits, overhead, or specialized contracting premiums."],"source_ids":["S3","S5","S6","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"Official sources support the general need to label uncertainty, limitations, knowledge gaps, and unmeasurable risks, but the search found no actual exact, terminating, unrestricted gender-policy verdict or documented collapse of timeout into compliance. The defining problem claim remains unverified.","source_ids":["S2","S3","S5","S6"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Victoria's Commission and defined entities, and NSW agencies, Treasury reviewers, and government decision-makers, are externally documented owners or users of gender-impact-assessment processes.","source_ids":["S5","S6"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim can be tested by comparing ordinary outputs with enforced scope plus distinct VIOLATION, CLEAR-WITHIN-CONTRACT, UNKNOWN, and OUT-OF-SCOPE labels, measuring false clearance and interpretation accuracy. Prior art makes this incremental rather than foundational.","source_ids":["S1","S2","S3"]},"bounded_next_evidence_step":{"status":"YES","reason":"A preregistered audit of public artifacts and output schemas is non-deployed, finite, and can directly falsify the diagnosed problem before any software pilot or institutional integration.","source_ids":["S1","S2","S5","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The proposed next study uses public artifacts and synthetic policies, makes no decisions about people, infers no gender identities, and can be overseen using the consultation, limitation-reporting, and contextual-review practices in official guidance. This does not clear live deployment.","source_ids":["S3","S4","S5","S6"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four scopes distinguish laboratory evidence, institution-specific startup, operational launch, and recurring maintenance. Current BLS occupational wages support the order of magnitude of labor-intensive bands, while low confidence appropriately reflects unknown staffing, integration, compliance, and governance requirements.","source_ids":["S3","S5","S6","S8"]}},"next_evidence_step":"Preregister a non-deployed prevalence study of 40 public manuals, reports, and output schemas: 20 automated fairness/compliance-analysis artifacts and 20 official or mixed-method gender-impact-assessment artifacts. Two independent coders should compare whether each defines its input class and predicate, distinguishes timeout/unsupported/inconclusive/no-violation-found, and makes an exact class-wide clearance claim. Falsify advancement if no artifact promises an exact unrestricted verdict and none collapses an inconclusive state into compliance; if at least one does, use its disclosed output contract to construct the later seeded finite-state label experiment.","blocking_evidence":["An observed target institution, report, or product that makes the defining exact terminating class-wide gender-equity claim or treats an inconclusive analysis as compliance.","A governed disparity predicate and execution semantics that affected-party reviewers judge adequate for one narrow synthetic use case.","Evidence that enforceable fragment membership and proposed labels improve interpretation beyond controls already recommended by NIST and used in mixed-method gender-impact assessment.","Evidence that a restricted fragment covers materially relevant policy workflows at an acceptable UNKNOWN and false-alarm rate.","A willing institutional authority that can own the guarantee contract, public claims, escalation, versioning, and rollback.","Operational cost evidence tied to a specified institution, supported policy family, integration environment, and staffing model."],"research_disposition":"PROBLEM_PREVALENCE_STUDY","world_novelty_boundary":"This was a bounded search across formal fairness verification, model checking, AI assurance guidance, and official gender-impact-assessment practice. It found substantial adjacent and overlapping prior art but no exact implementation of the full proposed governance integration. That absence is not a world-novelty claim, and the remaining integration claim should not be described as novel without a broader systematic literature, patent, standards, product, and institutional-practice review."}