{"schema_version":1,"research_id":"eoa_inverse_innovation_exp03_external48_20260801","source_assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","cell_id":"computability_boundary_mapping__gender_studies","selection_stratum":"REJECTION_LOW_BAND_AUDIT","search_queries":["site:dl.acm.org FairSquare probabilistic verification fairness programs paper","site:link.springer.com VeriFair verifying fairness of machine learning systems paper CAV","site:nist.gov AI RMF uncertainty limitations documentation validity reliability","site:eige.europa.eu gender impact assessment guide stakeholders participation","FairSquare probabilistic verification of program fairness Albarghouthi 2017 pdf","VeriFair verifying fairness of machine learning systems Bastani 2019 pdf","policy as code formal verification government rules Catala language paper","three valued model checking unknown result abstraction verification paper","software verification tool unknown timeout result formal methods documentation","site:nyc.gov Local Law 144 automated employment decision tools bias audit requirements impact ratio","automated gender equity audit policy workflow universal guarantee exact compliance software","\"gender equity\" \"formal verification\" policy audit","\"gender impact assessment\" automated tool compliance verdict","\"gender bias audit\" automated policy compliance software guarantee","\"policy verification\" \"gender\" model checking"],"sources":[{"source_id":"S1","title":"FairSquare: Probabilistic Verification of Program Fairness","publisher":"Proceedings of the ACM on Programming Languages","url":"https://pages.cs.wisc.edu/~aws/papers/oopsla17.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2017-10-01","accessed_at":"2026-08-02","claims_supported":["FairSquare encodes formal fairness definitions as probabilistic program properties and automatically attempts to certify them.","Formal fairness verification of executable decision programs predates this candidate."]},{"source_id":"S2","title":"Probabilistic Verification of Fairness Properties via Concentration","publisher":"Proceedings of the ACM on Programming Languages / arXiv","url":"https://arxiv.org/abs/1812.02573","source_class":"PRIMARY_RESEARCH","publication_date":"2018-12-02","accessed_at":"2026-08-02","claims_supported":["VeriFair verifies formal fairness specifications using adaptive sampling and explicitly provides probabilistic rather than absolute guarantees.","Its termination guarantee has exceptional boundary cases, illustrating the need to qualify verification guarantees."]},{"source_id":"S3","title":"SMTChecker and Formal Verification","publisher":"Solidity Documentation","url":"https://docs.solidity.org/en/latest/smtchecker.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"","accessed_at":"2026-08-02","claims_supported":["A production formal-verification interface distinguishes proved failure from an unknown timeout result.","The documentation warns that specifications may omit unintended effects and that some verification problems cannot be solved automatically in general."]},{"source_id":"S4","title":"Catala: A Programming Language for the Law","publisher":"ACM SIGPLAN / arXiv","url":"https://arxiv.org/abs/2103.03198","source_class":"PRIMARY_RESEARCH","publication_date":"2021-03-04","accessed_at":"2026-08-02","claims_supported":["Catala provides an executable language for a restricted class of computational statutory rules.","The work combines legal and programming expertise and reports formally verified compiler components."]},{"source_id":"S5","title":"AI Risk Management Framework Core","publisher":"National Institute of Standards and Technology","url":"https://airc.nist.gov/airmf-resources/airmf/5-sec-core/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2023-01-26","accessed_at":"2026-08-02","claims_supported":["NIST calls for documenting system scope, knowledge limits, uncertainty, generalization limits, unmeasurable risks, and human oversight.","NIST recommends mixed-method evaluation, independent review, consultation with affected communities, and ongoing reassessment."]},{"source_id":"S6","title":"Gender Impact Assessment: General Considerations","publisher":"European Institute for Gender Equality","url":"https://eige.europa.eu/gender-mainstreaming/toolkits/gender-impact-assessment/general-considerations?language_content_entity=en","source_class":"OFFICIAL_GUIDANCE","publication_date":"","accessed_at":"2026-08-02","claims_supported":["Gender-impact assessment requires institutional support, gender expertise, data, training, monitoring, and stakeholder consultation.","EIGE warns that data gaps and oversimplification can undermine gender-impact assessment and recommends making gaps explicit."]},{"source_id":"S7","title":"Automated Employment Decision Tools","publisher":"New York City Department of Consumer and Worker Protection","url":"https://home4.nyc.gov/site/dca/about/automated-employment-decision-tools.page","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"","accessed_at":"2026-08-02","claims_supported":["New York City requires covered automated employment decision tools to undergo recurring bias audits and publish audit information.","A real regulatory audience therefore exists for gender-related algorithmic audit results, although the rule does not promise universal verification."]},{"source_id":"S8","title":"Halting Problem","publisher":"National Institute of Standards and Technology Dictionary of Algorithms and Data Structures","url":"https://xlinux.nist.gov/dads/HTML/haltingProblem.html","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2023-04-07","accessed_at":"2026-08-02","claims_supported":["No algorithm can terminate and correctly decide halting for every arbitrary program, although heuristic or partial solutions remain possible.","An unrestricted executable-policy guarantee would require a problem-specific reduction before undecidability could be asserted."]}],"problem_evidence":{"support":"WEAK","rationale":"Official regimes require gender-related impact or bias assessments, and real digital workflows support those assessments. Research also documents ambiguity in interpreting missing or narrow audit evidence. However, the bounded search found no institution or product promising an exact, terminating verdict about every possible gender-disparate outcome of an unrestricted executable policy. Existing public systems are predominantly guided workflows, statistical audits, or mixed-method assessments. The candidate's precise computability failure mode therefore remains plausible but externally unverified.","source_ids":["S5","S6","S7","S8"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Regulators and official guidance establish demand for gender-impact assessment, recurring bias audits, documented limits, accountable roles, and consultation with affected communities. This supports an identifiable authorizer and adjacent demand, but no adopter was found requesting computability-boundary labels or an UNKNOWN-routing guarantee contract specifically.","source_ids":["S5","S6","S7"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"FairSquare","similarity":"It formally encodes fairness properties of executable decision programs and automatically attempts to certify whether they hold, closely matching the candidate's computational fairness-analysis core.","remaining_difference":"It does not establish the proposed institution-facing contract that enforces a policy-language fragment, separates timeout and out-of-scope cases, and reserves guarantee changes to an affected-party governance body.","source_ids":["S1"]},{"name":"VeriFair","similarity":"It verifies formal fairness specifications at scale while explicitly qualifying its results as probabilistic and identifying exceptional termination cases.","remaining_difference":"It targets machine-learning fairness properties rather than unrestricted policy workflows and does not test whether governance labels prevent false institutional clearance.","source_ids":["S2"]},{"name":"Solidity SMTChecker unknown-result handling","similarity":"It operationally distinguishes proved failure from timeout-driven UNKNOWN and preserves soundness by reporting uncertainty rather than clearance.","remaining_difference":"It is not a gender-equity assessment, does not formalize social-policy semantics, and lacks affected-party authorization of the specification and guarantee boundary.","source_ids":["S3"]},{"name":"Catala law-as-code","similarity":"It restricts formalization to computational legal rules, provides executable semantics, and links legal-domain and programming expertise.","remaining_difference":"It does not verify gender-disparate outcomes or provide the proposed tri-state equity-audit and escalation contract.","source_ids":["S4"]}],"distinctive_claim_remaining":"The remaining testable claim is an integration claim: for executable gender-impact policy audits, enforceable fragment membership plus explicit PROVED-IN-FRAGMENT, VIOLATION, UNKNOWN, and OUT-OF-SCOPE labels—governed by an equity body with affected-party participation—will reduce false-clearance interpretations relative to ordinary audit labels without eliminating practically useful coverage. The underlying verification, restriction, and unknown-result mechanisms are not distinctive individually.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Primary research demonstrates executable fairness specifications and fairness verifiers, official product documentation demonstrates sound UNKNOWN handling, and law-as-code research demonstrates restricted executable policy languages. These components make a synthetic finite-state prototype credible. No source demonstrates the complete gender-policy governance integration, semantic adequacy for nonbinary and intersectional harms, or useful coverage on real institutional policies.","source_ids":["S1","S2","S3","S4","S5","S6"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Gender-impact and algorithmic-bias audits affect consequential institutional decisions, so preventing false clearance could matter. The exact failure mode's prevalence and realized harms remain unmeasured.","source_ids":["S6","S7"]},"stakeholder_pull":{"score":3,"rationale":"Official audit and impact-assessment requirements create adjacent demand, but no adopter request or procurement signal was found for computability-boundary labels specifically.","source_ids":["S5","S6","S7"]},"incremental_advantage":{"score":3,"rationale":"Explicit UNKNOWN and enforced scope are technically meaningful compared with an interface that collapses inconclusive states, but such handling is already established in formal-verification practice and its interpretation benefit in gender audits is untested.","source_ids":["S3","S5"]},"distinctiveness_plausibility":{"score":2,"rationale":"Fairness verification, qualified probabilistic guarantees, restricted law-as-code, UNKNOWN handling, and participatory governance all have close prior art. Only their specific institutional integration and comparative interpretation claim remain potentially distinctive.","source_ids":["S1","S2","S3","S4","S5","S6"]},"technical_implementability":{"score":4,"rationale":"A non-deployed finite-state demonstrator with exhaustive ground truth and seeded witnesses is readily plausible using demonstrated verification approaches. Real-policy semantics and fragment coverage remain substantially harder.","source_ids":["S1","S2","S3","S4"]},"adoption_authority_feasibility":{"score":3,"rationale":"Regulators, institutional leadership, independent assessors, gender experts, and affected communities are credible governance participants. No specific organization has agreed to own the proposed guarantee contract.","source_ids":["S5","S6","S7"]},"evidence_readiness":{"score":4,"rationale":"The candidate admits a safe synthetic comparison with exhaustive truth, seeded violations, label variants, and interpretation outcomes. Existing tools and guidance provide reusable design patterns even though no partner or corpus is secured.","source_ids":["S1","S2","S3","S5","S6"]},"safety_net_benefit":{"score":4,"rationale":"Scope enforcement, explicit uncertainty, human escalation, documented knowledge limits, and affected-community consultation directly guard against turning analyzer failure into clearance. They cannot validate the substantive justice of the encoded predicate.","source_ids":["S3","S5","S6"]},"scalability":{"score":2,"rationale":"Formal verification components can be reused, but policy-specific formalization, evolving assumptions, gender expertise, missing data, and stakeholder governance constrain transfer across policy families and institutions.","source_ids":["S4","S5","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Preregister and conduct a non-deployed artifact study and synthetic finite-state comparison, covering corpus acquisition and coding, formal semantics, seeded witnesses, exhaustive reference results, label-interface development, reviewer interpretation testing, affected-party consultation, analysis, and reporting.","confidence":"MODERATE","assumptions":["The corpus contains only public documents and synthetic policies.","One provisional disparity predicate and a small finite-state language are studied.","Labor includes formal-methods, gender-research, study-design, software, coordination, and governance-review effort.","No sensitive identity data, production integration, or consequential decisions are included.","The band is a reasoned resource-equivalent estimate; sources establish relevant work components but not direct dollar prices."],"source_ids":["S1","S2","S3","S5","S6"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Prepare one narrow institutional implementation, including policy-language formalization, fragment-membership enforcement, analyzer integration, versioned guarantee records, interface and accessibility work, legal and semantic review, security, training, escalation design, and predeployment evaluation.","confidence":"LOW","assumptions":["One institution and one policy family are in scope.","Existing solver or model-checking infrastructure can be reused.","Formal results remain advisory and identity inference is prohibited.","Custom semantic work and stakeholder coordination dominate cost.","No direct implementation-price source was located."],"source_ids":["S3","S4","S5","S6"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Launch a governed service for one bounded policy class, including production engineering, secure infrastructure, compliance review, independent evaluation, reviewer staffing and training, affected-party oversight, monitoring, incident response, documentation, and rollback capability.","confidence":"LOW","assumptions":["Launch follows successful synthetic and shadow-mode evaluation.","UNKNOWN cases receive accountable human review.","The institution must maintain traceability between policy versions, semantics, evidence, and guarantee labels.","The range excludes sector-wide rollout and automated sanctions.","No direct deployment-price source was located."],"source_ids":["S5","S6","S7"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain one bounded program through policy and model updates, guarantee versioning, data and semantic review, software and security maintenance, human escalation, stakeholder governance, monitoring, audits, and periodic comparative reevaluation.","confidence":"LOW","assumptions":["The supported fragment remains narrow.","Changing policies, evidence, and gender concepts require recurring interdisciplinary review.","Human review and affected-party participation remain funded functions.","Major expansion to additional institutions or policy families is excluded.","No direct recurring-cost source was located."],"source_ids":["S5","S6","S7"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"Real gender-impact and algorithmic-bias audit regimes exist, but the search did not verify the candidate's defining problem: a product or institution promising a terminating, exact, class-wide verdict for unrestricted executable policies or treating timeout as compliance.","source_ids":["S5","S6","S7"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Government regulators, institutional leadership, gender experts, independent assessors, and affected communities are externally supported governance actors, although no specific partner commitment exists.","source_ids":["S5","S6","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"Against close prior art, the remaining claim is narrowly testable: enforced fragment membership and explicit guarantee labels reduce false-clearance interpretation relative to ordinary audit labels while retaining useful coverage.","source_ids":["S1","S2","S3","S4"]},"bounded_next_evidence_step":{"status":"YES","reason":"A fixed public-document corpus and synthetic finite-state vignette experiment can compare labels against exhaustive truth without touching live policies or making consequential decisions.","source_ids":["S1","S2","S3","S5","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The next step can use public artifacts, synthetic policies, provisional categories, and non-consequential reviewer judgments. Advancement to deployment remains prohibited until semantic legitimacy and authorizer ownership are separately established.","source_ids":["S5","S6"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four scopes include labor, software, data, compliance, coordination, security, governance, and evaluation. Bands are broad and explicitly low-confidence where no direct price evidence exists.","source_ids":["S1","S3","S4","S5","S6"]}},"next_evidence_step":"Preregister a non-deployed artifact-and-vignette study: sample 40 public gender-impact or gender-bias audit reports and product/user documents from regulated contexts, independently code whether scope exclusions, timeouts, and inconclusive evidence are distinguishable from clearance, then randomize 60 professional and affected-party reviewers to ordinary audit labels versus explicit PROVED-IN-FRAGMENT, VIOLATION, UNKNOWN, and OUT-OF-SCOPE labels on 12 synthetic finite-state cases with exhaustive truth. Falsify advancement if no sampled artifact makes an unqualified class-wide claim or collapses an inconclusive state into clearance, or if explicit labels fail to reduce false-clearance interpretations without preserving useful coverage.","blocking_evidence":["Observed prevalence of exact, terminating, class-wide gender-equity claims or of timeout, out-of-scope, and inconclusive states being reported as compliance.","A stakeholder-governed disparity predicate that does not materially erase nonbinary, intersectional, contextual, or changing identities.","Comparative evidence that the proposed labels reduce false-clearance interpretation relative to ordinary audit labels.","Evidence that the enforceable fragment covers materially important policies at acceptable UNKNOWN and false-alarm rates.","A named equity-governance body willing and authorized to own scope, guarantee changes, escalation, and rollback.","Direct implementation and recurring-cost evidence from a defined institutional setting."],"research_disposition":"PROBLEM_PREVALENCE_STUDY","world_novelty_boundary":"This was a bounded search of primary fairness-verification research, formal-verification documentation, law-as-code research, and official gender-impact, algorithmic-audit, and AI-risk guidance. It found substantial component-level prior art but no exact integrated match. It was not an exhaustive patent, proprietary-system, non-English, or unpublished-practice search; therefore neither the remaining integration claim nor the absence of an exact match establishes world novelty."}