{"schema_version":1,"research_id":"eoa_inverse_innovation_exp04_external_evaluation_20260802","source_assessment_id":"negative_space_design__tech_ethics_ai_governance:RETRIEVAL_FIRST:v0","cell_id":"negative_space_design__tech_ethics_ai_governance","search_queries":["site:nist.gov AI RMF Playbook Govern resource allocation independent testing AI red teaming","site:gov.uk AI assurance portfolio techniques pre-deployment testing AI official","site:gao.gov AI governance workforce resources testing report","site:oecd.org AI pilots governance pre deployment testing capacity","site:gov.uk AI assurance industry temperature check resources funding capacity barriers","site:whitehouse.gov OMB AI governance board pre-deployment testing agencies 2025 memorandum","site:gao.gov artificial intelligence agencies resources governance challenges pre-deployment testing 2025","AI red teaming costs resource intensive external red teaming primary research","UNFPA costed evaluation plan ring fenced funds annual resource ceiling evaluation PDF","site:anthropic.com responsible scaling policy evaluations red teaming pause external review","\"ring-fenced\" \"AI assurance\" budget","\"percentage\" \"AI\" portfolio capacity red teaming evaluation budget","site:bls.gov Occupational Employment and Wage Statistics data scientists May 2025 median annual wage","site:bls.gov information security analysts median pay 2025","site:bls.gov management analysts median annual wage 2025","site:unfpa.org costed evaluation plan indicative budgets evaluation USD table","site:learn.microsoft.com Azure AI governance sandbox initial experiments stakeholder consultation pre deployment official"],"sources":[{"source_id":"S1","title":"AI RMF Core","publisher":"National Institute of Standards and Technology","url":"https://airc.nist.gov/airmf-resources/airmf/5-sec-core/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2023-01-26","accessed_at":"2026-08-02","claims_supported":["Organizational practices should enable AI testing, incident identification, and information sharing.","Independent review can improve testing effectiveness and mitigate internal bias and conflicts.","Organizations tailor risk-management activity to resources, capabilities, and risk tolerance; the framework does not prescribe a fixed portfolio reserve."]},{"source_id":"S2","title":"Industry temperature check: barriers and enablers to AI assurance","publisher":"Centre for Data Ethics and Innovation and UK Department for Science, Innovation and Technology","url":"https://www.gov.uk/government/publications/industry-temperature-check-barriers-and-enablers-to-ai-assurance/industry-temperature-check-barriers-and-enablers-to-ai-assurance","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2022-12-07","accessed_at":"2026-08-02","claims_supported":["Stakeholders reported that lack of financial resources, skills, senior-management buy-in, and guidance impedes AI assurance.","The report describes standards and assurance work as potentially labor-intensive and costly, particularly for smaller organizations.","Finance-sector respondents reported inadequate resources and skills for AI assurance, but the study did not measure saturated pilot portfolios or demand for a fixed 15% reserve."]},{"source_id":"S3","title":"OMB Memorandum M-25-21: Accelerating Federal Use of AI through Innovation, Governance, and Public Trust","publisher":"Executive Office of the President, Office of Management and Budget","url":"https://www.whitehouse.gov/wp-content/uploads/2025/02/M-25-21-Accelerating-Federal-Use-of-AI-through-Innovation-Governance-and-Public-Trust.pdf","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2025-04-03","accessed_at":"2026-08-02","claims_supported":["CFO Act agencies must have AI Governance Boards chaired at Deputy Secretary level with CAIO, budget, legal, privacy, civil-rights, cybersecurity, procurement, and evaluation representation.","CAIOs advise agency heads and CFOs on resourcing, oversee independent review, and track AI spending and high-impact use cases.","High-impact federal AI requires predeployment testing, impact assessment, cost analysis, and independent review, while limited pilots may receive centrally tracked CAIO certification.","The memorandum identifies a concrete authorizer but does not require a percentage reserve and emphasizes speedy, value-preserving adoption."]},{"source_id":"S4","title":"Artificial Intelligence: Federal Efforts Guided by Requirements and Advisory Groups","publisher":"U.S. Government Accountability Office","url":"https://www.gao.gov/products/gao-25-107933","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2025-09-09","accessed_at":"2026-08-02","claims_supported":["GAO identified 94 government-wide AI requirements and ten executive-branch oversight or advisory groups.","Federal agencies constitute an identifiable class of governed adopters, although GAO reports no fixed pilot-capacity reserve or outcome evidence for one."]},{"source_id":"S5","title":"OpenAI's Approach to External Red Teaming for AI Models and Systems","publisher":"OpenAI authors via arXiv","url":"https://arxiv.org/abs/2503.16431","source_class":"PRIMARY_RESEARCH","publication_date":"2025-03-20","accessed_at":"2026-08-02","claims_supported":["External red teaming can discover novel risks, stress-test mitigations, seed quantitative evaluations, and inform deployment risk assessment.","Campaign design requires choices about scope, participant expertise, access, confidentiality, timing, criteria, and outputs.","Red teaming has limitations and belongs within a broader evaluation program rather than serving as sufficient assurance alone."]},{"source_id":"S6","title":"Costed Evaluation Plan: Guidance, Tools and Templates","publisher":"United Nations Population Fund Independent Evaluation Office","url":"https://www.unfpa.org/sites/default/files/admin-resource/CEPlan%20Guidance%2C%20Tools%20and%20Templates-307.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-07-03","accessed_at":"2026-08-02","claims_supported":["UNFPA ring-fences evaluation funds inside an annual programme ceiling and prohibits other uses absent formal approval by designated authorities.","The reserve is not additional funding, making this a close structural analogue for protected capacity within an existing ceiling.","Indicative evaluation budgets range from roughly $40,000 to $150,000 depending on portfolio size and context, with explicit warning that expertise and logistics alter cost."]},{"source_id":"S7","title":"Anthropic's Responsible Scaling Policy","publisher":"Anthropic","url":"https://www.anthropic.com/responsible-scaling-policy","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"2026-04-29","accessed_at":"2026-08-02","claims_supported":["Anthropic institutionalizes recurring evaluations, internal and external red teaming, risk reporting, independent-review authority, and potential pauses.","The policy uses capability thresholds and scheduled safeguards, not a fixed share of an annual multi-pilot portfolio."]},{"source_id":"S8","title":"Portfolio of AI Assurance Techniques","publisher":"UK Department for Science, Innovation and Technology","url":"https://www.gov.uk/guidance/portfolio-of-ai-assurance-techniques","source_class":"OFFICIAL_GUIDANCE","publication_date":"2023-06-07","accessed_at":"2026-08-02","claims_supported":["AI assurance already uses impact assessment, audits, conformity assessment, performance testing, formal verification, and other techniques across the lifecycle.","The portfolio documents established practices but does not allocate organizational pilot capacity or establish protected release rules."]},{"source_id":"S9","title":"National Employment and Wage Data by Occupation, May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.t01.htm","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026-05-15","accessed_at":"2026-08-02","claims_supported":["Annual mean wages were $132,510 for information-security analysts, $126,800 for data scientists, $111,490 for software quality-assurance analysts and testers, and $153,930 for computer and information research scientists.","These wage benchmarks support labor-equivalent cost estimates but exclude benefits, overhead, outside experts, community compensation, compute, and displaced pilot value."]},{"source_id":"S10","title":"Guidance to Set Up Your Organization's AI Governance Process","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai/govern","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-04-10","accessed_at":"2026-08-02","claims_supported":["Microsoft recommends sandbox environments before production validation, cross-functional stakeholder consultation, measurement plans, independent reviews, and red-team testing.","The guidance supplies deployable workflow and tooling analogues but no annual fixed-percentage nondeployment reserve."]}],"problem_evidence":{"support":"WEAK","rationale":"The problem is plausible and adjacent evidence is visible: S2 documents resource, skill, funding, and management-priority barriers to AI assurance, while S1 and S3 treat independent predeployment testing as important. However, no opened source measures the candidate's specific failure state—organization-years where live-use pilots saturate capacity, eligible nondeployment inquiries go unfunded, and material findings consequently arrive later. The prevalence and causal importance of portfolio saturation therefore remain unverified.","source_ids":["S1","S2","S3"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Federal CFO Act agency AI Governance Boards and CAIOs are identifiable authorizers with explicit responsibility for AI investment, resourcing, independent review, and predeployment testing. Industry stakeholders also report inadequate assurance resources and skills. This supports an adopter and a general need for assurance capacity, but no stakeholder expressly requests or endorses a fixed 15% nondeployment reserve.","source_ids":["S2","S3","S4"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"UNFPA ring-fenced evaluation funding","similarity":"Protected evaluation resources sit inside an annual ceiling, cannot fund delivery activities, and can be released only through formal authority—the closest structural match.","remaining_difference":"It governs international-development programme evaluation, is context-costed rather than a fixed 15%, and does not reserve AI-pilot capacity for predeployment harm inquiry.","source_ids":["S6"]},{"name":"NIST/OMB independent AI testing and governance","similarity":"Existing public governance requires or recommends predeployment testing, independent review, documented risk decisions, capable authorizers, and resource allocation.","remaining_difference":"Controls are applied by use case and risk level; neither source establishes an exclusive minimum organization-year share protected before pilot selection.","source_ids":["S1","S3"]},{"name":"Microsoft sandbox, measurement, and independent-review workflow","similarity":"Provides product-supported sandbox experimentation, stakeholder consultation, quantitative and qualitative evaluation, red teaming, and independent review before or around deployment.","remaining_difference":"It is lifecycle and workload guidance, not portfolio ring-fencing with an annual share and controlled release.","source_ids":["S10"]},{"name":"Anthropic Responsible Scaling Policy","similarity":"Institutionalizes recurring evaluations, external review, red teaming, risk reporting, and authority to pause or constrain deployment.","remaining_difference":"It is capability-threshold governance for frontier models, not a percentage reserve across an organization's annual pilot portfolio.","source_ids":["S7"]},{"name":"UK Portfolio of AI Assurance Techniques","similarity":"Collects established technical and procedural methods that could occupy the proposed reserve.","remaining_difference":"It is a catalogue of practices, not a resource-allocation or release-control mechanism.","source_ids":["S8"]}],"distinctive_claim_remaining":"Within the bounded opened-source search, the remaining claim is that an organization-year rule reserving at least 15% of total AI-pilot capacity, before individual selection, exclusively for nondeployment harm inquiry and subject to independent, documented release criteria will increase independently adjudicated material predeployment issues per proposed system without reducing validated pilots by more than 10%, compared with ordinary risk-tiered assurance lacking a portfolio reserve. This is falsifiable but not yet supported, and it is not a world-novelty claim.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"The component workflow is feasible: existing guidance supports separate testing, sandboxing, independent review, stakeholder input, documented metrics, and governance boards with budget and legal representation. A records-only shadow allocation is technically simple if portfolio, staffing, approval, inquiry, and outcome records exist. Feasibility is reduced by undefined capacity units, heterogeneous pilot effort, confidential or personal data, possible inability to reconstruct inquiries that were never conducted, gaming of 'inquiry' labels, and the absence of evidence that 15% is appropriate. Prospective affected-community work would additionally require consent, compensation, privacy, accessibility, and safeguarding.","source_ids":["S1","S3","S5","S8","S10"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Earlier discovery could prevent exposure and lock-in, but the incidence of reserve-constrained missed findings and the effect size are unknown.","source_ids":["S1","S3","S5"]},"stakeholder_pull":{"score":3,"rationale":"Authorizers and general demand for adequately resourced assurance are identifiable, but no source expresses pull for the fixed-share mechanism.","source_ids":["S2","S3","S4"]},"incremental_advantage":{"score":2,"rationale":"The reserve could make assurance less vulnerable to delivery pressure, but no comparative evidence shows advantage over risk-tiered reviews, mandatory per-system testing, or adequately funded independent assurance.","source_ids":["S1","S3","S7","S10"]},"distinctiveness_plausibility":{"score":3,"rationale":"No exact combination appeared in the bounded search, although ring-fenced evaluation funding, percentage allocation, independent AI assurance, and controlled pauses are established adjacent practices.","source_ids":["S1","S6","S7","S8"]},"technical_implementability":{"score":4,"rationale":"Separate accounting, sandbox testing, review workflows, and portfolio records are ordinary organizational capabilities; measurement definitions and data access are the main obstacles.","source_ids":["S3","S6","S10"]},"adoption_authority_feasibility":{"score":3,"rationale":"AI Governance Boards, CAIOs, CFOs, and equivalent portfolio bodies can plausibly authorize a reserve, but a fixed share may conflict with existing efficiency, risk-proportionality, and mission priorities.","source_ids":["S1","S3","S4"]},"evidence_readiness":{"score":2,"rationale":"Public sources establish component practices but provide neither portfolio-saturation prevalence nor outcome data for fixed reserves; proprietary records and prospective testing are needed.","source_ids":["S2","S3","S8"]},"safety_net_benefit":{"score":3,"rationale":"Protected inquiry capacity could provide an additional pre-exposure safeguard, provided it supplements rather than replaces mandatory safety and incident-response work.","source_ids":["S1","S3","S5"]},"scalability":{"score":3,"rationale":"A percentage and accounting rule can scale administratively, but absolute cost, inquiry demand, portfolio heterogeneity, and access to qualified independent reviewers vary substantially.","source_ids":["S2","S6","S9"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"A records-only audit of one organization across one to three completed portfolio years: define capacity units, reconstruct proposals and approvals, identify eligible unmet inquiries, simulate a 15% reserve, and obtain independent materiality review.","confidence":"MODERATE","assumptions":["Approximately 6-12 person-weeks across an analyst, AI-assurance reviewer, portfolio owner, and legal/privacy review.","Records are already digitized and no new affected-person data collection is required.","Loaded labor and limited specialist review are inferred from national wage benchmarks; this is not a vendor quote."],"source_ids":["S6","S9"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Design the portfolio policy, capacity taxonomy, separate accounting codes, inquiry eligibility and materiality rubric, release workflow, reporting templates, privacy review, and staff training.","confidence":"MODERATE","assumptions":["A medium organization can adapt existing portfolio, governance, and sandbox systems rather than procure a new platform.","Approximately 0.5-1.5 FTE-years of cross-functional work plus limited counsel or external assurance.","No major integration with procurement or financial systems is required."],"source_ids":["S3","S6","S9","S10"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Operate the reserve for the first annual portfolio, including dedicated red teaming or shadow evaluation, independent adjudication, community participation where appropriate, governance review, and opportunity cost of displaced live-pilot capacity.","confidence":"LOW","assumptions":["Illustrative medium portfolio with roughly 15-30 technical FTE-equivalents, making a 15% reserve approximately 2-5 FTE-equivalents.","External subject-matter experts, participant compensation, secure environments, and compute remain bounded.","The host does not have unusually expensive frontier-model testing or a very large pilot portfolio."],"source_ids":["S5","S6","S9"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Recurring protected inquiry capacity, portfolio administration, independent review, secure data handling, evaluation maintenance, reporting, and documented release decisions for a medium-size organization.","confidence":"LOW","assumptions":["The reserve continues to represent approximately 2-5 assurance FTE-equivalents plus external expertise and compute.","Costs scale with total pilot capacity; very small organizations may fall below this band and large or frontier-model organizations may exceed it substantially.","The figure is resource-equivalent cost, including opportunity cost, not necessarily incremental cash spending."],"source_ids":["S2","S6","S9"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"Sources establish inadequate assurance resources and the importance of predeployment testing, but not saturated live-pilot portfolios, zero protected inquiry capacity, or later findings caused by that allocation pattern.","source_ids":["S1","S2","S3"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"OMB identifies CFO Act agency AI Governance Boards, Deputy Secretaries, CAIOs, CFOs, and cross-functional officials with AI investment, resourcing, review, and spending responsibilities.","source_ids":["S3","S4"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The 15% protected share, exclusive eligible uses, release compliance, material issues per proposed system, and no-more-than-10% validated-pilot loss can be measured against no-reserve or existing-assurance comparators.","source_ids":["S3","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A one-organization, one-to-three-year proprietary-record audit can test prevalence, measurement feasibility, unmet inquiry demand, and opportunity-cost bounds without changing deployments.","source_ids":["S3","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The records-only step can be authorized by the existing portfolio body and halted if lawful access, privacy, security, or independent adjudication is unavailable; it entails no experimental exposure or deployment change.","source_ids":["S3"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"Labor and adjacent evaluation benchmarks support order-of-magnitude bands, but total portfolio capacity is unspecified and a 15% resource-equivalent reserve can vary by more than an order of magnitude across organizations.","source_ids":["S6","S9"]}},"next_evidence_step":"With approval from one organization's AI portfolio authority and privacy/legal function, preregister a records-only audit covering one to three completed annual portfolios. Define a common capacity denominator such as loaded labor-months plus directly attributable compute and vendor spend; enumerate all proposed pilots, actual live-use commitments, planned and requested nondeployment inquiries, approval timing, findings, and validated-pilot outcomes; and have an uninvolved reviewer apply a preregistered materiality rubric. Compare (A) actual allocation without a protected minimum, (B) a mechanical 15% shadow allocation applied before selections, and, if available, (C) a historical year or business unit with independently funded assurance. Do not impute extra detected issues from inquiries that never occurred; use the audit only to test problem prevalence, measurement feasibility, eligible unmet demand, and maximum displaced-pilot exposure. Stop progression if no audited year is effectively saturated, all eligible inquiries were already timely resourced, capacity cannot be defined reproducibly, lawful records are unavailable, or the shadow reserve necessarily reduces validated pilots by more than 10%. If those falsifiers do not fire, preregister a prospective stepped-wedge or matched-portfolio pilot comparing fixed-reserve, existing risk-tiered assurance, and no-fixed-reserve conditions; causal impact cannot be established by the records audit alone.","blocking_evidence":["No public evidence quantifies how often annual AI-pilot portfolios saturate live-use capacity while eligible nondeployment inquiries remain unfunded.","No evidence links lack of a protected portfolio share to later material findings after controlling for risk tier, proposal mix, organizational maturity, and existing assurance.","The 15% threshold has no empirical dose-response or optimization basis.","No comparative outcome evidence shows superiority to mandatory per-system testing, risk-weighted assurance budgets, or independent assurance funded outside the pilot portfolio.","Capacity units, inquiry eligibility, issue materiality, and validated-pilot status are not standardized.","Actual portfolio size and resource-equivalent opportunity cost are unknown, so recurring cost could fall outside the stated illustrative band.","Proprietary portfolio, assurance, staffing, vendor, and incident records are required for prevalence testing.","A prospective or naturally occurring comparison is required to estimate detected-issue and validated-pilot effects.","Live affected-community inquiry would require context-specific consent, compensation, privacy, accessibility, and safeguarding review.","Proprietary policies, non-English sources, paywalled standards, and fuller patent searches may contain an exact prior-art match."],"research_disposition":"PROBLEM_PREVALENCE_STUDY","world_novelty_boundary":"The search covered 17 English-language ordinary-web queries and ten independently opened direct sources across U.S. and UK government guidance and data, UN evaluation guidance, first-party AI policies and documentation, and primary research. It found adjacent but not exact public prior art. World novelty, patentability, freedom to operate, market size, realized impact, proprietary organizational policies, paywalled standards, non-English materials, unpublished practices, and comprehensive patent-family or claims searches remain unmeasured.","arm":"RETRIEVAL_FIRST","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":true,"material_progress_observed":true,"progress_targets":["Obtain authorized access to proprietary portfolio, assurance-demand, staffing, approval, finding, and outcome records for at least one organization and preferably multiple completed years.","Operationalize and test inter-rater reliability for pilot-capacity units, eligible nondeployment inquiry, issue materiality, validated-pilot status, saturation, diversion, and compliant release.","Establish whether saturated portfolios with unmet eligible inquiry demand visibly occur; falsify the problem if they do not.","Estimate the 15% reserve's opportunity cost and threshold sensitivity at 5%, 10%, 15%, and risk-weighted alternatives before any live allocation change.","Secure written confirmation of the portfolio authorizer, independent release-concurrence role, legal/privacy basis, data-retention rules, and community-participation safeguards.","If the retrospective prevalence gate passes, run a preregistered prospective stepped-wedge or matched-portfolio comparison against existing risk-tiered assurance and no-fixed-reserve practice.","Measure independently adjudicated material issues per proposed system, validated-pilot count, time to finding, diversion frequency, idle capacity, and inquiry completion; stop if issue detection does not improve or validated pilots fall by more than 10%.","Conduct targeted follow-up searching of proprietary policies, paywalled standards, non-English sources, and relevant patent claims before making any distinctiveness or adoption claim."],"reason":"Bounded web research verifies the importance and feasibility of independent predeployment assurance, identifies credible authorizers, and finds close structural and AI-governance analogues. It does not verify the candidate's specific saturation problem, the 15% threshold, incremental efficacy, or organization-specific cost. Those questions require proprietary records and prospective or naturally occurring comparisons, so further ordinary web search cannot close the decisive evidence gaps."}}