{"schema_version":1,"research_id":"eoa_inverse_innovation_exp04_external_evaluation_20260802","source_assessment_id":"negative_space_design__tech_ethics_ai_governance:RETRIEVAL_FIRST:v0","cell_id":"negative_space_design__tech_ethics_ai_governance","search_queries":["site:airc.nist.gov airrmf core independent review testing information sharing","site:gov.uk industry temperature check barriers enablers AI assurance financial resources skills","site:whitehouse.gov OMB M-25-21 AI governance testing impact assessment","site:gao.gov AI accountability framework monitoring human oversight","site:deloitte.com innovation portfolios public sector organizations 70 20 10 published","site:anthropic.com responsible scaling policy evaluations external review red teaming policy 2025","site:bls.gov/news.release/ocwage.t01.htm May 2025 occupational wage 2026","site:learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai/govern sandbox independent reviews red team"],"sources":[{"source_id":"S1","title":"AI RMF Core","publisher":"National Institute of Standards and Technology","url":"https://airc.nist.gov/airmf-resources/airmf/5-sec-core/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2023-01-26","accessed_at":"2026-08-02","claims_supported":["AI systems should be tested before deployment, with documented metrics, benchmarks, uncertainty, and repeatable evaluation processes.","Independent review can improve testing and reduce internal bias or conflicts of interest; affected communities and external actors may be consulted.","NIST calls for regular allocation of risk-management resources according to assessed risks and organizational capabilities, but prescribes no fixed portfolio reserve."]},{"source_id":"S2","title":"Industry temperature check: barriers and enablers to AI assurance","publisher":"Centre for Data Ethics and Innovation and UK Department for Science, Innovation and Technology","url":"https://www.gov.uk/government/publications/industry-temperature-check-barriers-and-enablers-to-ai-assurance/industry-temperature-check-barriers-and-enablers-to-ai-assurance","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2022-12-07","accessed_at":"2026-08-02","claims_supported":["Stakeholders reported that insufficient financial resources, skills, senior-management support, guidance, and internal or external demand impede AI assurance.","Standards and assurance work can be labor-intensive and costly, especially for smaller organizations.","The engagement supports a resource-constrained assurance problem and demand for operational guidance, but did not measure saturated pilot portfolios or demand for a 15% reserve."]},{"source_id":"S3","title":"OMB Memorandum M-25-21: Accelerating Federal Use of AI through Innovation, Governance, and Public Trust","publisher":"Executive Office of the President, Office of Management and Budget","url":"https://www.whitehouse.gov/wp-content/uploads/2025/02/M-25-21-Accelerating-Federal-Use-of-AI-through-Innovation-Governance-and-Public-Trust.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-04-03","accessed_at":"2026-08-02","claims_supported":["Federal agencies must designate Chief AI Officers, establish AI governance boards, track AI spending and use cases, and allocate appropriate resources and responsibilities.","For high-impact AI, agencies must conduct predeployment testing and impact assessments, obtain independent review, document costs and risk acceptance, and consult affected groups where appropriate.","The memorandum identifies agency heads, Chief AI Officers, governance boards, and risk-accepting officials as credible authorizers, but it exempts qualifying pilots from some minimum practices and does not require a fixed annual nondeployment reserve."]},{"source_id":"S4","title":"Costed Evaluation Plan Guidance, Tools and Templates","publisher":"United Nations Population Fund","url":"https://www.unfpa.org/sites/default/files/admin-resource/CEPlan%20Guidance%2C%20Tools%20and%20Templates-307.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-04","accessed_at":"2026-08-02","claims_supported":["UNFPA places evaluation funding within an annual country-office resource ceiling, ring-fences it exclusively for evaluation, and requires formal authority for diversion or reallocation.","This is a close structural analogue for protected capacity and controlled release in a neighboring domain, although it is not an AI-pilot rule.","Indicative evaluation budgets range from roughly $40,000 to $150,000 depending on portfolio and context, supporting the order of magnitude for a bounded records-and-review study."]},{"source_id":"S5","title":"Innovation portfolios for public sector organizations","publisher":"Deloitte Insights","url":"https://www.deloitte.com/us/en/insights/industry/government-public-sector-services/innovation-portfolios-public-sector-organizations.html","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2018","accessed_at":"2026-08-02","claims_supported":["Portfolio management commonly allocates explicit shares across core, adjacent, and transformational work, including the 70-20-10 pattern and public-sector examples with protected higher-risk categories.","Portfolio investments may be justified by learning value, and longer-horizon work may require agency-level earmarking and protected time.","Percentage allocation is established practice, but the protected categories pursue innovation rather than nondeployment AI-harm inquiry."]},{"source_id":"S6","title":"Anthropic’s Responsible Scaling Policy","publisher":"Anthropic","url":"https://www.anthropic.com/responsible-scaling-policy","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"2026-07-08","accessed_at":"2026-08-02","claims_supported":["Anthropic maintains recurring risk evaluations, Risk Reports, external-review provisions, oversight arrangements, and authority to pause development when appropriate.","The policy demonstrates that technically sophisticated AI organizations can institutionalize evaluations, independent challenge, and deployment restraint.","It is capability-threshold governance for frontier models, not a stated minimum share of an annual multi-pilot portfolio reserved for harm inquiry."]},{"source_id":"S7","title":"National Employment and Wage Data by Occupation, May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.t01.htm","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026-05-15","accessed_at":"2026-08-02","claims_supported":["National annual mean wages include $111,490 for software quality-assurance analysts and testers and $153,930 for computer and information research scientists.","These benchmarks support labor-equivalent cost estimates but exclude benefits, overhead, outside experts, community compensation, compute, data preparation, and displaced pilot value."]},{"source_id":"S8","title":"Guidance to Set Up Your Organization's AI Governance Process","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai/govern","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-04-10","accessed_at":"2026-08-02","claims_supported":["Microsoft recommends cross-functional risk assessment, documented governance policies, sandbox environments for initial experiments, validation, model review, benchmark datasets, and monitoring.","The guidance supplies deployable workflow and tooling analogues for nonproduction inquiry but no annual fixed-percentage reserve.","Implementation can build on ordinary governance, portfolio accounting, sandbox, review, and documentation systems rather than requiring a novel technical platform."]}],"problem_evidence":{"support":"MODERATE","rationale":"Official guidance establishes that predeployment testing, independent review, stakeholder input, and adequate risk resources matter, while UK stakeholder research directly reports resource, skill, management-support, and guidance barriers. This makes insufficient assurance capacity visible and consequential. However, no opened source measures how often annual AI-pilot portfolios are fully committed to live-use work, whether saturation causes later material findings, or whether a protected percentage is the missing cause.","source_ids":["S1","S2","S3"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"U.S. federal agency heads, Chief AI Officers, governance boards, CFOs, and risk-accepting officials are identifiable authorizers with explicit duties to resource, test, independently review, and govern AI. UK industry participants expressed demand for resources and concrete operational guidance. No source expresses demand for the candidate's fixed 15% reserve, so pull is for the underlying assurance function rather than the proposed allocation mechanism.","source_ids":["S2","S3"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"UNFPA ring-fenced evaluation funding","similarity":"Evaluation resources are protected inside an annual organizational ceiling, unavailable for delivery work, tracked separately, and subject to formal release or reallocation authority.","remaining_difference":"The mechanism funds international-development programme evaluation rather than reserving a minimum share of AI-pilot capacity for predeployment harm inquiry.","source_ids":["S4"]},{"name":"Percentage-based innovation portfolio allocation","similarity":"Organizations allocate explicit annual portfolio shares to protected learning or transformational categories that would otherwise compete with core delivery.","remaining_difference":"The allocated categories pursue innovation and opportunity discovery, not red-teaming, shadow evaluation, or affected-community harm inquiry.","source_ids":["S5"]},{"name":"Independent and sandboxed AI assurance","similarity":"NIST, OMB, and Microsoft specify predeployment testing, independent review, cross-functional governance, stakeholder input, sandboxes, documentation, and approval decisions.","remaining_difference":"These practices are applied per system or risk tier and do not reserve a stated fraction of organization-year pilot capacity before projects are selected.","source_ids":["S1","S3","S8"]},{"name":"Anthropic Responsible Scaling Policy","similarity":"The policy institutionalizes recurring evaluations, external review, risk reporting, oversight, and possible pauses in development or deployment.","remaining_difference":"It governs frontier-model capability thresholds rather than a fixed annual reserve spanning an organization's pilot portfolio and affected-community inquiries.","source_ids":["S6"]}],"distinctive_claim_remaining":"In a bounded contrast to the opened sources, the remaining claim is that an organization-year rule reserving at least 15% of total AI-pilot capacity exclusively for nondeployment harm inquiry before individual pilots are approved, protecting that capacity from live-use commitments, and allowing release only under documented independent-concurrence criteria will increase independently adjudicated material predeployment issues detected per proposed system without reducing validated pilots by more than 10%.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"The component workflows already exist separately: annual portfolio allocation, ring-fenced accounting, formal reallocation approval, sandbox experimentation, predeployment testing, independent review, affected-party consultation, and documented risk acceptance. A records-only shadow allocation is technically straightforward. Feasibility remains uncertain around defining a common capacity unit, obtaining lawful access to portfolio and assurance records, independently adjudicating issue materiality, preventing relabeling, compensating affected participants, and absorbing the 15% opportunity cost.","source_ids":["S1","S3","S4","S5","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Earlier detection could prevent exposure and lock-in for consequential systems, but prevalence, causal effect, and realized harm reduction are unmeasured.","source_ids":["S1","S2","S3"]},"stakeholder_pull":{"score":3,"rationale":"Authorizers and expressed demand for assurance resources and guidance are visible, but no stakeholder requests a fixed 15% reserve.","source_ids":["S2","S3"]},"incremental_advantage":{"score":2,"rationale":"Protection from portfolio competition is plausible, but no evidence shows it outperforms flexible risk-tiered assurance or project-specific funding.","source_ids":["S1","S3","S4"]},"distinctiveness_plausibility":{"score":3,"rationale":"The exact AI-specific combination was not found, although all major components and close ring-fencing structures are established.","source_ids":["S4","S5","S6","S8"]},"technical_implementability":{"score":4,"rationale":"Accounting controls, sandboxes, reviews, testing, and governance workflows are available; the main difficulty is organizational measurement rather than new technology.","source_ids":["S3","S4","S8"]},"adoption_authority_feasibility":{"score":3,"rationale":"Federal agency governance boards and Chief AI Officers are identifiable authorities, but budget law, bargaining, procurement, and organization-specific delegations must be checked.","source_ids":["S3"]},"evidence_readiness":{"score":2,"rationale":"The first study is bounded, but requires proprietary historical records, a partner organization, and independent materiality adjudication.","source_ids":["S3","S4"]},"safety_net_benefit":{"score":3,"rationale":"The reserve could preserve inquiry before exposure, provided it supplements rather than displaces mandatory safety, privacy, security, accessibility, and incident-response work.","source_ids":["S1","S3"]},"scalability":{"score":3,"rationale":"A percentage rule and standardized records can scale across portfolios, but small portfolios, continuous deployment, procurement, and uneven assurance capacity require adaptations.","source_ids":["S2","S4","S5"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"One completed organization-year records study: reconstruct capacity, perform a preregistered shadow allocation, code eligible inquiries, independently adjudicate material findings, and report feasibility and counterfactual outcomes.","confidence":"MODERATE","assumptions":["Approximately 0.4-1.0 combined FTE-year across analyst, assurance reviewer, portfolio owner, and legal/privacy support.","Existing portfolio and assurance records are usable without new data collection from exposed people.","External reviewer, secure data handling, and overhead are included.","UNFPA's $40,000-$150,000 indicative evaluation budgets and BLS professional wages are order-of-magnitude anchors, not prices for this exact study."],"source_ids":["S4","S7"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Design the capacity denominator, eligibility and materiality rules, separate accounting codes, approval and release workflow, audit trail, templates, training, and participation safeguards for one organization.","confidence":"MODERATE","assumptions":["Existing governance and portfolio systems can be configured rather than replaced.","Scope covers policy and workflow setup, not the reserved inquiry work itself.","Roughly 0.3-1.0 combined FTE-year plus legal, privacy, and change-management review."],"source_ids":["S3","S7","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Operate the reserve during the first annual cycle for a medium-sized portfolio, including red-teaming, shadow evaluations, independent review, community inquiry, and the opportunity cost of displaced pilot capacity.","confidence":"LOW","assumptions":["A medium portfolio's annual pilot resource-equivalent is roughly $2 million-$6 million, making 15% equal to $300,000-$900,000.","Some reserve work uses internal staff; specialized experts and affected-community participation require additional support.","The reserve is a reallocation within the portfolio, not necessarily new cash.","Portfolio size and inquiry intensity can move actual cost outside this band."],"source_ids":["S1","S2","S3","S7"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain a 15% reserve, execute eligible inquiries, compensate external or affected participants, administer release decisions, audit classification and diversion, and report outcomes each year for a medium-sized organization.","confidence":"LOW","assumptions":["Annual pilot capacity remains roughly $2 million-$6 million in resource-equivalent terms.","Startup design work is excluded, while recurring administration and independent review are included.","Costs scale directly with portfolio capacity, so very small or frontier-scale organizations fall in different bands."],"source_ids":["S2","S3","S7"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official and stakeholder evidence confirms that predeployment assurance matters and that resources, skills, and management support are barriers; the narrower saturation prevalence remains unmeasured.","source_ids":["S1","S2","S3"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Federal agency heads, Chief AI Officers, CFOs, risk-accepting officials, and AI governance boards have explicit governance, resourcing, review, and tracking responsibilities and could authorize a portfolio rule within applicable budget authority.","source_ids":["S3"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The candidate specifies a 15% reserve, eligible uses, protection and release rules, an issue-detection outcome, a validated-pilot comparator, and numerical falsifiers; its advantage over flexible assurance is testable.","source_ids":["S1","S3","S4","S5"]},"bounded_next_evidence_step":{"status":"YES","reason":"A single completed organization-year can be reconstructed and shadow-allocated without changing any deployment or budget decision.","source_ids":["S3","S4"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The records-only study is reversible and can halt for unlawful access, sensitive-data exposure, or inability to adjudicate materiality. Any later live adoption remains contingent on budget authority, privacy, human-subject, labor, procurement, accessibility, and participation safeguards.","source_ids":["S1","S3"]},"credible_cost_scope_and_range":{"status":"YES","reason":"All four estimates state their included scope, portfolio-size assumptions, labor anchors, opportunity-cost treatment, and low-confidence scaling limitations.","source_ids":["S4","S7"]}},"next_evidence_step":"With one willing organization, preregister a records-only study of one completed annual AI-pilot portfolio. Define the capacity denominator, eligible nondeployment inquiries, materiality rubric, and missing-data rules before examining outcomes. Reconstruct (A) the actual flexible or risk-tiered allocation and (B) a 15% protected shadow allocation applied before proposal selection; if feasible, add another completed year without a reserve as a descriptive comparator. Have two reviewers, including one independent of pilot selection, adjudicate whether existing red-team, shadow-evaluation, audit, or affected-community records contain material predeployment issues and measure agreement. Compare material issues per proposed system, validated-pilot count, inquiry completion, and simulated release-rule compliance. Falsify efficacy if the shadow reserve does not increase independently adjudicated material issues per proposed system or reduces validated pilots by more than 10%; falsify the mechanism's feasibility if capacity cannot be consistently measured, eligible demand cannot occupy the reserve, or ordinary project work cannot be distinguished from inquiry. Treat any positive result as feasibility and measurement evidence, not causal confirmation, and make no live deployment or budget changes.","blocking_evidence":["No opened source estimates the prevalence of organization-years with effectively saturated live-use pilot capacity and no protected inquiry capacity.","No comparative evidence shows that a fixed reserve improves material issue detection over flexible risk-tiered, per-project, or three-lines assurance.","The 15% threshold and 10% validated-pilot tolerance lack empirical calibration.","Historical portfolio, staffing, evaluation, issue, and approval records are proprietary and may be incomplete or legally restricted.","A reliable cross-project capacity unit and an independently reproducible materiality rubric have not been demonstrated.","Demand, willingness to adopt, and budget authority for this exact reserve have not been expressed by a named organization.","Costs for live affected-community inquiry, specialized red teams, and displaced pilot value vary materially by portfolio.","Continuous deployment, procurement, model updates, and very small portfolios fall outside the proposed annual-pilot denominator."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The exact combination was not found in eight opened direct sources from multiple publishers, but this is only a bounded ordinary-web contrast. World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured. Proprietary organizational policies, full patent-family and claims searches, paywalled standards, non-English materials, unpublished practice, and implementations after the accessed dates may contain the same mechanism.","arm":"RETRIEVAL_FIRST","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":true,"material_progress_observed":true,"progress_targets":["Secure a partner and lawful access to one completed annual AI-pilot portfolio with proposal, capacity, assurance, finding, and approval records.","Preregister a stable capacity denominator, eligible-inquiry taxonomy, materiality rubric, comparators, missing-data rules, and the 15% and 10% falsifiers.","Demonstrate acceptable independent-reviewer agreement on material findings and distinguish inquiry from relabeled development or mandatory compliance work.","Complete the actual-versus-shadow comparison and quantify issue-detection change, validated-pilot displacement, eligible-demand utilization, and sensitivity to alternative reserve shares.","Obtain written confirmation from the relevant portfolio and assurance authorities that a live reserve could be authorized without displacing mandatory safeguards or violating budget, privacy, labor, procurement, or participation rules."],"reason":"Bounded web research has established the underlying problem, credible authorities, adjacent prior art, and implementable component workflows, but cannot determine portfolio saturation, calibrate the threshold, or test incremental efficacy. Those questions require proprietary organizational records, independent adjudication, and eventually live or quasi-experimental evidence."}}