{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"bounded_rivalry_governance__criminology_forensic:P1:v0","cell_id":"bounded_rivalry_governance__criminology_forensic","search_queries":["site:nist.gov forensic science proficiency testing blind quality control errors laboratory","site:nist.gov forensic science human factors errors quality management OSAC","site:ojp.gov NIJ forensic laboratory needs funding backlog quality assurance","forensic science blind proficiency testing laboratory Houston study","forensic laboratory error reporting culture nonconformity underreporting quality management study","forensic science errors corrective action reporting incentives laboratory quality study","site:nist.gov OSAC proficiency testing forensic standard blind tests quality assurance","forensic proficiency testing declared test limitations examiner performance primary study","site:bja.ojp.gov forensic science improvement grants 2025 funding laboratory quality training equipment","site:bls.gov forensic science technicians median pay 2025","site:anab.ansi.org forensic testing laboratories proficiency testing ISO 17025","site:ilac.org P9 proficiency testing ISO IEC 17025 laboratory policy","Texas Forensic Science Commission authority proficiency testing forensic laboratories quality oversight","forensic science commission strategic plan laboratory quality incidents disclosure blind proficiency testing","regional forensic laboratory consortium proficiency testing collaborative challenge competition","forensic laboratory interlaboratory comparison competition leaderboard proficiency testing","\"Implementation of a Blind Quality Control Program in a Forensic Laboratory\" DOI","\"Implementing blind proficiency testing in forensic laboratories\" DOI","\"How do latent print examiners perceive proficiency testing\" DOI","\"Quality issue management and disclosure in forensic science\" publication date","10.1111/1556-4029.14259 publication date Hundl 2020","10.1111/1556-4029.70121 publication date","NIST OSAC Technical Series 0004 publication date human factors validation performance testing","BLS May 2025 occupational employment wage forensic science technicians release date"],"sources":[{"source_id":"S1","title":"Needs Assessment of Forensic Laboratories and Medical Examiner/Coroner Offices: A Report to Congress","publisher":"National Institute of Justice","url":"https://nij.ojp.gov/library/publications/needs-assessment-forensic-laboratories-and-medical-examinercoroner-offices","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2019","accessed_at":"2026-08-03","claims_supported":["Forensic providers expressed needs for stronger quality assurance and fewer preventable nonconformities.","The field faces insufficient funding, increased workloads, and interagency-communication challenges.","Training, equipment, and development resources are genuinely scarce."]},{"source_id":"S2","title":"Research and Evaluation in Publicly Funded Forensic Laboratories","publisher":"National Institute of Justice","url":"https://nij.ojp.gov/topics/forensics/research-and-evaluation-publicly-funded-forensic-laboratories","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2024-09-23","accessed_at":"2026-08-03","claims_supported":["NIJ has an identifiable program funding evaluations of laboratory protocols for accuracy, reliability, efficiency, and cost-effectiveness.","Eligible applicants include accredited state, regional, county, municipal, and tribal public forensic laboratories.","The program seeks practical evidence that can inform laboratory policy."]},{"source_id":"S3","title":"Implementation of a Blind Quality Control Program in a Forensic Laboratory","publisher":"Journal of Forensic Sciences / Wiley","url":"https://onlinelibrary.wiley.com/doi/full/10.1111/1556-4029.14259","source_class":"PRIMARY_RESEARCH","publication_date":"2020","accessed_at":"2026-08-03","claims_supported":["Houston Forensic Science Center implemented blind quality control in six forensic disciplines.","From 2015 through 2018, 901 of 973 submitted blind samples were completed and only 51 were recognized as blind samples.","Known-source samples inserted by an organizationally separate quality division can test a laboratory's broader workflow, not merely an analyst's final answer."]},{"source_id":"S4","title":"How Do Latent Print Examiners Perceive Proficiency Testing? An Analysis of Examiner Perceptions, Performance, and Print Quality","publisher":"Science & Justice / Elsevier","url":"https://pubmed.ncbi.nlm.nih.gov/32111284/","source_class":"PRIMARY_RESEARCH","publication_date":"2020-03","accessed_at":"2026-08-03","claims_supported":["A study of 322 latent-print examiners found conventional proficiency items were perceived as relatively easy.","Low error rates and high objective print-quality measures supported concerns that conventional tests were insufficiently challenging or representative.","Lower-quality, case-like items were perceived as more difficult, supporting matched realistic challenge construction."]},{"source_id":"S5","title":"Quality Issue Management and Disclosure in Forensic Science: A Survey of Practice and Perceptions","publisher":"Journal of Forensic Sciences / Wiley","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12424106/","source_class":"PRIMARY_RESEARCH","publication_date":"2025-07-07","accessed_at":"2026-08-03","claims_supported":["The 43-participant practitioner survey found inconsistent terminology and practices that impede cross-agency comparison and benchmarking of quality issues.","Only 55.8% agreed that enough was being done with collected quality-issue data.","Respondents described workload, management buy-in, reputation, resource, and authority conflicts, while valuing standardization, benchmarking, transparency, and destigmatization.","The authors conclude that negative quality culture can impede reporting and that a just culture is important for improvement."]},{"source_id":"S6","title":"Human Factors in Validation and Performance Testing of Forensic Science, OSAC Technical Series Publication 0004","publisher":"National Institute of Standards and Technology","url":"https://www.nist.gov/system/files/documents/2023/10/26/OSACTechSeriesPub_HF%20in%20Validation%20and%20Performance%20Testing%20of%20Forensic%20Science_March2020.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2020-03","accessed_at":"2026-08-03","claims_supported":["Blind studies can test whether performance observed in open studies resembles actual practice.","Realistic blind simulations can be burdensome and expensive and may require cooperation from submitting agencies.","Blind programs have been implemented at Houston Forensic Science Center, the Netherlands Forensic Institute, the U.S. Army Defense Forensic Science Center, and elsewhere.","Blinding checks, context management, realistic specimens, and discipline-specific feasibility are important design constraints."]},{"source_id":"S7","title":"Texas Forensic Science Commission: About Us","publisher":"Texas Judicial Branch","url":"https://www.txcourts.gov/fsc/about-us/","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["An identifiable state oversight body has authority over accreditation and practices intended to improve forensic-analysis quality.","The Commission may initiate educational investigations to advance forensic integrity and reliability.","Its membership includes scientists, a prosecutor, and a defense attorney, and it conducts collaborative education and development initiatives."]},{"source_id":"S8","title":"Table 1. National Employment and Wage Data from the Occupational Employment and Wage Statistics Survey by Occupation, May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.t01.htm","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026-05-15","accessed_at":"2026-08-03","claims_supported":["The May 2025 national mean wage for forensic science technicians was $38.08 per hour or $79,200 annually.","The national median hourly wage was $34.65.","These wages provide an official labor-cost anchor, but not prices for challenge materials, secure systems, legal review, or independent scoring."]}],"problem_evidence":{"support":"MODERATE","rationale":"The problem's general mechanism is visible: laboratories face scarce resources and quality-assurance needs; conventional proficiency tests may be unrealistically easy; and a recent practitioner survey documents inconsistent issue classification, workload and buy-in barriers, negative quality culture, and explicit tension between quality improvement and management concerns about reputation and resources. Evidence does not establish the candidate's stronger prevalence claim that regional development allocations routinely depend on low reported-error counts or that laboratories strategically withhold lessons, so that portion remains unverified.","source_ids":["S1","S4","S5"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"NIJ expressly funds evaluations of protocols in accredited public laboratories, and the Texas Forensic Science Commission is a concrete oversight body with authority to improve forensic practices and conduct educational initiatives. Practitioners surveyed in S5 value benchmarking and better use of quality-issue data. No source expresses demand for a competitive league, rankings, or resource-linked prizes; pull is for quality improvement, realistic testing, and benchmarking rather than this incentive design specifically.","source_ids":["S2","S5","S7"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Houston Forensic Science Center blind quality-control program","similarity":"Uses known-source mock submissions, organizationally separate challenge administration, realistic workflow insertion, known answers, and system-level evaluation across multiple forensic disciplines—the technical core of the proposed challenge packets and independent scoring.","remaining_difference":"It is a quality-control program, not a governed interlaboratory rivalry that publicly ranks teams or allocates scarce development credits across complementary achievements.","source_ids":["S3","S6"]},{"name":"Conventional forensic proficiency testing and interlaboratory comparison","similarity":"Already compares participant performance against predetermined answers and supports accreditation and quality assurance.","remaining_difference":"The proposed league adds balanced defect-characterization metrics, explicit false-positive and escalation penalties, equal resource caps, held-out reproducibility, portfolio awards, and a causal claim about redirecting resource competition; current tests may also be easier and less representative than casework.","source_ids":["S4","S6"]},{"name":"Cross-agency quality-issue classification and benchmarking","similarity":"Practitioners already seek standardized issue data for comparison, trend detection, improvement, and resource planning.","remaining_difference":"Standardized benchmarking is noncompetitive and does not test whether attaching controlled rivalry and development credits increases candid, reproducible weakness discovery.","source_ids":["S5"]}],"distinctive_claim_remaining":"Holding packets, time, tools, blinding, and scoring constant, framing the exercise as a safeguarded portfolio-award contest will produce more verified, reproducible, appropriately escalated defect characterizations than a non-ranked independent review, without increasing false-positive findings, leakage, rule violations, workload displacement, or suppression of ordinary quality reporting.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"Multi-discipline blind challenge programs and independent known-answer scoring are demonstrably implementable, and official guidance identifies workable blinding and context-management practices. The proposal wisely removes live-case consequences from the first step. Implementation becomes harder across laboratories and disciplines because realistic packet construction is burdensome, tests can be recognized, different workflows reduce comparability, and quality-issue terminology is not standardized. No external evidence validates the proposed composite score, portfolio-credit allocation, anti-collusion screen, incidental-case quarantine, or legal treatment of contest records.","source_ids":["S3","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Better discovery and communication of forensic-process weaknesses could affect evidence reliability and justice outcomes, while current quality and resource constraints are documented. Realized impact and effect size remain unmeasured.","source_ids":["S1","S5","S7"]},"stakeholder_pull":{"score":3,"rationale":"Public laboratories, NIJ, and an oversight commission are identifiable stakeholders with expressed quality-improvement and evaluation mandates, but none requests rivalry, rankings, or prize-linked allocation.","source_ids":["S2","S5","S7"]},"incremental_advantage":{"score":2,"rationale":"Blind QC, proficiency testing, realistic packets, known-answer scoring, and benchmarking already cover most operational value. The unproven increment is the resource-linked rivalry framing versus an otherwise identical non-ranked review.","source_ids":["S3","S4","S5","S6"]},"distinctiveness_plausibility":{"score":2,"rationale":"The incentive-redirection claim is contrastive and testable, but the intervention substantially collides with established proficiency-testing and blind-QC practice. No close evidence was found for allocating development credits through such a league, but world novelty was not assessed.","source_ids":["S3","S4","S6"]},"technical_implementability":{"score":4,"rationale":"Known-source packets, separated quality staff, blind insertion, access controls, fixed time, and independent scoring are feasible and partly demonstrated. Cross-discipline realism and blinding remain labor-intensive.","source_ids":["S3","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"A commission or laboratory consortium can plausibly authorize a synthetic educational study, and NIJ can fund partnered evaluations. Authority to use resulting ranks for actual resource allocation would require separate procurement, labor, accreditation, records, and due-process review.","source_ids":["S2","S7"]},"evidence_readiness":{"score":4,"rationale":"The remaining claim supports a preregistered matched-packet experiment with measurable outcomes and an active comparator. Standardizing defect classes and obtaining enough independent teams are the main preparation gaps.","source_ids":["S4","S5","S6"]},"safety_net_benefit":{"score":4,"rationale":"Synthetic or known-source packets, no stakes, no public rankings, independent safeguards, and stopping before real awards sharply limit case and employment harm. Ranking pressure, data leakage, stigma, and incidental real-case information still require formal controls.","source_ids":["S5","S6","S7"]},"scalability":{"score":2,"rationale":"A few laboratories can run the protocol, but realistic packet construction, workflow differences, scarce quality staff, and the expense of blinding make multi-discipline regional scaling difficult. Common issue classification is not yet mature.","source_ids":["S1","S5","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"One no-stakes, single-discipline matched-packet study with 8 volunteer teams, two experimental conditions, preregistration, packet preparation, independent scoring, safeguard review, and a short analysis report.","confidence":"MODERATE","assumptions":["Each team uses two analysts for approximately 12-16 hours.","Approximately 300-600 total specialist hours cover analysts, packet design, scoring, study management, and safeguard review.","Existing laboratory tools and a basic controlled file environment are reused.","Resource-equivalent labor is valued from the $38.08 mean forensic-technician hourly wage, with an assumed allowance for senior staff and overhead."],"source_ids":["S6","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Design and validate reusable packets and rubrics for up to three disciplines; create governance, appeals, blinding, access-log, conflict, and quarantine procedures; configure a secure submission environment; and conduct scorer training.","confidence":"LOW","assumptions":["Existing laboratory instruments and buildings are used.","One to three staff-equivalents contribute over roughly three to six months.","No actual development credits or equipment purchases are included.","Legal, records, labor, and accreditation reviews are limited to the controlled program."],"source_ids":["S3","S5","S6","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"First regional scored round involving approximately 6-10 laboratories and up to three disciplines, including validated packets, independent administration and audit, appeals, monitoring, evaluation, and administration of non-case development credits.","confidence":"LOW","assumptions":["Four to eight staff-equivalents are distributed among participating laboratories, designers, scorers, security, and governance roles.","Existing instruments are used; the value of separately awarded instrument upgrades is excluded and must be budgeted by the authorizer.","Participant labor displacement and loaded overhead are counted as resource-equivalent costs.","Realistic packet construction remains the largest uncertain nonlabor cost."],"source_ids":["S1","S3","S6","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"One or two regional rounds per year, refreshed held-out packets, independent scoring and audit, security and records administration, appeals, post-round evaluation, and rule revision.","confidence":"LOW","assumptions":["The program covers 6-10 laboratories and up to three disciplines.","Packet content must be refreshed to control leakage and construction-signature learning.","Two to six ongoing staff-equivalents plus participating-team time are required.","Actual training places, validation time, and upgrade-credit values are excluded from operating cost and reported separately."],"source_ids":["S5","S6","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"External evidence supports scarce laboratory resources, persistent quality-assurance needs, shortcomings in ordinary proficiency tests, inconsistent issue reporting, and reputation/resource conflicts. The frequency of strategic underreporting tied specifically to regional allocations remains uncertain but is not required to justify a bounded mechanism test.","source_ids":["S1","S4","S5"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"The Texas Forensic Science Commission is a concrete oversight authority with quality-improvement and educational functions, while NIJ funds partnered evaluations in accredited regional and other public laboratories. Neither has endorsed this proposal.","source_ids":["S2","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining increment is explicitly the causal effect of governed, resource-linked rivalry relative to an otherwise matched non-ranked review, with verified findings, false positives, calibration, reproducibility, leakage, violations, pressure, and workload displacement as outcomes.","source_ids":["S3","S4","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A single-discipline, no-stakes, matched-packet randomized study can test the claim without live evidence, public rankings, employment consequences, or actual resource awards.","source_ids":["S3","S4","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"For the evidence step only, synthetic or known-source materials, volunteer participation, no consequential rankings, no case action, independent review, and immediate stopping on leakage or unequal exposure bound foreseeable harm. This gate does not authorize actual awards or operational case access.","source_ids":["S6","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The bands state participant count, discipline count, labor, packet, scoring, governance, and excluded prize assumptions; official wages anchor labor and NIST confirms substantial simulation burden. Vendor, legal, and secure-platform quotations remain unavailable, so confidence falls after the first study.","source_ids":["S6","S8"]}},"next_evidence_step":"Recruit 8 volunteer teams from at least four accredited laboratories for a preregistered, no-stakes, single-discipline matched-packet study. Randomize teams to (A) the proposed governed contest framing or (B) a non-ranked independent-review framing; use identical de-identified synthetic/known-source packets, analysis-hour limits, tools, scorer access, feedback timing, and confidentiality rules. Keep laboratory and analyst identities from scorers, withhold all rankings and resource awards, and include a held-out reproduction packet. Primary outcomes are correctly characterized seeded defects and held-out reproducibility; secondary outcomes are false-positive findings, confidence calibration, appropriate escalation, hours, routine-work displacement, leakage indicators, rule violations, appeals, and participant-reported pressure and candor. Falsify the incremental claim if arm A does not outperform arm B on verified defect characterization, or if any gain is accompanied by materially worse false positives, reproducibility, leakage, violations, workload displacement, pressure, or ordinary quality reporting. Stop and void the study if blinding, packet equivalence, access containment, voluntary participation, or quarantine controls fail.","blocking_evidence":["No comparative field evidence shows that competitive framing improves verified weakness discovery beyond an otherwise identical non-ranked review.","No prevalence data connect reported-error counts or turnaround records to regional allocations of validation time, training places, or upgrade credits.","The composite score's construct validity, weights, and resistance to overcalling and packet-signature gaming are unknown.","The legal and administrative treatment of rankings, contest records, incidental case information, labor displacement, appeals, and later resource awards has not been reviewed for a named jurisdiction.","Cross-discipline packet equivalence and common quality-issue classification are not established.","Cost estimates lack quotations for realistic packet production, secure administration, independent scoring, and legal review."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This assessment only compared the proposal with eight directly reviewed sources located through bounded web searches. It did not conduct a systematic literature review, patent search, product-market census, standards-completeness review, freedom-to-operate analysis, or worldwide practice survey. World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain written participation and protocol authority from a named oversight body and at least four accredited laboratories.","Complete jurisdiction-specific legal, records, accreditation, labor, case-quarantine, and adverse-event review for the no-stakes study.","Preregister and execute the matched contest-versus-non-ranked study with independent scoring and held-out reproduction.","Demonstrate improved verified defect characterization without worse false positives, leakage, rule violations, workload displacement, participant pressure, or ordinary quality reporting.","Validate the scoring rubric across packet difficulty and laboratory resource levels, including tests for packet-construction signatures and concentrated-specialist effects.","Replace broad cost assumptions with recorded staff time and quotations before considering a consequential pilot.","Do not attach actual training, validation, or equipment credits until the no-stakes mechanism test and safeguard review pass."],"reason":"Bounded web research establishes a meaningful quality-management problem, credible authorizers and funders, and technically feasible adjacent practice, but it also shows substantial collision with existing blind QC and proficiency testing. The only important remaining advantage—the causal benefit and safety of adding governed, resource-linked rivalry—cannot be resolved by further web search. It requires proprietary operational participation and a live but no-stakes comparative study; therefore the research loop should stop for empirical work."},"proposal_index":1}