{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"representation_independent_interface_contract__criminology_forensic:P3:v0","cell_id":"representation_independent_interface_contract__criminology_forensic","search_queries":["site:nist.gov forensic science blind proficiency testing laboratory recommendations","blind proficiency testing forensic laboratory case management software clues study","Houston Forensic Science Center blind proficiency testing program publication","ANAB ISO 17025 proficiency testing forensic laboratories blind","site:nist.gov OSAC blind proficiency testing standard forensic laboratory 2024","site:ojp.gov blind proficiency testing forensic laboratory LIMS","blind proficiency testing software LIMS forensic laboratory product","forensic laboratory blind quality control LIMS hidden flag","\"Development and Implementation of an Effective Blind Proficiency Testing Program\" full text","site:anab.ansi.org AR 3125 proficiency testing forensic requirements pdf","site:fbi.gov quality assurance standards forensic DNA proficiency testing 2025 pdf","site:nationalacademies.org forensic science blind proficiency testing report 2009","Perceptions of blind proficiency testing among latent print examiners full text 2022","site:forensicstats.org blind proficiency testing survey laboratory directors LIMS features","forensic blind proficiency testing participant privacy legal employment consent","blind proficiency forensic laboratory ethics labor legal covert testing","doi 10.1016/j.fsisyn.2020.01.002 Implementing blind proficiency testing forensic laboratories","doi 10.1111/1556-4029.14259 full text","doi 10.1111/1556-4029.14269 Wiley","doi 10.1016/j.scijus.2022.12.005 full text pdf","Hundl blind quality control program Table 4 costs 2015 2018 HFSC","\"Implementation of a Blind Quality Control Program\" \"Cost\" digital forensics 2490","HFSC blind quality control program annual cost personnel hours"],"sources":[{"source_id":"S1","title":"Implementation of a Blind Quality Control Program in a Forensic Laboratory","publisher":"Journal of Forensic Sciences / Wiley","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC7317955/","source_class":"PRIMARY_RESEARCH","publication_date":"2019-12-24","accessed_at":"2026-08-03","claims_supported":["HFSC implemented blind quality control across six forensic disciplines and processed 901 completed cases through 2018.","Fifty-one completed blind cases were recognized by analysts, with clues including case numbering, handwriting, packaging, chain-of-custody information, submission details, and unrealistic materials.","HFSC's replacement LIMS could mask blind samples from everyone except the Quality Division and included a request portal that could submit on behalf of an authorized officer.","The program found process deficiencies that led to preventive changes, demonstrating potential quality-system impact.","Historical supply expenses ranged from negligible amounts to about $28,901 per discipline-year, while the authors estimated that maintaining the overall program required at least two dedicated full-time staff."]},{"source_id":"S2","title":"Implementing Blind Proficiency Testing in Forensic Laboratories: Motivation, Obstacles, and Recommendations","publisher":"Forensic Science International: Synergy / Elsevier","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC7552087/","source_class":"PRIMARY_RESEARCH","publication_date":"2020","accessed_at":"2026-08-03","claims_supported":["Laboratory directors and quality managers expressed wider interest in blind testing and general agreement that its benefits justified investment.","Participants identified realistic case creation, cost, submission, tracking, accidental release, metrics, and institutional culture as implementation barriers.","The paper specifically recommends a QA-visible, examiner-hidden LIMS flag and notes that laboratories may need to select a suitable LIMS or build an internal tracker.","Synthetic results and profiles must be prevented from reaching prosecutors, submitting agencies, NDIS, or NGI.","Implementation requires a senior institutional champion rather than unilateral action by a QA manager or discipline chief."]},{"source_id":"S3","title":"Development and Implementation of an Effective Blind Proficiency Testing Program","publisher":"Journal of Forensic Sciences / Wiley","url":"https://pubmed.ncbi.nlm.nih.gov/31922611/","source_class":"PRIMARY_RESEARCH","publication_date":"2020-01-10","accessed_at":"2026-08-03","claims_supported":["Harris County Institute of Forensic Sciences independently adopted blind testing in 2015.","Its stated objective was to test the whole system from evidence receipt through report release.","The program was feasible but required careful planning, trial and error, continuing assessment, and substantial logistical work."]},{"source_id":"S4","title":"Perceptions of Blind Proficiency Testing Among Latent Print Examiners","publisher":"Science & Justice / Elsevier","url":"https://www.sciencedirect.com/science/article/pii/S1355030622001733","source_class":"PRIMARY_RESEARCH","publication_date":"2023","accessed_at":"2026-08-03","claims_supported":["A survey of 338 practicing latent-print examiners found overall ambivalent views but likely receptiveness to blind testing.","Examiners working in laboratories with blind-testing experience viewed it significantly more positively.","The study identifies examiner acceptance and workload as adoption considerations rather than assuming uniform stakeholder support."]},{"source_id":"S5","title":"OSAC 2022-S-0012 Standard for Proficiency Testing in Friction Ridge Examination","publisher":"National Institute of Standards and Technology","url":"https://www.nist.gov/system/files/documents/2021/10/04/OSAC%202022-S-0012%20Standard%20for%20Proficiency%20Testing%20in%20Friction%20Ridge%20Examination_OPEN%20COMMENT_STRP%20Version.pdf","source_class":"STANDARD","publication_date":"2021-10-04","accessed_at":"2026-08-03","claims_supported":["The proposed standard states that blind testing is more robust than non-blind testing.","It requires avoiding subtle cues that could reveal expected results.","Test validation, acceptable-performance criteria, methods, and allowable deviations are to be documented before administration.","It requires evaluation of both individual performance and the forensic service provider's quality system."]},{"source_id":"S6","title":"Quality Assurance Standards for Forensic DNA Testing Laboratories","publisher":"Federal Bureau of Investigation","url":"https://le.fbi.gov/file-repository/forensic-qas-070125.pdf","source_class":"STANDARD","publication_date":"2025-07-01","accessed_at":"2026-08-03","claims_supported":["The FBI requires documented quality systems, proficiency-test records, discrepancy handling, corrective action, and defined technical-leader authority for covered DNA laboratories.","Covered analytical and interpretation software must undergo documented functional and reliability testing, with regression testing for major revisions, before implementation.","Reports, case files, DNA records, databases, and personally identifiable information are subject to confidentiality and controlled-release procedures.","The required external DNA proficiency program remains open rather than blind, so the candidate would supplement rather than replace that mandatory program."]},{"source_id":"S7","title":"AR 3125: Accreditation Requirements for Forensic Testing and Calibration Laboratories","publisher":"ANSI National Accreditation Board","url":"https://anab.ansi.org/standard/ar-3125/","source_class":"STANDARD","publication_date":"2023","accessed_at":"2026-08-03","claims_supported":["ANAB identifies AR 3125 as the forensic-specific supplement to ISO/IEC 17025:2017.","An accredited laboratory's management and quality structure is a credible authorization pathway, but this page does not mandate the proposed capsule or blind-testing architecture."]},{"source_id":"S8","title":"WK93533: New Practice for Intralaboratory Blind Quality Control Programs Specific to Seized-Drugs Analysis","publisher":"ASTM International","url":"https://www.astm.org/membership-participation/technical-committees/workitems/workitem-wk93533","source_class":"STANDARD","publication_date":"2025-01-21","accessed_at":"2026-08-03","claims_supported":["ASTM is developing minimum requirements for seized-drug intralaboratory blind quality-control programs.","The work-item rationale explicitly states that minimum implementation requirements were absent, evidencing a standards gap but also active adjacent standardization.","The work item concerns program requirements, not a representation-independent software interface or cross-implementation conformance oracle."]}],"problem_evidence":{"support":"STRONG","rationale":"Observed practice directly confirms that blind assignments can be revealed by participant-visible representation details: HFSC documented recognition from numbering, handwriting, packaging, chain of custody, submission patterns, and materials, and responded by adding LIMS masking. Blind programs also found process defects and are considered more representative than declared tests. The narrower allegation that nominally compatible software replacements currently diverge on submission immutability or lifecycle states was not directly documented and remains an evidence gap.","source_ids":["S1","S2","S3","S5"]},"stakeholder_evidence":{"support":"STRONG","rationale":"HFSC and Harris County are identifiable adopters, while the CSAFE-convened laboratory leaders expressed interest and identified the need for examiner-hidden LIMS tracking. OSAC calls blind testing more robust, ASTM is developing program requirements, and surveyed examiners appeared receptive, although acceptance was not uniformly strong. A particular laboratory has not committed to adopting this capsule contract.","source_ids":["S1","S2","S3","S4","S5","S8"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"HFSC blind-QC workflow with a QA-only LIMS flag and masked request portal","similarity":"It already hides the test classification from analysts, preserves an ordinary-case appearance, routes synthetic cases through normal workflows, restricts knowledge to quality staff, and tracks results.","remaining_difference":"The publication does not describe a representation-independent abstract state machine, explicit submission-finality and disclosure laws, a reusable cross-implementation black-box suite, mutation testing, or a systematic transcript-level leakage oracle.","source_ids":["S1"]},{"name":"Harris County whole-pipeline blind proficiency program","similarity":"It independently demonstrates that disguised proficiency cases can exercise the workflow from evidence receipt to report release and that institutional implementation is feasible.","remaining_difference":"No evidence was found that it defines implementations through one role-sensitive behavioral contract or accepts replacement software through shared lifecycle and leakage conformance tests.","source_ids":["S3"]},{"name":"OSAC/FBI/ANAB/ASTM proficiency, validation, and accreditation frameworks","similarity":"These frameworks require documented quality systems, predefined evaluation rules, proficiency records, software validation in covered contexts, confidentiality, and accountable laboratory leadership; OSAC and ASTM directly address blind testing.","remaining_difference":"They do not specify the proposed opaque capsule interface, complete role-indexed lifecycle, cross-LIMS substitutability rule, or ordinary-case-versus-exercise observable-transcript comparator.","source_ids":["S5","S6","S7","S8"]}],"distinctive_claim_remaining":"Compared with an examiner-hidden LIMS flag plus procedural walkthroughs, a predeclared role-sensitive state-machine contract and implementation-independent lifecycle/leakage oracle will reject replacements that otherwise appear field-compatible but disclose exercise identity, permit forbidden post-submission changes, or produce divergent abstract states. This is falsified in the bounded sandbox if the contract accepts a seeded faulty adapter, two passing implementations diverge on any mandatory abstract state, or participant-role reviewers identify exercises from declared software-visible transcripts materially above chance after tolerated variables are removed.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"HFSC proves that QA-only flags, masked requests, synthetic workflows, and restricted program administration are technically and operationally achievable, and FBI standards provide analogous authority, confidentiality, validation, and recordkeeping controls. The proposed capsule, generated operation sequences, transcript normalization, and cross-implementation leakage oracle have not been built or validated in a forensic LIMS. Site-specific workflow access, notification isolation, labor/privacy review, discovery policy, and the treatment of legitimate safety warnings remain unresolved.","source_ids":["S1","S2","S3","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Blind testing can exercise an entire quality system and has already identified preventable process defects; preventing software cues and invalid state transitions could protect the validity of that assessment. Realized impact of the capsule itself is unmeasured.","source_ids":["S1","S2","S3","S5"]},"stakeholder_pull":{"score":4,"rationale":"Two laboratories adopted blind programs, laboratory leaders reported broader interest and a specific hidden-LIMS-feature need, and examiners appear broadly receptive. No adopter has requested this exact contract architecture.","source_ids":["S1","S2","S3","S4"]},"incremental_advantage":{"score":3,"rationale":"A shared state and leakage oracle could improve replacement assurance beyond hidden flags and walkthroughs, but no comparative evidence yet shows that it catches materially more defects than existing QA validation and workflow testing.","source_ids":["S1","S2","S5","S6"]},"distinctiveness_plausibility":{"score":2,"rationale":"HFSC already implements the central hidden-classification LIMS behavior. The remaining distinction is the formal, cross-implementation lifecycle and leakage contract, which was not found in the reviewed sources but is narrower than the original proposal.","source_ids":["S1","S2","S5","S6","S8"]},"technical_implementability":{"score":4,"rationale":"Existing masking, request-portal, role separation, predefined evaluation, and software-validation practices make a sandbox prototype credible. Exhaustive observable-channel coverage and concurrency semantics are technically difficult.","source_ids":["S1","S2","S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"Laboratory executives, quality management, and designated technical leaders are identifiable authorizers, and accreditation systems provide governance. Local labor, privacy, discovery, prosecutor, and submitting-agency approvals are not established.","source_ids":["S1","S2","S6","S7"]},"evidence_readiness":{"score":3,"rationale":"A synthetic sandbox comparison is well bounded and does not require live cases, but a meaningful independent adapter and realistic participant-visible transcripts require proprietary LIMS and workflow access.","source_ids":["S1","S2","S5","S6"]},"safety_net_benefit":{"score":4,"rationale":"An isolated conformance gate could prevent accidental release, database upload, unauthorized disclosure, or mutable submissions before production use. It cannot address physical, scheduling, social, or administrator cues outside the interface.","source_ids":["S1","S2","S5","S6"]},"scalability":{"score":3,"rationale":"One parameterized oracle could be reused, but each laboratory has different LIMS, clients, submission partners, case conventions, policies, and observable channels. Existing literature reports substantial personnel and logistical requirements, especially for smaller laboratories.","source_ids":["S1","S2","S3"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Eight-to-twelve-week synthetic sandbox: formal lifecycle specification, transparent reference model, one thin adapter, role clients, seeded mutants, transcript comparator, threat review, and predeclared analysis plan.","confidence":"MODERATE","assumptions":["No live cases, personnel records, national databases, or production notifications are used.","Existing staff can provide bounded quality-policy and workflow workshops.","Estimate includes resource-equivalent engineering, QA, forensic-domain, security, and project-management time rather than only cash purchases.","The adapter targets one limited discipline and a sandbox API."],"source_ids":["S1","S2","S5","S6"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Production-grade contract and adapter design for one laboratory, observable-channel inventory, access controls, immutable audit history, notification isolation, security assessment, documentation, and institutional policy review.","confidence":"LOW","assumptions":["A configurable or accessible LIMS integration point exists.","One laboratory and one or two disciplines are in scope.","No wholesale LIMS replacement is included.","Local privacy, labor, discovery, and records review can be handled within the institution."],"source_ids":["S1","S2","S6","S7"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Validated integration, regression and reliability testing, staff training, independent acceptance review, staged rollout, rollback tooling, and initial operation for one laboratory.","confidence":"LOW","assumptions":["The synthetic-only gate succeeds before any production integration.","At least one QA lead, technical lead, software maintainer, security reviewer, and submitting-agency liaison participate.","The band excludes purchasing or replacing an enterprise LIMS and excludes substantive proficiency-material production at scale."],"source_ids":["S1","S2","S3","S6"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Two dedicated or resource-equivalent QA/operations staff, adapter maintenance, contract-version review, recurring conformance and leakage runs, audit support, security maintenance, and modest synthetic materials.","confidence":"MODERATE","assumptions":["HFSC's historical estimate of at least two dedicated full-time staff for a blind program is used as the principal labor anchor.","Historical supply costs are not treated as complete costs and are escalated only qualitatively to a broad 2026 band.","One laboratory operates the capability across a limited number of disciplines.","Major LIMS replacement, enterprise licensing, and live-case examiner labor are excluded."],"source_ids":["S1","S2"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Multiple primary implementations document that superficial workflow and system details reveal blind assignments and that blind programs can uncover quality-system defects.","source_ids":["S1","S2","S3","S5"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"HFSC and Harris County are demonstrated adopters of blind programs; laboratory quality leadership, executive management, and designated technical leaders are credible authorizers for a synthetic evaluation.","source_ids":["S1","S2","S3","S6","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim is narrower than blind testing itself: whether a reusable state-and-leakage oracle rejects seeded or independently implemented defects missed by existing hidden-flag procedures. It has explicit behavioral comparators and falsifiers.","source_ids":["S1","S2","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A synthetic, one-laboratory, one-discipline sandbox can compare an incumbent-style adapter, transparent reference model, independent adapter, and seeded mutants without touching personnel or live casework.","source_ids":["S1","S2","S5","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The proposed next step can remain synthetic and disconnected; it must halt on any live-service contact. Production use still requires institution-specific labor, privacy, discovery, records, and quality approvals, but those do not block the isolated evidence step.","source_ids":["S2","S6","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The ranges explicitly separate prototype, startup, launch, and recurring scope. HFSC provides actual supply-cost observations and a two-FTE recurring labor anchor, although software-specific startup estimates remain low-confidence.","source_ids":["S1","S2"]}},"next_evidence_step":"With one partner laboratory, preregister a 10-week disconnected sandbox trial for one discipline. Encode the prepared-to-closed lifecycle in a transparent reference model; build an incumbent-style hidden-flag adapter and one independently implemented capsule adapter; seed at least 20 mutants covering unauthorized disclosure, role escalation, adjudication-before-lock, overwrite-after-lock, non-append corrections, repeated disclosure, error-state mutation, notification leakage, identifier/route leakage, and submission/disclosure races. Execute at least 1,000 generated role-operation sequences plus 100 matched ordinary/exercise transcript pairs. Compare abstract post-states, allowed errors, role-visible fields, URLs, notifications, ordering, and coarse timing using tolerances fixed before execution. Required comparators are the reference model, the incumbent-style adapter, the independent adapter, and seeded mutants. Provisional success requires zero mandatory-state or authorization divergence between unmutated implementations, zero live-service contacts, zero forbidden disclosures or overwrites, detection of at least 18 of 20 seeded mutants, and blinded participant-role classification of matched transcripts no better than 55% accuracy with a confidence interval including chance. Falsify the incremental claim if a seeded material defect passes, two conforming implementations reach different mandatory states, participant reviewers identify exercises above the threshold from declared software channels, or required authorized workflows cannot be represented without exposing the hidden classification.","blocking_evidence":["No field evidence shows how often nominally compatible forensic workflow implementations diverge on submission finality, correction history, disclosure, or other proposed lifecycle laws.","No laboratory or LIMS vendor has committed proprietary sandbox access, an independent adapter, or staff time for the comparative test.","No empirical evidence establishes that the proposed transcript normalization and 55% distinguishability threshold are valid for actual participant clients.","Institution-specific labor, employee-notice, privacy, public-records, discovery, prosecutor, and submitting-agency requirements remain unreviewed.","No measured software-engineering cost exists for the capsule, adapter, conformance generator, or leakage comparator.","No reviewed source demonstrates that passing this contract preserves blindness against physical materials, scheduling, case facts, or human communication outside the software boundary."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The eight-source search establishes substantial collision with operational blind-proficiency programs, especially HFSC's QA-only LIMS masking and request portal. It did not find the complete role-sensitive abstract-state-machine and cross-implementation leakage-oracle combination. This is only a bounded prior-art contrast: world novelty, patentability, freedom to operate, market size, and realized impact were not measured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure a forensic laboratory partner and disconnected LIMS sandbox for one discipline.","Obtain institution-specific quality, privacy, labor, records, discovery, security, and submitting-agency review for the sandbox protocol.","Implement the transparent reference model, incumbent-style comparator, and independently authored adapter.","Preregister observable channels, tolerated differences, acceptance thresholds, concurrency cases, and at least 20 seeded defects.","Run generated lifecycle sequences and blinded transcript-distinguishability testing, reporting every divergence without post hoc threshold changes.","Produce measured engineering and operating effort to replace the current low-confidence software cost estimates."],"reason":"Web evidence verifies the leakage problem, credible authorizers, stakeholder interest, feasibility of hidden LIMS masking, substantial prior-art collision, and a narrow falsifiable remainder. The remaining question is comparative performance of the formal capsule and oracle on proprietary workflows and implementations; it cannot be resolved by further bounded web search and requires partnered sandbox construction and testing. Under the stopping rule, this empirical-research stop is non-repairable."},"proposal_index":3}