{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"predictive_residual_processing__criminology_forensic:P4:v0","cell_id":"predictive_residual_processing__criminology_forensic","search_queries":["site:nationalacademies.org proactive policing effects crime displacement legitimacy report 2018","site:ojp.gov place based crime prevention displacement measurement evaluation guide","hot spots policing systematic review displacement diffusion Campbell 2019","police initiated activity measurement crime trends evaluation official","site:bjs.ojp.gov nation's two crime measures police reported victimization survey official","site:cops.usdoj.gov assessing responses to problems process impact evaluation displacement PDF","site:ojp.gov Smart Policing Initiative evaluation partnership expressed need police agencies analysts","police activity changes recorded crime detection measurement bias research police generated data","site:justice.gov law enforcement data collection racial disparities community oversight official guidance","site:justice.gov title VI law enforcement data algorithm policing civil rights guidance","site:nist.gov privacy framework aggregate data de-identification small cells official","site:ojp.gov law enforcement program evaluation privacy aggregate data IRB"],"sources":[{"source_id":"S1","title":"The Nation’s Two Crime Measures, 2015–2024","publisher":"Bureau of Justice Statistics","url":"https://bjs.ojp.gov/library/publications/nations-two-crime-measures-2015-2024","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-03","accessed_at":"2026-08-03","claims_supported":["Police-reported incident data and victimization-survey data have different purposes, methods, and offense coverage.","The NCVS includes victimizations not reported to police, while NIBRS represents incidents reported by law-enforcement agencies.","Multi-source interpretation is necessary for a more comprehensive account of crime than either stream alone."]},{"source_id":"S2","title":"Measurement Error in Calls-For-Service as an Indicator of Crime","publisher":"Criminology; cataloged by the National Institute of Justice","url":"https://nij.ojp.gov/library/publications/measurement-error-calls-service-indicator-crime","source_class":"PRIMARY_RESEARCH","publication_date":"1997-01-01","accessed_at":"2026-08-03","claims_supported":["An observational study covering 60 neighborhoods found that calls-for-service records substantially undercounted crime encountered by officers.","Discrepancies varied by crime type and systematically across space.","Calls for service require caution when used as a crime indicator."]},{"source_id":"S3","title":"Spatial Displacement and Diffusion of Benefits Among Geographically-Focused Policing Initiatives","publisher":"Campbell Collaboration","url":"https://www.campbellcollaboration.org/review/geographically-focused-policing/","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2011-06-15","accessed_at":"2026-08-03","claims_supported":["Spatial, temporal, target, method, offense-type, and offender displacement are recognized possible consequences of focused interventions.","The systematic review included 44 geographically focused policing studies.","Monitoring adjacent places is relevant even though the review's aggregate evidence favored diffusion of benefits over displacement."]},{"source_id":"S4","title":"Assessing Responses to Problems: An Introductory Guide for Police Problem-Solvers","publisher":"Office of Community Oriented Policing Services, U.S. Department of Justice","url":"https://portal.cops.usdoj.gov/resourcecenter/RIC/Publications/cops-w0012-pub.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2002","accessed_at":"2026-08-03","claims_supported":["Process evaluation compares the planned response with what actually occurred and is needed to interpret impact results.","Impact evaluation requires before-and-after measurement plus a design that addresses alternative explanations.","Simple pre/post comparisons are weak for causal attribution; interrupted time series and control comparisons are established stronger practices.","High-consequence evaluations should involve professional or outside evaluators."]},{"source_id":"S5","title":"An Ex Post Facto Evaluation Framework for Place-Based Police Interventions","publisher":"Evaluation Review; cataloged by the Office of Justice Programs","url":"https://www.ojp.gov/library/publications/ex-post-facto-evaluation-framework-place-based-police-interventions","source_class":"PRIMARY_RESEARCH","publication_date":"2012-01-01","accessed_at":"2026-08-03","claims_supported":["A place-based police program was retrospectively evaluated with treated and comparison street segments, propensity-score matching, and growth-curve models.","Few agencies were reported to make advance commitments to rigorous evaluation despite broad adoption of place-based strategies.","Real-world place-based evaluation frameworks and displacement assessment are established prior art."]},{"source_id":"S6","title":"Smart Policing Initiative: Overview","publisher":"Bureau of Justice Assistance","url":"https://www.bja.ojp.gov/program/smart-policing-initiative-spi/overview","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2012-02-27","accessed_at":"2026-08-03","claims_supported":["BJA sponsors a program in which law-enforcement agencies and research partners collect and analyze data to determine what works.","The program expressly supports effective, efficient, economical, evidence-based policing and researcher-agency partnerships.","Its guidance encourages data beyond calls, offenses, arrests, and complaints, including external health, school, service, and justice sources.","BJA, participating agencies, research partners, and community stakeholders form an identifiable funder-adopter-authorizer pathway."]},{"source_id":"S7","title":"NIST SP 800-188: De-Identifying Government Datasets—Techniques and Governance","publisher":"National Institute of Standards and Technology","url":"https://csrc.nist.gov/pubs/sp/800/188/final","source_class":"STANDARD","publication_date":"2023-09-14","accessed_at":"2026-08-03","claims_supported":["Government agencies can use de-identification to reduce privacy risk while retaining useful statistical analysis.","Agencies should assess release risk and select an explicit data-sharing model, including protected enclaves or synthetic data.","Disclosure review, measurable de-identification standards, and re-identification testing are available governance controls.","Simple masking is not necessarily adequate de-identification."]},{"source_id":"S8","title":"Conduct of Law Enforcement Agencies","publisher":"Civil Rights Division, U.S. Department of Justice","url":"https://www.justice.gov/crt/conduct-law-enforcement-agencies","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Federal law permits review of state or local law-enforcement practices that systematically violate rights.","Federal-funding statutes prohibit specified forms of discrimination by recipient agencies.","DOJ remedies frequently require transparency, data collection, independent oversight, community partnerships, and improved force review.","Rights, discriminatory-policing, force, and oversight signals cannot safely be treated as ordinary suppressible monitoring noise."]}],"problem_evidence":{"support":"STRONG","rationale":"The problem is visible in authoritative data architecture and empirical measurement research: police-recorded incidents omit unreported victimizations, calls-for-service error varies by offense and place, focused interventions can have displacement or diffusion effects, and implementation fidelity must be separated from impact. The exact prevalence of evaluators being overwhelmed by repetitive dashboards was not measured.","source_ids":["S1","S2","S3","S4","S5"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"BJA is an identifiable funder and convenor, while participating law-enforcement agencies, research partners, and oversight bodies are plausible adopters or authorizers. BJA expressly seeks efficient data use, rigorous performance measurement, external data, and research partnerships. No source expresses demand for this exact residual-first interface, and no named agency has committed data or authority to a trial.","source_ids":["S6","S8"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Process-plus-impact evaluation of police problem responses","similarity":"Already freezes or documents the planned response, compares it with implementation, measures outcomes, examines alternatives, and separates implementation failure from impact.","remaining_difference":"It ordinarily reports complete evaluation measures and does not use a synchronized action-conditioned predictor to suppress expected content, reconstruct windows, audit suppressed data, or trigger decompression.","source_ids":["S4"]},{"name":"Ex post facto matched evaluation of a place-based police intervention","similarity":"Uses place-time treatment definitions, comparison places, longitudinal crime outcomes, and displacement analysis for retrospective evaluation.","remaining_difference":"It estimates program-associated effects rather than operating a continuing residual communication and attention-allocation layer with protected bypasses and raw-window audits.","source_ids":["S5"]},{"name":"Smart Policing Initiative performance measurement and research partnership","similarity":"Combines agency data, external sources, research partners, evidence-based decision support, and organizational/community collaboration.","remaining_difference":"It is an institutional program and evaluation practice, not a specified action-copy residual architecture or residual-first interface.","source_ids":["S6"]},{"name":"Geographically focused policing displacement analysis","similarity":"Explicitly examines changes outside targeted places and recognizes multiple displacement dimensions.","remaining_difference":"It tests intervention outcomes but does not predict the intervention's direct measurement footprint or govern selective residual presentation.","source_ids":["S3"]}],"distinctive_claim_remaining":"Against a full multi-indicator dashboard and a conventional process-plus-impact evaluation, a frozen model conditioned only on the authorized aggregate program schedule can present precision- and consequence-weighted residuals while preserving complete reconstructibility, independent raw-window audits, rights-critical bypasses, and automatic fallback, thereby reducing evaluator time without materially lowering detection of displacement, reporting divergence, service withdrawal, or harm and without increasing unsupported causal interpretations.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"The component methods—process/impact evaluation, matched retrospective comparison, external-source use, aggregate de-identification governance, civil-rights review, and independent oversight—are established. The integrated residual interface, footprint model calibration, suppression thresholds, protected-signal bypass performance, reviewer behavior, and maintenance burden have not been demonstrated. Access to intervention schedules and multi-source historical data will be agency-specific.","source_ids":["S4","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Measurement error and intervention spillovers can materially distort program interpretation, while rights and force monitoring carry high consequences. Realized impact is unmeasured.","source_ids":["S1","S2","S3","S8"]},"stakeholder_pull":{"score":3,"rationale":"BJA and its agency-research partnerships express demand for efficient, rigorous, multi-source performance analysis, but not for the proposed residual implementation specifically.","source_ids":["S6"]},"incremental_advantage":{"score":2,"rationale":"Existing process, impact, matched-comparison, displacement, and multi-source evaluation practices cover much of the substantive task. The residual interface may save attention, but no comparative evidence establishes that advantage.","source_ids":["S3","S4","S5","S6"]},"distinctiveness_plausibility":{"score":3,"rationale":"The integrated action-conditioned residual, reconstructibility, raw-audit, rights-bypass, and fallback package was not found in the eight sources, although its components are adjacent to established evaluation practice. This is not a world-novelty finding.","source_ids":["S4","S5","S6"]},"technical_implementability":{"score":3,"rationale":"A retrospective aggregate prototype is technically plausible with established evaluation and de-identification methods, but model calibration, source synchronization, rare-event coverage, and reliable fallback require testing.","source_ids":["S4","S5","S7"]},"adoption_authority_feasibility":{"score":3,"rationale":"BJA, participating agencies, independent evaluators, and oversight bodies provide a credible pathway. A specific agency's records authority, public-record obligations, labor rules, community governance, and data-sharing approvals remain unresolved.","source_ids":["S6","S8"]},"evidence_readiness":{"score":2,"rationale":"The next test can use historical or synthetic data, but decisive evidence requires nonpublic schedules, linked administrative sources, audit labels, and blinded reviewer participation.","source_ids":["S4","S5","S6"]},"safety_net_benefit":{"score":4,"rationale":"If correctly implemented, full rights-signal bypasses, independent oversight, de-identification review, raw audits, and fallback could expose omissions that ordinary dashboards miss. Their reliability is not yet validated.","source_ids":["S7","S8"]},"scalability":{"score":3,"rationale":"The architecture could transfer across place-based programs, but every jurisdiction would need local source mapping, calibration, privacy review, legal authority, and community oversight, limiting turnkey scaling.","source_ids":["S6","S7","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"A 12- to 16-week preregistered retrospective or synthetic shadow study for one completed program, including data preparation, one frozen predictor and rival, a residual-first prototype, blinded reviewer sessions, and independent full-window auditing.","confidence":"LOW","assumptions":["Roughly 6–10 person-months across an evaluator, crime analyst, data engineer, statistician, privacy reviewer, and community/rights reviewer.","One agency or synthetic replay; no production integration or live operational use.","Existing aggregate dashboard and historical records are available without substantial acquisition fees.","The band is a resource-equivalent estimate, not a vendor quotation or market-price finding."],"source_ids":["S4","S6","S7"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Read-only deployment preparation for one agency and one program family: governed data enclave, source contracts, model and dashboard engineering, de-identification review, audit sampling, access controls, fallback tests, training, and independent evaluation setup.","confidence":"LOW","assumptions":["Approximately 2–5 full-time-equivalent staff-years plus legal, civil-rights, privacy, security, and community-governance support.","Existing agency data platforms can export versioned aggregate records.","No individual scoring, patrol allocation, or live decision automation is included.","Costs rise materially if source systems require replacement or new statutory authority."],"source_ids":["S6","S7","S8"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"A multi-program or multi-district operational launch with production-quality data pipelines, redundancy, model registry, monitoring, audit staff, incident/fallback procedures, independent oversight, and formal validation.","confidence":"LOW","assumptions":["Approximately 6–15 full-time-equivalent staff-years across engineering, analysis, evaluation, governance, security, training, and oversight.","Several administrative and external data sources require recurring integration and quality assurance.","The launch remains advisory and aggregate; no enforcement automation is authorized.","No authoritative 2026 procurement benchmark for this exact system was located."],"source_ids":["S6","S7","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Annual operation for one agency after launch: data quality and synchronization, model revalidation, privacy review, random and risk-stratified audits, evaluator staffing, community oversight, training, incident response, and periodic full-report fallback exercises.","confidence":"LOW","assumptions":["Approximately 2–6 recurring full-time-equivalent roles plus infrastructure and contracted independent review.","Program definitions and sources change slowly enough to avoid rebuilding the platform annually.","Audit intensity is sufficient for common failures but cannot guarantee discovery of extremely rare harms.","Resource equivalence excludes downstream costs of separate formal causal evaluations."],"source_ids":["S6","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official and empirical sources directly establish divergent crime measures, spatially varying calls-for-service error, displacement/diffusion concerns, and the need to distinguish planned implementation from impact.","source_ids":["S1","S2","S3","S4"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"BJA's Smart Policing Initiative identifies a credible funder and an established pathway involving law-enforcement agencies, research partners, external stakeholders, and evidence-based performance measurement; independent oversight is also an established DOJ remedy.","source_ids":["S6","S8"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal can be compared directly with full dashboards and conventional process-plus-impact evaluation on evaluator time, material-change recall, protected-signal presentation, reconstruction, false escalation, and causal overinterpretation.","source_ids":["S4","S5"]},"bounded_next_evidence_step":{"status":"YES","reason":"A preregistered, retrospective, read-only, one-program comparison can be bounded by fixed windows, reviewers, seeded/documented events, audit procedures, comparators, and explicit stopping rules.","source_ids":["S4","S5","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"No fundamental stop applies to a properly authorized retrospective aggregate study that is isolated from live enforcement. De-identification review, independent oversight, rights-signal completeness, and local data authority must be satisfied as launch prerequisites.","source_ids":["S7","S8"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The four ranges have explicit staffing and scope assumptions, but no direct wage, procurement, integration, or agency-specific data-remediation benchmark was among the eight sources.","source_ids":["S6","S7"]}},"next_evidence_step":"Run a preregistered 12- to 16-week offline study on 150–250 historical or synthetic place-time monitoring windows from one completed program. Freeze the schedule, geographic units, transformations, minimum cell sizes, model, rival model, thresholds, bypass list, audit sample, and holdout before scoring. Randomize at least eight blinded evaluators to (A) the existing full dashboard, (B) a conventional process-plus-impact dashboard, or (C) the residual-first interface, with crossover and balanced window order. Include documented or seeded reporting outages, intensity changes, adjacent and temporal displacement, officer-initiated activity changes, complaint/injury shifts, service withdrawal, benign external shocks, source disagreement, and version mismatch. Independently audit every protected-signal window plus a random sample of all other windows. Primary tests are noninferiority of material-change recall within five percentage points and at least a 20% reduction in median review time versus the better comparator. Secondary tests are protected-signal presentation, full-window reconstruction, false escalation, unsupported causal interpretations, subgroup/place miss patterns, fallback behavior, and maintenance time. Falsify the incremental claim upon any protected-signal omission or privacy breach, recall worse than the noninferiority margin, systematic independent-audit misses, reconstruction failure above 1% of windows, increased unsupported causal attribution, failed heartbeat/version fallback, or failure to reduce total reviewer-plus-maintenance effort. Results may authorize only a larger retrospective study.","blocking_evidence":["No comparative evidence shows that residual-first presentation saves total evaluator effort after model maintenance and audits.","No empirical evidence establishes noninferior detection of displacement, service withdrawal, reporting divergence, or rights-related harm.","Protected-signal bypass, heartbeat, synchronization, reconstruction, and fallback reliability have not been tested.","No named agency has committed historical schedules, multi-source data, blinded reviewers, legal approval, or independent audit access.","Local privacy, public-records, retention, collective-bargaining, data-sharing, and oversight requirements remain jurisdiction-specific.","The action-conditioned model's ability to separate direct measurement footprint from external change and omitted context is unknown.","Reviewer susceptibility to treating residuals as causal proof is unmeasured.","Agency-specific 2026 implementation and recurring costs lack procurement or labor benchmarks."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search establishes only that conventional process/impact evaluation, matched place-based evaluation, displacement analysis, multi-source performance measurement, de-identification governance, and rights oversight are existing practices. It did not measure world novelty, patentability, freedom to operate, market size, or realized impact, and absence of the exact integrated architecture from eight sources is not evidence of worldwide novelty.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure a named agency-evaluator-oversight partnership with lawful access to one completed program's frozen schedule and multi-source aggregate records.","Preregister comparator interfaces, sample size, noninferiority margin, time-saving threshold, protected bypasses, privacy thresholds, audits, and falsifiers.","Demonstrate zero protected-signal omissions and zero privacy-threshold breaches in the bounded study.","Show material-change recall no more than five percentage points below the best comparator while reducing total reviewer-plus-maintenance effort by at least 20%.","Demonstrate reliable reconstruction, heartbeat/version-mismatch detection, and full-report fallback, with reconstruction failure below 1% of windows.","Show no increase in unsupported causal interpretations and no systematic miss pattern by geography, exposure, source availability, or protected-group disparity check.","Replace resource-equivalent cost assumptions with agency-specific labor, integration, governance, audit, and recurring operating estimates."],"reason":"Web evidence verifies the measurement problem, credible institutional pathway, adjacent prior art, and feasibility of a bounded retrospective test, but the proposal's incremental benefit and safety performance depend on proprietary historical data, blinded human review, independent audits, and interface testing. Those questions cannot be resolved by further bounded web search."},"proposal_index":4}