{"schema_version":1,"research_id":"eoa_inverse_innovation_exp04_external_evaluation_20260802","source_assessment_id":"negative_space_design__behavioral_economics:SENTINEL:v0","cell_id":"negative_space_design__behavioral_economics","search_queries":["site:cdc.gov NHSN PSC Annual Survey Data Quality Dashboard missing survey PDF","site:cdc.gov/nssp data quality tools facility stopped sending data","site:hl7.org/fhir R4 Observation dataAbsentReason","site:who.int data quality missing values true zero facility reports","site:grafana.com/docs alerting missing data No Data MissingSeries state reason","site:osr.statisticsauthority.gov.uk regulatory guidance dashboards data quality limitations user needs","missingness visualization controlled study decision making Song Fu Saket Stasko","site:bls.gov OEWS software developers hourly mean wage 2025","site:who.int \"Missing data should be clearly differentiated from true zero values\""],"sources":[{"source_id":"S1","title":"Data Quality Assurance, Module 2: Desk Review of Data Quality","publisher":"World Health Organization","url":"https://cdn.who.int/media/docs/default-source/data-quality-pages/2021_-dqa_module-2_desk-review-of-data-quality.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2020-12","accessed_at":"2026-08-02","claims_supported":["WHO requires missing data to be differentiated from true zero values in district and facility reports.","WHO explains that assigning zero to missing entries makes no-event periods indistinguishable from reporting failures.","WHO recommends routine completeness and timeliness monitoring and managerial interpretation of apparent zeros."]},{"source_id":"S2","title":"FHIR R4 Observation: Detailed Descriptions","publisher":"Health Level Seven International","url":"https://hl7.org/fhir/R4/observation-definitions.html","source_class":"STANDARD","publication_date":"2019","accessed_at":"2026-08-02","claims_supported":["FHIR defines dataAbsentReason as the reason an expected observation value is missing.","FHIR requires dataAbsentReason to appear only when the observation value is absent.","FHIR demonstrates that cause-coded absence can be represented in a standardized data model, while noting that use-case agreements remain necessary for interpretation."]},{"source_id":"S3","title":"Quick Reference Guide to the NHSN PSC Annual Survey Data Quality Dashboard","publisher":"U.S. Centers for Disease Control and Prevention","url":"https://www.cdc.gov/nhsn/pdfs/pscmanual/PSC-Survey-Data-Quality-Dashboard-QRG.pdf","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2025","accessed_at":"2026-08-02","claims_supported":["A deployed surveillance dashboard distinguishes incomplete or missing surveys, facilities not operational, required surveys, insufficient comparison data, confirmed issues, and resolved issues.","The dashboard preserves prior-year context and provides review, edit, submit, confirm, and resolution workflows.","CDC cautions that a flagged discrepancy is not necessarily incorrect, supporting human review rather than automatic enforcement."]},{"source_id":"S4","title":"Data Quality Tools","publisher":"U.S. Centers for Disease Control and Prevention, National Syndromic Surveillance Program","url":"https://www.cdc.gov/nssp/php/onboarding-toolkits/data-quality.html","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-01-13","accessed_at":"2026-08-02","claims_supported":["NSSP operates tools for monitoring data flow, timeliness, completeness, validity, processing exceptions, backlogs, and facility activity.","Site administrators can define anomaly rules, alert users, inspect incomplete messages, and request assistance from site inspectors.","CDC states that these reports flag potential problems so corrective action can be taken, identifying operational users and an expressed need."]},{"source_id":"S5","title":"Handle Missing Data in Grafana Alerting","publisher":"Grafana Labs","url":"https://grafana.com/docs/grafana/latest/alerting/guides/missing-data/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"undated; current documentation","accessed_at":"2026-08-02","claims_supported":["Grafana distinguishes query failure, No Data, and MissingSeries states.","Grafana warns that missing observations can create silent monitoring failures when alerts do not fire.","The product supports state-specific behavior, state-reason annotations, delayed-data handling, alert routing, and preservation of prior state."]},{"source_id":"S6","title":"Regulatory Guidance: Dashboards","publisher":"Office for Statistics Regulation, UK Statistics Authority","url":"https://osr.statisticsauthority.gov.uk/guidance/regulatory-guidance-dashboards/pages/6/","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2025-08-21; updated 2026-03-04","accessed_at":"2026-08-02","claims_supported":["Dashboard producers should communicate uncertainty, data limitations, quality information, and material risks of misinterpretation.","Producers should identify user needs, manage errors, protect confidential data, ensure accessibility, and provide sufficient maintenance resources.","The guidance places implementation responsibility with dashboard-producing organizations and their governance processes."]},{"source_id":"S7","title":"Understanding the Effects of Visualizing Missing Values on Visual Data Exploration","publisher":"Hayeong Song, Yu Fu, Bahador Saket, and John Stasko; arXiv","url":"https://arxiv.org/abs/2109.08723","source_class":"PRIMARY_RESEARCH","publication_date":"2021-09-17","accessed_at":"2026-08-02","claims_supported":["A controlled study compared omission of records containing missing values with explicit estimated-value and error-bar representations.","Missing-value presentation changed participants' decision workflow in a hypothetical investment task.","The study does not test regulatory closure decisions or cause-matched actions."]},{"source_id":"S8","title":"Toward Systematic Considerations of Missingness in Visual Analytics","publisher":"Maoyuan Sun et al.; arXiv","url":"https://arxiv.org/abs/2108.04931","source_class":"PRIMARY_RESEARCH","publication_date":"2022-07-27","accessed_at":"2026-08-02","claims_supported":["The research identifies system failure, network interruption, intentional hiding, and bias as distinct causes of missingness that can affect decisions.","It distinguishes observed, inferred, and ignored missingness and argues that visualization should move missingness into users' awareness.","It is conceptual prior art and does not validate the proposed regulator-facing comparator."]}],"problem_evidence":{"support":"STRONG","rationale":"The underlying problem visibly exists: WHO documents the specific missing-versus-zero ambiguity; Grafana documents silent failures when sources stop reporting; CDC operates monitoring workflows for missing reports, feed problems, completeness, and backlogs; and two research sources establish that missingness and its presentation can affect decisions. Evidence that ambiguous absence actually causes false regulatory case closure, or systematically advantages under-reporting entities, was not found, so the exact consequence remains an important field-evidence gap.","source_ids":["S1","S3","S4","S5","S7","S8"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Identifiable adopters and authorizers include CDC NHSN dashboard owners, NSSP site administrators and inspectors, and analogous official-statistics dashboard producers. Their published materials express needs for accurate data, rapid identification of processing and reporting problems, corrective action, user-centered design, uncertainty communication, accessibility, confidentiality, and maintenance. No named regulator committed to adopting or funding this exact cause-matched interface or its evaluation.","source_ids":["S3","S4","S6"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"CDC NHSN PSC Annual Survey Data Quality Dashboard","similarity":"A deployed surveillance dashboard distinguishes several reporting and comparison states, retains prior-period values, directs review or submission, and records confirmation and resolution.","remaining_difference":"The source reports neither regulator-like false no-incident closure outcomes nor a controlled comparison of one cause-matched action against an equally salient generic missing-data warning.","source_ids":["S3"]},{"name":"Grafana No Data and MissingSeries handling","similarity":"It operationally separates absence and failure states, records state reasons, preserves previous state, and permits state-specific alerts and routing.","remaining_difference":"It addresses infrastructure monitoring rather than regulated-entity oversight and does not test closure, data-request, or investigation choices.","source_ids":["S5"]},{"name":"CDC NSSP data-quality monitoring","similarity":"It monitors feed flow, timeliness, completeness, validity, exceptions, backlogs, and facility anomalies and supports alerts and investigation.","remaining_difference":"Public documentation does not establish the proposed empty-state panel or its behavioral advantage over a generic warning.","source_ids":["S4"]},{"name":"WHO missing-versus-zero rule and FHIR DataAbsentReason","similarity":"Together they establish the core representational practice of separating missing values from true zeros and coding a reason for absence.","remaining_difference":"They do not validate a single cause-matched oversight action or the proposed behavioral endpoints.","source_ids":["S1","S2"]},{"name":"Missingness-aware visualization research","similarity":"The research catalogs missingness causes and provides experimental evidence that presentation can change decision workflow.","remaining_difference":"It does not test regulator-like entity-period decisions, under-reporting incentives, false closure, or cause-matched actions.","source_ids":["S7","S8"]}],"distinctive_claim_remaining":"In regulator-like entity-period decisions, adding a diagnosed absence cause and exactly one matched oversight action to an equally salient generic missing-data panel—with audit context, available underlying actions, and non-focal spacing held constant—materially reduces false no-incident closure and increases predefined appropriate follow-up without materially increasing inappropriate escalation or delay.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Technical feasibility is supported by standardized absent-reason fields and deployed products that classify missing, incomplete, error, delayed, and missing-series conditions, maintain context, and trigger workflows. A synthetic prototype and randomized simulation require no live enforcement authority and can preserve permissions and audit history. The main unresolved technical and safety dependency is reliable cause classification: labels derived from weak metadata could create automation bias, excessive escalation, confidentiality leakage, or discrimination against low-capacity reporters. Live use additionally requires the adopting body's legal, privacy, accessibility, records, labor-workflow, and enforcement-authority review; no cross-jurisdiction legal clearance is established.","source_ids":["S2","S3","S4","S5","S6"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Avoiding false reassurance in oversight could materially improve attention allocation and data integrity, but the prevalence and realized consequence of false closure are unmeasured.","source_ids":["S1","S4","S5"]},"stakeholder_pull":{"score":3,"rationale":"Operational dashboard owners visibly need data-quality diagnosis and corrective workflows, but no organization has requested or funded this exact intervention.","source_ids":["S3","S4","S6"]},"incremental_advantage":{"score":3,"rationale":"The matched-action comparison is plausibly useful beyond a generic warning, yet deployed systems already provide closely related state-specific workflows and no direct superiority evidence exists.","source_ids":["S3","S5"]},"distinctiveness_plausibility":{"score":2,"rationale":"Most components are established practice; only the tightly controlled behavioral comparison and endpoints remain distinct in the bounded search.","source_ids":["S1","S2","S3","S4","S5","S7","S8"]},"technical_implementability":{"score":4,"rationale":"Cause fields, state machines, context retention, routing, and resolution histories are standard implementable patterns; accurate diagnosis and integration remain nontrivial.","source_ids":["S2","S3","S4","S5"]},"adoption_authority_feasibility":{"score":3,"rationale":"Dashboard product and data-governance teams can authorize a synthetic test, while live oversight actions remain with legally authorized officials and require local review.","source_ids":["S3","S4","S6"]},"evidence_readiness":{"score":4,"rationale":"The residual claim supports a bounded randomized simulation with explicit comparators, outcomes, and falsifiers; representative participants and validated case ground truth must still be obtained.","source_ids":["S3","S7"]},"safety_net_benefit":{"score":3,"rationale":"The intervention may reduce advantages from nonreporting and preserve auditability, but could disproportionately escalate entities with weak reporting infrastructure unless monitored by capacity and cause.","source_ids":["S1","S4","S6","S8"]},"scalability":{"score":4,"rationale":"A reusable cause taxonomy and panel component could scale across entity-period dashboards, subject to domain-specific mappings, metadata quality, governance, accessibility, and maintenance.","source_ids":["S2","S4","S5","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"One preregistered randomized synthetic-case study: prototype and comparator, independent cause/action adjudication, approximately 100–250 regulator-like participants, accessibility review, analysis, and reporting.","confidence":"MODERATE","assumptions":["Uses synthetic records and an existing research platform.","Requires roughly 3–6 person-months across UX, behavioral research, data quality, accessibility, and analysis.","Participant recruitment may require specialist panels or partner staff time.","No live-system integration or enforcement action is included."],"source_ids":["S6","S7"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Integrate the state taxonomy and audited panel into one existing dashboard, including metadata mapping, permissions, provenance, accessibility, security/privacy review, logging, tests, and governance approval.","confidence":"LOW","assumptions":["One dashboard and one jurisdiction are in scope.","Existing identity, audit, and data-pipeline infrastructure can be reused.","No major source-system replacement is required.","The range is resource-equivalent, not a vendor quote or procurement estimate."],"source_ids":["S2","S3","S4","S6"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Time-bounded live pilot preparation and launch across a limited unit: training, support, parallel-run validation, monitoring, incident response, evaluation, rollback capability, and governance oversight.","confidence":"LOW","assumptions":["Launch follows favorable simulation evidence and explicit legal/operational authorization.","The pilot runs in parallel with existing review rules and does not automate sanctions.","Approximately 4–10 staff equivalents contribute part-time over 6–12 months.","Costs vary sharply with procurement, security accreditation, and legacy integration."],"source_ids":["S3","S4","S6"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Annual maintenance for one deployed dashboard: state-mapping review, quality monitoring, accessibility regression testing, audit sampling, user support, incident response, and outcome surveillance.","confidence":"LOW","assumptions":["Approximately 0.5–2 combined staff equivalents plus ordinary hosting and tooling.","No major redesign, new jurisdiction, or source-system migration is included.","State labels and action mappings are reviewed whenever reporting rules change.","The range is resource-equivalent and lacks procurement or wage validation."],"source_ids":["S4","S6"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official guidance, deployed monitoring documentation, and primary research support missing-versus-zero ambiguity, silent failure, multiple missingness causes, and decision effects; exact false-closure prevalence remains unknown.","source_ids":["S1","S4","S5","S7","S8"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"CDC dashboard owners, NSSP site administrators and inspectors, and official-statistics dashboard producers are identifiable operational actors with documented data-quality and dashboard-governance responsibilities, although none has committed to this proposal.","source_ids":["S3","S4","S6"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The residual claim isolates cause diagnosis plus one matched action against an equally salient generic warning and specifies beneficial and adverse behavioral outcomes.","source_ids":["S3","S5","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A preregistered synthetic-case randomized simulation can hold context and action availability constant, compare two panels, and use predefined response keys without changing real cases.","source_ids":["S3","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The proposed first step is non-production, synthetic, reversible, and leaves all real authority untouched. Cause accuracy, escalation bias, privacy, accessibility, and local legal authority remain mandatory halt checks before live use rather than stops on the simulation.","source_ids":["S3","S6","S8"]},"credible_cost_scope_and_range":{"status":"YES","reason":"All four ranges are tied to bounded resource-equivalent scopes and explicit staffing/integration assumptions. Confidence is only moderate or low because no procurement quote or validated 2026 wage basis was found.","source_ids":["S3","S4","S6","S7"]}},"next_evidence_step":"With one authorized dashboard owner, preregister a randomized synthetic-case simulation covering confirmed no events, report not received, delayed feed, permission block, and system failure. Randomize regulator-like participants between (A) an equally salient generic missing-data panel with the same audit context and underlying actions and (B) that panel plus a validated cause label and exactly one matched action. Hold entity facts, period, prior values, provenance, non-focal spacing, and action availability constant. Have independent oversight and data-quality reviewers establish the response key before recruitment. Use false no-incident closure as the primary endpoint; secondary endpoints are correct cause-consistent follow-up, cause comprehension, context retrieval, decision time, confidence, inappropriate escalation, and delay. Stratify by cause, participant expertise, and reporter capacity; include cases with and without under-reporting incentives. Predefine a minimally important effect and power the test accordingly. Falsify the claim if the matched-action condition fails to improve the primary endpoint by that threshold, worsens inappropriate escalation or delay beyond a prespecified non-inferiority margin, or performs worse for low-capacity reporters. Halt if cause labels cannot be independently validated or accessibility/context retrieval fails.","blocking_evidence":["No representative estimate of false no-incident closure under the current or generic-warning interface was found.","No regulator or dashboard owner has committed staff, participants, data definitions, or funding.","No controlled evidence compares cause-matched actions with an equally salient generic missing-data warning on the proposed outcomes.","The accuracy, coverage, and auditability of the five-state cause classifier have not been demonstrated in a target system.","The minimally important effect size, sample-size calculation, and acceptable inappropriate-escalation margin are unset.","Local legal, privacy, records, accessibility, labor-workflow, and enforcement-authority review is required before any live pilot.","No procurement evidence or validated 2026 labor-price source supports the numerical cost bands.","World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This bounded open-web review establishes substantial collision with public-health dashboards, technical alerting products, a health-data standard, official guidance, and missingness-visualization research. It does not measure world novelty, patentability, freedom to operate, market size, or realized impact. Unindexed patents, internal regulator systems, procurement documents, non-English materials, proprietary evaluations, and unpublished studies may contain the exact comparison or closer prior art.","arm":"SENTINEL","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":true,"material_progress_observed":true,"progress_targets":["Secure a named dashboard owner and data-governance authorizer for a synthetic, non-production study.","Audit a target data pipeline to estimate the prevalence of ambiguous empty entity-periods and validate whether the five causes are observable and mutually usable.","Define and independently adjudicate cause-specific correct actions, including legitimate alternatives and permission boundaries.","Preregister the comparator, minimally important effect, power analysis, adverse-outcome margins, subgroup analysis, and falsifiers.","Recruit representative regulator-like decision-makers and run the randomized simulation.","Complete accessibility, confidentiality, cause-misclassification, automation-bias, and disparate-escalation reviews before considering live deployment.","Obtain jurisdiction-specific legal and operational authorization and a rollback plan before any production pilot.","Validate cost assumptions with partner staffing data or procurement quotes."],"reason":"Bounded web research has answered the public-evidence questions: the problem class is real, credible operational actors exist, implementation is plausible, and prior art substantially collides with nearly every component. The only meaningful residual claim is causal and comparative; resolving it requires representative participants, validated case ground truth, proprietary workflow information, and controlled testing rather than more general web search."}}