{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"negative_space_design__astronomy_astrophysics:P4:v0","cell_id":"negative_space_design__astronomy_astrophysics","search_queries":["astronomy time series plot gaps connect points missing observations misleading continuity upper limits","astronomical light curve plotting missing data gaps upper limits quality flags standard","visualization missing data connected line chart gaps study","official astronomy figure guidelines accessibility color plots AAS","site:ivoa.net time series data model quality upper limit nondetection standard","site:docs.astropy.org time series masked values plot gaps","site:matplotlib.org stable gallery masked_demo line gaps official","site:rubinobservatory.org light curve flags nondetection data quality documentation","Completeness of the Gaia-verse I gaps observations bias study MNRAS 2020 official","BLS software developers median wage 2025 official occupational employment statistics"],"sources":[{"source_id":"S1","title":"Where's My Data? Evaluating Visualizations with Missing Data","publisher":"IEEE Computer Society","url":"https://cmci.colorado.edu/visualab/papers/song_VIS_2018.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2019-01-01","accessed_at":"2026-08-03","claims_supported":["Missing-data encodings in line charts affect perceived quality, credibility, and confidence.","The study directly compared seven line-chart treatments, including absent data and disconnected geometry.","The tested visualization treatments did not significantly change trend-detection accuracy, while data-absent and disconnected treatments reduced perceived quality.","A controlled reader experiment is an appropriate method for evaluating the proposal, but existing results do not establish its claimed advantage for astronomy-trained readers."]},{"source_id":"S2","title":"Instrumental Noise in Kepler and K2 #1: Data Gaps and Quality Flags","publisher":"Lightkurve Project","url":"https://lightkurve.github.io/lightkurve/tutorials/2-creating-light-curves/2-2-kepler-noise-1-data-gaps-and-quality-flags.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Astronomical light curves visibly contain gaps arising from multiple instrument and operational causes.","Kepler files contain cadence-level quality flags, and Lightkurve masks flagged cadences by default.","Some unusable cadences are represented as NaNs or absent timestamps.","Data on both sides of gaps must be treated carefully, and some quality events are not evident from gaps alone.","Lightkurve and its maintainers constitute an identifiable potential implementation venue."]},{"source_id":"S3","title":"Masked Demo","publisher":"Matplotlib Development Team","url":"https://matplotlib.org/3.0.2/gallery/lines_bars_and_markers/masked_demo.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2018-11-11","accessed_at":"2026-08-03","claims_supported":["Mainstream plotting software already supports masked values that break a line at data gaps.","The basic intervention of terminating connecting geometry is technically simple and established rather than novel.","A renderer can implement gaps without modifying source measurements."]},{"source_id":"S4","title":"IVOA Spectrum Data Model Version 1.2","publisher":"International Virtual Observatory Alliance","url":"https://www.ivoa.net/documents/SpectrumDM/20231215/REC-SpectrumDM-1.2-20231215.pdf","source_class":"STANDARD","publication_date":"2023-12-15","accessed_at":"2026-08-03","claims_supported":["An astronomy interoperability standard separately represents statistical upper and lower limits.","The standard provides point-level quality values and permits named meanings for quality codes.","Quality value 1 can indicate bad data or no data in a sample interval, while additional codes can distinguish other bad or dubious states.","The proposal can build on existing machine-readable status and limit fields, but the standard does not prescribe protected blank plotting bands."]},{"source_id":"S5","title":"Graphics Guide","publisher":"American Astronomical Society Journals","url":"https://journals.aas.org/graphics-guide/","source_class":"OFFICIAL_GUIDANCE","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["AAS Journals are an identifiable authorizer of astronomy publication figures.","AAS calls for high-quality, accessible scientific visualizations and warns against trusting software defaults.","AAS advises against using color as the only distinguishing delimiter and supports differentiated line styles, symbols, and hatching.","AAS supports interactive time-series figures and access to underlying data, but does not express demand for this specific evidence-gap specification."]},{"source_id":"S6","title":"DASCH: DR7 Lightcurve Reduction Guide","publisher":"Center for Astrophysics | Harvard & Smithsonian, DASCH","url":"https://dasch.cfa.harvard.edu/dr7/reduce-lightcurve/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["A production astronomical archive exposes quality flags and standard rejection workflows for light curves.","DASCH warns that automated rejection cannot fully clean data and that its behavior can change.","Some apparent upper limits arise from incorrect astrometry or missing compiled detections, demonstrating that absence and nondetection causes can be confounded.","Recoverable flags and provenance are necessary guardrails for any simplified rendering."]},{"source_id":"S7","title":"Completeness of the Gaia-verse – I. When and where were Gaia's eyes on the sky during DR2?","publisher":"Oxford University Press for the Royal Astronomical Society","url":"https://academic.oup.com/mnras/article/497/2/1826/5873016","source_class":"PRIMARY_RESEARCH","publication_date":"2020-07-17","accessed_at":"2026-08-03","claims_supported":["Astronomical observation and processing gaps can materially affect completeness and scientific interpretation.","The authors identified 94 Gaia photometric gaps longer than 1 percent of a day and distinguished data-taking, processing, and efficiency effects.","The paper used white gaps, coverage-like maps, hatching, and textual annotations, providing close prior art for visually encoding absence and its cause.","Seven of thirty cited astrometric microlensing events fell within identified gaps, illustrating potential scientific consequences.","Gap boundaries are approximate and distinguishing a true gap from reduced detection efficiency can be difficult."]},{"source_id":"S8","title":"Table 1. National employment and wage data from the Occupational Employment and Wage Statistics survey by occupation, May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.t01.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["May 2025 mean wages were $71.20 per hour for software developers, $53.60 for software quality-assurance analysts and testers, and $63.91 for astronomers.","These wage benchmarks support order-of-magnitude 2026 resource-equivalent estimates when combined with explicit overhead, staffing, and duration assumptions."]}],"problem_evidence":{"support":"MODERATE","rationale":"Astronomy documentation and primary research verify that time-series gaps, cadence exclusions, missing timestamps, quality flags, upper limits, and multiple causes of absence are real and scientifically consequential. Generic visualization research verifies that missing-data rendering affects reader judgments. However, no source audits the prevalence of the proposal's exact failure mode—astronomy figures joining valid points across unsupported intervals—or shows that readers presently mistake those joins for measured continuity. The problem is externally credible but its plot-level prevalence and error rate remain unmeasured.","source_ids":["S1","S2","S4","S6","S7"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"AAS Journals can authorize figure requirements and explicitly seek readable, accessible, non-color-only graphics. Lightkurve and DASCH are identifiable software/archive adopter classes that already maintain gap, flag, and rejection behavior. Their documentation expresses a need to interpret gap causes and retain quality context, but none requests, commits to, or funds the proposed specification. Adoption pull is therefore credible but indirect.","source_ids":["S2","S5","S6"]},"prior_art":{"proximity":"ESTABLISHED_PRACTICE","closest_analogues":[{"name":"Matplotlib masked-line gaps","similarity":"Directly removes connecting line geometry at masked or missing samples and is documented specifically for gappy data.","remaining_difference":"It supplies neither cadence-sensitive interval classification nor mandatory labels, upper-limit preservation, provenance recovery, or comparative reader testing.","source_ids":["S3"]},{"name":"Lightkurve gap, NaN, quality-flag, and shaded-interval workflows","similarity":"Astronomy-specific tooling already exposes multiple gap causes, masks rejected cadences, plots gappy light curves, and sometimes highlights intervals.","remaining_difference":"It does not establish a uniform protected-blank specification with a controlled vocabulary and a mandatory recovery path across archives and publications.","source_ids":["S2"]},{"name":"Song and Szafir missing-data line-chart treatments","similarity":"Direct experimental prior art compares absent data, disconnected lines, annotations, gradients, and other treatments using reader tasks.","remaining_difference":"The participants and tasks were not astronomy-specific; the study did not compare a classified protected gap against an aligned observation-coverage raster, and it did not preserve astronomy-specific nondetections and quality ledgers.","source_ids":["S1"]},{"name":"Gaia gap and coverage visualizations","similarity":"Primary astronomy research detects bounded observation and processing gaps and represents them using white regions, coverage maps, hatching, and explanatory annotations.","remaining_difference":"It is a study-specific analysis rather than a reusable measurement-layer rule, and it does not test reader interpretation against conventional joined light curves.","source_ids":["S7"]},{"name":"IVOA quality and statistical-limit fields","similarity":"Provides standardized machine-readable primitives for distinct quality states and evidence-bearing upper limits.","remaining_difference":"It standardizes data representation, not visual geometry, protected whitespace, reveal behavior, or effectiveness criteria.","source_ids":["S4"]},{"name":"DASCH flags, rejections, and anomalous upper-limit diagnosis","similarity":"Preserves quality decisions and distinguishes cases in which an apparent upper limit reflects a pipeline or catalog problem.","remaining_difference":"It is a reduction and diagnostic workflow rather than a general plotting specification or tested negative-space intervention.","source_ids":["S6"]}],"distinctive_claim_remaining":"Given identical archived astronomy time series and a preregistered cadence rule, a classified protected gap in the primary measurement geometry will increase astronomy-trained readers' correct identification of measurement support and absence cause versus both a conventional joined plot and an aligned coverage-raster rival, while remaining noninferior for cadence comparison, event-boundary judgment, upper-limit recovery, model-versus-measurement discrimination, and accessible status retrieval. This comparative performance claim is falsifiable; the constituent practices of breaking lines, encoding quality, retaining upper limits, and showing coverage are already established.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Technical feasibility is strong: masked or NaN values can break lines in Matplotlib, astronomy tools already carry cadence quality information, and IVOA fields can represent quality and statistical limits. A read-only renderer over frozen data is reversible and low risk. Workflow feasibility depends on mapping archive-specific flags into a governed vocabulary, defining defensible cadence-sensitive thresholds for irregular sampling, retaining excluded rows through static exports, and separating models from measurements. AAS guidance supports accessibility guardrails. No general legal prohibition was found, but data licenses, attribution, human-subjects review for reader testing, journal policy, and the responsible scientist's authority must be checked locally. No maintainer has approved integration.","source_ids":["S2","S3","S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Observation gaps can affect completeness, event interpretation, and quality cuts, but the frequency and severity of errors caused specifically by joined plotting geometry are unknown.","source_ids":["S2","S7"]},"stakeholder_pull":{"score":2,"rationale":"Credible authorizers and adopter classes exist, yet no source expresses demand, funding, or implementation commitment for this exact specification.","source_ids":["S2","S5","S6"]},"incremental_advantage":{"score":2,"rationale":"The proposed comparison is meaningful, but no evidence shows superiority over coverage rasters or status glyphs. Existing experiments found no trend-accuracy advantage and found lower perceived quality for absent or disconnected treatments.","source_ids":["S1","S7"]},"distinctiveness_plausibility":{"score":2,"rationale":"The integrated rule is more specific than any single source, but its principal components are established plotting behavior, standards, and astronomy practice. Only the registered comparative effect remains distinctive and unverified.","source_ids":["S1","S2","S3","S4","S7"]},"technical_implementability":{"score":5,"rationale":"Existing plotting libraries, masks, quality fields, upper-limit fields, and archive metadata make a read-only prototype straightforward.","source_ids":["S2","S3","S4","S6"]},"adoption_authority_feasibility":{"score":4,"rationale":"Pipeline maintainers, responsible scientists, archives, and journals have clear authority over renderers and figures; implementation can begin off-production. Willingness and vocabulary governance are not established.","source_ids":["S2","S5","S6"]},"evidence_readiness":{"score":2,"rationale":"Archived data and established experimental methods support a bounded study, but there is no prevalence audit, astronomy-reader experiment, preregistered threshold, or production validation.","source_ids":["S1","S2","S7"]},"safety_net_benefit":{"score":3,"rationale":"Frozen-data alternative rendering is reversible and can retain provenance, upper limits, and non-color cues. Benefits depend on export and recovery controls that have not been tested.","source_ids":["S4","S5","S6"]},"scalability":{"score":4,"rationale":"Once a status mapping and rule are defined, common plotting libraries and standardized fields could support reusable implementation. Archive-specific flags and irregular cadences limit automatic portability.","source_ids":["S2","S3","S4","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Prevalence audit, renderer prototype, ledger-to-geometry validation, and a preregistered four-arm interpretation study using at most 24 archived time series and 36–48 astronomy-trained readers.","confidence":"MODERATE","assumptions":["Eight to sixteen weeks of combined visualization engineering, astronomy curation, experimental design, analysis, and project management.","Participant honoraria, accessibility review, and human-subjects/ethics determination are included.","BLS base wages are converted to resource-equivalent cost using roughly 1.5–2.0 times wage for benefits, overhead, and administration.","No new observations or proprietary data acquisition are required."],"source_ids":["S1","S2","S3","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Integrate the verified rule into one plotting package or archive viewer, including status mapping, tests, documentation, accessible static export, and rollback controls.","confidence":"MODERATE","assumptions":["Approximately 0.5–1.5 full-time-equivalent years across software, astronomy/data-quality, design, and QA roles.","One well-documented archive schema is targeted first.","Existing plotting infrastructure and metadata remain usable."],"source_ids":["S2","S3","S4","S5","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Productionize across two or three heterogeneous archives or publication workflows, govern a status vocabulary, validate exports, migrate templates, train users, and monitor interpretation defects.","confidence":"LOW","assumptions":["Two to five full-time-equivalent years distributed across a 6–12 month launch.","Archive-specific quality flags require adapters and scientific review.","Launch includes accessibility, documentation, release engineering, governance, and support but no telescope or instrument changes.","Institutional overhead is included."],"source_ids":["S2","S4","S5","S6","S8"]},"annual_recurring":{"band_2026_usd":"10K_TO_50K","scope":"Maintain mappings, regression fixtures, documentation, accessibility checks, and periodic review for one mature implementation.","confidence":"LOW","assumptions":["Approximately 0.1–0.25 technical or scientific full-time equivalent per year after stabilization.","Major archive-schema migrations, new research studies, and cross-ecosystem expansion are excluded.","BLS wages are treated as a floor before benefits and institutional overhead."],"source_ids":["S4","S5","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent research and astronomy documentation establish real observation gaps, multiple absence causes, quality exclusions, and consequences of ignoring coverage. The exact prevalence of misleading joined plots remains an evidence gap but does not negate the supported problem class.","source_ids":["S1","S2","S6","S7"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"AAS Journals are a figure-policy authorizer, while Lightkurve and DASCH maintain astronomy time-series rendering and quality workflows. Specific willingness to adopt is unverified.","source_ids":["S2","S5","S6"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The protected classified gap can be compared against conventional joined plots, aligned coverage rasters, and dense status-glyph plots using prespecified accuracy and noninferiority outcomes.","source_ids":["S1","S2","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A frozen-data, at-most-24-series study with fixed renderers, trained readers, geometry audits, comparators, primary endpoints, and explicit falsifiers is bounded and reversible.","source_ids":["S1","S2","S3","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The authorized first step changes only alternative renderings of frozen archived data. It can halt on status mismatch, lost provenance, hidden upper limits, inaccessible encoding, or geometry-ledger disagreement. Local ethics review and data-license checks are required before recruiting readers but are not an intrinsic stop.","source_ids":["S4","S5","S6"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four estimates state implementation scope, staffing assumptions, exclusions, and uncertainty, and are anchored to current government wage benchmarks for software, QA, and astronomy roles.","source_ids":["S8"]}},"next_evidence_step":"Preregister a read-only study using no more than 24 frozen archived time series selected before rendering to cover detections, valid nondetections, rejected samples, planned non-observation, instrument/data-taking gaps, and unavailable processing outputs. For identical data, generate four treatments: conventional joined baseline, protected classified evidence gaps, an aligned coverage raster, and explicit status glyphs. Independently verify every vector or pixel segment and displayed status against the acquisition/quality ledger. Recruit 36–48 astronomy-trained readers who did not author the figures; randomize treatment and series; record correct support-boundary identification, absence classification, measured-versus-modeled continuity, upper-limit and rejection recovery, cadence and event-boundary answers, time, confidence, and accessibility failures. The primary claim passes only if protected gaps show a preregistered practically meaningful accuracy improvement over both baseline and coverage raster, with multiplicity control, while meeting preregistered noninferiority margins on cadence, event-boundary, model-comparison, retrieval-time, and accessibility outcomes. Falsify the intervention if either principal comparator matches or exceeds it, if readers misread planned non-observation as failure, if upper limits or rejected-sample provenance become less recoverable, or if results materially depend on a post hoc gap threshold.","blocking_evidence":["No audit quantifies how often published or operational astronomy plots join measurements across unsupported intervals or how often readers infer measured continuity from them.","No astronomy-trained reader study establishes superiority over an aligned coverage raster or explicit status-glyph rival.","No Lightkurve, archive, journal, or other maintainer has committed to adoption, funded work, or accepted responsibility for the status vocabulary.","Cross-archive mappings among no exposure, nondetection, rejection, processing unavailability, and other states have not been validated for semantic equivalence.","Cadence-sensitive gap thresholds have not been shown robust across irregular, multiband, heterogeneous, or highly fragmented time series.","Static-export recovery of upper limits, excluded samples, model separation, and accessible non-color cues has not been tested."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This evaluation measured neither world novelty nor patentability, freedom to operate, market size, or realized impact. The search establishes that line breaking at missing values, astronomy quality flags, explicit upper limits, coverage displays, shaded or white gap regions, and experimental comparison of missing-data encodings all predate the proposal. It did not establish whether the exact combination of a preregistered cadence rule, protected primary-layer blankness, classified cause, recoverable exclusions, model separation, and comparative astronomy-reader criterion has appeared elsewhere.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Complete a reproducible prevalence audit of recent astronomy light-curve figures and operational plotting defaults.","Secure a named archive, plotting-package, or journal partner with authority over status semantics and alternative rendering.","Freeze and publish the crosswalk from source-ledger states to detection, valid nondetection, rejection, no exposure, and unavailable-output categories.","Implement four ledger-verified comparator renderers without changing underlying data, masks, models, or conclusions.","Run the preregistered astronomy-trained reader study and demonstrate superiority over both the joined baseline and coverage raster under a practically meaningful margin.","Demonstrate noninferiority for cadence and event-boundary interpretation and zero material loss of upper-limit, exclusion, provenance, model-separation, and accessibility recovery."],"reason":"Bounded web research resolves feasibility and shows substantial established prior art, but it cannot determine the remaining contrastive claim. The decisive evidence requires a controlled study with astronomy-trained readers and archive-specific ledgers, plus live partner decisions about status governance. Under the evaluation rule, evidence requiring fieldwork or live testing mandates STOP_EMPIRICAL_RESEARCH_NEEDED; all STOP recommendations are non-repairable in this controller output."},"proposal_index":4}