{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"layer_decay_and_expiration_management__physics:P2:v0","cell_id":"layer_decay_and_expiration_management__physics","search_queries":["site:usqcd.org lattice QCD gauge configurations data management archive storage","ILDG lattice gauge configurations metadata standard official","lattice QCD configuration thinning autocorrelation archive checkpoint reproducibility","NERSC lattice gauge connection gauge configurations archive metadata","lattice QCD gauge configuration data preservation policy configuration retention storage","site:docs.nersc.gov HPSS archive restore purge retention data integrity","lattice QCD Markov chain configurations autocorrelation thinning primary research","lattice QCD reproducibility gauge configurations metadata provenance FAIR primary study"],"sources":[{"source_id":"S1","title":"USQCD Long-term Data Management Strategy and Plan","publisher":"USQCD Collaboration","url":"https://www.usqcd.org/documents/USQCD_tape_DM.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2021-01-12","accessed_at":"2026-08-03","claims_supported":["Gauge ensembles are costly, long-lived scientific resources, while short-term campaign data are frequently deleted and moved between disk and tape.","USQCD describes its prior preservation approach as ad hoc and expressly seeks a coherent data-management strategy.","USQCD already distinguishes short- and long-term lifecycles, tiering, provenance retention, project data managers, storage proposals, site responsibilities, and decisions not to migrate some data.","The USQCD Scientific Program Committee, Executive Committee, facility sites, project principal investigators, and data managers are identifiable authorizers or operators.","The plan requires projects to balance storage and recomputation costs and acknowledges significant recurring tape costs."]},{"source_id":"S2","title":"Specification of ILDG Standards","publisher":"Deutsches Elektronen-Synchrotron DESY / International Lattice Data Grid","url":"https://hpc.desy.de/ildg/specifications/","source_class":"STANDARD","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["ILDG already defines ensemble-level and configuration-level QCDml metadata schemas.","ILDG specifies a binary gauge-configuration file format and catalogue interfaces.","A proposed registry should extend or map to existing identifiers and formats rather than introduce an incompatible metadata system."]},{"source_id":"S3","title":"Lattice Gauge Ensembles and Data Management","publisher":"arXiv; authors representing 16 lattice-QCD collaborations","url":"https://arxiv.org/abs/2502.08303","source_class":"PRIMARY_RESEARCH","publication_date":"2025-02-12","accessed_at":"2026-08-03","claims_supported":["Sixteen collaborations recently reported ensemble-generation, publication, data-management, and storage requirements, demonstrating continuing community attention and potential adopters.","Data-management practices vary across collaborations, making a cross-workflow lifecycle controller a coordination proposition rather than a settled uniform practice."]},{"source_id":"S4","title":"Provenance for Lattice QCD Workflows","publisher":"arXiv; Auge, Bali, Klettke, Ludäscher, Söldner, Weishäupl, and Wettig","url":"https://arxiv.org/abs/2303.12640","source_class":"PRIMARY_RESEARCH","publication_date":"2023-03-22","accessed_at":"2026-08-03","claims_supported":["Lattice-QCD workflows generate and analyze data at multi-petabyte scale.","QCDml contains some configuration-generation provenance, but the authors find that a full workflow provenance concept is incomplete and propose a W3C PROV-based multilayer model.","The paper reports a real case of silent corruption affecting stored configurations, supporting integrity validation and restore testing.","The proposed provenance model supplies adjacent machinery for dependency and lineage tracing but does not itself establish safe deletion decisions."]},{"source_id":"S5","title":"JLab Lattice QCD Facility—Filesystems","publisher":"Thomas Jefferson National Accelerator Facility","url":"https://www.jlab.org/lqcd/filesystems","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["An identifiable LQCD facility already uses disk-to-tape migration, pinning, quotas, age/access criteria, and automatic deletion.","The cache policy deletes the oldest unpinned files only after tape backup, illustrating established tiering and a limited dependency-preservation override.","Existing facility automation is substantially similar to the proposed operational lifecycle, although it is storage-centric rather than estimator-dependency-aware."]},{"source_id":"S6","title":"NERSC Data Policy","publisher":"National Energy Research Scientific Computing Center","url":"https://docs.nersc.gov/policies/data-policy/policy/","source_class":"OFFICIAL_GUIDANCE","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["NERSC distinguishes temporary scratch, community storage, and long-term HPSS tape and applies purge periods to scratch.","HPSS normally stores one tape copy, lacks an off-site backup by default, and therefore retains residual permanent-loss risk.","Projects must contact NERSC for firm retention commitments, confirming facility authority and the need for explicit retention arrangements.","Default allocations and differing availability characteristics make storage-state transitions operationally consequential."]},{"source_id":"S7","title":"Review of Particle Physics: Lattice Quantum Chromodynamics","publisher":"Particle Data Group / Lawrence Berkeley National Laboratory","url":"https://pdg.lbl.gov/2022/reviews/rpp2022-rev-lattice-qcd.pdf","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2022-08-11","accessed_at":"2026-08-03","claims_supported":["Gauge configurations are generated through correlated Markov chains, commonly using HMC.","Autocorrelation lengths can be difficult to estimate accurately, creating uncertainty about equilibration and statistical errors.","Uniform trajectory age or stride is not a sufficient scientific qualification rule because separation requirements depend on measured autocorrelation behavior."]},{"source_id":"S8","title":"The International Lattice Data Grid (ILDG 2.0)","publisher":"arXiv; Francesco Di Renzo","url":"https://arxiv.org/abs/2401.14752","source_class":"PRIMARY_RESEARCH","publication_date":"2024-01-26","accessed_at":"2026-08-03","claims_supported":["ILDG has operated for roughly two decades to share valuable, expensive gauge configurations.","A modernization effort is already pursuing larger FAIR datasets after earlier service availability and usage degraded.","Cataloguing, accessibility, usability, and citability are established objectives, creating substantial prior-art overlap with the proposal's registry and archive aspects."]}],"problem_evidence":{"support":"STRONG","rationale":"The problem is visible at scientific and infrastructure levels: expensive gauge ensembles and multi-petabyte workflows require preservation, short-term data are routinely migrated or deleted, autocorrelation and equilibration status are scientifically consequential, and silent corruption has occurred. USQCD explicitly characterized preservation as ad hoc and called for coherent lifecycle management. The sources do not measure how often stale configurations actually enter estimators or how many deletions break reconstruction, so those specific failure rates remain unknown.","source_ids":["S1","S3","S4","S5","S6","S7"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"USQCD identifies the SPC, EC, facility sites, project PIs, and data managers as allocation and lifecycle authorities; JLab and NERSC visibly operate relevant storage controls; and 16 collaborations have expressed current data-management and storage requirements. This establishes credible authorizers and general need, but no source records a commitment to adopt the proposed per-configuration dependency-aware controller.","source_ids":["S1","S3","S5","S6"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"USQCD Long-term Data Management Strategy and Plan","similarity":"Already separates short- and long-term data, uses disk/tape tiers, requires project DMPs and named data managers, preserves provenance and publication data, weighs storage against recomputation, and authorizes eventual non-migration or deletion.","remaining_difference":"It is primarily project- and dataset-level policy; it does not specify a per-configuration state machine separating statistical ensemble authority from retention value, a machine-checked analysis/restart dependency gate, quarantine, tombstones, and sampled executable restores as one controller.","source_ids":["S1"]},{"name":"ILDG/QCDml catalogues and ILDG 2.0","similarity":"Provide standardized ensemble and configuration identity, metadata, binary formats, catalogues, distributed storage, and FAIR sharing infrastructure.","remaining_difference":"They catalogue and distribute configurations but the reviewed sources do not define estimator-membership states, age-weighted operational ranking, dependency-gated deletion, reversible quarantine, or disposition tombstones.","source_ids":["S2","S8"]},{"name":"W3C PROV-based lattice-QCD workflow provenance model","similarity":"Connects configuration generation and measurement through explicit provenance and can support lineage queries needed for dependency checks.","remaining_difference":"It reports incomplete coverage of important provenance questions and does not turn the graph into an authoritative deletion gate or lifecycle state machine.","source_ids":["S4"]},{"name":"JLab automated LQCD cache management","similarity":"Implements quota pressure, age-based tape migration, pin overrides, backup checks, and deletion of old disk copies.","remaining_difference":"Its policy is storage-centric and does not test statistical authority, analysis-manifest dependencies, restart ancestry, quarantine restoration, or persistent configuration tombstones.","source_ids":["S5"]},{"name":"NERSC scratch-to-HPSS lifecycle","similarity":"Distinguishes temporary, active, and archival storage with purge rules, allocations, retention contacts, and explicit residual backup risks.","remaining_difference":"It is general facility policy rather than configuration-aware scientific governance and provides no evidence of per-configuration estimator or restart dependency decisions.","source_ids":["S6"]}],"distinctive_claim_remaining":"For a completed Markov-chain ensemble, a per-configuration controller that represents scientific ensemble membership and operational retention as separate states, and gates irreversible disposition on explicit analysis/restart dependencies, will identify more false-active and unsafe-to-delete configurations than fixed-stride thinning or whole-ensemble retention while preserving successful reconstruction and archive restore. This is contrastive and falsifiable, but no effect size or superiority evidence was found.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"Configuration identifiers, metadata schemas, catalogues, provenance graphs, storage tiers, pinning, migration, and purge automation already exist, so a read-only registry and shadow classifier are technically credible. Remaining gaps are substantial: notebook and external-copy dependencies may be invisible; statistical classifications are uncertain; archive formats and executability can drift; facility interfaces differ; quarantine temporarily increases storage; and actual deletion requires project and facility authorization. The proposed isolated, non-mutating first step contains these risks.","source_ids":["S1","S2","S4","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"The intervention addresses scientifically expensive, multi-petabyte data whose misclassification, corruption, or premature loss can affect inference and reproducibility, although incident prevalence and realized savings are unmeasured.","source_ids":["S1","S4","S7"]},"stakeholder_pull":{"score":3,"rationale":"USQCD and multiple collaborations explicitly need data-management and storage planning, but expressed demand is for broad DMP and FAIR infrastructure rather than this exact controller.","source_ids":["S1","S3","S8"]},"incremental_advantage":{"score":3,"rationale":"Separating statistical authority from storage state and testing dependencies before disposition is plausibly better than age-only deletion, but superiority over current manifests, pinning, and whole-ensemble policies has not been tested.","source_ids":["S1","S5","S7"]},"distinctiveness_plausibility":{"score":2,"rationale":"Most components have close established analogues in USQCD policy, ILDG/QCDml, provenance research, and facility tiering. Distinctiveness survives mainly in the specific integration and per-configuration decision claim.","source_ids":["S1","S2","S4","S5","S6","S8"]},"technical_implementability":{"score":4,"rationale":"Standards, catalogues, provenance representations, archival tiers, and automated lifecycle controls demonstrate feasible building blocks; completeness of the dependency graph is the principal technical uncertainty.","source_ids":["S2","S4","S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"Scientific leads, project data managers, USQCD allocation bodies, and facility operators have identifiable roles, but joint authority and policy changes would be required for live disposition.","source_ids":["S1","S5","S6"]},"evidence_readiness":{"score":4,"rationale":"A completed ensemble permits a bounded, read-only shadow study with authoritative manifests, independent reviewers, archived copies, and explicit comparators. Required data are nevertheless local or proprietary and unavailable through web research.","source_ids":["S1","S2","S4","S7"]},"safety_net_benefit":{"score":4,"rationale":"Dependency vetoes, quarantine, isolated restores, preservation holds, and retained tombstones directly reduce irreversible-loss risk; sampled restores cannot prove universal recoverability.","source_ids":["S4","S6"]},"scalability":{"score":3,"rationale":"Standards and automated storage controls can scale across large collections, but per-configuration dependency discovery, human exception review, and cross-facility integration may become bottlenecks.","source_ids":["S2","S4","S5","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"A read-only shadow study on one completed ensemble: reconcile a bounded stratified sample, obtain two independent classifications, audit known dependencies, restore archive copies into scratch, and evaluate one documented observable.","confidence":"MODERATE","assumptions":["Approximately 3–8 person-weeks across a physicist, data steward, and research-software engineer.","Existing manifests, catalogues, archive access, and scratch capacity are available.","No source data are migrated, hidden, rewritten, or deleted."],"source_ids":["S1","S2","S4","S6"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Build the configuration registry, metadata adapters, lifecycle-state model, read-only dependency ingestion, review interface, audit log, and facility-specific archive connector for one collaboration.","confidence":"LOW","assumptions":["Existing QCDml or equivalent identifiers cover most configurations.","One or two storage systems and a limited set of manifest formats are integrated.","The figure is a resource-equivalent estimate, not a vendor quotation."],"source_ids":["S1","S2","S4","S5"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Production hardening across a campaign: authorization workflows, migration and quarantine jobs, monitoring, recovery runbooks, validation, training, and staged integration with analysis and restart workflows.","confidence":"LOW","assumptions":["Launch covers one multi-ensemble collaboration rather than all USQCD or ILDG sites.","No new tape library is purchased; existing facility storage is used.","Independent validation is required before enabling irreversible disposition."],"source_ids":["S1","S4","S5","S6"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Ongoing data stewardship, exception review, software maintenance, dependency audits, periodic restore drills, scratch/quarantine capacity, and incremental archive allocation for one collaboration.","confidence":"LOW","assumptions":["Roughly 0.25–1.0 combined FTE plus existing institutional storage allocations.","Storage volume and restore frequency remain within current facility-scale infrastructure.","Major media migrations or geographically independent replicas are excluded and could raise cost materially."],"source_ids":["S1","S4","S6"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official plans, current collaboration reports, provenance research, and facility policies jointly establish costly accumulated configurations, lifecycle pressure, integrity risk, and scientifically meaningful configuration status.","source_ids":["S1","S3","S4","S5","S6","S7"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"USQCD's SPC and EC, project PIs and data managers, and JLab/NERSC facility operators are identifiable authorizers; current lattice collaborations are credible adopters.","source_ids":["S1","S3","S5","S6"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The surviving claim compares a per-configuration two-dimensional lifecycle plus dependency gate against fixed-stride thinning and whole-ensemble retention using false-active findings, missed dependencies, disposition disagreement, and recoverability outcomes.","source_ids":["S1","S4","S5","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"One completed, non-publication-critical ensemble can be studied in shadow mode with stratified configurations, independent classification, known-analysis dependency checks, isolated restores, named comparators, and predeclared falsifiers.","source_ids":["S1","S2","S4","S6","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The evidence step is read-only and non-authoritative, restores only copies into isolated scratch, and leaves all source configurations and estimators unchanged. Live migration or deletion remains outside the authorized step.","source_ids":["S1","S6"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The bands are scoped resource-equivalent estimates grounded in multi-petabyte workflow complexity, existing infrastructure, and documented recurring storage burden, but no site-specific labor rates, configuration counts, storage inventory, or integration assessment was publicly available.","source_ids":["S1","S4","S5","S6"]}},"next_evidence_step":"Partner with one lattice-QCD collaboration to study 100–300 configurations from one completed, non-publication-critical ensemble, stratified across burn-in, production, restart checkpoints, anomalies, and superseded trajectories. Team A uses the proposed registry and dependency rules in read-only shadow mode; an independent physicist/data-steward panel classifies the same configurations without the operational-value score. Compare both with (1) fixed-stride thinning and (2) whole-ensemble retention. Measure authoritative-membership agreement, false-active candidates, known dependency recall, unresolved identities, reviewer time, disposition differences, archive retrieval latency, checksum/parse success, and reproduction of one documented observable. Falsify the incremental claim if the controller misses any known analysis or restart dependency, cannot reproduce the authoritative manifest, changes classifications materially under reasonable reviewer judgments, fails any sampled end-to-end restore, or produces no decision-relevant improvement over either comparator. Do not migrate or delete source data.","blocking_evidence":["No public source measures the prevalence of stale or superseded configurations entering lattice-QCD estimators.","No public evaluation shows that the integrated controller outperforms immutable manifests, existing pinning/tiering policies, fixed-stride selection, or whole-ensemble retention.","Completeness of analysis, notebook, external-copy, and restart dependencies can only be tested against collaboration-controlled records.","Archive readability and executable reconstruction require live tests on real archived copies.","Site-specific configuration counts, staff effort, storage allocations, and integration complexity are needed to validate costs.","No adopter has committed to the proposed per-configuration workflow or authorized production disposition."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This evaluation establishes neither world novelty nor absence of similar unpublished collaboration tooling. Patentability, freedom to operate, market size, and realized impact were not assessed. The searched evidence shows substantial collision with USQCD lifecycle policy, ILDG/QCDml cataloguing, provenance models, and facility tiering; only the integrated per-configuration separation of statistical authority from retention value, coupled to dependency-gated reversible disposition and tested restoration, remains unverified as a distinctive claim.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure a collaboration and named scientific lead/data steward for the read-only shadow study.","Reconcile 100–300 sampled configuration identities with the authoritative ensemble manifest.","Demonstrate 100% recall of seeded and already-known analysis and restart dependencies before any disposition proposal.","Complete stratified end-to-end restores and reproduce one documented observable with predefined integrity tolerances.","Quantify improvement or non-improvement against fixed-stride thinning and whole-ensemble retention.","Produce a site-specific labor, integration, storage, quarantine, and recurring-operations estimate.","Obtain an explicit authority map and approval conditions for any later live migration or deletion trial."],"reason":"Bounded web research verified the problem, credible authorities, implementable building blocks, and a testable residual claim, while also finding substantial prior-art overlap. The remaining questions—dependency completeness, classification reliability, archive recoverability, comparative advantage, adopter acceptance, and site-specific cost—require collaboration-controlled data and live shadow/restore testing rather than further bounded web search."},"proposal_index":2}