{"schema_version":1,"research_id":"eoa_inverse_innovation_exp03_external48_20260801","source_assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","cell_id":"layer_decay_and_expiration_management__data_science","selection_stratum":"DEPLOYABLE_PRIORITY","search_queries":["site:learn.microsoft.com Azure Machine Learning archive assets restore archived data models environments components","site:docs.datahub.com deprecation status datasets search documentation DataHub","site:docs.open-metadata.org data asset lifecycle retention soft delete restore lineage","empirical study stale Jupyter notebooks outdated datasets reproducibility analytical artifacts","data scientists old stale notebooks accumulation study interview cleanup notebooks provenance reuse","machine learning technical debt dead experimental code paths undeclared consumers paper","site:learn.microsoft.com Microsoft Purview retention disposition review audit records management official","Federal Reserve SR 11-7 model risk management inventory documentation official PDF","site:bls.gov Occupational Employment and Wage Statistics data scientists May 2025 software developers compliance officers","site:learn.microsoft.com/en-us/cli/azure/ml/model archiving model hide restore","site:learn.microsoft.com/en-us/cli/azure/ml/component archive restore Azure ML"],"sources":[{"source_id":"S1","title":"Managing Messes in Computational Notebooks","publisher":"ACM CHI / Microsoft Research","url":"https://www.microsoft.com/en-us/research/wp-content/uploads/2019/01/Managing_Exploratory_Messes_in_Computational_Notebooks-2.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2019-05-04","accessed_at":"2026-08-02","claims_supported":["Computational notebooks commonly develop disorder, deleted or overwritten state, dispersed dependencies, and other forms of clutter.","The study observed analysts preserving and recovering old versions and identified duplicated, unrecognized parent-child notebook relationships.","Twelve professional analysts evaluated dependency-aware notebook-cleaning and version-recovery tools."]},{"source_id":"S2","title":"Strategies for Reuse and Sharing among Data Scientists in Software Teams","publisher":"IEEE/ACM ICSE-SEIP","url":"https://doi.org/10.1145/3510457.3513042","source_class":"PRIMARY_RESEARCH","publication_date":"2022-05-21","accessed_at":"2026-08-02","claims_supported":["Interviews with 17 and a survey of 132 professional data scientists found multiple forms of past-analysis reuse.","A shared analysis store could be out of date, while shared code was commonly reused more frequently than new code was contributed.","Participants reported reuse obstacles involving incentives, modularity, interoperability, discoverability, and the effort required to clean artifacts."]},{"source_id":"S3","title":"Hidden Technical Debt in Machine Learning Systems","publisher":"NeurIPS","url":"https://proceedings.neurips.cc/paper/2015/file/86df7dcfd896fcaf2674f757a2463eba-Paper.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2015","accessed_at":"2026-08-02","claims_supported":["ML systems accumulate maintenance debt involving data dependencies, undeclared consumers, changing external conditions, and dead experimental paths.","Data dependencies can be difficult to detect without dedicated tooling, making dependency chains difficult to untangle."]},{"source_id":"S4","title":"Azure CLI: az ml command reference","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/cli/azure/ml?view=azure-cli-latest","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"","accessed_at":"2026-08-02","claims_supported":["Azure Machine Learning exposes archive and restore operations for data assets, feature sets, feature-store entities, models, jobs, components, environments, and deployment templates.","Many of the relevant lifecycle commands are generally available, demonstrating that reversible lifecycle primitives already span several analytical artifact classes."]},{"source_id":"S5","title":"Azure CLI: az ml data","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/cli/azure/ml/data?view=azure-cli-latest","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"","accessed_at":"2026-08-02","claims_supported":["An archived data asset is hidden by default from list queries but remains referenceable and usable in workflows.","Data-asset containers or individual versions can be archived and subsequently restored."]},{"source_id":"S6","title":"Azure CLI: az ml component","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/cli/azure/ml/component?view=azure-cli-latest","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"","accessed_at":"2026-08-02","claims_supported":["Versioned pipeline components can be archived, hidden from ordinary listing, remain usable by reference, and be restored.","The same reversible discovery-demotion pattern used for data assets applies to executable analytical components."]},{"source_id":"S7","title":"How to Delete a Data Asset","publisher":"OpenMetadata","url":"https://docs.open-metadata.org/v1.12.x/how-to-guides/guide-for-data-users/delete","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"","accessed_at":"2026-08-02","claims_supported":["OpenMetadata recommends removing outdated or test catalog assets to maintain relevant search results.","Deletion can destroy ownership, lineage, usage, test, profiling, and other graph metadata that is difficult to recreate.","OpenMetadata distinguishes read-only soft deletion from permanent hard deletion."]},{"source_id":"S8","title":"Learn about records management","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/purview/records-management","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2025-09-26","accessed_at":"2026-08-02","claims_supported":["Microsoft Purview implements retention labels, retention schedules, disposition review, deletion evidence, activity logging, and specialized permissions.","Record labels can block deletion and other actions, illustrating that disposition authority and retention controls are established system capabilities."]}],"problem_evidence":{"support":"MODERATE","rationale":"Primary studies directly document notebook clutter, duplicated or unrecognized versions, costly cleanup, frequent reuse of prior analysis code, and shared stores that may be out of date. ML-systems research independently documents hard-to-detect data dependencies and undeclared consumers. OpenMetadata explicitly identifies outdated catalog assets as a search-relevance problem and warns that deletion can destroy difficult-to-recreate lineage and ownership metadata. The bounded evidence does not establish the candidate's full prevalence claim across datasets, features, models, notebooks, and reports, nor does it measure how often a superseded artifact is reused as current evidence.","source_ids":["S1","S2","S3","S7"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Professional data scientists report recurring reuse, cleanup, discoverability, and interoperability burdens, while current platform and records-management products assign lifecycle functions to platform operators, records managers, reviewers, and control roles. This establishes plausible users and authorizers, but no source documents a purchase commitment or adoption request for the proposed unified service.","source_ids":["S1","S2","S4","S8"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Azure Machine Learning multi-asset archive and restore","similarity":"Azure ML already supplies version-aware archive and restore operations across data, feature, model, job, component, environment, and related asset classes. For data and components, archival removes assets from default listing while preserving references and later restoration, closely matching reversible discovery demotion.","remaining_difference":"The documentation does not establish semantic supersession detection, retention-class and hold evaluation, verified inbound-dependency completeness as an archive gate, scheduled restore drills, exception aging, or one policy spanning notebooks and reports.","source_ids":["S4","S5","S6"]},{"name":"Microsoft Purview records lifecycle and disposition review","similarity":"Purview supplies retention labels, protected records, role-limited disposition review, audit history, holds or suspended deletion, relabeling, archival choices, and proof of disposition. These closely match the candidate's authority, retention, exception, and audit mechanisms.","remaining_difference":"Purview's documented scope is records and Microsoft 365 content rather than dependency-aware analytical registries; the reviewed documentation does not rank analytical discovery, infer model or dataset supersession, or test reconstruction fidelity.","source_ids":["S8"]},{"name":"OpenMetadata soft and hard deletion","similarity":"OpenMetadata connects catalog search cleanliness with deletion of outdated or test assets, represents ownership and lineage metadata, and provides reversible read-only soft deletion versus permanent hard deletion.","remaining_difference":"The documented workflow is user-initiated deletion, not a retention-authorized lifecycle service with supersession evidence, dependency gates, restore drills, holds, and disposition-specific authority.","source_ids":["S7"]}],"distinctive_claim_remaining":"The remaining testable claim is not generic archiving, catalog deprecation, retention, or restoration. It is that one cross-artifact control plane can combine adjudicated supersession evidence, ordinary-search demotion, retention and hold authority, verified inbound-dependency gates, reversible quarantine, scheduled restoration verification, and durable disposition records across datasets, features, models, notebooks, and reports, producing better search currency or lower triage effort than configured existing lifecycle tools without additional safety failures.","confidence":"HIGH"},"implementation_evidence":{"support":"STRONG","rationale":"Commercially documented systems already implement most individual primitives: multi-class asset identity and versioning, archive and restore, default-list suppression, retained references, soft versus hard deletion, protected retention states, disposition roles, and audit evidence. This strongly supports technical feasibility for a one-class shadow pilot. It does not verify reliable semantic staleness classification, complete cross-system lineage, or economical organization-wide integration.","source_ids":["S4","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Research supports meaningful cleanup, reuse, discoverability, dependency, and maintenance burdens, and official catalog documentation recognizes outdated assets as a search-quality issue. The exact rate of harmful stale reuse remains unmeasured.","source_ids":["S1","S2","S3","S7"]},"stakeholder_pull":{"score":3,"rationale":"Analysts demonstrably reuse prior work and incur cleaning and discovery costs, while platform and records-management roles already perform adjacent lifecycle work. No adopter commitment for this integrated mechanism was found.","source_ids":["S1","S2","S4","S8"]},"incremental_advantage":{"score":3,"rationale":"The proposal adds dependency-gated, policy-authorized disposition and restore verification to capabilities already widely represented in Azure ML, OpenMetadata, and records-management tooling. The incremental operational benefit has not been demonstrated.","source_ids":["S4","S5","S6","S7","S8"]},"distinctiveness_plausibility":{"score":2,"rationale":"A substantial portion of the claimed composition already exists across closely adjacent first-party systems, especially Azure ML's multi-asset reversible archive and Purview's governed disposition. Distinctiveness is confined to their cross-artifact integration with dependency gates and tested restoration.","source_ids":["S4","S5","S6","S7","S8"]},"technical_implementability":{"score":4,"rationale":"Existing generally available archive and restore operations across multiple ML asset classes make a shadow integration technically credible. Semantic supersession, heterogeneous identifiers, and lineage completeness remain difficult components.","source_ids":["S3","S4","S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"Existing records systems demonstrate specialized permissions and protected retention states, so the proposed authority partition is credible. Coordinating artifact owners, platform teams, privacy, legal, audit, and model-risk functions across systems remains a material adoption burden.","source_ids":["S7","S8"]},"evidence_readiness":{"score":4,"rationale":"Existing archive, restore, and listing controls make an offline comparison possible without production deletion. Search exposure, adjudication accuracy, reviewer time, dependency blocks, audit completeness, and restoration fidelity are measurable within one artifact class.","source_ids":["S4","S5","S6","S7"]},"safety_net_benefit":{"score":4,"rationale":"Reversible archival, continued referenceability, restoration, soft deletion, protected retention states, and disposition evidence are established safety mechanisms. Their benefit depends on actually testing recovery and preventing incomplete lineage from authorizing destructive action.","source_ids":["S3","S5","S6","S7","S8"]},"scalability":{"score":3,"rationale":"Azure ML demonstrates that a common archive-and-restore interface can span several ML artifact types, but notebooks, reports, external stores, semantic validity rules, dependencies, and retention obligations remain heterogeneous.","source_ids":["S2","S4","S7","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"A 30-day offline shadow study for one non-production artifact class: registry extraction, identity resolution, owner adjudication, simulated policies, comparison against the existing catalog or archive workflow, restoration of a stratified sample, safety review, and analysis.","confidence":"MODERATE","assumptions":["Existing registry, archive, and access-control capabilities are reused.","The study covers roughly one hundred metadata records and a small stratified restore sample rather than production-scale data movement.","No hard deletion or production search change occurs.","Part-time effort is required from a platform engineer, data scientist, artifact owners, and governance or compliance reviewer.","The band includes loaded labor, coordination, software, storage and retrieval, compliance review, and evaluation; official wage data is used only as a labor-cost anchor."],"source_ids":["S4","S5","S7"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"A production-capable service for a small number of artifact classes, including connectors, identity mapping, lifecycle metadata, discovery integration, policy and hold evaluation, dependency checks, reversible disposition, audit records, restoration procedures, access controls, security review, and training.","confidence":"MODERATE","assumptions":["An existing catalog, lineage graph, identity provider, and tiered object store are available.","Several engineering and governance roles contribute for multiple months.","Metadata remediation is bounded and no enterprise platform replacement is undertaken.","Archive retrieval and transaction charges are included, but infrastructure is expected to be smaller than loaded labor and integration costs.","Hard deletion remains human-authorized and outside initial automation."],"source_ids":["S3","S4","S5","S7","S8"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Controlled organizational launch across datasets, features, models, jobs or notebooks, and reports, including connector development, lineage remediation, policy harmonization, migration, security and privacy validation, restore drills, support processes, owner onboarding, incident preparation, and comparative evaluation.","confidence":"LOW","assumptions":["Artifact systems, identifiers, retention duties, and restore methods are heterogeneous.","Multiple loaded engineering, data-science, compliance, program-management, and support roles are required.","Historical ownership and dependency gaps require substantial remediation.","The scope is one organization, not a multi-enterprise commercial product.","The range includes storage, retrieval, audit, monitoring, software, coordination, training, and evaluation resources."],"source_ids":["S2","S3","S4","S7","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Annual service operation for several artifact classes, including infrastructure, catalog and lineage maintenance, archive storage and retrieval, policy and exception review, metadata stewardship, restore drills, monitoring, incident response, compliance evidence, user support, and outcome evaluation.","confidence":"LOW","assumptions":["Human review continues for contested, sensitive, held, or irreversible dispositions.","At least several part-time operational and governance roles are required.","Archive retrieval latency and transaction costs must be budgeted as well as storage.","Exception aging, tombstones, and lifecycle metadata require continuing stewardship.","Potential storage savings are not deducted from the resource-equivalent band."],"source_ids":["S3","S4","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Primary evidence supports cluttered and duplicated analytical work, out-of-date shared analysis stores, frequent reuse, and hard-to-detect dependencies; catalog documentation independently identifies outdated assets as a search-relevance problem. Local prevalence and harmful reuse frequency remain blocking measurements.","source_ids":["S1","S2","S3","S7"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Azure ML and OpenMetadata establish credible platform-operator implementations, while records-management documentation establishes specialized retention and disposition roles. These are credible adopters and authorizers even though no commitment to this candidate was found.","source_ids":["S4","S7","S8"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"Despite substantial prior-art collision, the remaining claim can be tested directly: compare an integrated supersession, policy, dependency, reversible-disposition, and restore-verification workflow with existing catalog badges or native archive operations on search exposure, triage time, and safety failures.","source_ids":["S4","S5","S6","S7","S8"]},"bounded_next_evidence_step":{"status":"YES","reason":"A one-class, non-production, offline replay with a fixed registry sample and a small archive-and-restore sample is bounded, reversible, and produces a clear comparative result without live deletion.","source_ids":["S4","S5","S6","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The proposed first step excludes hard deletion and production discovery changes. Existing products demonstrate reversible archive, restore, soft deletion, protected retention, and role-limited disposition mechanisms that can implement the stated halts and rollbacks.","source_ids":["S5","S6","S7","S8"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four scopes enumerate labor, integration, governance, software, storage, retrieval, compliance, coordination, support, and evaluation. Existing capabilities reduce first-evidence build effort, while official labor data and storage constraints support broad rather than point estimates.","source_ids":["S4","S5","S7","S8"]}},"next_evidence_step":"Run a 30-day offline shadow comparison on one non-production artifact class using a fixed, stratified sample of approximately 100 registry records. Have owners adjudicate current, superseded, held, and indeterminate status without seeing the lifecycle score. Compare the existing catalog or native archive workflow with the proposed policy-and-dependency-gated workflow on stale items remaining in ordinary search, reviewer minutes, ownership and inbound-dependency coverage, exception rate, and audit completeness. Archive and restore approximately 20 policy-eligible copies, checking content or semantic fidelity and retrieval time. No hard deletion or live search change is allowed. Falsify incremental advantage if the proposed workflow does not improve stale-search exposure or reviewer effort, or if it permits a known dependency or hold to pass, produces an unauthorized exposure, lacks an audit record, or fails a restoration check.","blocking_evidence":["A representative local audit measuring superseded-but-active artifacts and actual reuse of obsolete artifacts as current evidence.","Baseline and treatment measurements of ordinary-search exposure and owner-review time.","Adjudicated precision and recall for supersession labels, separated from artifact age.","Measured ownership and inbound-dependency coverage, including undeclared consumers discovered during review.","Stratified restoration fidelity and retrieval-time results for the target artifact class.","Counts, causes, and aging of legal, privacy, audit, model-risk, and business exceptions.","A head-to-head comparison with configured native archive, catalog deprecation, and records-retention capabilities to establish whether the integration adds material value.","Resource measurements from the shadow study sufficient to refine connector, stewardship, review, storage, retrieval, and support costs."],"research_disposition":"PRIOR_ART_DIFFERENTIATION_STUDY","world_novelty_boundary":"This bounded search found substantial close prior art: reversible multi-asset lifecycle operations in Azure Machine Learning, governed retention and disposition in Microsoft Purview, and soft versus hard catalog deletion in OpenMetadata. It did not find one documented system combining all candidate mechanisms across datasets, features, models, notebooks, and reports with dependency-gated disposition and scheduled restoration verification. That absence is not a world-novelty claim; vendor configurations, internal enterprise systems, patents, and additional literature were not exhaustively searched."}