{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp04_retrieval_first_paired20_20260802","cell_id":"invariant_mode_decomposition_design__library_information_science","hypothesis_id":"H1","search_queries":["metadata crosswalk validation record-level semantic loss rare collections","metadata mapping SVD reconstruction error schema matching","library metadata migration quality assurance crosswalk validation records","ontology alignment matrix factorization low-rank SVD mapping","patent metadata schema mapping singular value decomposition crosswalk","ETL migration validation rare records stratified sampling semantic defects","digital library metadata migration acceptance testing sample record classes","site:dublincore.org interoperability levels metadata crosswalk loss quality"],"sources":[{"source_id":"S1","title":"Interoperability Levels for Dublin Core Metadata","publisher":"Dublin Core Metadata Initiative","url":"https://www.dublincore.org/specifications/dublin-core/interoperability-levels/","source_class":"STANDARD","claims_supported":["DCMI distinguishes shared-term mapping from formal semantic, record-level syntactic, and profile-constraint interoperability.","A field correspondence alone does not establish that transformed records preserve formal semantics or record constraints."]},{"source_id":"S2","title":"Metadata Quality Control for Content Migration: The Metadata Migration Project at the University of Houston Libraries","publisher":"International Conference on Dublin Core and Metadata Applications","url":"https://dcpapers.dublincore.org/files/articles/952137124/dcmi-952137124.pdf","source_class":"PRIMARY_RESEARCH","claims_supported":["University of Houston Libraries used automated extraction and reports, audits, and iterative human remediation during repository migration preparation.","Required remediation varied by collection, and programmatic analysis exposed anomalies remaining after earlier standardization.","The project traced problematic terms back to individual digital objects; it also reports contextual errors caused by allowing only one mapping from an alternate vocabulary to LCSH."]},{"source_id":"S3","title":"WMS Data Verification Checklist","publisher":"OCLC","url":"https://help.oclc.org/Librarian_Toolbox/WMS_data_checklists/WMS_data_verification_checklist","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","claims_supported":["OCLC's first-party migration checklist requires inspection of individual records and named fields rather than relying only on aggregate totals.","Verification is partitioned across patron, monograph, serial, circulation, and policy data, with discrepancies reported using record identifiers and examples."]},{"source_id":"S4","title":"A Workflow for GLAM Metadata Crosswalk","publisher":"arXiv","url":"https://arxiv.org/abs/2405.02113","source_class":"PRIMARY_RESEARCH","claims_supported":["The paper presents a replicable cultural-heritage metadata-crosswalk workflow using RML to transform tabular data into RDF.","The workflow incorporates domain experts and digital humanists when abstracting and formalizing domain knowledge, but its disclosed method does not select mappings through singular-mode residual gates."]},{"source_id":"S5","title":"A Rational Approach to Legacy Data Validation When Transitioning Between Electronic Health Record Systems","publisher":"Journal of the American Medical Informatics Association","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC11741007/","source_class":"PRIMARY_RESEARCH","claims_supported":["This neighboring-domain migration method combines exhaustive mapped-value testing with manual validation samples calculated separately for each data type using confidence levels and error limits.","It explicitly accounts for validation time and cost and reports that a document subset omitted from the original migration scope was discovered after go-live.","Per-type statistical validation is a close precedent for governed-class acceptance thresholds, although it does not use SVD or reconstruction residual structure."]},{"source_id":"S6","title":"US11726969B2: Matching Metastructure for Data Modeling","publisher":"United States Patent and Trademark Office, mirrored by Google Patents","url":"https://patents.google.com/patent/US11726969B2/en","source_class":"GOVERNMENT_OR_REGULATOR","claims_supported":["The patent represents mappings between source and target schema objects, including many-to-one correspondences, semantic-equivalence assertions, confidence properties, and transformation rule stacks.","It addresses reusable and automatable schema alignment for ETL and migration, but the opened specification contains no singular-value, residual, or rare-class validation mechanism."]}],"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Per-data-type statistical migration validation in electronic health-record conversion","similarity":"Very close to the proposed governance lever: mapped values are tested and migrated records are sampled separately by data type under explicit confidence and error limits, while review time and cost are measured.","remaining_difference":"It validates source-defined data types through statistical sampling rather than decomposing a crosswalk into singular modes, measuring record-by-field reconstruction residuals, or requiring residual tolerances for rare metadata classes.","source_ids":["S5"]},{"name":"OCLC WMS record- and category-level migration verification","similarity":"Uses individual-record inspection, separates several library-record categories, checks named target fields and behaviors, and records concrete migration defects.","remaining_difference":"It is a checklist-based acceptance practice without a learned modal representation, held-out reconstruction test, class-specific numerical residual threshold, or mode-retention stopping rule.","source_ids":["S3"]},{"name":"University of Houston automated metadata migration analysis and remediation","similarity":"Automates repository-wide metadata analysis, exposes collection-dependent and contextual mapping anomalies, and traces defects to individual objects for remediation.","remaining_difference":"It analyzes vocabulary values and rule outcomes rather than reconstructing records from retained singular modes or gating acceptance on out-of-sample class residuals.","source_ids":["S2"]},{"name":"GLAM RML metadata-crosswalk workflow","similarity":"Provides a systematic, expert-governed workflow for transforming heterogeneous cultural-heritage metadata across schemas.","remaining_difference":"The mapping is explicitly modeled through RML and expert knowledge, not low-rank singular modes selected by residual performance on governed record classes.","source_ids":["S4"]},{"name":"Matching metastructure patent for schema mapping","similarity":"Formalizes cross-schema semantic correspondences, many-to-one mappings, mapping confidence, and executable transformation rules for migration and ETL.","remaining_difference":"It structures and maintains mappings but does not disclose singular-mode compression, reconstruction residual analysis, rare-class gates, or defect-versus-review-cost evaluation.","source_ids":["S6"]}],"overlapping_components":["Schema and metadata crosswalk definition","Many-to-one and one-to-many mapping representation","Record-level migration verification","Collection- or data-type-stratified validation","Automated anomaly and exception reporting","Explicit validation error limits","Manual-review time and cost accounting","Semantic, syntactic, and constraint-level interoperability checks","Traceability from detected defects to individual records"],"remaining_contrastive_claim":"No located source selects retained singular modes of a library metadata crosswalk by held-out record-by-field reconstruction residuals, requires every governed rare record class to meet its own tolerance, and evaluates semantic-defect reduction subject to a 20% manual-review ceiling.","claim_falsifier":"A pre-existing paper, patent, product manual, or documented library implementation showing crosswalk SVD or equivalent low-rank modes retained through held-out class-specific reconstruction thresholds—especially with rare-class semantic-defect and review-cost results—would falsify the remaining contrastive claim.","problem_support":"MODERATE","recommendation":"RESEARCH","world_novelty_boundary":"Across this bounded eight-query ordinary-web search, established practice covers lossy or context-sensitive crosswalks, record-level inspection, category-stratified migration validation, explicit error limits, automated exception reports, semantic-conformance layers, and review-cost accounting; the unlocated boundary is their combination with SVD-based crosswalk reduction and rare-class residual gates, so this is not a claim of world novelty."}