{"closest_prior_art":[{"name":"Brown University Library Blacklight search-relevancy regression tests","overlap":"Uses predetermined patron-like queries against Solr to systematically verify expected retrieval and detect regressions from indexing-schema or term-weighting changes; the library reports using the tests to identify and resolve retrieval deficiencies.","remaining_difference":"Does not describe a frozen bulk bibliographic transformation, explicit result/facet tolerance envelope, uncertainty, record-level transformation provenance, a cataloging-authority publication gate, or version-linked forecast-versus-realized calibration.","source_ids":["SRC1"]},{"name":"Apache Solr Schema Designer temporary collection and Query Tester","overlap":"Creates temporary schema and collection resources, indexes sample documents, previews how unpublished schema changes affect matching, sorting, faceting, and highlighting, reports changes before publication, and notes that published-schema changes may require full reindexing.","remaining_difference":"Uses limited sample data and interactive queries rather than a bounded full-record shadow migration with a preregistered use-case suite, decision thresholds, uncertainty, accountable acceptance/revision choices, and post-publication backtesting.","source_ids":["SRC2"]},{"name":"OpenSearch Search Relevance Workbench pairwise query-set comparison","overlap":"Runs fixed query sets against two search configurations, exposes top results, and quantifies result-set and rank similarity with Jaccard and rank-biased-overlap metrics.","remaining_difference":"Does not tie the configurations to bibliographic crosswalk provenance, test collection leakage or facet tolerances, bind failures to cataloging authority, or calibrate forecasts against later production outcomes; the documented feature is experimental and not recommended for production use.","source_ids":["SRC3"]},{"name":"Established offline information-retrieval evaluation practice","overlap":"Treats relevance as contextual, quantifies precision and recall across collections and queries, and recommends sample-query planning plus in-house, TREC, or A/B testing when tuning search systems.","remaining_difference":"Does not supply the proposal's metadata-transformation-specific shadow index, preset migration gate, provenance, uncertainty boundary, or forecast-error trace.","source_ids":["SRC4"]}],"contrastive_claim_falsifier":"On the preregistered offline migrations, schema validation plus stratified manual sampling matches or exceeds the shadow gate's accept/revise/delay decision accuracy and out-of-tolerance defect prevention at lower review cost, or shadow-recommended corrections increase tolerance violations; either outcome falsifies the claimed advantage for this setting.","contrastive_claim_remaining":"Beyond existing normalization previews, schema sandboxes, and query-regression comparisons, the remaining testable claim is that applying a proposed bulk bibliographic transformation to a frozen corpus, replaying retrieval-critical queries against current and shadow indexes, and binding preset tolerance and uncertainty breaches to an authorized prepublication decision will improve migration decisions over schema validation plus stratified sampling; version-linked realized outcomes will additionally reveal when that preview should be recalibrated or demoted.","experiment_id":"eoa_inverse_innovation_exp13_second_slot_policy60_20260806","gates":{"adequate_source_search":{"rationale":"The bounded search covered the proposal directly, older and neighboring terminology, library products and practices, and component combinations including schema sandboxes, regression queries, pairwise relevance comparison, facets, and migration acceptance testing. Four opened direct sources from three publishers were retained; phrase misses were not treated as novelty evidence.","source_ids":["SRC1","SRC2","SRC3","SRC4"],"status":"PASS"},"bounded_next_test":{"rationale":"The proposed comparison is limited to one frozen, non-sensitive record set, 30–50 preregistered tasks, seeded and natural defects, blinded review, and no production change. SRC1 demonstrates executable library query-regression tests, while SRC2 and SRC3 demonstrate temporary indexing and query-set comparison capabilities.","source_ids":["SRC1","SRC2","SRC3"],"status":"PASS"},"distinct_testable_claim":{"rationale":"Adjacent sources establish most technical components, but not the combined claim that a corpus-specific metadata-transformation forecast with preset tolerances and an accountable publication gate improves decisions over schema validation plus sampling. That incremental claim has explicit comparative outcomes and rejection criteria.","source_ids":["SRC1","SRC2","SRC3"],"status":"PASS"},"no_obvious_safety_or_authority_stop":{"rationale":"The first step is offline, frozen, bounded, de-identified, and non-production; publication authority remains human and rollback discards only staged resources. SRC2 documents temporary collections and explicit configuration permissions, and SRC3 itself confines its experimental comparison feature away from production use.","source_ids":["SRC2","SRC3"],"status":"PASS"},"supported_problem":{"rationale":"Brown reports unexpected retrieval regressions from indexing-schema changes and actual deficiencies found and resolved through query tests. Solr documentation shows schema changes can alter matching, sorting, and facets and that post-publication changes may require full reindexing. This supports the mechanism, though the retained evidence does not directly measure the frequency or cost of bulk library crosswalk failures.","source_ids":["SRC1","SRC2","SRC4"],"status":"PASS"}},"prior_art_disposition":"ADJACENT_PRIOR_ART","problem_evidence":{"finding":"Retrieval behavior can regress when indexing fields, schemas, or weighting change, including failures to return expected records and changes to matching, ordering, and facets; published schemas can be costly to change without reindexing. Evidence is direct for library discovery and search-schema changes but only indirect for bulk bibliographic crosswalk migrations and downstream propagation costs.","source_ids":["SRC1","SRC2","SRC4"],"status":"PARTLY_SUPPORTED"},"research_id":"eoa_inverse_innovation_exp13_light_screen_20260806","schema_version":1,"screen_id":"E13P140","screen_survival":true,"search_lanes":{"component_combination":{"no_result_note":null,"queries":["library discovery regression testing query set compare search results Solr relevance","library catalog migration test environment query testing facets before go live","catalog discovery shadow index compare queries metadata mapping"],"source_ids":["SRC1","SRC2","SRC3","SRC4"]},"direct_problem_and_intervention":{"no_result_note":null,"queries":["library metadata migration shadow index query replay retrieval testing bulk crosswalk","library discovery metadata migration preflight search results facets testing","MARC crosswalk migration discovery regression test query replay"],"source_ids":["SRC1","SRC2"]},"products_practices_and_standards":{"no_result_note":null,"queries":["site:knowledge.exlibrisgroup.com Primo test normalization rules preview records search facets","VuFind search regression testing relevance test queries","Blacklight relevance testing query regression library discovery","library catalog migration acceptance testing known item topical searches metadata"],"source_ids":["SRC1","SRC2","SRC4"]},"synonyms_and_historical_terms":{"no_result_note":null,"queries":["bibliographic data conversion acceptance testing catalog search results","OPAC migration parallel index pre-production search testing metadata","metadata crosswalk quality assurance retrieval performance library catalog"],"source_ids":["SRC1","SRC2","SRC4"]}},"sources":[{"claims_supported":["Predetermined phrases can mimic user searches and provide systematically analyzable retrieval results.","Library search tests can detect regressions caused by Solr indexing-schema or term-weighting changes.","Brown used such tests to identify and resolve title-weighting and identifier-search deficiencies."],"publisher":"Brown University Library Digital Technologies","source_id":"SRC1","source_type":"OTHER","title":"Search relevancy tests","url":"https://library.brown.edu/create/digitaltechnologies/search-relevancy-tests/"},{"claims_supported":["A temporary collection can index sample documents while schema changes are tested before publication.","The Query Tester exposes effects on matching, sorting, faceting, and highlighting.","Published schema changes may require full reindexing, and configuration access can be permission-controlled."],"publisher":"Apache Software Foundation","source_id":"SRC2","source_type":"OFFICIAL_GUIDANCE","title":"Schema Designer — Apache Solr Reference Guide","url":"https://solr.apache.org/guide/solr/latest/indexing-guide/schema-designer.html"},{"claims_supported":["Pairwise experiments compare two search configurations using a query set.","Returned documents and result-list similarity can be evaluated across queries.","Jaccard and rank-biased-overlap metrics quantify set and rank differences; the feature is experimental and not recommended for production."],"publisher":"OpenSearch Project","source_id":"SRC3","source_type":"OFFICIAL_GUIDANCE","title":"Comparing query sets — OpenSearch Documentation","url":"https://docs.opensearch.org/3.1/search-plugins/search-relevance/compare-query-sets/"},{"claims_supported":["Search relevance is contextual and precision and recall can quantify retrieval effectiveness across collections and queries.","Faceting and query configuration affect retrieval behavior.","Sample-query planning, in-house testing, TREC testing, and A/B testing are established evaluation approaches."],"publisher":"Apache Software Foundation","source_id":"SRC4","source_type":"OFFICIAL_GUIDANCE","title":"Relevance — Apache Solr Reference Guide","url":"https://solr.apache.org/guide/solr/latest/getting-started/relevance.html"}],"world_novelty_boundary":"This bounded public-web screen found adjacent technical practices but no retained source describing the exact library-specific combination of frozen bulk transformation, retrieval tolerance envelope, binding cataloging-authority gate, uncertainty/provenance reporting, and forecast-error backtesting. That absence cannot establish world novelty, patentability, market size, expert acceptance, or realized value."}