{"actors":["Discovery-services librarian who owns relevance evaluation","Search engineer who configures ranking and indexing","Reference librarians who define representative information-seeking tasks","Accessibility and metadata specialists who review failure patterns","Library technology governance committee that approves production changes"],"affected_objective":"Select discovery-layer ranking changes on evidence calibrated to the full configuration and evaluation search, while protecting reliable access across query types and user tasks.","arm":"ORDINARY_DIVERSE_P2","authority_safety":{"authorized_first_step":"The discovery-services librarian may authorize an offline, read-only replay of de-identified historical queries against nonproduction ranking configurations, with no change to live results or user accounts.","decision_authority":"Only the library technology governance committee may approve a production ranking experiment or deployment after privacy, accessibility, metadata, and confirmation reviews.","excluded_actions":["No production ranking, indexing, autocomplete, or interface change from an exploratory result","No use of patron identities, session histories, or unsuppressed rare queries","No removal of unfavorable queries, metrics, segments, configurations, or windows after results are viewed","No tuning against the locked confirmation set","No claim that a metric improvement establishes accessibility, fairness, source quality, or user benefit","No automatic suppression or promotion of individual resources based on the pilot"],"halt_rollback":"Halt if attempted configurations or evaluations are missing from the record, the locked query set is accessed during tuning, privacy suppression fails, or an exploratory result enters a production approval packet as confirmed. Quarantine affected reports, revoke decision eligibility, preserve the current production ranking, and restart only with a newly isolated confirmation set and frozen protocol."},"baseline":"A discovery team iterates over ranking weights, query-processing options, metadata boosts, and result metrics on a reusable relevance set, then presents the best-performing configuration and favorable metric slice as evidence for deployment without preserving the complete search path.","candidate_id":"multiple_testing_discipline__library_information_science__ORDINARY_DIVERSE_P2","causal_chain":["A discovery-layer redesign creates many candidate ranking configurations and many evaluation opportunities across queries, task classes, metrics, cutoffs, and user segments.","Repeated reuse of the same relevance evidence allows ordinary variation and configuration-specific overfitting to generate an attractive score.","Selective reporting isolates the winning configuration and favorable metric from the alternatives that produced it.","Governance reviewers may interpret that selected score as confirmatory evidence and authorize a ranking change that does not persist across fresh information-seeking tasks.","A preregistered evaluation family and append-only run ledger expose every configuration, metric, cutoff, exclusion, and subgroup examined.","A false-discovery-rate rule treats offline winners as screened candidates rather than deployment-ready findings.","Explicit status labels keep exploratory and adjusted-discovery configurations out of production approval.","A locked query-task set and bounded prospective shadow comparison test the chosen configuration without further tuning.","Only a configuration that passes the frozen confirmation criteria and specialist review can become eligible for a separately authorized production experiment."],"cell_id":"multiple_testing_discipline__library_information_science","consequence":"A chance-favored or evaluation-set-specific ranking configuration can be deployed, changing which records users encounter while contradictory metrics and unsuccessful configurations remain absent from the decision record.","diversity_from_prior_proposals":"This opportunity concerns discovery-system ranking selection rather than collection management: it targets configuration overfitting and selective relevance reporting, uses a run ledger with discovery-rate screening and locked query-task confirmation, and gates search-system deployment rather than withdrawal, transfer, or acquisition decisions.","experiment_id":"eoa_inverse_innovation_exp13_second_slot_policy60_20260806","intervention":"Establish a ranking-evaluation protocol that freezes the discovery family before tuning: all eligible configurations, query-task classes, relevance metrics, cutoffs, language or accessibility segments, and exclusion rules. Automatically append every run and result to an immutable ledger. Apply a predeclared false-discovery-rate procedure to offline comparisons and label surviving configurations adjusted-discovery candidates, not confirmed improvements. Select at most one candidate under a frozen rule, evaluate it once on a sequestered query-task set, and conduct a non-user-facing shadow comparison reviewed by reference, accessibility, and metadata specialists. Preserve rejected, null, selected, and confirmed results; require committee authorization before any later production experiment.","mechanism_mapping":[{"counterfactual_removal":"Without the registry, unsuccessful configurations and unfavorable metric slices could disappear, restoring the illusion that the selected ranking was the only comparison.","mechanism_slug":"claim_registry","role":"Maintains the complete configuration-and-evaluation family, run history, owners, specifications, and result statuses."},{"counterfactual_removal":"Without discovery-rate control, each offline score could retain an ordinary single-comparison threshold despite the number of configurations and metric slices searched.","mechanism_slug":"false_discovery_rate_control","role":"Screens offline ranking candidates while calibrating their status to the declared family of comparisons."},{"counterfactual_removal":"Without a hierarchy, the team could replace a disappointing primary relevance measure with whichever secondary cutoff or segment favors deployment.","mechanism_slug":"metric_hierarchy","role":"Predeclares primary, secondary, accessibility, and diagnostic measures and constrains promotion based on secondary results."},{"counterfactual_removal":"Without sequestered queries and tasks, the evidence used to tune a ranking configuration would also be used to certify it.","mechanism_slug":"holdout_validation","role":"Provides a locked confirmation set that is evaluated once after candidate selection."},{"counterfactual_removal":"Without the staged follow-up, an adjusted offline discovery could be treated as sufficient evidence for a live ranking change.","mechanism_slug":"confirmatory_follow_up","role":"Requires one frozen confirmation and specialist review before the candidate can seek authorization for a separate production experiment."}],"nearest_rivals":["A reproducible relevance benchmark, which can rerun every configuration but does not calibrate the winning result to the number of configurations and metric slices searched","A representative query-sampling intervention, which addresses whether evaluation tasks reflect users but not false discoveries created by repeated configuration search","An effect-size or practical-relevance threshold, which asks whether a ranking difference matters but not how many chances existed to find a favorable difference","A search-quality monitoring dashboard, which observes production behavior but can itself enable repeated untracked looks"],"negative_tests":{"intervention_falsifier":"The intervention fails as a multiplicity guardrail if unregistered runs influence candidate selection, the chosen configuration changes after confirmation data are exposed, adjusted-discovery results appear as confirmed in approval materials, or a candidate reaches deployment review without the locked evaluation.","problem_falsifier":"The problem is absent if exactly one ranking configuration, query-task set, primary metric, cutoff, exclusion rule, and analysis are fixed before evaluation, all outcomes are retained, and no iterative tuning or alternative slicing creates additional discovery opportunities.","risks":["A broad family or conservative screening rule may discard ranking candidates worth inexpensive investigation.","Engineers may perform unlogged local experiments outside the evaluation service.","Correlated metrics and configurations may make the selected discovery-rate procedure poorly calibrated.","The locked query-task set may become indirectly familiar through repeated organizational use.","Historical queries may encode legacy interface behavior and may not represent later information-seeking tasks.","Aggregate relevance metrics may conceal accessibility, language, discipline, or metadata-specific failures.","Rare queries or linked session data may create privacy risk unless minimized and suppressed.","Specialist review may become ceremonial if its findings cannot block promotion."],"strongest_counterevidence":"If versioned evaluation records show that one configuration and one primary analysis were fixed before access to results, every alternative run was retained, the confirmation set remained isolated, and production approval depended on an unchanged fresh-data test, hidden multiplicity would not explain the selection process."},"next_evidence_step":"Run one offline shadow pilot for a single discovery-layer component using de-identified, privacy-screened historical queries. Before execution, freeze the configuration space, task taxonomy, metric hierarchy, cutoffs, exclusions, false-discovery-rate rule, candidate-selection rule, and confirmation criteria. Log every run automatically, select at most one adjusted-discovery candidate, and evaluate it once on a newly sequestered query-task set followed by a nonproduction shadow comparison. End after that confirmation review; report registration completeness, status transitions, contamination incidents, and specialist vetoes, with no live ranking change.","observable_state":"The discovery team can repeatedly tune ranking weights, stemming, field boosts, deduplication, and query expansion against a reusable relevance set while examining several metrics, cutoffs, task classes, and user segments; the deployment memo retains the selected configuration but lacks a complete machine-readable inventory of attempted runs and slices.","prior_art_status":"UNSEARCHED","problem":"When evaluating a library discovery layer, staff can try many ranking configurations and inspect each across multiple relevance metrics, result cutoffs, query classes, language or accessibility segments, and exclusion rules. The configuration with the most attractive score can then be presented as a confirmed improvement even though it was selected from a large, incompletely reported opportunity set and tuned on the same evidence used to justify it.","proposal_index":2,"remaining_contrastive_claim":"The proposal addresses evidentiary inflation caused specifically by selecting a discovery-system configuration from many ranking and evaluation alternatives; rerunnability, representative task sampling, or a minimum score difference alone would not expose the opportunity set, assign exploratory status, or require untouched confirmation before deployment eligibility.","revision_record":{"claim_changes":[],"conceptual_changes":[],"evidence_changes":[],"operational_changes":[],"parent_version":null,"progress_targets_addressed":["Instantiate multiple-testing discipline in discovery-system relevance evaluation","Create an affected problem, intervention, and causal path materially independent of collection-weeding decisions","Preserve claim-family definition, attempted-look inventory, multiplicity-aware screening, status labeling, fresh confirmation, and discovery-record retention","Bound authority through a read-only offline pilot, privacy safeguards, deployment exclusions, falsifiers, and rollback conditions"]},"schema_version":1,"structural_mapping":[{"archetype_element":"Many attempted claims","domain_realization":"Ranking configurations crossed with query-task classes, metrics, cutoffs, language or accessibility segments, time windows, and exclusion rules"},{"archetype_element":"Ordinary single-claim interpretation","domain_realization":"The best configuration and favorable score are presented without accounting for the other configurations and evaluation slices examined."},{"archetype_element":"Inflated false-positive opportunity","domain_realization":"Repeated tuning and metric inspection create multiple chances for noise or evaluation-set overfitting to resemble a ranking improvement."},{"archetype_element":"Claim family","domain_realization":"The frozen set of eligible configurations, query-task classes, metrics, cutoffs, segments, and analytic rules for one evaluation cycle"},{"archetype_element":"Multiplicity inventory and discovery record","domain_realization":"An append-only ledger containing every run, parameter specification, evaluation slice, null result, rejection, selection, and status transition"},{"archetype_element":"Error-risk policy and multiplicity-aware rule","domain_realization":"False-discovery-rate screening for offline candidate generation, with no screened candidate treated as deployment-ready"},{"archetype_element":"Exploratory-confirmatory boundary","domain_realization":"Tuning occurs only on discovery evidence; the selected candidate is tested once on a sequestered query-task set and in a nonproduction shadow comparison."},{"archetype_element":"Claim-status labeling","domain_realization":"Exploratory, adjusted-discovery, rejected, confirmation-failed, and confirmed labels appear in the ledger and approval materials."},{"archetype_element":"Confirmation before costly action","domain_realization":"A ranking candidate cannot become eligible for a separately authorized production experiment until locked confirmation and reference, accessibility, and metadata reviews are complete."}],"title":"Multiplicity-Aware Confirmation Gate for Discovery Ranking Changes","version":0}