{"actors":["Library discovery-service manager","Metadata and cataloging librarians","Reference librarians","Patrons using the discovery interface","Library privacy and accessibility reviewers"],"affected_objective":"Reliable discovery of relevant library resources across different search tasks, interfaces, languages, accessibility modes, and collection segments.","arm":"COMMON_P1","authority_safety":{"authorized_first_step":"Conduct a read-only retrospective analysis of an approved, de-identified sample of discovery-session events; report only strata meeting a predeclared minimum cell-size rule.","decision_authority":"The discovery-service manager may authorize analysis and propose configuration changes; metadata, accessibility, privacy, and collection owners retain approval authority for changes within their respective systems, and designated library governance approves any production intervention.","excluded_actions":["No patron-level scoring, identification, or outreach based on search logs.","No use of protected characteristics or inferred sensitive attributes without separate governance approval and a defined legitimate purpose.","No automatic suppression, re-ranking, metadata rewriting, or vendor configuration change during the first evidence step.","No conclusion that a flagged stratum reflects patron deficiency rather than metadata, interface, collection, or measurement conditions.","No replacement of professional review with an excursion alert."],"halt_rollback":"Halt the analysis if de-identification fails, small cells could expose patrons, event definitions cannot be applied consistently across interfaces, or a stratum cannot be interpreted without sensitive inference. Because the first step is read-only, rollback consists of deleting derived tables under the approved retention procedure and withholding operational recommendations."},"baseline":"The library reviews a rolling collection-wide discovery-success indicator, such as the share of sessions ending in a record view, availability check, request, or full-text link, and investigates individual failures mainly through complaints. A stable aggregate is treated as evidence that discovery is functioning consistently, without routinely showing the distributions beneath it.","candidate_id":"ensemble_and_population_level_equilibrium_versus_individual_level_heterogeneity__library_information_science__COMMON_P1","causal_chain":["Discovery sessions differ by search task, query formulation, interface, accessibility mode, language, and collection segment.","The headline success indicator aggregates binary session outcomes using one session as one unit within a fixed reporting window.","Opposing movements and persistent tail failures can coexist with a stable aggregate rate because the aggregation discards subgroup distributions, repeated reformulation paths, and local collection context.","Managers interpreting the stable rate as a typical patron experience do not route patterned excursions to the responsible metadata, interface, accessibility, or collection owner.","Those local conditions persist even while the library-wide indicator remains stable.","Pairing the aggregate with predeclared strata, trajectory measures, and excursion review exposes where intervention is warranted without treating ordinary individual variation as failure.","Authorized owners can then address the implicated local condition while monitoring both the aggregate indicator and the affected distribution."],"cell_id":"ensemble_and_population_level_equilibrium_versus_individual_level_heterogeneity__library_information_science","consequence":"Persistent zero-result, repeated-reformulation, or abandonment patterns can remain operationally invisible for particular search contexts, delaying access to relevant holdings while a valid library-wide success measure remains stable.","diversity_from_prior_proposals":"Not assessed against prior proposals because runtime isolation forbids inspecting them; this candidate is specifically grounded in the micro-to-macro translation from patron discovery sessions to a stable library-wide retrieval indicator.","experiment_id":"eoa_inverse_innovation_exp13_second_slot_policy60_20260806","intervention":"Adopt a two-level discovery assurance rule: every stable library-wide discovery-success report must include a privacy-preserving distributional companion showing predeclared session strata, outcome quantiles or rates, reformulation trajectories, uncertainty, and sustained local excursions. Each excursion is mapped through a micro-macro crosswalk to a review owner—metadata, interface, accessibility, or collection—while the aggregate indicator remains in place and production changes require human approval.","mechanism_mapping":[{"counterfactual_removal":"Without the joint display, reviewers can continue reading the stable headline as a description of typical sessions.","mechanism_slug":"distributional_dashboard","role":"Displays the aggregate indicator beside stratum-level outcomes and trajectory summaries under the same time window."},{"counterfactual_removal":"Without representation checks, apparent subgroup differences could reflect missing interfaces, bot traffic, or uneven event capture rather than patron experience.","mechanism_slug":"stratified_sampling_review","role":"Verifies that included sessions and event definitions cover the predeclared search contexts consistently."},{"counterfactual_removal":"Without decomposition, the team cannot distinguish persistent between-stratum structure from within-stratum or temporal variation.","mechanism_slug":"variance_decomposition_table","role":"Partitions observed variation by search context, collection segment, interface, and reporting period with uncertainty shown."},{"counterfactual_removal":"Without an explicit translation rule, local failures cannot be related coherently to the stable library-wide indicator.","mechanism_slug":"micro_macro_crosswalk","role":"Documents the session unit, success-event rule, exclusions, weighting, reporting window, and information lost during aggregation."},{"counterfactual_removal":"Without a sustained-excursion rule, either stable headline metrics suppress review or transient noise produces excessive escalation.","mechanism_slug":"subgroup_excursion_alert","role":"Routes only predeclared, sufficiently supported, persistent local departures for human review while retaining macro monitoring."}],"nearest_rivals":["Variability characterization: it would describe differences among discovery sessions but would not make the stable aggregate claim and its interpretation the organizing problem.","Aggregation bias detection and correction: it would argue that the headline indicator is invalid or improperly weighted; this candidate allows the aggregate to be valid and challenges only its extension to individual contexts.","Equilibrium restoration: it would seek to recover a deteriorated library-wide indicator; here the macro indicator may remain stable throughout.","Generic usability analytics: it may segment interface behavior but need not specify the aggregation rule, level-of-analysis boundary, or joint macro-and-distribution monitoring."],"negative_tests":{"intervention_falsifier":"Using the same retrospective data, the intervention is falsified as a useful diagnostic if predeclared, adequately observed strata and trajectories add no stable, actionable information beyond the aggregate indicator, or if alerts cannot be assigned reproducibly to an authorized review owner.","problem_falsifier":"The problem is falsified if the aggregate discovery indicator is not stable under the declared window, if it is itself invalid because event capture or weighting is biased, or if adequately powered predeclared strata show no consequential heterogeneity beneath it.","risks":["Search logs may expose intellectual interests or identities through rare queries or small cells.","Thin strata may create false precision, unstable alerts, or subgroup fishing.","Behavioral proxies such as record views or link clicks may not represent successful information seeking.","Context labels can stigmatize patrons if interpreted as user deficits.","Local optimization could improve a flagged metric while degrading relevance, privacy, accessibility, or the aggregate service.","Vendor-controlled event definitions may make cross-interface comparisons inconsistent."] ,"strongest_counterevidence":"The most damaging evidence would be that apparent excursions disappear after harmonizing event instrumentation and excluding bots, indicating measurement inconsistency rather than meaningful session-level heterogeneity."},"next_evidence_step":"For one library discovery service, pre-register the ensemble definition, success-event rule, exclusions, reporting window, a small set of operationally meaningful non-identifying strata, minimum cell sizes, and excursion criteria; then analyze eight weeks of approved de-identified logs read-only. Compare the stable aggregate series with stratum outcomes and reformulation trajectories, audit a bounded sample of flagged sessions through privacy-preserving query categories, and produce a go/no-go memo without changing production systems.","observable_state":"Across consecutive reporting windows, the library-wide share of discovery sessions reaching a defined success event remains within its existing operating band, yet de-identified session data show that some predeclared contexts repeatedly have more zero-result searches, reformulations, or exits. The current report exposes the aggregate rate but not the session distribution, aggregation rule, or clustered excursions.","prior_art_status":"UNSEARCHED","problem":"A library discovery service can have a genuinely stable collection-wide success indicator while materially different search contexts experience persistent failure patterns. Treating the stable aggregate as a description of each patron session prevents the library from distinguishing harmless variation from patterned metadata, interface, accessibility, or collection-level barriers.","proposal_index":1,"remaining_contrastive_claim":"The candidate is specifically warranted when a valid, stable discovery-wide indicator is being overextended to patron sessions or search contexts; if the indicator is invalid, the system is globally deteriorating, or no member-level consequence matters, a neighboring approach is more appropriate.","revision_record":{"claim_changes":[],"conceptual_changes":[],"evidence_changes":[],"operational_changes":[],"parent_version":null,"progress_targets_addressed":["Initial closed-book construction of one complete domain-grounded candidate.","Explicit preservation of the stable macro indicator, session-level heterogeneity, aggregation translation, level boundary, and targeted feedback.","Inclusion of bounded authority, privacy safeguards, falsifiers, risks, and a read-only first evidence step."]},"schema_version":1,"structural_mapping":[{"archetype_element":"Ensemble Frame","domain_realization":"All eligible human discovery sessions on one library service during each declared reporting window, with bots, staff testing, and sessions lacking consistent instrumentation excluded by rule."},{"archetype_element":"Macro Equilibrium Indicator","domain_realization":"The rolling library-wide proportion of eligible sessions reaching a predefined discovery-success event, interpreted as stable only within an established operating band and fixed instrumentation regime."},{"archetype_element":"Microstate Variability Profile","domain_realization":"Distributions of zero-result events, reformulation counts, time-to-success, exits, and success events across predeclared non-identifying search contexts and session trajectories."},{"archetype_element":"Aggregation Translation Rule","domain_realization":"Each eligible session contributes one binary success outcome to the headline proportion; pooling removes query-path shape, collection locality, interface context, tails, and repeated difficulty."},{"archetype_element":"Level-of-Analysis Boundary","domain_realization":"The headline supports a claim about the eligible session population during the window, not a claim that any patron, query, interface, or collection segment has the same outcome or probability."},{"archetype_element":"Heterogeneity Relevance Test","domain_realization":"Variation prompts review only when it is sufficiently observed, sustained across windows, consequential to access, and plausibly linked to an addressable metadata, interface, accessibility, or collection condition."},{"archetype_element":"Subgroup and Locality Map","domain_realization":"Excursions are localized by operational context such as interface, accessibility mode, query category, language of query where explicitly supplied, material type, and collection segment, subject to privacy thresholds."},{"archetype_element":"Multi-Level Feedback Design","domain_realization":"The library monitors the headline and distribution together and routes supported excursions to the responsible owner while checking that local correction does not degrade aggregate discovery or other strata."},{"archetype_element":"Representative Case Guardrail","domain_realization":"No single complaint, rare query, or successful session is treated as representative; human review uses multiple privacy-preserving cases sampled across the relevant distribution."}],"title":"Distribution-Aware Assurance for Stable Library Discovery Metrics","version":0}