{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp11_mechanism_context_external20_20260804","cell_id":"layer_decay_and_expiration_management__history_historiography","judge_id":"J2","item_assessments":[{"opaque_id":"layer_decay_and_expiration_management__history_historiography__A","supported_problem":4,"external_distinctiveness":4,"testability":5,"researchability":5,"evidence_quality":5,"fatal_issue":null},{"opaque_id":"layer_decay_and_expiration_management__history_historiography__C","supported_problem":3,"external_distinctiveness":4,"testability":4,"researchability":4,"evidence_quality":4,"fatal_issue":null},{"opaque_id":"layer_decay_and_expiration_management__history_historiography__B","supported_problem":4,"external_distinctiveness":3,"testability":5,"researchability":5,"evidence_quality":5,"fatal_issue":null}],"pairwise_comparisons":[{"pair_id":"A_C","left_id":"layer_decay_and_expiration_management__history_historiography__A","right_id":"layer_decay_and_expiration_management__history_historiography__C","preference":"LEFT","confidence":"MODERATE","rationale":"Both retain a credible incremental claim, but A has stronger direct evidence for version ambiguity in digital historical scholarship and a cleaner, tightly bounded shadow pilot with objective retrieval, dependency, and restoration outcomes. C is ethically important and somewhat distinctive, yet its central equal-visibility and mistaken-current problem is less directly evidenced, while distributional classification and community-governance requirements make the initial causal test harder to interpret."},{"pair_id":"A_B","left_id":"layer_decay_and_expiration_management__history_historiography__A","right_id":"layer_decay_and_expiration_management__history_historiography__B","preference":"LEFT","confidence":"MODERATE","rationale":"B offers the sharper quantitative experiment, but Wikidata statement ranks and SEP's revision, retirement, and archival model closely reproduce both its problem and causal lever. A's individual controls are established, yet their integration around heterogeneous historical-project derivatives, dependency-gated discovery demotion, quarantine, and sampled restoration leaves a more meaningful externally contrastive question. A also avoids B's costly claim-level annotation as an initial requirement."},{"pair_id":"C_B","left_id":"layer_decay_and_expiration_management__history_historiography__C","right_id":"layer_decay_and_expiration_management__history_historiography__B","preference":"RIGHT","confidence":"MODERATE","rationale":"C is less closely anticipated as a complete package, but its strongest outcome claim lacks direct empirical support and its broad notion of interpretive overlays creates difficult relevance and disparity judgments. B is closer to established practice, yet clearly isolates the remaining increment—demotion versus label-only treatment—behind a baseline prevalence gate, quantitative efficacy threshold, workload measurement, representation safeguards, and reversible rollback. That makes B the more worthwhile immediate research candidate of this pair."}],"overall_top_choice":"layer_decay_and_expiration_management__history_historiography__A","overall_rationale":"A has the best balance of external distinctiveness, supported problem, operational feasibility, and safety-conscious falsifiability. The evidence does not establish novelty of its components, but it does leave a credible integrated claim in a specifically bounded historical-project corpus. Its pilot can measure wrong-version selection and retrieval time while independently testing dependencies and recoverability, without deletion or public indexing changes. B is more experimentally precise but substantially closer to deployed same-lever systems; C is distinctive and normatively valuable but begins from weaker direct outcome evidence and a more confounded governance-heavy test.","blinding_limitations":"The assessment used only the supplied preserved proposals and external-evaluation records. Source claims and search completeness could not be independently verified, and differences in search success, source selection, proposal breadth, and evaluator framing may affect apparent distinctiveness and evidence quality."}