{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp11_mechanism_context_external20_20260804","cell_id":"layer_decay_and_expiration_management__systems_cybernetics","judge_id":"J3","item_assessments":[{"opaque_id":"layer_decay_and_expiration_management__systems_cybernetics__A","supported_problem":2,"external_distinctiveness":3,"testability":4,"researchability":4,"evidence_quality":4,"fatal_issue":null},{"opaque_id":"layer_decay_and_expiration_management__systems_cybernetics__C","supported_problem":3,"external_distinctiveness":2,"testability":4,"researchability":4,"evidence_quality":4,"fatal_issue":null},{"opaque_id":"layer_decay_and_expiration_management__systems_cybernetics__B","supported_problem":2,"external_distinctiveness":2,"testability":3,"researchability":4,"evidence_quality":3,"fatal_issue":null}],"pairwise_comparisons":[{"pair_id":"A_vs_C","left_id":"layer_decay_and_expiration_management__systems_cybernetics__A","right_id":"layer_decay_and_expiration_management__systems_cybernetics__C","preference":"RIGHT","confidence":"MODERATE","rationale":"C has materially stronger external support for the underlying harm: official sources specifically document temporary control-system changes persisting and contributing to severe incidents. Its overlay-specific phenotype remains unproven and management-of-change plus Simplex are close prior art, but the offline three-arm authority-shadowing test preserves a meaningful, explicit increment. A is somewhat more externally distinctive and has an exceptionally clean prevalence falsifier, yet its controller-family problem is supported mainly by analogy from model registries, legacy systems, and configuration databases."},{"pair_id":"A_vs_B","left_id":"layer_decay_and_expiration_management__systems_cybernetics__A","right_id":"layer_decay_and_expiration_management__systems_cybernetics__B","preference":"LEFT","confidence":"MODERATE","rationale":"A offers the sharper bounded experiment: complete enrollment, identical-event replay, explicit stale-eligibility denominator, zero-baseline problem falsifier, 50% effect threshold, and zero dependency-loss condition. B addresses a credible multi-model substrate, but supervisory mismatch detection and eligibility gating are already longstanding control practices, while its rare-regime comparison, restoration fidelity, and instability proxies are less fully operationalized."},{"pair_id":"C_vs_B","left_id":"layer_decay_and_expiration_management__systems_cybernetics__C","right_id":"layer_decay_and_expiration_management__systems_cybernetics__B","preference":"LEFT","confidence":"HIGH","rationale":"C is supported by more direct, authoritative evidence that temporary control changes can persist and cause harm, and it specifies a concrete overlay, actuator contribution, matched comparators, noninferiority bound, restoration test, and halt criteria. B remains researchable, but its precise stale-artifact failure mode lacks direct incident or prevalence evidence and lies closer to established multiple-model supervisory selection."}],"overall_top_choice":"layer_decay_and_expiration_management__systems_cybernetics__C","overall_rationale":"C is the strongest overall research candidate because it combines the best-supported meaningful problem with a bounded, reversible, and genuinely falsifiable test of an overlay-specific increment. Its novelty is constrained by established management-of-change and safe authority-switching practices, so the research value is moderate rather than transformative; nevertheless, it has a firmer empirical motivation than A or B and a more consequential test than merely auditing registry eligibility.","blinding_limitations":"Assessment used only the supplied preserved proposals and their separate bounded public-web scrutiny. Search coverage, terminology, and source mixes differ across records, proprietary practices may be absent, and none of the evidence establishes prevalence, world novelty, or realized impact. No treatment identity or earlier outcome was inferred."}