{"schema_version":1,"research_id":"eoa_inverse_innovation_exp04_external_evaluation_20260802","source_assessment_id":"layer_decay_and_expiration_management__education_pedagogy:DISCARDED_AUDIT:H2:v0","cell_id":"layer_decay_and_expiration_management__education_pedagogy","search_queries":["knowledge tracing forgetting time decay learner model prior performance recency primary research","student model time decay mastery learning adaptive sequencing forgetting knowledge tracing","Recent-Performance Factors Analysis recency weighting student model paper","Does time matter modeling effect of time in Bayesian knowledge tracing","DAS3H modeling student learning and forgetting distributed practice","Attentive Knowledge Tracing exponential decay paper","Logistic Knowledge Tracing recent learning memory paper","half-life regression student memory personalized practice primary paper","mastery threshold over-practice Bayesian knowledge tracing","deep knowledge tracing comparative decay functions"],"sources":[{"source_id":"S1","title":"Optimal and Worst-Case Performance of Mastery Learning Assessment with Bayesian Knowledge Tracing","publisher":"International Educational Data Mining Society","url":"https://files.eric.ed.gov/fulltext/ED558215.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2013","accessed_at":"2026-08-02","claims_supported":["False-negative mastery judgments push learners too slowly and consume instructional time on already-mastered knowledge components.","False-positive judgments create the countervailing risk of premature advancement.","Mastery-decision evaluation must measure both over-practice and advancement errors."]},{"source_id":"S2","title":"Does Time Matter? Modeling the Effect of Time with Bayesian Knowledge Tracing","publisher":"International Educational Data Mining Society","url":"https://educationaldatamining.org/EDM2011/wp-content/uploads/proc/edm11_proceedings.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2011","accessed_at":"2026-08-02","claims_supported":["Conventional knowledge tracing systematically overpredicted performance after a gap of at least one day.","A knowledge-tracing variant incorporating forgetting significantly improved overall prediction.","Elapsed time can make an accumulated learner-state estimate stale."]},{"source_id":"S3","title":"Move Your Lamp Post: Recent Data Reflects Learner Knowledge Better than Older Data","publisher":"Journal of Educational Data Mining","url":"https://eric.ed.gov/?id=EJ1115280","source_class":"PRIMARY_RESEARCH","publication_date":"2015","accessed_at":"2026-08-02","claims_supported":["Recent-Performance Factors Analysis explicitly gives newer learner observations greater authority than older observations.","An exponentially decayed proportion of successes had the best predictive accuracy among evaluated recency representations.","The recency-weighted model outperformed existing logistic-regression learner models."]},{"source_id":"S4","title":"Logistic Knowledge Tracing: A Constrained Framework for Learner Modeling","publisher":"IEEE Transactions on Learning Technologies","url":"https://digitalcommons.memphis.edu/facpubs/8158/","source_class":"PRIMARY_RESEARCH","publication_date":"2021","accessed_at":"2026-08-02","claims_supported":["Adaptive-learning systems use learner models to make pedagogical decisions.","Across six datasets, modeling recent learning was generally important, with memory terms especially important for fact learning.","No single learner-model specification was best in every dataset, implying that decay must be context-calibrated."]},{"source_id":"S5","title":"Context-Aware Attentive Knowledge Tracing","publisher":"Association for Computing Machinery","url":"https://doi.org/10.1145/3394486.3403282","source_class":"PRIMARY_RESEARCH","publication_date":"2020-08-23","accessed_at":"2026-08-02","claims_supported":["Attentive Knowledge Tracing downweights distant historical responses with a monotonic exponential-decay mechanism.","The model improved future-response AUC by up to six percentage points on benchmark datasets.","Decay-weighted learner histories were proposed for personalized feedback and recommendations."]},{"source_id":"S6","title":"DAS3H: Modeling Student Learning and Forgetting for Optimally Scheduling Distributed Practice of Skills","publisher":"EC-TEL / arXiv","url":"https://arxiv.org/abs/1905.06873","source_class":"PRIMARY_RESEARCH","publication_date":"2019-05-14","accessed_at":"2026-08-02","claims_supported":["DAS3H represents past successes and attempts in temporal windows and permits skill-specific learning and forgetting curves.","Temporal features consistently improved predictive AUC, and DAS3H outperformed comparison models on three educational datasets.","The learner-state estimates were designed to support adaptive skill-practice scheduling."]},{"source_id":"S7","title":"A Trainable Spaced Repetition Model for Language Learning","publisher":"Association for Computational Linguistics","url":"https://aclanthology.org/P16-1174/","source_class":"PRIMARY_RESEARCH","publication_date":"2016-08-07","accessed_at":"2026-08-02","claims_supported":["Half-life regression operationalizes time-dependent memory decay from learner histories.","On Duolingo data it reduced recall-prediction error by more than 45 percent against several baselines.","A deployment study reported a 12 percent increase in daily engagement, demonstrating platform authority and operational feasibility for decay-informed sequencing."]},{"source_id":"S8","title":"Is There a Better Way to Forget? Modelling Memory Decay in Deep Knowledge Tracing","publisher":"Knowledge-Based Systems (Elsevier)","url":"https://doi.org/10.1016/j.knosys.2025.114884","source_class":"PRIMARY_RESEARCH","publication_date":"2026-01-15","accessed_at":"2026-08-02","claims_supported":["Multiple decay functions have already been implemented and compared in deep knowledge-tracing models.","The widely used Ebbinghaus curve underperformed sigmoid and inverse decay functions in some scenarios.","Removing or changing forgetting components affects both prediction performance and computation, so a decay curve cannot safely be assumed universal."]}],"problem_evidence":{"support":"STRONG","rationale":"Mastery systems have a documented false-negative failure mode that produces unnecessary practice and consumes time on already-mastered skills. Independent learner-model studies also show that prediction depends materially on recency and elapsed time. The exact prevalence of remediation assignments contradicted by a latest validated performance is not quantified, but the underlying decision error is directly supported.","source_ids":["S1","S2","S3","S4"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"The affected learners bear the documented costs of over-practice and premature advancement, while operational research shows that a learning platform can deploy decay-informed sequencing with measurable engagement effects. Evidence demonstrates consequential use and an adopter, but not explicit learner demand for this particular evidence-expiration rule.","source_ids":["S1","S6","S7"]},"prior_art":{"proximity":"ESTABLISHED_PRACTICE","closest_analogues":[{"name":"Recent-Performance Factors Analysis","similarity":"Directly assigns exponentially declining weight to older learner-performance observations so recent evidence better represents current mastery.","remaining_difference":"It evaluates predictive accuracy rather than the proposed joint endpoint of contradicted remediation assignments and non-inferior premature advancement; it does not specify protected longitudinal evidence.","source_ids":["S3"]},{"name":"Logistic Knowledge Tracing with recency and memory features","similarity":"Provides an implemented framework for combining recent-performance and memory-decay features in learner models used for pedagogical decisions.","remaining_difference":"It is a general model-selection framework rather than a fixed expiration policy tied to the stated sequencing outcome.","source_ids":["S4"]},{"name":"Context-Aware Attentive Knowledge Tracing","similarity":"Uses exponential decay to reduce the authority of distant assessment interactions when predicting current performance.","remaining_difference":"Decay is an attention weight rather than explicit observation expiry, renewal, or longitudinal-review protection.","source_ids":["S5"]},{"name":"DAS3H","similarity":"Separates historical attempts into temporal windows, models skill-specific forgetting, and targets adaptive scheduling.","remaining_difference":"It emphasizes distributed-practice scheduling and memory retention rather than specifically suppressing remediation caused by stale low scores.","source_ids":["S6"]},{"name":"Half-Life Regression","similarity":"Uses fitted decay curves and renewed practice evidence to drive deployed adaptive review decisions.","remaining_difference":"It focuses on recall decay and review timing, often lowering confidence as old positive evidence ages, rather than specifically retiring obsolete negative mastery evidence.","source_ids":["S7"]}],"distinctive_claim_remaining":"Only the narrowly framed safety-constrained evaluation remains distinctive: demonstrate fewer next-activity remediation assignments contradicted by a learner's latest validated performance while showing no increase in premature advancement, with explicit expiry and a protected longitudinal-review copy. The causal lever itself—age-decaying learner evidence—is not distinctive.","confidence":"HIGH"},"implementation_evidence":{"support":"STRONG","rationale":"Recency-weighted logistic models, temporal-window models, exponential-decay attention, half-life regression, and several deep-model decay functions have all been implemented and evaluated on real learner data. The evidence supports technical feasibility and warns that decay shape and skill context require calibration. It does not establish the proposed pedagogical outcome or validate hard expiration.","source_ids":["S2","S3","S4","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Avoiding unnecessary remediation can preserve instructional time, but the size and prevalence of the exact stale-negative-evidence problem are not quantified.","source_ids":["S1","S3"]},"stakeholder_pull":{"score":3,"rationale":"Adaptive platforms demonstrably use learner models for sequencing and have deployed decay-informed scheduling, but direct learner or institutional demand for this exact rule was not found.","source_ids":["S4","S7"]},"incremental_advantage":{"score":2,"rationale":"Recency and temporal features often improve prediction, but the hypothesis supplies no evidence that its simple decay-and-expire rule beats established time-aware alternatives on the stated joint sequencing endpoint.","source_ids":["S3","S4","S5","S6","S8"]},"distinctiveness_plausibility":{"score":1,"rationale":"The causal mechanism substantially overlaps longstanding recency-weighted, forgetting-aware, and temporal-window learner models.","source_ids":["S2","S3","S4","S5","S6","S7"]},"technical_implementability":{"score":5,"rationale":"Timestamps, decay functions, temporal features, and learner-state updates are routine and have multiple validated implementations.","source_ids":["S3","S4","S5","S6","S7","S8"]},"adoption_authority_feasibility":{"score":4,"rationale":"An adaptive-platform operator controls the learner model and next-activity selector, and Duolingo provides operational precedent. Institutional governance and instructor override requirements reduce certainty.","source_ids":["S4","S7"]},"evidence_readiness":{"score":4,"rationale":"Existing interaction logs permit retrospective replay and shadow-policy evaluation, and mature baseline implementations exist. A randomized test of actual remediation and premature-advancement outcomes remains necessary.","source_ids":["S1","S3","S4","S6"]},"safety_net_benefit":{"score":3,"rationale":"The intervention could prevent over-practice, but poorly calibrated decay can also erase useful difficulty evidence and increase premature advancement. Longitudinal retention and conservative shadow testing are necessary safeguards.","source_ids":["S1","S4","S8"]},"scalability":{"score":5,"rationale":"Per-observation timestamping and decay are computationally lightweight, while even richer decay-aware models have been evaluated on large platform datasets.","source_ids":["S4","S5","S6","S7"]}},"score_confidence":"HIGH","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Retrospective replay on one course: construct equal-weight, R-PFA/LKT, and candidate decay baselines; label contradicted-remediation and premature-advancement outcomes; run sensitivity and subgroup analyses.","confidence":"MODERATE","assumptions":["Existing timestamped item-level response and assignment logs are available.","One data scientist and one learning-science analyst work for approximately 8 to 12 weeks.","No learner-facing system change is required."],"source_ids":["S1","S3","S4"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Integrate timestamped evidence authority, calibrated decay parameters, longitudinal-review protection, audit logging, instructor override, and shadow-mode assignment comparison into one adaptive course platform.","confidence":"LOW","assumptions":["The platform already has a functioning learner model and next-activity service.","Work includes engineering, data science, privacy review, quality assurance, and monitoring instrumentation.","No new assessment-content production is required."],"source_ids":["S4","S5","S7","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Run a preregistered, learner-level controlled pilot across several course sections with non-inferiority monitoring for premature advancement and rollback authority.","confidence":"LOW","assumptions":["Approximately 1,000 to 5,000 learners are available.","The trial uses existing instructional content and platform infrastructure.","Independent analysis and instructor support are included."],"source_ids":["S1","S7"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Monitor sequencing errors and subgroup performance, recalibrate decay by skill and context, review protected evidence, maintain audit logs, and respond to instructor appeals.","confidence":"LOW","assumptions":["Deployment remains limited to one platform and a modest course portfolio.","One fractional data-science/learning-science team maintains the system.","Infrastructure cost is minor relative to personnel and governance."],"source_ids":["S4","S6","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"False-negative mastery decisions are documented to cause over-practice and wasted instructional time, and recency studies show older observations can be less representative of current learner knowledge.","source_ids":["S1","S3"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Adaptive-learning platform operators control learner-model updates and activity sequencing; Duolingo has operationally tested a decay-informed scheduling model.","source_ids":["S4","S7"]},"distinct_testable_incremental_claim":{"status":"NO","reason":"The claim is testable, but its causal lever is not distinct: R-PFA already tested exponentially decayed learner observations against non-recency baselines, while LKT, AKT, and DAS3H provide additional close implementations. Only the exact downstream error-pair endpoint remains unevaluated.","source_ids":["S3","S4","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A bounded retrospective replay and shadow-policy study can estimate both contradicted remediation and premature advancement before changing any learner's assignments.","source_ids":["S1","S3","S4"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"No categorical authority or safety stop was identified. Premature advancement is a material risk, but it can be bounded through offline testing, non-inferiority limits, instructor override, protected history, and rollback.","source_ids":["S1","S8"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The next evidence step and deployment phases have bounded staffing and integration scopes, although dollar ranges are bottom-up estimates rather than externally priced benchmarks.","source_ids":["S3","S4","S7"]}},"next_evidence_step":"Using one timestamp-rich mastery course, preregister a retrospective replay comparing equal-weight evidence, R-PFA/LKT recency baselines, and the hypothesis's fixed decay/expiry rule. Define a remediation as contradicted only when a later independently validated performance clears the mastery criterion; measure the paired change in contradicted-remediation assignments and impose a non-inferiority bound on premature advancement. Proceed to shadow mode only if the candidate beats the established recency baselines, not merely equal weighting.","blocking_evidence":["No direct study was found showing that expiring stale low-performance observations reduces actual remediation assignments without increasing premature advancement.","No evidence establishes that the proposed fixed decay/expiry rule outperforms established R-PFA, LKT, AKT, or temporal-window alternatives.","Decay effects vary by skill, dataset, and functional form, so a universal aging curve is unsupported.","The prevalence and instructional-time burden of remediation specifically caused by stale low observations have not been measured."],"research_disposition":"KNOWN_PRACTICE_DIFFUSION","world_novelty_boundary":"No world-novelty claim is supportable for age-discounting learner evidence: exponentially recency-weighted performance models were directly evaluated by 2015, and forgetting-aware knowledge tracing and deployed decay-based scheduling are mature prior art. At most, novelty could lie in the exact safety-constrained outcome protocol, explicit hard expiry, and separation of active instructional authority from a protected longitudinal record; those elements are not yet evidenced as an effective differentiated intervention.","arm":"DISCARDED_AUDIT","candidate_version":0,"controller_recommendation":{"action":"STOP_DIFFUSION","repairable":false,"material_progress_observed":false,"progress_targets":[],"reason":"The hypothesis does not satisfy the preregistered strict differentiated-opportunity endpoint as written. The problem is credible, an adopter exists, implementation is straightforward, and a bounded study is possible, but the central age-decay lever is established learner-modeling practice and therefore fails the distinct-incremental-claim gate. The appropriate disposition is diffusion or comparative evaluation of known practice, not differentiated-opportunity advancement."}}