{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"layer_decay_and_expiration_management__cognitive_science:P2:v0","cell_id":"layer_decay_and_expiration_management__cognitive_science","search_queries":["adaptive working memory training task difficulty performance longitudinal model nonstationary older observations recency primary study","cognitive training adaptive difficulty algorithms working memory longitudinal performance primary research","temporal knowledge tracing forgetting old observations current ability model primary paper","cognitive assessment adaptive testing drift task version scoring measurement invariance longitudinal official guidance","site:dl.acm.org knowledge tracing forgetting temporal decay student model historical observations paper","Bayesian knowledge tracing forgetting model temporal decay primary research paper","adaptive cognitive training user model recency weighting performance history algorithm","online change point detection cognitive performance longitudinal adaptive task model paper","site:nist.gov AI RMF data provenance monitoring human oversight lifecycle official","site:eur-lex.europa.eu GDPR Article 5 storage limitation official","site:learn.microsoft.com event sourcing pattern append-only reconstruct state official","site:aws.amazon.com s3 lifecycle pricing glacier deep archive official pricing 2026","\"Does working memory training have to be adaptive?\" 2015 publication date Memory Cognition","\"Adapting training in real time\" \"1897451\" publication date","\"Time-dependant Bayesian knowledge tracing\" publication date","\"Rethinking and Improving Student Learning and Forgetting Processes\" AAAI 2025 authors","\"Adapting training in real time: An empirical test of adaptive difficulty schedules\" PMC","\"Adapting training in real time\" 38536336 authors publication date journal","site:pmc.ncbi.nlm.nih.gov PMC10013456 adaptive difficulty schedules"],"sources":[{"source_id":"s1","title":"Does Working Memory Training Have to Be Adaptive?","publisher":"Psychological Research / Springer Nature","url":"https://eprints.bournemouth.ac.uk/29211/1/vonBastian.Eschen.2016.PsychRes.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2016-03","accessed_at":"2026-08-03","claims_supported":["A randomized study of 130 young adults compared adaptive, randomized, self-selected, and active-control working-memory training.","Adaptive difficulty did not outperform alternative difficulty-variation procedures on training or transfer, showing that added adaptation machinery requires direct comparative validation.","Working-memory training and its transfer effects remain empirically contested."]},{"source_id":"s2","title":"Adapting Training in Real Time: An Empirical Test of Adaptive Difficulty Schedules","publisher":"Military Psychology / Society for Military Psychology","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC10013456/","source_class":"PRIMARY_RESEARCH","publication_date":"2021-04-14","accessed_at":"2026-08-03","claims_supported":["Naval Air Warfare Center researchers tested adaptive difficulty schedules because controlled evidence about effective adaptive-training features was lacking.","Waiting an entire scenario to adapt could leave difficulty misaligned with ability; within-scenario adaptation improved some performance outcomes.","The Office of Naval Research funded the study, identifying a real funder and applied training stakeholder interested in adaptive-training best practices."]},{"source_id":"s3","title":"Rethinking and Improving Student Learning and Forgetting Processes for Attention Based Knowledge Tracing Models","publisher":"Association for the Advancement of Artificial Intelligence","url":"https://ojs.aaai.org/index.php/AAAI/article/view/34998","source_class":"PRIMARY_RESEARCH","publication_date":"2025-04-11","accessed_at":"2026-08-03","claims_supported":["Knowledge-tracing systems predict future performance from expanding historical interaction sequences.","Existing uniform time-decay and fixed-window approaches can conflate forgetting with relevance or fail to represent continuous forgetting.","Relative-forgetting attention is a close predictive-model prior art that already modulates historical evidence by time and relevance."]},{"source_id":"s4","title":"Time-Dependant Bayesian Knowledge Tracing—Robots That Model User Skills Over Time","publisher":"Frontiers in Robotics and AI","url":"https://www.frontiersin.org/journals/robotics-and-ai/articles/10.3389/frobt.2023.1249241/full","source_class":"PRIMARY_RESEARCH","publication_date":"2024-02-26","accessed_at":"2026-08-03","claims_supported":["Time-dependent Bayesian knowledge tracing continuously updates noisy user-skill estimates and uses them to select teaching actions.","The model was evaluated in simulation and a participant study, demonstrating technical feasibility for time-aware user-state estimation.","Temporal user models are a strong comparator that may solve current-state estimation without separate evidence leases."]},{"source_id":"s5","title":"Event Sourcing Pattern","publisher":"Microsoft Azure Architecture Center","url":"https://learn.microsoft.com/en-us/azure/architecture/patterns/event-sourcing","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-03-28","accessed_at":"2026-08-03","claims_supported":["Append-only event stores can preserve full history while current read models are maintained as separate projections.","Historical states can be reconstructed by replay, with snapshots available to bound rehydration cost.","Event sourcing provides auditability but adds schema-evolution, testing, operational, and personal-data deletion complexity.","Separating immutable history from an inference-active projection is already an established software architecture pattern."]},{"source_id":"s6","title":"AI Risk Management Framework Core","publisher":"National Institute of Standards and Technology","url":"https://airc.nist.gov/airmf-resources/airmf/5-sec-core/","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2023-01-26","accessed_at":"2026-08-03","claims_supported":["NIST calls for continuous lifecycle risk management, inventory, monitoring, periodic review, documentation, testing, and clearly differentiated human oversight roles.","Executive leadership and governing authorities are responsible for deployment-risk decisions.","The framework supports the proposed governance workflow but is voluntary and does not specifically request evidence leases."]},{"source_id":"s7","title":"Regulation (EU) 2016/679, General Data Protection Regulation","publisher":"European Union / EUR-Lex","url":"https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2016-04-27","accessed_at":"2026-08-03","claims_supported":["Article 5 establishes storage limitation for identifiable personal data.","Longer retention for scientific or historical research is conditional on appropriate safeguards.","Inferential expiry cannot itself supply legal authority to retain, archive, quarantine, or destroy participant-linked evidence."]},{"source_id":"s8","title":"Amazon S3 Pricing","publisher":"Amazon Web Services","url":"https://aws.amazon.com/s3/pricing/","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"undated; current pricing page","accessed_at":"2026-08-03","claims_supported":["Commercial object storage offers differentiated standard, infrequent-access, archive, and deep-archive tiers.","Lifecycle transitions, object monitoring, minimum storage durations, metadata overhead, and archive restoration can incur separate charges.","Storage tiering is technically available, but engineering, governance, review, and validation labor—not raw storage—is likely to dominate this proposal's cost."]}],"problem_evidence":{"support":"MODERATE","rationale":"Primary studies confirm that adaptive systems use historical performance to choose difficulty or teaching actions and that delayed or poorly specified adaptation can misalign difficulty with current ability. Recent knowledge-tracing research explicitly reports problems with ever-growing interaction histories, uniform decay, and fixed windows. However, no source measures how often individually stale trial layers distort longitudinal working-memory recommendations, and one working-memory RCT found no advantage from adaptivity over simpler difficulty variation. The exact prevalence and consequence asserted by the proposal therefore remain unverified.","source_ids":["s1","s2","s3","s4"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Naval Air Warfare Center researchers, with Office of Naval Research funding, explicitly described the need for controlled evidence and best practices for adaptive-training schedules. NIST identifies organizational leaders and oversight teams as credible authorizers for lifecycle risk controls. Neither source documents a commitment to adopt per-evidence inference leases, so pull is for improved adaptive-training evidence and governance rather than for this exact intervention.","source_ids":["s2","s6"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Relative-forgetting attention in LefoKT","similarity":"Weights historical learner interactions using time- and relevance-sensitive forgetting in expanding sequences, directly addressing stale or overgeneralized historical influence.","remaining_difference":"It is a predictive model, not a governed per-layer lifecycle with explicit active, archived, quarantined, held, and deleted states or reconstruction drills.","source_ids":["s3"]},{"name":"Time-Dependant Bayesian Knowledge Tracing","similarity":"Continuously updates a user's skill state from noisy temporal observations and uses that state to choose teaching actions.","remaining_difference":"It models temporal state directly and averages recent observations; it does not expose evidence-layer leases, disposition authority, legal holds, or archive restoration.","source_ids":["s4"]},{"name":"Event sourcing with materialized projections","similarity":"Separates immutable historical events from a current-state projection and supports audit, replay, versioning, and reconstruction.","remaining_difference":"It is domain-general data architecture and does not specify cognitive-evidence relevance, semantic staleness, or predictive lease-renewal rules.","source_ids":["s5"]}],"distinctive_claim_remaining":"On a governed multi-session cognitive-training dataset, explicit context-sensitive inference leases can produce a smaller active evidence projection that predicts held-out current-session performance or next-block difficulty at least as well as a cumulative estimator and better than a fixed-recency rival, while preserving sampled historical recommendation reconstruction and without unacceptable recommendation instability, false-staleness rates, or steward review burden.","confidence":"MODERATE"},"implementation_evidence":{"support":"STRONG","rationale":"Time-aware user-state estimation, append-only histories, replayable projections, lifecycle oversight, storage limitation, and commercial storage tiers are all technically demonstrated or officially documented. An offline shadow registry is therefore implementable. Important conditional gaps remain: reliable context/version metadata, complete dependency discovery, consent and jurisdiction-specific retention authority, unbiased lease rules, restore fidelity, and acceptable human-review load have not been demonstrated for the proposed dataset.","source_ids":["s3","s4","s5","s6","s7","s8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Better alignment between current ability and task difficulty could improve training efficiency and interpretability, but the frequency and magnitude of stale-evidence errors are not measured, and working-memory adaptivity itself has mixed evidence.","source_ids":["s1","s2"]},"stakeholder_pull":{"score":3,"rationale":"A military training research unit and ONR have funded research on responsive adaptive schedules, and governance frameworks identify responsible authorizers; demand for the exact lease manager is not documented.","source_ids":["s2","s6"]},"incremental_advantage":{"score":2,"rationale":"Strong temporal models already address forgetting and changing skill states, while event sourcing already separates historical records from current projections. Added advantage must come from governance, auditability, and safer disposition rather than clearly superior inference.","source_ids":["s3","s4","s5"]},"distinctiveness_plausibility":{"score":3,"rationale":"The combined evidence-layer lease, authority-gated disposition, and reconstruction test was not found as one cognitive-training practice, although its predictive and data-architecture components have close established analogues.","source_ids":["s3","s4","s5"]},"technical_implementability":{"score":4,"rationale":"The proposal can be prototyped as an offline projection over versioned events using established temporal modeling and replay architectures. Metadata and dependency completeness are the principal technical uncertainties.","source_ids":["s4","s5","s8"]},"adoption_authority_feasibility":{"score":3,"rationale":"Model owners, cognitive scientists, data stewards, and organizational leadership are identifiable decision-makers, but participant permissions, institutional policy, and applicable law can constrain retention and destruction.","source_ids":["s6","s7"]},"evidence_readiness":{"score":3,"rationale":"The proposed replay has clear comparators and measurable outputs, but requires access to a governed multi-session dataset with version/context changes plus expert adjudication of staleness and dependencies.","source_ids":["s1","s2","s3","s4"]},"safety_net_benefit":{"score":4,"rationale":"Shadow-only execution, immutable source preservation, quarantine, tombstones, and reconstruction tests substantially bound first-step harm. Residual privacy risk remains because archived and quarantined participant data still exist.","source_ids":["s5","s6","s7"]},"scalability":{"score":3,"rationale":"Projection and tiered-storage patterns scale technically, but semantic review, dependency tracing, small-object overhead, and restore testing may scale poorly without automation and sampling.","source_ids":["s5","s8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"A sandbox replay on one completed dataset: reproduce the baseline, implement provisional leases, add cumulative and fixed-recency comparators, adjudicate a sample of stale classifications, and test historical reconstruction.","confidence":"MODERATE","assumptions":["One governed dataset is already accessible and documented.","One data scientist or ML engineer works for roughly two to four months with part-time cognitive-science, data-steward, and software support.","No participant recruitment, live adaptation, source deletion, or production integration is included.","Compute and storage are modest relative to labor."],"source_ids":["s1","s2","s3","s4","s8"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Production-grade evidence registry, version/context metadata, active and archive projections, dependency inventory, access controls, audit logs, quarantine, tombstones, and restore tooling for one platform.","confidence":"LOW","assumptions":["The existing platform exposes stable trial identifiers and model-version bindings.","Two to five engineering, data, governance, and validation staff contribute over six to twelve months.","Migration is limited to one task family and does not require rebuilding the core estimator.","Security and privacy review are required because the records are participant-linked."],"source_ids":["s5","s6","s7","s8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Controlled rollout for one task family, including historical backfill, parallel validation, lease-policy approval, operator training, incident and rollback procedures, and acceptance testing.","confidence":"LOW","assumptions":["Launch remains nonclinical and does not trigger medical-device validation.","Historical metadata require partial manual remediation.","The live system runs in shadow before any recommendation uses leased activation.","Independent privacy, fairness, and reconstruction reviews are included."],"source_ids":["s5","s6","s7"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Registry operation, monitoring, policy revalidation, exception review, archive storage and restoration drills, incident response, and periodic predictive benchmarking for one task family.","confidence":"LOW","assumptions":["A fraction of an engineer, model owner, cognitive scientist, and data steward is required throughout the year.","Review queues are sampled or risk-prioritized rather than manually inspecting every trial layer.","Data volume is moderate; labor dominates tiered-storage charges.","Major task redesigns and new jurisdictions are excluded."],"source_ids":["s5","s6","s7","s8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Primary evidence shows that adaptation timing can leave training difficulty misaligned and that expanding learner histories challenge uniform decay and fixed-window models, although domain-specific prevalence remains unknown.","source_ids":["s2","s3","s4"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"The Naval Air Warfare Center and ONR are identifiable applied-training stakeholders, while NIST describes organizational leadership and oversight roles with authority over AI lifecycle risk. No adoption commitment to this specific design was found.","source_ids":["s2","s6"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim specifies a lifecycle-managed estimator, cumulative and fixed-recency comparators, held-out performance, recommendation stability, false-staleness, review burden, and reconstruction outcomes.","source_ids":["s1","s3","s4","s5"]},"bounded_next_evidence_step":{"status":"YES","reason":"A single-dataset offline replay can be conducted without live adaptation or source deletion and has explicit comparators, outcome measures, halt conditions, and falsifiers.","source_ids":["s1","s2","s3","s4"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step is shadow-only and preserves source data. It remains permissible only after confirming dataset authorization, participant permissions, retention rules, and role-based access; no inference lease may authorize destruction.","source_ids":["s6","s7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"All four bands have bounded scopes and explicit staffing, duration, platform, and data assumptions. Official architecture and storage sources support the main complexity and cost drivers, though confidence is low beyond the first evidence step.","source_ids":["s5","s8"]}},"next_evidence_step":"With the data steward's written authorization, select one completed multi-session working-memory dataset containing at least two documented task or scoring contexts. In an isolated sandbox, reproduce the existing cumulative estimator and freeze its outputs; implement provisional per-layer inference states without changing or deleting source records. Replay three preregistered comparators: (A) cumulative history, (B) a fixed recent-trial or fixed-time window, and (C) context-sensitive inference leases. If feasible, add a strong temporal-model sensitivity analysis based on time-dependent knowledge tracing or relative-forgetting attention. Evaluate held-out current-session trial prediction, next-block recommendation agreement with a task-owner rubric, calibration, recommendation transition rate, active-set size, reasons for state changes, expert-reviewed false-staleness rate, dependency-recall against seeded known dependencies, steward minutes per 1,000 layers, and exact reconstruction of stratified historical recommendations. Falsify the intervention if C fails to improve on both simple comparators on a preregistered primary metric, is inferior to a strong temporal model without compensating governance benefit, exceeds the agreed instability or review-burden ceiling, misses a seeded dependency, or fails any sampled reconstruction. Halt on identity mismatch, unauthorized data exposure, or baseline nonreproduction; discard only derived shadow artifacts.","blocking_evidence":["No externally measured prevalence or effect size shows that stale historical evidence materially changes current working-memory recommendations.","No direct comparison exists between inference leases and strong temporal, forgetting, or change-point user models.","No adopter has committed data, staff, policy authority, or deployment resources to this exact intervention.","Dataset-specific consent, retention, legal-hold, deletion, and research-archive rules have not been mapped.","Semantic-staleness precision, dependency completeness, reconstruction fidelity, recommendation stability, and steward review burden are unmeasured.","Post-pilot production costs depend on undocumented data volume, architecture, metadata quality, jurisdictions, and integration debt."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search establishes only adjacent predictive models, governance guidance, event-sourcing practice, and storage services. It does not measure world novelty, patentability, freedom to operate, market size, or realized impact, and absence of an exact match among eight sources is not evidence that none exists.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure written authorization and one governed multi-session dataset with task/scoring-version changes.","Preregister cumulative, fixed-recency, lifecycle-managed, and strong temporal-model comparators with quantitative falsifiers.","Measure whether old evidence actually changes current estimates or recommendations after conditioning on recent evidence and context.","Obtain blinded task-owner adjudication of semantic staleness and measure review burden and subgroup error rates.","Demonstrate complete seeded-dependency detection and exact reconstruction for stratified historical recommendations.","Replace broad deployment estimates with an architecture-specific staffing and data-volume cost model after the replay."],"reason":"Bounded web research verifies that the problem is plausible, the components are implementable, and close temporal-model and event-sourcing prior art exists. It cannot establish dataset-level prevalence, incremental predictive or governance advantage, false-staleness rates, dependency completeness, reconstruction success, or operational burden. Those questions require proprietary governed data, expert adjudication, and offline execution, so the next decision point is empirical rather than another bounded web search."},"proposal_index":2}