{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"layer_decay_and_expiration_management__cognitive_science:P4:v0","cell_id":"layer_decay_and_expiration_management__cognitive_science","search_queries":["site:act-r.psy.cmu.edu ACT-R reference manual production utility conflict resolution compilation","site:soar.eecs.umich.edu Soar manual excise production firing count","computational cognitive modeling reproducibility executable models archived environments paper","cognitive architecture model lifecycle obsolete production rules provenance","site:soar.eecs.umich.edu soar manual excise productions firing count pwatch official","site:soar.eecs.umich.edu Soar User Manual excise production command","site:act-r.psy.cmu.edu ACT-R model repository reproducibility model files official","OSF computational cognitive modeling reproducibility expressed need executable code","ACT-R model repository submission requirements model code reproducibility","ACT-R community model repository maintained model code documentation","computational cognitive science model repository standards reproducibility code expressed need","Soar workshop model maintenance rules obsolete productions debugging","\"The importance of standards for sharing of computational models and data\" publication date","\"A systematic methodology for cognitive modelling\" 1996 publication date","\"Towards Replication in Computational Cognitive Modeling\" publication date"],"sources":[{"source_id":"S1","title":"ACT-R 7.30+ Reference Manual","publisher":"ACT-R Research Group, Carnegie Mellon University","url":"https://act-r.psy.cmu.edu/actr7.x/reference-manual.pdf","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["ACT-R productions are executable condition-action pairs whose matching and conflict resolution centrally control model behavior.","ACT-R exposes production inspection, debugging, tracing, add-production hooks, parameters, and deterministic replay controls that can support read-only instrumentation.","Production compilation can add productions, so executable rule inventories can change during model execution or development."]},{"source_id":"S2","title":"production — Soar Command-Line Reference","publisher":"Soar Group, University of Michigan","url":"https://soar.eecs.umich.edu/reference/cli/cmd_production/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Soar already provides production inventories, firing counts, match inspection, per-rule watching, memory-use analysis, and production excision.","Soar can excise never-fired rules, but its documentation cautions that memory use is heuristic and does not establish scientific or behavioral dispensability.","Existing architecture-level instrumentation makes a sandbox prototype technically plausible while constituting close functional prior art."]},{"source_id":"S3","title":"A Systematic Methodology for Cognitive Modelling","publisher":"Elsevier, Artificial Intelligence","url":"https://www.sciencedirect.com/science/article/pii/0004370295001123","source_class":"PRIMARY_RESEARCH","publication_date":"1996-08","accessed_at":"2026-08-03","claims_supported":["Computational cognitive-model development has historically been ad hoc and can conflate empirically justified mechanisms with pragmatic implementation details.","The Sceptic executable specification approach is prior art for exposing theoretical commitments and investigating alternative assumptions in Soar-like models."]},{"source_id":"S4","title":"Reproducibility in Computational Neuroscience Models and Simulations","publisher":"Frontiers in Neuroinformatics / PubMed Central","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC5016202/","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2016-10","accessed_at":"2026-08-03","claims_supported":["Model reproducibility requires version control, documentation, modularity, shared code, repositories, and standards.","Executable source, model parameters, results, and environmental dependencies must remain aligned for replication and reuse.","Preserving executable model history matters, but this review does not establish the prevalence of obsolete reachable production rules."]},{"source_id":"S5","title":"Towards Replication in Computational Cognitive Modeling: a Machine Learning Perspective","publisher":"Springer Nature, Computational Brain & Behavior","url":"https://link.springer.com/article/10.1007/s42113-019-00055-w","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2019-08-14","accessed_at":"2026-08-03","claims_supported":["Reproducibility challenges directly affect computational cognitive modeling.","Fine-grained logging, documentation, metadata, and error analysis can impose substantial administrative and interpretive effort.","The authors call for standardized, preferably automatically summarized metadata and recommend cautious, gradual adoption of new practices."]},{"source_id":"S6","title":"The Importance of Standards for Sharing of Computational Models and Data","publisher":"Springer Nature, Computational Brain & Behavior / PubMed Central","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC7241435/","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2019-10-07","accessed_at":"2026-08-03","claims_supported":["Researchers in cognitive science, psychology, and neuroscience expressed a need for community standards describing computational models.","A NIH BRAIN Initiative-supported group began developing a model-description framework, providing identifiable stakeholders and potential institutional support.","Computational graphs with explicit nodes, parameters, and edges are established adjacent practice for dependency-oriented model descriptions."]},{"source_id":"S7","title":"Reusability Standards","publisher":"Open Modeling Foundation","url":"https://www.openmodelingfoundation.org/standards/reusability/","source_class":"STANDARD","publication_date":"2025-04-11","accessed_at":"2026-08-03","claims_supported":["OMF standards require detailed model metadata, provenance, qualified software dependencies, and execution instructions.","Ideal practice includes archival-quality containers, automated testing, and related-research metadata.","OMF identifies framework scaffolding and automated compliance support as potential infrastructure, evidencing adopter interest in model-governance tooling but not specifically rule expiration."]},{"source_id":"S8","title":"CoMSES Net Computational Model Library","publisher":"CoMSES Net","url":"https://www.comses.net/","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["CoMSES operates an identifiable repository and peer-review workflow for preserving model code, digital artifacts, and research context.","The organization explicitly promotes model documentation, preservation, discoverability, reuse, and reproducibility.","Repository curation is an established adoption pathway for archival packages, although CoMSES is not specifically an ACT-R or Soar authority."]}],"problem_evidence":{"support":"MODERATE","rationale":"The general problem—difficulty distinguishing theoretical mechanisms from implementation detail while preserving executable reproducibility—is directly supported. ACT-R and Soar documentation confirms that productions remain executable participants in matching and conflict resolution and that users need tracing, firing-count, inspection, and excision tools. However, no searched source measured how often superseded yet reachable productions contaminate current cognitive-model predictions. The proposal's exact prevalence and impact therefore remain unverified.","source_ids":["S1","S2","S3","S4","S5","S6"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Identifiable stakeholders exist: ACT-R and Soar model owners can authorize sandbox instrumentation; the University of Michigan Soar group already maintains rule-analysis tools; an NIH BRAIN Initiative-supported group expressed a need for computational-model description standards; OMF explicitly contemplates framework scaffolding; and CoMSES operates model-preservation workflows. None explicitly requested production-rule expiration leases, theoretical-authority labels, or shadow-retirement governance, so specific adopter pull remains uncertain.","source_ids":["S2","S5","S6","S7","S8"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Soar production inspection and excision commands","similarity":"Directly inventories rules and supplies firing counts, matching traces, memory-use heuristics, per-rule watches, and excision including never-fired productions.","remaining_difference":"It does not attach hypothesis ownership, review leases, scientific preservation classes, dependency-gated shadow retirement, or restore-tested historical packages.","source_ids":["S2"]},{"name":"ACT-R procedural tracing, hooks, conflict resolution, and production compilation","similarity":"Provides the architecture-level signals and hooks needed to observe when productions are added, selected, and fired.","remaining_difference":"The reference tooling debugs execution but does not govern whether a reachable production retains current theoretical authority or manage reversible lifecycle disposition.","source_ids":["S1"]},{"name":"Sceptic systematic cognitive-model specification","similarity":"Separates theory from implementation detail and makes alternative theoretical assumptions executable and inspectable.","remaining_difference":"It is a specification methodology rather than a lifecycle system for accumulated elements in an evolving model family.","source_ids":["S3"]},{"name":"Versioned, documented, containerized model preservation","similarity":"Established standards preserve code, provenance, dependencies, execution instructions, tests, and archival environments.","remaining_difference":"These practices preserve whole releases; they do not classify and experimentally shadow-retire individual reachable cognitive commitments while retaining tested historical reconstruction.","source_ids":["S4","S5","S6","S7","S8"]}],"distinctive_claim_remaining":"Compared with source history, architecture-native tracing/excision, whole-model testing, and containerized releases, an element-level lifecycle harness that binds each reachable production or memory element to a current hypothesis/task owner and review lease, then gates one-at-a-time shadow retirement through dependency traces and current-plus-historical replay, will improve attribution of current model behavior and identify safely demotable commitments without reducing benchmark or reconstruction fidelity. This is falsifiable by owner-verified classifications, attribution coverage, review time, behavioral deltas, dependency failures, and historical restore results.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"ACT-R and Soar already expose matching, firing, hooks, counts, watchpoints, and excision primitives, while OMF standards establish provenance, dependency, CI, and container practices. A read-only inventory and trace layer is therefore technically credible. Harder unresolved work includes identity resolution across dynamically created productions and chunks, indirect dependencies, deterministic environments, rare-condition coverage, and semantic judgments about current theoretical authority. Legal feasibility depends on model licenses and repository permissions; scientific authority must remain with the model owner. The proposed sandbox, shadow-only first step, quarantine, and rollback materially limit safety risk.","source_ids":["S1","S2","S4","S5","S7","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Better attribution and reproducibility could materially improve complex model families, but the frequency and prediction-level consequences of stale reachable elements are not measured.","source_ids":["S1","S3","S4","S5"]},"stakeholder_pull":{"score":2,"rationale":"Model-standard and repository communities express adjacent needs, but no named cognitive-architecture lab has requested or committed to this specific harness.","source_ids":["S6","S7","S8"]},"incremental_advantage":{"score":3,"rationale":"The combined theoretical-authority, shadow-retirement, and historical-restore workflow goes beyond existing commands and whole-release preservation, but its advantage over disciplined modules, branches, tests, and containers requires a comparative pilot.","source_ids":["S1","S2","S3","S7"]},"distinctiveness_plausibility":{"score":3,"rationale":"No exact integrated practice appeared in the bounded search, but most components are established and world novelty was not investigated.","source_ids":["S1","S2","S3","S6","S7"]},"technical_implementability":{"score":4,"rationale":"Official architecture tooling supplies most observation and experimental-control primitives; semantic classification and dynamic dependency completeness remain difficult.","source_ids":["S1","S2","S7"]},"adoption_authority_feasibility":{"score":3,"rationale":"A theory lead or model owner can authorize a sandbox pilot without changing architecture governance, but broader adoption requires repository curators and multiple task owners to agree on labels and preservation policy.","source_ids":["S5","S6","S7","S8"]},"evidence_readiness":{"score":3,"rationale":"A bounded sandbox comparison is well specified, but it requires access to a real multi-task model, owner judgments, historical packages, and live re-execution.","source_ids":["S1","S2","S4","S7"]},"safety_net_benefit":{"score":4,"rationale":"Read-only instrumentation, one-at-a-time shadow disabling, unchanged authoritative code, quarantine, tombstones, and historical replay directly address accidental scientific loss; effectiveness still needs testing.","source_ids":["S2","S4","S7"]},"scalability":{"score":3,"rationale":"Instrumentation and metadata generation can be automated, but semantic review, rare-path validation, archive maintenance, and cross-architecture adapters may scale poorly.","source_ids":["S5","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"A six-to-ten-week sandbox study of one existing model family, two to four task configurations, 30–60 review candidates, and one historical package.","confidence":"MODERATE","assumptions":["Existing model code, benchmarks, architecture runtime, and compute are available without license fees.","Approximately 0.25–0.6 full-time-equivalent engineering/modeling effort plus 20–50 hours of model-owner review.","No authoritative code is modified and no production hosting is required."],"source_ids":["S1","S2","S5"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Production-quality adapter for one cognitive architecture, metadata schema, trace store, review UI, access controls, CI replay, quarantine, archival packaging, and documentation.","confidence":"LOW","assumptions":["One architecture and one laboratory are in scope.","Existing source control, CI, storage, and container infrastructure can be reused.","The estimate includes engineering, model-owner design time, security review, and migration of one model family but excludes redesign of the cognitive architecture."],"source_ids":["S1","S2","S7"]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"Launch across one laboratory's active model family, including inventory reconciliation, task-owner validation, historical-package curation, training, and monitored shadow-retirement cycles.","confidence":"LOW","assumptions":["Five to ten users and no more than three substantial model branches are included.","Historical dependencies are available and legally preservable.","The launch remains advisory; automated authoritative retirement is excluded."],"source_ids":["S5","S7","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Ongoing metadata curation, review cycles, archive restore drills, dependency and security updates, adapter maintenance, compute, and storage for one laboratory.","confidence":"LOW","assumptions":["Approximately 0.25–1.0 full-time-equivalent combined curator/engineer/model-owner effort.","Model and archive volume remains moderate and uses institutional compute and storage.","Major architecture migrations or reconstruction of missing historical environments are excluded."],"source_ids":["S5","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent methodological and reproducibility literature supports difficulty identifying theoretical commitments and preserving executable models, while official architecture documentation confirms that productions participate in matching and conflict resolution. Exact stale-rule prevalence remains a gap but does not negate the supported underlying problem.","source_ids":["S1","S2","S3","S4","S5","S6"]},"externally_credible_adopter_or_authorizer":{"status":"UNCERTAIN","reason":"ACT-R or Soar model owners have authority over a sandbox model, and OMF, CoMSES, and NIH-supported model-standard stakeholders express adjacent needs. No external source identifies a committed adopter for rule-level lifecycle governance.","source_ids":["S2","S6","S7","S8"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The claim is contrastive against architecture-native inspection/excision and whole-release versioning, with measurable attribution, review-time, behavioral-fidelity, dependency, and restoration outcomes.","source_ids":["S1","S2","S3","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"One model family, two to four tasks, 30–60 candidates, one historical package, three comparator conditions, and explicit stop/falsification criteria bound the proposed study.","source_ids":["S1","S2","S4","S5"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"A read-only sandbox and shadow-only intervention can be authorized by the model owner; the authoritative model remains unchanged. Halt conditions address behavioral drift, missing dependencies, licensing restrictions, and failed restoration.","source_ids":["S1","S2","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four resource-equivalent bands are scoped by model families, users, labor, infrastructure reuse, and exclusions. Confidence is low to moderate because no adopter-specific labor rates, archive size, or integration complexity are available.","source_ids":["S5","S7","S8"]}},"next_evidence_step":"With a consenting ACT-R or Soar model owner, preregister a sandbox study of one model family containing two to four current task configurations and one preserved published configuration. Inventory all executable productions, chunks, overrides, and modules; collect matching, selection, firing, and dependency traces; and select 30–60 candidates using missing ownership, explicit supersession, retired-task binding, or validation-age criteria. Compare (A) existing source history, comments, modules, and regression tests; (B) those practices plus the read-only inventory and traces; and (C) the full lifecycle metadata and one-at-a-time shadow-disable review. Counterbalance reviewers where practical. Measure classification time, inter-reviewer agreement, proportion of candidates confirmed as superseded-but-reachable, current trace and prediction deltas, unexplained dependency failures, benchmark fidelity, historical-package restore success, and curator effort. Preserve the authoritative model unchanged. Treat the problem as falsified for this model if a complete owner-verified inventory finds no superseded element participating in any current trace. Treat incremental advantage as falsified if condition C produces neither a prespecified 20% classification-time reduction nor improved agreement/attribution over A and B, or if review effort exceeds twice baseline without additional verified findings. Treat the intervention as unsafe or ineffective if shadow disabling produces unexplained reviewed-benchmark changes, misses known dependencies, or prevents exact historical reconstruction. Halt immediately on instrumentation-induced behavior change, unresolved licensing, missing publication dependencies, or inability to restore the untouched baseline.","blocking_evidence":["Measured prevalence of superseded-but-reachable executable elements in a real cognitive-model family.","Model-owner validation that lifecycle labels correspond to current theoretical and task commitments rather than mere firing frequency.","Comparative evidence that the harness improves attribution or review efficiency beyond existing modules, source control, traces, and regression tests.","Live shadow-disable evidence showing that nominated retirements preserve current benchmark behavior and rare-condition coverage.","Successful restoration and execution of a historical package under realistic dependency drift.","A named cognitive-modeling laboratory or repository willing to authorize and resource the pilot."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The bounded search found no exact integrated production-rule expiration harness, but it was not a systematic literature review, patent search, source-code census, standards census, or commercial-product survey. World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured. The defensible boundary is only that the searched sources establish extensive adjacent practice but not the complete contrastive workflow.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure a named ACT-R or Soar model owner with authority to provide one multi-task model family and historical executable package.","Pre-register candidate-selection rules, baseline comparators, outcome measures, 20% efficiency criterion, fidelity requirements, and halt conditions.","Produce an owner-verified executable-element inventory and quantify superseded-but-reachable prevalence.","Run counterbalanced baseline-versus-harness classification and attribution comparisons.","Complete one-at-a-time shadow-disable replays across current benchmarks and one historical configuration.","Demonstrate end-to-end historical restoration and document labor, compute, storage, disagreement, and failure modes.","Obtain an explicit adopter decision on whether observed benefit justifies recurring curation cost."],"reason":"Web evidence supports the general problem, adjacent stakeholder need, technical plausibility, and extensive prior art, but cannot establish model-specific prevalence, semantic classification validity, incremental workflow benefit, safe shadow retirement, reconstruction fidelity, or actual adopter commitment. Those questions require proprietary model access, expert fieldwork, and live sandbox execution, so bounded web research is no longer sufficient."},"proposal_index":4}