{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","source_assessment_id":"layer_decay_and_expiration_management__cognitive_science:P4:v0","cell_id":"layer_decay_and_expiration_management__cognitive_science","proposal_index":4,"qualification":"EMPIRICAL_PARTNER_CANDIDATE","criteria":{"specific_differentiated_claim":{"status":"YES","reason":"The proposal retains a contrastive, falsifiable claim that element-level lifecycle metadata plus dependency-gated, one-at-a-time shadow retirement and historical replay improves attribution or review efficiency beyond source history, modules, native tracing/excision, regression tests, and whole-release preservation without reducing benchmark or reconstruction fidelity."},"credible_problem_signal":{"status":"YES","reason":"Independent methodological and reproducibility evidence supports difficulty separating theoretical commitments from implementation detail and preserving executable models, while architecture documentation confirms that reachable productions participate in behavior; the unresolved issue is model-specific prevalence and impact, not whether the underlying problem is credible."},"identifiable_partner_or_adopter":{"status":"YES","reason":"An ACT-R or Soar laboratory maintaining a multi-task model family is a concrete partner class: its theory lead or model owner can authorize sandbox access, validate semantic bindings, and decide adoption, while a repository curator can supply and govern the historical executable package."},"partner_access_is_necessary":{"status":"YES","reason":"Public research cannot establish superseded-but-reachable prevalence, validate current hypothesis and task bindings, compare reviewer performance, exercise proprietary or unpublished model dependencies, replay rare paths, or obtain an adopter decision. These require a real model family, historical package, owner judgments, and live sandbox execution."},"safe_authorized_first_step":{"status":"YES","reason":"The first step is a consenting-owner-authorized, read-only sandbox inventory and trace study with one-at-a-time shadow disabling. It leaves the authoritative model and published packages unchanged and includes explicit halt and rollback conditions for behavior changes, unresolved identities, missing dependencies, licensing problems, or failed restoration."},"bounded_decisive_empirical_design":{"status":"YES","reason":"The study is bounded to one model family, two to four current tasks, 30–60 candidates, one historical configuration, and three comparator conditions. Prespecified prevalence, 20% efficiency, agreement and attribution, fidelity, effort, dependency, restoration, safety, and falsification thresholds can decisively reject the problem or incremental claim for the selected model."},"no_material_negative_gate":{"status":"YES","reason":"No verified gate is materially NO: the problem, incremental claim, bounded test, safety and authority, and cost scope pass. Adopter commitment remains uncertain, but this is precisely a partner-resolved uncertainty rather than an authority stop or evidence of collision with established practice."},"not_merely_more_research":{"status":"YES","reason":"The next step specifies the partner, required model and archive access, comparator conditions, candidate-selection rules, measurements, quantitative decision thresholds, falsifiers, and halt conditions; it is a concrete field experiment rather than open-ended research."}},"uncertainty_types":["PROBLEM_PREVALENCE","ADOPTER_PULL","INCREMENTAL_EFFECT","WORKFLOW_FIT","DATA_ACCESS","COST_SCOPE"],"partner_profile":"A consenting ACT-R or Soar laboratory with an actively maintained model family spanning two to four task configurations, at least one preserved published executable configuration, a theory lead or model owner able to validate current commitments and authorize sandbox instrumentation, task owners available for review, and a curator able to provide reconstruction assets and make an adoption decision.","required_access":"Read-only access to the model source, executable productions and memory elements, task configurations, regression benchmarks, matching and firing traces, dependency information, preserved published package and runtime dependencies, plus model-owner time for binding validation, candidate classification, comparator review, and the final adoption decision. Shadow disabling must be permitted only in an isolated sandbox.","bounded_empirical_test":"Preregister a six-to-ten-week sandbox study of one model family. Inventory executable elements and collect traces; select 30–60 candidates using missing ownership, explicit supersession, retired-task binding, or validation age; compare existing practices alone, practices plus inventory and traces, and the full lifecycle workflow. Counterbalance reviewers where practical, shadow-disable candidates one at a time, replay all current benchmarks and one historical configuration, restore the untouched historical package, and record classification time, agreement, attribution coverage, confirmed stale prevalence, behavioral deltas, dependency failures, fidelity, restoration success, and labor.","success_condition":"At least one superseded-but-reachable element is owner-verified in current execution, and the full lifecycle condition achieves the prespecified 20% classification-time reduction or a prespecified improvement in reviewer agreement or causal attribution over both comparators, without unexplained benchmark changes, missed known dependencies, loss of historical reconstruction fidelity, or review effort exceeding twice baseline; the authorized adopter judges the observed benefit worth continued use.","falsification_condition":"Falsify the problem for the model if an owner-verified complete inventory finds no superseded element participating in any current trace. Falsify incremental advantage if the full workflow yields neither the 20% time reduction nor improved agreement or attribution over the comparators, or costs more than twice baseline without additional verified findings. Treat the intervention as unsafe or ineffective if instrumentation changes behavior, shadow disabling causes unexplained reviewed-benchmark changes, known dependencies are missed, or exact historical reconstruction fails.","rationale":"This is a narrow empirical-partner case: adjacent tooling establishes substantial component-level prior art, but the integrated claim remains differentiated and testable. The decisive uncertainties concern real-model prevalence, semantic validity, incremental benefit, workflow burden, reconstruction fidelity, and adopter willingness. They cannot be resolved through ordinary public research and require authorized access to a partner's executable model, expert judgments, and historical package under a reversible comparator-based pilot."}