{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"layer_decay_and_expiration_management__psychology:P3:v0","cell_id":"layer_decay_and_expiration_management__psychology","search_queries":["site:intestcom.org guidelines test use norms up to date scoring version psychological assessment","site:efpa.eu test review model norms age sample psychological tests pdf","psychological assessment outdated norms clinical misclassification Flynn effect study","psychometric scoring software norm version identifier registry lifecycle superseded norms","site:pearsonassessments.com Q-global norms update scoring report exact version norms","site:pearsonassessments.com Q-global historical norms scoring software updates norms","site:pearsonclinical.com Q-global scoring norms update manual","psychometric test publisher obsolete norms lifecycle scoring software superseded","site:bls.gov Occupational Employment and Wage Statistics software developers May 2025 median wage psychologists 2025","site:bls.gov/ooh software developers quality assurance analysts testers median pay 2025 psychologists","site:apa.org ethics code obsolete tests outdated results 9.08","site:parinc.com NEO PI 3 normative update outdated norms current population","APA Ethics Code 2017 Standard 9.08 obsolete tests outdated test results official","site:apa.org/ethics/code \"9.08 Obsolete Tests\"","site:apa.org ethics code 2017 pdf 9.08 obsolete tests","Trahan Stuebing Fletcher Hiscock 2014 Flynn Effect meta-analysis DOI","Woodcock Reading Mastery Test impact normative changes 2005 DOI"],"sources":[{"source_id":"S1","title":"Guidelines for Practitioner Use of Test Revisions, Obsolete Tests, and Test Disposal","publisher":"International Test Commission","url":"https://members.intestcom.org/files/guideline_test_disposal.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2015-06-24","accessed_at":"2026-08-03","claims_supported":["Revisions range from norm updates to complete new test versions.","Choosing old versus revised tests requires case-specific professional judgment rather than an arbitrary age rule.","Changed norms can affect diagnosis, eligibility, public support, and other life-altering decisions.","Older versions can remain necessary for longitudinal, forensic, scholarly, and retrospective purposes.","Test developers, users, and publishers are relevant actors, while test security constrains disposal and archival handling."]},{"source_id":"S2","title":"ITC Guidelines on Test Use, Final Version 1.2","publisher":"International Test Commission","url":"https://www.intestcom.org/files/guideline_test_use.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2013-10-08","accessed_at":"2026-08-03","claims_supported":["Users should prevent conclusions based on outdated or population-irrelevant norms.","Reports should clearly identify the norms, scale types, and equations used.","Appropriate comparison groups should be selected for the intended context.","Users should periodically review population changes and reevaluate tests when form, content, mode, or purpose changes.","Test security, copyright, qualification, and professional responsibility constrain automated scoring workflows."]},{"source_id":"S3","title":"Model for the Review, Description and Evaluation of Psychological and Educational Tests, Version 5.0","publisher":"European Federation of Psychologists’ Associations","url":"https://www.efpa.eu/sites/default/files/2025-08/efpa_test_review_model_v2025_5.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-08","accessed_at":"2026-08-03","claims_supported":["Norms are especially sensitive to societal, educational, diagnostic, and job-content change and require periodic renorming or evidence of continued appropriateness.","Data-collection dates, reference-population representativeness, test version, mode, and regular-update provisions are expected review metadata.","The model rates norms at least 20 years old as inadequate while explicitly noting the lack of consensus on universal validity periods and requiring contextual judgment.","The model is intended for expert reviewers and can support test-review programs operated by associations, national bodies, regulators, developers, and publishers."]},{"source_id":"S4","title":"The Flynn Effect: A Meta-analysis","publisher":"Psychological Bulletin / U.S. National Library of Medicine","url":"https://pubmed.ncbi.nlm.nih.gov/24979188/","source_class":"PRIMARY_RESEARCH","publication_date":"2014-09","accessed_at":"2026-08-03","claims_supported":["Across 285 studies and 14,031 participants, different normative bases produced a mean change of 2.31 standard-score points per decade.","For modern Stanford-Binet and Wechsler comparisons, the estimated change was 2.93 IQ points per decade.","The observed temporal change supports the existence and potential consequence of norm obsolescence, but heterogeneity limits universal age-only rules."]},{"source_id":"S5","title":"The Woodcock Reading Mastery Test: Impact of Normative Changes","publisher":"Assessment / U.S. National Library of Medicine","url":"https://pubmed.ncbi.nlm.nih.gov/16123255/","source_class":"PRIMARY_RESEARCH","publication_date":"2005-09","accessed_at":"2026-08-03","claims_supported":["In 899 referred first- through third-grade students, rescoring identical test content under revised versus updated norms produced systematic differences averaging 5 to 9 standard-score points.","The result demonstrates that changing only the normative layer can materially alter standardized outputs.","The authors identify consequences for reading-disability identification and for longitudinal research that changes norms mid-study."]},{"source_id":"S6","title":"NEO Personality Inventory-3 (Normative Update)","publisher":"PAR, Inc.","url":"https://www.parinc.com/products/NEO-PI-3-NU","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["PAR collected updated norms in 2024 and states that the legacy norms no longer accurately represent the current U.S. population.","PAR removed the legacy version from new purchase while allowing existing inventory to remain usable and visibly labeling it as legacy.","The scoring platform selects normative comparison groups, retains a path for legacy records, and permits old data to be entered under the updated package, illustrating both lifecycle demand and potential lineage ambiguity."]},{"source_id":"S7","title":"Pearson Scoring Software","publisher":"Pearson Assessments","url":"https://www.pearsonassessments.com/professional-assessments/products/search/scoring.html","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["Pearson centrally administers, scores, and reports assessments through Q-global and Q-interactive.","Pearson has retired desktop scoring products, will end support in 2027, and publishes mappings from remaining legacy assistants to current digital alternatives.","Central digital scoring, product retirement states, successor mappings, and ongoing updates are implemented adjacent practices, although the page does not document immutable norm-package identifiers, dependency-gated retirement, or restoration tests."]},{"source_id":"S8","title":"Occupational Employment and Wages—May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-05-15","accessed_at":"2026-08-03","claims_supported":["May 2025 national mean annual wages were $148,100 for software developers and $111,490 for software quality-assurance analysts and testers.","The wage data provide a public labor-cost anchor for 2026 resource-equivalent estimates; benefits, overhead, psychometric expertise, licensing, hosting, and security costs must be added as assumptions."]}],"problem_evidence":{"support":"STRONG","rationale":"Official guidance explicitly requires relevant, identifiable, periodically reviewed norms and recognizes the tension between adopting revisions and retaining older versions for continuity. EFPA treats norm currency as a review dimension. A large meta-analysis documents temporal score drift, and a study of 899 children shows 5–9-point differences caused solely by the normative layer. PAR independently states that a legacy norm set became unrepresentative. What remains unmeasured is the prevalence of ambiguous norm-package use or failed reconstruction in actual organizations.","source_ids":["S1","S2","S3","S4","S5","S6"]},"stakeholder_evidence":{"support":"STRONG","rationale":"Test publishers and assessment-program owners are identifiable adopters and authorizers. PAR has already withdrawn a legacy norm version from new purchase while maintaining labeled legacy inventory, and Pearson actively retires legacy scoring software and maps users to successors. ITC and EFPA explicitly address developers, publishers, expert reviewers, professional bodies, and regulators. This establishes operational pull for lifecycle management, although no source explicitly requests the proposal's complete registry-plus-gate implementation.","source_ids":["S1","S2","S3","S6","S7"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"ITC obsolete-test revision and disposal workflow","similarity":"Directly governs transitions between revised and older tests, professional authorization, continuity exceptions, retention of older materials, and secure disposal.","remaining_difference":"It is guidance for case-by-case practice, not an executable package registry with immutable identities, context-gated prospective selection, dependency tracing, quarantine, and restore verification.","source_ids":["S1"]},{"name":"ITC test-use identification and review requirements","similarity":"Requires accurate scoring, explicit identification of norms and equations, context-appropriate norm choice, and periodic reevaluation.","remaining_difference":"It specifies user duties but does not supply centralized lifecycle states, expiring prospective authority, or enforcement across manuals, spreadsheets, scripts, and scoring platforms.","source_ids":["S2"]},{"name":"EFPA Test Review Model Version 5.0","similarity":"Captures collection dates, population and version fit, norm age, obsolescence, expert review, and publisher update provisions in a structured model.","remaining_difference":"It evaluates tests and norms rather than controlling selection for each scoring run or preserving executable historical dependencies through tested archives.","source_ids":["S3"]},{"name":"PARiConnect NEO normative-update transition","similarity":"Maintains separately labeled legacy and updated packages, uses digital norm-group selection, stops new sale of outdated norms, and preserves limited legacy access.","remaining_difference":"The public documentation does not show immutable package-level lineage on every output, expiring review leases, expert override records, inbound-dependency checks, quarantine, tombstones, or restoration drills.","source_ids":["S6"]},{"name":"Pearson Q-global and scoring-software retirement program","similarity":"Centralizes scoring, publishes retirement status and successor alternatives, and supports platform updates unavailable to legacy desktop software.","remaining_difference":"The documentation concerns product/platform lifecycle rather than explicit lifecycle governance of every norm and scoring transformation, including historical reconstruction dependencies.","source_ids":["S7"]}],"distinctive_claim_remaining":"For one authorized test family containing multiple norm and scoring packages, an identity-resolved registry that expires prospective authority and gates selection by declared context—while dependency-constraining retirement and testing restoration—will reduce context-mismatched package selection and increase exact historical-score reconstruction relative to the existing file/manual or ordinary platform presentation, without increasing valid-use blocks or expert-override burden beyond predeclared tolerances.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Digital scoring platforms, labeled legacy inventories, successor mappings, structured norm metadata, expert review, and secure disposal workflows demonstrate that the constituent technical and governance functions are feasible. A registry, role-based gate, audit log, archive, and fixed-vector restoration test use conventional software patterns. The difficult parts are organizational: resolving copies to immutable identities, discovering historical inbound dependencies, licensing restricted test content, defining context rules without implying validity, and maintaining qualified human authority. No source demonstrates the entire integrated workflow or its performance in practice.","source_ids":["S1","S2","S3","S6","S7"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Norm-layer differences can change standardized scores materially and can affect diagnostic, educational, financial-support, and research decisions; exact historical reconstruction also matters.","source_ids":["S1","S4","S5"]},"stakeholder_pull":{"score":4,"rationale":"Publishers already renorm, label legacy inventories, retire scoring products, and operate centralized scoring platforms, while professional bodies explicitly call for current and identifiable norms.","source_ids":["S1","S2","S3","S6","S7"]},"incremental_advantage":{"score":3,"rationale":"The combined gate could add enforcement and reconstruction guarantees beyond documentation and ordinary versioning, but its advantage over publisher-controlled platforms has not been measured.","source_ids":["S2","S3","S6","S7"]},"distinctiveness_plausibility":{"score":2,"rationale":"Nearly every constituent practice is established or officially recommended. Distinctiveness rests on integration at executable norm-package level, not on a new lifecycle principle, and no world-novelty search was conducted.","source_ids":["S1","S2","S3","S6","S7"]},"technical_implementability":{"score":4,"rationale":"Registry, permissions, selection rules, archival tiers, audit records, and test-vector restoration are technically conventional; identity resolution and undocumented dependencies remain material risks.","source_ids":["S6","S7","S8"]},"adoption_authority_feasibility":{"score":4,"rationale":"Test publishers, assessment-program owners, and designated psychometric reviewers possess plausible authority. Practitioner overrides and records execution can remain separated from validity decisions.","source_ids":["S1","S2","S3","S6","S7"]},"evidence_readiness":{"score":3,"rationale":"A synthetic, read-only comparator experiment is bounded and measurable, but real baseline prevalence and package/dependency data are proprietary or must be collected from a partner.","source_ids":["S1","S2","S6","S7"]},"safety_net_benefit":{"score":5,"rationale":"Keeping superseded packages out of ordinary selection while retaining authorized reconstruction paths directly addresses both stale-use harm and irreversible-deletion harm.","source_ids":["S1","S2","S6"]},"scalability":{"score":3,"rationale":"The software pattern can scale across test families, but metadata harmonization, licensing, local copies, population-specific validity judgments, and expert review do not scale automatically.","source_ids":["S2","S3","S6","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Build a read-only mock registry for one synthetic or authorized test family; prepare 24–40 scoring and reconstruction scenarios; recruit 20–30 qualified evaluators; run the randomized comparator pilot and one fixed-vector archive restore.","confidence":"MODERATE","assumptions":["Roughly 250–500 combined hours of software, psychometric, research, and coordination effort.","Synthetic packages or an existing partner license avoid purchasing norm-sample data.","Evaluator honoraria, secure hosting, and overhead fit within the band.","BLS wage anchors are converted to loaded resource equivalents rather than treated as contractor prices."],"source_ids":["S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Production-quality registry for one test family, metadata migration, immutable identifiers, role-based review states, read-only scoring-gate integration, audit logging, archive packaging, and security/licensing review.","confidence":"MODERATE","assumptions":["Approximately 0.5–1.5 software-engineer years plus part-time psychometric, records, QA, and security effort.","Existing identity, scoring, and storage infrastructure can be integrated rather than rebuilt.","No acquisition or renorming of the underlying reference sample is included."],"source_ids":["S6","S7","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Launch across several test families or one medium assessment program, including inventory reconciliation across manuals/scripts/software, user training, override workflow, historical dependency migration, monitoring, restore drills, and staged production validation.","confidence":"MODERATE","assumptions":["Two to five full-time-equivalent staff-years across engineering, QA, psychometrics, security, records, training, and program management.","Commercial test-material licensing and integration are negotiated within the band.","Existing reports lacking package identifiers require partial manual dependency mapping.","This excludes new normative-sample collection and test renorming, which could exceed the band."],"source_ids":["S1","S2","S6","S7","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Registry stewardship, periodic psychometric reviews, help desk and override review, security and platform maintenance, archive storage, dependency updates, and scheduled restoration tests for a bounded portfolio.","confidence":"MODERATE","assumptions":["Approximately 0.5–1.5 recurring FTE equivalents plus hosting, audit, and training costs.","Norm validation studies and major software migrations are separately funded projects.","Archive volume is modest relative to assessment-response storage because the retained objects are scoring packages and documentation."],"source_ids":["S3","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official guidance recognizes outdated or mismatched norms and the need to retain older versions for legitimate continuity, while research demonstrates material score changes from normative layers.","source_ids":["S1","S2","S3","S4","S5","S6"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"PAR and Pearson are identifiable publishers already managing norm and scoring-product transitions; ITC and EFPA identify publishers, expert reviewers, professional bodies, and regulators as responsible actors.","source_ids":["S1","S3","S6","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal can be compared against existing file/manual or ordinary platform presentation on mismatched selection, reconstruction, valid-use blocks, time, and override burden. Its incremental claim is the integrated enforcement-and-restoration path, not generic renorming or versioning.","source_ids":["S1","S2","S3","S6","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"One synthetic or authorized test family, a read-only mock gate, qualified evaluators, predeclared scenarios, and a fixed-vector restoration exercise bound the first test without altering live assessments or reports.","source_ids":["S1","S2","S3"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The bounded step uses synthetic or explicitly licensed packages, qualified evaluators, read-only operation, expert rather than automated validity authority, and no client data or production scoring. Test-security and licensing review remains mandatory before real deployment.","source_ids":["S1","S2","S3"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four bands specify included scope and exclusions and are anchored to current public software and QA wage data, with explicit assumptions for psychometric, security, licensing, and overhead resources.","source_ids":["S8"]}},"next_evidence_step":"Partner with one test publisher, university assessment program, or other authorized owner and use one synthetic or licensed historical test family containing at least three norm/scoring packages. Build a read-only mock registry and gate. Randomly assign 20–30 qualified evaluators to the current file/manual or ordinary-platform presentation versus the registry presentation across 24–40 predeclared scenarios varying edition, population, assessment date, prospective versus longitudinal purpose, supersession state, and reconstruction dependency. Primary comparator outcomes are context-mismatched package selections and exact reconstruction of fixed historical scores; secondary outcomes are completion time, valid-use blocks, overrides, identity errors, and preservation decisions. Also restore one copied archived package end to end and score a predeclared test vector. Advance only if the registry produces at least a 10-percentage-point absolute reduction in mismatched selections, at least a 15-point increase in exact reconstruction, no more than a 5-point increase in valid-use blocks, no more than 10% unnecessary expert overrides, zero package-identity errors, and exact fixed-vector restoration. Falsify or redesign the intervention if either primary outcome fails to improve directionally, any safety/noninferiority tolerance is exceeded, evaluators mistake lifecycle status for validity, or the archived transformation cannot be restored exactly.","blocking_evidence":["No organization-level estimate was found for how often current scoring uses an unintended norm package or how often historical scores cannot be reconstructed.","The completeness and accuracy of inbound-dependency discovery for reports, publications, longitudinal datasets, spreadsheets, and local scripts require proprietary inventory data or field audit.","The comparative effect of the integrated registry and gate on qualified users requires a live human-subject workflow experiment.","Publisher-specific licensing, trade-secret, test-security, privacy, records-retention, and jurisdictional requirements require partner legal and security review.","No evidence yet shows that package metadata can reliably encode intended context without users mistaking lifecycle status for a guarantee of psychometric validity.","No evidence yet establishes portability of the proposed identifier and lifecycle model across independent publishers or assessment programs."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search assessed only eight direct sources and supports adjacent prior art plus a remaining integration claim. World novelty, patentability, freedom to operate, market size, and realized impact were not measured. Absence of a publicly documented identical system is not evidence that no publisher or assessment program operates one internally.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure an authorized publisher or assessment-program partner and one bounded test family.","Audit baseline package identities, selectable states, and reconstruction dependencies without changing production scoring.","Preregister the randomized comparator, outcome definitions, tolerances, and falsifiers.","Demonstrate reduced context-mismatched selection and improved exact reconstruction without excess valid-use blocks or override burden.","Complete an exact fixed-vector archive restore and document test-security, licensing, privacy, retention, and authority approval."],"reason":"Bounded web research verifies the problem, credible authorities, adjacent practices, technical plausibility, and a distinct comparator claim. The remaining decisive evidence—baseline prevalence in a real workflow, dependency-map completeness, qualified-user performance, and restoration reliability—requires proprietary data, partner authorization, and live testing rather than additional bounded web search."},"proposal_index":3}