Literature source inventory¶
Part of Inverse Innovation with the Encyclopedia of Abstractions · Literature source inventory · Last revised August 2026
Review date: 2026-08-03
Coverage: Primary and near-primary sources inspected for the integrated
literature review. This is a curated evidence inventory, not an exhaustive
bibliography. More complete field-specific bibliographies are linked at the end.
Evidence labels¶
- Peer reviewed — journal or archival conference publication.
- Preprint — public manuscript not treated here as independently validated.
- System/project — official artifact or repository used to understand an implementation; claims are limited to what the artifact exposes.
- Review/framework — useful for taxonomy and citation chaining, but central empirical claims were checked against primary studies when possible.
Historical search-space precedents¶
- Zwicky (1967), “The Morphological Approach to Discovery, Invention, Research and Construction.” Primary methodological chapter. DOI. Defines the morphological field as a systematic way to enumerate configurations across declared parameters. It is the clearest historical precedent for treating a Cartesian product as an explicit discovery space.
- Ritchey (2015), “Principles of Cross-Consistency Assessment in General Morphological Modelling.” Methodological development. Paper. Explains how cross-consistency assessment removes incompatible pairs from a raw morphological field. It clarifies both the similarity to the EoA matrix and the difference between enumeration and proposal-specific scrutiny.
- Altshuller Foundation, TRIZ historical materials. Primary institutional archive. English index. Documents the extraction of generalized inventive principles from selected patent evidence. The companion review uses AutoTRIZ for a contemporary implemented comparison and these materials for the historical method boundary.
1. Cognitive foundations and computational analogy¶
-
Gentner (1983), “Structure-Mapping.” Peer-reviewed theory. Paper. Defines analogy as alignment of relational systems, with structural consistency and systematicity. It motivates EoA role-and-relation representations but does not solve retrieval, adaptation, or problem discovery.
-
Gick and Holyoak (1980), “Analogical Problem Solving.” Peer-reviewed controlled experiments. Paper. Distant source stories can enable target solutions, but spontaneous retrieval is unreliable and improves sharply with a hint. This is a foundational reason to externalize the EoA's retrieval structure.
-
Gick and Holyoak (1983), “Schema Induction and Analogical Transfer.” Peer-reviewed controlled experiments. Paper. Comparing multiple analogs helps people induce portable schemas, and schema quality predicts transfer. It supports using several varied instantiations rather than a single polished exemplar for each archetype.
-
Gentner, Rattermann, and Forbus (1993), retrievability versus inferential soundness. Peer-reviewed controlled experiments. Paper. Surface similarity strongly affects retrieval while relational similarity more strongly affects judged inference quality. It directly motivates separate high-recall generation and structurally demanding scrutiny stages.
-
Falkenhainer, Forbus, and Gentner (1989), SME. Peer-reviewed implemented model. Paper. The Structure-Mapping Engine builds consistent correspondences and candidate inferences from represented source and target descriptions. Its boundary is equally important: it assumes both representations already exist.
-
Forbus, Gentner, and Law (1995), MAC/FAC. Peer-reviewed implemented model. Paper. A cheap content filter reduces a large case memory before expensive structural matching. It is the classical analogue of the EoA's cost–recall question, but it starts from a supplied target/query.
-
Hummel and Holyoak (1997), LISA. Peer-reviewed implemented cognitive model. Paper. Integrates access, mapping, inference, and schema induction through dynamic role binding. Relevant to mechanistic modeling, but evaluated in represented tasks rather than open-domain innovation.
-
Doumas, Hummel, and Sandhofer (2008), DORA. Peer-reviewed implemented model. Paper. Shows how explicit predicate-like relational representations can be learned from initially nonrelational inputs. Later work extends the account to controlled cross-domain generalization; this is not evidence of autonomous problem discovery from literature.
-
Aamodt and Plaza (1994), case-based reasoning. Review/framework grounded in implemented systems. Paper. The retrieve–reuse–revise–retain cycle provides a mature precedent for repair and learning from outcomes, ordinarily within a supplied problem and domain-specific case base.
-
Veloso and Carbonell (1993), derivational analogy in PRODIGY. Peer-reviewed implemented planner. Paper. Stores annotated problem-solving traces, justifications, abandoned alternatives, and failures, and repairs a replay when assumptions fail. It is a strong precedent for preserving EoA trajectories rather than only final proposals.
-
Kittur et al. (2019), “Scaling Up Analogical Innovation with Crowds and AI.” Peer-reviewed human–AI research program. Paper. Decomposes target-first analogical innovation into abstraction, search, and application and tests scalable hybrid methods. It is a mandatory comparator for retrieval and human–AI division of labor, not a solution-archetype-first matrix.
-
Christensen and Schunn (2007), analogy in engineering design. Peer-reviewed observational study. Paper. Observes analogies serving problem-identification, solution, and explanation functions in real design work. It makes analogical problem finding cognitively plausible, while the observed problem-identification analogies were mainly within-domain.
2. Design by analogy, bio-inspired design, and opportunity discovery¶
-
Gero (1990), design prototypes. Peer-reviewed framework. Paper. Separates function, expected and actual behavior, structure, context, and design knowledge. It warns against treating shared function as proof of a shared mechanism.
-
Stone and Wood (2000), Functional Basis. Peer-reviewed engineering representation. Paper. Standardizes function–flow descriptions so artifacts can be searched independently of physical form. It is a close representational ancestor of reusable solution descriptions, but less causally rich than many EoA mechanisms.
-
Linsey, Markman, and Wood (2012), WordTree. Peer-reviewed design study. Paper. Re-represents problem language to find more distant analogies. It is target-problem-first and human-led.
-
Vattam et al. (2011), DANE. Peer-reviewed implemented design environment. Paper. Uses structure–behavior–function causal models for function-indexed retrieval and understanding of biological systems. Its explanatory depth comes with expensive human knowledge engineering.
-
Siddharth and Chakrabarti (2018), Idea-Inspire 4.0. Peer-reviewed controlled evaluation. Paper. Uses SAPPhIRE's multilevel causal representation to support analogical transfer. It is among the nearest mechanism-explicit tools but remains substantially human- and problem-led.
-
Helms and Goel (2012), analogical problem evolution. Peer-reviewed book chapter/study. Record. Biological analogs can change the engineering problem formulation itself. This contradicts a claim that problem formulation from source solutions is absent from design research.
-
A system-of-systems solution-driven BID process (2019). Peer-reviewed design process. Paper. Explicitly proceeds from a biological solution through principle extraction and reframing to problem search, problem definition, and application. It is decisive prior art for the direction of work.
-
AskNatureNet (Chen et al., 2024). Peer-reviewed implemented knowledge network. Paper. Represents 1,797 biological sources, 3,101 functions, and 643 applications in a 5,541-node, 5,082-edge network and supports bidirectional source–function–application retrieval. It retrieves rather than independently scrutinizes complete opportunity proposals.
-
AskNatureGPT (Chen et al., 2025/2026). Peer-reviewed generative system. Paper. Fine-tunes GPT-3.5 on 305 successful bio-inspired cases; from a biological source, benefit, and model it identifies an engineering application/problem and generates a concept. Two learned relevance classifiers, held-out cases, ablations, human ratings, and design cases make it the strongest directional precedent. It lacks a domain-general matrix, output-specific prior-art search, independent adversarial criticism, substantive negative trajectories, cost/adopter scrutiny, and built-prototype evidence.
-
Murphy et al. (2014), function-based patent analogy. Peer-reviewed retrieval system. Paper. Uses functional vectors to find literal through far-field patent analogies for a supplied need. It demonstrates scalable abstraction-assisted retrieval, not target problem generation.
-
Yoon et al. (2015), function-based technology-opportunity discovery. Peer-reviewed system. Paper. Builds a knowledge base from 223,603 patents and defines four routes from existing technologies/products toward new technologies, products, or applications. This is direct large-scale prior art for technology-first opportunity search.
-
Luo, Yan, and Wood (2017), InnoGPS. Peer-reviewed system. Paper. Maps more than five million patent records into a technology space for neighborhood exploration, path finding, and near/far-domain opportunity search. It is a navigator and inspiration system, not an evidence-preserving proposal-adjudication pipeline.
-
Qiao, Zhang, and Chen (2025), technology–function link prediction. Peer-reviewed system. Paper. Extracts technology–function pairs using subject–action–object parsing, identifies explicit applications, and predicts implicit application links. It strengthens the case that systematic application discovery already exists in patent analytics.
-
Kang et al. (2022), scientific analogical search. Peer-reviewed system and user studies. Paper. Extracts purpose and mechanism representations from roughly 1.7 million scientific papers and helps researchers retrieve distant analogies. It is the strongest corpus-scale comparator for structural retrieval and downstream human adaptation, but it starts from a target problem.
-
Jiang and Luo, AutoTRIZ (2024/2025). Conference paper/preprint with artifacts. Paper. Uses LLM modules and a fixed TRIZ knowledge base to convert a user problem into interpretable generalized principles and multiple solutions. It is an important abstraction-mediated baseline with the opposite task direction.
-
Reinsberger et al. (2026), deliberate exploratory search. Peer-reviewed longitudinal field research. Paper. Four organizational technology–market-linking projects involved 306 interviews, 89 proposals, and 18 search agents; needs and solutions co-evolved through search and validation. It is the strongest real-world comparator for adopter evidence, but not a computational matrix.
-
ViMimic (2025), analogy-to-innovation design system. Peer-reviewed human-in-the-loop system. Author manuscript. Uses LLMs for structured analogy retrieval, archetypal solutions, problem finding, and cross-domain semantic/morphological mapping. An 18-participant evaluation reported gains in novelty, functionality, mapping efficiency, and diversity. Its product-design setting and interactive workflow differ from a fixed, fully recorded archetype × domain experiment.
-
Divago (Pereira and Cardoso, 2006). Peer-reviewed computational creativity system. Paper. Maps represented domains and generates conceptual blends under constraints. It establishes structured cross-domain concept generation but not unmet- problem discovery, prior-art research, or deployment governance.
3. LLM analogy, robustness, and causal controls¶
-
SCAN (Czinczoll et al., 2022). Peer-reviewed benchmark. Paper. Tests systematic mapping in complex cross-domain analogies and finds low performance in the evaluated models. It is recognition/mapping evidence, not innovation.
-
SCAR (2023). Peer-reviewed benchmark. Paper. Uses 400 scientific analogies across 13 fields and requires abducting the common relation, separating semantic association from structural explanation.
-
Webb, Holyoak, and Lu (2023), emergent analogical reasoning. Peer-reviewed controlled tasks. Paper. Reports strong LLM performance on several newly constructed text analogy tasks. It proves meaningful behavioral competence, not open-corpus retrieval or a settled human-like mechanism.
-
Lewis and Mitchell (2024), counterfactual analogy tasks. Peer-reviewed conference paper/preprint. Paper. Finds large performance drops when familiar symbolic conventions are changed while abstract rules are preserved. The follow-up debate shows that tools and task representation materially change the evaluated system.
-
Webb, Holyoak, and Lu (2025), counterfactual reply. Peer-reviewed. Paper. Code-augmented GPT-4 recovers on some counterfactual formal tasks. This is valuable contrary evidence: brittleness is not equivalent to an absolute inability to remap unfamiliar symbols.
-
ARN (2024). Peer-reviewed TACL benchmark. Paper. Separates near, far, and disanalogous narratives. GPT-4 fell below random in the reported zero-shot far-analogy condition; few-shot prompting closed only part of the human gap.
-
AnaloBench (2024). Peer-reviewed benchmark. Paper. Uses 340 human-written analogies and tests long scenarios and retrieval from large contexts. Model scale yielded limited improvement in the difficult retrieval settings.
-
Qin et al. (2025), “Relevant or Random.” Peer-reviewed causal control study. Paper. Random self-generated examples can rival or exceed relevant examples on some tasks, and exemplar correctness is a major driver. It shows why a prompt labeled “analogical” is not sufficient evidence that analogy caused a gain.
-
Stevenson et al. (2026), children, adults, and LLMs. Peer-reviewed TACL study. Paper. Humans generalized from Latin to Greek and novel symbol systems substantially more stably than the tested LLMs, which degraded with distance. It is strong evidence of a remaining far-transfer robustness gap in compact formal tasks.
-
Musker et al. (2024), LLMs as models of analogical reasoning. Preprint. Paper. Finds that advanced models reproduce several human semantic-structure effects but are not identical cognitive models. It counters an overly negative reading of benchmark failures.
-
Strategy Science human–LLM comparison (2026). Peer-reviewed controlled comparison. Paper. Across 199 people and eight LLMs, the reported pattern is lower-recall, higher-precision human retrieval versus higher-recall, lower-precision model retrieval, including coherent-looking but causally wrong matches. This is a direct argument for broad generation followed by independent causal scrutiny.
4. Scientific discovery systems and synthetic transfer training¶
-
SciMON (2023). Peer-reviewed/preprint system. Paper. Generates research ideas from a problem/background and retrieved literature, then iteratively compares candidates with related work for novelty. It establishes prior-art-aware ideation and refinement without an analogy-specific ontology.
-
Scideator (2024). Peer-reviewed/preprint human–LLM system. Paper. Recombines purpose, mechanism, and evaluation facets, retrieves distant material, and provides a novelty checker. It is a strong comparator for structured scientific ideation and human interaction, but it remains problem/research-area initiated.
-
Si et al. (2024), human evaluation of LLM research ideas. Preprint. Paper. In a study involving more than 100 NLP researchers, LLM ideas were judged more novel but slightly less feasible; model self-evaluation and idea diversity were important weaknesses. This warns against using generator ratings as opportunity evidence.
-
AI co-scientist (2025/2026). Preprint and peer-reviewed Nature report. Preprint. Uses multi-agent generation, critique, tournament ranking, and iterative evolution with literature access, and reports experimental validation in selected biomedical cases. It proves that critique-and-revision discovery pipelines are already real; it does not isolate analogical transfer or solution-first problem finding.
-
Shen, Druckmann, and Zou (2026), cross-domain scientific analogy. Preprint with code/data. Paper. For a supplied biomedical problem, explicitly extracts objects and relations, generates distant problem analogies, searches the analogous domain for a solution, maps it back, checks literature novelty, and implements one proposal for each of four problems. Reported diversity, novelty, and implementation results make it the nearest contemporary cross-domain transfer system. It is target-first, biomedical, sampled rather than a declared matrix, and identifies feasibility and end-to-end grounding as remaining bottlenecks.
-
MOOSE-Chem and MOOSE-Chem2 (2024–2025). Peer-reviewed systems. MOOSE-Chem and MOOSE-Chem2. Reconstruct or search for chemistry hypotheses from backgrounds and inspirations under temporal controls. They are strong unseen-paper reconstruction and structured-search precedents, but domain-specific and not solution-archetype-first.
-
ResearchBench (2025/2026). Peer-reviewed Findings of ACL benchmark. Paper. Decomposes 1,386 recent papers from 12 disciplines into inspiration retrieval, hypothesis composition, and ranking with temporal controls. It is a benchmark rather than a transfer curriculum and exposes ranking and position-bias weaknesses.
-
MOOSE-Star (2026). ICML 2026 paper/preprint. Paper. TOMATO-Star contains 108,717 decomposed scientific-paper trajectories; models are trained for inspiration retrieval and hypothesis composition with hard negatives and temporal tests. It decisively rules out claiming that EoA would be the first discovery- trajectory training resource. Its trajectories are reconstructed from papers and citations, not ontology-mediated bidirectional transfer with rejected attempts and real prospective utility.
-
RLAD (2025/2026). ICLR 2026 paper/preprint. Paper. Jointly trains an abstraction proposer and solver, rewarding the intermediate abstraction for downstream solution success. It is a close architectural precedent for learned problem→abstraction→solution reasoning, but “abstraction” is a same-problem mathematical strategy rather than a reusable cross-domain mechanism.
-
ANALOGYKB (2024). Peer-reviewed dataset and training study. Paper. Contains more than one million analogies spanning 943 relations and 103 analogous relation pairs; training improves recognition and generation. Its relations are much more compact than multi-role scientific mechanisms.
-
ParallelPARC / ProPara-Logy (2024). Peer-reviewed synthetic-data study. Paper. Generates paragraph- level scientific-process analogies and hard distractors; silver data improves trained models while people retain an advantage. It is a strong template for EoA hard negatives and human-validated gold subsets.
-
Emergent Analogical Reasoning in Transformers (2026). ICML 2026 spotlight/preprint. Paper. Controlled synthetic tasks show learned relational geometry and functor-like mappings, with emergence sensitive to data diversity, optimization, scale, and OOD composition. The mechanistic evidence is valuable but comes from toy worlds, not language-grounded opportunities.
-
ScienceAgentBench (2025). Peer-reviewed benchmark. Paper. Contains 102 executable research tasks from 44 papers in four disciplines; the best reported agent completed roughly one-third. It is a sobering bridge between fluent proposals and actual research execution.
-
DiscoveryBench (2025). Peer-reviewed benchmark. Paper. Includes 264 real-world and 903 synthetic data-driven discovery tasks; the best reported system solved only a minority. It does not test analogy, but it reinforces the need to distinguish plausible language from reliable discovery.
Confidence and exclusions¶
High-confidence conclusions rely on multiple primary sources and are unlikely to change with one additional paper: solution-driven design and technology-first application search predate the EoA; retrieval, mapping, adaptation, criticism, and repair are separable; and recent systems already train on decomposed scientific-discovery trajectories.
The “full configured workflow was not located” conclusion is only moderate confidence. Terminology is fragmented, some full texts were inaccessible, commercial tools are opaque, patents were not exhaustively searched, and 2026 work is moving quickly. The inventory excludes ordinary transfer learning, organizational reverse innovation, inverse design optimization, educational question generation, and generic recommender systems unless they perform structural problem–solution transfer.