Literature search log¶
Part of Inverse Innovation with the Encyclopedia of Abstractions · Literature search log · Last revised August 2026
Search date: 2026-08-03
Coverage cutoff: Material discoverable through 2026-08-03
Purpose: Preserve the search strategy, terminology changes, contrary cases,
and unresolved gaps behind the integrated review. This is a scoping search, not
a claim of exhaustive bibliographic recall.
Search method¶
Searches were decomposed into four literatures that use different terminology: cognitive analogy, engineering design and computational creativity, LLM-enabled ideation and discovery, and synthetic analogy curricula. Broad discovery queries were followed by title/DOI searches, citation chaining from review papers, and inspection of primary-source abstracts or full texts. Central claims were not based on news coverage, vendor pages, Wikipedia, or search-result summaries when an original paper was available.
Primary sources were sought through publisher and conference repositories, including ACL Anthology, arXiv, DRS Digital Library, AAAI, Design Society, Cambridge Core, ScienceDirect metadata pages, institutional repositories, and official project pages. Some publisher full texts were access-restricted; in those cases the record is marked as abstract-level unless an author manuscript or institutional copy was available.
Query families¶
The following are representative query families rather than an export of every minor wording variation.
Cognitive analogy and far transfer¶
analogical transfer structural alignment far transfer relational similarityGick Holyoak analogical problem solving convergence radiation problemGentner structure mapping systematicity analogical retrievalsurface similarity structural similarity analogical retrieval failureschema induction multiple analogs transfer problem solvingnegative analogical transfer causal structure source target
Computational and engineering design¶
data driven design by analogy encoding retrieval mapping evaluationsolution driven bio inspired design problem search application searchLLM solution-driven bio-inspired concept generation problem searchAskNatureGPT AskNatureNet bidirectional retrieval solution driventechnology push application search existing technology new applicationtechnology opportunity discovery function patents existing productcomputational creativity problem finding automated systemTRIZ LLM design analogy automated inventionFritz Zwicky morphological analysis morphological box innovation searchgeneral morphological analysis cross-consistency assessment combinatorial fieldpatent analogy mining problem solved concept cross technologyarchetypal solutions cross-domain analogy design LLM
LLM analogy, creativity, and discovery¶
large language model analogical reasoning cross-domain transfer benchmarkLLM far analogy narrative system similarityLLM robustness analogy unfamiliar symbols shuffled alphabetLLM strategic decision analogy causal mapping precision recallLLM cross-domain analogical creativity problem reformulationLLM scientific idea generation prior literature novelty evaluationmulti-agent hypothesis generation critique tournament refinementLLM problem finding innovation opportunity generation prior art
Synthetic analogy curricula¶
synthetic data train analogical reasoning relational structure language modelmillion scale analogy knowledge base train generation recognitionLLM generated natural language analogies hard distractors trainingsynthetic relational tasks emergence analogical reasoning transformerheld out domain analogy compositional generalization hard negatives
Search pivots and important discoveries¶
Pivot 1: “inverse innovation” was not the useful historical term¶
The phrase is dominated by a different management concept. The relevant design literatures use solution-driven, biology-push, technology-push, application search, function finding, and technology opportunity discovery. Searching those terms revealed much closer precedents than generic queries for reverse or cross-domain innovation.
Pivot 2: solution-to-problem direction has clear prior art¶
Solution-driven biologically inspired design explicitly starts from a biological solution, abstracts its principle, searches for an applicable problem, defines that problem, and applies the principle. A 2024 system by Chen et al. uses an LLM to automate problem search, analogical transfer, concept generation, and concept evaluation. Its expanded 2025 journal version, AskNatureGPT, is the closest directional precedent discovered. Any historical-priority claim for the mere solution-to-problem direction is therefore untenable.
Pivot 3: technology opportunity discovery is a second close lineage¶
Patent-based technology opportunity discovery starts with an existing technology or product and searches for new technologies, products, functions, or application areas. A 2015 framework extracted functions from 223,603 patents and automated cross-field opportunity discovery through functional similarity. A 2025 IEEE framework distinguishes explicit and implicit application opportunities and uses link prediction over a technology-function network. These systems are not framed as cognitive analogy, but they materially overlap the Encyclopedia program's solution-first opportunity-search direction.
Pivot 4: current LLM evidence is mixed rather than simply negative¶
Advanced LLMs approach or match humans on some constrained semantic or abstract analogy tasks. They also remain brittle on distant narrative analogies, retrieval from long contexts, causal matching, and transfer to unfamiliar symbol systems. The strongest synthesis is that an LLM can be a high-recall analogy or candidate generator, but unscaffolded output should not be treated as reliable structural transfer. Independent filtering, explicit mechanisms, negative cases, and empirical tests are substantive parts of the method, not administrative polish.
Pivot 5: the synthetic-training direction has direct precedents¶
ANALOGYKB, ParallelPARC, and controlled synthetic relational tasks show that constructed analogy data can improve recognition or generation and can reveal mechanisms of relational transfer. They support the plausibility of an Encyclopedia-derived curriculum, but do not establish that opportunity trajectories will teach robust, open-domain problem-to-archetype-to-solution transfer. That remains a testable proposal.
Pivot 6: 2026 work rules out a broad cross-domain-transfer claim¶
Shen, Druckmann, and Zou (May 2026) provide a close operational precedent. Their pipeline decomposes a supplied biomedical problem into objects and relations, constructs analogies to distant fields, searches those fields for solutions, maps the solutions back, checks novelty against the literature, and implements four selected proposals. The work is target-problem-first rather than solution-archetype-first, but it means the Encyclopedia program must not claim to be the first LLM system to perform explicit cross-domain solution transfer.
Pivot 7: discovery-trajectory training is already an active research line¶
MOOSE-Star (ICML 2026) trains models on 108,717 decomposed scientific-paper trajectories and evaluates inspiration retrieval and hypothesis composition. RLAD (ICLR 2026) learns a problem-to-abstraction-to-solution architecture by rewarding abstractions for downstream solving value. These are not Encyclopedia-style inverse trajectories, but they make a general claim such as "first synthetic data for discovery or abstraction-mediated solving" untenable. The narrower proposal is to train on explicit, bidirectional, ontology-grounded transfer trajectories containing role alignments, hard negatives, criticism, prior-art evidence, repair, and terminal dispositions.
Pivot 8: the best LLM systems argument is architectural, not categorical¶
Recent controlled comparisons find a useful asymmetry: people tend toward lower recall and higher precision in analogy retrieval, while LLMs can produce more candidates but more causally incoherent matches. Other studies find strong performance on compact analogy tasks alongside severe degradation on unfamiliar symbols, long contexts, or distant narratives. This supports a generator-plus- filter architecture and causal controls; it does not support either "LLMs lack analogy" or "LLMs have solved human-like far transfer."
Citation chaining hubs¶
The principal mapping sources used to locate earlier systems were:
- Jiang et al., Data-Driven Design-by-Analogy: State of the Art and Future Directions (47 data-driven design-by-analogy studies organized around encoding, retrieval, mapping, and evaluation).
- Chen et al., AskNatureNet and AskNatureGPT for solution-driven BID, bidirectional retrieval, and related LLM design tools.
- Recent LLM analogy benchmark papers, especially SCAN, ARN, AnaloBench, ANALOGYKB, ParallelPARC, and Stevenson et al. (2026), for the debate over human-like far transfer.
- Recent scientific-ideation systems, especially SciMON, Scideator, the Si et al. human comparison, and Co-Scientist, for novelty search, critique, iteration, and empirical validation.
- Shen, Druckmann, and Zou (2026), MOOSE-Star, RLAD, and ResearchBench for the newest cross-domain solution-transfer and discovery-training boundaries.
Negative and exclusion searches¶
Searches for ordinary transfer learning, domain adaptation, style transfer, cross-lingual transfer, organizational “reverse innovation,” and generic product recommendation were excluded unless they involved structural problem-solution transfer. Problem-first TRIZ and design-by-analogy tools were retained as methodological comparators but not counted as solution-to-problem precedents.
No discovered source combined all of the following in one evaluated system: a broad domain-general ontology of solution structures, a systematic solution-structure-by-human-domain matrix, generation of complete target-domain problem/solution proposals, external prior-art research on every proposal, independent strict and empirical-partner gates, iterative repair with preserved trajectories, and resource accounting. This is a provisional scoped-search gap, not evidence that no such system exists.
Material search limitations¶
- The search crossed fields with fragmented terminology and cannot prove historical priority.
- No subscription bibliographic database was used to guarantee exhaustive citation coverage.
- Some publisher pages exposed only abstracts or article previews. Author manuscripts and institutional repositories were used when available.
- Commercial innovation platforms may implement undisclosed workflows that cannot be evaluated from public technical evidence.
- Recent 2025-2026 papers are still changing versions and citation networks.
- Patent literature was sampled through representative method papers rather than exhaustively searched as patent claims.
- The review evaluates published systems, not whether any generated Encyclopedia candidate is legally novel, commercially valuable, or effective.
Searches still worth repeating before public release¶
- Backward and forward citation search from AskNatureGPT, the 2015 function-based technology-opportunity framework, and ViMimic.
- Focused IEEE Xplore and Web of Science/Scopus searches for technology-push application discovery and automated function finding.
- Patent-family search for systems that claim archetype-by-domain innovation matrices or automated solution-first problem discovery.
- A final version check for rapidly evolving 2026 analogy and scientific-agent papers.
- Outreach to authors of the nearest systems for missed precedents and corrections after the public report is drafted.