Skip to content

Literature search log

Part of Inverse Innovation with the Encyclopedia of Abstractions · Literature search log · Last revised August 2026

Search date: 2026-08-03
Coverage cutoff: Material discoverable through 2026-08-03
Purpose: Preserve the search strategy, terminology changes, contrary cases, and unresolved gaps behind the integrated review. This is a scoping search, not a claim of exhaustive bibliographic recall.

Search method

Searches were decomposed into four literatures that use different terminology: cognitive analogy, engineering design and computational creativity, LLM-enabled ideation and discovery, and synthetic analogy curricula. Broad discovery queries were followed by title/DOI searches, citation chaining from review papers, and inspection of primary-source abstracts or full texts. Central claims were not based on news coverage, vendor pages, Wikipedia, or search-result summaries when an original paper was available.

Primary sources were sought through publisher and conference repositories, including ACL Anthology, arXiv, DRS Digital Library, AAAI, Design Society, Cambridge Core, ScienceDirect metadata pages, institutional repositories, and official project pages. Some publisher full texts were access-restricted; in those cases the record is marked as abstract-level unless an author manuscript or institutional copy was available.

Query families

The following are representative query families rather than an export of every minor wording variation.

Cognitive analogy and far transfer

  • analogical transfer structural alignment far transfer relational similarity
  • Gick Holyoak analogical problem solving convergence radiation problem
  • Gentner structure mapping systematicity analogical retrieval
  • surface similarity structural similarity analogical retrieval failure
  • schema induction multiple analogs transfer problem solving
  • negative analogical transfer causal structure source target

Computational and engineering design

  • data driven design by analogy encoding retrieval mapping evaluation
  • solution driven bio inspired design problem search application search
  • LLM solution-driven bio-inspired concept generation problem search
  • AskNatureGPT AskNatureNet bidirectional retrieval solution driven
  • technology push application search existing technology new application
  • technology opportunity discovery function patents existing product
  • computational creativity problem finding automated system
  • TRIZ LLM design analogy automated invention
  • Fritz Zwicky morphological analysis morphological box innovation search
  • general morphological analysis cross-consistency assessment combinatorial field
  • patent analogy mining problem solved concept cross technology
  • archetypal solutions cross-domain analogy design LLM

LLM analogy, creativity, and discovery

  • large language model analogical reasoning cross-domain transfer benchmark
  • LLM far analogy narrative system similarity
  • LLM robustness analogy unfamiliar symbols shuffled alphabet
  • LLM strategic decision analogy causal mapping precision recall
  • LLM cross-domain analogical creativity problem reformulation
  • LLM scientific idea generation prior literature novelty evaluation
  • multi-agent hypothesis generation critique tournament refinement
  • LLM problem finding innovation opportunity generation prior art

Synthetic analogy curricula

  • synthetic data train analogical reasoning relational structure language model
  • million scale analogy knowledge base train generation recognition
  • LLM generated natural language analogies hard distractors training
  • synthetic relational tasks emergence analogical reasoning transformer
  • held out domain analogy compositional generalization hard negatives

Search pivots and important discoveries

Pivot 1: “inverse innovation” was not the useful historical term

The phrase is dominated by a different management concept. The relevant design literatures use solution-driven, biology-push, technology-push, application search, function finding, and technology opportunity discovery. Searching those terms revealed much closer precedents than generic queries for reverse or cross-domain innovation.

Pivot 2: solution-to-problem direction has clear prior art

Solution-driven biologically inspired design explicitly starts from a biological solution, abstracts its principle, searches for an applicable problem, defines that problem, and applies the principle. A 2024 system by Chen et al. uses an LLM to automate problem search, analogical transfer, concept generation, and concept evaluation. Its expanded 2025 journal version, AskNatureGPT, is the closest directional precedent discovered. Any historical-priority claim for the mere solution-to-problem direction is therefore untenable.

Pivot 3: technology opportunity discovery is a second close lineage

Patent-based technology opportunity discovery starts with an existing technology or product and searches for new technologies, products, functions, or application areas. A 2015 framework extracted functions from 223,603 patents and automated cross-field opportunity discovery through functional similarity. A 2025 IEEE framework distinguishes explicit and implicit application opportunities and uses link prediction over a technology-function network. These systems are not framed as cognitive analogy, but they materially overlap the Encyclopedia program's solution-first opportunity-search direction.

Pivot 4: current LLM evidence is mixed rather than simply negative

Advanced LLMs approach or match humans on some constrained semantic or abstract analogy tasks. They also remain brittle on distant narrative analogies, retrieval from long contexts, causal matching, and transfer to unfamiliar symbol systems. The strongest synthesis is that an LLM can be a high-recall analogy or candidate generator, but unscaffolded output should not be treated as reliable structural transfer. Independent filtering, explicit mechanisms, negative cases, and empirical tests are substantive parts of the method, not administrative polish.

Pivot 5: the synthetic-training direction has direct precedents

ANALOGYKB, ParallelPARC, and controlled synthetic relational tasks show that constructed analogy data can improve recognition or generation and can reveal mechanisms of relational transfer. They support the plausibility of an Encyclopedia-derived curriculum, but do not establish that opportunity trajectories will teach robust, open-domain problem-to-archetype-to-solution transfer. That remains a testable proposal.

Pivot 6: 2026 work rules out a broad cross-domain-transfer claim

Shen, Druckmann, and Zou (May 2026) provide a close operational precedent. Their pipeline decomposes a supplied biomedical problem into objects and relations, constructs analogies to distant fields, searches those fields for solutions, maps the solutions back, checks novelty against the literature, and implements four selected proposals. The work is target-problem-first rather than solution-archetype-first, but it means the Encyclopedia program must not claim to be the first LLM system to perform explicit cross-domain solution transfer.

Pivot 7: discovery-trajectory training is already an active research line

MOOSE-Star (ICML 2026) trains models on 108,717 decomposed scientific-paper trajectories and evaluates inspiration retrieval and hypothesis composition. RLAD (ICLR 2026) learns a problem-to-abstraction-to-solution architecture by rewarding abstractions for downstream solving value. These are not Encyclopedia-style inverse trajectories, but they make a general claim such as "first synthetic data for discovery or abstraction-mediated solving" untenable. The narrower proposal is to train on explicit, bidirectional, ontology-grounded transfer trajectories containing role alignments, hard negatives, criticism, prior-art evidence, repair, and terminal dispositions.

Pivot 8: the best LLM systems argument is architectural, not categorical

Recent controlled comparisons find a useful asymmetry: people tend toward lower recall and higher precision in analogy retrieval, while LLMs can produce more candidates but more causally incoherent matches. Other studies find strong performance on compact analogy tasks alongside severe degradation on unfamiliar symbols, long contexts, or distant narratives. This supports a generator-plus- filter architecture and causal controls; it does not support either "LLMs lack analogy" or "LLMs have solved human-like far transfer."

Citation chaining hubs

The principal mapping sources used to locate earlier systems were:

  • Jiang et al., Data-Driven Design-by-Analogy: State of the Art and Future Directions (47 data-driven design-by-analogy studies organized around encoding, retrieval, mapping, and evaluation).
  • Chen et al., AskNatureNet and AskNatureGPT for solution-driven BID, bidirectional retrieval, and related LLM design tools.
  • Recent LLM analogy benchmark papers, especially SCAN, ARN, AnaloBench, ANALOGYKB, ParallelPARC, and Stevenson et al. (2026), for the debate over human-like far transfer.
  • Recent scientific-ideation systems, especially SciMON, Scideator, the Si et al. human comparison, and Co-Scientist, for novelty search, critique, iteration, and empirical validation.
  • Shen, Druckmann, and Zou (2026), MOOSE-Star, RLAD, and ResearchBench for the newest cross-domain solution-transfer and discovery-training boundaries.

Negative and exclusion searches

Searches for ordinary transfer learning, domain adaptation, style transfer, cross-lingual transfer, organizational “reverse innovation,” and generic product recommendation were excluded unless they involved structural problem-solution transfer. Problem-first TRIZ and design-by-analogy tools were retained as methodological comparators but not counted as solution-to-problem precedents.

No discovered source combined all of the following in one evaluated system: a broad domain-general ontology of solution structures, a systematic solution-structure-by-human-domain matrix, generation of complete target-domain problem/solution proposals, external prior-art research on every proposal, independent strict and empirical-partner gates, iterative repair with preserved trajectories, and resource accounting. This is a provisional scoped-search gap, not evidence that no such system exists.

Material search limitations

  • The search crossed fields with fragmented terminology and cannot prove historical priority.
  • No subscription bibliographic database was used to guarantee exhaustive citation coverage.
  • Some publisher pages exposed only abstracts or article previews. Author manuscripts and institutional repositories were used when available.
  • Commercial innovation platforms may implement undisclosed workflows that cannot be evaluated from public technical evidence.
  • Recent 2025-2026 papers are still changing versions and citation networks.
  • Patent literature was sampled through representative method papers rather than exhaustively searched as patent claims.
  • The review evaluates published systems, not whether any generated Encyclopedia candidate is legally novel, commercially valuable, or effective.

Searches still worth repeating before public release

  • Backward and forward citation search from AskNatureGPT, the 2015 function-based technology-opportunity framework, and ViMimic.
  • Focused IEEE Xplore and Web of Science/Scopus searches for technology-push application discovery and automated function finding.
  • Patent-family search for systems that claim archetype-by-domain innovation matrices or automated solution-first problem discovery.
  • A final version check for rapidly evolving 2026 analogy and scientific-agent papers.
  • Outreach to authors of the nearest systems for missed precedents and corrections after the public report is drafted.