From Solution Archetypes to Problems¶
Part of Inverse Innovation with the Encyclopedia of Abstractions · From Solution Archetypes to Problems · Last revised August 2026
Review date: 2026-08-03
Coverage cutoff: Material discoverable through 2026-08-03
Review type: Cross-disciplinary scoping review with primary-source checking;
not a registered systematic review and not a historical-priority determination
Program covered: Encyclopedia of Abstractions inverse-innovation Experiments
1–6; Experiment 7 had not been run when this review was frozen
Executive answer¶
The Encyclopedia experiments appear to have found a serious and researchable way to use abstractions for systematic cross-domain opportunity generation. The literature review also makes the defensible claim substantially narrower than “the first cross-domain innovation system” or “a breakthrough that taught LLMs analogy.” Those broader claims are contradicted by prior work.
The core ingredients have long histories. Cognitive science explains how people align relational structures, why distant analogies are difficult to retrieve, and why mapping an analogy is not the same as adapting it. Case-based reasoning stores and repairs prior solution paths. Engineering design uses functions, causal models, biological strategies, patents, TRIZ principles, and conceptual blends to transfer solutions. Solution-driven bio-inspired design explicitly starts with a biological solution and searches for problems. Patent-based technology-opportunity systems start with existing technologies and search for new applications at corpus scale. Recent LLM systems generate and criticize scientific hypotheses, search the literature, and in selected cases conduct experiments.
There are particularly close precedents. AskNatureGPT starts with a biological source and generates an engineering application/problem and design concept. Yoon et al. mine functions from 223,603 patents to discover technology and product opportunities. Kang et al. retrieve analogies from roughly 1.7 million scientific papers using purpose–mechanism representations. Most importantly, Shen, Druckmann, and Zou (2026) give an LLM a biomedical problem, explicitly map it to distant domains, import solutions, check related literature, and implement four selected proposals. These works mean that neither solution-to-problem search nor LLM-mediated cross-domain solution transfer can be claimed as new in itself.
The narrower finding is nevertheless meaningful. Within this scoped search, no reported system was found that combines all of the following:
- a curated, domain-general ontology of reusable solution archetypes and mechanisms;
- a predeclared crossing of those structures with a broad set of target domains;
- multiple complete, operational proposals for each declared pairing;
- output-specific external prior-art research;
- criticism separated from proposal generation;
- practicality, impact, deployability, cost, adopter, and empirical-partner assessment;
- a falsifiable next test;
- iterative revision with rejected and terminal trajectories preserved; and
- raw artifacts, provenance, denominators, and resource accounting.
That is a provisional integration and experimental-design gap, not a historical-priority claim. It may reflect an overlooked paper, patent, dissertation, or proprietary workflow. And even if the combination is new, uniqueness does not establish usefulness. The next scientific questions are comparative: whether the ontology improves structural transfer, whether the matrix finds opportunities adaptive search would miss, whether scrutiny reduces false positives, whether later proposals justify their cost, and whether domain readers recognize real problems and worthwhile tests.
The review also changes the synthetic-data proposal. Training on discovery trajectories is no longer merely speculative: MOOSE-Star trains on 108,717 decomposed scientific-paper trajectories, and RLAD learns problem→abstraction→solution reasoning by rewarding abstractions for downstream solution value. The EoA opportunity is not to be first to train on reasoning or discovery traces. It is to test whether explicit, bidirectional, ontology-grounded transfer trajectories—including role alignments, hard negatives, prior-art evidence, criticism, repair, and terminal failures—teach a more portable form of abstraction-mediated transfer under held-out-domain and held-out-pair tests.
1. What question this review asked¶
The experiments began with a simple reversal of the normal innovation prompt:
Instead of beginning with a problem and searching for a solution, begin with a general solution structure and ask what problem it could solve in each domain.
The idea is plausible because the Encyclopedia of Abstractions supplies a large, curated vocabulary of domain-general structures rather than a loose request to “be creative.” The experiments then added multiple proposals, critique, revision, external research, practical assessment, strict and empirical-partner gates, and complete artifact retention.
The review did not ask only whether anyone had used the identical name. It asked five more useful questions:
- What is known about analogical retrieval, structural mapping, transfer, and adaptation?
- Do established design systems already work from solutions toward problems or applications?
- Have computational systems performed cross-domain opportunity search at scale?
- What can contemporary LLMs actually do reliably in analogy and scientific ideation?
- What prior work bears directly on training models with the preserved EoA trajectories?
The search was intentionally adversarial toward broad novelty claims. It used the vocabulary of each neighboring field, followed citation chains, preferred primary papers and official artifacts, and retained contradictory findings. The full protocol is in REVIEW_PROTOCOL.md, the query and pivot record is in SEARCH_LOG.md, and 55 central sources are annotated in SOURCE_INVENTORY.md.
2. Terminology: the name is useful locally, but not as a literature-search key¶
The phrase inverse innovation is useful within this project for the reversal from solution archetype to target problem. It is not an established unambiguous field label. Reverse innovation already refers to innovations developed in lower-income or emerging markets and later introduced into wealthier markets. Inverse design usually refers to computational optimization from desired performance toward a structure. Using either term without a local definition would send readers into the wrong literature.
Neighboring fields use several different labels for the relevant direction:
- solution-driven or biology-push bio-inspired design;
- technology-push innovation and outward technology transfer;
- application search, function finding, and requirement identification;
- technology-opportunity discovery;
- design by analogy and analogical problem evolution;
- problem finding, need finding, or need–solution co-evolution;
- literature-based discovery, conceptual blending, and bisociation.
For the public report, a precise descriptive phrase should accompany the local name on first use:
Solution-archetype-first cross-domain opportunity search: a method that starts with an explicit reusable solution structure, crosses it with target domains to formulate candidate problems and operational responses, and then scrutinizes the candidates for prior art, structural validity, usefulness, and testability.
This definition makes the task direction and the empirical obligations visible.
3. Cognitive foundations: why the method could work, and why it will fail often¶
3.1 Structural similarity is not merely topic similarity¶
Structure-mapping theory treats analogy as alignment of relational systems. A good analogy preserves roles and relations—particularly interconnected causal or higher-order relations—even if the objects differ. The Structure-Mapping Engine made the idea computational: propose local correspondences, assemble globally consistent mappings, prefer systematic relational structures, and project candidate inferences.
For the Encyclopedia, this supports representing an archetype as more than a name and paragraph. A useful transfer unit needs roles, causal or functional relations, enabling conditions, constraints, intervention, expected result, value criterion, and failure conditions. Without those, a model can produce metaphorical resemblance while missing the mechanism that makes the source solution work.
But structure mapping begins after a usable source and target representation have been supplied. It does not discover a target problem, retrieve the right case from millions of candidates, prove that a projected mechanism will survive adaptation, or establish that anyone wants the result. It is a component, not a complete innovation engine.
3.2 Deep retrieval is harder than recognizing a good match¶
The most important classical result for this program is the separation between retrieval and evaluation. In the convergence studies of Gick and Holyoak, a distant source could unlock the target solution, but spontaneous access was unreliable and improved greatly when participants were told to use the source. Gentner, Rattermann, and Forbus found that surface similarity strongly influenced which prior situation was retrieved, while deeper relational similarity more strongly influenced judgments of which inference was sound.
MAC/FAC embodies the compromise: a cheap, broad content stage filters the case library, then an expensive structural matcher reranks a small set. That architecture is strikingly relevant to the EoA scaling problem. Exhaustive structural scrutiny is costly, but an aggressive semantic prefilter may throw away the distant match that matters.
The Experiment 6 finding that proposals two through four rescued additional cells is an empirical version of this recall problem. The preregistered rule did not support a blanket “always generate four” conclusion, yet later proposals still recovered five P1-failure cells. Experiment 7 should therefore evaluate selectors in terms of opportunity recall and avoided scrutiny cost, not only the fraction of proposals rejected early.
3.3 Comparison, varied examples, and hard negatives can make structure portable¶
Gick and Holyoak's schema-induction work showed that comparing multiple analogs can make their common schema explicit and improve later transfer. Later behavioral and computational work similarly supports relational labels, varied positive examples, near misses, and contrasts that share surface features while differing structurally.
This suggests that an archetype should eventually be taught or presented with a small contrast set:
- several structurally valid examples from distant domains;
- a semantically close but structurally wrong “surface twin”;
- a near miss that violates one enabling condition;
- a repaired example showing what modification restored validity; and
- an explicit statement of what must remain invariant across implementations.
That contrast set would serve three purposes: clarify the ontology, improve generation, and create hard evaluation cases. It would also expose “archetypes” that are simply attractive prose generalizations of one source case.
3.4 Adaptation is a separate failure point¶
Correct retrieval and mapping do not guarantee a valid target intervention. Case-based reasoning decomposes work into retrieve, reuse, revise, and retain. PRODIGY's derivational analogy goes further: it stores problem-solving justifications, abandoned alternatives, and failures, tests those justifications during replay, and repairs or abandons the transferred path when they no longer hold.
This is a strong precedent for the experiment's criticism–revision loop and for preserving failed trajectories. An EoA proposal should state which source relations were preserved, which were changed, what target constraint forced the change, why the adapted mechanism should still work, and what observation would show that the analogy was decorative rather than causal.
3.5 Farther is not necessarily better¶
Far analogies can produce novelty and escape local fixation, but distance is not a quality score. A surprising analogy may omit a critical boundary condition; a nearer source may preserve constraints required for implementation. The 2026 scientific-analogy preprint reports a novelty–applicability tension in some conditions, which is exactly why novelty, structural depth, feasibility, importance, and evidential support should be scored separately.
The cognitive literature therefore predicts the overall pattern seen in the experiments: a broad generator can produce legitimate remote mappings, many outputs will be shallow or already known, retrieval and adaptation will fail for different reasons, and independent criticism can add real value.
4. Prior systems already work in the “reverse” direction¶
4.1 Solution-driven bio-inspired design¶
Bio-inspired design supplies the clearest conceptual contrary case. The field distinguishes a problem-driven direction, where an engineering need prompts a search in biology, from a solution-driven direction, where a biological system is studied first and its principle is matched to engineering applications.
A published solution-driven system-of-systems process explicitly moves from biological solution identification through principle extraction and reframing to problem search, problem definition, and application. Research on analogical problem evolution shows that a source solution can change the problem formulation itself.
Mechanism-rich tools also predate LLMs. DANE represents biological cases with structure–behavior–function causal models. Idea-Inspire uses SAPPhIRE—State change, Action, Part, Phenomenon, Input, oRgan, and Effect—to support structured retrieval and transfer. These systems show why “same function” is insufficient: causal roles and physical phenomena constrain whether a transfer works.
4.2 AskNatureNet and AskNatureGPT¶
AskNatureNet connects biological sources, functions, and applications and supports both forward and reverse retrieval. It contains 5,541 nodes and 5,082 edges: 1,797 biological sources, 3,101 functions, and 643 applications. It is not a broad human-domain ontology, but its source→function→application path is plainly close to the EoA direction.
AskNatureGPT is the strongest directional and generative precedent found. Trained on 305 successful cases, it receives a biological source, benefits, and biological model, then identifies an engineering application/problem and generates a natural-language design concept. Two separately trained relevance classifiers screen the result. The study includes held-out cases, multiple outputs, ablations, human novelty and feasibility ratings, and design cases.
The differences from EoA are real but narrower than the original intuition. The source space is biology rather than domain-general solution structures; the training set contains successful cases rather than a factorial matrix with failures; the classifiers test relation to source benefits/models rather than functioning as independent adversarial critics; and the paper does not run proposal-specific prior-art searches, adopter/cost analysis, iterative repair, built-prototype trials, or resource accounting. The right comparison is not “they did ordinary design and we reverse it.” Both reverse direction. The EoA program adds breadth, a different unit of abstraction, systematic coverage, and a much heavier evidence pipeline.
4.3 Technology-opportunity discovery¶
Patent analytics provides a second independent lineage. Instead of beginning with a problem, technology-opportunity systems begin with an existing technology, product, capability, or patent and search for new functions or applications.
Yoon et al. (2015) extracted functions from 223,603 patents and defined four opportunity paths from existing technologies and products toward modifiable technologies, producible products, related products, and technologies applicable to a product. InnoGPS maps more than five million patent records into a technology space for neighborhood exploration and near- or far-domain inspiration. Qiao et al. (2025) extract technology–function pairs, map explicit applications, and use network link prediction to infer implicit applications.
These systems operate at a scale far beyond the EoA experiments. Their usual output is a ranked technology, function, application area, or network link—not a complete target-domain problem and intervention with adopter analysis, cost, falsifier, external critique, repair history, and terminal disposition. But that is a difference in output and governance, not a license to ignore them as prior systematic solution-first opportunity discovery.
4.4 Real organizations deliberately search for applications¶
Reinsberger et al. followed four organizational technology–market-linking projects involving 306 interviews, 89 innovation proposals, and 18 search agents. Existing technologies and emerging needs co-evolved into need–solution pairs through domain crossing, evaluation, and learning.
This field evidence matters because it is closer to adoption than any LLM rubric. It confirms that solution-to-need search can be an intentional process and that the problem and solution should be allowed to change together. It also sets a higher bar for “useful”: a model's plausible adopter story is not the same as an interview, a committed partner, a prototype, or a decision to proceed.
5. Systems that cover adjacent parts of the pipeline¶
5.1 Design-by-analogy retrieval and representation¶
The Functional Basis standardizes verb–object and function–flow descriptions so designs can be retrieved independently of physical form. DANE and Idea-Inspire encode causal structure in more depth. Patent analogy systems and Analogy Mining for Specific Design Needs retrieve semantically distant sources for supplied targets.
Kang et al. scale purpose–mechanism retrieval to roughly 1.7 million scientific papers and test its effect on scientists' ideation. This is a particularly important baseline for any future claim that EoA's explicit ontology improves retrieval or creative adaptation. A fair test would compare curated archetype retrieval with purpose–mechanism embeddings and with unstructured semantic retrieval under equal budgets.
5.2 TRIZ and abstraction-mediated solving¶
TRIZ has long mapped specific problems into generalized contradictions and general solution principles before returning to specific solutions. It shows that a corpus of generalized solution strategies can guide invention without modern generative models. AutoTRIZ uses LLM modules and a fixed TRIZ knowledge base to automate much of this flow and generate multiple interpretable solutions.
The dominant direction is problem→generalized problem→general solution→specific solution, so AutoTRIZ is not a close task-direction match. It is nonetheless a mandatory baseline for the proposition that a curated abstraction layer adds value beyond an unstructured prompt.
5.3 Zwicky morphological analysis¶
Zwicky's morphological approach declares the parameters of a problem and the possible values of each parameter, then constructs the combinatorial field of resulting configurations. Later general morphological analysis uses cross-consistency assessment to eliminate incompatible parameter pairs and reduce the raw field to a smaller internally consistent space.
The EoA archetype × domain matrix is recognizably a two-axis morphological field. Its distinctive move is not the Cartesian product itself, but what it places in each cell: a generated target problem, a structural mapping, an operational intervention, contrastive prior-art research, and a falsifiable next test. The external scrutiny stage performs some of the practical work that cross-consistency assessment performs in morphological analysis, but on a fully developed proposal rather than on parameter pairs alone. This is a close methodological precedent and should be treated as such.
5.4 Computational creativity and conceptual blending¶
Computational creativity did not begin with LLMs. Divago and COINVENT formalize cross-domain mapping, generalization, blending, constraint checking, and, in some cases, argumentation over generated blends. These systems establish that structured cross-domain concept generation and automated evaluation have long precedents.
Their goals are typically coherent, novel, or valuable artifacts inside a creative space—not discovering unmet problems, validating stakeholders, or researching output-specific prior art. They are adjacent ancestors rather than the same end-to-end system.
5.5 Contemporary human-in-the-loop design¶
ViMimic uses LLMs for structured analogy retrieval, “archetypal solutions,” problem finding, and cross-domain semantic and morphological mapping. In an 18-participant product-design study, the authors reported improvements in novelty, functionality, mapping efficiency, and diversity. It is close in language and intent, but interactive, product-focused, and not a declared archetype-by-human-domain search with proposal-specific external scrutiny.
6. What current LLM research actually says about cross-domain transfer¶
6.1 The answer is neither “cannot” nor “solved”¶
LLMs can perform impressive analogical tasks. Webb, Holyoak, and Lu reported human-level or better behavior on several newly constructed text analogy tasks. Other studies find that advanced models reproduce multiple human effects of semantic and structural similarity.
The capability is also brittle. ARN finds severe difficulty with far narrative analogies. AnaloBench shows limited gains from scale on long scenarios and retrieval from large contexts. Stevenson et al. (2026) compare children, adults, and multiple LLMs across familiar, near, and far symbol systems: human performance remains relatively stable while every tested model degrades with distance.
Counterfactual-task research adds an important qualification. Models can fail when familiar conventions are remapped, but code augmentation can recover performance on some formal tasks. The relevant unit is therefore the whole system—model, prompt, tools, representation, retrieval, and verifier—not an essentialist claim about what the base model “has.”
6.2 High recall and low precision is a productive but dangerous profile¶
A 2026 Strategy Science comparison tested 199 people and eight LLMs. Its reported pattern is highly relevant: humans produced fewer analogies with higher precision, whereas models retrieved more possibilities with lower precision, including plausible-sounding but causally wrong matches.
That profile explains why the EoA pipeline can be useful even if the base model has not acquired robust human-like transfer. A model can widen the candidate pool; explicit archetypes can focus the search; critics can test structural and operational coherence; prior-art search can remove rediscoveries; and an empirical-partner lane can separate plausible but research-dependent ideas from strictly deployable ones. The filtering architecture is part of the scientific contribution because the generator is expected to overproduce.
6.3 Prompting gains do not prove analogy caused the gain¶
Some methods ask a model to generate relevant examples or analogies before solving a problem and report higher accuracy. But Qin et al. find that random self-generated examples can match or beat relevant examples on some tasks and that exemplar correctness is a major driver. More context, an extra computation pass, or answer leakage can masquerade as an analogy effect.
Future causal tests of EoA should therefore include:
- an irrelevant-archetype control;
- an equal-length unstructured context control;
- a surface-similar but structurally wrong source;
- an archetype name without its mechanism;
- a mechanism description without the archetype label;
- a corrupted mechanism with one critical relation reversed; and
- equal model, sampling, retrieval, and critic budgets.
6.4 The nearest contemporary cross-domain system¶
Shen, Druckmann, and Zou (2026) is the source most likely to change an incautious paper title or abstract. Their pipeline:
- receives a biomedical target problem;
- extracts objects and relations;
- generates analogies to distant-domain problems with explicit mappings;
- searches the analogous domain for solution methods;
- transfers a selected solution back to the biomedical target;
- searches Semantic Scholar and reranks results for novelty checking;
- validates part of the judging process with human annotations; and
- implements selected proposals for four biomedical problems.
The paper reports substantially greater domain and solution diversity, greater judged novelty than its baselines, and quantitative gains in the four case implementations. This is not merely fluent ideation.
Its limits define the comparison. It is target-problem-first, uses biomedical targets, samples candidates rather than exhausting a declared ontology×domain matrix, and implements a small selected subset. The authors identify feasibility evaluation and grounded end-to-end execution as remaining bottlenecks. The EoA program can still test the reverse direction and a different coverage and scrutiny architecture, but should cite this system prominently and should not claim to have established cross-domain LLM creativity in general.
7. Automated scientific ideation already includes research, critique, and iteration¶
Systems such as SciMON and Scideator retrieve literature, generate structured ideas, and check novelty. The AI co-scientist uses multiple agents for generation, critique, ranking, and iterative evolution, with experimental validation in selected biomedical cases. These systems mean that “generate, criticize, revise, research, and test” is not a new workflow by itself.
Their dominant direction is from a supplied research goal, problem, or area to hypotheses. Their representations are generally literature snippets, research facets, or generated hypotheses rather than a large curated ontology of reusable solution structures. EoA's claim boundary remains task direction, representation, coverage accounting, and the integration of practical opportunity gates—not the invention of agentic critique.
The execution evidence also sets a sobering bar. In ScienceAgentBench, the best reported agent completed only about one-third of 102 executable research tasks. Fluent opportunity documents and even good literature searches remain far from reliable autonomous research execution.
8. What is distinctive about the EoA program after the literature review¶
The nearest-systems matrix compares the configured features directly. The defensible distinction has five layers.
8.1 The experimental unit is a curated solution structure¶
Prior systems commonly use solved cases, biological strategies, patents, functions, contradictions, products, purpose–mechanism snippets, or free-form examples. The EoA uses deliberately authored, domain-general solution archetypes connected to mechanisms and prime abstractions. Whether that representation is better is unproven, but it is a meaningful independent variable.
8.2 The search space is declared rather than opportunistically sampled¶
The archetype×domain matrix creates a denominator. It records unproductive cells, duplicates, rediscoveries, failures, and research-dependent candidates as well as survivors. That enables yield and resource measurements that a portfolio of selected examples cannot provide.
The matrix is also expensive and possibly wasteful. Its value must be compared with adaptive search, retrieval-led search, and cheap prefilters. Systematic does not automatically mean efficient.
8.3 Complete proposals precede expensive scrutiny¶
Experiments 5 and 6 moved beyond short hypotheses: each initial proposal had to state a target problem, mechanism mapping, intervention, users or adopters, operating conditions, expected value, risks, and test. This reduces the chance that an attractive phrase survives only because critical details were never specified.
8.4 Prior art and opportunity quality are separate gates¶
The experiments discovered that conceptual and operational quality can coexist with extensive prior art, and that a novel-looking proposal can lack a real adopter or feasible deployment path. External search was therefore moved earlier and the rubric expanded beyond novelty to impact, deployability, cost, adopter pull, and empirical-partner suitability.
This is still bounded scrutiny, not proof of novelty or market demand. The value is procedural: the pipeline makes distinct failure reasons visible.
8.5 The rejected trajectories are treated as data¶
The program retains proposals, research queries, retrieved evidence, critiques, revisions, controller decisions, terminal labels, and resource use. That enables retrospective policy evaluation such as Experiment 7 and creates a possible training resource. Most published innovation systems foreground successful outputs; a complete matrix makes survivorship and search cost measurable.
Taken together, these layers justify describing the work as a reproducible, ontology-grounded experimental pipeline for generating and boundedly scrutinizing cross-domain opportunity hypotheses. They do not justify calling the outputs inventions, validated interventions, or evidence that world novelty has been established.
9. How the literature changes the interpretation of Experiments 1–6¶
What the experiments do support¶
- A reusable solution structure can guide a model to formulate operational opportunities in unrelated target domains.
- The result is repeatable across multiple archetypes and domains rather than a single anecdote.
- Multiple complete proposals can recover additional strict or empirical-partner candidates beyond the first proposal.
- Proposal-specific prior-art research eliminates or reclassifies many superficially promising ideas and therefore belongs earlier in the pipeline.
- Critique and revision can improve some candidates, while explicit stopping rules prevent indefinite polishing.
- A separately defined empirical-partner lane captures plausible research opportunities that would be unfairly rejected for lacking preexisting validation.
- Complete trajectory retention permits retrospective selector and cost-policy experiments without regenerating the matrix.
What the experiments do not yet support¶
- Historical priority for solution-to-problem innovation or cross-domain analogy.
- A conclusion that the ontology caused better outcomes than equally budgeted unstructured, patent-function, purpose–mechanism, TRIZ, or random-archetype baselines.
- Proof that strict survivors are novel in the world, patentable, safe, economically valuable, or effective.
- A representative estimate of yield over the full Cartesian product; the archetypes were selected rather than sampled to estimate population yield.
- Independent expert or adopter validation across the candidate portfolio.
- A claim that the model has learned robust general-purpose cross-domain transfer rather than performing useful scaffolded generation with substantial external filtering.
- A conclusion that four proposals per cell is universally cost-optimal.
The most scientifically interesting result may therefore be architectural. The experiments show a way to turn a high-recall, uneven generator into an auditable search process. The next work should test which parts of that architecture actually contribute value.
10. Implications for Experiment 7¶
Experiment 7 should remain the planned retrospective selective-scrutiny benchmark. The literature strengthens that choice.
The MAC/FAC analogy predicts that cheap filtering can make scale feasible but can discard distant structural matches. The human–LLM precision/recall asymmetry predicts that a model selector may prefer polished but causally weak candidates. The Experiment 6 labels allow those risks to be measured without paying for new research.
A credible frozen benchmark should:
- give fresh selectors only the four sealed proposals within each Experiment 6 cell;
- conceal all prior-art evidence, evaluations, controller decisions, and final labels;
- freeze top-one, top-two, adaptive, heuristic, and random policies before joining predictions to outcomes;
- report proposal- and cell-level recall for strict survivors and separately for strict-or-empirical-partner candidates;
- report avoided external-research calls, tokens, wall time, and implied cost;
- audit loss by archetype, domain, proposal position, and opportunity type;
- inspect the false negatives qualitatively for structural reasons; and
- proceed to a small prospective confirmation only if a frozen selector preserves a predeclared amount of yield.
The literature also suggests selector ablations: proposal text alone; explicit role mapping; archetype name and mechanism; novelty-without-web judgment; deployability judgment; and a deliberately random or semantically superficial selector. This can reveal whether cheap triage works because it recognizes structure or simply favors conventional, well-written proposals.
Experiment 7 should not be retroactively treated as part of the confirmatory Experiment 6 verdict. It is a new policy-learning study on a sealed labeled dataset.
11. Synthetic data: a promising direction, but the novelty boundary has moved¶
11.1 Direct precedents now exist¶
ANALOGYKB contains more than one million compact analogies and shows training gains in analogy recognition and generation. ParallelPARC generates paragraph-level scientific-process analogies and hard distractors; silver training data improves models while a human gap remains. Controlled synthetic relational studies show analogical behavior emerging under particular data, optimization, scale, and OOD conditions.
Discovery-specific training has moved rapidly. MOOSE-Star constructs TOMATO-Star from 108,717 decomposed scientific papers and trains inspiration retrieval and hypothesis composition with temporal testing and hard negatives. ResearchBench decomposes 1,386 recent papers across 12 disciplines. RLAD learns an intermediate natural-language abstraction whose reward comes from downstream mathematical solving.
These systems eliminate broad claims such as “the first synthetic dataset for scientific discovery,” “the first training on discovery trajectories,” or “the first learned problem-to-abstraction-to-solution model.”
11.2 What an EoA trajectory could add¶
A minimum EoA training unit should be a visible structured record rather than unverifiable hidden chain-of-thought:
target problem → structural decomposition → selected prime/archetype → source
exemplars → role alignment → invariant mechanism → adaptation operators →
proposal → prior-art evidence → critic objections → revision → terminal
disposition → falsifiable next test
The distinctive properties would be:
- bidirectionality: train both archetype→problem and problem→archetype paths;
- ontology grounding: use reusable human-authored solution structures rather than only paper citations or free-form inspirations;
- adaptation labels: show what changed between source and target and why;
- substantive negatives: include duplicates, known mechanisms, non-problems, structural mismatches, impractical proposals, research-needed candidates, and harmful or untestable outputs;
- evidence: attach prior-art results and observable reasons for disposition;
- repair: include successful and unsuccessful critique–revision transitions; and
- prospective outcomes: eventually add expert responses, prototypes, experiments, adoption decisions, and real failures.
This would combine ideas that are currently distributed across MOOSE-Star's decomposed discovery trajectories, RLAD's downstream-utility objective, ParallelPARC's hard negatives, case-based repair memory, and the EoA's explicit ontology and opportunity gates.
11.3 The evaluation must prevent easy leakage¶
Random train/test splits would not establish cross-domain transfer. Near- duplicate domains, repeated archetype wording, or source–target pairs seen in training could let a model memorize templates. A serious evaluation suite should include:
- chronological holdout;
- leave-target-domain-out;
- leave-archetype-out;
- leave-archetype×domain-pair-out;
- leave-source–target-domain-pair-out;
- held-out combinations of relations;
- entity-disjoint and counterfactual-alias tests;
- surface-similar structural negatives;
- “none of the available archetypes applies” cases; and
- a prospective locked set of naturally occurring external problems.
The central test is whether training improves valid transfer on new relations and domain pairings, not whether it makes descriptions sound more analogical.
11.4 Necessary baselines¶
Compare the EoA-trained model with:
- the same base model without additional training;
- an equal-token generic synthetic-data control;
- random-archetype and irrelevant-mechanism training;
- ordinary scientific-paper trajectory training in the MOOSE-Star style;
- retrieval-only purpose–mechanism systems;
- the target-first Shen et al. prompting pipeline;
- source exemplars without explicit archetypes;
- archetype names without descriptions;
- descriptions with corrupted critical relations; and
- with and without critic, prior-art, repair, and negative-terminal supervision.
Evaluation should ascend from structural mapping and hard-negative rejection to adaptation validity, independent problem reality, expert usefulness, executable tests, and ultimately prospective outcomes. No single LLM judge should define success.
12. Compositional transfer as a future research direction¶
Experiments 1–6 principally evaluate first-order transfer: one solution archetype is instantiated in one target domain. Many consequential real-world problems require multiple mechanisms—for example, a primary intervention plus measurement, feedback control, incentive alignment, failure containment, or maintenance. This motivates a separate line of research on higher-order compositional transfer.
Exhaustively crossing every pair or triple of archetypes with every domain would be combinatorially expensive and dominated by arbitrary combinations. A better approach would adapt MOOSE-Star's sequential composition architecture. After a first-order proposal is criticized, its most important residual failure or constraint would become the query for a complementary archetype. The added archetype would need to perform a specific causal role—enabling, monitoring, constraining, repairing, sequencing, or amplifying the first—and search would stop when marginal benefit no longer justified the added complexity.
The existing failed and empirical-partner trajectories provide unusually useful inputs because they record unresolved obstacles that can drive this retrieval. Scientific-paper inspiration sequences could also be mapped to Encyclopedia mechanisms to build an empirical archetype-composition graph, then used to look for combinations documented in one field but absent in another.
The central evaluation must distinguish real composition from decorative complexity. Each component should pass a removal test; the composition should outperform the best component alone; assumptions and timescales must be compatible; and added benefit must be weighed against cost and failure surface. Equal-budget controls should include ordinary revision without a second archetype, a random archetype, and a surface-related but structurally irrelevant archetype.
This possibility expands the Encyclopedia's potential search space, but it has not been demonstrated by the completed experiments. It should appear in the paper as a forward research direction rather than a present capability claim.
13. Claims recommended for the public research report¶
Strong, supportable framing¶
We developed and iteratively tested a reproducible, solution-archetype-first method for generating and boundedly scrutinizing cross-domain opportunity hypotheses. Across a predeclared matrix, the method preserved complete proposals, prior-art evidence, critiques, terminal outcomes, and resource use, enabling yield and policy analysis beyond selected success stories.
The method combines established ideas from structural analogy, solution-driven design, technology-opportunity discovery, computational creativity, case-based repair, and LLM scientific ideation. Its proposed contribution is their integration around a curated domain-general ontology and a declared cross-domain experiment—not the invention of analogy or solution-to-problem search.
The experiments demonstrate systematic hypothesis generation and bounded scrutiny. They do not establish world novelty, commercial value, scientific effectiveness, patentability, or autonomous invention.
Claims to avoid¶
- “The first system for cross-domain innovation.”
- “The first solution-to-problem innovation method.”
- “No one has previously used AI to transfer solutions between domains.”
- “LLMs cannot perform analogies, and the Encyclopedia fixes them.”
- “LLMs now possess human-level far transfer.”
- “Strict survivors are new inventions.”
- “The full Cartesian product will yield the observed sample rate.”
- “Training on the trajectories will create general abstract reasoning.”
A balanced statement of significance¶
The work can be significant without being historically unprecedented. It turns a uniquely large authored ontology into an experimental instrument; repeatedly tests a counter-normal search direction; makes the denominator, failure modes, and resource costs visible; and creates data that can support retrospective policy analysis and future transfer training. A single researcher coordinating frontier models was able to construct and test an integration that previously would have required a multidisciplinary team. That is an important fact about the changing accessibility of research, even though the claims must still meet ordinary evidential standards.
14. Limitations and confidence¶
This was a broad scoping review rather than a database-complete systematic review. Terminology is fragmented across cognitive science, engineering design, innovation management, biomimetics, patent analytics, HCI, computational creativity, NLP, and automated science. No subscription-index search guaranteed coverage of Scopus, Web of Science, IEEE Xplore, ProQuest dissertations, or commercial patent systems. Some full texts were accessible only through author manuscripts or abstracts. Search was predominantly in English. Commercial and unpublished systems may combine features not described publicly.
Recent 2025–2026 papers are moving targets, and several important sources are preprints without independent replication. Reported evaluations differ in models, prompts, tools, corpora, domains, outcome measures, and contamination controls. The feature matrix records what a source demonstrated or reported; “not demonstrated” does not mean the authors or system could not perform it.
Confidence by conclusion:
- High confidence: relational mapping, retrieval, adaptation, and evaluation are separable; surface similarity powerfully affects retrieval; solution-driven design and technology-first application search predate the EoA; recent systems already perform explicit cross-domain solution transfer and train on discovery trajectories.
- Moderate confidence: the full configured EoA workflow was not present in the inspected literature. This is a scoped search result, not proof of absence.
- Low confidence without future evidence: that the EoA representation is superior to alternatives, that the full matrix is economically optimal, that generated survivors are genuinely novel or valuable, or that EoA training will create robust far-transfer capability.
Before a public release, the highest-value literature work would be a backward and forward citation search from AskNatureGPT, Shen et al., Yoon et al., ViMimic, MOOSE-Star, and the technology–function network literature; a focused patent and dissertation search; and an invitation to authors of the nearest systems to identify missed precedents or correct the comparison.
Conclusion¶
The literature review does not deflate the experiments; it locates them.
The EoA program did not invent analogy, problem finding, solution-driven design, technology-push application search, agentic criticism, or synthetic discovery training. Strong systems already occupy every one of those neighborhoods. The 2026 literature in particular has advanced farther in explicit cross-domain scientific transfer and discovery-trajectory training than a broad initial claim would acknowledge.
What remains potentially distinctive is the configured integration: a large curated ontology of domain-general solution structures used as the row variable of a declared human-domain matrix; multiple operational attempts per cell; early proposal-specific prior-art research; separate structural, practical, strict, and empirical-partner judgments; bounded revision; falsifiable next tests; and preservation of the entire search history and cost.
That integration has now produced enough evidence for a serious public research report, provided it is framed as a transparent experimental program rather than a declaration of autonomous invention. Experiment 7 can answer the next narrow operational question—how much of the yield can be retained with selective scrutiny—before the full paper models the scientific and economic consequences of scaling.