Inverse Innovation with the Encyclopedia of Abstractions¶
Part of Inverse Innovation with the Encyclopedia of Abstractions · The report
This is the full report
Nothing here is abridged. The report's four appendices are published as their own pages — candidate dossiers, artifact index, expert review template, and publication QA. A single file containing the report and all four appendices is also available for offline reading.
Methods, negative results, surviving opportunities, and directions for human-machine research
Last revised: August 2026
Status: research-grade working paper; not peer reviewed
Companion materials: candidate dossiers · artifact index · expert review form · machine-readable data
Abstract¶
Can a system search for useful problems by starting with a reusable solution pattern? This report describes thirteen linked experiments—twelve prospective and one blinded retrospective policy benchmark—and five explicitly post-hoc analyses using the Encyclopedia of Abstractions as a structured source of solution archetypes and mechanisms. Instead of beginning with a known problem and asking for a solution, the pipeline paired an archetype with a domain, generated possible problems that could exhibit the archetype's causal structure, translated the archetype into interventions, and subjected the proposals to criticism, revision, external prior-art search, practical scoring, and—in later experiments—a separate lane for candidates that could only be resolved with an empirical partner.
The program produced evidence for a limited but meaningful claim. A contemporary large language model, when given explicit structural scaffolding and a sufficiently severe scrutiny pipeline, generated researchable cross-domain opportunity hypotheses in a repeatable matrix. The work does not show autonomous invention, world novelty, commercial value, or successful deployment. Those stronger claims would require independent domain experts, deeper searches, field data, and prospective tests.
Several results shaped that conclusion. In Experiment 2, relevant mechanism context outperformed both archetype-only context and an equal-length irrelevant-mechanism control on the prespecified paired internal-quality comparison. Experiment 11 repeated the same 20-cell, three-arm comparison after eight-source external scrutiny and did not confirm that advantage: relevant mechanisms lost 9–11 to archetype-only context, beat the irrelevant control 12–8, and achieved only 52.5% pooled preference (p = 0.446 on the frozen omnibus test). In Experiment 3, 33 of 320 closed-book cells passed the internal pipeline, but external scrutiny found an established or substantially colliding system for 40 of 47 assessed selections; only two survived a stricter post-hoc composite. Experiment 4 favored proposal-first over retrieval-first construction, and Experiments 5 and 6 showed that later proposals can recover opportunities missed by a first proposal, although Experiment 6 missed its frozen rescue floor by one cell. Experiment 7 showed that a cheap proposal-only selector could not safely replace downstream web scrutiny. Experiment 9 then supplied the first probability sample of the larger archetype corpus: all 24 sampled previously untested generated archetypes produced at least one light-screen survivor across three fixed domains, with 60 of 72 cells surviving. That endpoint is coarse researchability, not novelty or deployment readiness. Five explicitly post-hoc analyses found closer adjacent precedent for 8 of 10 sentinel candidates without collapsing their remaining claims, mixed rather than single-archetype yield, no general disposition advantage for governance/process proposals, and a two-level substrate pattern: archetype character shifted proposals between governance and computation while 139 of 150 balanced E9 proposals remained in those two channels. Experiment 12 intervened on that pattern. Sixty-seven of 72 constrained outputs complied with the prohibition on governance and computation—60 with an unambiguously allowed substrate and seven mixed but allowed-primary—but ordinary generation won 68–4 in blinded quality and retained a higher light-screen yield, 60/72 versus 49/72. Experiment 13 then tested the resulting portfolio policy on 60 new cells. Ordinary second proposals produced 41 incremental survivors and substrate-diverse second proposals produced 32; the latter nevertheless added ten opportunities that ordinary diversification missed and therefore met the frozen bounded-complement rule at its maximum permitted 15-point deficit. Experiment 14 separated applicability retrieval from route verification. The new graph achieved 93.2% balanced accuracy when verifying whether a supplied target archetype fit a graph-derived case, but its current diagnostic search retrieved the intended archetype in the top five for only 2/77 positive cases, versus 8/77 for the existing solution-oriented index. Experiment 15 then tested five frozen condition-level route aggregators on 40 new archetypes and 78 eligible positives. The primary aggregator retrieved only 1/78 targets in its top five, versus 6/78 for the solution index and 12/78 for diagnostic maximum-hit. Post-hoc localization found that single-literal routes occupied 96.9% of the primary arm's positive-case top-five slots and that contradicted near misses often ranked above matched positives. The graph is therefore a promising internal verification substrate, not yet an end-to-end problem-to-archetype retriever. The dominant channels are not a hard expressive ceiling, but ordinary diversification remains the stronger default in the tested bundle.
Across Experiments 3–6, this report preserves 59 candidate instances for human review: 27 strict or post-hoc strict-like candidates and 32 empirical-partner candidates. Every one has adjacent prior art. Their unresolved value lies in narrower compositions, governance arrangements, comparisons, or domain instantiations—not in components appearing from nowhere. The principal contribution is therefore a documented search-and-scrutiny method, a body of positive and negative evidence about its operation, and a transparent candidate corpus that specialists can accept, revise, reject, or use as material for further research.
Executive summary¶
The short version¶
The experiment began with a reversal: if an abstraction represents a recurring solution structure, perhaps the system can ask, “What important problem in this domain would have exactly this structure?” A solution archetype such as computability-boundary mapping can be crossed with fields as different as art conservation, engineering design, and public policy. The model must identify concrete actors and a measurable failure, map each part of the abstraction to the target situation, propose an intervention, state why ordinary alternatives are insufficient, and specify what evidence would prove it wrong.
That generation step is easy to make look impressive. The difficult part is distinguishing useful transfer from fluent renaming, rediscovery, and untestable speculation. The experimental program gradually moved its effort downstream:
- generate a structurally explicit proposal;
- criticize it and allow bounded revision;
- search for the problem, nearby methods, and the complete proposed composition;
- compare it with actual rivals rather than with a weak straw baseline;
- demand an operational next test, authority, stop conditions, and rough resource bands; and
- preserve failures rather than reporting only attractive survivors.
The strongest conclusion is not “the system invented 59 things.” It is that a general-purpose language model plus a structured abstraction library can populate a cross-domain search space with candidates that remain coherent after progressively stronger automated scrutiny. Some warrant specialist attention. Most raw ideas do not.
What was learned¶
- Structural context helped internal quality, but the stronger external replication did not confirm an endpoint advantage. Experiment 1 did not show a large, consistent omnibus gain. Experiment 2's narrower paired test favored relevant mechanisms over both controls on terminal internal quality. Experiment 11 reused those exact 60 terminal proposals, added uniform eight-source research, and found no frozen-gate support for relevant mechanisms after scrutiny (Experiment 1 findings; Experiment 2 results; Experiment 11 report).
- Prior-art search is part of generation quality, not a final clerical check. Experiment 3's closed-book funnel produced many plausible candidates, but external research sharply reduced the defensible yield. Of 47 assessed selected cells, 25 substantially collided with prior art and 15 described established practice; six were adjacent and one indeterminate (Experiment 3 final report).
- Proposal-first search was more productive than retrieval-first search in the tested workflow. In Experiment 4, proposal-first produced strict success in 5 of 20 arms, retrieval-first in 1 of 20. The paired estimate favored proposal-first by 20 percentage points, but the sample was too small to establish a stable effect (Experiment 4 report).
- One proposal is not enough. Experiment 5 found all four strict successes among proposals 1–4 and none at proposal 5. Experiment 6 replicated four-proposal generation at larger scale. It found five additional strict-success cells beyond proposal 1, narrowly missing the frozen requirement of six rescues (Experiment 5 report; Experiment 6 report).
- A cheap prefilter was not reliable enough. Experiment 7's blinded proposal-only rankings were moderately consistent with one another, but consistency did not predict downstream survival. The primary two-of-four selection retained 60% of strict proposals and failed all four prespecified retention requirements (Experiment 7 report).
- Coarse researchability was broadly distributed in the sampled archetype corpus. Experiment 9 applied one-shot generation and the same four-source light screen to all 16 hand-curated archetypes, 24 probability-sampled previously untested generated archetypes, and ten reference archetypes. All 24 random generated archetypes were productive in at least one of three fixed domains; 60/72 cells survived. Because the screen was deliberately inexpensive, this is evidence of breadth, not a claim that 83.3% of cells contain novel or valuable opportunities (Experiment 9 report).
- Post-hoc audits narrowed several possible overinterpretations. A ten-candidate prior-art sentinel found closer adjacent precedent in eight cases but no likely substantial collision with the surviving contrastive claims. Yield decomposition showed mixed, protocol-dependent breadth rather than one universal archetype. Blinded coding of 422 externally evaluated proposals found that governance/process and computational/information proposals dominated generation, but neither category had a clear general survival advantage. A corrected join and a new blinded 50-archetype classification then found a two-level substrate constraint: structural/framed character shifted allocation between computation and governance, yet 92.7% of the balanced E9 proposals remained in those two channels (sentinel audit; yield decomposition; proposal-type analysis; substrate analysis).
- The pipeline can leave its dominant substrates, but the forced escape was costly. Experiment 12 prohibited governance and computation as the primary causal substrate in the 72 probability-sampled E9 cells. Sixty-seven outputs complied; 49 survived the light screen compared with 60 ordinary controls, and eight cells survived only under the constraint. Yet blinded quality favored ordinary generation 68–4. Max effort improved both ordinary and constrained proposals in a small follow-up but did not preferentially rescue the constrained lane (Experiment 12 report).
- A deliberately different second lane can add portfolio coverage without becoming the default. Experiment 13 generated common first proposals and paired ordinary-diverse and substrate-diverse second proposals in 60 new cells. All 60 substrate proposals complied. Ordinary P2s produced 41 incremental survivors versus 32 for substrate P2s, but ten cells survived only through the substrate lane. Blinded quality was much closer than in Experiment 12, 32 ordinary wins to 28 substrate wins. The substrate lane met the frozen bounded-complement rule exactly at its allowed 15-point yield deficit; it did not earn outright preference (Experiment 13 report).
- Explicit applicability routes verified structure well, but did not retrieve it. Experiment 14 used 40 stratified route-bearing archetypes to construct 120 remedy-free positive and one-literal-near-miss cases. After blinded eligibility auditing, the DNF verifier accepted 77/77 positives and rejected 32/37 near misses, but diagnostic Recall@5 was only 2/77, below the solution index's 8/77. This supports the graph as an internal verification layer while exposing candidate generation as the bottleneck (Experiment 14 report).
- Naive route-aware embedding aggregation made retrieval worse. Experiment 15 excluded all E14 archetypes, constructed and outcome-blind audited another 120 cases, and compared two baselines with five prespecified condition-level aggregators. On 78 eligible positives, the primary route-aware arm achieved Recall@5 of 1/78, versus 6/78 for the solution index and 12/78 for diagnostic maximum-hit. A post-hoc diagnostic found severe short-route bias and showed that cosine similarity often ranked an explicit one-literal contradiction above its matched positive. Logical route structure requires signed condition evidence and calibration, not arithmetic over raw similarity scores (Experiment 15 report).
- The broadest useful output is a research portfolio. The 59 preserved candidates are invitations to expert judgment and bounded evidence collection. Twenty-seven have a strict or post-hoc strict-like status under their original protocols; 32 need a data-owning partner before the central claim can be judged. These endpoint classes must not be merged into a single “success rate.”
What the work does not establish¶
It does not establish that any candidate is new to the world, patentable, safe to deploy, economically valuable, demanded by customers, or better than the best existing system. It does not show that the model learned a general human-like faculty of far transfer. It does not estimate the yield of the Encyclopedia's full archetype-by-domain Cartesian product without strong sampling assumptions. It also does not show that the method is the first solution-first innovation system; several research traditions and recent systems occupy neighboring territory.
The practical recommendation¶
Do not scale the generator alone. Experiment 9 shows that the search signal is not confined to the small archetype set chosen for early experiments, but its high light-screen yield also shows why a coarse screen cannot be the terminal endpoint. Experiments 12 and 13 show that alternative-substrate generation can add portfolio diversity, while ordinary diversification remains the more productive default. Experiments 14 and 15 add a second architectural rule: use explicit DNF routes to verify a candidate after retrieval, but do not mistake raw embedding similarity—whether maximum-hit or condition-aggregated—for logical route satisfaction. If this program continues, use breadth sampling to allocate attention, preserve multiple proposals where missed opportunities matter, buy an alternative-substrate second slot only as a bounded complement when the value of missed opportunities justifies its lower researched yield, develop signed and calibrated candidate retrieval on a new frozen benchmark, spend deeper retrieval effort on the proposal-specific problem and composition, and send only bounded candidates to specialists. The next scientific step for the opportunity portfolio remains prospective human validation: can appropriate experts understand these candidates, find the evidence record adequate, and identify a nontrivial fraction worth testing?
1. The inverse-innovation idea¶
From problem-first to solution-first search¶
Most innovation narratives begin with a recognized need. Someone observes an undesirable state, studies its causes, and searches for interventions. The present method reverses the first move. It begins with a known pattern of intervention and asks where an unmet problem with the corresponding structure might exist.
This is not simply a request for an analogy. A plausible candidate must identify:
- concrete actors or material systems;
- an observable state, not merely a topic;
- a consequential failure or missed objective;
- a mapping from the source abstraction's parts and causal relations to the target;
- an intervention that changes those relations;
- a serious comparator;
- evidence that the problem and stakeholders are real;
- a remaining distinction after prior-art search; and
- a test whose outcome could stop the idea.
We call the resulting process solution-archetype-first cross-domain opportunity search. “Inverse innovation” is the shorter project name, not a claim to the established development-economics meaning of reverse innovation.
What the Encyclopedia contributes¶
The Encyclopedia of Abstractions is a large, structured corpus of abstractions, mechanisms, solution archetypes, and relations among them. A prime abstraction is intended to capture a compact recurring structure. A mechanism describes how a transition or effect occurs. A solution archetype packages a reusable arrangement of mechanisms that can address a class of problems. A domain is an area of human knowledge or practice, such as aviation, psychology, or museum studies.
The experiments treated an archetype–domain pair as a search cell. The archetype constrained the kind of causal structure sought; the domain supplied potential actors, observables, institutions, and failure modes. This makes the search space explicit. With five archetypes and 64 domains, for example, Experiment 3 had a declared 320-cell matrix rather than a hand-picked collection of favorable anecdotes.
The core workflow¶
flowchart LR
A["Solution archetype<br/>mechanisms and constraints"] --> C["Archetype × domain cell"]
B["Target domain<br/>actors, states, institutions"] --> C
C --> D["Generate distinct problem–proposal hypotheses"]
D --> E["Criticize and revise"]
E --> F["Search problem, prior art, and complete composition"]
F --> G["Strict practical evaluation"]
G --> H["Reject, research with a partner, or send to experts"]
The diagram's important feature is the sequence. Retrieval and testing are not decorations applied after the creative act; they determine whether a fluent proposal remains a defensible research hypothesis.
2. Claims, non-claims, and terminology¶
The supported claim¶
The experiments support a bounded claim:
Given explicit structural representations and proposal-specific external scrutiny, a contemporary large language model can generate cross-domain opportunity hypotheses across a declared matrix and a probability sample of previously untested archetypes; some hypotheses meet prespecified internal researchability and practical-screening criteria.
This is evidence about a configured human–AI research process. It is not evidence that an unaided model spontaneously performs robust transfer in arbitrary settings.
Stronger claims that remain open¶
Independent evidence would be required to conclude that:
- a candidate is absent from all prior art;
- a proposed intervention outperforms the best domain baseline;
- an adopter would fund, authorize, or use it;
- its benefit exceeds its cost and externalities;
- the method transfers across models, prompts, corpus versions, or researchers; or
- trajectories generated by this process can train a model to perform cross-domain transfer without the external pipeline.
Terminology used in this report¶
| Term | Meaning here |
|---|---|
| Cell | One declared solution-archetype × domain combination. A cell can contain several proposals. |
| Proposal | A particular problem–intervention hypothesis within a cell. It is not automatically a unique mechanism or opportunity family. |
| Stage success | A proposal cleared one internal stage. It says nothing about later scrutiny. |
| Established practice | External search found the proposal's substantive intervention already used or well documented. |
| Substantial collision | Prior art covered enough of the claimed composition that the remaining distinction did not support the tested strict lane. |
| Adjacent prior art | Relevant components or nearby systems exist, but a narrower composition, comparison, or domain instantiation remains unresolved. All 59 dossier candidates have this status. |
| Strict success | A candidate met the strict final rules in Experiment 4, 5, or 6. It is an internal researched-candidate endpoint, not proof of novelty or impact. |
| Post-hoc strict-like survivor | One of two Experiment 3 candidates identified after the sealed experiment by combining verified-pipeline and prior-art conditions. It was not a preregistered endpoint. |
| Empirical-partner candidate | An Experiment 6 candidate whose central question could not be decided from public sources and requires a specific data-owning or operational partner. It is disjoint from strict success. |
| Candidate dossier | A plain-language, evidence-linked presentation prepared after the experiments for expert review. Dossier ordering is not an experimental outcome. |
3. Related work and the contribution boundary¶
The motivating intuition has substantial intellectual ancestry. Analogical reasoning research distinguishes resemblance at the level of surface features from alignment at the level of relations. Structure-mapping theory, retrieval models such as MAC/FAC, and studies of analogical encoding all suggest why explicit relational representation can help—and why noticing and adapting a distant source are separate problems. Design-by-analogy, biomimetic search, patent recombination, and computational creativity likewise use existing solutions to stimulate or construct new possibilities.
Two older methods are especially close. TRIZ generalized recurring ways of resolving technical contradictions from patent evidence into reusable inventive principles; contemporary systems such as AutoTRIZ automate parts of that problem-to-principle-to-solution workflow. Zwicky's morphological analysis declares problem parameters and their possible values, constructs a combinatorial field or “morphological box,” and then uses cross-consistency assessment to remove incompatible configurations (Zwicky, 1967; Ritchey, 2015). The experiments' archetype-by-domain matrix is therefore not historically unprecedented as a search-space form. It is a particularly simple two-axis morphological field followed by generative construction and evidence-bearing scrutiny rather than cross-consistency assessment alone.
The Encyclopedia differs descriptively from classical TRIZ in source breadth, catalog size, mechanism detail, and typed relations, while the present pipeline differs from morphological analysis in what occupies a cell and how a proposed configuration is tested. Those architectural differences do not establish superior breadth or yield. Only random corpus sampling and fair end-to-end comparisons could do that.
Recent AI systems come still closer. AskNatureGPT retrieves biological strategies for design problems. Yoon and colleagues infer technology opportunities from functions in patents. A purpose–mechanism knowledge base supports cross-domain analogical search (Kang et al.). Recent preprints evaluate or train scientific analogy, hypothesis generation, and agentic literature search, including Shen, Druckmann, and Zou, MOOSE-Star, and RLAD. The companion literature review and nearest-systems matrix provide the fuller comparison.
Accordingly, this report makes no “first system” or “breakthrough” claim. Its more defensible contribution is the integration of several choices in one auditable program:
- a general abstraction corpus rather than a single source domain;
- a declared archetype-by-domain matrix;
- multiple diverse complete proposals per cell;
- preserved proposal, critique, revision, retrieval, and evaluation artifacts;
- proposal-specific searches for the problem, components, and full composition;
- separate quality, prior-art, practical, and empirical-partner gates;
- explicit negative tests, authority, safety, and resource estimates;
- frozen thresholds for later experiments; and
- publication of failures and resource records alongside survivors.
The distinction is architectural rather than absolute. Neighboring systems may contain individual pieces, and future comparison may reveal closer precedents.
The related-work review and candidate prior-art records were produced primarily through agentic web retrieval. That process favored digitally accessible, well-indexed material. TRIZ appeared in the companion review but was omitted from the first main-report synthesis, while Zwicky's older book-form precedent was missed until an adversarial review. These are observed retrieval-and-synthesis failures, not merely hypothetical limitations. Older books, non-English literature, trade practice, proprietary systems, and work described under different vocabulary remain plausible blind spots. Accordingly, ADJACENT_PRIOR_ART is always bounded by the documented searches; it is not a claim that closer precedent does not exist.
4. How the program evolved¶
The thirteen linked experiments were not thirteen replications of one frozen protocol. They were an iterative research program. Each experiment answered a narrower question exposed by the preceding work. Four diagnostic analyses and one sentinel audit were added to probe report-level vulnerabilities; they are explicitly post hoc and do not change any frozen verdict. Experiments 9 and 11–15 were frozen and executed prospectively; Experiment 7, although blinded, benchmarked a selector against already-known Experiment 6 outcomes and is therefore retrospective; the E9 substrate follow-up was designed only after E9 generation had finished and remains post hoc despite prespecified blinded coding.
| Experiment | Main question | Scale | Principal finding | Consequence |
|---|---|---|---|---|
| 1 | Does adding mechanism context produce a large general lift? | 60 cells × 3 conditions | No large consistent omnibus lift; context effects were heterogeneous. | Replace the broad comparison with a tighter paired design. |
| 2 | Do relevant mechanisms beat archetype-only and irrelevant mechanisms under iterative propose–criticize–revise–test? | 20 cells × 3 arms | Relevant mechanisms won 15/20 against each comparator and passed corrected paired tests. | Treat mechanism context as useful, then test at matrix scale. |
| 3 | Can five archetypes generate researched opportunities across all 64 domains? | 320 cells | Closed-book success was common; external prior art eliminated or narrowed most assessed selections. | Move external scrutiny earlier and make it proposal-specific. |
| 4 | Should retrieval precede or follow proposal construction? | 20 paired cells | Proposal-first yielded 5 strict arms; retrieval-first yielded 1; uncertainty remained large. | Keep proposal-first generation, then scrutinize complete proposals. |
| 5 | How many diverse complete proposals should a cell receive? | 20 cells × 5 proposals | Four strict candidates appeared in positions 1–4; position 5 added none. | Replicate four-proposal portfolios on new archetypes and domains. |
| 6 | Do four proposals rescue enough strict cells in a larger generalization set? | 60 cells × 4 proposals | Five rescues versus a frozen floor of six: NOT_SUPPORTED; 15 strict proposals and 32 partner candidates still resulted. |
Distinguish scientific threshold failure from practical yield and test cheaper selection. |
| 7 | Can proposal-only judgments prefilter four candidates without losing valuable ones? | 60 cells, 180 blinded rankings | Primary top-two rule retained 9/15 strict proposals and failed all frozen thresholds. | Do not replace prior-art scrutiny with proposal-only ranking. |
| 9 | Is coarse productivity distributed beyond the archetypes selected for earlier experiments? | 50 archetypes × 3 fixed domains | All 24 probability-sampled untested generated archetypes were productive; 60/72 random-sample cells survived the light screen. | Treat breadth as supported under a coarse endpoint; do not equate it with strict novelty or deployability. |
| 11 | Does Experiment 2's relevant-mechanism advantage survive uniform external research? | 20 cells × 3 preserved arms | R lost 9–11 to S, beat D 12–8, and failed the pooled and omnibus frozen gates. | Retain the internal-quality result but withdraw any general researched-yield advantage claim. |
| 12 | Can the generator leave governance/computation when those substrates are prohibited, and does Max effort rescue the cost? | 72 matched High-effort cells; triggered 18-cell 2×2 Max follow-up | 67/72 constrained outputs complied, but ordinary won blinded quality 68–4 and survived the light screen 60–49; Max did not preferentially rescue the constrained lane. | Treat dominant substrates as a soft but consequential default; use alternative-substrate generation for bounded diversification, not wholesale replacement. |
| 13 | If one additional proposal can be funded, should it use ordinary diversity or deliberately seek another causal substrate? | 12 probability-sampled untested archetypes × 5 probability-sampled domains; 60 common P1s and 120 paired P2s | Ordinary P2s produced 41 incremental survivors and substrate P2s 32; ten were substrate-only, 19 ordinary-only, and blinded quality was 32–28. | Keep ordinary diversity as the default and use an alternative-substrate P2 as a bounded complement when unique coverage merits the measured yield cost. |
| 14 | Can explicit applicability routes improve problem-to-archetype retrieval and reject structurally incomplete near misses? | 40 stratified archetypes × 3 constructed cases; 114 eligible after blinded audit | Diagnostic Recall@5 was 2/77 versus 8/77 for the solution index; target-route sensitivity was 77/77 and near-miss specificity 32/37. | Preserve DNF verification, but redesign route-aware candidate retrieval and validate it on independently sourced cases. |
| 15 | Does condition-level DNF aggregation improve candidate retrieval on new archetypes and cases? | 40 new stratified archetypes × 3 constructed cases; 115 eligible after blinded audit | Primary route-aware Recall@5 was 1/78, versus 6/78 for the solution index and 12/78 for diagnostic maximum-hit: NOT_SUPPORTIVE. |
Reject raw cosine aggregation as DNF retrieval; investigate signed evidence and route-length/count calibration without changing the E15 verdict. |
The post-hoc prior-art sentinel, Experiment 8 yield decomposition, Experiment 10 proposal-type analysis, E10B score-to-substrate join, and E9 substrate follow-up reuse existing proposals or outcomes. Their methods were frozen before the relevant searching, outcome join, or new substrate labels, but the questions themselves were chosen after seeing the main program. They are sensitivity and diagnostic analyses, not new confirmatory experiments.
This sequence matters when interpreting apparent contradictions. Experiment 6's negative primary verdict does not mean later proposals had zero value; it means the observed rescue count did not meet a threshold set in advance. Conversely, the existence of 59 dossier candidates does not retrospectively turn every experiment positive.
5. Unified method and protocol differences¶
Common anatomy of a candidate¶
Later protocols required a proposal to identify a problem, actors, observable state, consequence, affected objective, intervention, baseline, causal chain, structural mapping, nearest rivals, authority and safety constraints, and negative tests. External evaluation then searched for evidence of the problem, stakeholder pull, implementation feasibility, and prior art. It assigned nine 1–5 opportunity scores and rough 2026 USD resource-equivalent bands for first evidence, initial deployment, operational launch, and recurring cost.
The nine common dimensions were meaningful impact, stakeholder pull, incremental advantage, distinctiveness plausibility, technical implementability, adoption/authority feasibility, evidence readiness, safety/net benefit, and scalability. A high score did not authorize deployment. The required output was a next evidence step.
Generate, criticize, revise, test¶
Experiment 2 formalized the iterative loop rather than treating ideation as one shot. Each trajectory could be criticized, revised, and retested up to five revisions after the original proposal. A controller stopped when the candidate passed, failed without useful progress, required unavailable research under the closed-book rule, or exhausted the revision budget. This better approximated the iterative nature of human innovation while preserving bounded cost and an inspectable trajectory (protocol; analysis plan).
Experiments 4–6 shifted effort from repeatedly repairing a single weak idea toward generating several substantially different complete proposals and researching each one. That change separated within-proposal refinement from between-proposal search.
Prior-art scrutiny¶
Search was deliberately contrastive. It did not ask only whether the proposal's name appeared online. Evaluators searched for:
- the target problem and affected stakeholder;
- each load-bearing component;
- the full composition or causal sequence;
- the strongest realistic baseline;
- implementation and authority constraints; and
- evidence capable of resolving the remaining distinction.
The external labels were conservative. ESTABLISHED_PRACTICE and SUBSTANTIAL_COLLISION removed a proposal from strict consideration. ADJACENT_PRIOR_ART meant that public search left a narrower, testable contrast; it did not mean no prior art. NO_CLOSE_MATCH_FOUND was permitted by the schemas but did not occur among the 59 dossiers.
Strict and empirical-partner lanes¶
The strict lane favored candidates whose problem, comparison, implementation, and next test could be defended using the available record. Experiment 6 added a disjoint empirical-partner lane for cases where public evidence could not answer the decisive question but a named kind of partner plausibly could. The lane was calibrated on seven cases; all seven adjudications matched the calibration labels, although the small all-agreement set could not provide a meaningful chance-corrected reliability estimate (calibration result).
Post-hoc dossier ordering¶
For this report only, the 59 candidates are ordered with the three opportunity profiles first used after Experiment 2: balanced, deployment-heavy, and impact-heavy (method). The nine common ordinal scores are reused. A tenth input, pilot speed/cost, is approximated from the first-evidence cost band: under $10,000 = 5; \(10,000–\)50,000 = 4; \(50,000–\)250,000 = 3; \(250,000–\)1 million = 2; above $1 million = 1. Because elapsed pilot time was not independently scored, this is an affordability proxy, not a full pilotability measure.
For each profile, the weighted 1–5 mean is multiplied by 20 to create the familiar 0–100 display; because the underlying minimum is 1, the numerical floor is 20 rather than zero. The balanced profile supplies the reading order; rank ranges across the three profiles expose sensitivity. Ranks 1–10, 11–25, 26–45, and 46–59 are shown as broad bands. These values are an editorial navigation aid. They neither alter original endpoint labels nor measure economic value.
Reproducibility and preservation¶
From Experiment 2 onward, the project increasingly preserved schemas, prompts, source snapshots, raw model events, controller states, judgments, design freezes, completion manifests, analysis scripts, and resource accounting. The artifact index links the canonical materials. The public harmonization script does not rewrite experimental data; it reads sealed analysis outputs and candidate artifacts into a 59-row JSONL index.
Units, samples, and denominators¶
The experiments distinguish several units that are easy to blur:
- an archetype is the source solution structure;
- a domain is the target field;
- a cell is one archetype–domain pairing;
- a proposal is one complete problem–intervention hypothesis in a cell;
- a version is the original or a revision of that proposal;
- a trajectory is the sequence of versions, critiques, tests, and controller decisions; and
- a candidate instance is a proposal retained by a specified endpoint.
Experiment 3's 33/320 is a cell yield under a closed-book adopter-pipeline rule. Experiment 6's 15 strict proposals belong to 11 cells. The latter can be reported as 15/240 proposals or 11/60 cells, but the fractions answer different questions. The empirical-partner count is proposal-level and disjoint from strict proposals. This report gives denominators whenever practical and avoids adding outcomes from unlike levels.
Most samples were purposive rather than probabilistic. Experiment 1 balanced five archetypes across domain groups. Experiment 2 selected 20 informative cells for a paired condition test. Experiment 3 enumerated the project's complete 5 × 64 declared matrix but did not thereby sample the universe of real problems. Experiments 4 and 5 selected new cells to test pipeline changes. Experiment 6 deliberately used five different archetypes crossed with 12 domains to probe generalization beyond the original five. Experiment 7 reused all 60 Experiment 6 cells because it was a blinded retrospective selector study. Experiment 9 added a seeded probability sample of 24 previously untested generated archetypes from a 1,092-item corpus, but crossed them with only three purposively fixed domains. Experiment 11 reused Experiment 2's 20 selected cells and 60 terminal proposals to isolate the effect of equal external scrutiny. Experiment 12 reused E9's probability-sampled 72-cell stratum and sealed ordinary outputs as matched controls for a prospective prompt intervention; its domains therefore retain E9's purposive limitation. Experiment 13 excluded all E9-exposed archetypes, probability-sampled 12 previously untested generated archetypes, independently sampled five domains outside the E6/E9/E12 domain union, and crossed them to form 60 new cells. This improves both sides of the sampling design within the declared corpus and domain taxonomy, but still does not sample real-world problems. Experiment 14 hash-sampled eight unique route-bearing archetypes from each of five prespecified Applicability Graph strata. Experiment 15 repeated that stratified sampling design after excluding all 40 E14 archetypes. Both experiments deliberately constructed 120 cases from the graph rather than sampling problems from the world; this supports controlled route tests but sharply limits external-validity claims.
Model sessions, role separation, and blinding¶
Later runs used fresh or retained sessions according to the role. In Experiments 5 and 6, one retained session per cell generated the proposal portfolio so the model could avoid repeating its earlier ideas; external evaluation used fresh candidate-version sessions so the evaluator did not inherit the generator's discussion. Experiment 7 used fresh proposal-only sessions and blinded candidate identifiers. Experiment 9 used one fresh generator and one fresh light-screen evaluator for each cell. Experiment 2 concealed arm identity from quality judges and revealed keys only after the judgments were sealed. Experiment 11 preserved that concealment: fresh evaluators researched opaque proposals, two fresh judges compared the researched arms within each cell, a third resolved disagreement, and arm keys were revealed only after adjudications were sealed. Experiment 12 used two fresh blinded substrate-classification passes and two fresh randomized A/B quality passes, with separate blinded adjudication of every disagreement before private keys were joined. Its web screens used the unchanged E9 protocol. Experiment 13 used isolated fresh calls for P1 and both P2 arms, then replaced proposal, pair, arm, archetype, and domain identities with opaque keys. Two fresh passes classified substrate and P1–P2 independence, two randomized A/B passes compared P2 quality, fresh calls adjudicated every disagreement, and 180 separate opaque evaluators applied the unchanged four-source screen before key reveal. Experiment 14 used fresh isolated case-writing calls, two independent opaque case-audit passes, fresh adjudication of all material disagreements, deterministic retrieval from frozen indexes, and one fresh opaque DNF verifier per eligible case. The verifier never saw archetype names, target identity, or intended case type; private keys were joined only in analysis. Experiment 15 used the same fresh case-writing and duplicate blinded-audit procedure on a disjoint sample, sealed eligibility before deterministic retrieval, and made no model call during ranking.
This is role separation, not full independence. The same model family sometimes occupied several roles, and the model may have encountered similar source material during training. Independent sessions reduce direct conversational leakage but not shared priors. The exact runtime settings, service tier, timeouts, retry budgets, and worker counts are preserved in run manifests linked from the artifact index.
Validation, retries, and nonresponses¶
Model outputs were constrained by JSON Schemas and checked by deterministic code. Transport or schema failures could be retried within declared limits; substantive revisions followed the controller rules rather than silently replacing an unfavorable judgment. Run records preserve prompts, responses, events, errors, elapsed time, and exposed token telemetry. Experiment 3's external tranche retained one terminal nonresponse after three 1,800-second timeouts instead of imputing a favorable or unfavorable answer.
Amendments corrected transport and processing issues without changing sealed scientific fields. Completion seals and manifests record hashes or references for the applicable design and evidence. The presence of a valid JSON object guarantees structural conformance, not truth; substantive claims still depend on sources and judgment.
Statistical interpretation¶
The program used simple statistics matched to the small paired or binomial questions. Experiment 2 used paired direction counts and sign tests with Holm correction. Experiment 4 used the paired McNemar test and reported the absolute arm difference. Experiments 5 and 6 reported marginal cell rescue with Wilson intervals. Experiment 7 reported proposal and cell recall with Wilson intervals, along with fixed threshold pass/fail decisions. Experiment 9 reported Wilson intervals for sampled-archetype productivity and cell survival. Experiment 11 used paired direction counts, Holm-adjusted one-sided sign tests, and an exact within-cell label-randomization omnibus test. Experiment 12 reported Wilson intervals, exact paired sign or McNemar tests, and a 50,000-draw paired bootstrap clustered by archetype; its small Max subset remained descriptive. Experiment 13 reported the paired survivor table, Wilson intervals, a 20,000-draw archetype-clustered bootstrap, and an exact sign-flip test over 12 archetype-level summed differences; cell-level McNemar and sign tests remained sensitivity analyses. Experiment 14 reported Recall@k, reciprocal rank, a paired exact McNemar test, a 20,000-draw archetype-clustered bootstrap, sensitivity, specificity, and balanced accuracy. Experiment 15 reused Recall@k, reciprocal rank, paired exact McNemar tests, and 20,000-draw archetype-clustered intervals, with one confirmatory route-aware arm and four prespecified but non-rescuing sensitivity arms. These analyses quantify sampling uncertainty conditional on their stated sampling frame; they do not repair domain-taxonomy, model, or evaluator dependence.
Frozen thresholds were treated as decision rules, not as natural constants. A rule such as “at least six rescued cells” embodies a prior judgment about what would justify the next scale-up. Reporting the observed count and interval remains necessary when the rule fails. Exploratory comparisons are labeled as such and are not allowed to replace the frozen primary verdict.
Provenance of load-bearing decision thresholds¶
The table distinguishes a threshold being present before execution from a public preregistration. These were internal design freezes in a rapidly adapting solo research program. Where the record preserves no contemporaneous numerical derivation, the report says so rather than supplying a retrospective rationale as though it had been frozen.
| Experiment | Frozen rule | Timing and evidence | Preserved rationale and limitation |
|---|---|---|---|
| 1 | Condition C had to exceed A by a paired median of at least 10 points, alongside diversity, safety, and survivor conditions. | Present in the experiment specification used before execution; no separate signed design-freeze manifest was created. | The rule demanded a large, practically visible omnibus gain before scaling. The record does not preserve a more detailed power or cost derivation for 10 points. |
| 5 | At least four rescued cells, at least three additional strict proposals beyond P1, and at least one strict proposal at P3–P5. | Sealed at 2026-08-03T04:38:21Z, before scientific execution, in the design freeze. |
The rule required both cell rescue and evidence that search depth beyond the first two proposals mattered. It was a minimum-practical-effect rule, not a significance test. |
| 6 | At least six rescued cells, at least six strict proposals beyond P1, and at least one new strict cell at P4. | Frozen before execution in the method and hashed design manifest; execution authorization was recorded at 2026-08-03T11:41:12Z. |
Experiment 5 rescued 3/19 P1 failures (15.8%); applying that observed rate to Experiment 6's 54 P1 failures would project about 8.5 rescues. A floor of six was below that projection. This numerical comparison explains the rule after the fact; the sealed method itself does not record that derivation. |
| 7 | The top-two policy had to retain at least 9/11 strict cells, 25/31 broad cells, 11/15 strict proposals, and 33/47 broad proposals. | Design frozen at 2026-08-04T01:00:01Z; all 180 selector outputs were frozen before label join at 2026-08-04T02:02:41Z (prediction freeze). |
These were high-recall safeguards: saving half the searches was not sufficient if the selector discarded too much known yield. Designers knew aggregate Experiment 6 outcomes, so this was a blinded retrospective policy benchmark, not a prospective preregistration. |
| 9 | At least 10/24 random generated archetypes productive, their productive fraction no more than 0.20 below the reference block, and at least two random-sample survivors in each fixed domain. | Frozen before execution in the design freeze; a documented pre-run correction repaired the candidate-pool implementation without changing the scientific question. | These were broad-distribution floors at a deliberately light screening endpoint. They were not novelty or strict-opportunity thresholds. |
| 11 | R had to be net-positive against both S and D, reach at least 65% pooled pairwise preference, pass an exact omnibus test at p ≤ 0.05, and avoid excess safety stops. |
Protocol, analysis, blinded inputs, and arm key were sealed in the design freeze before external evaluation. | The joint gate demanded a large, consistent researched-record advantage rather than a marginal difference in strict counts. It was intentionally stronger than simply repeating Experiment 2's internal score comparison. |
| 12 | Run Max only if at least 24/72 constrained outputs complied and either constrained decisive quality wins were at most 40% or ordinary usable yield led by at least 15 points. | Method, treatment prompt, inputs, schemas, and the 18-cell Max subset were sealed before treatment generation. A post-run audit correction records that the mutable status README was mistakenly included in the original hash set; all load-bearing design sources still verify. | The first clause distinguished instruction failure from a viable but difficult task. The second allocated expensive Max calls only when a meaningful cost appeared. It was a resource trigger, not a confirmatory hypothesis test. |
| 13 | Prefer the substrate P2 only for a gain of at least five points and more unique survivors; otherwise call it a bounded complement only with at least 48/60 compliant proposals, at least six substrate-only survivors, and a yield deficit no worse than 15 points. | The sampling frame, selected cells, prompts, schemas, endpoints, statistics, and decision rule were sealed before any E13 model output in the design freeze. | The rule required at least 10% unique coverage while capping the cost of diversification. The observed deficit landed exactly on the permitted 15-point boundary, so the operational gate passed but is not a robust equivalence result. |
| 14 | Diagnostic Recall@5 had to exceed solution Recall@5 by at least ten points, produce at least eight more diagnostic-only successes, pass exact McNemar at p ≤ 0.05, and avoid a loss greater than five points in either positive subtype. DNF verification separately required sensitivity and specificity of at least 0.70 and balanced accuracy of at least 0.75. |
The method, sample, graph and MCP snapshots, prompts, schemas, endpoints, statistics, and gates were sealed before any E14 scientific model output in the design freeze. | The retrieval rule demanded a practically decisive improvement before replacing the existing index. The verification rule tested whether explicit conjunctions added structural discrimination rather than merely semantic resemblance. Retrieval failed; verification passed. |
External evidence record¶
Later external evaluations required eight structured sources per terminal candidate, including the problem, stakeholder, implementation, prior-art, and resource evidence needed for the scoring record. Source objects preserve title, publisher, URL, source class, publication or coverage date where available, access date, and the claims supported. Evaluators also preserved emitted queries. Source count is a coverage discipline, not a guarantee of eight independent or equally strong facts.
Search used public web services only where the user explicitly authorized sending candidate summaries and derived queries. The public report includes those URLs, candidate text, and project records; it does not knowingly include a partner's private operational dataset. Future partnered studies may introduce human-subject, proprietary, security, or regulated data obligations and must obtain the appropriate institutional, legal, and participant approvals before data transfer or intervention.
6. Results by research question¶
Does structural mechanism context improve proposal quality?¶
Experiment 1 provided the first caution. Sixty cells were run in three conditions and reviewed, but the prespecified condition-C versus condition-A paired median gain was 3.12 points, below the frozen 10-point threshold. Hard-gate and promotion counts moved in the expected direction—34/60 and 22/60 in condition A, 38/60 and 27/60 in B, 41/60 and 30/60 in C—but the evidence did not support a large, consistent omnibus lift. The protocol also carried context-length and review-blinding limitations described in its quality audit.
Experiment 2 narrowed the question and improved the controls. Each of 20 cells had three trajectories: relevant mechanism context (R), solution-archetype context without the relevant mechanisms (S), and an equal-length irrelevant-mechanism decoy (D). All three used the bounded propose–criticize–revise–test loop. Terminal relevant-mechanism proposals later passed a post-hoc quality entry gate in every cell. On paired blinded judgments, R beat S in 15 cells, lost in four, and tied in one; it beat D in 15 and lost in five. The Holm-adjusted sign-test values were 0.01921 and 0.020695 respectively (machine-readable results).
Experiment 11 subjected the same 60 terminal proposals to a stronger endpoint. Treatment labels were hidden while one fresh external evaluator per proposal searched six lanes and retained exactly eight sources; two fresh judges then compared the three opaque researched records within each cell, with a third judge used for nine disagreement cells. Relevant mechanism context (R) beat archetype-only context (S) in 9 cells and lost in 11; it beat the irrelevant control (D) in 12 and lost in 8. Pooled R preference was 52.5%, below the frozen 65% floor, and the exact omnibus randomization result was p = 0.4459. Strict researched counts were close—12/20 for R, 11/20 for S, and 10/20 for D—and no arm produced a safety or authority stop (Experiment 11 results).
The combined conclusion is narrower than the earlier report. Experiment 2 remains evidence that relevant mechanisms improved internal terminal quality under its rubric. Experiment 11 did not support an advantage after external scrutiny and comparative judgment. It does not show that mechanisms are useless—the R-versus-D direction remained positive, and the interval is wide—but it prevents the internal-quality result from being generalized into a researched-survivor claim.
Can the method operate across a declared matrix?¶
Experiment 3 crossed five solution archetypes with 64 domains, producing 320 cells. Its internal funnel was permissive by later standards: 297 of 320 cells completed the first generation stage successfully, including 284 of the 300 cells in the prospectively designated tranche. Thirty-three of 320 passed the complete closed-book pipeline.
That result showed operational scale, but it overstated defensible opportunity yield because the pipeline had deliberately withheld external research. Forty-eight candidates were then selected for external assessment; 47 returned an assessment and one did not. Among the assessed set:
| Prior-art disposition | Candidates | Share of 47 assessed |
|---|---|---|
| Substantial collision | 25 | 53.2% |
| Established practice | 15 | 31.9% |
| Adjacent prior art | 6 | 12.8% |
| Indeterminate | 1 | 2.1% |
| No close match found | 0 | 0% |
Twenty-one of 47 passed the externally verified pipeline as it was then defined, but only two combined verified-pipeline status with the later strict-like prior-art condition. Because that two-candidate composite was constructed post hoc, the report preserves the candidates while refusing to treat “2/320” as a preregistered success rate. The more general lesson is firm: internal coherence and closed-book quality cannot substitute for retrieval.
Experiment 9 asked a different matrix question: whether even coarse opportunity productivity was confined to the archetypes selected by the researchers. It applied the same one-shot proposal and four-source light screen to all 16 hand-curated archetypes, a seeded probability sample of 24 previously untested generated archetypes from a 1,092-item generated corpus, and the ten earlier archetypes as a same-protocol reference block. An archetype counted as productive if at least one of its accounting/auditing, chemistry/materials, or computer-science cells survived.
All 24 random generated archetypes were productive (Wilson 95% interval 86.2%–100%), and 60 of their 72 cells survived (83.3%; 73.1%–90.2%). The fixed-domain counts were 23/24 in accounting/auditing, 16/24 in chemistry/materials, and 21/24 in computer science. The hand-curated census produced 16/16 productive archetypes and 34/48 surviving cells; the earlier-archetype reference produced 9/10 and 20/30. All three frozen breadth conditions passed (Experiment 9 results).
This unusually high rate must be read against the endpoint: among the random cells, the screener assigned 60 ADJACENT_PRIOR_ART, 11 SUBSTANTIAL_COLLISION, and one ESTABLISHED_PRACTICE. A four-source light screen is meant to eliminate obvious failures, not establish novelty, economic value, or deployment readiness. Experiment 9 supports broad coarse researchability across the sampled generated corpus and three purposive domains. It does not estimate strict opportunity yield over the full Cartesian product.
Should retrieval come before or after proposal construction?¶
Experiment 4 compared two arms within 20 cells. The proposal-first arm created a complete problem–intervention hypothesis and then searched for the problem, rivals, components, and full combination. The retrieval-first arm searched broadly first, generated 100 short hypotheses, screened 53 out, subjected 47 to a critic, and developed 15 finalists.
Proposal-first produced strict success in 5 of 20 arms; retrieval-first produced 1 of 20. Thirteen of the 15 retrieval-first finalists collided substantially during full evaluation. The paired difference favored proposal-first by 20 percentage points, but McNemar's exact test was not significant at conventional levels (p = 0.2188). The sample therefore supplied a design recommendation, not a stable effect estimate.
Why might proposal-first help? Broad retrieval can anchor the search on already-named problems and established solution vocabularies. A complete proposal gives research a contrastive object: the evaluator can ask whether this precise composition exists and whether it improves on this specific rival. The result does not imply that retrieval should be delayed until the end. It supports proposal construction followed immediately by scrutiny, with revision after evidence.
How many complete proposals should a cell receive?¶
Experiment 5 generated five diverse complete proposals in each of 20 new cells. One strict success appeared at each of positions 1, 2, 3, and 4; position 5 added none. Portfolio success therefore rose from 1/20 cells after proposal 1 to 4/20 after proposal 4, a three-cell rescue among 19 first-proposal failures. The 95% Wilson interval for the rescue rate was wide, 5.5%–37.6%, so “four rather than five” was a replication choice rather than a universal optimum.
Experiment 6 tested that choice in 60 new cells created from five new archetypes and 12 domains. Its strict yield curve was:
| Proposal position | Strict proposals at position | Newly successful cells | Cumulative strict-success cells |
|---|---|---|---|
| 1 | 6 | 6 | 6/60 (10.0%) |
| 2 | 2 | 1 | 7/60 (11.7%) |
| 3 | 1 | 0 | 7/60 (11.7%) |
| 4 | 6 | 4 | 11/60 (18.3%) |
The four-proposal portfolio contained 15 strict proposals in 11 cells. It rescued five of the 54 cells that failed at proposal 1, a 9.3% rescue rate (95% Wilson interval 4.0%–19.9%). The frozen rule required at least six rescues and also specified later-position conditions; five rescues missed the primary floor by one. The correct confirmatory verdict is NOT_SUPPORTED.
That verdict and the practical calculation answer different questions. Scientifically, the prespecified replication threshold failed. Operationally, researching proposals 2–4 found nine additional strict proposals and five additional strict cells. A program that values a missed candidate may still rationally pay for later positions. The data do not determine that value; they make the tradeoff visible.
Can public evidence resolve every promising proposal?¶
No. In Experiment 6, 32 proposals met the empirical-partner criteria and were disjoint from the 15 strict proposals. The combined broad lane contained 47 proposals across 31 of 60 cells. Later positions added 19 of those 31 cells: the combined cell curve rose from 12/60 after proposal 1 to 31/60 after proposal 4.
This broad yield must not be interpreted as 51.7% innovation success. The partner label means that a decisive observable—often an internal error rate, workflow trace, operational baseline, or counterfactual comparison—was unavailable in public sources. These candidates are better understood as structured requests for collaboration. Their value is that they specify what a partner would need to measure and what result should terminate the idea.
Can a cheap selector reduce research cost without losing the good candidates?¶
Experiment 7 retrospectively hid Experiment 6 outcomes and asked three fresh proposal-only selectors to rank the four candidates in every cell. Their rankings showed meaningful agreement: mean pairwise Spearman correlation was 0.634 and the median was 0.8. Yet consensus did not supply the needed validity:
| Consensus rank | Strict candidates | Strict-or-partner candidates |
|---|---|---|
| 1 | 3 | 12 |
| 2 | 6 | 13 |
| 3 | 2 | 8 |
| 4 | 4 | 14 |
Rank 2 contained the most strict candidates, while rank 4 contained the most broad-lane candidates. Downstream survival was not monotonic in proposal-only rank, and the highest-ranked proposal was not reliably privileged.
The frozen primary policy selected the top two proposals per cell—120 of 240 proposals. It retained:
- 9/15 strict proposals (60.0%; 95% Wilson interval 35.7%–80.2%);
- 8/11 strict cells (72.7%; 43.4%–90.3%);
- 25/47 strict-or-partner proposals (53.2%; 39.2%–66.7%); and
- 23/31 strict-or-partner cells (74.2%; 56.8%–86.3%).
All four frozen recall thresholds failed, so prospective confirmation was not warranted. A proposal-only selector may still order a queue or support triage when misses are acceptable. It cannot, on this evidence, replace the expensive search stage while claiming to preserve the opportunity yield.
Did a stronger prior-art search overturn the sampled survivors?¶
The post-hoc sentinel audit hash-selected ten candidates across Experiments 3–6 and ran four targeted search lanes per case: historical lineage, trade or standards practice, non-English terminology, and the full proposed composition. It found closer adjacent precedent for 8 of 10 candidates and no closer precedent for two. None met the audit's rule for a likely substantial collision, and all ten remaining contrastive claims survived.
This result is reassuring only in a narrow sense. It does not validate all 59 dossiers or estimate a corpus-wide false-novelty rate. It shows that the sampled labels were robust to this extra search while also demonstrating that the first-pass records often did not contain the strongest neighbors. Future strict scrutiny should make the four search lanes explicit before assigning a novelty-like disposition (audit method and case results).
Was observed yield concentrated in one solution archetype?¶
Experiment 8 decomposed existing outcomes by archetype and domain without changing the original endpoints. Across 13 protocol–outcome strata, seven met the frozen diagnostic for breadth across tested archetypes and six were concentrated. Strict success appeared under 2 of 5 archetypes in Experiment 5 and 4 of 5 in Experiment 6; Experiment 6 empirical-partner candidates appeared under all 5. The 11 strict-success cells in Experiment 6 were somewhat concentrated by the frozen diagnostic, whereas its 15 strict proposals were broadly distributed.
The defensible conclusion is mixed, protocol-dependent breadth. No single archetype explains the program, but sparse events and purposive samples make archetype rankings unstable. The selected 47-candidate Experiment 3 external subset is not a denominator for the full 320-cell matrix, and domain rows cannot be read as domain effects (Experiment 8 report).
Did proposal substrate or evidence dependency track disposition?¶
Experiment 10 assembled the 422 primary externally evaluated proposals from Experiments 3–6 and removed their outcomes before classification. Two fresh passes agreed on proposal substrate for 399/422 cases (94.5%; Cohen's κ = 0.9004) and evidence dependency for 373/422 (88.4%; κ = 0.8284); a third blinded pass adjudicated the 68 rows with any disagreement.
The generated mix was highly uneven: 208 proposals (49.3%) were governance/process interventions and 189 (44.8%) were computational/information interventions; only 25 fell into the other three substrate classes. Yet strict-success rates for those two large classes were similar across the pooled heterogeneous records—13/208 and 13/189—and Experiment 6's exploratory governance-versus-other empirical-partner comparison was also inconclusive (odds ratio 1.27; two-sided Fisher p = 0.572). Evidence dependency was dominated by private operational, laboratory/technical, and field/human-institutional tests, but small counts and protocol differences do not support a universal evidence-class ranking.
This is a diagnosis of what the pipeline generated and where evidence bottlenecks occurred, not a causal test of which proposal type is intrinsically more innovative. It weakens the simpler story that the portfolio exists mainly because evaluators favored governance ideas (Experiment 10 report).
Did archetype character explain the proposal substrate?¶
The motivating internal note initially omitted Deadweight Loss Reduction from the ten archetypes used in Experiments 3–6 and therefore overstated their mean structural/framed score as +0.462. The corrected ten-archetype mean is +0.366, close to the +0.374 corpus reference. E10B then used the ten archetypes—not 422 repeated proposals—as its inferential units. Across the 340 complete-census version-zero proposals from Experiments 5 and 6, source score was associated with governance share within the governance/computational channels (Spearman ρ = −0.666; exact exploratory permutation p = 0.0422; leave-one-archetype-out range −0.822 to −0.601), but not with escape from those channels (ρ = −0.058; p = 0.85).
The stronger E9 follow-up retained all 150 proposals from 50 archetypes crossed with three fixed domains and removed archetype, domain, sampling, score, screen, and prior-art fields before new classification. Two passes agreed on 140/150 labels (93.3%; Cohen's κ = 0.885), and a third blinded pass adjudicated all ten disagreements before the key was joined. Across all 50 archetypes, source score versus governance allocation had ρ = −0.624 (seeded permutation p = 0.00001); in E9's probability-sampled 24-archetype stratum, ρ = −0.791 (p = 0.00002). Each domain's direction was negative, while 139/150 proposals (92.7%) remained governance/process or computational/information.
Together these results support a two-level substrate constraint in the tested pipeline: archetype character modulates which dominant channel appears, but does not explain why the pipeline remains concentrated in those two channels. This is an association within the complete bundle of corpus, prompt, model, rubric, and protocol—not evidence that the base model has an intrinsic substrate limitation. The question was selected after E9 finished, so prespecified blinded coding strengthens the result without making it prospective (combined analysis and validation).
Can the pipeline be forced to leave governance and computation?¶
Experiment 12 converted the post-hoc substrate observation into a prospective prompt intervention. It reused the 72 random-generated E9 cells and their sealed ordinary/High proposals as matched controls. A fresh gpt-5.6-sol High-effort call for each cell had to make the primary causal intervention physical/material or measurement/instrumentation; governance and computation could appear only as removable support. Two independent blinded classifier passes and adjudication found that 67 of 72 constrained outputs complied: 40 were physical/material and 27 measurement/instrumentation. Two proposals disguised governance as a physical intervention and three abstained.
Compliance did not imply equal quality. In randomized blinded A/B comparison, the ordinary proposal won 68 cells and the constrained proposal won four, with no ties (p = 4.62 × 10⁻¹⁶, exact two-sided sign test). After the authorized four-source web screen, 49 constrained proposals and 60 ordinary controls survived. The paired difference was −15.3 percentage points; an archetype-clustered bootstrap interval ran from −30.6 to 0 points. Among the 27 discordant cells, eight cells survived only under the constraint and 19 only under ordinary generation (p = 0.0522, exact paired McNemar/sign test).
The cost was uneven. Constrained and ordinary light-screen yield tied at 16/24 in chemistry/materials, while the constraint trailed 14–21 in computer science and 19–23 in accounting/auditing. Framed archetypes produced only 6/12 constrained survivors versus 11/12 ordinary survivors, with no constrained quality win. Measurement/instrumentation was a frequent escape route—27 proposals and 21 survivors—but produced only one constrained quality win. These subgroup counts are descriptive because domains were purposive and the experiment was not powered for multiple heterogeneity tests.
The frozen trigger launched an 18-cell, four-arm Max follow-up. Max beat High 14–4 for ordinary proposals and 12–6 for constrained proposals, but ordinary beat constrained 16–2 at both effort levels. Light-screen survival was 12/18 ordinary/High, 11/18 constrained/High, and 9/18 for each Max arm. Higher effort improved the judged articulation and coherence of both arms without preferentially rescuing substrate denial or increasing this coarse researched yield.
The causal conclusion is bounded but important. Governance and computation are not a hard expressive ceiling: the tested generator can leave them when instructed. They are a consequential default, because forced alternatives were usually more contrived or less faithful and survived external screening less often. The eight constrained-only survivors show portfolio-diversification value, so the practical response is a limited alternative-substrate lane rather than a blanket ban (Experiment 12 report and artifacts).
Which second proposal should a search portfolio buy?¶
Experiment 13 turned Experiment 12's diversification suggestion into a prospective policy comparison on 60 new cells. Twelve probability-sampled previously untested generated archetypes were crossed with five independently sampled domains. Each cell received a common first proposal, an ordinary-diverse second proposal, and a second proposal whose primary causal substrate had to be physical/material or measurement/instrumentation. All calls were isolated and closed-book during generation. Two blinded classification passes, two blinded P1–P2 independence passes, two randomized P2 quality passes, fresh adjudication of every disagreement, and 180 opaque four-source web screens were completed before the private keys were joined.
The substrate instruction was fully effective: all 60/60 proposals complied, comprising 24 physical/material and 36 measurement/instrumentation interventions. Ordinary generation remained concentrated in the earlier house style: 54/60 ordinary P2s were governance/process or computational/information. Yet the sharp E12 quality gap did not recur in this more appropriate second-slot comparison. Blinded quality favored ordinary P2s 32–28, with no ties (p = 0.699, cell-level exact sign-test sensitivity).
The researched incremental endpoint still favored ordinary diversification. A P2 counted only if blinded judgment found it complete, structurally faithful, materially independent from P1, and its light screen survived; substrate P2s also had to comply. Ordinary P2s yielded 41/60 (68.3%) incremental survivors and substrate P2s 32/60 (53.3%). The paired table was 22 both, 10 substrate-only, 19 ordinary-only, and 9 neither. The substrate-minus-ordinary difference was −15.0 points; the archetype-clustered 95% bootstrap interval was −31.7 to +3.3 points, and the exact sign-flip sensitivity test over 12 archetype aggregates gave p = 0.1953. Those uncertainty results do not establish equivalence or a stable population difference.
The unique-survivor result is the reason the answer is not simply “always buy ordinary P2.” P1 survived in 43 cells. Among its 17 failures, ordinary P2 rescued five and substrate P2 rescued eight. Across all cells, ten researched opportunities appeared only in the substrate lane. The frozen policy therefore returned USE_AS_BOUNDED_PORTFOLIO_COMPLEMENT: compliance exceeded 48, substrate-only survivors exceeded six, and the yield deficit was no worse than 15 points. The last condition passed exactly at its boundary. This is evidence for a deliberately limited diversification lane under the tested policy—not for replacing ordinary diversity or for claiming that any light-screen survivor is a validated invention (Experiment 13 report and artifacts).
Can the Applicability Graph retrieve and verify matching archetypes?¶
Experiment 14 tested the new trigger-logic layer on a controlled benchmark. Forty unique route-bearing archetypes were hash-sampled across five strata: single-literal, fully grounded prime-only, fully grounded routes containing a domain-specific abstraction, partially open routes, and alternative-route archetypes. A fresh case writer produced a direct positive, a cross-domain positive, and a one-literal near miss for each. Two opaque audit passes plus fresh adjudication retained 114/120 cases—77 positives and 37 near misses spanning all 40 archetypes—before retrieval began.
The existing solution-oriented index retrieved the intended archetype in its top five for 8/77 positives (10.4%). The diagnostic maximum-hit search retrieved it for 2/77 (2.6%), a difference of −7.8 percentage points. The paired table contained two shared successes, zero diagnostic-only successes, six solution-only successes, and 69 failures by both arms (p = 0.03125, exact two-sided McNemar). The diagnostic arm failed every directional part of its frozen superiority gate. Its target appeared anywhere in the top 50 for only 13/77 positives, versus 22/77 for the solution index.
Verification told a different story. Each eligible case received an opaque literal-by-literal assessment of the target's complete DNF routes; when retrieval had missed the target, a shadow copy permitted fidelity measurement but was barred from the reranking pool. The verifier accepted 77/77 positives and correctly rejected 32/37 near misses, yielding 93.2% balanced accuracy and passing all frozen verification floors. Because the diagnostic top 12 contained the target for only 4/77 positives, reranking could not repair retrieval.
The frozen verdict was therefore VERIFICATION_ONLY. Within this graph-derived benchmark, explicit routes are useful for deciding whether an archetype under consideration actually fits. The current maximum-single-hit search is not an effective way to decide which archetype to consider. Four of five false acceptances occurred in the partially open stratum, an exploratory localization suggesting that incompletely grounded conditions need particular curation. The benchmark establishes internal operational fidelity only: its cases came from the graph, and related model-family systems constructed, audited, and verified them (Experiment 14 report and artifacts).
Does explicit route aggregation repair candidate retrieval?¶
Experiment 15 tested the next design without tuning it against E14. It excluded
all 40 E14 archetypes, hash-sampled eight new archetypes from each of the same
five strata, and froze five condition-level aggregation formulas before any new
case existed. Every formula scored conditions separately, combined them within
a route, and took the best alternative route. The confirmatory
SENTENCE_LOWER_HALF arm matched every condition against the full scenario and
deterministic sentence-like segments, then averaged the lowest half of the
condition scores. Four minimum or mean variants were sensitivity arms and
could not rescue the verdict.
The same outcome-blind case procedure retained 115/120 cases—78 positives
and 37 one-literal near misses across all 40 archetypes—before retrieval. The
primary arm retrieved only 1/78 positives (1.3%) in its top five. The
solution index retrieved 6/78 (7.7%), and diagnostic maximum-hit retrieved
12/78 (15.4%). Against the solution baseline, the primary arm had zero
exclusive successes and five baseline-only successes (p = 0.0625); against
diagnostic maximum-hit, it had one exclusive success and 12 baseline-only
successes (p = 0.00342). It also exceeded the allowed five-point loss in both
direct and transfer cases. The frozen verdict was NOT_SUPPORTIVE.
| Frozen E15 arm | Recall@1 | Recall@5 | Recall@10 | Recall@50 |
|---|---|---|---|---|
| Solution-oriented index | 2/78 | 6/78 | 8/78 | 23/78 |
| Diagnostic maximum-hit | 4/78 | 12/78 | 15/78 | 22/78 |
| Sentence-aware lower-half mean (primary) | 0/78 | 1/78 | 1/78 | 19/78 |
| Sentence-aware route mean (secondary) | 0/78 | 2/78 | 7/78 | 33/78 |
The secondary sentence-mean arm's 42.3% Recall@50 shows that condition-level similarity contained some broad candidate signal, but its 2.6% Recall@5 was not operationally competitive and cannot change the primary result. Baseline order also reversed across E14 and E15: maximum-hit lost in E14 and led in E15. Descriptively pooling the two disjoint benchmarks gives both baselines 14/155 top-five successes, with nine exclusive successes apiece. This was not a prespecified pooled endpoint; it shows instability and low absolute recall, not equivalence.
A clearly post-hoc failure localization found two mechanisms. Single-literal routes occupied 378/390 (96.9%) primary-arm top-five slots on positive cases, demonstrating severe route-length and alternative-count bias. More fundamentally, cosine similarity did not distinguish entailment from explicit contradiction. Among 36 eligible matched transfer-positive/near-miss pairs, the primary arm ranked the contradicted near miss better in 22 and the positive better in 14; the median near miss ranked 28.5 places higher. Arithmetic over raw similarity scores therefore did not implement logical conjunction. A future design would need signed condition evidence and calibration for route length, route count, and sentence-search multiplicity. Those are hypotheses, not retroactive repairs to E15 (Experiment 15 report and artifacts).
7. Negative results and what they changed¶
Negative evidence is central to this program because fluent generation makes false confidence cheap.
The expected large context gain did not appear¶
Experiment 1 rejected the idea that progressively richer context would produce a large uniform score gain across the calibration matrix. That failure led to an equal-length irrelevant control and paired iterative trajectories in Experiment 2. The improvement came from making the causal question narrower.
The researched context comparison did not confirm Experiment 2's advantage¶
Experiment 11 applied the same external research budget and blinded comparative endpoint to all 60 Experiment 2 terminal proposals. Relevant mechanisms lost 9–11 to archetype-only context, beat the irrelevant control 12–8, and reached only 52.5% pooled preference; the exact omnibus result was p = 0.4459. The internal-quality finding remains part of the record, but it cannot now be described as an externally researched endpoint advantage. This negative replication is consequential because it narrows one of the program's most tempting causal claims.
Substrate denial changed outputs but did not preserve their average quality¶
Experiment 12 rejected both extreme interpretations of the earlier substrate pattern. The model was not trapped: 67/72 outputs complied with the non-governance, non-computational rule. But greater effort did not erase the cost, and the constrained lane remained far behind in blinded quality at both High and Max. This changes the engineering target. The next problem is not stronger wording of the prohibition; it is learning when an alternative substrate is structurally appropriate and generating it without turning the source archetype into a decorative physical analogy.
Alternative-substrate diversity did not replace ordinary diversity¶
Experiment 13 gave the substrate lane the more defensible role suggested by Experiment 12: an additional proposal rather than a replacement. It generated ten unique researched survivors and rescued more P1 failures, but ordinary P2 still produced nine more incremental survivors overall. The constrained lane met the frozen bounded-complement gate only at the maximum permitted deficit. The result supports portfolio routing, not a general claim that deliberately unusual substrates are equally productive. A future implementation should therefore make this lane optional and budget-aware rather than mechanically spending it on every cell.
Applicability verification did not solve candidate retrieval¶
Experiment 14 prevented a strong verifier result from being mistaken for an end-to-end system result. The target route was classified accurately when supplied, yet the diagnostic search failed to place that target in the top five for 75 of 77 positive cases and underperformed the older solution-oriented index. Route logic and candidate generation are separate engineering problems. This failure motivated the prospectively frozen condition-level aggregation test in Experiment 15.
Route-aware similarity did not implement route logic¶
Experiment 15 rejected the most direct embedding-only repair. Five frozen aggregators represented conditions separately and respected the route graph's AND/OR shape, yet the confirmatory arm fell below both baselines. The post-hoc diagnostic explains why the design was structurally faithful in form but not in semantics: uncalibrated aggregation favored short routes, and cosine similarity treated a contradiction as highly related evidence rather than negative evidence. The engineering target is now narrower. Candidate retrieval needs an entailment-sensitive or otherwise signed condition scorer and explicit opportunity-count calibration; changing the averaging rule alone is not enough.
Closed-book success was a poor proxy for external distinctiveness¶
Experiment 3's 33 internally successful cells looked encouraging until external search found established or substantially colliding practice for 40 of 47 assessed selections. This was not wasted work. It revealed that prior art had to shape the scrutiny pipeline rather than merely annotate its output.
Retrieval-first search did not improve yield in the tested form¶
Experiment 4's retrieval-first arm consumed search and screening effort before a complete contrastive proposal existed, yet yielded only one strict candidate. The negative result does not condemn research-first workflows generally. It rejects this particular broad-hypothesis implementation as the default scaling path.
The four-proposal replication missed its frozen rescue floor¶
Experiment 6 missed by one cell. Changing the threshold after observing five rescues would erase the value of having frozen it. The report therefore keeps the negative verdict and separately reports the actual marginal yield. This is an example of why “not supported” should not be translated into “useless.”
Proposal-only judgment was reliable without being valid enough¶
Experiment 7's selectors often agreed with one another, but their preferred proposals did not monotonically correspond to researched survival. A candidate can read as unusually specific or useful while hiding a prior-art collision; another can read as awkward while containing a defensible empirical contrast. Agreement among language models cannot substitute for an external criterion.
No candidate cleared a no-close-prior-art label¶
All 59 dossiers sit beside prior art. This narrows the interpretation from invention ex nihilo to recombination, transfer, governance composition, and comparison. It is also consistent with a mature world in which useful primitives are widely distributed. The unresolved question is often whether a particular arrangement or control produces an advantage in a new setting.
The sentinel audit makes this negative result more informative. Eight of ten sampled dossiers acquired closer neighbors under explicitly historical, trade, non-English, and compositional searches even though none crossed the audit's substantial-collision rule. Surviving a label is not the same as having retrieved the best precedent.
8. The surviving opportunity portfolio¶
What is in the portfolio¶
The candidate appendix contains 59 experiment-level instances:
| Source | Endpoint | Dossiers |
|---|---|---|
| Experiment 3 | Post-hoc strict innovation-like survivor | 2 |
| Experiment 4 | Strict success | 6 |
| Experiment 5 | Strict success | 4 |
| Experiment 6 | Strict success | 15 |
| Experiment 6 | Empirical-partner candidate | 32 |
| Total | 59 |
These are not 59 independently verified inventions. They are 59 proposals that earned one of the listed statuses in their own protocol and have enough preserved evidence to support specialist review. Canonical titles were checked for exact duplication; none were identical. Conceptual overlap can still exist across candidates, and expert reviewers may merge or split opportunity families differently.
The appendix intentionally does not absorb every attractive earlier output. Experiment 2's 20 relevant-mechanism proposals passed a quality entry gate and received a post-hoc opportunity assessment, but they predated the later strict researched-candidate endpoint. Experiment 3 had 21 externally verified adopter-pipeline candidates, of which only two met the later post-hoc strict-like composite. Those broader sets remain available through the artifact index; silently mixing them into the 59 would make the portfolio look larger by changing its definition.
How to read a dossier¶
Each dossier begins with a short alias and preserves the original title. It then explains the problem, intervention, structural transfer, reason for advancement, nearest approaches, remaining contrast, smallest decisive test, resource bands, risks, and questions for appropriate experts. Direct links lead to the original proposal and evaluation.
The appendix is ordered by a post-hoc balanced score, with sensitivity to deployment-heavy and impact-heavy profiles. Near ranks should be read as ties. A strict endpoint remains strict wherever it appears; a partner candidate does not become strict because it ranks highly. Conversely, a low-ranked dossier is not disproven—its combination of evidence burden, implementation difficulty, or modest expected impact merely makes it a lower-priority expert review under the chosen weights.
Recurring opportunity forms¶
Although the domains vary widely, several families recur:
- Identity and lifecycle governance: preserve historical records while expiring their authority for current decisions.
- Representation-independent contracts: maintain a decision-relevant meaning across formats, tools, organizations, or physical transformations.
- Residual monitoring: model what should happen, then treat structured deviations as evidence requiring investigation.
- Bounded rivalry: use controlled competition among methods or candidate explanations without allowing the tournament itself to damage the system.
- Ritualized commitment and closeout: make transitions, obligations, uncertainty, or recovery explicit through repeated governed observances.
- Computability and assurance boundaries: distinguish what a system can prove, what depends on outside oracles or human choices, and what remains unknown.
These recurrences are not automatically new mechanisms missing from the Encyclopedia. Some are domain-specific instantiations of documented archetypes; some combine familiar components. The corpus may nevertheless reveal where the Encyclopedia needs a better domain example, a more precise composition, or a candidate mechanism for later curation.
What expert review should decide¶
The next reviewer is not being asked, “Do you like this idea?” A useful review should decide:
- whether the problem occurs at a consequential frequency or scale;
- whether the proposed baseline is the strongest fair comparator;
- whether a close system was missed;
- whether the intervention changes the causal structure as claimed;
- whether the test can be run with valid data and authority;
- what safety, legal, ethical, and distributional constraints were omitted; and
- whether to reject, revise, merge, observe, or run the bounded next test.
The expert review template turns those questions into a comparable record.
9. Resource use and scale economics¶
Measured Experiments 9 and 11–15 resources¶
Experiment 9's breadth probe made 300 valid transport attempts—one generator and one four-source light screen for each of 150 cells. It recorded 52,876,667 input tokens, of which 40,308,736 were cached, 1,252,720 output tokens, 376,597 reasoning-output tokens, and 40,022 seconds (11.1 hours) of summed agent time. With concurrent workers, its proposal and screen phases took approximately 3 hours 44 minutes of execution wall time. The experiment's relatively high throughput reflects its deliberately coarse endpoint.
Experiment 11 made 231 transport attempts to complete 60 eight-source external evaluations, 40 primary comparative judgments, and nine tiebreak judgments. It recorded 90,356,643 input tokens, of which 75,299,072 were cached, 1,169,103 output tokens, 466,307 reasoning-output tokens, and 32,682 seconds (9.1 hours) of summed agent time. Ninety-one attempts were invalid or failed and three lacked exposed usage, largely because a transport adapter rejected otherwise repairable structured outputs; accepted scientific artifacts still passed the canonical validator, and both recovery amendments are preserved. These figures show why a tightly controlled comparison can be expensive even when it generates no new proposals.
Experiment 12 completed 281 scientific model calls with no failed transport calls. It recorded 44,256,895 input tokens, of which 29,580,032 were cached, 1,337,531 output tokens, 660,600 reasoning-output tokens, and 43,522 seconds (12.1 hours) of summed call time. The 102 authorized web screens consumed most input tokens, while the 36 Max proposal calls accounted for 4.1 summed hours and unusually large constrained-output reasoning. This reinforces the broader cost finding: producing a different-looking proposal is cheaper than determining whether it remains distinct and useful after research (resource summary).
Experiment 13 completed 458 scientific calls with no failed transport calls: 180 proposal calls, 90 batched blinded-measurement calls, eight disagreement-adjudication calls, and 180 authorized opaque web screens. It recorded 80,429,441 input tokens, of which 54,841,600 were cached, 1,715,925 output tokens, 618,283 reasoning-output tokens, and 51,426 seconds (14.3 hours) of summed call time. Its 13.3-hour first-to-last telemetry span includes the generation phases, a pause for payload-specific web authorization, and the final screen run; it is not continuous active execution time. The comparison demonstrates the cost of measuring a portfolio policy properly: the 120 competing P2s were only one part of the workload, while blinding, duplicate measurement, adjudication, and complete web scrutiny supplied the evidence needed to interpret their incremental value (Experiment 13 results).
Experiment 14 completed 215 scientific model calls with no failed final calls: case construction, duplicate blinded case audit plus adjudication, and 114 blinded DNF verifications. It recorded 3,584,592 input tokens, of which 1,880,832 were cached, 912,450 output tokens, 236,247 reasoning-output tokens, and 18,501 seconds (5.14 hours) of summed call time. Three concurrent workers reduced its first-to-last telemetry span to 1.79 hours. The much lower input total reflects local deterministic index retrieval and the absence of public-web research; this was a controlled applicability benchmark, not a prior-art screen (Experiment 14 results).
Experiment 15 recorded 83 successful model-response attempts to obtain 40 accepted case bundles, 24 accepted blinded-audit batches, and two accepted adjudication batches. Seventeen responses were rejected by schema or deterministic validation and retried; there were no failed model transports in the canonical run. It recorded 1,329,255 input tokens, of which 457,728 were cached, 255,170 output tokens, 110,461 reasoning-output tokens, and 5,618 seconds (1.56 hours) of summed call time. Three concurrent workers reduced the first-to-last telemetry span to 32.5 minutes. All seven retrieval arms were local and deterministic, and no public-web research was used. The lower cost came from not repeating E14's 114 per-case DNF verifier calls (Experiment 15 results).
Measured Experiment 6 resources¶
Experiment 6 is the best basis for scaling calculations because it used the four-proposal pipeline on 60 cells and recorded exposed usage for all 591 scientific calls. Summed call time is not the same as elapsed calendar time because three workers ran concurrently.
| Resource | Recorded amount |
|---|---|
| Scientific calls | 591 |
| Input tokens | 205,117,798 |
| Cached input tokens | 165,694,464 |
| Uncached input tokens | 39,423,334 |
| Output tokens | 5,433,719 |
| Reasoning output tokens | 1,782,582 |
| Summed call time | 102,071 seconds (28.4 hours) |
External evaluation dominated the recorded workload: 243 calls, 25.3 million uncached input tokens, 2.22 million output tokens, and 16.7 summed hours. Proposal generation itself grew across positions because later calls had to preserve and avoid earlier ideas. The full breakdown is in resource accounting.
What Experiment 7 says about cost cutting¶
The three-replication selector used 180 calls, 3.57 million uncached input tokens, and 294,142 output tokens. Under the top-two policy it would have avoided evaluation of 120 proposals, but retained too few downstream positives. Net resource savings are therefore inseparable from the value assigned to false negatives. A selector that is inexpensive but discards six of 15 strict candidates is not automatically efficient.
One selector replication was much cheaper than three and may be appropriate for non-destructive queue ordering. Position-based rules are even cheaper. Neither has been prospectively shown to meet a high-recall requirement.
Scaling scenarios, not forecasts¶
The measured per-cell averages from Experiment 6 are approximately 9.85 scientific calls, 657,000 uncached input tokens, 90,600 output tokens, and 28.4 minutes of summed call time. A naïve linear projection to 70,000 cells would therefore imply roughly 689,500 calls, 46.0 billion uncached input tokens, 6.34 billion output tokens, and 3.8 years of summed call time. With three continuously occupied workers, the last figure is roughly 1.26 years before failures, rate limits, maintenance, and changing search conditions. Experiment 9 demonstrates a much cheaper breadth pass—two calls per cell—but its light-screen survivors are not interchangeable with Experiment 6 strict or empirical-partner outcomes. It is best understood as a possible allocation layer, not a validated substitute for deep scrutiny.
Those numbers are order-of-magnitude scenarios, not a budget quote. Caching behavior, context length, model and search pricing, concurrency limits, model upgrades, and protocol improvements would materially change them. Subscription usage also cannot be converted honestly into API expense without the applicable product contract. The defensible economic conclusion is simpler: exhaustive search is technically conceivable for a well-resourced organization, but external scrutiny—not raw idea generation—is the dominant cost, and the current evidence does not justify an immediate 70,000-cell run.
For an API deployment, a reader can insert the prices applicable to the chosen model and contract into the transparent approximation: 46,000 × uncached-input price per million tokens + 6,340 × output price per million tokens, then add cached-input charges, web-search or retrieval fees, storage, retries, engineering, monitoring, and expert review. This formula is deliberately not populated with a public list price that may not correspond to the experimental product or remain current when the report is read.
Where the cost actually sits¶
The projection above applies the Experiment 6 pipeline to every cell. No one would run it that way. The realistic shape is the two-stage funnel this program's own results suggest: an Experiment 9-grade breadth pass across the whole matrix, then the Experiment 6 pipeline on a selected fraction. Costing that shape from the same measured per-cell averages, and using a deep stage of 10,000 cells purely as a round illustrative figure:
| Stage | Cells | Scientific calls | Uncached input | Output + reasoning | Summed call time |
|---|---|---|---|---|---|
| Breadth pass, Experiment 9 protocol | 70,912 | 141,800 | 5.9 billion | 770 million | 5,260 hours |
| Deep pipeline, Experiment 6 protocol | 10,000 | 98,500 | 6.6 billion | 1.20 billion | 4,730 hours |
| Two-stage total | 240,300 | 12.5 billion | 1.97 billion | 9,980 hours | |
| Undifferentiated deep pass, for comparison | 70,912 | 698,500 | 46.6 billion | 8.53 billion | 33,510 hours |
Staging reduces uncached input by a factor of 3.7 and summed call time by 3.4. At three continuously occupied workers the two-stage total implies roughly 139 days of execution rather than the 1.26 years implied by the undifferentiated projection. That is the difference between an infeasible program and an expensive but unremarkable one.
The table conceals this program's largest unvalidated dependency. It assumes something reduces 70,912 cells to about 10,000, a retention near 14%. The Experiment 9 light screen does not do this: it passed 60 of 72 random-stratum cells, an 83.3% rate that would forward roughly 59,000 cells to deep scrutiny and recover almost none of the saving. Reaching a 14% retention requires a far stricter selector, and Experiment 7 is the only prospective evidence this program holds about cheap selection—where it failed all four prespecified retention requirements. The economics of a full-corpus run therefore rest on a component that has been tested once and did not work. Improving that selector is a cheaper research target than improving the generator, and it determines whether the rest of this arithmetic is reachable at all.
Even granting the funnel, the model-side figures are not the binding constraint. Suppose the deep stage is ranked and its strongest 500 candidates are sent to appropriate specialists. At 1.5 to 3 hours of genuine review per dossier—reading the evidence record, judging whether the comparator is the strongest fair one, and checking for a missed close system—that is 750 to 1,500 reviewer-hours, or roughly 19 to 38 person-weeks of specialist attention, distributed across the twenty or more domains the candidates occupy. For comparison, the initial Experiments 1–13 program, including its design freezes, blinded coding, adjudication, and first complete writeup, was executed in approximately six calendar days by one person, as the preserved run timestamps show; Experiments 14 and 15 were added later. Reviewing a selected 500 outputs would cost on the order of twenty times the human effort that produced that initial evidence base.
That is this section's central finding, and it inverts the intuition the token counts invite. Generation was never the expensive part. No experiment in this report was limited by compute; the limits were on human judgment, and the largest acknowledged limitation—that no candidate has been assessed by a domain expert—is the direct consequence. A reader deciding where to spend next should read the token tables as evidence that the cheap half of the problem is solved and the expensive half has not been started.
What the cost structure does not settle¶
It is tempting to close the loop with an expected-value argument: if one candidate in several hundred proved worth a large sum, a full run would repay itself. This report cannot support that argument, for two reasons that are unmeasured here rather than merely uncertain.
First, no candidate's value has been estimated. None has been reviewed by a domain expert, and every dossier sits beside adjacent prior art. The distribution of candidate value is not simply unknown—nothing in this program was designed to sample it.
Second, value realized is not value captured. These candidates propose interventions in other parties' domains, generally requiring their data, authority, and operating consent. Value created by an adopting organization does not accrue to the operator of the search absent a specific vehicle, such as a defensible claim, an operating entity, a consulting relationship, or a research collaboration—and universal adjacent prior art makes the first of those harder rather than easier. An economic case must therefore state a capture rate, not only a value estimate.
A third consideration cuts the other way, and is worth naming because it bears on selector design rather than on the budget. If the value of innovation opportunities is heavy-tailed, as it appears to be in adjacent settings such as patent-value and venture-portfolio distributions, then expected return is dominated by rare extreme outcomes rather than by the median candidate. Under that assumption the objective at scale is not to maximize average candidate quality but to maximize independent draws while avoiding filters that truncate the tail. Every selection stage in the present pipeline—the light screen, the strict lane, the post-hoc composite ordering—was tuned to remove weak candidates, and none was designed or tested for tail preservation. That assumption is imported from outside this program rather than established by it, and it is offered as a design consideration for future work, not as a finding.
10. Interpretation and implications¶
A search capability, not an invention oracle¶
The program demonstrates a way to turn a corpus of abstractions into a systematic search instrument. That matters because ordinary prompting often produces isolated analogies with no declared denominator, no preserved failures, and no way to distinguish transfer from lucky recall. Here, the system can be asked to cover a matrix, generate several alternatives per cell, and expose where candidates fail.
The severe qualification is equally important. The generator is best understood as a hypothesis proposer embedded in a research workflow. Its apparent creativity includes retrieval from pretrained knowledge, recombination of familiar components, and translation into new settings. The external evidence record repeatedly showed that fluency and specificity were compatible with rediscovery. The method becomes scientifically interesting not when the model says something surprising, but when a proposal survives a comparator-aware attempt to disprove it and still points to a feasible decisive observation.
Cross-domain transfer can be scaffolded¶
The experiments do not isolate an internal psychological process in the model. They show that transfer-like output can be scaffolded by an explicit source representation, a target domain packet, structural mapping requirements, diversity pressure, and iterative evaluation. Experiment 2's irrelevant-mechanism control is evidence that relevant structure contributed to quality. Experiments 3–7 then show that this contribution is insufficient without search and selection.
This suggests a useful engineering perspective. Rather than asking whether a model “has” analogical transfer as a single latent faculty, one can decompose the task into source representation, target retrieval, mapping, adaptation, evaluation, and evidence acquisition. Different tools—and humans—can own different stages. The Encyclopedia is especially valuable at the source-representation stage because it supplies explicit mechanisms and archetypes rather than requiring the model to infer every source structure from prose. Experiment 11, however, shows that this plausible scaffold cannot yet be credited with a general researched-endpoint advantage over archetype-only context.
The output may be valuable before deployment¶
A candidate can create value without becoming a startup, product, or intervention. It can:
- reveal a problem that a domain expert recognizes but has not formalized;
- supply a new comparator or measurement design;
- connect specialists who use similar structures under different names;
- identify an archival, governance, or safety control worth adding to an existing system;
- inspire a better human reformulation; or
- show that an apparent Encyclopedia gap is actually an established domain instantiation.
This is why expert response should preserve revisions rather than reduce every dossier to approve/reject. A domain specialist may move the idea to a more consequential setting, replace a weak baseline, simplify the intervention, or identify a known term that collapses the novelty claim while preserving practical value.
The most useful candidates may be prompts for co-creation¶
The “counterfactual discovery” discussion that emerged from the library recommendation candidate illustrates this possibility. The original proposal concerned feedback between discovery exposure and later recommendation signals in a digital library. A human reader immediately recognized a potentially wider relevance to music, film, e-commerce, and streaming systems, where prominent placement can manufacture some of the popularity later used as evidence of relevance. That extension is not an experimental result or a validated product idea. It shows how a structurally explicit candidate can become shared working material for a human who supplies context, priorities, and reframing.
The same interaction may occur in the opposite direction. Experts can identify why a transfer fails: the supposedly equivalent variables may not be manipulable, the authority structure may be wrong, or the proposed signal may be an artifact. Those failures are useful training and curation data if preserved.
11. Limitations and threats to validity¶
No independent domain-expert validation¶
The largest limitation is simple: the 59 dossiers have not been adjudicated by the diverse specialists needed to assess them. External sources establish that relevant problems, methods, and institutions exist; they do not substitute for tacit operational knowledge. No candidate should enter a live consequential workflow on the strength of this report.
Bounded search cannot establish absence¶
Search coverage varied with query construction, indexing, accessible language, paywalls, terminology, and the web's representation of practice. An evaluator can fail to find a close system that exists in a proprietary workflow, patent, non-English publication, local policy, or different vocabulary. ADJACENT_PRIOR_ART is therefore a statement about a bounded evidence record, not about the world.
Model-family and evaluator dependence¶
The program used related contemporary language-model systems for generation, criticism, retrieval orchestration, evaluation, and later editorial synthesis. Fresh sessions, blinded keys, deterministic validators, and external sources reduce some dependence but do not create fully independent human judgment. Models can share training data, stylistic preferences, blind spots, and an attraction to elaborate governance structures. Experiments 14 and 15 are especially exposed: the graph, constructed cases, audits, and E14 verifier all passed through related model-family systems. Future work should cross models and include human evaluators.
Adaptive research program¶
The thirteen linked experiments were designed sequentially in response to earlier results. That is appropriate for method development but creates researcher degrees of freedom across the program. Later experiments used design freezes and prespecified thresholds; earlier exploratory choices and the current 59-candidate synthesis should not be mistaken for one preregistered study. Twelve of the thirteen were prospective; Experiment 7 is the exception, and is described throughout as a blinded retrospective policy benchmark rather than a prospective preregistration. Experiments 9 and 11–15 were prospective tests added after the original seven-experiment arc. Experiment 12's question arose from the post-hoc substrate analyses, Experiment 13's policy comparison arose from E12, Experiment 14's retrieval benchmark arose after the Applicability Graph was built, and Experiment 15's frozen retrieval alternatives responded to E14's failure; their respective samples, inputs, endpoints, and decision rules were frozen before scientific generation or measurement.
Post-hoc hardening analyses¶
The prior-art sentinel, Experiments 8 and 10, E10B, and the E9 substrate follow-up were proposed after the relevant data or report questions existed. Their methods, samples, or blinded inputs were sealed before the relevant case work, outcome join, or new coding, which protects those steps from silent alteration. It does not make the questions prospective. The sentinel has only ten stratified cases; yield decomposition inherits purposive archetype and domain selection; the proposal-type analysis pools heterogeneous protocols; and the substrate analyses use a derived source score rather than a randomized treatment. E10B also followed an exploratory cross-tab. E9's balanced domain exposure and random 24-archetype stratum strengthen generalization within the declared frame, but its three domains remain purposive and its substrate question was chosen after generation. Experiments 12 and 13 supply prospective interventions on the bundle; they do not retrospectively make the observational question confirmatory or identify the base model as the sole cause.
Heterogeneous endpoints¶
The 59 candidates come from protocols with different samples, generation strategies, scrutiny stages, and endpoints. The report preserves their labels rather than estimating one pooled success probability. The post-hoc ranking harmonizes common dimensions for reading order only.
Candidate-instance dependence¶
Several proposals can come from one cell, archetype, domain, or conceptual family. Exact canonical titles are unique, but semantic independence has not been established. Counts describe candidate instances, not necessarily 59 unique mechanisms or market opportunities.
Sampling and domain taxonomy¶
The domains are a useful project taxonomy, not a universally accepted partition of human knowledge. Experiments sampled archetypes and domains for diversity, and Experiment 6 intentionally used a new set. Experiment 9 strengthens the archetype-side evidence by probability-sampling 24 previously untested generated archetypes, but its three domains were fixed purposively and its endpoint was a light screen. Experiment 12 inherits both strengths and limitations: its archetypes represent the sampled generated corpus, while its domain differences do not estimate all-domain effects. Experiment 13 independently probability-sampled both 12 previously untested generated archetypes and five domains outside the E6/E9/E12 union, strengthening generalization within the declared frames. Its 12 archetypes remain a small cluster sample, its five domains are not a probability sample of all human activity, and the cells are not sampled real-world problems. Experiments 14 and 15 sampled the declared route-bearing graph by stratum, but their scenarios were generated to instantiate those routes and therefore do not estimate performance on naturally occurring problems. Yield cannot be extrapolated to the full Cartesian product without modeling archetype, domain, proposal position, scrutiny depth, and selection effects.
Construct-derived applicability benchmark¶
Experiments 14 and 15 constructed positive and near-miss cases from the exact graph under test. Outcome-blind auditing, disjoint archetype samples, and label concealment reduce trivial leakage, and the near misses make the tests stricter than paraphrase recognition, but both benchmarks remain endogenous. E14's 93.2% balanced accuracy shows that the written predicates can be operationalized consistently by related models under controlled conditions; E15 shows that raw embedding similarity does not recover that entailment relation. Neither result shows that the predicates are complete, necessary, sufficient, or causally correct in real settings. Independently collected cases and specialist adjudication are required for that claim.
Prompt, corpus, and time dependence¶
The Encyclopedia was actively evolving, and experiments used source snapshots where specified. A more complete corpus may change the quality or specificity of transfer. Model behavior, search results, prices, and public prior art also change. Reproduction should preserve both the relevant corpus and a coverage date.
Rubric validity¶
The opportunity dimensions capture practical considerations but remain ordinal expert judgments produced by models. Weights reflect policy preferences. A high score can reward well-documented, governable ideas and disadvantage speculative high-upside science. Sensitivity profiles expose some but not all value judgments.
Cost measurement¶
Token telemetry was incomplete in some earlier experiments. Summed call time includes parallel calls and is not wall-clock duration. Cached and uncached tokens have different economic meanings. Cost bands are rough resource equivalents, not quotes or audited budgets.
No direct end-to-end baseline¶
Experiment 11 now provides an otherwise matched externally researched comparison of Experiment 2's relevant, archetype-only, and irrelevant-mechanism arms; it did not support a relevant-mechanism advantage under its frozen gate. The program still lacks an end-to-end comparison against ordinary free ideation, a different analogy library, or a human innovation team. The Encyclopedia's contribution to researched survivor yield, relative to those alternatives, therefore remains incompletely isolated.
12. Directions for human–machine research¶
1. Prospective expert evaluation¶
Recruit specialists for a stratified sample of strict and empirical-partner dossiers, including intentionally lower-ranked cases. Give them the original evidence record, not only the plain-language summary. Measure comprehension, missed prior art, problem prevalence, baseline adequacy, revision magnitude, recommendation, and inter-reviewer agreement. The most informative outcome is not an approval count alone but the transition from model proposal to expert-revised candidate.
2. Prospective field tests with stop rules¶
For candidates that retain support, run the smallest reversible study specified in the dossier. Preregister the comparator, outcome, authority, safety constraints, and termination thresholds. Treat a clean rejection as a successful test of the pipeline's falsifiability.
3. High-recall, cost-aware allocation¶
Experiment 7 ruled out the tested proposal-only top-two selector as a safe replacement for research. It did not exhaust sequential allocation. A future policy could perform a cheap search on all four proposals, use evidence-bearing signals to allocate deeper research, reserve one “wild-card” slot, and audit a random sample of discarded candidates. The objective should explicitly price false negatives rather than optimize call count alone.
4. Cross-model and corpus ablations¶
Repeat a frozen subset with different model families, no Encyclopedia context, archetype-only context, mechanisms without an archetype, and alternative analogy corpora. Hold the evaluator and retrieval budget fixed where possible. This would separate corpus value, prompt structure, model capability, and search intensity.
5. Compositional and higher-order problems¶
The present matrix mostly searched for first-order problems addressable by one archetype. Real systems often require several interacting solution structures—for example, a representation-independent contract plus residual monitoring and bounded rivalry. A compositional program could first diagnose several structural bottlenecks, then search a constrained graph of compatible archetypes. The combinatorial space is much larger, so composition should follow evidence of a multi-causal problem rather than enumerate every pair blindly.
This direction connects naturally to systems such as MOOSE-Star, which iteratively expands and filters research hypotheses. The Encyclopedia's possible advantage is breadth and explicit mechanism structure; the neighboring work offers ideas for tree search, novelty checks, and evidence-aware refinement. Comparative experiments should test that claim rather than assume it.
6. Synthetic data for transfer training¶
The preserved corpus can support a forward-looking training hypothesis. Each trajectory contains a source archetype, target domain, structural mapping, candidate problem, intervention, critique, revision, search evidence, comparator, disposition, and negative test. Positive and negative trajectories could teach several separable skills:
- retrieve a structurally relevant abstraction from a problem;
- infer a plausible target problem from a solution structure;
- distinguish relational transfer from surface analogy;
- adapt a mechanism without violating domain constraints;
- generate contrastive searches and falsifiers; and
- recognize collisions, uncertainty, and the need for external evidence.
This report does not show that such training will create a general cross-domain-transfer capability. A credible study would prevent leakage by holding out entire archetypes, domains, and archetype families; compare against equal-volume ordinary scientific and design data; evaluate both generation and rejection; and test whether gains transfer to human-authored problems outside the Encyclopedia taxonomy. Training examples should use preserved observable artifacts and concise rationales, not depend on undisclosed private chain-of-thought.
7. Encyclopedia curation feedback¶
Candidate families can be compared with the corpus to distinguish an absent mechanism from an absent example, a domain-specific instantiation, a recurring composition, or a naming problem. Only expert-reviewed patterns should enter the Encyclopedia. Failed mappings are also valuable: they can clarify an archetype's boundary conditions and improve future generation packets.
8. Human inspiration as an outcome¶
Measure whether a dossier helps a specialist produce a stronger idea than either the specialist or the model produces alone. A useful design would compare unaided experts, experts with ordinary brainstormed ideas, and experts with structurally mapped dossiers. Outcomes could include conceptual distance from the input, prior-art survival, test quality, revision effort, and the expert's own assessment of usefulness.
9. Public interface and living review¶
The eventual website can expose the report, raw data downloads, candidate filters, source records, and a versioned expert-review layer. Reviews should be attributable, dated, and non-destructive: new evidence can change a disposition without erasing what the pipeline originally produced.
10. Substrate-aware diversification¶
Experiments 12 and 13 support routing rather than prohibition. E13 supplied the first prospective second-slot policy test: the alternative-substrate lane added ten unique researched survivors but produced nine fewer incremental survivors overall and met its bounded-complement rule exactly at the allowed 15-point deficit. A future system should therefore generate ordinary diversity by default and route a limited alternative-substrate lane where missed-opportunity cost, natural substrate fit, or strategic portfolio breadth justifies the extra screen. The next study should test a cheap evidence-bearing router prospectively, preserve the ordinary P2 whenever recall matters, and prespecify archetype and domain strata rather than extracting favorable stories after the run.
11. Route-aware retrieval and external applicability validation¶
Experiments 14 and 15 identify candidate generation, not literal-level verification, as the immediate Applicability Graph bottleneck. E15 represented each condition separately and respected the route graph's AND/OR form, but raw cosine aggregation still failed because it favored short routes and could not distinguish satisfied from contradicted evidence. A next design should treat condition matching as a signed inference problem, calibrate scores across route length, route count, and query segmentation, and freeze those choices before another new sample. The stronger external-validity study should collect naturally occurring problems without using the graph to write them, ask independent specialists which routes genuinely apply, and measure both retrieval recall and false acceptance under that independent criterion. Given two low-recall internal benchmarks, external cases may be more informative than another large graph-generated run unless a materially different scorer exists.
13. Conclusion¶
The experiments began with an unusually simple reversal: start from a solution archetype and search backward for the problem it could solve. What made the project substantive was not the reversal alone, but the decision to operationalize it as a declared matrix and repeatedly try to break the resulting ideas.
The evidence supports cautious optimism. Relevant mechanisms improved internal terminal quality in Experiment 2, but Experiment 11 did not confirm an advantage after equal external research. Matrix-scale generation worked. Multiple complete proposals recovered candidates that a one-shot process missed. Experiment 9 found coarse researchability across every probability-sampled previously untested generated archetype, but its light screen cannot establish strict novelty or deployability. External search removed or narrowed most apparently promising ideas, and a cheap proposal-only selector failed to preserve enough of the researched yield. Post-hoc hardening exposed denser adjacent prior art, strong evidence-access bottlenecks, and a reproducible substrate house style. Experiment 12 then showed that this style is not a hard ceiling: 67/72 constrained outputs left governance and computation, although ordinary proposals retained a large quality and screen advantage. Experiment 13 tested the practical response on new sampled cells. Its substrate-diverse P2s added ten unique researched survivors and won 28/60 blinded quality comparisons, but ordinary P2s remained more productive, 41 incremental survivors to 32. Experiment 14 then showed that explicit applicability routes can support discriminating verification without automatically solving retrieval: the verifier passed its frozen gate, while the diagnostic candidate generator performed worse than the existing solution index. Experiment 15 tested the obvious condition-level aggregation repair on a disjoint frozen sample and rejected it: the primary route-aware arm found only 1/78 targets in its top five, below both baselines. The resulting system is therefore neither a novelty machine nor a reason to enumerate tens of thousands of cells without restraint. It is a configurable search-and-scrutiny pipeline that can produce concrete, falsifiable, evidence-linked opportunities for humans to judge, with a promising explicit verification layer and an unresolved candidate-retrieval bottleneck now localized to signed condition inference and score calibration rather than route representation alone.
The 59 dossiers are the testable public residue of that process. Their ultimate value will be decided outside this report: by experts who find missed precedents, by partners who supply unavailable measurements, by experiments that reject weak causal chains, and perhaps by a small number of cases that survive and become useful. Making that judgment possible—without hiding the failures—is the principal achievement documented here.
Data availability, authorship, and corrections¶
Data and code¶
The artifact index links the design freezes, prompts, schemas, raw records, source snapshots, analysis outputs, and scripts needed to audit each experiment and the post-hoc hardening work. The report's candidate index and fact table are generated from preserved artifacts by build_inverse_innovation_public_report_data.py. The Experiment 9 breadth probe, Experiment 11 external context comparison, Experiment 12 substrate-denial test, Experiment 13 second-slot policy test, Experiment 14 Applicability Graph test, and Experiment 15 route-aware retrieval test include prospective design freezes, raw records, and machine-readable results. The sentinel audit, yield decomposition, proposal-type analysis, and substrate analysis each include frozen methods or coding plans and machine-readable results. The plain-language dossiers are editorial derivatives; original proposals and evaluations remain authoritative.
Publication quality control¶
The publication QA register reports all eleven release gates from the approved report design as passed, partial, pending, or not applicable, with evidence for each status. Automated count, lineage, artifact-target, and local-link checks are distinguished from semantic review. In particular, valid URLs and complete source objects do not mean that every external citation has received independent claim-level verification.
Contributions and AI assistance¶
The human author of the Encyclopedia of Abstractions originated the inverse-innovation concept, developed the Encyclopedia that made the matrix possible, made the sequential research decisions, approved external search, selected interpretation boundaries, and retains responsibility for the report. OpenAI and Anthropic language-model systems were used, at different stages, as generators, critics, revisers, search orchestrators, evaluators, analysts, and drafting assistants. Deterministic scripts handled schema validation, aggregation, hashing, and many frozen decisions. Exact run-level model and prompt records are available where captured in the experiment artifacts.
This is an independent project report, not a statement by the author's employer, OpenAI, Anthropic, or any cited organization. The report has not been peer reviewed.
Revision history and corrections¶
The report's revision history is preserved in the repository. Corrections should add a dated entry to CHANGELOG.md, preserve the earlier report artifact when a release is sealed, and distinguish factual correction from changed interpretation. Candidate dispositions should be amended with new evidence rather than silently overwritten.
References and further reading¶
The literature review contains the annotated bibliography and search strategy. Particularly close starting points include:
- AskNatureGPT and biological-strategy retrieval: Design Science, DOI 10.1080/09544828.2025.2481536
- Patent function-based opportunity discovery: Technological Forecasting and Social Change, DOI 10.1016/j.techfore.2015.04.012
- Purpose–mechanism knowledge for analogical design: ACM Transactions on Computer-Human Interaction, DOI 10.1145/3530013
- TRIZ as abstraction-mediated inventive problem solving: AutoTRIZ and the Altshuller Foundation's historical description
- Morphological search and cross-consistency assessment: Zwicky (1967) and Ritchey (2015)
- Emergent analogical reasoning in transformers: arXiv:2605.11258
- MOOSE-Star, an agentic scientific-discovery framework: arXiv:2603.03756
- Reinforcement learning for analogical discovery: arXiv:2510.02263
For experiment-specific references, use the linked external evaluations in the candidate dossiers.