Revision history¶
Part of Inverse Innovation with the Encyclopedia of Abstractions · Revision history · Last revised August 2026
All notable report-package changes should be recorded here. Experimental artifacts are immutable; corrections to their interpretation should point to the original record.
August 2026 — prospective route-aware retrieval test¶
- Added Experiment 15, a prospectively frozen comparison of two baselines and five condition-level DNF retrieval aggregators on 40 new stratified archetypes, with no overlap with Experiment 14.
- Constructed 120 cases and preserved two independent outcome-blind audits plus fresh adjudication of all material disagreements. The sealed analysis retained 115 eligible cases: 78 positives and 37 one-literal near misses.
- Recorded the frozen
NOT_SUPPORTIVEverdict. The primary route-aware arm retrieved 1/78 positive targets in its top five, versus 6/78 for the existing solution index and 12/78 for diagnostic maximum-hit; neither paired gate nor either subtype noninferiority condition passed. - Added an explicitly post-hoc failure-localization analysis. Single-literal routes occupied 96.9% of primary-arm positive top-five slots, and contradicted near misses frequently ranked above their matched positives. These findings motivate signed condition inference and score calibration; they do not alter the frozen result.
- Added resource accounting for 83 successful model-response attempts, 66 accepted scientific outputs, 17 schema-rejected retries, and local deterministic retrieval with no public-web research.
- Integrated the method, negative result, failure localization, artifacts, validation, website chapter, and fifteenth experiment archive into the public package.
August 2026 — Applicability Graph retrieval and verification test¶
- Added Experiment 14, a prospectively frozen construct-based benchmark of the new Applicability Graph across 40 stratified route-bearing archetypes and 120 generated cases.
- Preserved two independent outcome-blind case audits, fresh adjudication of all material disagreements, sealed retrieval and verifier inputs, and 114 fresh opaque DNF verifications.
- Recorded the frozen
VERIFICATION_ONLYverdict: the graph accepted all 77 eligible positives and rejected 32/37 one-literal near misses (93.2% balanced accuracy), while diagnostic Recall@5 was 2/77 versus 8/77 for the existing solution-oriented index. - Added the explicit interpretation boundary that the graph currently works as an internal candidate verifier, not as an end-to-end problem-to-archetype retriever, and that graph-derived cases plus same-family model judgments do not establish external semantic truth.
- Added Experiment 14 facts, artifact links, validation, resource accounting, a website chapter, and a fourteenth experiment archive.
- Corrected the machine fact table's experiment-type scope: Experiment 7 now appears under a dedicated retrospective list rather than the prospective list. This aligns the data with the narrative correction already recorded below.
August 2026 — initial public report package¶
- Created the public report package and approved report design.
- Harmonized 59 candidate records from Experiments 3–6.
- Added a disclosed post-hoc dossier ordering with balanced, deployment-heavy, and impact-heavy sensitivity profiles.
- Drafted the main report, full candidate dossiers, artifact index, and expert review template.
- Replaced eight overlong editorial dossier drafts under hard field-length limits; the original and replacement records remain preserved.
- Passed the automated count, endpoint, prompt-leakage, artifact-target, and 510-link local audit.
- Status remains a working draft pending human-author review and independent domain-expert evaluation.
August 2026 — adversarial hardening pass¶
- Added TRIZ and Zwicky morphological analysis to the main contribution boundary and expanded the companion literature review.
- Disclosed the observed older-literature retrieval and synthesis blind spot without treating bounded web search as exhaustive.
- Replaced ambiguous Experiment 7
strict/broadcount pairs with an explicit table and corrected the false statement that rank 4 contained the most strict candidates. - Added provenance for the load-bearing Experiment 1, 5, 6, and 7 decision thresholds, distinguishing internal freezes, retrospective explanation, and prospective preregistration.
- Added a publication QA register that separates automated conformance from semantic and independent verification.
- Completed a frozen ten-case prior-art sentinel audit. Eight cases acquired closer adjacent precedent; two did not; no sampled remaining claim met the audit's likely-substantial-collision rule.
- Added Experiment 8 as an explicitly post-hoc decomposition of existing yield. Its 13 compatible strata support mixed, protocol-dependent archetype breadth rather than one universal winner.
- Added Experiment 10 as an explicitly post-hoc analysis of all 422 primary externally evaluated proposals from Experiments 3–6. Two fresh outcome-blinded classifiers and a third blinded adjudicator separated proposal substrate from evidence dependency before outcomes were joined.
- Integrated the hardening results into the abstract, executive summary, methods, results, limitations, conclusion, fact table, artifact index, and downloadable publication package without changing any primary experimental verdict.
August 2026 — prospective breadth and context extension¶
- Added Experiment 9, a frozen 150-cell archetype-breadth probe covering all 16 hand-curated archetypes, 24 probability-sampled previously untested generated archetypes, and ten reference archetypes across three fixed domains.
- Recorded Experiment 9's supportive broad-distribution verdict: all 24 random generated archetypes were productive at least once and 60/72 random-sample cells survived the four-source light screen. The report explicitly limits this to coarse researchability rather than novelty or deployment readiness.
- Added Experiment 11, a frozen externally researched replication of Experiment 2's 20-cell relevant-mechanism, archetype-only, and irrelevant-mechanism comparison.
- Recorded Experiment 11's negative frozen verdict: R lost 9–11 to S, beat D 12–8, reached 52.5% pooled preference, and produced an exact omnibus
p = 0.4459. - Revised the report's mechanism-context claim, resource accounting, limitations, conclusion, fact table, artifact index, website chapters, and experiment downloads without changing any earlier experimental record or the 59-dossier portfolio definition.
August 2026 — archetype substrate correction and follow-up¶
- Corrected the internal archetype-score analysis to include Deadweight Loss Reduction (−0.50). The ten experimental archetypes have mean +0.366, not +0.462.
- Added E10B as an explicitly post-hoc archetype-level join to the 422 Experiment 10 substrate labels. In the 340-proposal complete-census primary universe, source score tracked governance-versus-computational allocation but not escape from those two channels.
- Added a prespecified, outcome- and archetype-blinded substrate classification of all 150 completed E9 proposals. Two passes agreed on 140/150 labels (κ = 0.885), and all ten disagreements were adjudicated before key join.
- Recorded the supportive post-hoc two-level substrate result: all-50 ρ = −0.624, random-24 ρ = −0.791, and 139/150 proposals in governance/process or computational/information.
- Preserved the timing boundary: the E9 generation run was already complete when the substrate question was formulated, so the analysis remains post hoc even though its new coding and gate were frozen before labels existed.
- Integrated the finding, source correction, methods, limitations, validation record, artifact links, website chapter, and downloadable package without changing E9's prospective breadth verdict or the 59-candidate portfolio.
August 2026 — prospective substrate-denial intervention¶
- Added Experiment 12, a prospective paired intervention on the 72-cell probability-sampled E9 stratum that prohibited governance/process and computational/information as the primary causal substrate.
- Recorded 67/72 compliant constrained outputs, a 68–4 ordinary advantage in blinded quality, and four-source light-screen survival of 49/72 constrained versus 60/72 ordinary proposals.
- Preserved the paired portfolio decomposition: 41 both survived, eight constrained-only, 19 ordinary-only, and four neither.
- Ran the prespecified 18-cell Max follow-up after both cost clauses triggered. Max beat High under both constraints but did not preferentially rescue substrate denial; ordinary beat constrained 16–2 at both effort levels.
- Added phase-level accounting for 281 scientific calls and 102 explicitly authorized web screens, a standalone experiment report, deterministic validation, public-report facts, artifact links, a website chapter, and an Experiment 12 download archive.
- Documented that the original design seal mistakenly hashed the mutable status README. Its changed hash is not treated as design drift; all load-bearing frozen sources continue to verify.
August 2026 — prospective second-slot portfolio policy¶
- Added Experiment 13, a prospectively frozen 60-cell comparison of ordinary-diverse and alternative-substrate second proposals after a common P1.
- Probability-sampled 12 previously untested generated archetypes and five domains outside the E6/E9/E12 domain union, producing 180 isolated proposals.
- Preserved duplicate opaque substrate, independence, and quality judgments; fresh adjudication of every disagreement; and 180 explicitly authorized four-source web screens before arm-key reveal.
- Recorded the frozen bounded-complement verdict: ordinary P2 produced 41 incremental survivors, substrate P2 produced 32, and the paired table was 22 both, ten substrate-only, 19 ordinary-only, and nine neither.
- Documented that the substrate lane cleared its gate at the exact maximum permitted 15-point deficit. It adds unique portfolio coverage but does not replace ordinary diversity as the default.
- Added 458-call resource accounting, deterministic design and output validation, public-report facts, an Experiment 13 website chapter, and a thirteenth complete experiment archive.
August 2026 — editorial accuracy corrections¶
- Corrected the Experiment 12 phase decomposition in the standalone experiment report. The published split — 108 proposal-generation calls, 71 web screens, 66 primary blinded measurement or adjudication calls, 36 Max-follow-up measurement or adjudication calls — attached real recorded quantities to the wrong labels and coincidentally re-summed to the correct 281 total, which is why it survived arithmetic review. The recorded phase table gives 108 proposal-generation calls, 102 authorized web screens, 39 primary blinded measurement or adjudication calls, and 32 Max-follow-up blinded measurement or adjudication calls. The 102-screen figure already stated elsewhere in the same report, in the main report's resource section, and in this change log was correct throughout; no measured quantity changed and no verdict is affected.
- Corrected the abstract's description of Experiment 12 compliance. Sixty-seven of 72 constrained outputs complied with the prohibition; 60 of those used an unambiguously allowed substrate and seven were mixed with an allowed primary substrate. The abstract previously described all 67 as having "genuinely" left governance and computation, conflating the compliant total with the genuine-substrate subtotal in the experiment's own adjudication table. The main report body and this change log already used the correct term.
- Corrected the program-level experiment count. The abstract, the program-evolution section, the limitations section, and the package landing page described all eleven linked experiments as prospective. Experiment 7 is a blinded retrospective policy benchmark against already-known Experiment 6 outcomes, as the methods, threshold-provenance table, and results sections have always stated. The description is now ten prospective and one retrospective, with the exception named wherever the count appears.
- Not changed: the machine-readable fact table's
report_scope.prospective_experimentskey still enumerates Experiment 7 among its members. Renaming a key in a registered public download is a schema change and is deferred to a separate decision. - Added two Section 9 subsections deriving the program's cost structure from the already-published per-cell telemetry: a two-stage funnel projection, the observation that the funnel's assumed selector is the one component Experiment 7 tested and failed, and a reviewer-hour estimate showing that expert review of a selected 500 candidates would cost roughly twenty times the human effort that produced the entire evidence base. No new measurement was taken; the figures reproduce the Experiment 6 and Experiment 9 averages already reported. A companion subsection records why this cost structure does not support an expected-value argument, naming the unmeasured value distribution and the distinction between value realized and value captured.
- Repaired seven corrupted editorial risk bullets. Every instance was the fifth and final element of
risks_and_uncertainties, indicating a systematic failure of the field-length pass rather than seven independent drafting errors. Four had degenerated into filler padding, two of them terminating in the drafting model's own self-directed text ("must fix"), and three had absorbed fragments of an expert question plus a restatement of theexpert_typesfield. Each repair is a pure truncation: the retained sentence is a verbatim prefix of the original draft, with no rewording and nothing invented. The absorbed question fragments were verified to be already covered by properexpert_questionsentries, so no substantive content was lost. Affected records: EXP04 none; EXP05-STRICT-02; EXP06-STRICT-05, -10, -11, -12; EXP06-PARTNER-03, -16. The immutable raw drafting output underraw/dossier_drafts/is unchanged and remains authoritative for what the drafting model actually produced. - Recorded the QA implication rather than only the fix. Publication QA gate 4 reported that validators check field lengths but that no independent reader had performed sentence-level review of all 59 dossiers. These defects were visible to any reader of the affected sentence, which shows the gate's "PARTIAL" status was an understatement: for at least seven dossiers the human-reading layer had not been applied at all. The gate remains PARTIAL pending a full independent read.