Inverse Innovation with the Encyclopedia of Abstractions¶
Methods, negative results, surviving opportunities, and directions for human-machine research
Last revised: August 2026
Status: research-grade working paper; not peer reviewed
Companion materials: candidate dossiers · artifact index · expert review form · machine-readable data
Abstract¶
Can a system search for useful problems by starting with a reusable solution pattern? This report describes thirteen linked experiments—twelve prospective and one blinded retrospective policy benchmark—and five explicitly post-hoc analyses using the Encyclopedia of Abstractions as a structured source of solution archetypes and mechanisms. Instead of beginning with a known problem and asking for a solution, the pipeline paired an archetype with a domain, generated possible problems that could exhibit the archetype's causal structure, translated the archetype into interventions, and subjected the proposals to criticism, revision, external prior-art search, practical scoring, and—in later experiments—a separate lane for candidates that could only be resolved with an empirical partner.
The program produced evidence for a limited but meaningful claim. A contemporary large language model, when given explicit structural scaffolding and a sufficiently severe scrutiny pipeline, generated researchable cross-domain opportunity hypotheses in a repeatable matrix. The work does not show autonomous invention, world novelty, commercial value, or successful deployment. Those stronger claims would require independent domain experts, deeper searches, field data, and prospective tests.
Several results shaped that conclusion. In Experiment 2, relevant mechanism context outperformed both archetype-only context and an equal-length irrelevant-mechanism control on the prespecified paired internal-quality comparison. Experiment 11 repeated the same 20-cell, three-arm comparison after eight-source external scrutiny and did not confirm that advantage: relevant mechanisms lost 9–11 to archetype-only context, beat the irrelevant control 12–8, and achieved only 52.5% pooled preference (p = 0.446 on the frozen omnibus test). In Experiment 3, 33 of 320 closed-book cells passed the internal pipeline, but external scrutiny found an established or substantially colliding system for 40 of 47 assessed selections; only two survived a stricter post-hoc composite. Experiment 4 favored proposal-first over retrieval-first construction, and Experiments 5 and 6 showed that later proposals can recover opportunities missed by a first proposal, although Experiment 6 missed its frozen rescue floor by one cell. Experiment 7 showed that a cheap proposal-only selector could not safely replace downstream web scrutiny. Experiment 9 then supplied the first probability sample of the larger archetype corpus: all 24 sampled previously untested generated archetypes produced at least one light-screen survivor across three fixed domains, with 60 of 72 cells surviving. That endpoint is coarse researchability, not novelty or deployment readiness. Five explicitly post-hoc analyses found closer adjacent precedent for 8 of 10 sentinel candidates without collapsing their remaining claims, mixed rather than single-archetype yield, no general disposition advantage for governance/process proposals, and a two-level substrate pattern: archetype character shifted proposals between governance and computation while 139 of 150 balanced E9 proposals remained in those two channels. Experiment 12 intervened on that pattern. Sixty-seven of 72 constrained outputs complied with the prohibition on governance and computation—60 with an unambiguously allowed substrate and seven mixed but allowed-primary—but ordinary generation won 68–4 in blinded quality and retained a higher light-screen yield, 60/72 versus 49/72. Experiment 13 then tested the resulting portfolio policy on 60 new cells. Ordinary second proposals produced 41 incremental survivors and substrate-diverse second proposals produced 32; the latter nevertheless added ten opportunities that ordinary diversification missed and therefore met the frozen bounded-complement rule at its maximum permitted 15-point deficit. Experiment 14 separated applicability retrieval from route verification. The new graph achieved 93.2% balanced accuracy when verifying whether a supplied target archetype fit a graph-derived case, but its current diagnostic search retrieved the intended archetype in the top five for only 2/77 positive cases, versus 8/77 for the existing solution-oriented index. Experiment 15 then tested five frozen condition-level route aggregators on 40 new archetypes and 78 eligible positives. The primary aggregator retrieved only 1/78 targets in its top five, versus 6/78 for the solution index and 12/78 for diagnostic maximum-hit. Post-hoc localization found that single-literal routes occupied 96.9% of the primary arm's positive-case top-five slots and that contradicted near misses often ranked above matched positives. The graph is therefore a promising internal verification substrate, not yet an end-to-end problem-to-archetype retriever. The dominant channels are not a hard expressive ceiling, but ordinary diversification remains the stronger default in the tested bundle.
Across Experiments 3–6, this report preserves 59 candidate instances for human review: 27 strict or post-hoc strict-like candidates and 32 empirical-partner candidates. Every one has adjacent prior art. Their unresolved value lies in narrower compositions, governance arrangements, comparisons, or domain instantiations—not in components appearing from nowhere. The principal contribution is therefore a documented search-and-scrutiny method, a body of positive and negative evidence about its operation, and a transparent candidate corpus that specialists can accept, revise, reject, or use as material for further research.
Executive summary¶
The short version¶
The experiment began with a reversal: if an abstraction represents a recurring solution structure, perhaps the system can ask, “What important problem in this domain would have exactly this structure?” A solution archetype such as computability-boundary mapping can be crossed with fields as different as art conservation, engineering design, and public policy. The model must identify concrete actors and a measurable failure, map each part of the abstraction to the target situation, propose an intervention, state why ordinary alternatives are insufficient, and specify what evidence would prove it wrong.
That generation step is easy to make look impressive. The difficult part is distinguishing useful transfer from fluent renaming, rediscovery, and untestable speculation. The experimental program gradually moved its effort downstream:
- generate a structurally explicit proposal;
- criticize it and allow bounded revision;
- search for the problem, nearby methods, and the complete proposed composition;
- compare it with actual rivals rather than with a weak straw baseline;
- demand an operational next test, authority, stop conditions, and rough resource bands; and
- preserve failures rather than reporting only attractive survivors.
The strongest conclusion is not “the system invented 59 things.” It is that a general-purpose language model plus a structured abstraction library can populate a cross-domain search space with candidates that remain coherent after progressively stronger automated scrutiny. Some warrant specialist attention. Most raw ideas do not.
What was learned¶
- Structural context helped internal quality, but the stronger external replication did not confirm an endpoint advantage. Experiment 1 did not show a large, consistent omnibus gain. Experiment 2's narrower paired test favored relevant mechanisms over both controls on terminal internal quality. Experiment 11 reused those exact 60 terminal proposals, added uniform eight-source research, and found no frozen-gate support for relevant mechanisms after scrutiny (Experiment 1 findings; Experiment 2 results; Experiment 11 report).
- Prior-art search is part of generation quality, not a final clerical check. Experiment 3's closed-book funnel produced many plausible candidates, but external research sharply reduced the defensible yield. Of 47 assessed selected cells, 25 substantially collided with prior art and 15 described established practice; six were adjacent and one indeterminate (Experiment 3 final report).
- Proposal-first search was more productive than retrieval-first search in the tested workflow. In Experiment 4, proposal-first produced strict success in 5 of 20 arms, retrieval-first in 1 of 20. The paired estimate favored proposal-first by 20 percentage points, but the sample was too small to establish a stable effect (Experiment 4 report).
- One proposal is not enough. Experiment 5 found all four strict successes among proposals 1–4 and none at proposal 5. Experiment 6 replicated four-proposal generation at larger scale. It found five additional strict-success cells beyond proposal 1, narrowly missing the frozen requirement of six rescues (Experiment 5 report; Experiment 6 report).
- A cheap prefilter was not reliable enough. Experiment 7's blinded proposal-only rankings were moderately consistent with one another, but consistency did not predict downstream survival. The primary two-of-four selection retained 60% of strict proposals and failed all four prespecified retention requirements (Experiment 7 report).
- Coarse researchability was broadly distributed in the sampled archetype corpus. Experiment 9 applied one-shot generation and the same four-source light screen to all 16 hand-curated archetypes, 24 probability-sampled previously untested generated archetypes, and ten reference archetypes. All 24 random generated archetypes were productive in at least one of three fixed domains; 60/72 cells survived. Because the screen was deliberately inexpensive, this is evidence of breadth, not a claim that 83.3% of cells contain novel or valuable opportunities (Experiment 9 report).
- Post-hoc audits narrowed several possible overinterpretations. A ten-candidate prior-art sentinel found closer adjacent precedent in eight cases but no likely substantial collision with the surviving contrastive claims. Yield decomposition showed mixed, protocol-dependent breadth rather than one universal archetype. Blinded coding of 422 externally evaluated proposals found that governance/process and computational/information proposals dominated generation, but neither category had a clear general survival advantage. A corrected join and a new blinded 50-archetype classification then found a two-level substrate constraint: structural/framed character shifted allocation between computation and governance, yet 92.7% of the balanced E9 proposals remained in those two channels (sentinel audit; yield decomposition; proposal-type analysis; substrate analysis).
- The pipeline can leave its dominant substrates, but the forced escape was costly. Experiment 12 prohibited governance and computation as the primary causal substrate in the 72 probability-sampled E9 cells. Sixty-seven outputs complied; 49 survived the light screen compared with 60 ordinary controls, and eight cells survived only under the constraint. Yet blinded quality favored ordinary generation 68–4. Max effort improved both ordinary and constrained proposals in a small follow-up but did not preferentially rescue the constrained lane (Experiment 12 report).
- A deliberately different second lane can add portfolio coverage without becoming the default. Experiment 13 generated common first proposals and paired ordinary-diverse and substrate-diverse second proposals in 60 new cells. All 60 substrate proposals complied. Ordinary P2s produced 41 incremental survivors versus 32 for substrate P2s, but ten cells survived only through the substrate lane. Blinded quality was much closer than in Experiment 12, 32 ordinary wins to 28 substrate wins. The substrate lane met the frozen bounded-complement rule exactly at its allowed 15-point yield deficit; it did not earn outright preference (Experiment 13 report).
- Explicit applicability routes verified structure well, but did not retrieve it. Experiment 14 used 40 stratified route-bearing archetypes to construct 120 remedy-free positive and one-literal-near-miss cases. After blinded eligibility auditing, the DNF verifier accepted 77/77 positives and rejected 32/37 near misses, but diagnostic Recall@5 was only 2/77, below the solution index's 8/77. This supports the graph as an internal verification layer while exposing candidate generation as the bottleneck (Experiment 14 report).
- Naive route-aware embedding aggregation made retrieval worse. Experiment 15 excluded all E14 archetypes, constructed and outcome-blind audited another 120 cases, and compared two baselines with five prespecified condition-level aggregators. On 78 eligible positives, the primary route-aware arm achieved Recall@5 of 1/78, versus 6/78 for the solution index and 12/78 for diagnostic maximum-hit. A post-hoc diagnostic found severe short-route bias and showed that cosine similarity often ranked an explicit one-literal contradiction above its matched positive. Logical route structure requires signed condition evidence and calibration, not arithmetic over raw similarity scores (Experiment 15 report).
- The broadest useful output is a research portfolio. The 59 preserved candidates are invitations to expert judgment and bounded evidence collection. Twenty-seven have a strict or post-hoc strict-like status under their original protocols; 32 need a data-owning partner before the central claim can be judged. These endpoint classes must not be merged into a single “success rate.”
What the work does not establish¶
It does not establish that any candidate is new to the world, patentable, safe to deploy, economically valuable, demanded by customers, or better than the best existing system. It does not show that the model learned a general human-like faculty of far transfer. It does not estimate the yield of the Encyclopedia's full archetype-by-domain Cartesian product without strong sampling assumptions. It also does not show that the method is the first solution-first innovation system; several research traditions and recent systems occupy neighboring territory.
The practical recommendation¶
Do not scale the generator alone. Experiment 9 shows that the search signal is not confined to the small archetype set chosen for early experiments, but its high light-screen yield also shows why a coarse screen cannot be the terminal endpoint. Experiments 12 and 13 show that alternative-substrate generation can add portfolio diversity, while ordinary diversification remains the more productive default. Experiments 14 and 15 add a second architectural rule: use explicit DNF routes to verify a candidate after retrieval, but do not mistake raw embedding similarity—whether maximum-hit or condition-aggregated—for logical route satisfaction. If this program continues, use breadth sampling to allocate attention, preserve multiple proposals where missed opportunities matter, buy an alternative-substrate second slot only as a bounded complement when the value of missed opportunities justifies its lower researched yield, develop signed and calibrated candidate retrieval on a new frozen benchmark, spend deeper retrieval effort on the proposal-specific problem and composition, and send only bounded candidates to specialists. The next scientific step for the opportunity portfolio remains prospective human validation: can appropriate experts understand these candidates, find the evidence record adequate, and identify a nontrivial fraction worth testing?
1. The inverse-innovation idea¶
From problem-first to solution-first search¶
Most innovation narratives begin with a recognized need. Someone observes an undesirable state, studies its causes, and searches for interventions. The present method reverses the first move. It begins with a known pattern of intervention and asks where an unmet problem with the corresponding structure might exist.
This is not simply a request for an analogy. A plausible candidate must identify:
- concrete actors or material systems;
- an observable state, not merely a topic;
- a consequential failure or missed objective;
- a mapping from the source abstraction's parts and causal relations to the target;
- an intervention that changes those relations;
- a serious comparator;
- evidence that the problem and stakeholders are real;
- a remaining distinction after prior-art search; and
- a test whose outcome could stop the idea.
We call the resulting process solution-archetype-first cross-domain opportunity search. “Inverse innovation” is the shorter project name, not a claim to the established development-economics meaning of reverse innovation.
What the Encyclopedia contributes¶
The Encyclopedia of Abstractions is a large, structured corpus of abstractions, mechanisms, solution archetypes, and relations among them. A prime abstraction is intended to capture a compact recurring structure. A mechanism describes how a transition or effect occurs. A solution archetype packages a reusable arrangement of mechanisms that can address a class of problems. A domain is an area of human knowledge or practice, such as aviation, psychology, or museum studies.
The experiments treated an archetype–domain pair as a search cell. The archetype constrained the kind of causal structure sought; the domain supplied potential actors, observables, institutions, and failure modes. This makes the search space explicit. With five archetypes and 64 domains, for example, Experiment 3 had a declared 320-cell matrix rather than a hand-picked collection of favorable anecdotes.
The core workflow¶
flowchart LR
A["Solution archetype<br/>mechanisms and constraints"] --> C["Archetype × domain cell"]
B["Target domain<br/>actors, states, institutions"] --> C
C --> D["Generate distinct problem–proposal hypotheses"]
D --> E["Criticize and revise"]
E --> F["Search problem, prior art, and complete composition"]
F --> G["Strict practical evaluation"]
G --> H["Reject, research with a partner, or send to experts"]
The diagram's important feature is the sequence. Retrieval and testing are not decorations applied after the creative act; they determine whether a fluent proposal remains a defensible research hypothesis.
2. Claims, non-claims, and terminology¶
The supported claim¶
The experiments support a bounded claim:
Given explicit structural representations and proposal-specific external scrutiny, a contemporary large language model can generate cross-domain opportunity hypotheses across a declared matrix and a probability sample of previously untested archetypes; some hypotheses meet prespecified internal researchability and practical-screening criteria.
This is evidence about a configured human–AI research process. It is not evidence that an unaided model spontaneously performs robust transfer in arbitrary settings.
Stronger claims that remain open¶
Independent evidence would be required to conclude that:
- a candidate is absent from all prior art;
- a proposed intervention outperforms the best domain baseline;
- an adopter would fund, authorize, or use it;
- its benefit exceeds its cost and externalities;
- the method transfers across models, prompts, corpus versions, or researchers; or
- trajectories generated by this process can train a model to perform cross-domain transfer without the external pipeline.
Terminology used in this report¶
| Term | Meaning here |
|---|---|
| Cell | One declared solution-archetype × domain combination. A cell can contain several proposals. |
| Proposal | A particular problem–intervention hypothesis within a cell. It is not automatically a unique mechanism or opportunity family. |
| Stage success | A proposal cleared one internal stage. It says nothing about later scrutiny. |
| Established practice | External search found the proposal's substantive intervention already used or well documented. |
| Substantial collision | Prior art covered enough of the claimed composition that the remaining distinction did not support the tested strict lane. |
| Adjacent prior art | Relevant components or nearby systems exist, but a narrower composition, comparison, or domain instantiation remains unresolved. All 59 dossier candidates have this status. |
| Strict success | A candidate met the strict final rules in Experiment 4, 5, or 6. It is an internal researched-candidate endpoint, not proof of novelty or impact. |
| Post-hoc strict-like survivor | One of two Experiment 3 candidates identified after the sealed experiment by combining verified-pipeline and prior-art conditions. It was not a preregistered endpoint. |
| Empirical-partner candidate | An Experiment 6 candidate whose central question could not be decided from public sources and requires a specific data-owning or operational partner. It is disjoint from strict success. |
| Candidate dossier | A plain-language, evidence-linked presentation prepared after the experiments for expert review. Dossier ordering is not an experimental outcome. |
3. Related work and the contribution boundary¶
The motivating intuition has substantial intellectual ancestry. Analogical reasoning research distinguishes resemblance at the level of surface features from alignment at the level of relations. Structure-mapping theory, retrieval models such as MAC/FAC, and studies of analogical encoding all suggest why explicit relational representation can help—and why noticing and adapting a distant source are separate problems. Design-by-analogy, biomimetic search, patent recombination, and computational creativity likewise use existing solutions to stimulate or construct new possibilities.
Two older methods are especially close. TRIZ generalized recurring ways of resolving technical contradictions from patent evidence into reusable inventive principles; contemporary systems such as AutoTRIZ automate parts of that problem-to-principle-to-solution workflow. Zwicky's morphological analysis declares problem parameters and their possible values, constructs a combinatorial field or “morphological box,” and then uses cross-consistency assessment to remove incompatible configurations (Zwicky, 1967; Ritchey, 2015). The experiments' archetype-by-domain matrix is therefore not historically unprecedented as a search-space form. It is a particularly simple two-axis morphological field followed by generative construction and evidence-bearing scrutiny rather than cross-consistency assessment alone.
The Encyclopedia differs descriptively from classical TRIZ in source breadth, catalog size, mechanism detail, and typed relations, while the present pipeline differs from morphological analysis in what occupies a cell and how a proposed configuration is tested. Those architectural differences do not establish superior breadth or yield. Only random corpus sampling and fair end-to-end comparisons could do that.
Recent AI systems come still closer. AskNatureGPT retrieves biological strategies for design problems. Yoon and colleagues infer technology opportunities from functions in patents. A purpose–mechanism knowledge base supports cross-domain analogical search (Kang et al.). Recent preprints evaluate or train scientific analogy, hypothesis generation, and agentic literature search, including Shen, Druckmann, and Zou, MOOSE-Star, and RLAD. The companion literature review and nearest-systems matrix provide the fuller comparison.
Accordingly, this report makes no “first system” or “breakthrough” claim. Its more defensible contribution is the integration of several choices in one auditable program:
- a general abstraction corpus rather than a single source domain;
- a declared archetype-by-domain matrix;
- multiple diverse complete proposals per cell;
- preserved proposal, critique, revision, retrieval, and evaluation artifacts;
- proposal-specific searches for the problem, components, and full composition;
- separate quality, prior-art, practical, and empirical-partner gates;
- explicit negative tests, authority, safety, and resource estimates;
- frozen thresholds for later experiments; and
- publication of failures and resource records alongside survivors.
The distinction is architectural rather than absolute. Neighboring systems may contain individual pieces, and future comparison may reveal closer precedents.
The related-work review and candidate prior-art records were produced primarily through agentic web retrieval. That process favored digitally accessible, well-indexed material. TRIZ appeared in the companion review but was omitted from the first main-report synthesis, while Zwicky's older book-form precedent was missed until an adversarial review. These are observed retrieval-and-synthesis failures, not merely hypothetical limitations. Older books, non-English literature, trade practice, proprietary systems, and work described under different vocabulary remain plausible blind spots. Accordingly, ADJACENT_PRIOR_ART is always bounded by the documented searches; it is not a claim that closer precedent does not exist.
4. How the program evolved¶
The thirteen linked experiments were not thirteen replications of one frozen protocol. They were an iterative research program. Each experiment answered a narrower question exposed by the preceding work. Four diagnostic analyses and one sentinel audit were added to probe report-level vulnerabilities; they are explicitly post hoc and do not change any frozen verdict. Experiments 9 and 11–15 were frozen and executed prospectively; Experiment 7, although blinded, benchmarked a selector against already-known Experiment 6 outcomes and is therefore retrospective; the E9 substrate follow-up was designed only after E9 generation had finished and remains post hoc despite prespecified blinded coding.
| Experiment | Main question | Scale | Principal finding | Consequence |
|---|---|---|---|---|
| 1 | Does adding mechanism context produce a large general lift? | 60 cells × 3 conditions | No large consistent omnibus lift; context effects were heterogeneous. | Replace the broad comparison with a tighter paired design. |
| 2 | Do relevant mechanisms beat archetype-only and irrelevant mechanisms under iterative propose–criticize–revise–test? | 20 cells × 3 arms | Relevant mechanisms won 15/20 against each comparator and passed corrected paired tests. | Treat mechanism context as useful, then test at matrix scale. |
| 3 | Can five archetypes generate researched opportunities across all 64 domains? | 320 cells | Closed-book success was common; external prior art eliminated or narrowed most assessed selections. | Move external scrutiny earlier and make it proposal-specific. |
| 4 | Should retrieval precede or follow proposal construction? | 20 paired cells | Proposal-first yielded 5 strict arms; retrieval-first yielded 1; uncertainty remained large. | Keep proposal-first generation, then scrutinize complete proposals. |
| 5 | How many diverse complete proposals should a cell receive? | 20 cells × 5 proposals | Four strict candidates appeared in positions 1–4; position 5 added none. | Replicate four-proposal portfolios on new archetypes and domains. |
| 6 | Do four proposals rescue enough strict cells in a larger generalization set? | 60 cells × 4 proposals | Five rescues versus a frozen floor of six: NOT_SUPPORTED; 15 strict proposals and 32 partner candidates still resulted. |
Distinguish scientific threshold failure from practical yield and test cheaper selection. |
| 7 | Can proposal-only judgments prefilter four candidates without losing valuable ones? | 60 cells, 180 blinded rankings | Primary top-two rule retained 9/15 strict proposals and failed all frozen thresholds. | Do not replace prior-art scrutiny with proposal-only ranking. |
| 9 | Is coarse productivity distributed beyond the archetypes selected for earlier experiments? | 50 archetypes × 3 fixed domains | All 24 probability-sampled untested generated archetypes were productive; 60/72 random-sample cells survived the light screen. | Treat breadth as supported under a coarse endpoint; do not equate it with strict novelty or deployability. |
| 11 | Does Experiment 2's relevant-mechanism advantage survive uniform external research? | 20 cells × 3 preserved arms | R lost 9–11 to S, beat D 12–8, and failed the pooled and omnibus frozen gates. | Retain the internal-quality result but withdraw any general researched-yield advantage claim. |
| 12 | Can the generator leave governance/computation when those substrates are prohibited, and does Max effort rescue the cost? | 72 matched High-effort cells; triggered 18-cell 2×2 Max follow-up | 67/72 constrained outputs complied, but ordinary won blinded quality 68–4 and survived the light screen 60–49; Max did not preferentially rescue the constrained lane. | Treat dominant substrates as a soft but consequential default; use alternative-substrate generation for bounded diversification, not wholesale replacement. |
| 13 | If one additional proposal can be funded, should it use ordinary diversity or deliberately seek another causal substrate? | 12 probability-sampled untested archetypes × 5 probability-sampled domains; 60 common P1s and 120 paired P2s | Ordinary P2s produced 41 incremental survivors and substrate P2s 32; ten were substrate-only, 19 ordinary-only, and blinded quality was 32–28. | Keep ordinary diversity as the default and use an alternative-substrate P2 as a bounded complement when unique coverage merits the measured yield cost. |
| 14 | Can explicit applicability routes improve problem-to-archetype retrieval and reject structurally incomplete near misses? | 40 stratified archetypes × 3 constructed cases; 114 eligible after blinded audit | Diagnostic Recall@5 was 2/77 versus 8/77 for the solution index; target-route sensitivity was 77/77 and near-miss specificity 32/37. | Preserve DNF verification, but redesign route-aware candidate retrieval and validate it on independently sourced cases. |
| 15 | Does condition-level DNF aggregation improve candidate retrieval on new archetypes and cases? | 40 new stratified archetypes × 3 constructed cases; 115 eligible after blinded audit | Primary route-aware Recall@5 was 1/78, versus 6/78 for the solution index and 12/78 for diagnostic maximum-hit: NOT_SUPPORTIVE. |
Reject raw cosine aggregation as DNF retrieval; investigate signed evidence and route-length/count calibration without changing the E15 verdict. |
The post-hoc prior-art sentinel, Experiment 8 yield decomposition, Experiment 10 proposal-type analysis, E10B score-to-substrate join, and E9 substrate follow-up reuse existing proposals or outcomes. Their methods were frozen before the relevant searching, outcome join, or new substrate labels, but the questions themselves were chosen after seeing the main program. They are sensitivity and diagnostic analyses, not new confirmatory experiments.
This sequence matters when interpreting apparent contradictions. Experiment 6's negative primary verdict does not mean later proposals had zero value; it means the observed rescue count did not meet a threshold set in advance. Conversely, the existence of 59 dossier candidates does not retrospectively turn every experiment positive.
5. Unified method and protocol differences¶
Common anatomy of a candidate¶
Later protocols required a proposal to identify a problem, actors, observable state, consequence, affected objective, intervention, baseline, causal chain, structural mapping, nearest rivals, authority and safety constraints, and negative tests. External evaluation then searched for evidence of the problem, stakeholder pull, implementation feasibility, and prior art. It assigned nine 1–5 opportunity scores and rough 2026 USD resource-equivalent bands for first evidence, initial deployment, operational launch, and recurring cost.
The nine common dimensions were meaningful impact, stakeholder pull, incremental advantage, distinctiveness plausibility, technical implementability, adoption/authority feasibility, evidence readiness, safety/net benefit, and scalability. A high score did not authorize deployment. The required output was a next evidence step.
Generate, criticize, revise, test¶
Experiment 2 formalized the iterative loop rather than treating ideation as one shot. Each trajectory could be criticized, revised, and retested up to five revisions after the original proposal. A controller stopped when the candidate passed, failed without useful progress, required unavailable research under the closed-book rule, or exhausted the revision budget. This better approximated the iterative nature of human innovation while preserving bounded cost and an inspectable trajectory (protocol; analysis plan).
Experiments 4–6 shifted effort from repeatedly repairing a single weak idea toward generating several substantially different complete proposals and researching each one. That change separated within-proposal refinement from between-proposal search.
Prior-art scrutiny¶
Search was deliberately contrastive. It did not ask only whether the proposal's name appeared online. Evaluators searched for:
- the target problem and affected stakeholder;
- each load-bearing component;
- the full composition or causal sequence;
- the strongest realistic baseline;
- implementation and authority constraints; and
- evidence capable of resolving the remaining distinction.
The external labels were conservative. ESTABLISHED_PRACTICE and SUBSTANTIAL_COLLISION removed a proposal from strict consideration. ADJACENT_PRIOR_ART meant that public search left a narrower, testable contrast; it did not mean no prior art. NO_CLOSE_MATCH_FOUND was permitted by the schemas but did not occur among the 59 dossiers.
Strict and empirical-partner lanes¶
The strict lane favored candidates whose problem, comparison, implementation, and next test could be defended using the available record. Experiment 6 added a disjoint empirical-partner lane for cases where public evidence could not answer the decisive question but a named kind of partner plausibly could. The lane was calibrated on seven cases; all seven adjudications matched the calibration labels, although the small all-agreement set could not provide a meaningful chance-corrected reliability estimate (calibration result).
Post-hoc dossier ordering¶
For this report only, the 59 candidates are ordered with the three opportunity profiles first used after Experiment 2: balanced, deployment-heavy, and impact-heavy (method). The nine common ordinal scores are reused. A tenth input, pilot speed/cost, is approximated from the first-evidence cost band: under $10,000 = 5; \(10,000–\)50,000 = 4; \(50,000–\)250,000 = 3; \(250,000–\)1 million = 2; above $1 million = 1. Because elapsed pilot time was not independently scored, this is an affordability proxy, not a full pilotability measure.
For each profile, the weighted 1–5 mean is multiplied by 20 to create the familiar 0–100 display; because the underlying minimum is 1, the numerical floor is 20 rather than zero. The balanced profile supplies the reading order; rank ranges across the three profiles expose sensitivity. Ranks 1–10, 11–25, 26–45, and 46–59 are shown as broad bands. These values are an editorial navigation aid. They neither alter original endpoint labels nor measure economic value.
Reproducibility and preservation¶
From Experiment 2 onward, the project increasingly preserved schemas, prompts, source snapshots, raw model events, controller states, judgments, design freezes, completion manifests, analysis scripts, and resource accounting. The artifact index links the canonical materials. The public harmonization script does not rewrite experimental data; it reads sealed analysis outputs and candidate artifacts into a 59-row JSONL index.
Units, samples, and denominators¶
The experiments distinguish several units that are easy to blur:
- an archetype is the source solution structure;
- a domain is the target field;
- a cell is one archetype–domain pairing;
- a proposal is one complete problem–intervention hypothesis in a cell;
- a version is the original or a revision of that proposal;
- a trajectory is the sequence of versions, critiques, tests, and controller decisions; and
- a candidate instance is a proposal retained by a specified endpoint.
Experiment 3's 33/320 is a cell yield under a closed-book adopter-pipeline rule. Experiment 6's 15 strict proposals belong to 11 cells. The latter can be reported as 15/240 proposals or 11/60 cells, but the fractions answer different questions. The empirical-partner count is proposal-level and disjoint from strict proposals. This report gives denominators whenever practical and avoids adding outcomes from unlike levels.
Most samples were purposive rather than probabilistic. Experiment 1 balanced five archetypes across domain groups. Experiment 2 selected 20 informative cells for a paired condition test. Experiment 3 enumerated the project's complete 5 × 64 declared matrix but did not thereby sample the universe of real problems. Experiments 4 and 5 selected new cells to test pipeline changes. Experiment 6 deliberately used five different archetypes crossed with 12 domains to probe generalization beyond the original five. Experiment 7 reused all 60 Experiment 6 cells because it was a blinded retrospective selector study. Experiment 9 added a seeded probability sample of 24 previously untested generated archetypes from a 1,092-item corpus, but crossed them with only three purposively fixed domains. Experiment 11 reused Experiment 2's 20 selected cells and 60 terminal proposals to isolate the effect of equal external scrutiny. Experiment 12 reused E9's probability-sampled 72-cell stratum and sealed ordinary outputs as matched controls for a prospective prompt intervention; its domains therefore retain E9's purposive limitation. Experiment 13 excluded all E9-exposed archetypes, probability-sampled 12 previously untested generated archetypes, independently sampled five domains outside the E6/E9/E12 domain union, and crossed them to form 60 new cells. This improves both sides of the sampling design within the declared corpus and domain taxonomy, but still does not sample real-world problems. Experiment 14 hash-sampled eight unique route-bearing archetypes from each of five prespecified Applicability Graph strata. Experiment 15 repeated that stratified sampling design after excluding all 40 E14 archetypes. Both experiments deliberately constructed 120 cases from the graph rather than sampling problems from the world; this supports controlled route tests but sharply limits external-validity claims.
Model sessions, role separation, and blinding¶
Later runs used fresh or retained sessions according to the role. In Experiments 5 and 6, one retained session per cell generated the proposal portfolio so the model could avoid repeating its earlier ideas; external evaluation used fresh candidate-version sessions so the evaluator did not inherit the generator's discussion. Experiment 7 used fresh proposal-only sessions and blinded candidate identifiers. Experiment 9 used one fresh generator and one fresh light-screen evaluator for each cell. Experiment 2 concealed arm identity from quality judges and revealed keys only after the judgments were sealed. Experiment 11 preserved that concealment: fresh evaluators researched opaque proposals, two fresh judges compared the researched arms within each cell, a third resolved disagreement, and arm keys were revealed only after adjudications were sealed. Experiment 12 used two fresh blinded substrate-classification passes and two fresh randomized A/B quality passes, with separate blinded adjudication of every disagreement before private keys were joined. Its web screens used the unchanged E9 protocol. Experiment 13 used isolated fresh calls for P1 and both P2 arms, then replaced proposal, pair, arm, archetype, and domain identities with opaque keys. Two fresh passes classified substrate and P1–P2 independence, two randomized A/B passes compared P2 quality, fresh calls adjudicated every disagreement, and 180 separate opaque evaluators applied the unchanged four-source screen before key reveal. Experiment 14 used fresh isolated case-writing calls, two independent opaque case-audit passes, fresh adjudication of all material disagreements, deterministic retrieval from frozen indexes, and one fresh opaque DNF verifier per eligible case. The verifier never saw archetype names, target identity, or intended case type; private keys were joined only in analysis. Experiment 15 used the same fresh case-writing and duplicate blinded-audit procedure on a disjoint sample, sealed eligibility before deterministic retrieval, and made no model call during ranking.
This is role separation, not full independence. The same model family sometimes occupied several roles, and the model may have encountered similar source material during training. Independent sessions reduce direct conversational leakage but not shared priors. The exact runtime settings, service tier, timeouts, retry budgets, and worker counts are preserved in run manifests linked from the artifact index.
Validation, retries, and nonresponses¶
Model outputs were constrained by JSON Schemas and checked by deterministic code. Transport or schema failures could be retried within declared limits; substantive revisions followed the controller rules rather than silently replacing an unfavorable judgment. Run records preserve prompts, responses, events, errors, elapsed time, and exposed token telemetry. Experiment 3's external tranche retained one terminal nonresponse after three 1,800-second timeouts instead of imputing a favorable or unfavorable answer.
Amendments corrected transport and processing issues without changing sealed scientific fields. Completion seals and manifests record hashes or references for the applicable design and evidence. The presence of a valid JSON object guarantees structural conformance, not truth; substantive claims still depend on sources and judgment.
Statistical interpretation¶
The program used simple statistics matched to the small paired or binomial questions. Experiment 2 used paired direction counts and sign tests with Holm correction. Experiment 4 used the paired McNemar test and reported the absolute arm difference. Experiments 5 and 6 reported marginal cell rescue with Wilson intervals. Experiment 7 reported proposal and cell recall with Wilson intervals, along with fixed threshold pass/fail decisions. Experiment 9 reported Wilson intervals for sampled-archetype productivity and cell survival. Experiment 11 used paired direction counts, Holm-adjusted one-sided sign tests, and an exact within-cell label-randomization omnibus test. Experiment 12 reported Wilson intervals, exact paired sign or McNemar tests, and a 50,000-draw paired bootstrap clustered by archetype; its small Max subset remained descriptive. Experiment 13 reported the paired survivor table, Wilson intervals, a 20,000-draw archetype-clustered bootstrap, and an exact sign-flip test over 12 archetype-level summed differences; cell-level McNemar and sign tests remained sensitivity analyses. Experiment 14 reported Recall@k, reciprocal rank, a paired exact McNemar test, a 20,000-draw archetype-clustered bootstrap, sensitivity, specificity, and balanced accuracy. Experiment 15 reused Recall@k, reciprocal rank, paired exact McNemar tests, and 20,000-draw archetype-clustered intervals, with one confirmatory route-aware arm and four prespecified but non-rescuing sensitivity arms. These analyses quantify sampling uncertainty conditional on their stated sampling frame; they do not repair domain-taxonomy, model, or evaluator dependence.
Frozen thresholds were treated as decision rules, not as natural constants. A rule such as “at least six rescued cells” embodies a prior judgment about what would justify the next scale-up. Reporting the observed count and interval remains necessary when the rule fails. Exploratory comparisons are labeled as such and are not allowed to replace the frozen primary verdict.
Provenance of load-bearing decision thresholds¶
The table distinguishes a threshold being present before execution from a public preregistration. These were internal design freezes in a rapidly adapting solo research program. Where the record preserves no contemporaneous numerical derivation, the report says so rather than supplying a retrospective rationale as though it had been frozen.
| Experiment | Frozen rule | Timing and evidence | Preserved rationale and limitation |
|---|---|---|---|
| 1 | Condition C had to exceed A by a paired median of at least 10 points, alongside diversity, safety, and survivor conditions. | Present in the experiment specification used before execution; no separate signed design-freeze manifest was created. | The rule demanded a large, practically visible omnibus gain before scaling. The record does not preserve a more detailed power or cost derivation for 10 points. |
| 5 | At least four rescued cells, at least three additional strict proposals beyond P1, and at least one strict proposal at P3–P5. | Sealed at 2026-08-03T04:38:21Z, before scientific execution, in the design freeze. |
The rule required both cell rescue and evidence that search depth beyond the first two proposals mattered. It was a minimum-practical-effect rule, not a significance test. |
| 6 | At least six rescued cells, at least six strict proposals beyond P1, and at least one new strict cell at P4. | Frozen before execution in the method and hashed design manifest; execution authorization was recorded at 2026-08-03T11:41:12Z. |
Experiment 5 rescued 3/19 P1 failures (15.8%); applying that observed rate to Experiment 6's 54 P1 failures would project about 8.5 rescues. A floor of six was below that projection. This numerical comparison explains the rule after the fact; the sealed method itself does not record that derivation. |
| 7 | The top-two policy had to retain at least 9/11 strict cells, 25/31 broad cells, 11/15 strict proposals, and 33/47 broad proposals. | Design frozen at 2026-08-04T01:00:01Z; all 180 selector outputs were frozen before label join at 2026-08-04T02:02:41Z (prediction freeze). |
These were high-recall safeguards: saving half the searches was not sufficient if the selector discarded too much known yield. Designers knew aggregate Experiment 6 outcomes, so this was a blinded retrospective policy benchmark, not a prospective preregistration. |
| 9 | At least 10/24 random generated archetypes productive, their productive fraction no more than 0.20 below the reference block, and at least two random-sample survivors in each fixed domain. | Frozen before execution in the design freeze; a documented pre-run correction repaired the candidate-pool implementation without changing the scientific question. | These were broad-distribution floors at a deliberately light screening endpoint. They were not novelty or strict-opportunity thresholds. |
| 11 | R had to be net-positive against both S and D, reach at least 65% pooled pairwise preference, pass an exact omnibus test at p ≤ 0.05, and avoid excess safety stops. |
Protocol, analysis, blinded inputs, and arm key were sealed in the design freeze before external evaluation. | The joint gate demanded a large, consistent researched-record advantage rather than a marginal difference in strict counts. It was intentionally stronger than simply repeating Experiment 2's internal score comparison. |
| 12 | Run Max only if at least 24/72 constrained outputs complied and either constrained decisive quality wins were at most 40% or ordinary usable yield led by at least 15 points. | Method, treatment prompt, inputs, schemas, and the 18-cell Max subset were sealed before treatment generation. A post-run audit correction records that the mutable status README was mistakenly included in the original hash set; all load-bearing design sources still verify. | The first clause distinguished instruction failure from a viable but difficult task. The second allocated expensive Max calls only when a meaningful cost appeared. It was a resource trigger, not a confirmatory hypothesis test. |
| 13 | Prefer the substrate P2 only for a gain of at least five points and more unique survivors; otherwise call it a bounded complement only with at least 48/60 compliant proposals, at least six substrate-only survivors, and a yield deficit no worse than 15 points. | The sampling frame, selected cells, prompts, schemas, endpoints, statistics, and decision rule were sealed before any E13 model output in the design freeze. | The rule required at least 10% unique coverage while capping the cost of diversification. The observed deficit landed exactly on the permitted 15-point boundary, so the operational gate passed but is not a robust equivalence result. |
| 14 | Diagnostic Recall@5 had to exceed solution Recall@5 by at least ten points, produce at least eight more diagnostic-only successes, pass exact McNemar at p ≤ 0.05, and avoid a loss greater than five points in either positive subtype. DNF verification separately required sensitivity and specificity of at least 0.70 and balanced accuracy of at least 0.75. |
The method, sample, graph and MCP snapshots, prompts, schemas, endpoints, statistics, and gates were sealed before any E14 scientific model output in the design freeze. | The retrieval rule demanded a practically decisive improvement before replacing the existing index. The verification rule tested whether explicit conjunctions added structural discrimination rather than merely semantic resemblance. Retrieval failed; verification passed. |
External evidence record¶
Later external evaluations required eight structured sources per terminal candidate, including the problem, stakeholder, implementation, prior-art, and resource evidence needed for the scoring record. Source objects preserve title, publisher, URL, source class, publication or coverage date where available, access date, and the claims supported. Evaluators also preserved emitted queries. Source count is a coverage discipline, not a guarantee of eight independent or equally strong facts.
Search used public web services only where the user explicitly authorized sending candidate summaries and derived queries. The public report includes those URLs, candidate text, and project records; it does not knowingly include a partner's private operational dataset. Future partnered studies may introduce human-subject, proprietary, security, or regulated data obligations and must obtain the appropriate institutional, legal, and participant approvals before data transfer or intervention.
6. Results by research question¶
Does structural mechanism context improve proposal quality?¶
Experiment 1 provided the first caution. Sixty cells were run in three conditions and reviewed, but the prespecified condition-C versus condition-A paired median gain was 3.12 points, below the frozen 10-point threshold. Hard-gate and promotion counts moved in the expected direction—34/60 and 22/60 in condition A, 38/60 and 27/60 in B, 41/60 and 30/60 in C—but the evidence did not support a large, consistent omnibus lift. The protocol also carried context-length and review-blinding limitations described in its quality audit.
Experiment 2 narrowed the question and improved the controls. Each of 20 cells had three trajectories: relevant mechanism context (R), solution-archetype context without the relevant mechanisms (S), and an equal-length irrelevant-mechanism decoy (D). All three used the bounded propose–criticize–revise–test loop. Terminal relevant-mechanism proposals later passed a post-hoc quality entry gate in every cell. On paired blinded judgments, R beat S in 15 cells, lost in four, and tied in one; it beat D in 15 and lost in five. The Holm-adjusted sign-test values were 0.01921 and 0.020695 respectively (machine-readable results).
Experiment 11 subjected the same 60 terminal proposals to a stronger endpoint. Treatment labels were hidden while one fresh external evaluator per proposal searched six lanes and retained exactly eight sources; two fresh judges then compared the three opaque researched records within each cell, with a third judge used for nine disagreement cells. Relevant mechanism context (R) beat archetype-only context (S) in 9 cells and lost in 11; it beat the irrelevant control (D) in 12 and lost in 8. Pooled R preference was 52.5%, below the frozen 65% floor, and the exact omnibus randomization result was p = 0.4459. Strict researched counts were close—12/20 for R, 11/20 for S, and 10/20 for D—and no arm produced a safety or authority stop (Experiment 11 results).
The combined conclusion is narrower than the earlier report. Experiment 2 remains evidence that relevant mechanisms improved internal terminal quality under its rubric. Experiment 11 did not support an advantage after external scrutiny and comparative judgment. It does not show that mechanisms are useless—the R-versus-D direction remained positive, and the interval is wide—but it prevents the internal-quality result from being generalized into a researched-survivor claim.
Can the method operate across a declared matrix?¶
Experiment 3 crossed five solution archetypes with 64 domains, producing 320 cells. Its internal funnel was permissive by later standards: 297 of 320 cells completed the first generation stage successfully, including 284 of the 300 cells in the prospectively designated tranche. Thirty-three of 320 passed the complete closed-book pipeline.
That result showed operational scale, but it overstated defensible opportunity yield because the pipeline had deliberately withheld external research. Forty-eight candidates were then selected for external assessment; 47 returned an assessment and one did not. Among the assessed set:
| Prior-art disposition | Candidates | Share of 47 assessed |
|---|---|---|
| Substantial collision | 25 | 53.2% |
| Established practice | 15 | 31.9% |
| Adjacent prior art | 6 | 12.8% |
| Indeterminate | 1 | 2.1% |
| No close match found | 0 | 0% |
Twenty-one of 47 passed the externally verified pipeline as it was then defined, but only two combined verified-pipeline status with the later strict-like prior-art condition. Because that two-candidate composite was constructed post hoc, the report preserves the candidates while refusing to treat “2/320” as a preregistered success rate. The more general lesson is firm: internal coherence and closed-book quality cannot substitute for retrieval.
Experiment 9 asked a different matrix question: whether even coarse opportunity productivity was confined to the archetypes selected by the researchers. It applied the same one-shot proposal and four-source light screen to all 16 hand-curated archetypes, a seeded probability sample of 24 previously untested generated archetypes from a 1,092-item generated corpus, and the ten earlier archetypes as a same-protocol reference block. An archetype counted as productive if at least one of its accounting/auditing, chemistry/materials, or computer-science cells survived.
All 24 random generated archetypes were productive (Wilson 95% interval 86.2%–100%), and 60 of their 72 cells survived (83.3%; 73.1%–90.2%). The fixed-domain counts were 23/24 in accounting/auditing, 16/24 in chemistry/materials, and 21/24 in computer science. The hand-curated census produced 16/16 productive archetypes and 34/48 surviving cells; the earlier-archetype reference produced 9/10 and 20/30. All three frozen breadth conditions passed (Experiment 9 results).
This unusually high rate must be read against the endpoint: among the random cells, the screener assigned 60 ADJACENT_PRIOR_ART, 11 SUBSTANTIAL_COLLISION, and one ESTABLISHED_PRACTICE. A four-source light screen is meant to eliminate obvious failures, not establish novelty, economic value, or deployment readiness. Experiment 9 supports broad coarse researchability across the sampled generated corpus and three purposive domains. It does not estimate strict opportunity yield over the full Cartesian product.
Should retrieval come before or after proposal construction?¶
Experiment 4 compared two arms within 20 cells. The proposal-first arm created a complete problem–intervention hypothesis and then searched for the problem, rivals, components, and full combination. The retrieval-first arm searched broadly first, generated 100 short hypotheses, screened 53 out, subjected 47 to a critic, and developed 15 finalists.
Proposal-first produced strict success in 5 of 20 arms; retrieval-first produced 1 of 20. Thirteen of the 15 retrieval-first finalists collided substantially during full evaluation. The paired difference favored proposal-first by 20 percentage points, but McNemar's exact test was not significant at conventional levels (p = 0.2188). The sample therefore supplied a design recommendation, not a stable effect estimate.
Why might proposal-first help? Broad retrieval can anchor the search on already-named problems and established solution vocabularies. A complete proposal gives research a contrastive object: the evaluator can ask whether this precise composition exists and whether it improves on this specific rival. The result does not imply that retrieval should be delayed until the end. It supports proposal construction followed immediately by scrutiny, with revision after evidence.
How many complete proposals should a cell receive?¶
Experiment 5 generated five diverse complete proposals in each of 20 new cells. One strict success appeared at each of positions 1, 2, 3, and 4; position 5 added none. Portfolio success therefore rose from 1/20 cells after proposal 1 to 4/20 after proposal 4, a three-cell rescue among 19 first-proposal failures. The 95% Wilson interval for the rescue rate was wide, 5.5%–37.6%, so “four rather than five” was a replication choice rather than a universal optimum.
Experiment 6 tested that choice in 60 new cells created from five new archetypes and 12 domains. Its strict yield curve was:
| Proposal position | Strict proposals at position | Newly successful cells | Cumulative strict-success cells |
|---|---|---|---|
| 1 | 6 | 6 | 6/60 (10.0%) |
| 2 | 2 | 1 | 7/60 (11.7%) |
| 3 | 1 | 0 | 7/60 (11.7%) |
| 4 | 6 | 4 | 11/60 (18.3%) |
The four-proposal portfolio contained 15 strict proposals in 11 cells. It rescued five of the 54 cells that failed at proposal 1, a 9.3% rescue rate (95% Wilson interval 4.0%–19.9%). The frozen rule required at least six rescues and also specified later-position conditions; five rescues missed the primary floor by one. The correct confirmatory verdict is NOT_SUPPORTED.
That verdict and the practical calculation answer different questions. Scientifically, the prespecified replication threshold failed. Operationally, researching proposals 2–4 found nine additional strict proposals and five additional strict cells. A program that values a missed candidate may still rationally pay for later positions. The data do not determine that value; they make the tradeoff visible.
Can public evidence resolve every promising proposal?¶
No. In Experiment 6, 32 proposals met the empirical-partner criteria and were disjoint from the 15 strict proposals. The combined broad lane contained 47 proposals across 31 of 60 cells. Later positions added 19 of those 31 cells: the combined cell curve rose from 12/60 after proposal 1 to 31/60 after proposal 4.
This broad yield must not be interpreted as 51.7% innovation success. The partner label means that a decisive observable—often an internal error rate, workflow trace, operational baseline, or counterfactual comparison—was unavailable in public sources. These candidates are better understood as structured requests for collaboration. Their value is that they specify what a partner would need to measure and what result should terminate the idea.
Can a cheap selector reduce research cost without losing the good candidates?¶
Experiment 7 retrospectively hid Experiment 6 outcomes and asked three fresh proposal-only selectors to rank the four candidates in every cell. Their rankings showed meaningful agreement: mean pairwise Spearman correlation was 0.634 and the median was 0.8. Yet consensus did not supply the needed validity:
| Consensus rank | Strict candidates | Strict-or-partner candidates |
|---|---|---|
| 1 | 3 | 12 |
| 2 | 6 | 13 |
| 3 | 2 | 8 |
| 4 | 4 | 14 |
Rank 2 contained the most strict candidates, while rank 4 contained the most broad-lane candidates. Downstream survival was not monotonic in proposal-only rank, and the highest-ranked proposal was not reliably privileged.
The frozen primary policy selected the top two proposals per cell—120 of 240 proposals. It retained:
- 9/15 strict proposals (60.0%; 95% Wilson interval 35.7%–80.2%);
- 8/11 strict cells (72.7%; 43.4%–90.3%);
- 25/47 strict-or-partner proposals (53.2%; 39.2%–66.7%); and
- 23/31 strict-or-partner cells (74.2%; 56.8%–86.3%).
All four frozen recall thresholds failed, so prospective confirmation was not warranted. A proposal-only selector may still order a queue or support triage when misses are acceptable. It cannot, on this evidence, replace the expensive search stage while claiming to preserve the opportunity yield.
Did a stronger prior-art search overturn the sampled survivors?¶
The post-hoc sentinel audit hash-selected ten candidates across Experiments 3–6 and ran four targeted search lanes per case: historical lineage, trade or standards practice, non-English terminology, and the full proposed composition. It found closer adjacent precedent for 8 of 10 candidates and no closer precedent for two. None met the audit's rule for a likely substantial collision, and all ten remaining contrastive claims survived.
This result is reassuring only in a narrow sense. It does not validate all 59 dossiers or estimate a corpus-wide false-novelty rate. It shows that the sampled labels were robust to this extra search while also demonstrating that the first-pass records often did not contain the strongest neighbors. Future strict scrutiny should make the four search lanes explicit before assigning a novelty-like disposition (audit method and case results).
Was observed yield concentrated in one solution archetype?¶
Experiment 8 decomposed existing outcomes by archetype and domain without changing the original endpoints. Across 13 protocol–outcome strata, seven met the frozen diagnostic for breadth across tested archetypes and six were concentrated. Strict success appeared under 2 of 5 archetypes in Experiment 5 and 4 of 5 in Experiment 6; Experiment 6 empirical-partner candidates appeared under all 5. The 11 strict-success cells in Experiment 6 were somewhat concentrated by the frozen diagnostic, whereas its 15 strict proposals were broadly distributed.
The defensible conclusion is mixed, protocol-dependent breadth. No single archetype explains the program, but sparse events and purposive samples make archetype rankings unstable. The selected 47-candidate Experiment 3 external subset is not a denominator for the full 320-cell matrix, and domain rows cannot be read as domain effects (Experiment 8 report).
Did proposal substrate or evidence dependency track disposition?¶
Experiment 10 assembled the 422 primary externally evaluated proposals from Experiments 3–6 and removed their outcomes before classification. Two fresh passes agreed on proposal substrate for 399/422 cases (94.5%; Cohen's κ = 0.9004) and evidence dependency for 373/422 (88.4%; κ = 0.8284); a third blinded pass adjudicated the 68 rows with any disagreement.
The generated mix was highly uneven: 208 proposals (49.3%) were governance/process interventions and 189 (44.8%) were computational/information interventions; only 25 fell into the other three substrate classes. Yet strict-success rates for those two large classes were similar across the pooled heterogeneous records—13/208 and 13/189—and Experiment 6's exploratory governance-versus-other empirical-partner comparison was also inconclusive (odds ratio 1.27; two-sided Fisher p = 0.572). Evidence dependency was dominated by private operational, laboratory/technical, and field/human-institutional tests, but small counts and protocol differences do not support a universal evidence-class ranking.
This is a diagnosis of what the pipeline generated and where evidence bottlenecks occurred, not a causal test of which proposal type is intrinsically more innovative. It weakens the simpler story that the portfolio exists mainly because evaluators favored governance ideas (Experiment 10 report).
Did archetype character explain the proposal substrate?¶
The motivating internal note initially omitted Deadweight Loss Reduction from the ten archetypes used in Experiments 3–6 and therefore overstated their mean structural/framed score as +0.462. The corrected ten-archetype mean is +0.366, close to the +0.374 corpus reference. E10B then used the ten archetypes—not 422 repeated proposals—as its inferential units. Across the 340 complete-census version-zero proposals from Experiments 5 and 6, source score was associated with governance share within the governance/computational channels (Spearman ρ = −0.666; exact exploratory permutation p = 0.0422; leave-one-archetype-out range −0.822 to −0.601), but not with escape from those channels (ρ = −0.058; p = 0.85).
The stronger E9 follow-up retained all 150 proposals from 50 archetypes crossed with three fixed domains and removed archetype, domain, sampling, score, screen, and prior-art fields before new classification. Two passes agreed on 140/150 labels (93.3%; Cohen's κ = 0.885), and a third blinded pass adjudicated all ten disagreements before the key was joined. Across all 50 archetypes, source score versus governance allocation had ρ = −0.624 (seeded permutation p = 0.00001); in E9's probability-sampled 24-archetype stratum, ρ = −0.791 (p = 0.00002). Each domain's direction was negative, while 139/150 proposals (92.7%) remained governance/process or computational/information.
Together these results support a two-level substrate constraint in the tested pipeline: archetype character modulates which dominant channel appears, but does not explain why the pipeline remains concentrated in those two channels. This is an association within the complete bundle of corpus, prompt, model, rubric, and protocol—not evidence that the base model has an intrinsic substrate limitation. The question was selected after E9 finished, so prespecified blinded coding strengthens the result without making it prospective (combined analysis and validation).
Can the pipeline be forced to leave governance and computation?¶
Experiment 12 converted the post-hoc substrate observation into a prospective prompt intervention. It reused the 72 random-generated E9 cells and their sealed ordinary/High proposals as matched controls. A fresh gpt-5.6-sol High-effort call for each cell had to make the primary causal intervention physical/material or measurement/instrumentation; governance and computation could appear only as removable support. Two independent blinded classifier passes and adjudication found that 67 of 72 constrained outputs complied: 40 were physical/material and 27 measurement/instrumentation. Two proposals disguised governance as a physical intervention and three abstained.
Compliance did not imply equal quality. In randomized blinded A/B comparison, the ordinary proposal won 68 cells and the constrained proposal won four, with no ties (p = 4.62 × 10⁻¹⁶, exact two-sided sign test). After the authorized four-source web screen, 49 constrained proposals and 60 ordinary controls survived. The paired difference was −15.3 percentage points; an archetype-clustered bootstrap interval ran from −30.6 to 0 points. Among the 27 discordant cells, eight cells survived only under the constraint and 19 only under ordinary generation (p = 0.0522, exact paired McNemar/sign test).
The cost was uneven. Constrained and ordinary light-screen yield tied at 16/24 in chemistry/materials, while the constraint trailed 14–21 in computer science and 19–23 in accounting/auditing. Framed archetypes produced only 6/12 constrained survivors versus 11/12 ordinary survivors, with no constrained quality win. Measurement/instrumentation was a frequent escape route—27 proposals and 21 survivors—but produced only one constrained quality win. These subgroup counts are descriptive because domains were purposive and the experiment was not powered for multiple heterogeneity tests.
The frozen trigger launched an 18-cell, four-arm Max follow-up. Max beat High 14–4 for ordinary proposals and 12–6 for constrained proposals, but ordinary beat constrained 16–2 at both effort levels. Light-screen survival was 12/18 ordinary/High, 11/18 constrained/High, and 9/18 for each Max arm. Higher effort improved the judged articulation and coherence of both arms without preferentially rescuing substrate denial or increasing this coarse researched yield.
The causal conclusion is bounded but important. Governance and computation are not a hard expressive ceiling: the tested generator can leave them when instructed. They are a consequential default, because forced alternatives were usually more contrived or less faithful and survived external screening less often. The eight constrained-only survivors show portfolio-diversification value, so the practical response is a limited alternative-substrate lane rather than a blanket ban (Experiment 12 report and artifacts).
Which second proposal should a search portfolio buy?¶
Experiment 13 turned Experiment 12's diversification suggestion into a prospective policy comparison on 60 new cells. Twelve probability-sampled previously untested generated archetypes were crossed with five independently sampled domains. Each cell received a common first proposal, an ordinary-diverse second proposal, and a second proposal whose primary causal substrate had to be physical/material or measurement/instrumentation. All calls were isolated and closed-book during generation. Two blinded classification passes, two blinded P1–P2 independence passes, two randomized P2 quality passes, fresh adjudication of every disagreement, and 180 opaque four-source web screens were completed before the private keys were joined.
The substrate instruction was fully effective: all 60/60 proposals complied, comprising 24 physical/material and 36 measurement/instrumentation interventions. Ordinary generation remained concentrated in the earlier house style: 54/60 ordinary P2s were governance/process or computational/information. Yet the sharp E12 quality gap did not recur in this more appropriate second-slot comparison. Blinded quality favored ordinary P2s 32–28, with no ties (p = 0.699, cell-level exact sign-test sensitivity).
The researched incremental endpoint still favored ordinary diversification. A P2 counted only if blinded judgment found it complete, structurally faithful, materially independent from P1, and its light screen survived; substrate P2s also had to comply. Ordinary P2s yielded 41/60 (68.3%) incremental survivors and substrate P2s 32/60 (53.3%). The paired table was 22 both, 10 substrate-only, 19 ordinary-only, and 9 neither. The substrate-minus-ordinary difference was −15.0 points; the archetype-clustered 95% bootstrap interval was −31.7 to +3.3 points, and the exact sign-flip sensitivity test over 12 archetype aggregates gave p = 0.1953. Those uncertainty results do not establish equivalence or a stable population difference.
The unique-survivor result is the reason the answer is not simply “always buy ordinary P2.” P1 survived in 43 cells. Among its 17 failures, ordinary P2 rescued five and substrate P2 rescued eight. Across all cells, ten researched opportunities appeared only in the substrate lane. The frozen policy therefore returned USE_AS_BOUNDED_PORTFOLIO_COMPLEMENT: compliance exceeded 48, substrate-only survivors exceeded six, and the yield deficit was no worse than 15 points. The last condition passed exactly at its boundary. This is evidence for a deliberately limited diversification lane under the tested policy—not for replacing ordinary diversity or for claiming that any light-screen survivor is a validated invention (Experiment 13 report and artifacts).
Can the Applicability Graph retrieve and verify matching archetypes?¶
Experiment 14 tested the new trigger-logic layer on a controlled benchmark. Forty unique route-bearing archetypes were hash-sampled across five strata: single-literal, fully grounded prime-only, fully grounded routes containing a domain-specific abstraction, partially open routes, and alternative-route archetypes. A fresh case writer produced a direct positive, a cross-domain positive, and a one-literal near miss for each. Two opaque audit passes plus fresh adjudication retained 114/120 cases—77 positives and 37 near misses spanning all 40 archetypes—before retrieval began.
The existing solution-oriented index retrieved the intended archetype in its top five for 8/77 positives (10.4%). The diagnostic maximum-hit search retrieved it for 2/77 (2.6%), a difference of −7.8 percentage points. The paired table contained two shared successes, zero diagnostic-only successes, six solution-only successes, and 69 failures by both arms (p = 0.03125, exact two-sided McNemar). The diagnostic arm failed every directional part of its frozen superiority gate. Its target appeared anywhere in the top 50 for only 13/77 positives, versus 22/77 for the solution index.
Verification told a different story. Each eligible case received an opaque literal-by-literal assessment of the target's complete DNF routes; when retrieval had missed the target, a shadow copy permitted fidelity measurement but was barred from the reranking pool. The verifier accepted 77/77 positives and correctly rejected 32/37 near misses, yielding 93.2% balanced accuracy and passing all frozen verification floors. Because the diagnostic top 12 contained the target for only 4/77 positives, reranking could not repair retrieval.
The frozen verdict was therefore VERIFICATION_ONLY. Within this graph-derived benchmark, explicit routes are useful for deciding whether an archetype under consideration actually fits. The current maximum-single-hit search is not an effective way to decide which archetype to consider. Four of five false acceptances occurred in the partially open stratum, an exploratory localization suggesting that incompletely grounded conditions need particular curation. The benchmark establishes internal operational fidelity only: its cases came from the graph, and related model-family systems constructed, audited, and verified them (Experiment 14 report and artifacts).
Does explicit route aggregation repair candidate retrieval?¶
Experiment 15 tested the next design without tuning it against E14. It excluded
all 40 E14 archetypes, hash-sampled eight new archetypes from each of the same
five strata, and froze five condition-level aggregation formulas before any new
case existed. Every formula scored conditions separately, combined them within
a route, and took the best alternative route. The confirmatory
SENTENCE_LOWER_HALF arm matched every condition against the full scenario and
deterministic sentence-like segments, then averaged the lowest half of the
condition scores. Four minimum or mean variants were sensitivity arms and
could not rescue the verdict.
The same outcome-blind case procedure retained 115/120 cases—78 positives
and 37 one-literal near misses across all 40 archetypes—before retrieval. The
primary arm retrieved only 1/78 positives (1.3%) in its top five. The
solution index retrieved 6/78 (7.7%), and diagnostic maximum-hit retrieved
12/78 (15.4%). Against the solution baseline, the primary arm had zero
exclusive successes and five baseline-only successes (p = 0.0625); against
diagnostic maximum-hit, it had one exclusive success and 12 baseline-only
successes (p = 0.00342). It also exceeded the allowed five-point loss in both
direct and transfer cases. The frozen verdict was NOT_SUPPORTIVE.
| Frozen E15 arm | Recall@1 | Recall@5 | Recall@10 | Recall@50 |
|---|---|---|---|---|
| Solution-oriented index | 2/78 | 6/78 | 8/78 | 23/78 |
| Diagnostic maximum-hit | 4/78 | 12/78 | 15/78 | 22/78 |
| Sentence-aware lower-half mean (primary) | 0/78 | 1/78 | 1/78 | 19/78 |
| Sentence-aware route mean (secondary) | 0/78 | 2/78 | 7/78 | 33/78 |
The secondary sentence-mean arm's 42.3% Recall@50 shows that condition-level similarity contained some broad candidate signal, but its 2.6% Recall@5 was not operationally competitive and cannot change the primary result. Baseline order also reversed across E14 and E15: maximum-hit lost in E14 and led in E15. Descriptively pooling the two disjoint benchmarks gives both baselines 14/155 top-five successes, with nine exclusive successes apiece. This was not a prespecified pooled endpoint; it shows instability and low absolute recall, not equivalence.
A clearly post-hoc failure localization found two mechanisms. Single-literal routes occupied 378/390 (96.9%) primary-arm top-five slots on positive cases, demonstrating severe route-length and alternative-count bias. More fundamentally, cosine similarity did not distinguish entailment from explicit contradiction. Among 36 eligible matched transfer-positive/near-miss pairs, the primary arm ranked the contradicted near miss better in 22 and the positive better in 14; the median near miss ranked 28.5 places higher. Arithmetic over raw similarity scores therefore did not implement logical conjunction. A future design would need signed condition evidence and calibration for route length, route count, and sentence-search multiplicity. Those are hypotheses, not retroactive repairs to E15 (Experiment 15 report and artifacts).
7. Negative results and what they changed¶
Negative evidence is central to this program because fluent generation makes false confidence cheap.
The expected large context gain did not appear¶
Experiment 1 rejected the idea that progressively richer context would produce a large uniform score gain across the calibration matrix. That failure led to an equal-length irrelevant control and paired iterative trajectories in Experiment 2. The improvement came from making the causal question narrower.
The researched context comparison did not confirm Experiment 2's advantage¶
Experiment 11 applied the same external research budget and blinded comparative endpoint to all 60 Experiment 2 terminal proposals. Relevant mechanisms lost 9–11 to archetype-only context, beat the irrelevant control 12–8, and reached only 52.5% pooled preference; the exact omnibus result was p = 0.4459. The internal-quality finding remains part of the record, but it cannot now be described as an externally researched endpoint advantage. This negative replication is consequential because it narrows one of the program's most tempting causal claims.
Substrate denial changed outputs but did not preserve their average quality¶
Experiment 12 rejected both extreme interpretations of the earlier substrate pattern. The model was not trapped: 67/72 outputs complied with the non-governance, non-computational rule. But greater effort did not erase the cost, and the constrained lane remained far behind in blinded quality at both High and Max. This changes the engineering target. The next problem is not stronger wording of the prohibition; it is learning when an alternative substrate is structurally appropriate and generating it without turning the source archetype into a decorative physical analogy.
Alternative-substrate diversity did not replace ordinary diversity¶
Experiment 13 gave the substrate lane the more defensible role suggested by Experiment 12: an additional proposal rather than a replacement. It generated ten unique researched survivors and rescued more P1 failures, but ordinary P2 still produced nine more incremental survivors overall. The constrained lane met the frozen bounded-complement gate only at the maximum permitted deficit. The result supports portfolio routing, not a general claim that deliberately unusual substrates are equally productive. A future implementation should therefore make this lane optional and budget-aware rather than mechanically spending it on every cell.
Applicability verification did not solve candidate retrieval¶
Experiment 14 prevented a strong verifier result from being mistaken for an end-to-end system result. The target route was classified accurately when supplied, yet the diagnostic search failed to place that target in the top five for 75 of 77 positive cases and underperformed the older solution-oriented index. Route logic and candidate generation are separate engineering problems. This failure motivated the prospectively frozen condition-level aggregation test in Experiment 15.
Route-aware similarity did not implement route logic¶
Experiment 15 rejected the most direct embedding-only repair. Five frozen aggregators represented conditions separately and respected the route graph's AND/OR shape, yet the confirmatory arm fell below both baselines. The post-hoc diagnostic explains why the design was structurally faithful in form but not in semantics: uncalibrated aggregation favored short routes, and cosine similarity treated a contradiction as highly related evidence rather than negative evidence. The engineering target is now narrower. Candidate retrieval needs an entailment-sensitive or otherwise signed condition scorer and explicit opportunity-count calibration; changing the averaging rule alone is not enough.
Closed-book success was a poor proxy for external distinctiveness¶
Experiment 3's 33 internally successful cells looked encouraging until external search found established or substantially colliding practice for 40 of 47 assessed selections. This was not wasted work. It revealed that prior art had to shape the scrutiny pipeline rather than merely annotate its output.
Retrieval-first search did not improve yield in the tested form¶
Experiment 4's retrieval-first arm consumed search and screening effort before a complete contrastive proposal existed, yet yielded only one strict candidate. The negative result does not condemn research-first workflows generally. It rejects this particular broad-hypothesis implementation as the default scaling path.
The four-proposal replication missed its frozen rescue floor¶
Experiment 6 missed by one cell. Changing the threshold after observing five rescues would erase the value of having frozen it. The report therefore keeps the negative verdict and separately reports the actual marginal yield. This is an example of why “not supported” should not be translated into “useless.”
Proposal-only judgment was reliable without being valid enough¶
Experiment 7's selectors often agreed with one another, but their preferred proposals did not monotonically correspond to researched survival. A candidate can read as unusually specific or useful while hiding a prior-art collision; another can read as awkward while containing a defensible empirical contrast. Agreement among language models cannot substitute for an external criterion.
No candidate cleared a no-close-prior-art label¶
All 59 dossiers sit beside prior art. This narrows the interpretation from invention ex nihilo to recombination, transfer, governance composition, and comparison. It is also consistent with a mature world in which useful primitives are widely distributed. The unresolved question is often whether a particular arrangement or control produces an advantage in a new setting.
The sentinel audit makes this negative result more informative. Eight of ten sampled dossiers acquired closer neighbors under explicitly historical, trade, non-English, and compositional searches even though none crossed the audit's substantial-collision rule. Surviving a label is not the same as having retrieved the best precedent.
8. The surviving opportunity portfolio¶
What is in the portfolio¶
The candidate appendix contains 59 experiment-level instances:
| Source | Endpoint | Dossiers |
|---|---|---|
| Experiment 3 | Post-hoc strict innovation-like survivor | 2 |
| Experiment 4 | Strict success | 6 |
| Experiment 5 | Strict success | 4 |
| Experiment 6 | Strict success | 15 |
| Experiment 6 | Empirical-partner candidate | 32 |
| Total | 59 |
These are not 59 independently verified inventions. They are 59 proposals that earned one of the listed statuses in their own protocol and have enough preserved evidence to support specialist review. Canonical titles were checked for exact duplication; none were identical. Conceptual overlap can still exist across candidates, and expert reviewers may merge or split opportunity families differently.
The appendix intentionally does not absorb every attractive earlier output. Experiment 2's 20 relevant-mechanism proposals passed a quality entry gate and received a post-hoc opportunity assessment, but they predated the later strict researched-candidate endpoint. Experiment 3 had 21 externally verified adopter-pipeline candidates, of which only two met the later post-hoc strict-like composite. Those broader sets remain available through the artifact index; silently mixing them into the 59 would make the portfolio look larger by changing its definition.
How to read a dossier¶
Each dossier begins with a short alias and preserves the original title. It then explains the problem, intervention, structural transfer, reason for advancement, nearest approaches, remaining contrast, smallest decisive test, resource bands, risks, and questions for appropriate experts. Direct links lead to the original proposal and evaluation.
The appendix is ordered by a post-hoc balanced score, with sensitivity to deployment-heavy and impact-heavy profiles. Near ranks should be read as ties. A strict endpoint remains strict wherever it appears; a partner candidate does not become strict because it ranks highly. Conversely, a low-ranked dossier is not disproven—its combination of evidence burden, implementation difficulty, or modest expected impact merely makes it a lower-priority expert review under the chosen weights.
Recurring opportunity forms¶
Although the domains vary widely, several families recur:
- Identity and lifecycle governance: preserve historical records while expiring their authority for current decisions.
- Representation-independent contracts: maintain a decision-relevant meaning across formats, tools, organizations, or physical transformations.
- Residual monitoring: model what should happen, then treat structured deviations as evidence requiring investigation.
- Bounded rivalry: use controlled competition among methods or candidate explanations without allowing the tournament itself to damage the system.
- Ritualized commitment and closeout: make transitions, obligations, uncertainty, or recovery explicit through repeated governed observances.
- Computability and assurance boundaries: distinguish what a system can prove, what depends on outside oracles or human choices, and what remains unknown.
These recurrences are not automatically new mechanisms missing from the Encyclopedia. Some are domain-specific instantiations of documented archetypes; some combine familiar components. The corpus may nevertheless reveal where the Encyclopedia needs a better domain example, a more precise composition, or a candidate mechanism for later curation.
What expert review should decide¶
The next reviewer is not being asked, “Do you like this idea?” A useful review should decide:
- whether the problem occurs at a consequential frequency or scale;
- whether the proposed baseline is the strongest fair comparator;
- whether a close system was missed;
- whether the intervention changes the causal structure as claimed;
- whether the test can be run with valid data and authority;
- what safety, legal, ethical, and distributional constraints were omitted; and
- whether to reject, revise, merge, observe, or run the bounded next test.
The expert review template turns those questions into a comparable record.
9. Resource use and scale economics¶
Measured Experiments 9 and 11–15 resources¶
Experiment 9's breadth probe made 300 valid transport attempts—one generator and one four-source light screen for each of 150 cells. It recorded 52,876,667 input tokens, of which 40,308,736 were cached, 1,252,720 output tokens, 376,597 reasoning-output tokens, and 40,022 seconds (11.1 hours) of summed agent time. With concurrent workers, its proposal and screen phases took approximately 3 hours 44 minutes of execution wall time. The experiment's relatively high throughput reflects its deliberately coarse endpoint.
Experiment 11 made 231 transport attempts to complete 60 eight-source external evaluations, 40 primary comparative judgments, and nine tiebreak judgments. It recorded 90,356,643 input tokens, of which 75,299,072 were cached, 1,169,103 output tokens, 466,307 reasoning-output tokens, and 32,682 seconds (9.1 hours) of summed agent time. Ninety-one attempts were invalid or failed and three lacked exposed usage, largely because a transport adapter rejected otherwise repairable structured outputs; accepted scientific artifacts still passed the canonical validator, and both recovery amendments are preserved. These figures show why a tightly controlled comparison can be expensive even when it generates no new proposals.
Experiment 12 completed 281 scientific model calls with no failed transport calls. It recorded 44,256,895 input tokens, of which 29,580,032 were cached, 1,337,531 output tokens, 660,600 reasoning-output tokens, and 43,522 seconds (12.1 hours) of summed call time. The 102 authorized web screens consumed most input tokens, while the 36 Max proposal calls accounted for 4.1 summed hours and unusually large constrained-output reasoning. This reinforces the broader cost finding: producing a different-looking proposal is cheaper than determining whether it remains distinct and useful after research (resource summary).
Experiment 13 completed 458 scientific calls with no failed transport calls: 180 proposal calls, 90 batched blinded-measurement calls, eight disagreement-adjudication calls, and 180 authorized opaque web screens. It recorded 80,429,441 input tokens, of which 54,841,600 were cached, 1,715,925 output tokens, 618,283 reasoning-output tokens, and 51,426 seconds (14.3 hours) of summed call time. Its 13.3-hour first-to-last telemetry span includes the generation phases, a pause for payload-specific web authorization, and the final screen run; it is not continuous active execution time. The comparison demonstrates the cost of measuring a portfolio policy properly: the 120 competing P2s were only one part of the workload, while blinding, duplicate measurement, adjudication, and complete web scrutiny supplied the evidence needed to interpret their incremental value (Experiment 13 results).
Experiment 14 completed 215 scientific model calls with no failed final calls: case construction, duplicate blinded case audit plus adjudication, and 114 blinded DNF verifications. It recorded 3,584,592 input tokens, of which 1,880,832 were cached, 912,450 output tokens, 236,247 reasoning-output tokens, and 18,501 seconds (5.14 hours) of summed call time. Three concurrent workers reduced its first-to-last telemetry span to 1.79 hours. The much lower input total reflects local deterministic index retrieval and the absence of public-web research; this was a controlled applicability benchmark, not a prior-art screen (Experiment 14 results).
Experiment 15 recorded 83 successful model-response attempts to obtain 40 accepted case bundles, 24 accepted blinded-audit batches, and two accepted adjudication batches. Seventeen responses were rejected by schema or deterministic validation and retried; there were no failed model transports in the canonical run. It recorded 1,329,255 input tokens, of which 457,728 were cached, 255,170 output tokens, 110,461 reasoning-output tokens, and 5,618 seconds (1.56 hours) of summed call time. Three concurrent workers reduced the first-to-last telemetry span to 32.5 minutes. All seven retrieval arms were local and deterministic, and no public-web research was used. The lower cost came from not repeating E14's 114 per-case DNF verifier calls (Experiment 15 results).
Measured Experiment 6 resources¶
Experiment 6 is the best basis for scaling calculations because it used the four-proposal pipeline on 60 cells and recorded exposed usage for all 591 scientific calls. Summed call time is not the same as elapsed calendar time because three workers ran concurrently.
| Resource | Recorded amount |
|---|---|
| Scientific calls | 591 |
| Input tokens | 205,117,798 |
| Cached input tokens | 165,694,464 |
| Uncached input tokens | 39,423,334 |
| Output tokens | 5,433,719 |
| Reasoning output tokens | 1,782,582 |
| Summed call time | 102,071 seconds (28.4 hours) |
External evaluation dominated the recorded workload: 243 calls, 25.3 million uncached input tokens, 2.22 million output tokens, and 16.7 summed hours. Proposal generation itself grew across positions because later calls had to preserve and avoid earlier ideas. The full breakdown is in resource accounting.
What Experiment 7 says about cost cutting¶
The three-replication selector used 180 calls, 3.57 million uncached input tokens, and 294,142 output tokens. Under the top-two policy it would have avoided evaluation of 120 proposals, but retained too few downstream positives. Net resource savings are therefore inseparable from the value assigned to false negatives. A selector that is inexpensive but discards six of 15 strict candidates is not automatically efficient.
One selector replication was much cheaper than three and may be appropriate for non-destructive queue ordering. Position-based rules are even cheaper. Neither has been prospectively shown to meet a high-recall requirement.
Scaling scenarios, not forecasts¶
The measured per-cell averages from Experiment 6 are approximately 9.85 scientific calls, 657,000 uncached input tokens, 90,600 output tokens, and 28.4 minutes of summed call time. A naïve linear projection to 70,000 cells would therefore imply roughly 689,500 calls, 46.0 billion uncached input tokens, 6.34 billion output tokens, and 3.8 years of summed call time. With three continuously occupied workers, the last figure is roughly 1.26 years before failures, rate limits, maintenance, and changing search conditions. Experiment 9 demonstrates a much cheaper breadth pass—two calls per cell—but its light-screen survivors are not interchangeable with Experiment 6 strict or empirical-partner outcomes. It is best understood as a possible allocation layer, not a validated substitute for deep scrutiny.
Those numbers are order-of-magnitude scenarios, not a budget quote. Caching behavior, context length, model and search pricing, concurrency limits, model upgrades, and protocol improvements would materially change them. Subscription usage also cannot be converted honestly into API expense without the applicable product contract. The defensible economic conclusion is simpler: exhaustive search is technically conceivable for a well-resourced organization, but external scrutiny—not raw idea generation—is the dominant cost, and the current evidence does not justify an immediate 70,000-cell run.
For an API deployment, a reader can insert the prices applicable to the chosen model and contract into the transparent approximation: 46,000 × uncached-input price per million tokens + 6,340 × output price per million tokens, then add cached-input charges, web-search or retrieval fees, storage, retries, engineering, monitoring, and expert review. This formula is deliberately not populated with a public list price that may not correspond to the experimental product or remain current when the report is read.
Where the cost actually sits¶
The projection above applies the Experiment 6 pipeline to every cell. No one would run it that way. The realistic shape is the two-stage funnel this program's own results suggest: an Experiment 9-grade breadth pass across the whole matrix, then the Experiment 6 pipeline on a selected fraction. Costing that shape from the same measured per-cell averages, and using a deep stage of 10,000 cells purely as a round illustrative figure:
| Stage | Cells | Scientific calls | Uncached input | Output + reasoning | Summed call time |
|---|---|---|---|---|---|
| Breadth pass, Experiment 9 protocol | 70,912 | 141,800 | 5.9 billion | 770 million | 5,260 hours |
| Deep pipeline, Experiment 6 protocol | 10,000 | 98,500 | 6.6 billion | 1.20 billion | 4,730 hours |
| Two-stage total | 240,300 | 12.5 billion | 1.97 billion | 9,980 hours | |
| Undifferentiated deep pass, for comparison | 70,912 | 698,500 | 46.6 billion | 8.53 billion | 33,510 hours |
Staging reduces uncached input by a factor of 3.7 and summed call time by 3.4. At three continuously occupied workers the two-stage total implies roughly 139 days of execution rather than the 1.26 years implied by the undifferentiated projection. That is the difference between an infeasible program and an expensive but unremarkable one.
The table conceals this program's largest unvalidated dependency. It assumes something reduces 70,912 cells to about 10,000, a retention near 14%. The Experiment 9 light screen does not do this: it passed 60 of 72 random-stratum cells, an 83.3% rate that would forward roughly 59,000 cells to deep scrutiny and recover almost none of the saving. Reaching a 14% retention requires a far stricter selector, and Experiment 7 is the only prospective evidence this program holds about cheap selection—where it failed all four prespecified retention requirements. The economics of a full-corpus run therefore rest on a component that has been tested once and did not work. Improving that selector is a cheaper research target than improving the generator, and it determines whether the rest of this arithmetic is reachable at all.
Even granting the funnel, the model-side figures are not the binding constraint. Suppose the deep stage is ranked and its strongest 500 candidates are sent to appropriate specialists. At 1.5 to 3 hours of genuine review per dossier—reading the evidence record, judging whether the comparator is the strongest fair one, and checking for a missed close system—that is 750 to 1,500 reviewer-hours, or roughly 19 to 38 person-weeks of specialist attention, distributed across the twenty or more domains the candidates occupy. For comparison, the initial Experiments 1–13 program, including its design freezes, blinded coding, adjudication, and first complete writeup, was executed in approximately six calendar days by one person, as the preserved run timestamps show; Experiments 14 and 15 were added later. Reviewing a selected 500 outputs would cost on the order of twenty times the human effort that produced that initial evidence base.
That is this section's central finding, and it inverts the intuition the token counts invite. Generation was never the expensive part. No experiment in this report was limited by compute; the limits were on human judgment, and the largest acknowledged limitation—that no candidate has been assessed by a domain expert—is the direct consequence. A reader deciding where to spend next should read the token tables as evidence that the cheap half of the problem is solved and the expensive half has not been started.
What the cost structure does not settle¶
It is tempting to close the loop with an expected-value argument: if one candidate in several hundred proved worth a large sum, a full run would repay itself. This report cannot support that argument, for two reasons that are unmeasured here rather than merely uncertain.
First, no candidate's value has been estimated. None has been reviewed by a domain expert, and every dossier sits beside adjacent prior art. The distribution of candidate value is not simply unknown—nothing in this program was designed to sample it.
Second, value realized is not value captured. These candidates propose interventions in other parties' domains, generally requiring their data, authority, and operating consent. Value created by an adopting organization does not accrue to the operator of the search absent a specific vehicle, such as a defensible claim, an operating entity, a consulting relationship, or a research collaboration—and universal adjacent prior art makes the first of those harder rather than easier. An economic case must therefore state a capture rate, not only a value estimate.
A third consideration cuts the other way, and is worth naming because it bears on selector design rather than on the budget. If the value of innovation opportunities is heavy-tailed, as it appears to be in adjacent settings such as patent-value and venture-portfolio distributions, then expected return is dominated by rare extreme outcomes rather than by the median candidate. Under that assumption the objective at scale is not to maximize average candidate quality but to maximize independent draws while avoiding filters that truncate the tail. Every selection stage in the present pipeline—the light screen, the strict lane, the post-hoc composite ordering—was tuned to remove weak candidates, and none was designed or tested for tail preservation. That assumption is imported from outside this program rather than established by it, and it is offered as a design consideration for future work, not as a finding.
10. Interpretation and implications¶
A search capability, not an invention oracle¶
The program demonstrates a way to turn a corpus of abstractions into a systematic search instrument. That matters because ordinary prompting often produces isolated analogies with no declared denominator, no preserved failures, and no way to distinguish transfer from lucky recall. Here, the system can be asked to cover a matrix, generate several alternatives per cell, and expose where candidates fail.
The severe qualification is equally important. The generator is best understood as a hypothesis proposer embedded in a research workflow. Its apparent creativity includes retrieval from pretrained knowledge, recombination of familiar components, and translation into new settings. The external evidence record repeatedly showed that fluency and specificity were compatible with rediscovery. The method becomes scientifically interesting not when the model says something surprising, but when a proposal survives a comparator-aware attempt to disprove it and still points to a feasible decisive observation.
Cross-domain transfer can be scaffolded¶
The experiments do not isolate an internal psychological process in the model. They show that transfer-like output can be scaffolded by an explicit source representation, a target domain packet, structural mapping requirements, diversity pressure, and iterative evaluation. Experiment 2's irrelevant-mechanism control is evidence that relevant structure contributed to quality. Experiments 3–7 then show that this contribution is insufficient without search and selection.
This suggests a useful engineering perspective. Rather than asking whether a model “has” analogical transfer as a single latent faculty, one can decompose the task into source representation, target retrieval, mapping, adaptation, evaluation, and evidence acquisition. Different tools—and humans—can own different stages. The Encyclopedia is especially valuable at the source-representation stage because it supplies explicit mechanisms and archetypes rather than requiring the model to infer every source structure from prose. Experiment 11, however, shows that this plausible scaffold cannot yet be credited with a general researched-endpoint advantage over archetype-only context.
The output may be valuable before deployment¶
A candidate can create value without becoming a startup, product, or intervention. It can:
- reveal a problem that a domain expert recognizes but has not formalized;
- supply a new comparator or measurement design;
- connect specialists who use similar structures under different names;
- identify an archival, governance, or safety control worth adding to an existing system;
- inspire a better human reformulation; or
- show that an apparent Encyclopedia gap is actually an established domain instantiation.
This is why expert response should preserve revisions rather than reduce every dossier to approve/reject. A domain specialist may move the idea to a more consequential setting, replace a weak baseline, simplify the intervention, or identify a known term that collapses the novelty claim while preserving practical value.
The most useful candidates may be prompts for co-creation¶
The “counterfactual discovery” discussion that emerged from the library recommendation candidate illustrates this possibility. The original proposal concerned feedback between discovery exposure and later recommendation signals in a digital library. A human reader immediately recognized a potentially wider relevance to music, film, e-commerce, and streaming systems, where prominent placement can manufacture some of the popularity later used as evidence of relevance. That extension is not an experimental result or a validated product idea. It shows how a structurally explicit candidate can become shared working material for a human who supplies context, priorities, and reframing.
The same interaction may occur in the opposite direction. Experts can identify why a transfer fails: the supposedly equivalent variables may not be manipulable, the authority structure may be wrong, or the proposed signal may be an artifact. Those failures are useful training and curation data if preserved.
11. Limitations and threats to validity¶
No independent domain-expert validation¶
The largest limitation is simple: the 59 dossiers have not been adjudicated by the diverse specialists needed to assess them. External sources establish that relevant problems, methods, and institutions exist; they do not substitute for tacit operational knowledge. No candidate should enter a live consequential workflow on the strength of this report.
Bounded search cannot establish absence¶
Search coverage varied with query construction, indexing, accessible language, paywalls, terminology, and the web's representation of practice. An evaluator can fail to find a close system that exists in a proprietary workflow, patent, non-English publication, local policy, or different vocabulary. ADJACENT_PRIOR_ART is therefore a statement about a bounded evidence record, not about the world.
Model-family and evaluator dependence¶
The program used related contemporary language-model systems for generation, criticism, retrieval orchestration, evaluation, and later editorial synthesis. Fresh sessions, blinded keys, deterministic validators, and external sources reduce some dependence but do not create fully independent human judgment. Models can share training data, stylistic preferences, blind spots, and an attraction to elaborate governance structures. Experiments 14 and 15 are especially exposed: the graph, constructed cases, audits, and E14 verifier all passed through related model-family systems. Future work should cross models and include human evaluators.
Adaptive research program¶
The thirteen linked experiments were designed sequentially in response to earlier results. That is appropriate for method development but creates researcher degrees of freedom across the program. Later experiments used design freezes and prespecified thresholds; earlier exploratory choices and the current 59-candidate synthesis should not be mistaken for one preregistered study. Twelve of the thirteen were prospective; Experiment 7 is the exception, and is described throughout as a blinded retrospective policy benchmark rather than a prospective preregistration. Experiments 9 and 11–15 were prospective tests added after the original seven-experiment arc. Experiment 12's question arose from the post-hoc substrate analyses, Experiment 13's policy comparison arose from E12, Experiment 14's retrieval benchmark arose after the Applicability Graph was built, and Experiment 15's frozen retrieval alternatives responded to E14's failure; their respective samples, inputs, endpoints, and decision rules were frozen before scientific generation or measurement.
Post-hoc hardening analyses¶
The prior-art sentinel, Experiments 8 and 10, E10B, and the E9 substrate follow-up were proposed after the relevant data or report questions existed. Their methods, samples, or blinded inputs were sealed before the relevant case work, outcome join, or new coding, which protects those steps from silent alteration. It does not make the questions prospective. The sentinel has only ten stratified cases; yield decomposition inherits purposive archetype and domain selection; the proposal-type analysis pools heterogeneous protocols; and the substrate analyses use a derived source score rather than a randomized treatment. E10B also followed an exploratory cross-tab. E9's balanced domain exposure and random 24-archetype stratum strengthen generalization within the declared frame, but its three domains remain purposive and its substrate question was chosen after generation. Experiments 12 and 13 supply prospective interventions on the bundle; they do not retrospectively make the observational question confirmatory or identify the base model as the sole cause.
Heterogeneous endpoints¶
The 59 candidates come from protocols with different samples, generation strategies, scrutiny stages, and endpoints. The report preserves their labels rather than estimating one pooled success probability. The post-hoc ranking harmonizes common dimensions for reading order only.
Candidate-instance dependence¶
Several proposals can come from one cell, archetype, domain, or conceptual family. Exact canonical titles are unique, but semantic independence has not been established. Counts describe candidate instances, not necessarily 59 unique mechanisms or market opportunities.
Sampling and domain taxonomy¶
The domains are a useful project taxonomy, not a universally accepted partition of human knowledge. Experiments sampled archetypes and domains for diversity, and Experiment 6 intentionally used a new set. Experiment 9 strengthens the archetype-side evidence by probability-sampling 24 previously untested generated archetypes, but its three domains were fixed purposively and its endpoint was a light screen. Experiment 12 inherits both strengths and limitations: its archetypes represent the sampled generated corpus, while its domain differences do not estimate all-domain effects. Experiment 13 independently probability-sampled both 12 previously untested generated archetypes and five domains outside the E6/E9/E12 union, strengthening generalization within the declared frames. Its 12 archetypes remain a small cluster sample, its five domains are not a probability sample of all human activity, and the cells are not sampled real-world problems. Experiments 14 and 15 sampled the declared route-bearing graph by stratum, but their scenarios were generated to instantiate those routes and therefore do not estimate performance on naturally occurring problems. Yield cannot be extrapolated to the full Cartesian product without modeling archetype, domain, proposal position, scrutiny depth, and selection effects.
Construct-derived applicability benchmark¶
Experiments 14 and 15 constructed positive and near-miss cases from the exact graph under test. Outcome-blind auditing, disjoint archetype samples, and label concealment reduce trivial leakage, and the near misses make the tests stricter than paraphrase recognition, but both benchmarks remain endogenous. E14's 93.2% balanced accuracy shows that the written predicates can be operationalized consistently by related models under controlled conditions; E15 shows that raw embedding similarity does not recover that entailment relation. Neither result shows that the predicates are complete, necessary, sufficient, or causally correct in real settings. Independently collected cases and specialist adjudication are required for that claim.
Prompt, corpus, and time dependence¶
The Encyclopedia was actively evolving, and experiments used source snapshots where specified. A more complete corpus may change the quality or specificity of transfer. Model behavior, search results, prices, and public prior art also change. Reproduction should preserve both the relevant corpus and a coverage date.
Rubric validity¶
The opportunity dimensions capture practical considerations but remain ordinal expert judgments produced by models. Weights reflect policy preferences. A high score can reward well-documented, governable ideas and disadvantage speculative high-upside science. Sensitivity profiles expose some but not all value judgments.
Cost measurement¶
Token telemetry was incomplete in some earlier experiments. Summed call time includes parallel calls and is not wall-clock duration. Cached and uncached tokens have different economic meanings. Cost bands are rough resource equivalents, not quotes or audited budgets.
No direct end-to-end baseline¶
Experiment 11 now provides an otherwise matched externally researched comparison of Experiment 2's relevant, archetype-only, and irrelevant-mechanism arms; it did not support a relevant-mechanism advantage under its frozen gate. The program still lacks an end-to-end comparison against ordinary free ideation, a different analogy library, or a human innovation team. The Encyclopedia's contribution to researched survivor yield, relative to those alternatives, therefore remains incompletely isolated.
12. Directions for human–machine research¶
1. Prospective expert evaluation¶
Recruit specialists for a stratified sample of strict and empirical-partner dossiers, including intentionally lower-ranked cases. Give them the original evidence record, not only the plain-language summary. Measure comprehension, missed prior art, problem prevalence, baseline adequacy, revision magnitude, recommendation, and inter-reviewer agreement. The most informative outcome is not an approval count alone but the transition from model proposal to expert-revised candidate.
2. Prospective field tests with stop rules¶
For candidates that retain support, run the smallest reversible study specified in the dossier. Preregister the comparator, outcome, authority, safety constraints, and termination thresholds. Treat a clean rejection as a successful test of the pipeline's falsifiability.
3. High-recall, cost-aware allocation¶
Experiment 7 ruled out the tested proposal-only top-two selector as a safe replacement for research. It did not exhaust sequential allocation. A future policy could perform a cheap search on all four proposals, use evidence-bearing signals to allocate deeper research, reserve one “wild-card” slot, and audit a random sample of discarded candidates. The objective should explicitly price false negatives rather than optimize call count alone.
4. Cross-model and corpus ablations¶
Repeat a frozen subset with different model families, no Encyclopedia context, archetype-only context, mechanisms without an archetype, and alternative analogy corpora. Hold the evaluator and retrieval budget fixed where possible. This would separate corpus value, prompt structure, model capability, and search intensity.
5. Compositional and higher-order problems¶
The present matrix mostly searched for first-order problems addressable by one archetype. Real systems often require several interacting solution structures—for example, a representation-independent contract plus residual monitoring and bounded rivalry. A compositional program could first diagnose several structural bottlenecks, then search a constrained graph of compatible archetypes. The combinatorial space is much larger, so composition should follow evidence of a multi-causal problem rather than enumerate every pair blindly.
This direction connects naturally to systems such as MOOSE-Star, which iteratively expands and filters research hypotheses. The Encyclopedia's possible advantage is breadth and explicit mechanism structure; the neighboring work offers ideas for tree search, novelty checks, and evidence-aware refinement. Comparative experiments should test that claim rather than assume it.
6. Synthetic data for transfer training¶
The preserved corpus can support a forward-looking training hypothesis. Each trajectory contains a source archetype, target domain, structural mapping, candidate problem, intervention, critique, revision, search evidence, comparator, disposition, and negative test. Positive and negative trajectories could teach several separable skills:
- retrieve a structurally relevant abstraction from a problem;
- infer a plausible target problem from a solution structure;
- distinguish relational transfer from surface analogy;
- adapt a mechanism without violating domain constraints;
- generate contrastive searches and falsifiers; and
- recognize collisions, uncertainty, and the need for external evidence.
This report does not show that such training will create a general cross-domain-transfer capability. A credible study would prevent leakage by holding out entire archetypes, domains, and archetype families; compare against equal-volume ordinary scientific and design data; evaluate both generation and rejection; and test whether gains transfer to human-authored problems outside the Encyclopedia taxonomy. Training examples should use preserved observable artifacts and concise rationales, not depend on undisclosed private chain-of-thought.
7. Encyclopedia curation feedback¶
Candidate families can be compared with the corpus to distinguish an absent mechanism from an absent example, a domain-specific instantiation, a recurring composition, or a naming problem. Only expert-reviewed patterns should enter the Encyclopedia. Failed mappings are also valuable: they can clarify an archetype's boundary conditions and improve future generation packets.
8. Human inspiration as an outcome¶
Measure whether a dossier helps a specialist produce a stronger idea than either the specialist or the model produces alone. A useful design would compare unaided experts, experts with ordinary brainstormed ideas, and experts with structurally mapped dossiers. Outcomes could include conceptual distance from the input, prior-art survival, test quality, revision effort, and the expert's own assessment of usefulness.
9. Public interface and living review¶
The eventual website can expose the report, raw data downloads, candidate filters, source records, and a versioned expert-review layer. Reviews should be attributable, dated, and non-destructive: new evidence can change a disposition without erasing what the pipeline originally produced.
10. Substrate-aware diversification¶
Experiments 12 and 13 support routing rather than prohibition. E13 supplied the first prospective second-slot policy test: the alternative-substrate lane added ten unique researched survivors but produced nine fewer incremental survivors overall and met its bounded-complement rule exactly at the allowed 15-point deficit. A future system should therefore generate ordinary diversity by default and route a limited alternative-substrate lane where missed-opportunity cost, natural substrate fit, or strategic portfolio breadth justifies the extra screen. The next study should test a cheap evidence-bearing router prospectively, preserve the ordinary P2 whenever recall matters, and prespecify archetype and domain strata rather than extracting favorable stories after the run.
11. Route-aware retrieval and external applicability validation¶
Experiments 14 and 15 identify candidate generation, not literal-level verification, as the immediate Applicability Graph bottleneck. E15 represented each condition separately and respected the route graph's AND/OR form, but raw cosine aggregation still failed because it favored short routes and could not distinguish satisfied from contradicted evidence. A next design should treat condition matching as a signed inference problem, calibrate scores across route length, route count, and query segmentation, and freeze those choices before another new sample. The stronger external-validity study should collect naturally occurring problems without using the graph to write them, ask independent specialists which routes genuinely apply, and measure both retrieval recall and false acceptance under that independent criterion. Given two low-recall internal benchmarks, external cases may be more informative than another large graph-generated run unless a materially different scorer exists.
13. Conclusion¶
The experiments began with an unusually simple reversal: start from a solution archetype and search backward for the problem it could solve. What made the project substantive was not the reversal alone, but the decision to operationalize it as a declared matrix and repeatedly try to break the resulting ideas.
The evidence supports cautious optimism. Relevant mechanisms improved internal terminal quality in Experiment 2, but Experiment 11 did not confirm an advantage after equal external research. Matrix-scale generation worked. Multiple complete proposals recovered candidates that a one-shot process missed. Experiment 9 found coarse researchability across every probability-sampled previously untested generated archetype, but its light screen cannot establish strict novelty or deployability. External search removed or narrowed most apparently promising ideas, and a cheap proposal-only selector failed to preserve enough of the researched yield. Post-hoc hardening exposed denser adjacent prior art, strong evidence-access bottlenecks, and a reproducible substrate house style. Experiment 12 then showed that this style is not a hard ceiling: 67/72 constrained outputs left governance and computation, although ordinary proposals retained a large quality and screen advantage. Experiment 13 tested the practical response on new sampled cells. Its substrate-diverse P2s added ten unique researched survivors and won 28/60 blinded quality comparisons, but ordinary P2s remained more productive, 41 incremental survivors to 32. Experiment 14 then showed that explicit applicability routes can support discriminating verification without automatically solving retrieval: the verifier passed its frozen gate, while the diagnostic candidate generator performed worse than the existing solution index. Experiment 15 tested the obvious condition-level aggregation repair on a disjoint frozen sample and rejected it: the primary route-aware arm found only 1/78 targets in its top five, below both baselines. The resulting system is therefore neither a novelty machine nor a reason to enumerate tens of thousands of cells without restraint. It is a configurable search-and-scrutiny pipeline that can produce concrete, falsifiable, evidence-linked opportunities for humans to judge, with a promising explicit verification layer and an unresolved candidate-retrieval bottleneck now localized to signed condition inference and score calibration rather than route representation alone.
The 59 dossiers are the testable public residue of that process. Their ultimate value will be decided outside this report: by experts who find missed precedents, by partners who supply unavailable measurements, by experiments that reject weak causal chains, and perhaps by a small number of cases that survive and become useful. Making that judgment possible—without hiding the failures—is the principal achievement documented here.
Data availability, authorship, and corrections¶
Data and code¶
The artifact index links the design freezes, prompts, schemas, raw records, source snapshots, analysis outputs, and scripts needed to audit each experiment and the post-hoc hardening work. The report's candidate index and fact table are generated from preserved artifacts by build_inverse_innovation_public_report_data.py. The Experiment 9 breadth probe, Experiment 11 external context comparison, Experiment 12 substrate-denial test, Experiment 13 second-slot policy test, Experiment 14 Applicability Graph test, and Experiment 15 route-aware retrieval test include prospective design freezes, raw records, and machine-readable results. The sentinel audit, yield decomposition, proposal-type analysis, and substrate analysis each include frozen methods or coding plans and machine-readable results. The plain-language dossiers are editorial derivatives; original proposals and evaluations remain authoritative.
Publication quality control¶
The publication QA register reports all eleven release gates from the approved report design as passed, partial, pending, or not applicable, with evidence for each status. Automated count, lineage, artifact-target, and local-link checks are distinguished from semantic review. In particular, valid URLs and complete source objects do not mean that every external citation has received independent claim-level verification.
Contributions and AI assistance¶
The human author of the Encyclopedia of Abstractions originated the inverse-innovation concept, developed the Encyclopedia that made the matrix possible, made the sequential research decisions, approved external search, selected interpretation boundaries, and retains responsibility for the report. OpenAI and Anthropic language-model systems were used, at different stages, as generators, critics, revisers, search orchestrators, evaluators, analysts, and drafting assistants. Deterministic scripts handled schema validation, aggregation, hashing, and many frozen decisions. Exact run-level model and prompt records are available where captured in the experiment artifacts.
This is an independent project report, not a statement by the author's employer, OpenAI, Anthropic, or any cited organization. The report has not been peer reviewed.
Revision history and corrections¶
The report's revision history is preserved in the repository. Corrections should add a dated entry to CHANGELOG.md, preserve the earlier report artifact when a release is sealed, and distinguish factual correction from changed interpretation. Candidate dispositions should be amended with new evidence rather than silently overwritten.
References and further reading¶
The literature review contains the annotated bibliography and search strategy. Particularly close starting points include:
- AskNatureGPT and biological-strategy retrieval: Design Science, DOI 10.1080/09544828.2025.2481536
- Patent function-based opportunity discovery: Technological Forecasting and Social Change, DOI 10.1016/j.techfore.2015.04.012
- Purpose–mechanism knowledge for analogical design: ACM Transactions on Computer-Human Interaction, DOI 10.1145/3530013
- TRIZ as abstraction-mediated inventive problem solving: AutoTRIZ and the Altshuller Foundation's historical description
- Morphological search and cross-consistency assessment: Zwicky (1967) and Ritchey (2015)
- Emergent analogical reasoning in transformers: arXiv:2605.11258
- MOOSE-Star, an agentic scientific-discovery framework: arXiv:2603.03756
- Reinforcement learning for analogical discovery: arXiv:2510.02263
For experiment-specific references, use the linked external evaluations in the candidate dossiers.
Appendix A — Candidate dossiers¶
This appendix translates every strict, post-hoc strict-like, and empirical-partner candidate preserved for the public inverse-innovation report. It is designed so that a general reader can understand what each proposal actually means and a specialist can reach the underlying evidence.
These are research candidates, not validated inventions. All 59 have adjacent prior art. Strict success means that a proposal cleared its experiment's researched-candidate rules. Post-hoc strict innovation-like survivor is an exploratory Experiment 3 label. Empirical-partner candidate means that decisive evidence requires a named kind of data-owning or operational partner. None of these labels establishes world novelty, patentability, demand, safety, impact, or authorization to deploy.
Ordering method¶
The reading order is a post-hoc editorial aid, not an experimental endpoint. It reuses the common nine opportunity scores and adds a first-evidence affordability proxy to the balanced weighting profile documented in the Experiment 2 opportunity method. Deployment-heavy and impact-heavy profiles provide a rank range. Small differences should be treated as ties, and original endpoint labels always take precedence over rank. Full inputs are in the candidate index.
Quick index¶
| Rank | Candidate | Experiment | Endpoint | Domain | First evidence |
|---|---|---|---|---|---|
| 1 | Fairer Selection for a Festival Slate | 6 | Strict success | Film Media Production | \(10,000–\)50,000 |
| 2 | A Common Meaning for Evidence Records | 6 | Strict success | Criminology Forensic | \(10,000–\)50,000 |
| 3 | Truthful Limits for Safety Verification | 4 | Strict success | Engineering Design | \(10,000–\)50,000 |
| 4 | Keeping Test Norms Current and Traceable | 5 | Strict success | Psychology | \(10,000–\)50,000 |
| 5 | Protected Quiet Periods Between Neural Stimuli | 4 | Strict success | Neuroscience | under $10,000 |
| 6 | A Safer Accessibility Testing Challenge | 6 | Empirical-partner candidate | Human Computer Interaction | \(10,000–\)50,000 |
| 7 | Consistent Rules for Climate Signpost Alerts | 6 | Empirical-partner candidate | Futurism Foresight | \(10,000–\)50,000 |
| 8 | Remembering Why Financial Controls Exist | 6 | Empirical-partner candidate | Accounting Auditing | \(10,000–\)50,000 |
| 9 | Predictive Checks for Evidence Custody | 6 | Strict success | Criminology Forensic | \(50,000–\)250,000 |
| 10 | A Stable Contract for Justice Histories | 6 | Strict success | Criminology Forensic | \(10,000–\)50,000 |
| 11 | A Credible End to Financial Close | 6 | Strict success | Accounting Auditing | under $10,000 |
| 12 | Blind Testing for Shared-Service Cost Models | 6 | Empirical-partner candidate | Accounting Auditing | \(10,000–\)50,000 |
| 13 | Separate Release Effects from River Disturbances | 6 | Empirical-partner candidate | Environmental Climate | \(50,000–\)250,000 |
| 14 | A Storage-Independent Aircraft Maintenance Ledger | 6 | Empirical-partner candidate | Aviation Aeronautics | \(50,000–\)250,000 |
| 15 | Portable Rules for Drought-Stage Advice | 6 | Empirical-partner candidate | Environmental Climate | \(10,000–\)50,000 |
| 16 | Portable Rules for Delphi Rounds | 6 | Empirical-partner candidate | Futurism Foresight | \(10,000–\)50,000 |
| 17 | Stable Sampling Across Changing Data Systems | 6 | Empirical-partner candidate | Sociology Anthropology | \(10,000–\)50,000 |
| 18 | Renewing Nanomaterial Stewardship Commitments | 6 | Empirical-partner candidate | Nanotechnology | under $10,000 |
| 19 | Ending Mutual-Aid Mandates Truthfully | 6 | Empirical-partner candidate | Sociology Anthropology | under $10,000 |
| 20 | Safer Cleanup of Post-Production Assets | 3 | Post-hoc strict innovation-like survivor | Film Media Production | \(10,000–\)50,000 |
| 21 | Show What Actually Solved the Model | 5 | Strict success | Economics Finance | \(10,000–\)50,000 |
| 22 | Keep Wetland Records Stable Across Systems | 6 | Empirical-partner candidate | Environmental Climate | \(10,000–\)50,000 |
| 23 | Retire Outdated Course Guidance Safely | 4 | Strict success | Education Pedagogy | under $10,000 |
| 24 | Test Cold-Case Theories on Equal Terms | 6 | Empirical-partner candidate | Criminology Forensic | \(50,000–\)250,000 |
| 25 | Send Raman Surprises, Preserve Full Evidence | 6 | Empirical-partner candidate | Chemistry Materials | \(10,000–\)50,000 |
| 26 | Review Materials Results by Model Mismatch | 6 | Empirical-partner candidate | Chemistry Materials | \(50,000–\)250,000 |
| 27 | A Common Contract for Synchronized Takes | 6 | Empirical-partner candidate | Film Media Production | \(10,000–\)50,000 |
| 28 | Portable Lighting Cues With Behavioral Tests | 6 | Empirical-partner candidate | Film Media Production | \(10,000–\)50,000 |
| 29 | Common-Wafer Contest for One Fabrication Slot | 6 | Strict success | Nanotechnology | \(10,000–\)50,000 |
| 30 | A Stable Contract for Crystal Structures | 6 | Strict success | Chemistry Materials | \(10,000–\)50,000 |
| 31 | Show Reviewers Only Unexpected Proof Effects | 6 | Empirical-partner candidate | Mathematics | \(10,000–\)50,000 |
| 32 | A Recurring Reckoning for Airport Noise Promises | 6 | Strict success | Aviation Aeronautics | \(10,000–\)50,000 |
| 33 | Retiring Outdated Criminal-Record Labels | 6 | Strict success | Criminology Forensic | \(10,000–\)50,000 |
| 34 | Stress-Testing Flight-Control Downselects | 6 | Empirical-partner candidate | Aviation Aeronautics | \(50,000–\)250,000 |
| 35 | Tracking Aging Wetland Erosion Mats | 5 | Strict success | Environmental Climate | \(10,000–\)50,000 |
| 36 | Subtracting Recoater Self-Noise | 6 | Empirical-partner candidate | Chemistry Materials | \(50,000–\)250,000 |
| 37 | A Shared Pause for Audit Independence | 6 | Empirical-partner candidate | Accounting Auditing | \(10,000–\)50,000 |
| 38 | Honest Guarantees for Generative Art | 4 | Strict success | Art Aesthetics | \(50,000–\)250,000 |
| 39 | Turning Archival Loss into Stewardship | 6 | Empirical-partner candidate | Religious Studies Theology | \(10,000–\)50,000 |
| 40 | Reviewing Ritual Records by Their Differences | 6 | Strict success | Religious Studies Theology | \(10,000–\)50,000 |
| 41 | Renewing Ramp Stop-Work Support | 6 | Strict success | Aviation Aeronautics | under $10,000 |
| 42 | Expiring Old Training Evidence Safely | 5 | Strict success | Cognitive Science | \(50,000–\)250,000 |
| 43 | A Capped Prize for Catalyst Endurance | 6 | Empirical-partner candidate | Nanotechnology | \(50,000–\)250,000 |
| 44 | Send Foresight Surprises, Keep Full Records | 6 | Empirical-partner candidate | Futurism Foresight | \(10,000–\)50,000 |
| 45 | Residual-First Review of Forensic Timelines | 6 | Strict success | Criminology Forensic | \(50,000–\)250,000 |
| 46 | A Clinic for Cross-Field Lemma Handoffs | 6 | Empirical-partner candidate | Mathematics | \(10,000–\)50,000 |
| 47 | Choosing One Proof Route Fairly | 6 | Strict success | Mathematics | \(10,000–\)50,000 |
| 48 | Monitoring a Program’s Measurement Footprint | 6 | Empirical-partner candidate | Criminology Forensic | \(50,000–\)250,000 |
| 49 | Making Carbon-Budget Boundaries Visible | 6 | Strict success | Environmental Climate | \(10,000–\)50,000 |
| 50 | Reserve AI Capacity for Harm Inquiry | 4 | Strict success | Tech Ethics Ai Governance | \(50,000–\)250,000 |
| 51 | Audited Competition for Close-Review Priority | 6 | Empirical-partner candidate | Accounting Auditing | \(10,000–\)50,000 |
| 52 | Renewing Climate-Monitoring Responsibilities | 6 | Empirical-partner candidate | Environmental Climate | under $10,000 |
| 53 | Recurring Stewardship for Digital Legacies | 6 | Empirical-partner candidate | Human Computer Interaction | \(50,000–\)250,000 |
| 54 | Testing Audit Findings Before Awarding Leads | 6 | Empirical-partner candidate | Accounting Auditing | \(10,000–\)50,000 |
| 55 | Full-Cost Bidding for Preservation Access | 6 | Empirical-partner candidate | Religious Studies Theology | \(10,000–\)50,000 |
| 56 | Reviewing What Changed in Student Reasoning | 6 | Strict success | Religious Studies Theology | \(10,000–\)50,000 |
| 57 | Testing Hidden Patterns in Library Exposure | 3 | Post-hoc strict innovation-like survivor | Library Information Science | \(50,000–\)250,000 |
| 58 | A Governed Contest Between Ethnographic Explanations | 6 | Empirical-partner candidate | Sociology Anthropology | \(50,000–\)250,000 |
| 59 | Mapping Coupled City Budget Deadlocks | 4 | Strict success | Political Science | \(10,000–\)50,000 |
Band A — highest post-hoc review priority (ranks 1–10)¶
1. Fairer Selection for a Festival Slate¶
Canonical title: Campaign-Bounded Selection for a Film Festival Slate
In one sentence: A festival would limit lobbying and unequal access, standardize review, and choose qualifying films for their contribution to a declared slate, then compare that process with its ordinary workflow.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-01 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Bounded Rivalry Governance × Film Media Production |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 73.0/100 · rank range 1–6 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
The problem¶
A festival has fewer competition slots than eligible films, yet entrants may receive unequal access to programmers, deadlines, screening attention, exceptions, or opportunities to supply extra material. Private pitches, gifts, sponsor pressure, and intermediary relationships can influence visibility. Screeners may also carry uneven workloads or undisclosed conflicts. The process can therefore reward campaign resources and access rather than each film’s contribution to the festival’s stated purpose, while leaving entrants unable to challenge clear procedural errors.
What is proposed¶
For one competition section, publish the purpose, available slots, eligibility rules, accepted materials, review stages, conflicts, individual quality threshold, slate-level criteria, tie-breaks, and procedural appeal route before submissions open. Ban gifts, private selection pitches, undisclosed referrals, and entrant-initiated lobbying. Randomly assign each eligible film to two screeners with capped caseloads, record recusals, and obtain a third review when assessments diverge materially. Only films clearing the individual floor reach slate selection. The panel then chooses a feasible, nonredundant portfolio using disclosed considerations such as runtime balance, programmatic perspectives, and audience pathways. An independent reviewer audits sampled files. Appeals may correct process errors but may not replace curatorial judgment. Selection applies to one edition only.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a clearly delimited competition for scarce screenings. A published rulebook defines the permitted arena; material and contact limits dampen campaign escalation; workload controls, recusals, sanctions, and audits constrain manipulation; and portfolio selection directs rivalry toward contribution to the whole slate. Post-edition review supplies a revise-or-retire decision.
Why it advanced¶
Experiment 6 classified this candidate as STRICT_SUCCESS because the problem, plausible adopters, component feasibility, reversible shadow test, and contrast with ordinary selection were sufficiently supported for its strict researched-candidate bar. That status does not establish real-world effectiveness, distinctiveness, deployment authority, or economic impact; the proposal remains a partnered research program.
Prior art and the remaining open claim¶
Transparent rules, controlled submission materials, assigned viewers, second reviews, conflict recusals, authorized exceptions, and intentional slate composition already exist in festival practice. The open claim is narrower: compared with a festival’s current workflow, combining standardized access, random double review, workload ceilings, an individual floor, disclosed portfolio constraints, sampled audit, and process-only appeals will reduce treatment discrepancies and access-linked variation without unacceptable losses in curatorial fit, reviewer attention, or timeliness. That advantage has not been demonstrated.
Smallest decisive test¶
With one festival, preregister a shadow comparison on 60–120 opt-in short-film submissions. Apply ordinary selection and the proposed workflow to the same files without changing official outcomes. Compare missing reviews, time, agreement, recusals, expertise reassignments, threshold stability, portfolio reasons, audit errors, appeal corrections, slate overlap, concentration, and blind coherence ratings. Reject the claim if discrepancies do not fall, agreement worsens, expertise corrections exceed 20%, portfolio choices breach the floor, median labor rises over 50% without matching benefit, appeals reopen taste, or auditors cannot separate rule breaches from lawful curation.
Deployment and cost¶
The authorized first step is a consented shadow pilot, not a live selection change. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence and startup, and \(50,000–\)250,000 for operational launch and annual recurring work. These are bottom-up assessment bands, not vendor quotes; no festival partner has committed.
Risks and uncertainties¶
- Standard material limits may omit cultural, production, safety, or accessibility context needed to interpret a film responsibly.
- Portfolio criteria could become vague cover for favoritism or token selection, including admission below the declared quality floor.
- Double screening and audits could overload programmers, delay decisions, or force a smaller review pool.
- A communications firewall may still favor films already legible through credits, press coverage, or established intermediaries.
- Audit samples may miss selective misconduct or wrongly treat innocent intermediary concentration as evidence of wrongdoing.
Expert review¶
Useful reviewer backgrounds: Festival artistic director or programming lead, Festival operations and submissions manager, Independent procedural auditor, Film-sector conflict, privacy, and labor counsel, Curatorial evaluation or portfolio-design researcher.
- Can the festival define portfolio criteria precisely enough to guide decisions without disguising unrestricted discretion?
- What caseload ceiling permits substantive double review within the existing calendar and staffing budget?
- Can auditors reliably distinguish a correctable process violation from a lawful curatorial judgment?
- Which contextual materials must remain available so standardization does not systematically disadvantage particular films?
Evidence and provenance¶
Selected sources: S1: Film Festival Best Practices · S2: 2025 Sundance Film Festival Unveils Short Film Program Presented by Vimeo · S3: Submitting Your Project to the 2026 Sundance Film Festival: FAQ · S4: BFI London Film Festival Assistant Programmer Job Pack · S5: Festival Entry Regulations 2026 · S6: Terms of Submission—Official Selection 2026 · S7: Pre-Jury Transparency in Short Film Festivals Organized in Turkey within the Framework of Gatekeeping Theory · S8: Socioeconomic Factors of National Representation in the Global Film Festival Circuit: Skewed Toward the Large and Wealthy, but Small Countries Can Beat the Odds
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: In the post-hoc harmonized reading order, this candidate scored 73–74 across profiles, ranked between 1 and 6, and fell in band A. The ordering used cost-band affordability as a pilot proxy; it is neither an experimental endpoint nor an estimate of economic value.
2. A Common Meaning for Evidence Records¶
Canonical title: Representation-Independent Evidence Continuity Record
In one sentence: A forensic organization would define custody records by their permitted actions and observable meaning, then require any replacement system to pass the same behavior-based tests rather than merely matching fields.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-09 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Representation Independent Interface Contract × Criminology Forensic |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 73.0/100 · rank range 2–7 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-10, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Forensic evidence systems can encode custody, seals, sampling, consumption, and corrections in different tables, status codes, display orders, or editable notes. During migration or system integration, two implementations may contain similar fields yet report different current custodians, unresolved transfers, seal histories, sample ancestry, remaining quantities, or corrections. Without a representation-independent rule, an organization cannot determine whether a new implementation preserves the procedural meaning of the original record rather than merely copying its visible data.
What is proposed¶
Define an Evidence Continuity Record as an abstract state machine, not a database layout. Its public operations would register items, offer and accept transfers, record seal opening and resealing, derive child samples, record consumption, append linked corrections, query current state, render authorized history, and check named continuity conditions. Each operation would specify permissions, preconditions, results, errors, and side-effect limits. Invariants would prohibit unmatched completed transfers, negative quantities, unlinked samples, silent deletion, and more than one current accountable custodian. Storage tables, vendor codes, indexes, internal identifiers, and row order would remain non-contractual. A shared black-box suite would run examples, generated action sequences, authorization checks, and representation-leakage probes before an implementation could qualify as a substitute.
The cross-domain transfer¶
The representation-independent interface archetype becomes a behavioral contract for one evidence record. Abstract states and operations define what custody history means; an opaque boundary prevents clients from depending on vendor internals; and a common conformance oracle tests whether different implementations preserve the same authorized observations, errors, transitions, and invariants.
Why it advanced¶
Experiment 6 classified this as STRICT_SUCCESS because heterogeneous forensic systems, migration difficulty, relevant institutional authority, established software techniques, and a safe synthetic test were sufficiently supported for the strict researched-candidate bar. This does not validate the contract, authorize production migration, establish legal sufficiency, demonstrate distinctiveness, or measure economic impact.
Prior art and the remaining open claim¶
Custody standards, common schemas, ontologies, audit trails, access controls, two-sided transfers, and migration tools already cover much of the surrounding territory. None of that alone establishes behavioral equivalence between implementations. The remaining claim is that an opaque, vendor-neutral state-machine contract plus one behavior-only conformance and leakage suite will identify semantically acceptable substitutes more reliably than schema matching, audit-trail presence, or vendor-specific acceptance tests. It remains open until the suite detects seeded defects and survives review of material legal and scientific distinctions.
Smallest decisive test¶
With a laboratory quality authority and authorized reviewer, preregister 12–20 synthetic scenarios covering transfers, seals, samples, consumption, corrections, repeated calls, unauthorized queries, and exports. Run them against a transparent reference model, an incumbent-behavior adapter, and at least five mutated adapters containing specified semantic defects. Advance only if the first two agree on every mandatory assertion, expose no prohibited information, and the suite rejects every mutant. Revise or reject the contract if an incumbent sequence is unrepresentable, a mutant passes, or reviewers find a material contract-observable difference after both implementations pass.
Deployment and cost¶
Begin only with synthetic data and sandbox adapters; do not alter live evidence or infer admissibility. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup and annual work, and \(250,000–\)1 million for operational launch. They are assessment estimates, not vendor quotes.
Risks and uncertainties¶
- The abstract state may omit jurisdiction-specific signatures, native documents, instrument metadata, or another legally or scientifically material distinction.
- Tests may copy incumbent quirks and mistakenly preserve them as required behavior.
- A passing suite could be misrepresented as proof that recorded events are true or that evidence is admissible.
- Incomplete authorized export operations could make the opaque boundary obstruct legitimate audit or disclosure.
- Generated tests may miss failures that appear only in long, concurrent, or malformed action sequences.
Expert review¶
Useful reviewer backgrounds: Forensic laboratory quality manager, Evidence custodian or property-unit supervisor, Forensic LIMS architect or integration engineer, Evidence and disclosure counsel, Software conformance and property-based testing specialist.
- Which representation-specific artifacts are legally or scientifically material and therefore must be contract-observable?
- Can an adapter reproduce incumbent behavior without relying on undocumented or mutable vendor features?
- Do the proposed invariants cover every authorized transfer, correction, sampling, and consumption sequence in the bounded scope?
- What mutant set would provide a credible test that the suite distinguishes field similarity from semantic equivalence?
Evidence and provenance¶
Selected sources: S1: Evidence Management Steering Committee Report: Opportunities to Strengthen Evidence Management Processes · S2: Laboratory Information Management Systems in Forensic Science Service Provider Laboratories: Current State and Next Generation · S3: Landscape Study of Software-Based Evidence Management Systems for Law Enforcement · S4: OSAC 2021-N-0018: Standard for On-Scene Collection and Preservation of Physical Evidence, Version 2.0 · S5: ISO 22095:2020 — Chain of custody — General terminology and models · S6: CASE Ontology Design and Specification · S7: Privacy Impact Assessment for the Laboratory Information Management System · S8: Best Practices for Authenticating Digital Evidence
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading order scored this candidate 73–74, placed it between ranks 2 and 7, and assigned band A. That order used a cost-band affordability proxy and is not an experimental endpoint, legal finding, or measure of economic value.
3. Truthful Limits for Safety Verification¶
Canonical title: Assurance-Boundary Gate for Extensible Safety-Control Designs
In one sentence: A design organization would classify what a safety verifier can honestly decide before routing each controller to exact checking, bounded or approximate analysis, or accountable human review.
| Field | Record |
|---|---|
| Portfolio ID | EXP04-STRICT-02 |
| Experiment and endpoint | Experiment 4 · Strict success |
| Archetype × domain | Computability Boundary Mapping × Engineering Design |
| Proposal position or arm | PROPOSAL_FIRST |
| Post-hoc reading order | Balanced score 72.0/100 · rank range 1–4 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
An organization asks one verifier to return a correct, terminating SAFE or UNSAFE answer for every programmable safety controller, even when submissions include unrestricted scripts, unbounded variables, and supplier extensions. The actual computation model and guarantee are left unclear. Engineers may then treat timeouts as failures, keep adding computing power to a requirement that may be undecidable, or quietly restrict accepted designs while continuing to make a universal assurance claim. Release records may not reveal which bounds or assumptions support a verdict.
What is proposed¶
Place an assurance-boundary gate before verification and release. The gate would formalize the controller language, state representation, safety property, environmental assumptions, quantifiers, computation model, and requested guarantee. It would admit exact terminating verification only for mechanically enforceable finite or otherwise decidable fragments. Richer designs would go to sound over-approximation, explicitly bounded exploration, or accountable expert escalation. UNKNOWN, OUT_OF_SCOPE, TIMEOUT, and TOOL_FAILURE would remain separate from SAFE and UNSAFE. Claims of impossibility would require a checked reduction or equivalent proof, not repeated timeouts. Every result would record its fragment, bounds, abstraction, assumptions, guarantee, residual uncertainty, and triggers for reclassification when the model, tool, supplier interface, or environment changes.
The cross-domain transfer¶
Computability-boundary mapping becomes an intake and routing decision for controller assurance. The gate separates exact solvability, weaker machine-checkable evidence, and unresolved cases relative to an explicit model. Enforced language fragments and typed results prevent bounded searches, abstractions, tool failures, or nonanswers from inheriting a stronger universal safety claim.
Why it advanced¶
Experiment 4 classified this candidate as STRICT_SUCCESS because the boundary is technically grounded, component practices are mature, relevant authorities exist, and a reversible non-release comparison is clearly testable. Passing that strict researched-candidate bar does not show field prevalence, better safety outcomes, adopter acceptance, deployment authorization, economic impact, or broad distinctiveness.
Prior art and the remaining open claim¶
Finite-state model checking, abstraction, bounded exploration, UNKNOWN results, witnesses, and formal-assurance records are established practices. The narrower open claim concerns their mandatory ordering and governance: placing a model-relative solvability and scope classification plus an enforced status schema before an otherwise standards-conformant workflow will reduce over-strong or mislabeled verdicts and improve reviewer routing agreement, while losing no more than one conclusive decision in a 12-model pilot. The individual mechanisms are adjacent prior art, and comparative advantage has not been measured.
Smallest decisive test¶
Freeze 12 previously adjudicated controller models spanning finite, bounded, abstractable, unrestricted, unsafe, and timeout cases. Blind and randomize reviewer pairs to the strongest existing workflow or the proposed gate, using equal tools, evidence, and time. Compare adjudicated mislabeled verdicts, routing agreement, hours, nonanswer rates, and conclusive decisions. Advance only with zero mislabeled gate verdicts, improvement in mislabeling or agreement, at most one lost conclusive result, at most one fragment-membership disagreement, and cost below $50,000. Reject if weak or incomplete evidence becomes unqualified SAFE/UNSAFE, agreement fails to improve, or physical semantics cannot be reproduced.
Deployment and cost¶
The first step is a frozen, non-release pilot that cannot alter deployed logic or release status. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence and \(250,000–\)1 million for startup, operational launch, and annual recurring work. These are assessment estimates rather than vendor quotes or certification budgets.
Risks and uncertainties¶
- A correct proof may concern a model that omits important physical-controller or environmental behavior.
- An unsound abstraction could create false confidence; a coarse but sound abstraction could generate too many unresolved alarms.
- Designers may move excluded but safety-relevant behavior into informal escape hatches outside the decidable fragment.
- A finite analysis may exhaust resources, inviting staff to confuse infeasibility or timeout with a semantic verdict.
- The recorded boundary can become stale after language, verifier, environment, or supplier-interface changes.
Expert review¶
Useful reviewer backgrounds: Safety-control chief engineer, Independent formal-verification specialist, Design assurance or certification reviewer, Controller-language and toolchain engineer, Domain regulator or licensing specialist.
- Does the declared controller language actually include unrestricted computation, or is the accepted class already enforceably decidable?
- Can fragment membership and environmental assumptions be reproduced independently for every pilot model?
- Are the approximate-analysis modes sound, and how are spurious counterexamples labeled and escalated?
- Would the gate add information beyond the strongest existing standards-conformant workflow under equal time and tools?
Evidence and provenance¶
Selected sources: S1: IEC 61508-3:2010 — Functional safety of electrical/electronic/programmable electronic safety-related systems, Part 3: Software requirements · S2: Digital Instrumentation and Controls Research · S3: DOT/FAA/TC-19/22: Use of Virtual Machines in Avionics Systems and Assurance Concerns · S4: Explainable Verification for Rapid Certification · S5: What is Formal Methods? · S6: CBMC: What is loop unwinding? · S7: SV-COMP 2013 Competition Procedure and Verification Result Definitions · S8: National employment and wage data by occupation, May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: In the post-hoc harmonized reading order, this candidate scored 72–74, ranked between 1 and 4, and fell in band A. The ordering used affordability as a pilot proxy; it is not a preregistered endpoint, safety finding, or economic-value measure.
4. Keeping Test Norms Current and Traceable¶
Canonical title: Psychometric Norm Lifecycle Registry and Scoring Gate
In one sentence: A registry and scoring gate would keep context-mismatched norm packages out of new psychological scoring while preserving exact historical packages for authorized reconstruction.
| Field | Record |
|---|---|
| Portfolio ID | EXP05-STRICT-04 |
| Experiment and endpoint | Experiment 5 · Strict success |
| Archetype × domain | Layer Decay And Expiration Management × Psychology |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 72.0/100 · rank range 2–5 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A psychological test may accumulate several norm tables, scoring transformations, cutoff rules, and population-specific calibration files. Older packages can remain in manuals, spreadsheets, scripts, or software without a clear current, restricted, superseded, or archived status. Users may then score against an unintended population or edition, producing inconsistent standardized results. Simply deleting old packages is also unsafe because prior reports, longitudinal studies, audits, or research reproductions may depend on the exact historical transformation.
What is proposed¶
Create a lifecycle registry linked to a scoring gate. Give every norm package an immutable identifier and record its instrument version, reference population, collection period, intended uses, transformations, cutoffs, validation evidence, software dependencies, latest review, successor, and lifecycle state. A review lease or relevant context change would move a package to review-due, not declare it invalid. New scoring would require an active, context-matched package or a documented expert override. Superseded packages would disappear from ordinary prospective selection but remain retrievable by identifier for authorized reconstruction. Dependency checks would block removal while reports, studies, longitudinal series, or audits still need a package. Quarantine, archival tiers, disposition markers, and fixed-vector restore drills would make retirement reversible and test whether historical scoring remains executable.
The cross-domain transfer¶
Layer-decay management becomes governance for successive norm and scoring packages. The registry identifies each deposited layer, detects expired reviews and changed contexts, and separates active authority from historical retention. Dependency checks, quarantine, archives, deletion markers, and restoration tests prevent cleanup from destroying the scoring basis of earlier work.
Why it advanced¶
Experiment 5 classified this candidate as STRICT_SUCCESS because norm changes can matter, relevant professional guidance and adopter classes exist, the software pattern is feasible, and a bounded read-only comparison is testable. This strict researched-candidate result does not establish field effectiveness, cross-publisher portability, distinctiveness, deployment authorization, or economic value; partnered research remains necessary.
Prior art and the remaining open claim¶
Professional guidance already calls for current, population-relevant, identifiable norms; publishers already renorm tests, label legacy products, retire scoring software, and maintain successor platforms. The open claim is the added effect of integrating immutable package identity, expiration of prospective authority, context-gated selection, dependency-constrained retirement, and tested restoration. Compared with ordinary files, manuals, or platform presentation, that package should reduce context-mismatched selection and improve exact historical reconstruction without excessive valid-use blocks or expert overrides. This comparative claim remains untested.
Smallest decisive test¶
Use one licensed or synthetic test family with at least three packages and randomly assign 20–30 qualified evaluators to the current presentation or a read-only registry across 24–40 scenarios. Compare mismatched selections and exact reconstruction, plus time, blocks, overrides, identity errors, and preservation choices; restore one archived package against a fixed vector. Advance only with at least a 10-point mismatch reduction, 15-point reconstruction gain, no more than a 5-point increase in valid-use blocks, at most 10% unnecessary overrides, zero identity errors, and exact restoration. Redesign if either primary outcome fails or users confuse lifecycle status with validity.
Deployment and cost¶
Start with a read-only inventory and mock gate; do not change production scores, reports, or source files. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup and annual work, and \(250,000–\)1 million for operational launch. They are not vendor quotes.
Risks and uncertainties¶
- Users may mistake an active lifecycle state for evidence that a norm is psychometrically valid for the individual case.
- Incomplete population, instrument, or intended-use metadata could falsely authorize a mismatched package.
- Missing references from reports, spreadsheets, printed tables, or local scripts could allow destructive retirement.
- Archived transformations may become non-executable as software environments and formats age.
- Registry access or preservation policies could expose restricted test content or retain sensitive norm data longer than authorized.
Expert review¶
Useful reviewer backgrounds: Psychometrician responsible for test norms, Test publisher or assessment-program owner, Practicing psychologist or qualified assessment user, Scoring-platform and archival systems engineer, Test-security, privacy, and records counsel.
- Can package metadata express intended population and use precisely enough to gate scoring without implying validity?
- What events should trigger review-due status, and who may renew, restrict, supersede, or override a package?
- How completely can inbound dependencies from reports, studies, scripts, and longitudinal datasets be discovered?
- Can an archived package reproduce the fixed historical score in a controlled environment without exposing restricted materials?
Evidence and provenance¶
Selected sources: S1: Guidelines for Practitioner Use of Test Revisions, Obsolete Tests, and Test Disposal · S2: ITC Guidelines on Test Use, Final Version 1.2 · S3: Model for the Review, Description and Evaluation of Psychological and Educational Tests, Version 5.0 · S4: The Flynn Effect: A Meta-analysis · S5: The Woodcock Reading Mastery Test: Impact of Normative Changes · S6: NEO Personality Inventory-3 (Normative Update) · S7: Pearson Scoring Software · S8: Occupational Employment and Wages—May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading order scored this candidate 72–74, ranked it between 2 and 5, and placed it in band A. That order used cost-band affordability as a proxy and is neither an experimental endpoint nor a measure of economic value.
5. Protected Quiet Periods Between Neural Stimuli¶
Canonical title: Protected Null Epochs for Closed-Loop Neural Stimulation
In one sentence: Explicitly reserved, carefully labeled periods without stimulation could reveal whether apparent neural responses reflect the current stimulus or lingering effects from earlier ones.
| Field | Record |
|---|---|
| Portfolio ID | EXP04-STRICT-05 |
| Experiment and endpoint | Experiment 4 · Strict success |
| Archetype × domain | Negative Space Design × Neuroscience |
| Proposal position or arm | PROPOSAL_FIRST |
| Post-hoc reading order | Balanced score 71.0/100 · rank range 2–6 across three profiles · band A |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
The problem¶
A closed-loop neural experiment may deliver another stimulus whenever timing and safety rules allow, treating unused time as wasted. If the brain has not returned to a stable reference state, responses can overlap with residual activation, adaptation, or changing neural state. Researchers may then attribute a response to the current stimulus when part of it came from preceding trials. Sparse or ambiguously recorded quiet periods also make scheduling histories harder to compare and reproduce.
What is proposed¶
The experiment generator would omit selected low-priority stimuli at genuine block or state boundaries and protect the resulting null epochs from automatic backfilling. Recording, synchronization, contextual logging, and safety monitoring would continue throughout each pause. Codes would distinguish an intentional pause, a recovery-triggered pause, an operator decision, and equipment or command failure. Stimulation would resume at a predeclared time or when an approved recovery measure reached a bounded criterion, with a maximum pause and manual override. The comparator is a dense schedule paired with the best history-dependent response model; additional comparisons include a uniformly longer interval and randomized omission. The aim is to separate stimulus effects from sequence-history effects without sacrificing unacceptable condition coverage.
The cross-domain transfer¶
The negative-space archetype becomes deliberately protected time in the experiment schedule. The omitted event creates a measured reference interval connected to the next response, while explicit boundaries and state codes give the absence a clear meaning. The mapping is strong: the empty space is an active experimental condition, not merely slower scheduling or missing data.
Why it advanced¶
This candidate passed Experiment 4's strict researched-candidate bar because the problem is measurable, offline testing is feasible, the intervention has explicit safety and rollback boundaries, and its strongest rivals are directly comparable. STRICT_SUCCESS denotes passage of that research screen only; it does not establish real-world effectiveness, novelty, deployment permission, or economic value.
Prior art and the remaining open claim¶
Null events, washout periods, adaptive stimulation, history-dependent models, and audit logs already exist, so this is adjacent prior art. The narrower open claim is that recovery-bounded, explicitly coded null epochs outperform dense scheduling with model correction, uniformly longer intervals, and random omission. Any advantage must remain after matching stimulated-trial count and elapsed time and must be large enough to justify reduced condition coverage.
Smallest decisive test¶
Using one synchronized historical session, preregister a response model containing stimulus identity, recent history, elapsed time, and prestimulus state. Estimate recovery without held-out trials, then replay fixed-gap, recovery-threshold, and matched-random-omission policies. Compare held-out response error, stimulus-versus-history identifiability, baseline stability, coverage, and simulated duration against the dense schedule with its best history model and a uniformly longer interval. Reject the proposal if recovery-based spacing does not improve the preregistered identifiability measure, random omission or modeling matches it, coverage falls below its floor, or recovery estimates cannot support bounded resumption.
Deployment and cost¶
The authorized first step is offline analysis; live stimulation requires separate ethics, institutional-safety, and possibly device-regulatory approval. Estimated 2026 resource-equivalent bands are under $10,000 for first evidence, \(10,000–\)50,000 for initial deployment, \(50,000–\)250,000 for operational launch, and \(10,000–\)50,000 annually. These are assessment bands, not vendor quotes.
Risks and uncertainties¶
- Fewer stimulated trials may reduce statistical precision or leave required conditions underrepresented.
- Replacing omitted trials could lengthen sessions, increasing participant fatigue or animal burden.
- A recovery-triggered rule could preferentially sample particular neural states and introduce selection bias.
- Faulty state codes could make equipment failure look like intentional silence.
- An unreliable recovery signal could repeatedly suppress conditions or destabilize the closed loop.
Expert review¶
Useful reviewer backgrounds: Closed-loop neural-stimulation researcher, Neural time-series statistician, Research ethics or animal-care specialist, Neurotechnology safety engineer, Experimental-design methodologist.
- Can the available event-level data distinguish recovery dynamics from ordinary time drift and recent stimulus history?
- What recovery measure and maximum pause could be specified without using held-out outcomes?
- Which conditions must remain above a prespecified coverage floor after omissions?
- Does model-based correction match the proposed schedule after trial count and elapsed time are equalized?
- Which approvals would a live pilot require for the specific device, population, and protocol?
Evidence and provenance¶
Selected sources: S01: Learning to Control the Brain through Adaptive Closed-Loop Patterned Stimulation · S02: Influence of Inter-Stimulus Interval on Auditory Evoked Potentials · S03: Stochastic Designs in Event-Related fMRI · S04: A Deep Brain Stimulation Trial Period for Treating Chronic Pain · S05: Subcortical Short-Term Plasticity Elicited by Deep Brain Stimulation · S06: Neural Recording and Modulation · S07: IDE Approval Process · S08: Center for Research Informatics Price Guide, FY26
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: A post-hoc harmonized score placed the candidate between ranks 2 and 6 across three reading profiles, in band A. This ordering is not an experimental endpoint or economic-value measure; its pilot-speed input is only a cost-affordability proxy.
6. A Safer Accessibility Testing Challenge¶
Canonical title: Accessibility Barrier Discovery Portfolio Challenge
In one sentence: A sealed, batch-scored challenge would reward a complementary set of reproducible accessibility barriers instead of whichever reports arrive first.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-06 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Human Computer Interaction |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 71.0/100 · rank range 6–9 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
An accessibility bounty with a fixed reward pool can reward speed or report volume rather than useful discovery. Testers may submit scanner output, reserve findings before documenting them, split one barrier into several reports, or withhold reproduction details. Reviewers then spend time resolving duplicates while difficult, task-blocking barriers remain poorly described. Unequal automation, unsafe testing, and access to live or personal data can also shift costs onto users, maintainers, and less-resourced testers.
What is proposed¶
Run bounded rounds in a synthetic-data environment with equal access windows, scoped accounts, submission and request caps, sealed reports, and published rules. Reports must describe the affected task, starting state, interaction sequence, observed barrier, expected behavior, environment, reproducible evidence, and remediation-relevant trace. Authorization, privacy, fixture integrity, and reproducibility are mandatory gates. Verified reports are scored for task obstruction, clarity, reproducibility, distinctness, and repair usefulness. Awards are selected as a portfolio covering complementary journeys, interaction modes, and failure classes, rather than paid in filing order. Independent verification, separate appeals, predefined foul rules, recurring challenger entry, and post-round review address favoritism, sabotage, collusion, lock-in, and rubric gaming.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a controlled contest for a finite accessibility reward pool. Rivalry remains, but a rulebook, equal resource ceilings, sealed batches, safety gates, portfolio awards, appeals, and recurring access limit how contestants can gain advantage. The structural mapping is direct, although whether rivalry adds value over collaboration remains untested.
Why it advanced¶
This candidate cleared the separate EMPIRICAL_PARTNER_CANDIDATE lane because a controlled mock study can compare allocation systems without touching production. It did not pass the strict-success lane. Crucially, there is no field evidence that accessibility-specific rivalry improves coverage over a well-run paid panel, and no organization has committed authority, staff, funding, or an environment.
Prior art and the remaining open claim¶
Accessibility audits, disability-led participatory testing, automated scanning, usability studies, and first-valid bug bounties already cover much of the work. The open claim concerns the combined allocation system: with expertise, scope, safety rules, and reward resources held constant, sealed batches, complementarity-based awards, and equal caps should yield more distinct, independently reproducible, repair-usable barriers per reviewer hour than either first-valid allocation or a noncompetitive paid panel, without added harm or exclusion.
Smallest decisive test¶
Run a preregistered two-round crossover on an isolated prototype with synthetic accounts, four essential journeys, hidden seeded barriers, and about 12 compensated testers spanning assistive technologies and input modes. Compare the proposed challenge with a flat-fee panel, then rescore all reports under first-valid allocation. Measure distinct reproduced barriers, seed recall, coverage, repair usefulness, reviewer time, duplicates, fragmentation, burden, urgent-report delay, appeals, and safety events. Advance only if the portfolio condition beats both comparators by the prespecified material margin—suggested as 20% per reviewer hour—without worse safety, delay, or loss of manual work.
Deployment and cost¶
Begin with a nonmonetary mock using fictitious points and equal base compensation. Estimated 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup and operational launch, and \(250,000–\)1 million annually. No direct pricing or internal labor study confirms these bands; they are not vendor quotes.
Risks and uncertainties¶
- Submission caps may suppress slow, uncertain, manual, or assistive-technology-intensive findings.
- Batch sealing may delay urgent accessibility or security disclosure.
- Portfolio scoring may encode the sponsor's incomplete view of important journeys and interaction modes.
- Identity or collusion screens may wrongly combine independent testers or treat suspicious patterns as guilt.
- Synthetic journeys may omit contextual, longitudinal, or socially mediated barriers found in real use.
Expert review¶
Useful reviewer backgrounds: Disabled accessibility tester, Accessibility evaluation methodologist, Bug-bounty program designer, Security and privacy reviewer, Experimental economist or contest-design researcher.
- Can independent scorers apply the complementarity rubric consistently across assistive-technology configurations?
- Do equal caps disproportionately exclude manual or assistive-technology-intensive investigation?
- What urgent-disclosure path preserves safety without compromising sealed allocation?
- Does the competitive condition outperform an equally funded noncompetitive panel after reviewer time is counted?
- Which identity and collusion signals can support investigation without becoming automatic penalties?
Evidence and provenance¶
Selected sources: S1: Guidance on Web Accessibility and the ADA · S2: Selecting Web Accessibility Evaluation Tools · S3: Tips for Usability Testing with People with Disabilities · S4: WCAG Evaluation Methodology (WCAG-EM) 2.0 · S5: Detailed Platform Standards · S6: Fable: Accessibility Research, Powered by People with Disabilities · S7: Crowdsourced Security Vulnerability Discovery: Modeling and Organizing Bug-Bounty Programs · S8: Vulnerability Disclosure Policy
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading aid placed this candidate between ranks 6 and 9 across profiles, in band A. It does not change the partner-candidate endpoint or measure economic value; its pilot-speed input is a cost-affordability proxy.
7. Consistent Rules for Climate Signpost Alerts¶
Canonical title: Behavioral Contract for Adaptive-Policy Signposts and Alerts
In one sentence: A representation-neutral behavioral contract would require spreadsheets, dashboards, and monitoring services to interpret the same climate signpost history in the same way.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-24 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Futurism Foresight |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 71.0/100 · rank range 7–12 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-23, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Coastal adaptation signposts may be copied among policy documents, spreadsheets, dashboards, and provider pipelines. Each implementation can handle dates, revisions, missing readings, sustained thresholds, units, acknowledgments, and retirement differently. Two systems can therefore produce different review alerts from the same observations even when their indicator names and displayed thresholds match. A software or data-source change may silently advance, delay, repeat, or suppress a policy review while officials believe the governing commitment is unchanged.
What is proposed¶
Create an Adaptive Signpost Register with defined operations for registering a versioned signpost, ingesting observations, marking missing periods, revising or retracting data, evaluating status, issuing and acknowledging alerts, suspending evaluation, superseding a signpost, and replaying history. The contract specifies units, time rules, persistence, missingness, revision behavior, typed errors, side effects, and lifecycle states. It requires provenance, an immutable event history, one active version, deterministic replay, and a firm separation between an advisory alert and authority to act. Storage layouts, formulas, provider payloads, caches, and visual presentation stay hidden. A replacement is accepted only if a shared black-box test suite produces the same governed states, alerts, errors, and audit records.
The cross-domain transfer¶
The representation-independent-interface archetype becomes a behavioral boundary around climate signposts. Different technical implementations may store and display data differently, but must expose the same operations, states, errors, replay behavior, and alerts. The mapping is strong at the software layer, though it cannot settle whether an indicator or threshold is scientifically or politically appropriate.
Why it advanced¶
This proposal cleared the EMPIRICAL_PARTNER_CANDIDATE lane because two offline implementations and a synthetic event corpus can test behavioral equivalence safely. It did not enter the strict-success lane. There is no field measurement of signpost divergence, no integrated prototype comparison, and no coastal authority has committed specifications, staff, historical data, or adoption authority.
Prior art and the remaining open claim¶
Adaptive pathways already use signposts and triggers; observation schemas, event sourcing, conformance tests, and migration reviews also exist. The remaining claim is narrower: independently built implementations following one behavioral contract should agree on governed outputs and expose more seeded semantic mistakes than schema validation plus manual spot checks. The contract fails if important divergences survive, approved semantics require provider-specific internals, or the existing comparator finds the same defects with materially less effort.
Smallest decisive test¶
With a named authority, freeze two approved synthetic signpost specifications and 30 histories covering missingness, late and duplicate data, corrections, retractions, persistence, acknowledgments, suspensions, supersession, and replay. Seed at least 12 semantic mutations, including missing-as-zero, arrival-time ordering, one-sample breaches, destructive revisions, duplicate alerts, and hidden rounding. Separate people build an event-ledger model and spreadsheet adapter. Compare contract tests with schema validation plus current manual checks. Reject the proposal if any decision-relevant mutant is missed, conforming systems disagree, approved rules require implementation leakage, or the comparator achieves equivalent detection at materially lower staff effort.
Deployment and cost¶
The first step keeps synthetic observations and copied specifications offline; it must not issue live alerts or change policy thresholds. Estimated 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(50,000–\)250,000 annually. These are not vendor quotes.
Risks and uncertainties¶
- Formal consistency may create false confidence in a weak indicator or poorly chosen threshold.
- The contract may hide contested judgments about time, missingness, or revisions inside technical rules.
- An incomplete test generator may omit consequential event sequences.
- Overly strict tests may freeze harmless choices such as display rounding or notification format.
- A conforming replacement may still have unacceptable latency or resource use outside the initial test.
Expert review¶
Useful reviewer backgrounds: Climate adaptation-pathways specialist, Municipal resilience policy owner, Data-contract or API architect, Geospatial observation-data steward, Software conformance-testing specialist.
- Can each approved signpost be expressed without provider-specific storage or payload fields?
- Which states and alerts are advisory, and which authority separately initiates a policy review?
- Does replay preserve the interpretation of histories created under earlier specification versions?
- Which semantic mutations would materially alter a real review decision?
- How much staff effort does the contract suite require compared with current schema and migration checks?
Evidence and provenance¶
Selected sources: S1: Dynamic adaptive policy pathways: A method for crafting robust decisions for a deeply uncertain world · S2: Designing a monitoring system to detect signals to adapt to uncertain climate change · S3: Taking an adaptive approach: Thames Estuary 2100 · S4: Coastal hazards and climate change guidance · S5: A Summary of Adaptation Pathways Approaches · S6: OGC SensorThings API Part 1: Sensing · S7: Event sourcing pattern · S8: National employment and wage data by occupation, May 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized aid ranked the candidate from 7 to 12 across reading profiles, in band A. This is neither an experimental endpoint nor an economic-value estimate; the pilot-speed input reflects cost-band affordability, not independently measured elapsed time.
8. Remembering Why Financial Controls Exist¶
Canonical title: Why This Control Exists: Control-Memory and Repair Cycle
In one sentence: A carefully bounded reenactment would help control owners remember a control's verified origin, dependencies, and reciprocal duties, then route accepted changes through ordinary governance.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-26 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Accounting Auditing |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 71.0/100 · rank range 3–19 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-11, EXP06-PARTNER-27, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A financial-reporting control may survive staff turnover as a checklist while its rationale and dependencies disappear. Preparers and reviewers can still perform the documented step without knowing which past failure it addresses, whose information it protects, what upstream inputs it assumes, or what conduct must continue. Generic certifications and refresher training may reinforce procedure without restoring that shared understanding. The result can be mechanical compliance, repeated handoff failures, obsolete controls, and false reassurance that remediation is complete.
What is proposed¶
Once each quarter, a cross-functional group would examine one cleared control family through a synthetic or redacted transaction path. A steward first verifies the history, separates fact from interpretation, removes protected information, and identifies affected stakeholders. Participants use functional roles and accessible cards to trace where information was lost, duplicated, misclassified, or repaired. Upstream and accounting roles exchange reciprocal commitments about reliable inputs, validation, feedback, and escalation. Recognition is optional and consent-based; participants may revise or decline proposed commitments without escaping formal job duties. Accepted changes move into authorized narratives, responsibility maps, remediation plans, budgets, or escalation records with owners and dates. Debrief and independent review may revise, pause, repair, or retire the practice.
The cross-domain transfer¶
The ritualized-meaning archetype becomes a marked, recurring examination of why a control exists and what people owe one another to keep it workable. Reenactment, symbolic framing, witnessing, reciprocal exchange, renewal, and retirement are present. The mapping is structurally clear, but evidence that these ritual features outperform simpler instruction in this domain is absent.
Why it advanced¶
This candidate cleared the EMPIRICAL_PARTNER_CANDIDATE lane because a synthetic, content-matched trial can measure learning and harm without changing live controls. It did not pass the strict-success lane. No accounting-specific evidence shows added benefit over a walkthrough, and no controller, internal-control leader, audit committee, or funder has committed to a pilot.
Prior art and the remaining open claim¶
Root-cause reports, training, simulations, certifications, governance software, audit testing, process mining, and ordinary knowledge transfer are established comparators. Ritual features have adjacent workplace and educational evidence, but not for durable financial-control operation. The open claim is that consent-governed symbolic reenactment, reciprocal obligation exchange, and revisable renewal improve seven-day recall, responsibility mapping, and workflow-ready follow-through beyond a content-equivalent walkthrough or ordinary training, without added coercion, blame, religious conflict, or confusion about control evidence.
Smallest decisive test¶
Recruit 36–60 volunteers, stratified by tenure, into three equal-content and equal-time arms: the full cycle, a facilitated walkthrough without ritual features, and ordinary case-based training. Use only synthetic materials. Blind-score immediate and seven-day recall of facts, causal sequence, assertions, dependencies, duties, authority limits, escalation, and the boundary between learning artifacts and control evidence. Reject the proposal if the full cycle trails the best comparator by the prespecified 10-point advantage, adds no workflow-readiness benefit, produces blame or pressure above 10%, exposes opt-outs to supervisors, fails an accommodation, or causes anyone to treat ceremonial closure as proof of effectiveness.
Deployment and cost¶
The authorized first step is one synthetic 75-minute session with 8–12 volunteers; it cannot change or evaluate a live control. Estimated 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence and startup, and \(50,000–\)250,000 for operational launch and annual recurrence. No time study, wage benchmark, or quote validates these bands.
Risks and uncertainties¶
- A selective origin story may legitimize leadership or suppress disputed facts.
- Functional reenactment may still expose or symbolically blame identifiable employees.
- Visible passing, silence, or dissent may become legible to supervisors despite formal consent rules.
- Participants may mistake recognition or ceremonial closure for evidence that a control works or remediation is finished.
- The practice may become costly training theater while underlying resources, authority, and incentives remain unchanged.
Expert review¶
Useful reviewer backgrounds: Corporate controller or internal-control leader, Internal-audit specialist, Organizational learning researcher, Employment, privacy, and religious-accommodation adviser, Accessible facilitation and workplace-safety specialist.
- Can the selected control history be verified and presented without exposing confidential or identifiable information?
- Does the ritual arm improve seven-day recall beyond an otherwise identical walkthrough?
- Can employees decline symbolic participation without supervisors learning or inferring that choice?
- Do proposed commitments enter authorized remediation workflows with owners, resources, and dates?
- Can participants reliably distinguish the learning exercise from control evidence, approval, or certification?
Evidence and provenance¶
Selected sources: S1: Importance of Audits of Internal Controls · S2: Improving Organizational Performance and Governance: How the COSO Frameworks Can Help · S3: The Importance of a Comprehensive Risk Assessment by Auditors and Management · S4: IIA Response to IAASB Proposed Standard on Audits of Financial Statements of Less Complex Entities · S5: Resolving Internal Control Deficiencies and Restatements · S6: Work Group Rituals Enhance the Meaning of Work · S7: Experiential Learning: A Game Changer for Accountants · S8: Section 12: Religious Discrimination
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized aid placed this candidate between ranks 3 and 19 across profiles, in band A. The broad range reflects different reading priorities. It is not an experimental endpoint or economic-value measure, and affordability only proxies pilot speed.
9. Predictive Checks for Evidence Custody¶
Canonical title: Residual Assurance for Physical-Evidence Custody Transitions
In one sentence: A frozen workflow model would flag meaningful differences between expected and recorded evidence-custody transitions while preserving the complete ledger and requiring human review.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-05 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Predictive Residual Processing × Criminology Forensic |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 71.0/100 · rank range 5–17 across three profiles · band A |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-04, EXP06-PARTNER-14, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Forensic laboratories must preserve every handoff, storage move, checkout, return, and seal change for physical evidence. Routine entries can overwhelm reviewers, allowing a missing scan, unauthorized custodian, location mismatch, or broken seal to remain unnoticed. Static rules catch known violations but may miss problems whose meaning depends on the item’s prior state. Yet treating every unusual entry as suspected tampering can create needless quarantines and unsupported conclusions about evidence or personnel.
What is proposed¶
Keep the append-only custody ledger as the authoritative record, then add a separate assurance layer. For one evidence class and workflow, a frozen, versioned model predicts the next authorized custodian, location, seal state, and timing window from the last fully reconciled state. Each actual event is compared with that prediction and labeled as missing, extra, duplicated, late, reordered, unauthorized, or otherwise inconsistent. A gate weighs source reliability, uncertainty, integrity consequences, and reviewer capacity before routing the discrepancy, with its full context, to a named human reviewer. Missing evidence, seal changes, unauthorized access, recorder failure, legal holds, and reviewer requests bypass filtering. Heartbeats verify that silence is genuine, complete-ledger audits test quiet cases, and drift or reconstruction failures return affected items to manual reconciliation.
The cross-domain transfer¶
The predictive-residual archetype becomes a custody-state monitor: expected transitions provide the prediction, recorded scans provide the observation, and their directional difference becomes the residual. The mapping is structurally strong because the workflow has observable states and ordered transitions, but the model remains subordinate to the immutable ledger, physical reconciliation, and human judgment.
Why it advanced¶
This candidate passed Experiment 6’s strict researched-candidate bar because the problem is consequential, most technical components already exist, responsible laboratory authorities and a bounded retrospective test are identifiable, and strong safeguards are specified. STRICT_SUCCESS does not mean real-world validation, novelty, deployment authorization, or demonstrated economic impact.
Prior art and the remaining open claim¶
Barcode and RFID tracking, forensic laboratory systems, deterministic alerts, complete-ledger review, physical inventory, and process-conformance checking already provide adjacent parts. The remaining claim is narrower: combining a frozen custody predictor, structured directional residuals, verified heartbeats, mandatory full-context bypasses, audits of quiet ledgers, and automatic fallback can improve protected-discrepancy recall, acknowledgement speed, reconstruction fidelity, false escalation, and total review workload against optimized rules and full-ledger review.
Smallest decisive test¶
Run a preregistered, read-only three-arm crossover using synthetic ledgers or approved closed-case records from one laboratory workflow. Qualified blinded reviewers compare complete-ledger review, optimized deterministic rules, and the residual interface on seeded custody errors, outages, version changes, and legitimate exceptions. Measure protected-discrepancy recall, acknowledgement time, false escalation, unsupported tampering interpretations, reconstruction error, audit misses, fallback success, reviewer time, and maintenance effort. Reject the approach after any unexplained protected-class miss, irrecoverable reconstruction, failed silence-versus-outage distinction, or failure to improve the joint recall-latency-workload criterion over both comparators.
Deployment and cost¶
The first evidence study is estimated at \(50,000–\)250,000 in rough 2026 resource-equivalent terms. Initial deployment and operational launch are each estimated at \(250,000–\)1 million, with \(50,000–\)250,000 annually. These are planning bands, not quotes; workflow integration, validation, security, training, discovery rules, and accreditation review remain substantial.
Risks and uncertainties¶
- A repeatedly used but unauthorized workflow could become the model’s expected pattern.
- Failed heartbeats could make missing observations look like legitimate quiet periods.
- Reviewers could interpret an unusual transition as evidence of tampering rather than a discrepancy requiring reconciliation.
- Threshold changes intended to quiet the queue could conceal integrity-relevant events.
- One missing or reordered event could corrupt later state reconstruction until full resynchronization occurs unless fallback works correctly and promptly.
Expert review¶
Useful reviewer backgrounds: Forensic laboratory quality manager, Evidence custodian, Forensic LIMS and data-integration engineer, Process-conformance or state-estimation specialist, Forensic accreditation and legal-discovery specialist.
- Which custody discrepancies must have zero unexplained misses in the retrospective test?
- Can approved records or synthetic sequences represent realistic missing, delayed, duplicated, reordered, and emergency-exception events?
- How accurately can the complete custody state be reconstructed after recorder outages or version changes?
- Do quiet-ledger audits reveal material discrepancies that the predictor suppresses or never represents?
- What jurisdiction-specific retention, discovery, privacy, labor, cybersecurity, and accreditation rules govern residual metadata?
Evidence and provenance¶
Selected sources: S1: Evidence Management Survey Results · S3: Evidence Management Steering Committee Report: Opportunities to Strengthen Evidence Management Processes · S4: A Landscape Study of Laboratory Information Management Systems (LIMS) for Forensic Crime Laboratories · S5: RFID Technology in Forensic Evidence Management: An Assessment of Barriers, Benefits, and Costs · S6: Conformance Checking: Foundations, Milestones and Challenges · S7: State of Michigan Contract 071B4300089: Laboratory Information Management System · S8: Laboratory Information Management Systems in Forensic Science Service Provider Laboratories: Current State and Next Generation
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed this candidate between ranks 5 and 17 across three profiles, in band A. That score only sets reading order, uses an affordability proxy for pilot speed, and is neither an experimental endpoint nor an estimate of economic value.
10. A Stable Contract for Justice Histories¶
Canonical title: Observation-Bounded Justice Event Trajectory Contract
In one sentence: An opaque data contract would make differently structured justice datasets return the same study-defined event histories, corrections, observation limits, and uncertainty states.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-10 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Representation Independent Interface Contract × Criminology Forensic |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 70.0/100 · rank range 8–22 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-09, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Longitudinal justice studies often inherit meaning from a warehouse’s rows, null values, keys, joins, and default ordering. A booking row may be treated as an event, a missing field as proof of absence, or a later correction as if it had always been known. Rebuilding the warehouse or analysis layer can therefore change cohorts, event counts, or timing without changing the intended study definition, while inadequate observation may be silently converted into a negative finding.
What is proposed¶
Define each pseudonymous subject’s trajectory through an opaque contract rather than a table layout. The contract stores source-attributed event assertions, occurrence and recording times, equivalence links, non-erasing corrections, explicit periods of source observation, and versioned classification policies. Authorized operations add assertions, relate duplicates, append corrections, register observation periods, reconstruct what was known at a cutoff, and classify study windows. A window returns one of three answers: a qualifying event was recorded, none was recorded during adequate observation, or observation was insufficient. Invalid operations leave the abstract state unchanged. Warehouse keys, joins, normalization, nesting, caches, and ordering stay hidden. Relational, document, or graph implementations are substitutable only when shared black-box, generated-sequence, metamorphic, and representation-leakage tests produce equivalent histories, classifications, uncertainty states, and permitted errors.
The cross-domain transfer¶
The representation-independent interface archetype becomes a justice-trajectory contract. Its abstract state contains events, corrections, observation coverage, and policy versions; its public operations expose study meanings rather than storage details. The mapping is strong for bounded research datasets, but it cannot erase source-specific legal meaning or establish that underlying records and linkages are accurate.
Why it advanced¶
The candidate passed Experiment 6’s strict researched-candidate bar because documented longitudinal-data errors could alter study classifications, the required software primitives exist, and a synthetic cross-representation test is feasible. STRICT_SUCCESS remains a researched-candidate result, not proof of adoption, factual accuracy, universal applicability, novelty, or real-data performance.
Prior art and the remaining open claim¶
Common justice schemas, harmonized research tables, temporal databases, observation-period models, statistical plans, linkage systems, and immutable snapshots are adjacent prior art. The open claim concerns their narrower combination: one opaque contract with source attribution, non-erasing corrections, two kinds of time, explicit observation intervals, versioned policies, three-valued window answers, and a cross-representation oracle can preserve approved-study classifications better than direct warehouse queries, schema checks, and aggregate comparisons.
Smallest decisive test¶
Preregister 15–25 synthetic trajectories containing duplicates, split and consolidated episodes, late dispositions, corrections, observation gaps, overlapping coverage, timing conflicts, and policy changes. Two independent teams encode them in relational and document stores. Compare direct queries, schema and aggregate checks, a harmonized-table baseline, and the contract against fixed expected outputs and at least 12 seeded semantic mutants. Advance only with zero material classification or uncertainty divergences across valid implementations and at least 90% mutant detection, outperforming baseline checks. Redesign or reject it if inadequate observation becomes absence or passing implementations disagree materially.
Deployment and cost¶
A synthetic first study is estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms. Startup is estimated at \(50,000–\)250,000, operational launch at \(250,000–\)1 million, and annual operation at \(50,000–\)250,000. These are not vendor quotes; real-data use also requires study governance, privacy review, and authorization.
Risks and uncertainties¶
- The abstract event categories could erase legally or analytically important source distinctions.
- An equivalence rule could merge separate events or count one event more than once.
- Weak observation criteria could convert incomplete coverage into an apparent absence.
- Synthetic cases may omit irregular corrections, delayed entry, and overlapping-source behavior found in actual data.
- A passing conformance suite could be mistaken for evidence that source assertions or record linkages are correct and complete across all cases.
Expert review¶
Useful reviewer backgrounds: Criminology research methodologist, Administrative-data steward, Temporal data-modeling specialist, Software conformance-testing engineer, Privacy, ethics, and institutional-review specialist.
- Which source-specific distinctions must remain visible to preserve the approved study’s interpretation?
- What evidence is sufficient to call an observation interval adequate for a negative finding?
- Can independent implementers agree on event-equivalence and correction behavior before seeing test results?
- Which seeded mutants represent material study errors rather than harmless implementation differences?
- Would real pseudonymous records remain identifiable or require additional institutional and legal authorization?
Evidence and provenance¶
Selected sources: s1: Overview · s2: Criminal Justice Administrative Records System (CJARS): Data Documentation, 2023 Q3 · s3: Longitudinal linkage of administrative data: design principles and the total error framework · s4: NIEM 5.0 Justice Domain: j:Arrest · s5: OMOP Common Data Model v5.3: Observation Period · s6: Temporal tables · s7: Coded Private Information or Biospecimens Used in Research, Guidance (2018) · s8: Software Developers, Quality Assurance Analysts, and Testers
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed this candidate between ranks 8 and 22 across profiles, in band A. It is a reading-order aid using a cost-based pilot proxy, not an experimental endpoint, economic-value measure, or change to its STRICT_SUCCESS status.
Band B — high post-hoc review priority (ranks 11–25)¶
11. A Credible End to Financial Close¶
Canonical title: Close-Mode Release and Recovery Observance
In one sentence: An optional post-close observance would mark the end of exceptional work, preserve every residual task, and turn gratitude into authorized, reviewable recovery commitments.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-11 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Accounting Auditing |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 70.0/100 · rank range 8–26 across three profiles · band B |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-26, EXP06-PARTNER-27, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A financial close may be officially complete while staff still face vague cleanup, monitoring, and audit-support duties. Completion emails and celebrations can praise visible overtime without clarifying who owns remaining work or what recovery will occur. Fatigue and ambiguous responsibility may then carry into ordinary operations and the next close. Symbolic gratitude can also conceal unresolved staffing, pay, leave, or process needs, while employees who are remote, junior, quiet, or unwilling to celebrate remain unseen.
What is proposed¶
After an authorized reporting milestone, hold a governed and optional release observance that has no power to close books, waive controls, or cancel work. A steward states that boundary, then switches off a nonoperational close-mode marker. Participants may use a bounded reflection period, remain off-camera, join asynchronously, or decline. With prior consent, witnesses recognize specific work such as error prevention, documentation, coordination, boundary setting, and asking for help, without ranking hours or glorifying exhaustion. A shared board separates completed work from residual obligations, which enter the official tracker with owners, limits, evidence needs, and escalation paths. Only authorized recovery, backfill, compensation-review, meeting-relief, staffing, or process commitments are published. An immediate debrief tests coercion and credibility; a later independent review compares promises with schedules, recovery, and task completion.
The cross-domain transfer¶
The ritualized-commitment archetype becomes a visible transition out of exceptional close mode. Its marker, silence, witnessing, commitment renewal, and later accountability review enact shared meaning and reciprocal obligations. The structural mapping is plausible, but the ceremony cannot itself reduce workload, provide recovery, satisfy wage rules, or correct accounting processes.
Why it advanced¶
The candidate passed Experiment 6’s strict researched-candidate bar because the underlying fatigue and close-management problems are consequential, a low-cost synthetic comparison is feasible, and explicit consent and authority safeguards make the claim testable. STRICT_SUCCESS does not establish workplace benefit, adopter demand, novelty, or permission for live use.
Prior art and the remaining open claim¶
Workplace rituals, project-completion ceremonies, close-management software, task trackers, retrospectives, overtime and leave policies, fatigue programs, and recognition systems already exist. The narrower open claim is that, with these controls held constant, an optional governed release sequence can improve retained understanding of official completion versus residual work and the credibility of authorized recovery commitments without increasing coercion, privacy exposure, status confusion, overwork glorification, or unpaid extra-role effort.
Smallest decisive test¶
Run a preregistered synthetic comparison with 8–12 volunteer accounting and workforce participants. Counterbalance two equivalent close scenarios: ordinary status communication, task tracking, recovery policy, and retrospective; and those same controls plus the observance. Seed residual tasks, unauthorized recovery proposals, an accessibility need, and a status-confusion cue. Measure immediate and 72-hour understanding of official status, ownership, escalation, and authorized commitments, alongside anonymous safety ratings. Revise or reject the observance after any lost task, increased authority error, noncredible opt-out, pressured positivity, unsafe disclosure, or failure to improve retained boundary clarity.
Deployment and cost¶
The first synthetic study is estimated below $10,000 in rough 2026 resource-equivalent terms. Startup and operational launch are each estimated at \(10,000–\)50,000, with \(50,000–\)250,000 annually. These bands exclude potentially dominant costs such as paid recovery, overtime, backfill, staffing, compensation changes, and automation, and are not vendor quotes.
Risks and uncertainties¶
- Staff may mistake switching off the marker for official completion of books, controls, or audit obligations.
- The expectation of celebration or silence may become coercive despite a formal opt-out.
- Recognition may reward visible long hours while overlooking remote, junior, contingent, or upstream contributors.
- Managers may promise recovery, leave, pay, staffing, or backfill without authority or resources.
- Private health, family, grievance, or workload information could be exposed during reflection or recognition and then mishandled by others involved in the session or organization afterward.
Expert review¶
Useful reviewer backgrounds: Corporate controller, Human-resources or labor-relations specialist, Accounting close-process owner, Occupational fatigue and workplace-safety specialist, Accessibility, privacy, and organizational-behavior specialist.
- Can every participant decline the symbolic portion without practical or perceived retaliation?
- Which statements and visual cues reliably distinguish the observance from official close status?
- Who has authority and funding to approve each proposed recovery or workload commitment?
- Does recognition capture preventive and boundary-setting work without rewarding excessive hours?
- What evidence before the next close would show that promised recovery and residual-task ownership actually occurred?
Evidence and provenance¶
Selected sources: S1: The Agentic Close: From Month-End Sprint to Always-Ready Finance · S2: How organizations can streamline the month-end close · S3: 2025 Corporate Finance & Accounting Talent Study · S4: Work group rituals enhance the meaning of work · S5: Financial period close workspace · S6: Fatigue and Work · S7: Wages and the Fair Labor Standards Act · S8: National employment and wage data by occupation, May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed this candidate between ranks 8 and 26 across profiles, in band B. Its relatively inexpensive test improved deployment-heavy ordering, but the score is only a reading aid—not an endpoint, elapsed-time measure, or estimate of value.
12. Blind Testing for Shared-Service Cost Models¶
Canonical title: Holdout Cost-Driver Tournament for Shared-Service Allocations
In one sentence: A controller would compare competing shared-service allocation models on withheld data under fixed rules, without posting the results or using them in live decisions.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-03 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Accounting Auditing |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 11–21 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-01, EXP06-PARTNER-02, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Business units often favor shared-service cost models that reduce their own allocated share, because one enterprise rule shifts costs among all units. Sponsors may choose favorable historical periods, redefine usage, exclude transactions, add opaque complexity, or influence judges. Yet competing proposals can reveal better cost drivers and weak data. The challenge is to distinguish legitimate measurement improvement from burden shifting when no organization-specific evidence yet shows that sponsorship incentives actually distort the selection process.
What is proposed¶
For one reconciled shared-service pool, run an identity-blinded tournament whose sole prize is a one-year designation as the managerial allocation basis. Freeze eligibility, data sources, transformations, holdout periods, scoring weights, conflicts, appeals, and prohibited conduct before outputs are inspected. Each model must disclose beneficiaries, use governed data, reproduce its logic, and reconcile the full pool. Blinded tests score usage linkage, out-of-period stability, auditability, data burden, and sensitivity to discretionary assumptions; sponsor savings are disclosed to validators but not scored. Independent reviewers reperform leading models. Development effort and complexity are capped, suspicious cross-model exclusions prompt investigation, and common data remains available to later challengers. A central holdback supports correction of trial-year defects. The designation expires, and post-contest review may revise or retire the process.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a controlled competition among allocation models. The scarce prize, fixed arena, eligibility rules, foul schedule, resource caps, independent judging, challenger access, and winner review map clearly. The empirical premise is weak, however: no partner data yet show systematic self-favoring sponsorship, selection influence, or manipulation beyond ordinary technical disagreement.
Why it advanced¶
This candidate did not enter the strict-success lane. It qualified only as an EMPIRICAL_PARTNER_CANDIDATE because a bounded, nonposting partner study could test the premise safely, while essential field evidence remains missing. Its status is an invitation to investigate actual sponsor behavior and model performance, not evidence that the proposed problem or remedy exists in practice.
Prior art and the remaining open claim¶
Managerial-costing standards, standard allocation drivers, commercial allocation software, model-risk governance, holdout testing, independent validation, and controller judgment are adjacent prior art. The open comparison is narrower and conditional: when direct tracing is infeasible and sponsors have distributive exposure, a blinded, preregistered tournament may select a more stable, reproducible, usage-linked model than controller selection or a non-blinded panel while weakening the relationship between sponsor savings and rank.
Smallest decisive test¶
With a controller and data owners, preregister a closed-year shadow study comparing the incumbent basis, a conventional non-blinded panel choice, and the blinded tournament using identical candidate models. Freeze the holdout quarter, weights, and stress tests first. Measure reconciliation, reproducibility, stability, usage linkage, assumption sensitivity, burden, transfers to nonparticipants, direct-tracing feasibility, and correlation between sponsor savings and rank. Do not post allocations or use them for budgets, pay, tax, transfer pricing, or reporting. Stop if data access, common-pool reconciliation, validator independence, or target validity fails; require material comparative improvement before considering another study.
Deployment and cost¶
The first partner study is estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms. Startup is estimated at \(50,000–\)250,000, operational launch at \(250,000–\)1 million, and annual operation at \(50,000–\)250,000. Actual costs, authority, data availability, and recurring validation burden have not been observed and these bands are not vendor quotes.
Risks and uncertainties¶
- Sponsor identity may remain obvious from a model’s structure, defeating blinding.
- Withheld historical periods may not represent future service consumption or behavior.
- A composite score may conceal disputed judgments about causal linkage, stability, and administrative burden.
- Complexity caps may reject a justified model for genuinely heterogeneous services.
- Historical driver data may already reflect earlier allocation choices, missing usage, or incumbent control and therefore bias comparisons in ways difficult to detect without detailed organization-specific investigation.
Expert review¶
Useful reviewer backgrounds: Corporate controller or managerial-accounting policy owner, Shared-service cost-modeling specialist, Independent model validator or internal auditor, Operational data owner and data-governance specialist, Business-unit finance representative without judging authority.
- Do sponsored models favor their sponsors after legitimate service differences are controlled?
- Can sponsor identity be hidden well enough for blinded judging to be meaningful?
- What operational measure can serve as a defensible usage-linkage target?
- Is direct metering or decomposition feasible at an acceptable authorized burden?
- Are model rankings stable across preregistered holdouts, stress scenarios, and reasonable scoring weights?
Evidence and provenance¶
Selected sources: S1: Why do so many organisations struggle with cost allocation? · S2: Who should pay for support functions? · S3: Distorted Cost Allocation: An Encouragement or Discouragement? · S4: The Conceptual Framework for Managerial Costing · S5: TBM Modeling Allocation Methods · S6: Supervisory Guidance on Model Risk Management · S7: The TBM Taxonomy · S8: HHS Policy for Information Technology Portfolio Management (PfM)
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed this candidate between ranks 11 and 21 across profiles, in band B. This reading-order score uses an affordability proxy and is not an experimental endpoint, economic-value estimate, or upgrade from its EMPIRICAL_PARTNER_CANDIDATE status.
13. Separate Release Effects from River Disturbances¶
Canonical title: Command-Conditioned Residual Watch for Managed River Releases
In one sentence: Test whether a model of verified reservoir operations can filter predictable downstream changes from operators’ attention without hiding protected signals, coincident environmental events, or model failures.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-15 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Environmental Climate |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 13–19 across three profiles · band B |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
Reservoir releases can predictably alter downstream water level, temperature, dissolved oxygen, conductivity, and turbidity. These large, expected changes can flood ordinary dashboards and fixed alarms, leaving operators to decide whether each warning reflects the release or a separate problem such as contamination, bank failure, or unexpected sediment movement. Widening thresholds during releases reduces nuisance alarms but can also create a blind spot precisely when an external event may coincide with operations.
What is proposed¶
Before downstream readings arrive, copy each authorized gate or turbine command and compare it with measured actuator behavior. A frozen, versioned model would predict when and how the release should affect each monitoring station. The system would preserve every raw reading, calculate signed observed-minus-predicted residuals, and direct reliable or consequential unexplained changes to an operator queue. A residual would always retain enough command, model, station, timing, and sensor information to reconstruct the full observation. Critical ecological limits, dam-safety variables, compliance records, sensor faults, and public-warning conditions would bypass filtering. Unverified actuation, excessive uncertainty, model drift, missing heartbeats, incompatible versions, or failed reconstruction would restore full-signal review. Residuals could prompt investigation or separately reviewed recalibration, but could not alter gates or emergency actions.
The cross-domain transfer¶
The predictive-residual archetype is mapped strongly here: a copy of the reservoir’s own command predicts its delayed sensory consequences, and the difference between prediction and observation highlights what the command does not explain. The mapping adds actuator verification, uncertainty weighting, reconstruction, raw-data audits, and a fallback because an environmental predictor must never become grounds for erasing consequential observations.
Why it advanced¶
This entered the empirical-partner lane because release modeling, gate-operation modeling, continuous monitoring, quality control, and alarm-management practices make a bounded replay plausible. It did not clear the strict-success lane: no site has supplied synchronized records, and there is no measured evidence yet for event recall, protected-signal routing, reconstruction, fallback reliability, workload reduction, or lower total burden.
Prior art and the remaining open claim¶
Adjacent systems already model reservoir operations and downstream water quality, detect sensor anomalies, and manage alarms. The narrower open comparison is whether conditioning a frozen model on both the issued command and measured actuation can outperform an unchanged raw-threshold dashboard: at least 30% less routine review, 100% routing of protected signals, at least 95% detection and acknowledgement of scripted coincident departures, reconstruction within declared tolerances, and correct fallback for every injected fault. No exact implementation or comparative result was established.
Smallest decisive test¶
With one reservoir partner, preregister a historical release replay with a contiguous untouched holdout. Freeze the model, stations, uncertainty assumptions, protected classes, tolerances, and fallback rules. Blindly add conductivity, turbidity, and dissolved-oxygen departures; inject actuator mismatch, dropout, checksum conflict, travel-time shift, and persistent bias; and compare with the unchanged dashboard. Reject the claim if review falls by under 30%, any protected signal is suppressed, scripted-event detection is below 95%, any injected fault misses fallback, reconstruction exceeds tolerance, or total labor is not lower. Only a complete pass permits a read-only shadow release.
Deployment and cost¶
The first step is retrospective and then shadow-only; existing dashboards, staffing, operating procedures, and control authority remain unchanged. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for initial startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. These are assessment bands, not vendor quotes or site estimates.
Risks and uncertainties¶
- A contaminant or sediment pulse arriving during a release could be predicted away as an operational effect.
- A logged command may not match physical gate or turbine movement, corrupting the downstream prediction.
- Wrong travel time could turn one event into misleading positive and negative residuals at different times.
- Season, tributary inflow, channel change, stratification, or shared upstream-data errors could invalidate the model.
- Expected release effects may still breach ecological or compliance limits and therefore cannot disappear from review context or records, even if they are accurately predicted.
Expert review¶
Useful reviewer backgrounds: Reservoir operations engineer, River water-quality scientist, Hydrologic or hydraulic modeler, Environmental incident-response lead, Continuous-sensor quality specialist subject to relevant independence safeguards.
- Are command, measured-actuator, upstream-condition, tributary, weather, sensor, alarm, and acknowledgement records synchronized well enough for the preregistered replay?
- Which stage, dissolved-oxygen, conductivity, and turbidity conditions must always bypass residual filtering?
- What variable-specific reconstruction errors and travel-time errors are acceptable before full-signal fallback?
- Can blinded coincident events be designed so they represent consequential external disturbances without being trivially detectable?
- Does total operator, modeling, audit, and investigation labor remain below the raw-dashboard baseline at equal completeness?
Evidence and provenance¶
Selected sources: S1: Water Quality · S2: The Effects of Hydropower Releases from Lake Texoma on Downstream Water Quality · S3: HEC-ResSim Version 4.0, New Water Quality Feature · S4: HOW TO Add Real-Time Gate Operation Capability to a ResSim Model · S5: Guidelines and Standard Procedures for Continuous Water-Quality Monitors: Station Operation, Record Computation, and Data Reporting · S6: ContDataQC: An R Package and Shiny App for Quality Control of Continuous Water Quality Sensor Data · S7: Alarm Management for Hydropower Plants · S8: Guidelines for Collecting Data to Support Riverine Hydrodynamic and Water Quality Simulation Models
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized score placed this candidate between ranks 13 and 19, in band B. That score is only a post-hoc reading order using a cost-band affordability proxy; it is not an experimental endpoint, elapsed-time estimate, deployment finding, or measure of economic value.
14. A Storage-Independent Aircraft Maintenance Ledger¶
Canonical title: Opaque Maintenance-Obligation Ledger for Aircraft Configuration and Usage Evidence
In one sentence: Define and test maintenance-obligation behavior independently of database layout so backend changes cannot silently alter configuration, usage, credit, conflict, or due-status answers.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-18 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Aviation Aeronautics |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 7–28 across three profiles · band B |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
Aircraft maintenance clients may read shared tables and calculated columns directly, each making its own assumptions about record order, nulls, unit conversions, corrections, installations, counter resets, and maintenance credit. A database migration, cache redesign, or replay into a new backend can therefore change whether an obligation appears satisfied, remaining, overdue, or indeterminate even when the underlying evidence is logically unchanged. Different clients may also produce conflicting answers with no common behavioral standard for deciding which is correct.
What is proposed¶
Create an opaque aircraft maintenance-obligation ledger whose public operations establish a baseline, submit typed evidence, supersede rather than erase an erroneous event, produce an as-of snapshot, query an obligation, compare revisions, and explain the evidence behind a status. Its abstract state would include configuration, serialized components, usage, requirements, credited accomplishments, conflicts, supersession links, and immutable history. Preconditions would define identifiers, units, timing, authority, applicability, and correction links. Invariants would prohibit deleted history, double installation, unsupported credit, and definite answers from contradictory evidence. Relational, event-log, graph, or cached implementations would remain hidden and would have to pass the same generated operation-sequence tests. The ledger would report evidence and conflict states; authorized personnel would still decide maintenance credit, deferrals, aircraft status, and operational release.
The cross-domain transfer¶
The representation-independent contract archetype maps directly to the ledger. It replaces database rows as the effective interface with an abstract state, explicit transitions, stable errors, evidence explanations, and a common black-box oracle. Opacity alone is insufficient: each backend must map to the same ledger meaning, and client code must stop reconstructing maintenance status from hidden storage details.
Why it advanced¶
This qualified for a bounded partner study because regulated record integrity, digital interoperability, mature maintenance systems, and established migration audits make isolated replay feasible. It did not enter strict success: no dependency audit, controlled representation perturbation, independent model, mutant suite, leakage audit, or comparative result exists, and no operator has approved the complete semantics for even one obligation type.
Prior art and the remaining open claim¶
Maintenance standards already cover information exchange, electronic logbooks, allowable configuration, transfer records, and due-status business rules; commercial systems already manage configurations, utilization, compliance, and component status. Append-only logs, typed APIs, and migration audits also address parts of the problem. The narrower open claim is that an opaque evidence-state contract plus an independent reference model and shared sequence oracle will catch semantic defects and preserve identical observable behavior across two differently represented backends more reliably than schema checks, record counts, and sampled reports.
Smallest decisive test¶
With an operator or MRO and an authorized records reviewer, preregister one obligation type and at least 100 curated histories plus generated sequences. Compare an incumbent wrapper and independent immutable model with schema, count, and sampled-report checks. Seed defects for duplicated utilization, erased superseded evidence, insertion-order dependence, double installation, and silent conflict resolution. Require identical revisions, statuses, errors, and provenance explanations between nominal implementations, and rejection of every seeded defect. Reject the claim if the baseline catches the same defects, any mutant passes, valid implementations need schema exposure to agree, contradictory evidence must become definite, or representation-only perturbations reveal no actual client dependency.
Deployment and cost¶
Begin with de-identified or synthetic histories in a read-only isolated environment; do not write official records or connect results to dispatch or release workflows. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. They are not quotations or organization-specific estimates.
Risks and uncertainties¶
- The abstract state may omit maintenance-program, jurisdiction, configuration, or evidence-authority context required to interpret an obligation.
- The independent model may repeat the incumbent’s mistaken rule or be too simple to adjudicate complex histories.
- Generated cases may miss irregular but valid corrections, counter changes, or applicability sequences.
- Events labeled independent may actually have order-dependent maintenance meaning.
- Opaque storage could hinder investigation if evidence explanations and sanctioned audit access are inadequate, while deterministic outputs could be mistaken for release authority.
Expert review¶
Useful reviewer backgrounds: Authorized aircraft maintenance-records specialist, Maintenance-planning or continuing-airworthiness engineer, MRO systems architect, Aviation software assurance and test specialist, Regulatory compliance or airworthiness counsel.
- Which single obligation type has semantics complete enough to specify units, applicability, cutoffs, corrections, conflicts, and authority?
- Which clients currently read tables, calculated columns, row order, nulls, or local credit rules?
- What contract-level evidence explanation is necessary for an authorized reviewer to resolve every seeded divergence?
- Which event pairs truly commute, and which only appear independent until configuration applicability is considered?
- What regulatory acceptance, retention, privacy, cybersecurity, and controlled-data conditions would govern any move beyond shadow replay?
Evidence and provenance¶
Selected sources: S1: 14 CFR §121.380 Maintenance Recording Requirements and §121.380a Transfer of Maintenance Records · S2: AC 120-78B: Electronic Signatures, Electronic Recordkeeping, and Electronic Manuals · S3: Easy Access Rules for Continuing Airworthiness: ML.A.305 Aircraft Continuing-Airworthiness Record System · S4: Adopting Aircraft Electronic Records, First Edition · S5: ATA e-Business Program Standards · S6: Regulatory Compliance Module · S7: Case Study: Endeavor Air—Managing Legacy MRO Systems · S8: Copa Airlines Selects GE Aviation for Digital Records Management
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed this candidate from rank 7 to 28, in band B. This wide, post-hoc reading range is not an experimental endpoint or evidence of value; its pilot-speed input was only a cost-band affordability proxy, with elapsed time unscored.
15. Portable Rules for Drought-Stage Advice¶
Canonical title: Representation-Independent Contract for Drought-Stage Recommendations
In one sentence: Specify drought-stage recommendation behavior independently of any spreadsheet or dashboard, then test whether a second engine can reproduce it without relying on hidden formulas, fields, colors, or evaluation order.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-20 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Environmental Climate |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 14–20 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-19, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A drought plan’s stage calculation may live in one spreadsheet, dashboard, indicator service, or rules engine. Staff and connected procedures can start depending on cell locations, colors, proprietary scores, rounding, formula order, or undocumented treatment of missing data. Replacing or refactoring that implementation can then change an escalation, recovery, or indeterminate result even though the approved policy and hydrologic evidence were supposed to remain the same. The reason for the change may be impossible to reconstruct cleanly.
What is proposed¶
Define an opaque drought-stage recommender whose public state includes the current recommended stage, assessment time, evidence sufficiency, pending transition, policy and parameter versions, and stable reason categories. Operations would initialize a version, assess or advance with typed evidence, invalidate withdrawn inputs, query the recommendation, and export an immutable receipt. Inputs must meet declared rules for units, time windows, location, eligibility, freshness, and coverage. The same abstract history must produce the same result; missing required evidence cannot silently become a normal stage; hysteresis and persistence must follow the approved policy; and rejected calls cannot change state. Formulas, database structures, vendor fields, caches, colors, and evaluation order remain hidden. A shared black-box suite would test boundaries, histories, corrections, missingness, receipts, errors, and leakage. Official declarations, restrictions, allocations, and public messages remain with the designated authority.
The cross-domain transfer¶
The representation-independent contract archetype is a strong structural match. The proposal separates the policy-facing behavior—stages, sufficiency, transitions, reasons, errors, and receipts—from any spreadsheet or rules engine. A replacement is acceptable only under the same approved policy version and after common conformance tests; changing a threshold or hysteresis rule is a policy change, not an implementation substitution.
Why it advanced¶
This reached the empirical-partner lane because real programs use multi-indicator stages, formal trigger governance exists, and environmental-data standards demonstrate typed exchange and black-box conformance. It did not reach strict success: no real workflow dependency has been documented, no independent engine has been compared, and no authority has approved the proposed semantics for missingness, corrections, invalidation, or hysteresis.
Prior art and the remaining open claim¶
Drought programs already define indicators, stages, triggers, governance, and history-sensitive assessment; WaterML standardizes observation exchange, and environmental-data APIs have executable conformance tests. These are adjacent rather than empty territory. The remaining claim is narrower: for one frozen policy version, a stateful contract covering insufficiency, time order, hysteresis, corrections, invalidation, stable reasons and errors, and receipts will let an independently built engine substitute without any contracted-output divergence or client dependence on internal representation.
Smallest decisive test¶
In a six-to-eight-week non-live partner study, freeze one approved policy and completed assessment packet. Preregister evidence quantities, units, freshness, coverage, stages, transitions, missingness, corrections, reasons, receipts, boundary outcomes, and tolerances. Compare the incumbent through an adapter, an independent decision-table engine, and current manual or schema-only validation across threshold, unit, field-order, stale-data, repeated-snapshot, escalation, recovery, correction, and invalidation cases. Reject the problem premise if no hidden dependency appears and substitution requires no client change or output correction. Reject the intervention if any mutant passes, two passing engines disagree, or expressing the policy requires exposing the incumbent representation.
Deployment and cost¶
The first study is sandboxed, non-authoritative, and disconnected from declarations, restrictions, allocations, controls, and public communications. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(50,000–\)250,000 annually. These bands are not vendor quotes and lack external comparable-cost evidence.
Risks and uncertainties¶
- The contract could accidentally turn an undocumented spreadsheet quirk into approved drought policy.
- A finite suite may miss a threshold boundary or assessment history that changes a consequential recommendation.
- Equivalent software behavior would not show that the indicators, thresholds, or policy are scientifically, legally, or equitably appropriate.
- Stable reason categories may omit nuance required by hydrologists or decision-makers.
- Generated cases may miss correlated missing data or unusual corrections, and shadow outputs may still encourage unauthorized automation.
Expert review¶
Useful reviewer backgrounds: Drought-program policy owner, Operational hydrologist, Water-utility drought planner, Environmental software and conformance-test engineer, Legal, equity, and public-communications reviewer.
- Which document and authority determine the approved policy when spreadsheet behavior and written rules disagree?
- What exact evidence, freshness, coverage, hysteresis, and missing-data rules must the contract express?
- Do any current clients depend on cells, colors, vendor scores, rounding, reason text, or evaluation order?
- Which deliberately mutated engines would represent the most decision-relevant implementation errors?
- What outputs and labels are necessary to prevent a shadow recommendation from being treated as an official declaration?
Evidence and provenance¶
Selected sources: S1: Monitoring Drought · S2: Drought Response and Recovery: A Basic Guide for Water Utilities · S3: Building a Drought Planning Platform · S4: Drought Indicators & Assessment · S5: Water company drought plan guideline, 2025 · S6: What is the USDM? · S7: WaterML · S8: OGC API - Environmental Data Retrieval 1.0 Conformance Test Suite
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed this candidate between ranks 14 and 20, in band B. This is a post-hoc reading order, not an experimental endpoint, economic-value score, or pilot-duration estimate; affordability was proxied from the rough cost band.
16. Portable Rules for Delphi Rounds¶
Canonical title: Representation-Independent Contract for Delphi Elicitation Rounds
In one sentence: Test whether survey, spreadsheet, and analysis implementations can preserve the same Delphi-round lifecycle and anonymity boundary under one storage-independent behavioral contract.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-23 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Futurism Foresight |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 13–24 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-24, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A multiround public-health Delphi exercise may spread enrollment, anonymous responses, closure, aggregation, feedback, revision, and withdrawal across a survey platform, spreadsheets, email, and analysis scripts. Each tool can treat eligibility, missing answers, late submissions, replacements, withdrawal, and participant identity differently. When data moves or a tool changes, matching questionnaires and columns does not show that the same responses entered the aggregate, the same feedback population was used, or the same anonymity promises survived.
What is proposed¶
Define a Blind Iterative Elicitation State with operations to enroll an eligible participant through a separate identity service, issue an unlinkable study handle, open a round, accept or replace a response before closure, record withdrawal, close the round, calculate a declared aggregate, publish controlled feedback, begin the next round, and export an audit view. The contract would allow one active response per eligible handle, question, and round; forbid post-closure mutation; derive feedback only from the declared closed population; apply a sponsor-approved withdrawal policy; and block ordinary analysis from resolving identities. Preconditions, errors, side effects, versioning, and sanctioned identity recovery would be explicit while vendor fields, spreadsheet rows, email routing, and storage remain hidden. Every implementation would have to pass the same sequence and leakage tests before holding authoritative round state.
The cross-domain transfer¶
The representation-independent contract archetype maps well to the procedural state of Delphi rounds. It defines observable enrollment, response, withdrawal, closure, aggregate, feedback, and identity-access behavior without choosing a survey or storage format. The transfer is limited, however: technical equivalence cannot preserve participant experience, facilitator judgment, or the social context that may influence elicitation quality.
Why it advanced¶
This entered the empirical-partner lane because Delphi is established in government and health research, mature tools implement most lifecycle functions, and a synthetic two-implementation test is readily bounded. It was not a strict success: no migration incident evidence, committed sponsor, executable contract, adapter comparison, mutant result, or leakage result exists, and the governing legal and withdrawal rules remain unspecified.
Prior art and the remaining open claim¶
Existing Delphi products already support anonymous participation, multiple rounds, response revision, feedback, stopping rules, and study identifiers. Reporting guidance makes methodological choices more explicit, while vendor-neutral study-data standards support portable records. The narrower unresolved claim is that two independent implementations of one sponsor-approved synthetic protocol will agree on every contracted lifecycle and identity-access observation, reject every single-rule semantic mutant, and reveal no predeclared identity-linkage channel through allowed outputs. This is compositional proximity, not a novelty finding.
Smallest decisive test¶
Within six weeks, have a method owner preregister a synthetic two-round protocol covering eligibility changes, duplicates, replacement, missing and late answers, pre- and post-closure withdrawal, failed closure, two aggregation rules, feedback, and identity-resolution attempts. Build independent in-memory and spreadsheet-backed implementations, run identical scripted and generated histories, and add at least eight single-rule mutants. Reject the claim if an approved workflow cannot be expressed, implementations pass while disagreeing on any contracted observation, any mutant survives, or handles, timestamps, ordering, filenames, errors, or exports expose a predeclared linkage channel. Compare with reconstruction through the ordinary spreadsheet/export procedure; use no real experts or active records.
Deployment and cost¶
Start entirely offline with synthetic participants, non-live questions, isolated identity mappings, and no authoritative study migration. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(50,000–\)250,000 for operational launch, and \(10,000–\)50,000 annually. These labor-equivalent bands are not vendor quotes and exclude unverified licensing, security, procurement, and integration costs.
Risks and uncertainties¶
- The contract may encode one contested Delphi method as if it were a universal technical invariant.
- Generated histories may omit the unusual event orderings most likely to expose disagreement.
- Study handles could remain linkable through timing, ordering, filenames, metadata, or error messages.
- A simple reference implementation could acquire unwarranted authority over sponsor-approved methodological choices.
- Passing lifecycle tests would not show that the expert panel is representative, its judgments are accurate, or the participant experience is equivalent across tools.
Expert review¶
Useful reviewer backgrounds: Delphi methodologist, Public-health workforce foresight program owner, Research data-protection or privacy officer, Survey-platform and research-software engineer, Study sponsor or research-governance reviewer.
- Which withdrawal, replacement, eligibility, aggregation, feedback, and identity-recovery rules has the sponsor actually approved?
- Which ordinary outputs could link pseudonymous handles to people through timing, ordering, filenames, metadata, or errors?
- Do the generated sequences cover every meaningful pre- and post-closure failure path?
- Can two implementations represent the approved workflow without importing vendor-specific identifiers or spreadsheet behavior?
- Does the ordinary spreadsheet/export comparator detect the same seeded mutants, and at what review effort?
Evidence and provenance¶
Selected sources: s1: The Futures Toolkit HTML · s2: Guidance on Conducting and REporting DElphi Studies (CREDES) in palliative care: Recommendations based on a methodological systematic review · s3: ACCORD (ACcurate COnsensus Reporting Document): A reporting guideline for consensus methods in biomedicine developed via a modified Delphi · s4: eDelphi 2026 · s5: DelphiManager Frequently Asked Questions · s6: ODM v2.0 · s7: Regulation (EU) 2016/679 (General Data Protection Regulation) · s8: Software Developers, Quality Assurance Analysts, and Testers
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed this candidate between ranks 13 and 24, in band B. This post-hoc ordering is not a preregistered endpoint, validation, or economic-value measure; its pilot-speed input is only an affordability proxy derived from the rough cost band.
17. Stable Sampling Across Changing Data Systems¶
Canonical title: Opaque Probability-Sampling Frame Contract for Multi-Wave Social Research
In one sentence: This external-partner study candidate would test whether an opaque sampling contract can preserve eligible units, seeded selections, and inclusion probabilities when the same approved household frame is stored differently; no field study or adopter commitment yet supports it.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-25 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Sociology Anthropology |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 15–21 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A household study may store its recruitment frame as spreadsheets, rosters, or databases that use different ordering, identifiers, and blank-value conventions. Sampling scripts can quietly depend on those details. Reordering or migrating an otherwise unchanged frame may then alter eligibility, give duplicate records extra selection chances, change a seeded draw, or assign different inclusion probabilities. Researchers could mistake those technical effects for genuine field outcomes or an authorized change in sampling design.
What is proposed¶
Define the sampling frame through public behavior rather than a particular table. Custodians would map source records to unique abstract sampling units, keep direct identifiers behind a protected boundary, and freeze an immutable snapshot with explicit eligible, ineligible, unresolved, unknown, and withdrawn states. Authorized clients could draw a sample, query probabilities, verify a draw, supersede a snapshot, or export opaque recruitment handles. The contract would require identical results for the same snapshot, design version, and seed despite row permutation, lossless reserialization, or replacement of internal keys. Duplicate physical records could not create extra chances, invalid requests could not return partial samples, and drawing could never initiate contact. Shared black-box, generated-case, transformation, and leakage tests would gate backend replacement.
The cross-domain transfer¶
The transferred archetype is a representation-independent interface contract: multiple storage systems must expose the same abstract units and observable sampling behavior while hiding their internal representation. The structural mapping is strong, although the difficult mapping from real records to unique people, households, or addresses remains a substantive governance decision rather than a software detail.
Why it advanced¶
It cleared the separately calibrated external-partner lane because the operations, comparator, failure conditions, and synthetic test are unusually specific, and mature sampling and metadata tools make a prototype feasible. It did not enter strict success: representation-caused frame divergence has not been quantified, and no study or field office has offered data or committed to evaluation.
Prior art and the remaining open claim¶
Sampling standards already require accurate, unduplicated frames, retained design information, testing, documentation, and protection of restricted data. Metadata standards and sampling products already represent frames, selection probabilities, seeds, and reproducible draws. The narrower open claim is that independently built backends passing one opaque behavioral oracle will preserve abstract snapshots, selections, probabilities, errors, and non-contact side effects across representation-only changes. This is adjacent prior art, not a novelty finding.
Smallest decisive test¶
Run a two-week trial with 200 fictional units encoded independently in a shuffled flat file and normalized database. Compare the current file-coupled scripts with both adapters used only through the contract across at least 1,000 prespecified seeds. Apply row permutations, reserialization, key replacement, ineligible-record insertion, and duplicate consolidation; also insert deliberately defective adapters. Advance only if every legitimate transformation preserves all specified outputs, every mutant is caught, no source information leaks, and reviewers can map every valid state without hiding an identity judgment. Any changed draw or probability defeats the claim.
Deployment and cost¶
The first synthetic evidence step is assessed at roughly \(10,000-\)50,000. Startup, operational launch, and annual recurring work are each roughly \(50,000-\)250,000 in 2026 resource-equivalent terms, not vendor quotes. Live deployment would additionally require approved identity rules, privacy and ethics review, security controls, version governance, training, and a certified rollback path.
Risks and uncertainties¶
- The abstract unit model could erase meaningful distinctions among households, addresses, or memberships.
- Incorrect deduplication could merge distinct eligible units; insufficient deduplication could create extra selection chances.
- Predictable seeds, handles, output order, or error details could expose identities or reveal selection.
- A technically conforming frame could still omit important parts of the intended population.
- Immutable snapshots could preserve known mistakes if supersession is too difficult to use promptly.
Expert review¶
Useful reviewer backgrounds: Survey sampling statistician, Social-research data engineer, Frame custodian or field-operations lead, Research ethics and privacy specialist, Community or participant-governance representative.
- Can every valid source state map to a unique abstract unit without concealing a disputed identity or eligibility judgment?
- Which draw outputs must be bit-for-bit identical across backends, and which implementation differences may legitimately vary?
- Do current migrations or reorderings measurably change units, selections, or weights after authorized design changes are excluded?
- Could opaque handles, deterministic seeds, ordering, errors, or timing permit reidentification or prediction?
- Who may approve eligibility, deduplication, design-version, and live-backend changes?
Evidence and provenance¶
Selected sources: S1: Statistical Quality Standard A3: Developing and Implementing a Sample Design · S2: Statistics Canada Quality Guidelines: Coverage and frames · S3: Methodology: 2022–23 Survey of Asian Americans · S4: Survey Development — DDI Lifecycle 3.3 Technical Guide · S5: sampling: Survey Sampling, version 2.11 · S6: PROC SURVEYSELECT Statement, SAS/STAT 13.1 User's Guide · S7: Pseudonymisation
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed it between ranks 15 and 21, in band B. That score only sets reading order, uses affordability as a pilot-speed proxy, and measures neither experimental success nor economic value; its endpoint remains EMPIRICAL_PARTNER_CANDIDATE.
18. Renewing Nanomaterial Stewardship Commitments¶
Canonical title: The Visible Boundary: Nanomaterial Stewardship Renewal
In one sentence: This external-partner study candidate adds a brief, consent-governed stewardship observance to ordinary nanomaterial custody controls, but no participant study yet shows that it improves ownership accuracy beyond a strong checklist and briefing.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-30 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Nanotechnology |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 1–43 across three profiles · band B |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
The problem¶
Nanomaterials and related duties pass among synthesis, fabrication, measurement, safety, storage, and waste teams. Records may exist while people disagree about who owns an unresolved containment, labeling, transport, or disposal task, when it is due, or when to escalate it. Routine signatures can conceal that gap, especially after turnover. The result may be stranded anomalies, unfunded controls, ambiguous custody, poor traceability, or avoidable exposure and contamination risks.
What is proposed¶
At project launch, each custody transfer, and quarterly thereafter, hold a 12-minute “Visible Boundary” observance outside controlled work areas. An inert token moves along a printed lifecycle map from synthesis through disposal; no real material enters the session. Relevant participants may speak, write privately, observe, or pass without explanation. They identify the current stage and unresolved uncertainties, then propose, revise, or decline symbolic commitments. An accepted commitment counts only after the normal action system records an owner, resources, deadline, and escalation condition. A safety witness reads those entries back and states that the event does not certify safety or replace procedures, training, stop-work rights, custody records, or engineering controls. Anonymous feedback and independent reviews can revise, pause, or retire the practice.
The cross-domain transfer¶
The archetype transfers recurring ritualized meaning and commitment into a technical custody setting. The marked time, lifecycle symbol, witnessing, renewal, debrief, and retirement paths map clearly. The proposed effect, however, depends on a social mechanism—shared enactment improving accurate recall and ownership—that has not been demonstrated in nanomaterial work.
Why it advanced¶
It cleared the external-partner lane because a low-cost tabletop can compare the symbolic layer directly with a strengthened checklist and briefing, using delayed individual measurements and explicit safety failures. It did not meet strict success: no facility has committed, the proposed ownership problem lacks prevalence evidence, and the central comparative effect remains untested.
Prior art and the remaining open claim¶
Lifecycle controls, training, labeling, waste procedures, worker participation, action tracking, project rituals, and safety-culture programs already exist. Research suggests group rituals can increase perceived meaning, but it does not establish better safety-relevant recall or voluntary participation under laboratory hierarchies. The remaining claim is narrower: adding this governed enactment to the same strong checklist and briefing improves agreement about owners, deadlines, uncertainties, downstream parties, and stop conditions without pressure, added errors, or false safety assurance.
Smallest decisive test¶
With facility and independent safety approval, counter-order two fictitious custody-transfer scenarios for one volunteer cross-functional group. Condition A uses a strengthened checklist and conventional briefing; condition B adds the observance. Measure each participant’s answers about ownership, deadlines, uncertainty, downstream parties, and stop conditions immediately and 48-72 hours later. A blinded safety reviewer should code omissions, invalid commitments, and false closure, while anonymous responses assess accessibility, pressure, and certification confusion. Do not advance if B lacks the prespecified accuracy gain, creates more errors, makes refusal consequential, or is mistaken for safety certification.
Deployment and cost¶
The tabletop is assessed below $10,000. Startup, launch, and annual recurring effort are each roughly \(10,000-\)50,000 in 2026 resource-equivalent terms, not vendor quotes. Any operational use would require local labor, accessibility, privacy, records, and safety review, trained independent facilitation, protected opt-outs, and continued funding for the actual controls and commitments.
Risks and uncertainties¶
- Employees or junior researchers could experience visible participation as a loyalty test.
- Ceremonial completion could be mistaken for evidence that a material or process is safe.
- Token passing could become rote and suppress uncertainty instead of surfacing it.
- The session could substitute recognition for money, staff, engineering controls, or corrective action.
- Remote, disabled, multilingual, contract, night-shift, or waste-handling workers could be excluded from the shared account.
Expert review¶
Useful reviewer backgrounds: Nanomaterial laboratory scientist, Environment, health, and safety professional, Human-factors or organizational-behavior researcher, Labor and research-ethics specialist, Waste-management or downstream custody representative.
- Do actual transfers produce disagreement about owners, deadlines, escalation conditions, or downstream effects after current records are reviewed?
- What accuracy improvement over a strengthened checklist would justify the additional social and time burden?
- Can people decline every symbolic act without their refusal becoming visible or consequential?
- Which wording and facilitation practices best prevent ceremonial closure from being interpreted as safety certification?
- Who can independently halt or retire the practice if commitments remain unfunded or participants report pressure?
Evidence and provenance¶
Selected sources: S1: General Safe Practices for Working with Engineered Nanomaterials in Research Laboratories · S2: Nanomaterials · S3: Results of the 2019 Survey of Engineered Nanomaterial Occupational Health and Safety Practices · S4: Safety Management—Worker Participation · S5: Safe Science: Promoting a Culture of Safety in Academic Chemical Research · S6: Work Group Rituals Enhance the Meaning of Work · S7: The Point of No Return: Ritual Performance and Strategy Making in Project Organizations · S8: Occupational Health and Safety Specialists and Technicians
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc ordering ranged from rank 1 under deployment-heavy weights to 43 under impact-heavy weights, placing it in band B. This volatility reflects weighting and a cost proxy, not experimental success or economic value; the endpoint remains EMPIRICAL_PARTNER_CANDIDATE.
19. Ending Mutual-Aid Mandates Truthfully¶
Canonical title: Mutual-Aid Mandate Renewal and Release Assembly
In one sentence: This external-partner study candidate would add a witnessed renewal-and-release assembly to lawful closure and transition procedures, while prominently requiring evidence that the ritual layer improves shared understanding without coercing volunteers or stranding residents.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-32 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Sociology Anthropology |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 14–25 across three profiles · band B |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
The problem¶
A neighborhood mutual-aid network may outlive the emergency that created it. Public channels, volunteer titles, pooled funds, routes, and phone trees can still signal readiness even when capacity and authority are unclear. Residents may rely on services that no longer operate; volunteers may feel unable to leave; and founders may retain accounts after nominal handoffs. Ordinary meetings can cancel tasks but may not create a legitimate, witnessed separation between ending volunteer roles and completing remaining duties.
What is proposed¶
Hold a quarterly and trigger-based Mandate Renewal and Release Assembly while any function remains open. Beforehand, stewards inventory advertised services, capacity, dependencies, funds, data, decision authority, reliance, and possible recipients. A status board marks each function as under review, and function cards enter provisional lanes: renew, transfer, suspend, repair, or release. Card placement is not a decision. Willing volunteers, lawful resource holders, and safeguards for affected residents must support the operational disposition. Closure then updates notices, rosters, escalation paths, funds, credentials, data, referrals, and deadlines. Participants may speak, submit privately, remain silent, delegate, or leave. Retiring a role never extinguishes debts, safeguarding duties, contracts, or promised transitions, and later audits compare public claims with actual capacity and custody.
The cross-domain transfer¶
The ritualized commitment archetype becomes a recurring, witnessed review of collective mandates, using marked time, function cards, reflection, renewed consent, release, handoff, debrief, and later audit. The mapping is coherent, but the distinct contribution of ceremony is uncertain because sunset rules, transition plans, demobilization, closure checklists, and volunteer recognition already perform much of the work.
Why it advanced¶
It entered the external-partner lane because it separates voluntary labor from surviving legal and service duties, names the necessary authorities, and offers a bounded comparator study. It did not enter strict success: no network has committed, stale-capacity prevalence is unknown, and there is no evidence that the assembly outperforms ordinary closure and transition practice.
Prior art and the remaining open claim¶
Emergency demobilization, nonprofit dissolution, service transfer, volunteer disengagement, debriefing, recognition, and care-oriented organizational closure are established. The open comparison begins only after a valid sunset rule, authorized closure checklist, and safe transition plan exist. The question is whether the witnessed assembly further improves volunteers’, residents’, and custodians’ accuracy about service status, authority, refusal rights, and unfinished duties without added coercion, disclosure, delay, or stranded work. It is a combination-level claim amid adjacent prior art.
Smallest decisive test¶
With a named network and legal or fiscal sponsor, run a compensated, preregistered tabletop with 24-40 volunteers, resident representatives, custodians, and potential partners. Counterbalance matched, noncritical mock mandates between ordinary sunset, checklist, and transition procedures and the same package plus the assembly. Measure service status, authority, refusal paths, notices, custody, and unfinished duties immediately and seven days later, alongside pressure, disclosure, distress, time, and false-closure beliefs. Reject incremental advantage if accuracy does not materially improve, burdens rise without compensating benefit, or any simulated duty is stranded. Halt for pressure, distress, privacy failure, or disputed authority.
Deployment and cost¶
The tabletop is assessed below $10,000; startup, launch, and annual recurring effort are each roughly \(10,000-\)50,000 in 2026 resource-equivalent terms, not vendor quotes. Live costs could rise with legal complexity, critical services, debts, accessibility needs, data cleanup, and funded transitions. Ordinary legal and operational authority remains mandatory.
Risks and uncertainties¶
- Residents could lose needed support before a capable alternative and truthful notice are in place.
- Public discussion could expose a resident’s need or a volunteer’s exhaustion.
- Recognition, gratitude, or lament could pressure volunteers to renew their labor.
- Founders could curate the evidence or retain credentials despite a witnessed handoff.
- Symbolic release could be misused to avoid debts, delete data improperly, or leave contracts and safeguarding duties incomplete.
Expert review¶
Useful reviewer backgrounds: Mutual-aid organizer or former volunteer, Affected-resident representative, Nonprofit or fiscal-sponsor lawyer, Data-protection and safeguarding specialist, Service-transition or emergency-demobilization practitioner.
- Which functions are legally controlled, informally controlled, safety-critical, or dependent on outside partners?
- Can volunteers withdraw immediately while residents still receive a safe and truthful transition?
- What evidence would show that the assembly adds understanding beyond a sunset clause, closure checklist, and facilitated transition meeting?
- How will private reliance statements and participation choices be protected from founders, employers, or public records?
- Who verifies that credentials, funds, notices, data, property, and surviving duties actually moved after the witnessed decision?
Evidence and provenance¶
Selected sources: S1: More Than a COVID-19 Response: Sustaining Mutual Aid Groups During and Beyond the Pandemic · S2: Covid Aid ceasing as a charity, with Support Community to continue under independent member ownership · S3: Demobilizing the VRC (2 of 2) · S4: Dissolution · S5: The Wind Down—For Better Organizational and Project Endings · S6: Group Termination: Completing the Study of Group Development · S7: Disengagement · S8: Occupational Employment and Wages—May 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading order placed it between ranks 14 and 25, in band B. The calculation uses affordability as a pilot-speed proxy and does not establish experimental success, deployment merit, or economic value; its endpoint remains EMPIRICAL_PARTNER_CANDIDATE.
20. Safer Cleanup of Post-Production Assets¶
Canonical title: Dependency-aware lifecycle governance for post-production assets
In one sentence: This post-hoc survivor proposes dependency-aware lifecycle rules for media-production files, but every major component has adjacent prior art and the complete workflow has not been compared with a manual sweep or configured asset-management system.
| Field | Record |
|---|---|
| Portfolio ID | EXP03-POSTHOC-02 |
| Experiment and endpoint | Experiment 3 · Post-hoc strict innovation-like survivor |
| Archetype × domain | Layer Decay And Expiration Management × Film Media Production |
| Proposal position or arm | Not recorded |
| Post-hoc reading order | Balanced score 68.0/100 · rank range 15–20 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A production accumulates edit versions, visual-effects renders, proxies, sound stems, exports, and project files whose usefulness changes over time. Old artifacts can remain visible as if current, consume active storage, and slow retrieval. Yet deleting them by age, name, or staff intuition can break project relinking, reconstruction, rights evidence, or future remastering. The core tension is distinguishing safely demotable clutter from creative or evidentiary material whose future dependencies are incomplete or hidden.
What is proposed¶
Assign each artifact an explicit lifecycle and creative-authority state, then assess dependencies, rights, legal holds, ownership, reconstruction needs, and restoration context before moving or deleting anything. Approved artifacts could move to cheaper storage or reversible quarantine; persistent markers would prevent restored or superseded files from silently appearing current again. Restore drills would test checksums, relinking, codecs, plug-ins, and software context, while recurring reviews would revisit earlier classifications. Owners would approve exceptions, archive or IT staff would verify recoverability, and legal or rights staff would control relevant holds and irreversible destruction. The proposal aims to bound active storage and search noise without silently sacrificing usable production history, but automatic deletion based only on age, access, score, or filename is excluded.
The cross-domain transfer¶
The layer-decay archetype maps to media files whose authority and usefulness expire at different rates. Lifecycle states nominate refresh, demotion, quarantine, or deletion, while dependency and restoration checks prevent unsafe transitions. The analogy is imperfect because dormant creative work can regain value years later and may depend on proprietary tools, external vendors, or undocumented practice.
Why it advanced¶
After Experiment 3, the candidate survived a stricter combined opportunity screen because the problem is concrete, existing technical primitives permit a bounded replay, and safety outcomes are measurable. This was a post-hoc result, not a preregistered success. Prevalence, net benefit, complete dependency capture, and adopter acceptance remain unestablished.
Prior art and the remaining open claim¶
Media-asset systems already archive and restore files with permissions; media ontologies represent versions and provenance; preservation standards cover rights, events, relationships, and technical dependencies; workflow tools track stale data; and cloud systems automate storage tiers. The only remaining contrast is the operating-policy combination: creative-authority state, dependency and reconstruction coverage, rights and owner gates, reversible quarantine, anti-staleness markers, restore drills, and recurring revalidation. The search did not establish novelty or superiority over a configured media-asset and archive workflow.
Smallest decisive test¶
On one completed, non-litigated production, conduct a read-only replay using a documented random sample of 200 artifacts from two classes. Independent reviewers compare a manual sweep, verified off-the-shelf asset-management rules, and the full proposal. Using isolated copies, simulate tiering and quarantine and test checksums, restoration, relinking, and software context. Compare labor, retrieval time, authoritative-version errors, safely demotable bytes, disputed labels, dependency misses, and restore success. Do not advance unless the full workflow materially improves a user outcome and active-tier burden after labor without worsening safety. Any missed live dependency, hold conflict, failed integrity check, relink, or restore defeats advancement.
Deployment and cost¶
First evidence is assessed at roughly \(10,000-\)50,000. Startup is roughly \(50,000-\)250,000, operational launch \(250,000-\)1 million, and annual recurring work \(50,000-\)250,000 in 2026 resource-equivalent terms, not vendor quotes. Actual costs depend on artifact volume, integrations, storage, licensing, legal review, vendor access, and recurring classification labor.
Risks and uncertainties¶
- The workflow could destroy unique creative work, contractual evidence, or material needed for a later reconstruction.
- Dependency graphs may omit external vendors, proprietary project formats, plug-ins, codecs, fonts, licenses, or tacit creative steps.
- Lifecycle labels could falsely imply that authority or disposal eligibility is settled.
- Archive latency and retrieval charges could disrupt an unexpected reuse request.
- Quarantine records and anti-staleness markers could themselves become another unmanaged layer.
Expert review¶
Useful reviewer backgrounds: Post-production supervisor or assistant editor, Media archivist or preservation engineer, VFX and sound pipeline engineer, Production IT or media-asset-management administrator, Entertainment rights, contracts, or records counsel.
- Can the production identify a canonical artifact and its authoritative status across edit, VFX, sound, and delivery systems?
- How complete are dependency records for external vendors, software versions, plug-ins, codecs, fonts, and licenses?
- Which outcomes would justify the added review labor compared with a configured media-asset system?
- Who has authority to approve demotion, quarantine, exceptions, holds, and irreversible destruction for each artifact class?
- What restore and relink failures are acceptable, if any, before all destructive transitions must stop?
Evidence and provenance¶
Selected sources: S1: Learning to Love End-to-End Cloud Production in Love Hurts · S2: Archive and Restore · S3: Ontology for Media Creation: Versions Ontology v2.7 · S4: Managing Audiovisual Records · S5: Depends: Workflow Management for Research and Visual Effects · S6: PREMIS Data Dictionary, Version 3.0, Now Available · S7: Archive a Blob · S8: National Employment and Wage Data by Occupation, May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading order placed it between ranks 15 and 20, in band B. This ordering is not an experimental endpoint or economic-value measure. Its endpoint remains POST_HOC_STRICT_INNOVATION_LIKE_SURVIVOR, not a preregistered success.
21. Show What Actually Solved the Model¶
Canonical title: Oracle-Relative Guarantee Passports for Macroeconomic Policy Models
In one sentence: A passport would distinguish what a macroeconomic policy platform computed itself from results supplied by solvers, data services, or people, while preserving bounded and unresolved outcomes.
| Field | Record |
|---|---|
| Portfolio ID | EXP05-STRICT-01 |
| Experiment and endpoint | Experiment 5 · Strict success |
| Archetype × domain | Computability Boundary Mapping × Economics Finance |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 68.0/100 · rank range 16–30 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
Policy models can contain unrestricted agent programs, continuous calculations, external equilibrium solvers, live data, and human choices among possible equilibria. Yet a final policy ranking may simply say that the model converged. That label can hide who or what supplied the answer, whether the supplier could abstain, whether other trajectories remained unresolved, and whether a finite timeout was mistaken for universal stability. The result may therefore claim more computational certainty than the platform actually established.
What is proposed¶
Before a model enters formal policy comparison, give it a capability-relative guarantee passport. Freeze the model language, economic target, equilibrium and stability definitions, time domain, and quantifiers. Define what the base platform can compute, then register each solver, data source, real-number routine, interactive environment, and human selection step as a named external capability with explicit promises, errors, latency, abstention, accountability, and version behavior. Test whether conclusions survive withdrawal, invalid responses, and promise violations. Search fairly, within declared bounds, for trajectories that refute convergence or target attainment. Route every result as BASE_CERTIFIED, ORACLE_RELATIVE, COUNTEREXAMPLE_FOUND, BOUNDED_ONLY, UNKNOWN, OUT_OF_MODEL, or SYSTEM_FAILURE. A timeout remains UNKNOWN, and a human-selected equilibrium remains assisted. Any claimed unrestricted impossibility reduction must be independently checked before use.
The cross-domain transfer¶
The computability-boundary archetype maps strongly here. External solvers, live information, and adaptive human choices act like added computational capabilities. The passport records which question the base program answers, which answer depends on a named capability and its promises, and which questions remain unresolved. Checked countertrajectories can refute precise universal claims without pretending to decide every possible model.
Why it advanced¶
This candidate passed Experiment 5's strict researched-candidate bar because the problem is consequential, the classifications and offline pilot are technically feasible, and existing model-governance functions could own them. That endpoint is only a researched-candidate result: it does not establish real-world effectiveness, novelty, economic value, adopter commitment, or authorization for policy use.
Prior art and the remaining open claim¶
Model inventories, validation, lifecycle governance, model cards, and documentation of human or third-party limits already exist. Numerical platforms also disclose dependence on initial guesses, iteration limits, and equilibrium selection. The narrower untested claim is that named capability contracts, promise checks, withdrawal tests, and mechanically preserved result labels will catch more capability-laundering errors than an ordinary model inventory and validation package, without changing the economic question being asked. No world-novelty finding was made.
Smallest decisive test¶
Preregister an offline experiment using six synthetic models and stubbed external services. Randomize cases between the ordinary inventory and validation template with raw logs, and the passport workflow. Blinded reviewers classify each result's computational dependency and status. Compare classification accuracy, false BASE_CERTIFIED results, preservation of UNKNOWN, promise-violation detection, review time, and semantic fidelity. Independently check any reduction and countertrajectory. Reject the incremental claim if accuracy does not improve, any false BASE_CERTIFIED label appears, UNKNOWN is lost, withdrawal cannot localize dependency, or at least two independent macroeconomists find that formalization materially changes the intended question.
Deployment and cost¶
The authorized first step is an offline synthetic pilot with no live policy instruments, confidential feeds, forecasts, or production materials. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(250,000–\)1 million annually. These are assessment bands, not vendor quotes.
Risks and uncertainties¶
- The formal equilibrium or target definition may omit what policy staff actually mean to evaluate.
- Staff may treat capability labels as paperwork and still collapse assisted, bounded, and unknown results into one ranking.
- Analysts may make adaptive equilibrium choices outside the registered procedure.
- A solver or data service may change behavior without a visible version change.
- Some solver promises may depend on semantic or empirical conditions that cannot be checked mechanically or conservatively screened with confidence. A base-certified calculation may still use an empirically poor economic model and produce a misleading policy conclusion. A valid impossibility result for unrestricted models may be wrongly extended to finite or promised subclasses.
Expert review¶
Useful reviewer backgrounds: Macroeconomist who develops or validates policy models, Computability or formal-methods researcher, Numerical equilibrium-solver specialist, Central-bank model-governance lead, Independent policy-model reviewer.
- Can the proposed formal convergence property preserve the economic question used by policy staff without silently narrowing it?
- For each external solver or human step, which promise conditions can be checked before its output is accepted?
- Can the unrestricted model class and policy query support a correct, independently checkable undecidability reduction?
- Do downstream policy materials preserve ORACLE_RELATIVE and UNKNOWN labels, or collapse them into a single ranking?
- Against ordinary validation documentation, how many dependency-classification errors does the passport prevent, and at what review cost?
Evidence and provenance¶
Selected sources: S1: Annual Report 2025, Box 7: Macroeconomic modelling in times of uncertainty · S2: The Dynare Reference Manual, version 7.1: The model file · S3: No-Regret Learning in Games is Turing Complete · S4: Supervisory Guidance on Model Risk Management, SR 26-2 Attachment · S5: Model Cards for Model Reporting · S6: Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 · S7: ECB macroeconometric models for forecasting and policy analysis: Development, current practices and prospective challenges · S8: Employer Costs for Employee Compensation—March 2026
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed this candidate in Band B, ranking 16–30 across three weighting profiles. That score is only a post-hoc reading order; it is not an experimental endpoint or economic-value measure, and its pilot-speed input is a cost-band proxy.
22. Keep Wetland Records Stable Across Systems¶
Canonical title: Representation-Independent Evidence Ledger for Wetland Greenhouse-Gas Monitoring
In one sentence: A behavioral evidence ledger would test whether different storage systems produce the same wetland monitoring history, coverage, provenance, and greenhouse-gas aggregates without exposing their internal layouts.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-19 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Environmental Climate |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 68.0/100 · rank range 11–29 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-20, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Wetland observations, laboratory results, corrections, and quality decisions may live in workbooks, database tables, raster layers, or files. Analysis scripts can quietly depend on row order, worksheet names, sentinel values, grid resolution, or a database's duplicate rules. A storage migration or regridding can then change which observations count, how revisions apply, or how values are aggregated even though the nominal fields remain. That makes it hard to tell a scientific change from an implementation-induced discontinuity.
What is proposed¶
Define the monitoring record by its behavior, not by its files or tables. The ledger would accept observation batches, retain superseded versions and reasons, apply or revoke quality decisions, answer version-specific queries, compute only declared variable-appropriate aggregates, report coverage and provenance, and create repeatable snapshots. Submissions must include stable identity, units, spatial and temporal support, method, and authorized provenance. Rejected operations must leave state unchanged and return stable errors for malformed data, incompatible units, duplicate identities, unsupported aggregation, unauthorized revision, or unavailable coverage. Workbooks, relational stores, rasters, indexes, and caches remain hidden behind the interface. Every backend must pass the same examples, generated operation sequences, and representation-leakage audit. Breaking semantic changes require a new contract version rather than reinterpretation of an existing snapshot.
The cross-domain transfer¶
The representation-independent interface archetype maps directly to storage substitution: different physical systems should denote the same versioned evidence state and obey the same operations, invariants, errors, and side-effect rules. The mapping is weaker for scientific validity. Backend agreement cannot show that field measurements, quality policies, ecological assumptions, or aggregation conventions are themselves correct.
Why it advanced¶
This candidate did not enter the strict-success lane. It cleared a separately calibrated lane for a bounded external data-partner study because the components are implementable and the claim is testable. Crucially, there is no real dependency inventory, authorized replay, measured discrepancy rate, comparative result, committed wetland partner, or project-specific cost evidence.
Prior art and the remaining open claim¶
Observation schemas, provenance standards, greenhouse-gas repositories, automated quality checks, versioning, and ecological workflow tools already provide close constituent parts. The remaining claim is narrower: for one frozen wetland workflow, a contract covering identity, supersession, quality decisions, units, coverage, provenance, errors, snapshots, and variable-aware aggregation will make an incumbent adapter and an independent backend agree on every management-relevant public result without clients inspecting storage internals. That comparison has not been run.
Smallest decisive test¶
With written owner approval, copy one completed project's data into a read-only sandbox and freeze 80–200 real or redacted operations spanning measurement, correction, quality review, snapshot, and aggregation. Compare the unchanged workflow, a schema-normalized export, and the behavioral contract implemented by both an incumbent adapter and an independent in-memory model. Add generated duplicates, reordered calls, incompatible units, revoked decisions, incomplete coverage, and authorization failures. Pass only with zero unexplained differences across declared results and no internal access. Reject the problem locally if no internal dependence exists; reject the intervention if conforming implementations still produce a management-relevant divergence or require exposing the incumbent layout or algorithm.
Deployment and cost¶
The first study must remain read-only and cannot alter official balances, records, eligibility, crediting, compliance, or management decisions. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for launch, and \(50,000–\)250,000 annually; they are not vendor quotes.
Risks and uncertainties¶
- The contract may preserve an incorrect scientific convention because incumbent behavior is mistaken for intended meaning.
- Unit or spatial-support normalization may combine observations that should instead be rejected as incompatible.
- An adapter defect may be misdiagnosed as a difference between the underlying storage systems.
- Finite conformance tests may miss operation sequences that alter a management-relevant result.
- Opaque storage may impede legitimate scientific inspection unless provenance and approved diagnostic views are adequate. Stable identifiers and detailed provenance may disclose sensitive site information if access controls are weak. A broad contract may become expensive to govern, while a narrow one may omit decision-critical behavior.
Expert review¶
Useful reviewer backgrounds: Wetland greenhouse-gas monitoring scientist, Environmental data architect, Scientific provenance and standards specialist, Carbon-program monitoring or methodology lead, Data-governance and confidential-location reviewer.
- Which observation, correction, quality, coverage, and aggregation behaviors can change a management-relevant greenhouse-gas result?
- Do any current client scripts inspect worksheet coordinates, schemas, file paths, raster cells, or private status codes?
- Which variables are extensive or intensive, and what aggregation rules preserve their scientific meaning?
- Can two independently implemented backends pass the contract yet still disagree on an official or management-relevant output?
- What access, retention, and disclosure rules apply to site identities and provenance in the proposed sandbox?
Evidence and provenance¶
Selected sources: S1: 2013 Supplement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories: Wetlands · S2: VM0033 Methodology for Tidal Wetland and Seagrass Restoration, v2.1 · S3: Global Greenhouse Gas Watch: Data Management · S4: World’s Largest Carbon Program Pilots Digital Measuring of Forest Carbon · S5: Observations, Measurements, and Samples · S6: PROV-N: The Provenance Notation · S7: OpenGHG 0.19.0 Developer API · S8: Developing a modern data workflow for regularly updated data
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized review placed this candidate in Band B, ranking 11–29 depending on weighting. This is a reading-order aid, not its endpoint or a value estimate; the pilot-speed input reflects cost-band affordability rather than measured elapsed time.
23. Retire Outdated Course Guidance Safely¶
Canonical title: Lifecycle Registry for Accumulated Course Scaffolds
In one sentence: An artifact registry would help course teams find stale guidance across copied course offerings while protecting accessibility supports, assessment dependencies, and records that must remain recoverable.
| Field | Record |
|---|---|
| Portfolio ID | EXP04-STRICT-04 |
| Experiment and endpoint | Experiment 4 · Strict success |
| Archetype × domain | Layer Decay And Expiration Management × Education Pedagogy |
| Proposal position or arm | PROPOSAL_FIRST |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 18–27 across three profiles · band B |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
Repeated course copies accumulate hints, examples, rubrics, refreshers, prompts, accommodations, and corrective notes. Older items may remain visible after the syllabus, assessment, learner group, or linked resource changes, leaving students with contradictory or obsolete guidance. Staff cannot safely clean everything up: an old-looking item may still support an accommodation, assessment, grade review, accreditation record, or reconstruction of a prior offering. Current course tools expose some clutter and broken links but do not resolve these lifecycle decisions.
What is proposed¶
Give each instructional-support artifact an identity, owner, source course version, lifecycle state, validation date, dependencies, retention class, and provisional expiry trigger. A dashboard would flag suspected stale items and rank review priority using age, present access, contradiction, ownership, and reconstruction value, but a person would decide what happens. Authorized reviewers could refresh, retain, unpublish, archive, compact, hold, or place an item in reversible quarantine. Accessibility, assessment, grade-review, accreditation, legal, and records dependencies must be checked before removal. Retired items leave a marker pointing to a successor, archive, or explicit unavailable state. Temporary supports receive review dates when created. Periodic revalidation revisits active items and exception holds, while restore tests confirm that archived materials remain readable, attributable, and connected to the correct historical course.
The cross-domain transfer¶
The layer-decay archetype has a clear structural match: each course offering deposits another support layer, and old layers can look current. Lifecycle states, expiry reviews, dependency checks, quarantine, retirement markers, and restore tests bound the active layer without erasing protected history. Age is only a review signal, however; it cannot determine whether instructional content remains valid.
Why it advanced¶
This candidate passed Experiment 4's strict researched-candidate bar because institutions already perform course cleanup and retention work, relevant owners exist, and a read-only comparison is feasible. The endpoint does not mean the registry works in practice, is novel, saves money, improves learning, or has institutional authorization beyond a bounded study.
Prior art and the remaining open claim¶
Manual course audits, link validation, cleanup tools, content templates, version distribution, retention schedules, and archives are established. The open contrast is whether an artifact-level registry spanning course generations improves identification of current versus stale supports and surfaces assessment, accessibility, and records dependencies better than a careful manual audit aided by Canvas Link Validator and TidyUP. It must also preserve recovery and add no more than 20% reviewer time. Only that integrated, measured comparison remains open; world novelty was not assessed.
Smallest decisive test¶
With instructor, LMS, accessibility, privacy, and records approval, use four sandboxed snapshots of one course, excluding submissions, grades, and identifiable accommodation records. Sample at most 80 instructor-created supports and compare the existing manual audit plus Link Validator and TidyUP with the same evidence plus the registry. Two authorized reviewers classify currentness, visibility, dependencies, exceptions, and proposed state; seed up to eight synthetic defects. Proceed only with sensitivity of at least 0.85, false flags at most 0.10, Cohen's kappa at least 0.70, no protected-data exposure, and median review time within 20% of baseline. Reject incremental advantage if classification or dependency recall does not improve.
Deployment and cost¶
Begin with a read-only inventory; do not change visibility, content, grades, assessments, accommodations, or retention. Rough 2026 resource-equivalent bands are under $10,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(50,000–\)250,000 annually. These are assessment bands, not vendor prices.
Risks and uncertainties¶
- Reviewers may mistake age for staleness even when an old explanation remains correct.
- Low access counts may unfairly flag essential but rarely used accessibility or safety materials.
- Incomplete links to assessments, external tools, grade reviews, or accommodations may make a load-bearing artifact appear safe to retire.
- A composite score may create false precision or encode local pedagogical preferences.
- Quarantine could conflict with mandatory destruction rules or retain sensitive material too long. Lifecycle labels may confuse students if exposed without careful wording. Registry metadata and temporary holds may themselves become stale. Successful restore tests on sampled formats may conceal unreadable materials elsewhere in the archive.
Expert review¶
Useful reviewer backgrounds: Instructional designer, Instructor experienced with repeated LMS course copies, LMS administrator or integration specialist, Accessibility and disability-services representative, Academic records, privacy, or retention officer.
- Can artifacts be matched reliably across copied course shells without confusing distinct items or missing descendants?
- Which dependencies can the LMS and external tools expose, and which still require human review?
- What local rules distinguish ordinary instructional supports from protected educational or grade-evaluation records?
- Do reviewers agree on currentness, contradiction, dependencies, and lifecycle state at the required thresholds?
- Does the registry improve detection over Link Validator, TidyUP, and manual review without exceeding the 20% time limit?
Evidence and provenance¶
Selected sources: S1: Declutter your Learn.UQ course checklist · S2: Course Data Purge Process · S3: Postgraduate Students’ Experience of Using a Learning Management System to Support Their Learning: A Qualitative Descriptive Study · S4: How do I validate links in a course? · S5: TidyUP · S6: Course Content Distribution Comparison · S7: Grades and Records of Student Performance — Academic Policy 1480.10 · S8: Occupational Employment and Wage Statistics: Instructional Coordinators
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed the candidate in Band B, with ranks from 18 to 27 across weighting profiles. This post-hoc ordering does not alter its strict endpoint or measure economic value; pilot speed was represented only by a cost-band proxy.
24. Test Cold-Case Theories on Equal Terms¶
Canonical title: Parallel-Hypothesis Gate for Cold-Case Resources
In one sentence: A bounded hypothesis gate would compare rival cold-case explanations using equal records, preregistered predictions, capped resources, and rights-aware scoring without turning advancement into a finding of guilt.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-05 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Criminology Forensic |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 24–28 across three profiles · band B |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
Cold-case units may have several plausible explanations but limited analyst time, specialist review, and laboratory capacity. The incumbent team can control the chronology, evidence requests, and briefing order, allowing one theory to advance through ownership or presentation rather than discriminating evidence. Rival teams can also duplicate work, hoard information, or escalate intrusive requests. A resource-allocation ranking may then be mistaken for proof about a named person even though it only selects what to examine next.
What is proposed¶
For at least two distinct, minimally supported hypotheses, create separated teams with equal access to a read-only, citation-addressable record. Before new results arrive, each team registers its theory, contrary evidence, falsifiers, discriminating predictions, uncertainty, proposed tests, and expected privacy, evidence-consumption, and third-party burdens. Freeze eligibility, scoring, conflicts, appeals, and prohibited conduct in advance. Equal analyst-hour and submission caps limit escalation. An independent panel scores evidence coverage, falsifiability, distinctiveness, provenance, treatment of contradictions, feasibility, and rights-adjusted information value. At most two complementary hypotheses receive a bounded validation slot, not investigative authority. Advancement establishes neither guilt, probable cause, admissibility, nor permission for contact, search, surveillance, testing, charging, or public accusation. New evidence or failed predictions can reopen entry, and later review compares preregistered predictions with results.
The cross-domain transfer¶
The bounded-rivalry archetype maps to several hypotheses competing for scarce analysis and testing. The gate defines who may compete, equalizes inputs and budgets, penalizes harmful off-process behavior, permits more than one bounded award, and reopens competition after new evidence. The structural mapping is plausible, but rivalry could intensify team commitment or strategic withholding instead of improving reasoning.
Why it advanced¶
This candidate did not reach the strict-success lane. It cleared a separate empirical-partner lane because a masked retrospective comparison is measurable and can be kept from affecting cases. The central field evidence is missing: no agency partner, prevalence estimate, comparative trial, validated scoring rubric, measured local cost, or evidence that team separation is safer than collaboration exists.
Prior art and the remaining open claim¶
Multidisciplinary cold-case review, Analysis of Competing Hypotheses, explicit alternatives, evidence for and against each theory, ordered information exposure, provenance, and auditable forensic reasoning already exist. Research also warns that formal hypothesis layouts do not reliably reduce bias. The remaining claim is that separated teams, identical cutoff records, registered predictions, equal budgets, rights-adjusted scoring, and no more than two validation slots outperform both case conferences and ordinary ACH without increasing unsupported allegations, intrusion, evidence use, entrenchment, withholding, or panel inconsistency.
Smallest decisive test¶
With an authorized partner, preregister a three-condition retrospective crossover using 8–12 masked, time-split closed or synthetic cases: the proposed gate, a multidisciplinary conference, and a noncompetitive ACH worksheet. Give each condition identical cutoff records and analyst hours, rotate qualified participants, and prevent case recognition or outcome leakage. Before revealing later evidence, capture citations, contradictions, probabilities, predictions, actions, expected information value, privacy burden, and evidence consumption. Blinded reviewers assess calibration, discrimination, citation completeness, unsupported allegations, redundancy, reliability, time, and intrusion. Reject the claim if the gate fails to beat the better comparator or worsens calibration, agreement, withholding, entrenchment, allegations, privacy burdens, or evidence-consumption proposals.
Deployment and cost¶
Only a retrospective, no-case-impact simulation using masked copies is authorized initially; it cannot reopen a case or affect any person or evidence. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for startup and operational launch, and \(1–\)5 million annually. They are not vendor quotes.
Risks and uncertainties¶
- Assigned teams may become more committed to their hypotheses and resist contrary evidence.
- Teams may withhold urgent exculpatory information until scoring or relabel similar explanations to qualify separately.
- Scorers may reward narrative confidence or specificity rather than evidentiary value and calibration.
- Historical cutoff packets may omit context legitimately available to investigators at the time.
- Case recognition or later-outcome knowledge may leak to participants or reviewers. Resource caps may disadvantage a genuinely complex explanation. Even confidential rankings could stigmatize investigators or people named in hypotheses. Officials or the public may misrepresent advancement as evidence of guilt or authority for coercive action.
Expert review¶
Useful reviewer backgrounds: Cold-case or major-case review supervisor, Forensic scientist or laboratory evidence custodian, Investigator trained in competing-hypothesis analysis, Prosecutor, defense-disclosure, or criminal-procedure specialist, Privacy, research-ethics, and experimental-design reviewer.
- Can masked cutoff packets provide equal and sufficiently complete records without revealing case identities or later outcomes?
- Does the scoring rubric produce reliable rankings across independent panels and resist gaming through confidence or narrative style?
- Compared with case conferences and ACH, does the gate improve held-out calibration and rights-adjusted information value?
- Does team separation increase withholding, hypothesis entrenchment, unsupported allegations, privacy burden, or duplicative evidence requests?
- Which local authorities must approve record access, disclosure handling, laboratory recommendations, retention, and participant involvement?
Evidence and provenance¶
Selected sources: S1: National Best Practices for Implementing and Sustaining a Cold Case Investigation Unit · S2: NamUs Cold Case Advisory Process · S3: Conducting Effective Investigations: Practice Evidence · S4: Tunnel Vision and Confirmation Bias Among Police Investigators and Laypeople in Hypothetical Criminal Contexts · S5: Psychology of Intelligence Analysis · S6: Effects of Task Structure and Confirmation Bias in Alternative Hypotheses Evaluation · S7: OSAC 2022-S-0030 Standard Methodology in Bloodstain Pattern Analysis, Version 2.1 · S8: Disclosure Manual: Chapter 5 — Reasonable Lines of Enquiry and Third Parties
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized review placed this candidate in Band B, ranking 24–28 across weighting profiles. That ordering is neither its empirical-partner endpoint nor an economic-value measure, and the pilot-speed input is an affordability proxy rather than observed study duration.
25. Send Raman Surprises, Preserve Full Evidence¶
Canonical title: Versioned Residual Raman Monitoring for Remote Reaction Experiments
In one sentence: A shadow study would test whether transmitting reconstructable differences from predicted Raman spectra can reduce data and review burdens without hiding consequential chemical or instrument changes.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-11 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Chemistry Materials |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 12–39 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-12, EXP06-PARTNER-13, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Remote reaction monitoring can produce a dense stream of Raman spectra. Most of each spectrum may reflect predictable solvent bands, baseline behavior, or gradual temperature effects, yet every spectrum still consumes transmission, storage, and chemist attention. A small unexpected peak, shifted band, or changed intensity may therefore be reviewed late. Simply sending selected peaks or anomaly labels is also risky because it can hide unfamiliar changes and deny the chemist enough context to reconstruct or challenge the result.
What is proposed¶
Before each acquisition, matched, versioned models at the instrument and review station would predict the next wavelength-aligned spectrum from accepted prior data and declared reaction conditions. The instrument computer would subtract that frozen prediction from the complete observation and transmit the signed correction, uncertainty, context, heartbeat, and model checksum. The receiver could reconstruct the spectrum by adding the correction to its synchronized prediction. Reliable, persistent, or consequential differences would bring the corresponding full spectral window to a named chemist. Random complete spectra and scheduled full-state anchors would provide independent checks. Missing heartbeats, incompatible versions, poor detector quality, stale calibration, structured errors, excessive reconstruction error, or protected safety conditions would restore full-spectrum transmission. Independent safety instruments and the original reaction controls would remain outside this system.
The cross-domain transfer¶
The predictive-residual archetype maps directly onto the spectral stream: predict one complete spectrum, compare it with the actual detector output, communicate the signed mismatch, and reconstruct the observation at the receiver. Version checks, raw audits, full anchors, and automatic decompression address the central danger that both endpoints could share the same wrong expectation.
Why it advanced¶
This is an empirical-partner candidate, not a strict success or real-world validation. Real-time Raman acquisition and spectral compression already exist, and the proposed shadow comparison is technically bounded. It advanced because its measurements and rejection rules are unusually concrete, despite having no field evidence that the target workflow suffers a binding bandwidth or attention constraint.
Prior art and the remaining open claim¶
Adjacent prior art includes commercial real-time Raman systems with multivariate predictors, compressive Raman acquisition, and standardized lossless or near-lossless spectral compression. The open claim is narrower: for one instrument, probe, reaction family, and one-scan horizon, a checksum-gated reconstructive residual stream with raw audits, heartbeats, full anchors, protected bypasses, and forced fallback would use fewer total resources than both full-spectrum streaming and ordinary lossless compression while preserving blinded decisions, protected perturbations, and regional error limits.
Smallest decisive test¶
A laboratory partner would run twelve non-hazard-escalating reactions: eight ordinary runs and four containing approved spectral, calibration, detector, heartbeat, and version challenges. Every raw spectrum would remain authoritative. Blinded chemists would compare full streaming, lossless compression, and the frozen residual path on total bytes, storage, computation, audit and fallback traffic, review minutes, decisions, reconstruction error, and protected-event capture. Reject the claim if any protected challenge is missed, reconstruction changes a material decision, a heartbeat or checksum fault fails to force raw mode, an error limit is exceeded, or net savings are absent.
Deployment and cost¶
The first shadow study requires instrument and API access plus bench-chemist, instrumentation-owner, safety, and information-security approval. Its rough first-evidence band is \(10,000–\)50,000 in 2026 resource equivalents. Estimated startup is \(50,000–\)250,000; operational launch \(250,000–\)1 million; and annual recurring work \(50,000–\)250,000. These are assessment bands, not vendor quotes.
Risks and uncertainties¶
- A shared wrong predictor could suppress an unexpected low-amplitude peak at both endpoints while producing a plausible reconstruction.
- Instrument fouling or changing chemistry could gradually be treated as normal rather than as evidence of a new regime.
- Alignment error, quantization, or lost corrections could distort later reconstructed spectral shapes.
- A chemist-approved consequence table could still privilege anticipated bands and discount unfamiliar features.
- Random raw audits may miss rare failures; a clean small sample cannot demonstrate completeness or justify later raw-data deletion.
Expert review¶
Useful reviewer backgrounds: Raman spectroscopy and chemometrics specialist, Reaction or process chemist, Laboratory instrumentation engineer, Laboratory safety and data-integrity specialist, Statistical audit-sampling expert.
- Which low-amplitude or unfamiliar spectral changes must always bypass suppression for the selected reaction family?
- What regional and cumulative reconstruction limits would preserve the chemist's actual reaction-state decisions?
- Does the intended workflow have measured transmission, storage, queue-delay, or review burdens that ordinary lossless compression cannot address?
- What random-audit size and stratification would have adequate power to reveal rare suppressed changes?
- Which instrument faults and calibration conditions can be injected safely without changing the controlling record?
Evidence and provenance¶
Selected sources: S1: PAT — A Framework for Innovative Pharmaceutical Development, Manufacturing, and Quality Assurance · S2: Raman spectrometry as a tool to study minimization of batch age effects and make product quality decisions for biotherapeutic antibody production · S3: Relative Intensity Correction Standards for Fluorescence and Raman Spectroscopy · S4: Questions and Answers on Current Good Manufacturing Practice Requirements—Records and Reports · S5: Raman Rxn2 analyzer powered by Kaiser Raman technology · S6: Integration of a Raman spectroscopic platform based on online sampling to monitor chemical reaction processes · S7: Recent Trends in Compressive Raman Spectroscopy Using DMD-Based Binary Detection · S8: CCSDS 123.0-B-2: Low-Complexity Lossless and Near-Lossless Multispectral and Hyperspectral Image Compression
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized scores place this candidate between ranks 12 and 39, in ordering band B. That range is only a post-hoc reading aid using an affordability proxy; it is not an experimental endpoint, economic-value estimate, or change to partner-candidate status.
Band C — intermediate post-hoc review priority (ranks 26–45)¶
26. Review Materials Results by Model Mismatch¶
Canonical title: Residual-Governed Characterization Queue for Materials Discovery
In one sentence: A blinded partner study would test whether scientists can review model-relative characterization differences instead of every complete package while preserving protected findings and scientific decisions.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-12 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Chemistry Materials |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 10–36 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-11, EXP06-PARTNER-13, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A materials campaign may generate a large characterization package for every composition and processing condition. When scientists review all packages equally, routine confirmations consume the same attention as results that contradict predicted phases, structures, or properties. Fixed features and anomaly labels can shorten the queue, but they may erase unfamiliar peaks, minority phases, or small reliable contradictions. The actual field gap is fundamental: no evidence yet shows that a target laboratory's full-package queue is overloaded or delays important discrepancies.
What is proposed¶
For each candidate, a versioned campaign model would freeze predicted phase, structure, and property outputs before measurements are revealed. The system would compare those predictions with standardized results and preserve signed property errors, categorical disagreements, missing observations, unpredicted peaks or phases, uncertainty, and provenance. Reliable or consequential discrepancies would enter a primary queue with a named reviewer and possible actions such as replication, an orthogonal assay, or a model-scope review. Expected results would remain reconstructable and complete raw files available on demand. Random candidates, protected classes, and risk-selected cases would still receive full-package review. New chemical classes, incompatible versions, invalid measurements, safety flags, stale models, structured residuals, or excessive reconstruction error would suspend triage. Scientists—not the queue—would authorize experiments and model changes.
The cross-domain transfer¶
The residual-processing archetype becomes a scarce-attention system rather than a data link. A model predicts each candidate's standardized characterization; measured-minus-predicted differences become the review message and later learning signal. The mapping is strong for numerical properties but weaker for images, peaks, phases, and unfamiliar features, which cannot always be reconstructed additively and must remain accessible in full.
Why it advanced¶
This candidate cleared only the separately calibrated empirical-partner lane. Autonomous laboratories, active learning, human-guided phase mapping, and large provenance systems make its components plausible. It did not reach strict success because no paired field trial shows that residual presentations reduce total work without changing scientific actions, and the target queue constraint remains unmeasured.
Prior art and the remaining open claim¶
Adjacent systems already use Bayesian optimization, uncertainty-guided measurement selection, automated phase analysis, anomaly detection, and human model updates. The narrower open comparison is an audited review interface: frozen multimodal predictions plus numerical, categorical, missingness, and unfamiliar-feature residuals, with protected and sampled full-package review. On one material family, it claims at least 30% less total reviewer time, at least 95% action concordance, complete capture of protected challenges, no acceptance of missing or incompatible results, and no increase in follow-up demand.
Smallest decisive test¶
A partner laboratory would conduct a paired, blinded shadow study on 60–150 already authorized candidates from one family and fixed assay suite. It would compare complete packages, uncertainty-only ordering, fixed anomaly or feature thresholds, and the frozen residual presentation. Concealed challenges would include small property shifts, phase conflicts, unfamiliar peaks, bad assays, missing data, version mismatches, and out-of-scope chemistry. Reject operational adoption if workload falls less than 30% after audits and fallbacks, concordance is below 95%, any protected case is missed, an incompatible or missing package passes, or one discrepancy class is systematically lost.
Deployment and cost¶
A shadow trial needs laboratory leadership, characterization, model, safety, and data-steward authorization, plus review of trade-secret, export-control, retention, and publication rules. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. No vendor quote supports them.
Risks and uncertainties¶
- Standardization may remove an unfamiliar peak, morphology, or minority phase before the residual is calculated.
- A confident but misspecified campaign model may classify informative chemistry as routine.
- Underestimated assay uncertainty may give noisy results excessive influence over routing and model updates.
- Audits may miss rare discrepancies, especially when risk sampling reflects only anticipated failure modes.
- Routine packages removed from the main queue may contain context needed to recognize a later campaign-wide pattern.
Expert review¶
Useful reviewer backgrounds: Materials discovery scientist, X-ray diffraction or multimodal characterization specialist, Bayesian modeling and active-learning researcher, Scientific data-governance specialist, Experimental-design statistician.
- Is complete-package review currently exceeding a declared capacity or delaying model-relevant discrepancies in the proposed campaign?
- Which raw or standardized observations cannot be represented safely as model-relative residuals?
- What protected challenge set would cover unfamiliar peaks, minority phases, specimen errors, and out-of-scope chemistry?
- How should audit size and sampling be powered for rare but decision-changing discrepancies?
- Can independent reviewers reliably classify each mismatch as chemical evidence, measurement failure, specimen error, or model-scope failure?
Evidence and provenance¶
Selected sources: S1: Achieving AI-Driven Autonomous Laboratories · S2: Autonomous Methods · S3: Workshop Report on Autonomous Methodologies for Accelerating X-ray Measurements · S4: An autonomous laboratory for the accelerated synthesis of inorganic materials · S5: Author Correction: An autonomous laboratory for the accelerated synthesis of inorganic materials · S6: On-the-fly closed-loop materials discovery via Bayesian active learning · S7: Human-in-the-loop for Bayesian autonomous materials phase mapping · S8: NCAL: Data Management
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: Post-hoc harmonized profiles place this candidate from rank 10 to 36, in ordering band C. The spread is a reading-order aid, not a preregistered outcome or estimate of economic value, and it does not remove the need for field evidence.
27. A Common Contract for Synchronized Takes¶
Canonical title: Representation-Independent Synchronized-Take Contract for Dailies
In one sentence: A read-only pilot would test whether different picture-and-sound synchronizers can expose the same take membership, timing, gaps, and conflicts without clients depending on filenames or vendor internals.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-21 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Film Media Production |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 18–32 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-22, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Dailies systems often decide which picture and sound files form a take by using filenames, folder order, embedded timecode, recorder labels, sidecars, or one algorithm's behavior. A camera, recorder, metadata carrier, file split, or synchronization tool can therefore change grouping and frame-to-sample alignment even when the intended take is unchanged. Existing pipelines lack a shared behavioral test for deciding whether a replacement synchronizer truly preserves membership, timing, channel roles, gaps, conflicts, and ambiguity.
What is proposed¶
An opaque SynchronizedTake component would register picture and sound streams by semantic role, accept clock or slate observations, resolve alignment under an explicit policy, map stream positions into common take time, report shared coverage, validate completeness, freeze a resolved version, and create superseding versions. Its contract would require monotonic mapping within continuous segments, one public correspondence for every resolved position, explicit gaps and discontinuities, declared ambiguity and conflict errors, immutability after freezing, and no state change after rejection. Filenames, folders, metadata locations, waveform features, caches, databases, and matching algorithms would remain private. The same black-box fixtures, generated sequences, and representation-change tests would judge manual-anchor, timecode, waveform-assisted, and device-specific implementations. The component could not rewrite, delete, or silently relabel source media.
The cross-domain transfer¶
The interface-contract archetype maps cleanly to dailies: define a synchronized take by observable behavior and invariants, not by its file layout or synchronization algorithm. Each implementation translates private evidence into the same semantic streams and common-time mappings. Substitution depends on passing one oracle, although narrowly authorized diagnostic access may still be needed for physical clock provenance.
Why it advanced¶
This is an empirical-partner candidate, not a strict-success result. Existing synchronization products, timecode and audio standards, interchange schemas, and implementation-neutral timeline models support feasibility. It advanced because a copied-media, read-only comparison is bounded and reversible, while real failure prevalence, operator acceptance, independent implementations, and production authorization remain absent.
Prior art and the remaining open claim¶
Premiere and Resolve already synchronize multicamera material; FCPXML represents synchronized clips; BWF and SMPTE timecode standardize important carriers; and OpenTimelineIO separates timelines from proprietary files. The remaining claim is narrower than a new format: two independent resolvers would return equivalent semantic membership, covered intervals, mappings, gaps, conflicts, and lifecycle outcomes after renaming, relocation, registration reordering, clock-origin shifts, and lossless segmentation—and would outperform a metadata-schema-only approach without exposing private implementation details.
Smallest decisive test¶
With written production approval, duplicate one short, non-live take containing two picture streams, separate sound, aligned coverage, and a known gap or conflict. Freeze expected membership, mappings, errors, positional tolerance, and the no-write rule. Compare the incumbent result, a simple manual-anchor model, one adapter, and a schema-only baseline where possible. Rename and relocate files, reorder registration, shift the clock origin with compensation, and split a continuous stream losslessly. Reject the claim if independent resolvers diverge beyond tolerance, silently resolve conflicts, mutate after rejection, require private fields for ordinary work, or the schema baseline matches them with materially less effort.
Deployment and cost¶
A pilot requires post-production, sound, camera or data, dailies, editorial, rights, and security approval. It must use copied media and leave the active pipeline unchanged. Rough 2026 bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for launch, and \(50,000–\)250,000 annually; no measured effort or quote confirms them.
Risks and uncertainties¶
- The common-time model may omit drift, variable rates, discontinuities, or corrupted-clock behavior found in production media.
- Clock provenance or recorder-specific channel facts may be necessary production information rather than accidental representation.
- A shared test fixture may encode the same mistaken synchronization assumptions as the reference implementation.
- Opacity may make failures difficult to diagnose unless a narrowly governed read-only diagnostic surface is defined.
- Two mappings within a numerical tolerance may still differ in ways assistant editors or sound staff consider operationally unacceptable.
Expert review¶
Useful reviewer backgrounds: Dailies pipeline engineer, Production sound mixer or synchronization specialist, Assistant editor, Media interchange and timecode standards expert, Post-production security and rights specialist.
- Which device-specific facts must remain visible for authorized diagnosis, and which are accidental client dependencies?
- What frame or sample tolerance is acceptable for each downstream viewing, logging, and editorial task?
- How must the abstract model represent drift, discontinuities, variable frame rates, missing coverage, and multichannel sound?
- Can a canonical metadata baseline provide the same invariance and usability with less integration work?
- What production-derived fixtures would adequately represent actual synchronization failures without exposing protected media?
Evidence and provenance¶
Selected sources: S1: The 2030 Vision White Paper Section 3.3 · S2: 2030 Greenlight · S3: Create multi-camera source sequences in Premiere · S4: DaVinci Resolve – Edit · S5: Document Type Definition: Final Cut Pro XML Interchange Format 1.10 · S6: EBU Tech 3285 v2: Specification of the Broadcast Wave Format · S7: SMPTE ST 12-1:2014, Time and Control Code · S8: OpenTimelineIO 0.16.0 Architecture
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized review places this candidate between ranks 18 and 32, in ordering band C. This is only a reading-order aid based partly on an affordability proxy; it is neither a preregistered success nor evidence of production benefit.
28. Portable Lighting Cues With Behavioral Tests¶
Canonical title: Rig-Independent Cinematic Lighting Cue Contract
In one sentence: An isolated trial would test whether different lighting controllers and fixture allocations can reproduce an approved cue trajectory and safe lifecycle without exposing raw console addresses.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-22 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Film Media Production |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 23–30 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-21, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Film lighting cues are often encoded through one console's channels, fixture profiles, addresses, show-file structure, macros, and transition rules. Moving a scene, changing a console, or reallocating fixtures can alter fades, colors, holds, conflicts, or emergency transitions even when the intended look has not changed. File import and transport standards help move data or commands, but the production still lacks a shared test for whether another controller preserves both the photographed cue behavior and its output boundaries.
What is proposed¶
An opaque LightingCueProgram would expose semantic operations: validate a rig, arm without changing output, trigger, sample declared state, hold, resume, cancel, abort to a safe state, report status, freeze, and supersede. Its abstract model would name lighting roles, targets, timed trajectories, interpolation, concurrency and conflict rules, enrolled outputs, a safe state, and version history. Valid event sequences would have deterministic abstract trajectories; rejected or unarmed commands would leave output unchanged; abort would take precedence; and commands could reach only enrolled endpoints. Console syntax, channels, addresses, universes, profiles, show files, and drivers would remain private. A shared black-box suite would compare a reference player and console adapters after address renumbering, registration reordering, and capability-equivalent fixture reallocation, using visual, camera, photometric, timing, and safety tolerances fixed by accountable staff beforehand.
The cross-domain transfer¶
The representation-independent-contract archetype is instantiated as a semantic lighting state machine separated from console and rig details. Adapters translate the same cue operations into different physical implementations, and one conformance oracle governs substitution. The mapping is incomplete unless it captures photographed light: equal console values do not guarantee equal spectra, flicker, optics, placement, or camera response.
Why it advanced¶
This candidate entered only the empirical-partner lane. Show-file import, GDTF/MVR, DMX, ACN, and open control gateways establish adjacent technical pieces, while a simulator test is small and reversible. It lacks measured production prevalence, an independent implementation, approved visual tolerances, vendor or production commitment, and evidence that physical fixture differences will not dominate representation effects.
Prior art and the remaining open claim¶
Current prior art standardizes show-data import, device and scene descriptions, transport, and cross-protocol control. It does not establish a cinematographer-approved cue lifecycle and emitted-light substitution test. The remaining claim is that two independent adapters can hide addressing yet deliver equivalent semantic and photographed trajectories, identical errors and lifecycle behavior, no unenrolled output, and an approved abort transition after representation changes. Existing interchange matching those results at equal or lower effort would defeat the incremental claim.
Smallest decisive test¶
With DP, gaffer, electrical, and safety approval, test one disconnected cue containing arming, a timed fade, concurrent endpoints, hold and resume, a conflict, an unarmed rejection, and abort-to-safe. Freeze expected trajectories, error codes, enrolled outputs, camera and photometric tolerances, timing limits, and abort deadline. Compare a reference player, one console adapter, and the incumbent import-and-rebuild workflow, then repeat after address, registration, and fixture-allocation changes. Reject the claim for any unenrolled command, rejected-command output, lifecycle disagreement, tolerance failure, required console-native escape hatch, cheaper equivalent incumbent performance, or materially different photographed result despite nominal conformance.
Deployment and cost¶
The first test must remain on a simulator or isolated prelight bay with no performers, shooting rig, hazardous effects, machinery, or production automation attached. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for launch, and \(50,000–\)250,000 annually. They are not quotes.
Risks and uncertainties¶
- Normalized semantic values may produce different perceived and photographed light across fixture technologies.
- The contract may omit spectral output, flicker, optics, spatial placement, latency, or resource contention important to the scene.
- Console telemetry may conform even while the emitted light falls outside the approved camera or photometric tolerance.
- A software-defined safe state may conflict with stage electrical procedure or fail over an unreliable lighting transport.
- The contract may accidentally preserve the incumbent console's interpolation or tie-breaking behavior instead of the cinematographer's intent.
Expert review¶
Useful reviewer backgrounds: Director of photography or imaging scientist, Gaffer and lighting console programmer, Stage electrical and entertainment-control safety specialist, Fixture calibration and spectral-measurement specialist, Lighting-control standards and adapter engineer.
- Which semantic endpoint properties and photographed-light measurements define equivalence for the selected cue?
- What tolerances for spectrum, exposure, color, flicker, trajectory timing, and abort completion must be frozen before testing?
- Can the chosen GDTF descriptions represent the relevant fixtures accurately, including unsupported or vendor-specific features?
- Which abort behavior is safe under actual electrical procedure, and what can be guaranteed over a transport that may lose packets?
- Does the incumbent import-and-manual-rebuild workflow already meet the same criteria with equal or lower effort?
Evidence and provenance¶
Selected sources: S1: Importing Show Files · S2: GDTF FAQ · S3: GDTF & MVR Help Pages · S4: ANSI E1.11-2024: USITT DMX512-A Asynchronous Serial Digital Data Transmission Standard for Controlling Lighting Equipment and Accessories · S5: ANSI E1.17-2015 (R2025): Entertainment Technology—Architecture for Control Networks (ACN) · S6: Open Lighting Architecture Developer Documentation · S7: General Device Type Format (GDTF) Fixture Import for Eos Family · S8: Color Reproduction in LED Wall Virtual Production Stages
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: Post-hoc harmonized profiles place this candidate from rank 23 to 30, in ordering band C. The ranking is a reading-order aid using an affordability proxy, not a preregistered endpoint, deployment authorization, novelty determination, or measure of economic value.
29. Common-Wafer Contest for One Fabrication Slot¶
Canonical title: Common-Wafer Trial for Awarding a Scarce Nanofabrication Integration Slot
In one sentence: A nanofabrication facility would compare teams on blinded, equally resourced test wafers before awarding its single process-integration slot.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-03 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Bounded Rivalry Governance × Nanotechnology |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 20–29 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-07, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A shared nanofabrication facility has one integration bay and limited technician time. Teams now compete largely through proposals and results from samples they chose themselves. That can reward unusually favorable devices, extensive private characterization, or incomplete reporting of failed runs. It can also leave the facility and other users bearing contamination, waste, cleanup, and downtime. The facility may therefore select a persuasive process that cannot reproduce safely and reliably on shared equipment.
What is proposed¶
Replace proposal-only selection with a rule-bound common-wafer trial. Before entrants are known, publish eligibility, identical substrates, allowed process steps, equal tool and measurement budgets, safety gates, scoring, tie-breaks, confidentiality, penalties, and appeals. Code the samples and score complete attempt histories, functional yield across wafers, dimensional and electrical reproducibility, equipment compatibility, waste, cleanup, and recovery time. Unsafe performance cannot be offset by technical strength. Independently remeasure the leaders and audit their resource use. Give two teams bounded validation runs before awarding the single integration bay. Afterwards, compare trial scores with actual integration performance, examine whether the winner gained control over shared interfaces or rules, and reopen competition through a scheduled challenger window.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a deliberately governed contest for a genuinely scarce facility slot. Common specimens and resource caps define the arena; blinded scoring, audits, safety floors, spillover accounting, penalties, appeals, post-cycle review, and later challenger access constrain how teams may compete and what winning confers.
Why it advanced¶
It passed Experiment 6's strict researched-candidate bar because the allocation problem, responsible authorities, testing capabilities, safeguards, comparator, and falsifiable shadow study were sufficiently specified. That status concerns the researched candidate only; it does not establish field performance, novelty, deployment authority, adopter demand, or economic benefit.
Prior art and the remaining open claim¶
Proposal review, safety qualification, common-specimen comparisons, metering, contamination controls, equipment standards, and post-selection oversight already exist, so the proposal sits next to substantial prior art. The narrower open claim is that the full score—held-out reproducibility, all attempts, equal counted resources, tool compatibility, and recovery burden—will predict audited performance and avoid ranking reversals better than both proposal-only review and a simpler common-wafer technical score.
Smallest decisive test¶
With facility, safety, and data approval, run a non-awarding shadow study lasting at most 12 weeks and costing at most $50,000. Require at least eight archived entrants from two process families, comparable retained wafers or replicate measurements, and usable operational records. Freeze the rubric and audit split before revealing identities or existing ranks. Advance only if the full score improves held-out Kendall rank correlation over proposal review by at least 0.15, reduces audit reversals by at least 20%, is no worse than technical-only scoring, avoids a leader-changing process-family interaction, keeps essential missingness below 20%, and causes no safety or confidentiality breach.
Deployment and cost¶
The first shadow evidence step is estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms, not a vendor quote. Initial deployment is \(50,000–\)250,000; operational launch and annual operation are each \(250,000–\)1 million. Real use would also require local authority over recipes, records, appeals, reserves, and penalties.
Risks and uncertainties¶
- Common wafers may favor one process family or poorly represent integration conditions.
- A frozen score may encourage teams to optimize measured proxies instead of robust integration performance.
- Equal in-contest resource caps may still favor teams with stronger infrastructure outside the counted arena.
- Recipe submission and access-log audits may expose confidential know-how or identifiable personnel data.
- A remediation reserve could exclude less-capitalized teams or exceed the facility's legal authority to collect it. Managers have not established the needed terms yet, and the effects on those teams require field data. No source establishes how complete the facility's records are or how frequently proposal winners fail shared-condition replication. No facility has committed to host the study, and site-specific technician capacity and opportunity costs remain unknown.
Expert review¶
Useful reviewer backgrounds: Nanofabrication process-integration engineer, Facility operations and metrology manager, Environmental health and contamination-control specialist, Research-allocation and data-governance counsel.
- Can retained wafers from at least eight entrants be compared without introducing process-family-specific measurement bias?
- Which waste, cleanup, downtime, and compatibility measures can be reconstructed reliably from existing records?
- Would the proposed score predict successful integration better than technical yield and reproducibility alone?
- What authority does the facility have to inspect recipes, hear appeals, impose penalties, or require a remediation reserve?
Evidence and provenance¶
Selected sources: S1: User Proposal Review and Evaluation Process · S2: Safety in the NanoFab · S3: Safety and Policies · S4: Nanosensor Manufacturing Workshop: Finding Better Paths to Products · S5: Benchmarking the ACEnano Toolbox for Characterisation of Nanoparticle Size and Concentration by Interlaboratory Comparisons · S6: Fees · S7: FAQ—Individual Standards · S8: Report on the Nanoscience Research Centers
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized score is only a post-hoc reading-order aid: band C, ranking between 20 and 29 across profiles. Its pilot-speed input reflects cost-band affordability, not independently measured elapsed time or economic value, and it does not alter strict-success status.
30. A Stable Contract for Crystal Structures¶
Canonical title: Representation-Independent Contract for Ordered Periodic Material Structures
In one sentence: An opaque software contract would let materials tools exchange the same ordered periodic structure without depending on atom order, file layout, or a particular canonicalization method.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-08 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Representation Independent Interface Contract × Chemistry Materials |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 26–30 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
Materials software passes crystal structures among parsers, databases, simulation tools, and analysis programs. Clients often depend on incidental details such as atom-array order, coordinate wrapping, lattice orientation, cell convention, or serialized field layout. The same physical ordered structure can then receive different identifiers or downstream treatment, while an internal parser or storage change can break clients even when the intended material state has not changed. It also becomes difficult to distinguish a physical change from an encoding change.
What is proposed¶
Define an immutable ordered-periodic-material value by its periodic decorated points in physical space, including species, occupancies, units, tolerance policy, and provenance. Expose validated construction, composition, physical-equivalence tests, invariant summaries, explicit conversions, and serialization, but hide arrays, atom order, coordinate basis, caches, canonical labels, and format-specific fields. Specify preconditions, typed errors, no-mutation-on-failure behavior, and laws for atom permutation, origin shifts, periodic wrapping, unit conversion, and admissible cell changes. Every parser or store must pass the same public-only black-box and metamorphic tests before substitution. Version changes to promised behavior and probe error text, ordering, serialization, debug access, and timing for leaks. Version zero rejects disorder, partial occupancy, trajectories, surfaces, and inferred bonding rather than silently approximating them.
The cross-domain transfer¶
The representation-independent-interface archetype maps strongly here. The abstract component is the physical ordered periodic structure; concrete arrays, graphs, files, and database rows remain hidden. Behavioral laws, typed failures, shared conformance tests, leakage review, and versioning determine whether independently built implementations may substitute for one another.
Why it advanced¶
It passed Experiment 6's strict researched-candidate bar because close comparators, a bounded scope, scientific authority, reversible sandbox, measurable failure conditions, and a live two-adapter benchmark were identified. This endpoint does not show that production repositories suffer widespread coupling or that the contract is novel, performant, deployable, or adopted.
Prior art and the remaining open claim¶
Representation-insensitive structure matching, pymatgen, spglib, CIF, OPTIMADE, and unified materials-data systems already address important parts of the problem. The remaining claim is narrower: two independently structured adapters can obey one public behavioral oracle for equivalence, invariants, errors, immutability, and provenance while leaking fewer internal details and producing fewer semantic disagreements than the existing array baseline, a canonicalization path, or a schema-only round trip.
Smallest decisive test¶
With a platform partner and scientific-method lead, preregister the abstract state, tolerance rules, observables, and exclusions. Test two independently structured adapters on 12 copied fixtures, five seeded representation changes per fixture, fixed assertions, and 200 operation sequences. Compare them with the existing array/serializer behavior, spglib plus pymatgen, and an OPTIMADE/CIF round trip. Reject version zero if any meaning-preserving transformation changes a promised observable, any fixture triple exposes tolerance-driven non-transitivity, provenance is lost, hidden representation must be exposed, or the contract fails to reduce semantic divergences and leakage relative to the best comparator. Also record runtime and memory.
Deployment and cost¶
The read-only sandbox is estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms. Initial deployment is \(50,000–\)250,000; operational launch is \(250,000–\)1 million; annual operation is \(50,000–\)250,000. Production identifier changes, deduplication, parser replacement, or database migration are expressly outside the first test.
Risks and uncertainties¶
- Tolerance-based equivalence may be non-transitive, producing unstable identity groups.
- The initial abstract state may omit meaningful distinctions such as defects, chirality, magnetic order, isotope labels, or provenance.
- An over-specified suite could freeze incidental numerical or serialization behavior.
- An under-specified suite could pass simple structures while missing difficult-cell disagreements.
- Opaque access may obstruct legitimate diagnostics and encourage unsupported escape hatches. Real repository data have not yet shown how frequent or costly representation coupling is. Scientific reviewers have not approved the proposed version-zero semantics, and comparator performance remains unmeasured. Passing finite tests would not establish chemical identity, scientific equivalence, implementation correctness, or acceptable production performance.
Expert review¶
Useful reviewer backgrounds: Computational crystallographer, Materials-data platform architect, Scientific software testing specialist, Materials provenance and standards expert.
- Does the proposed abstract state preserve every scientifically relevant distinction in the selected ordered-periodic scope?
- Can the tolerance policy avoid non-transitive equivalence for realistic near-boundary structures?
- Which current clients depend on atom order, serializer layout, canonical labels, or other hidden details?
- Does the contract reduce semantic divergence and leakage without unacceptable runtime, memory, or diagnostic costs?
Evidence and provenance¶
Selected sources: S1: Identifying duplicate crystal structures: XtalComp, an open-source solution · S2: pymatgen.core package: structure_matcher module · S3: Spglib conventions of standardized unit cell · S4: OPTIMADE API specification v1.3.0 · S5: OPTIMADE, an API for exchanging materials data · S6: Core CIF dictionary · S7: NOMAD — Materials science data, managed and shared · S8: Employer Costs for Employee Compensation — March 2026
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized result is a post-hoc ordering aid: band C, ranks 26–30 across profiles. It is not an experimental endpoint or economic-value measure; its pilot-speed input is an affordability proxy, and the candidate remains a strict success only under Experiment 6's researched bar.
31. Show Reviewers Only Unexpected Proof Effects¶
Canonical title: Residual Impact Maps for Mathematical Theory Revisions
In one sentence: A complete proof-library check would be reconstructed from frozen theorem-level predictions and typed discrepancies, allowing reviewers to focus on surprises while protected changes always remain visible.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-17 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Mathematics |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 66.0/100 · rank range 27–34 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
Revising an axiom, definition, notation rule, or trusted dependency can affect hundreds of machine-checked theorems. A checker can revalidate the whole library, but maintainers may still face a large migration report. Dependency graphs can flag harmless reachability and miss effects from elaboration, automation, or undeclared coupling. Reading every expected result wastes attention, yet showing only predicted failures could hide an unexpected survival, a new dependency, a removed obligation, or a theorem that still passes for the wrong reason.
What is proposed¶
Before migration, freeze a versioned model predicting every theorem's status, changed obligations, and dependency differences. Independently run the trusted checker over the complete authorized corpus. Compare each prediction with the actual result and record typed residuals for unexpected failures or survivals, changed diagnostics, obligations, dependencies, trust assumptions, timeouts, missing results, or confidence disagreements. Reconstruct the full ledger from prediction plus residual, while directing review primarily to consequential mismatches. Statements, axioms, admitted facts, trust changes, removed obligations, checker failures, missing results, and out-of-scope theorems always appear in full. Random predicted-unaffected records and boundary cases receive independent audits. Version, coverage, reconstruction, drift, or audit failures automatically restore the complete theorem-by-theorem report for the affected component.
The cross-domain transfer¶
Predictive residual processing becomes a review codec for mathematical migrations. A frozen impact model supplies the expected theorem ledger; complete checking supplies reality; typed differences carry surprises. Checksums, audits, resynchronization, protected-event bypasses, drift monitoring, and raw-report fallback keep compression from becoming selective proof checking.
Why it advanced¶
It did not enter the strict-success lane. It cleared a separately calibrated empirical-partner-candidate lane because a reversible archived replay and decision thresholds are specified, but the decisive field evidence is missing: no measured report burden, predictable-record share, predictor calibration, reviewer study, audit rate, cost benchmark, maintainer funding, or partner authorization exists.
Prior art and the remaining open claim¶
Static dependency analysis, grouped checker reports, incremental proof checking, iCoq-style regression selection, source diffs, and trust checklists are established neighbors. The open comparison is whether complete independent checking plus frozen predictions, typed residuals, protected full records, random audits, exact reconstruction, and component fallback can reduce report volume and review time without lowering protected-event recall or blinded classification accuracy relative to full and dependency-grouped reports.
Smallest decisive test¶
With maintainer authorization, replay an archived, non-release-blocking migration containing at least 500 declarations. Freeze the model, residual types, protected classes, thresholds, and audit sample before opening the evaluation partition. Randomize blinded reviewers among full reports, dependency-grouped reports, and the residual interface; insert cases covering failures, survivals, dependency and obligation changes, axioms, timeouts, missing outputs, manifest gaps, and version mismatches. Require 100% protected-event recall, zero missing-as-success errors, exact ledger reconstruction, no consequential audit miss, classification no more than five percentage points below the best comparator, and at least 20% lower median review time or report volume. Success permits only another shadow study.
Deployment and cost¶
The archived replay is estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms. Initial deployment is \(50,000–\)250,000; operational launch is \(250,000–\)1 million; annual operation is \(50,000–\)250,000. The model may prioritize review but may never approve revisions, waive checker results, or edit proofs automatically.
Risks and uncertainties¶
- A shared predictor could create correlated blind spots across whole theory components.
- Unexpected theorem survival may conceal a weakened statement or unintended dependency.
- Automation and elaboration may create semantic coupling absent from declared dependency graphs.
- The residual schema may preserve checker status but omit mathematical context needed for judgment.
- Reviewers may anchor on predictions, while thresholds may be tuned to shrink the queue. No evidence yet quantifies present reviewer burden, repeated predictable content, or protected-event frequency. The theorem-level predictor has not been calibrated across the required outcome types, and no human comparison has tested whether residual presentation preserves decisions. Audit rates, repository confidentiality rules, operating costs, partner authorization, and maintainer funding remain unresolved.
Expert review¶
Useful reviewer backgrounds: Formal-mathematics library maintainer, Proof-assistant kernel and elaboration expert, Human-factors researcher for technical review, Statistical audit and anomaly-detection specialist.
- What fraction of a real foundational migration report is predictable repetition, and how much reviewer time does it consume?
- Can the predictor detect unexpected survivals, trust changes, removed obligations, timeouts, and missing outputs with calibrated uncertainty?
- What random and boundary-focused audit rate would detect rare consequential blind spots?
- Does the residual interface preserve blinded reviewer accuracy while reducing total review, model-maintenance, audit, and fallback cost?
Evidence and provenance¶
Selected sources: S1: Maintaining a Library of Formal Mathematics · S2: mathlib4: The Math Library of Lean 4 · S3: RFC: Add left actions and right actions to expression tree elaborator and make ^ be a right action · S4: Pull Request Review Guide · S5: iCoq: Regression Proof Selection for Large-Scale Verification Projects · S6: Practical Machine-Checked Formalization of Change Impact Analysis · S7: The Isabelle System Manual (Isabelle2022) · S8: Axioms — The Lean Language Reference
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized score is a post-hoc reading-order aid only: band C, ranks 27–34 across profiles. The pilot-speed input approximates affordability rather than measured duration. This ordering does not convert the candidate into a strict success or indicate economic value.
32. A Recurring Reckoning for Airport Noise Promises¶
Canonical title: Shared-Sky Accountability Observance for Airport Noise Commitments
In one sentence: A short, consent-based observance would connect remembered airport-noise experiences and past commitments to an authorized public decision and a concrete follow-up ledger.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-13 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Aviation Aeronautics |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 66.0/100 · rank range 23–38 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-12, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Airports can collect noise measurements, complaints, and public comments while losing a shared memory of what communities experienced and what institutions promised. Staff and resident turnover may scatter commitments across minutes and dashboards. Conventional meetings can repeatedly ask people to recount distress without creating a bounded occasion to acknowledge it, renew or revise a promise, record disagreement, and assign follow-through. Participants may then disagree about a commitment's origin, scope, owner, authority, resources, or next review date.
What is proposed¶
Embed a quarterly 30-minute observance in an existing airport-community forum. A rotating community-institution pair first documents its purpose, participant standing, story and recording provenance, access options, obligations, and retirement rule. A restrained threshold introduces 60 seconds of optional listening to consent-cleared low-intensity audio or viewing an acoustic trace; silence, private reflection, remote attendance, or leaving are equivalent choices. Paired community and institutional accounts then reconstruct one commitment, preserving corrections and disagreement. Consenting witnesses acknowledge the experience and exact promise. An authorized representative must renew, revise, or decline it; community members may dissent or remain silent. Closure updates a public ledger with authority limits, owner, next action, resource dependency, forum, review date, and unresolved disagreement, followed by a harm-and-meaning debrief and periodic independent audit.
The cross-domain transfer¶
The ritualized-commitment archetype becomes a marked, recurring accountability occasion rather than an operational aviation decision. Threshold, optional reflection, paired histories, witnessing, explicit institutional recommitment, ledger closure, debrief, rotating stewardship, harm audit, and governed retirement connect symbolic recognition to ordinary authorized follow-through.
Why it advanced¶
It passed Experiment 6's strict researched-candidate bar because the intervention, authority boundary, conventional comparator, consent protections, stopping rules, and matched rehearsal are explicit and testable. This does not show noise reduction, community-wide legitimacy, real commitment fulfillment, adopter willingness, novelty, deployment authorization, or economic impact.
Prior art and the remaining open claim¶
Noise dashboards, complaint systems, community roundtables, facilitated planning, written agreements, action ledgers, and one-time listening sessions already exist. The remaining claim is incremental: adding the governed threshold, optional sensory interval, paired provenance accounts, witnessing, an explicit renew-revise-decline choice, and immediate debrief will improve accurate, retained understanding of one commitment beyond an equally timed facilitated ledger review, without increasing coercion, distress, false representation, or confusion that acknowledgment equals mitigation.
Smallest decisive test¶
With a willing forum, run a preregistered tabletop and randomized matched rehearsal with 12–20 consenting community and institutional participants, using a fictional commitment and synthetic low-intensity media. Compare the full observance with an equally timed agenda-and-ledger review containing identical facts. Score eight commitment facts immediately and after 7–14 days, plus authority understanding, fairness, comfort, manipulation, representation, disclosure pressure, accessibility, and freedom to opt out. Proceed only if median recall improves by at least two facts, at least 80% understand what was not decided, coercion and representation scores are no worse, and no serious distress or retaliation concern occurs. The observer may require revision or stop.
Deployment and cost¶
First evidence, initial deployment, operational launch, and annual recurring operation are each estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms, not external quotes. The observance cannot change routes, schedules, funding, regulation, environmental findings, or airline operations, and it cannot replace complaints, consultation, mediation, or technical analysis.
Risks and uncertainties¶
- The observance could aestheticize residents' distress or turn it into institutional performance.
- Audio or repeated recollection could cause sensory or emotional harm despite opt-out choices.
- Selected recordings, histories, or visible participants could be treated as representing absent communities.
- Ceremonial acknowledgment might be mistaken for mitigation, legal acceptance, or resolution.
- An authorized-looking renewal could conceal missing funding, regulatory power, or operational feasibility. No site has shown that commitment-memory failures occur at the assumed rate, and no airport or community group has agreed to host the pilot. Live evidence is absent on nonretaliatory refusal, adverse events, comparative benefit over a disciplined ledger meeting, and whether real representatives can decide without bypassing other authorities. Site-specific costs are also unknown.
Expert review¶
Useful reviewer backgrounds: Airport community-engagement and noise-management lead, Affected-community representative with independent standing, Accessibility, trauma-informed facilitation, and emotional-safety specialist, Aviation governance and authority-boundary expert.
- Do participants currently fail to reconstruct commitment history, scope, authority, ownership, dependencies, and unresolved disagreement from ordinary records?
- Does the ritual sequence improve delayed factual recall beyond an equally timed facilitated ledger review?
- Can audio-free participation, silence, dissent, and exit remain genuinely nonretaliatory under the forum's power relationships?
- Which representative can renew, revise, or decline a real commitment, and which decisions must remain with airports, airlines, regulators, boards, or air-traffic authorities?
Evidence and provenance¶
Selected sources: S1: Neighborhood Environmental Survey · S2: Aircraft Noise: Military Helicopter Operators Should Improve Outreach to Affected Communities in the D.C. Area (GAO-26-107758) · S3: Advisory Circular 150/5050-4A: Community Involvement in Airport Planning · S4: What is a community roundtable? · S5: SFO Airport/Community Roundtable · S6: Community Engagement on Aircraft Noise · S7: CAP3041: Guidance for Airport Engagement and Complaints Handling Around Environmental Sustainability · S8: Being a Fair Neighbor—Development and Validation of the Aircraft Noise-Related Fairness Inventory (fAIR-In)
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized result is a post-hoc reading-order aid: band C, ranks 23–38 across profiles. It neither measures economic value nor changes the strict-success endpoint; the pilot-speed input is an affordability proxy, and elapsed pilot time was not independently scored.
33. Retiring Outdated Criminal-Record Labels¶
Canonical title: Label Sunset: Criminal-Record Authority Retirement Cycle
In one sentence: A voluntary quarterly exercise would help record custodians distinguish an ended authorization from deletion, preserve lawful history, and assign unresolved downstream corrections to named owners.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-14 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Criminology Forensic |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 66.0/100 · rank range 31–33 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A court or other competent authority may restrict or end the permitted use of a criminal-justice label, yet copies, vendor feeds, access permissions, derived flags, and staff assumptions can persist across disconnected systems. Correcting the source record does not necessarily tell every downstream custodian what changed. The result can be continued reliance on a superseded status; however, careless cleanup can also destroy records that must lawfully remain available for provenance, oversight, or defined exceptions.
What is proposed¶
Authorized custodians would hold a quarterly, voluntary “Label Sunset” cycle after counsel confirms which disposition categories are eligible. Using only invented records, participants would move a fictional label through a map of source, repository, vendor, and user systems. A records witness would verify a mock retirement instruction, remove the label from the active layer, and place it in an archive sleeve marked “retired—provenance preserved.” Participants could decline any part of the exercise. The group would recognize only verified batch-level reconciliation and route every unresolved exception category into protected workflows with an owner, resources, and deadline. An immediate debrief and independent audit could revise, pause, or end the cycle. Unlike an ordinary reconciliation briefing, the proposal adds a witnessed enactment, recommitment, recognition, and explicit exception handoff; it changes no legal status or production record itself.
The cross-domain transfer¶
The ritualized meaning-and-commitment archetype becomes a governed enactment of how a label spreads and loses operational authority. Its marked object, witnesses, voluntary commitment, closure, debrief, and retirement path map clearly to the domain, while real legal powers, system privileges, and individual remedies remain entirely outside the ritual.
Why it advanced¶
This was a STRICT_SUCCESS because it passed Experiment 6’s strict researched-candidate bar. That endpoint reflects the quality and testability of the researched proposal, not field validation, novelty, legal authorization, adoption, or demonstrated effects on record accuracy, employment, stigma, recidivism, or other real-world outcomes.
Prior art and the remaining open claim¶
Legal relief workflows, automated access and retention rules, reconciliation audits, privacy training, and record-clearing programs already perform much of the practical work. The narrower open claim is that adding a voluntary fictional propagation exercise, witnessed retirement, recommitment, aggregate recognition, exception routing, and debrief to a matched reconciliation briefing will improve accurate legal distinctions, delayed recall, and assignment of downstream exceptions without increasing coercion, stigma, privacy risk, or beliefs that the ceremony itself completed legal relief. There is no field evidence for that comparison.
Smallest decisive test¶
A court or repository unit would recruit 24–40 volunteers from at least four custodial functions and randomize intact teams to the 45-minute cycle or a content-, facilitator-, and time-matched briefing. Blind scoring would cover six invented scenarios before, immediately after, and two weeks later. Advance only if the cycle gains at least 0.4 standard deviations on the delayed composite or 15 percentage points in fully correct exception routing, with no material safety loss. Retire the mechanism if neither threshold is met, real records enter the exercise, more than 10% feel dissent is unsafe, or false-relief beliefs exceed the limit.
Deployment and cost¶
The authorized first step is a synthetic pilot with invented cases and no production queries or changes. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup and operational launch, and \(10,000–\)50,000 annually. These are assessment bands, not vendor quotes; no institution has committed funding or adoption.
Risks and uncertainties¶
- Participants may mistake symbolic retirement for sealing, deletion, expungement, or complete downstream correction.
- Small aggregate exception categories could expose protected case information.
- Staff may experience participation, affirmation, or dissent as an employment loyalty test.
- Recognition could create ceremonial closure while vendor copies, derived flags, or informal assumptions persist.
- The archive metaphor could encourage destruction or concealment of records that must lawfully remain preserved or accessible under an exception.
Expert review¶
Useful reviewer backgrounds: Criminal-records counsel, Court or repository data-governance lead, Privacy and information-governance specialist, Background-screening systems expert, Lived-experience or record-subject advocate.
- Can the six scenarios reliably distinguish correction, sealing, restricted use, deletion, retention, and lawful exceptions?
- Which custodian has authority to own each downstream exception without exposing protected case data?
- Would staff reasonably perceive passing, dissenting, or challenging the script as safe and nonretaliatory?
- Can the matched briefing contain exactly the same legal and system information so the enactment is the only material difference?
- What evidence would show that the ceremony is hiding incomplete vendor or repository reconciliation?
Evidence and provenance¶
Selected sources: S1: Criminal History Records: Additional Actions Could Enhance the Completeness of Records Used for Employment-Related Background Checks · S2: Fair Credit Reporting; Background Screening · S3: Technical and Operational Challenges of Implementing Clean Slate: Research Findings · S4: Making the Promise of Expungement a Reality: A Guide to Record Relief in the State Courts · S5: NIST Privacy Framework: A Tool for Improving Privacy through Enterprise Risk Management, Version 1.0 · S6: New Jersey P.L. 2019, c.269: Automated Clean Slate Process · S8: Occupational Employment and Wages — May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed it 31st–33rd, band C, across post-hoc profiles. That score only sets reading order and uses an affordability proxy; it is neither an experimental endpoint nor a measure of economic value, deployment readiness, or impact.
34. Stress-Testing Flight-Control Downselects¶
Canonical title: Envelope-Robustness Downselect Arena for Flight-Control Candidates
In one sentence: A shadow contest would compare whether hidden scenarios, safety gates, resource caps, and two-candidate advancement select flight controllers more consistently than public benchmarks or expert panels.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-04 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Aviation Aeronautics |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 65.0/100 · rank range 22–40 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
Aircraft programs must decide which competing flight-control candidates receive scarce hardware-in-the-loop and flight-test resources. If teams know every scored simulation case, they may tune to those cases, simulator artifacts, or inferred test structure, while escalating compute and engineering effort. A visible-score winner may therefore be less reliable under new conditions than its rank suggests. The actual frequency of this problem in flight-control programs is unknown, and general competition research supplies meaningful counterevidence against assuming severe leaderboard overfitting is routine.
What is proposed¶
The program would freeze eligibility, configuration deadlines, allowed resources, safety floors, scoring, audit rules, fouls, appeals, and judge authority before submissions. Teams could develop against a disclosed suite, but an independent custodian would rank frozen builds using undisclosed seeds and disturbance combinations. Every controller would first face noncompensable safety gates. Passing candidates would then be scored for hidden-condition robustness, simulated pilot-intervention demand, control activity, reproducibility, and declared integration burden under resource caps. Two materially different candidates, rather than one leaderboard winner, would receive hypothetical advancement in the first shadow study. Independent reproduction and later hardware-in-the-loop, pilot, and maintainer evidence would govern any real progression. The comparator includes a fully public single-winner benchmark and an unranked expert-panel choice using the same builds, scenario families, safety constraints, and nominal compute limits. No contest score authorizes flight.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a controlled engineering contest with a scarce prize, fixed legal moves, safety floors, resource limits, audits, appeals, multiple winners, and a later challenger window. The mapping is strong, but most individual governance features already appear in aviation challenges, technical competitions, or staged verification programs.
Why it advanced¶
This is an EMPIRICAL_PARTNER_CANDIDATE, not a strict success. It cleared a separately calibrated lane for a bounded external data-partner study because the comparison is testable, but it lacks proprietary frozen controllers, program-specific scenarios, validated resource accounting, downstream hardware evidence, and any partner commitment.
Prior art and the remaining open claim¶
Robust flight-control challenges, protected leaderboards, aviation contest rules, safety stop authority, formal protests, and staged simulation-to-hardware validation already exist. The remaining claim concerns their combination: holding builds and approved conditions constant, a hidden, safety-gated, resource-capped arena that advances two complementary candidates will produce more stable selections on independently seeded validation and later authorized hardware or human evaluation than either a public-suite winner or an expert-panel choice. That comparative effect has not been measured, and portfolio discretion could simply advance a favored lower scorer.
Smallest decisive test¶
With an aircraft-program partner, preregister an offline shadow study using three to six frozen controllers. Apply the proposed arena, the disclosed single-winner benchmark, and an unranked panel to identical approved materials, then test their selections on a second inaccessible scenario set. Measure selection agreement, rank stability, safety failures, reproducibility, intervention demand, actuator activity, resources, disputes, and sensitivity to lawful seeds and weights. Do not seek hardware authority unless the arena is more stable or predictive. Falsify it if innocuous seeds reverse choices, results do not beat both comparators, portfolio selection adds an inferior redundant candidate, or resource accounting systematically favors incumbents. Void results after leakage, unverifiable builds, compensable safety violations, or disputed simulator validity.
Deployment and cost¶
The first step is offline shadow evaluation with no contract, integration, certification, personnel, or flight consequences. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. They are not facility or supplier quotes.
Risks and uncertainties¶
- Secret scenarios may contain the same modeling errors as the disclosed simulator while making those errors harder to challenge.
- Existing tools and pretrained models could let incumbents exceed the practical resource cap without recorded spending.
- A discretionary definition of “complementary” could justify advancing a favored lower-scoring controller.
- Teams may tune to the inferred scenario generator rather than to operationally relevant robustness.
- Simulation rank may fail to predict hardware behavior, pilot workload, maintainability, integration effort, or certification findings.
Expert review¶
Useful reviewer backgrounds: Flight-control engineer, Aircraft verification and validation lead, Test pilot or pilot-in-the-loop specialist, Airworthiness and certification specialist, Simulation-model validation expert.
- Which approved scenario variations are independent enough to test robustness without leaving the simulator’s validity envelope?
- Can legacy tools, reusable models, supplier labor, compute, and sponsor support be counted fairly across entrants?
- How should complementarity be defined before scores are known so it cannot become discretionary favoritism?
- What rank-stability or downstream-prediction improvement would justify the added contest administration?
- Which results may inform a later hardware study without being mistaken for certification evidence or flight authorization?
Evidence and provenance¶
Selected sources: S1: AC 25.1309-1B—System Design and Analysis · S2: AC 25.1329-1B—Approval of Flight Guidance Systems · S3: NASA Systems Engineering Handbook, Section 5.0—Product Realization · S5: Robust Flight Control Design Challenge: Problem Formulation and Manual—The Research Civil Aircraft Model · S6: The Ladder: A Reliable Leaderboard for Machine Learning Competitions · S7: A Meta-Analysis of Overfitting in Machine Learning · S8: DARPA Lift Challenge—Competitors, Rules, Scoring, Safety, and Protest Procedure
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized profiles rank this candidate from 22nd to 40th, band C. The wide range reflects different review weights. It is a reading-order aid using a cost proxy, not an experimental endpoint, economic-value estimate, or partner commitment.
35. Tracking Aging Wetland Erosion Mats¶
Canonical title: Lifecycle-Governed Retirement of Wetland Erosion-Control Mats
In one sentence: A site register would identify successive erosion-control layers, trigger review as they age, block unsafe removal, and preserve a record after authorized retrieval.
| Field | Record |
|---|---|
| Portfolio ID | EXP05-STRICT-03 |
| Experiment and endpoint | Experiment 5 · Strict success |
| Archetype × domain | Layer Decay And Expiration Management × Environmental Climate |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 64.0/100 · rank range 35–37 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
Wetland and streambank repairs can leave several generations of blankets, coir mats, synthetic netting, and anchors in the same location. Older layers may be exposed, buried, torn, rooted through, or missing from project records. Crews may cover an unidentified layer again, leaving fragmenting mesh in habitat, or remove material that still holds vegetation and sediment. The field prevalence of such sequential legacy layers is not yet established for a candidate site, and hydrology or geomorphology may matter more than installed material.
What is proposed¶
At one bounded restoration reach, managers would give each mapped installation a stable identity, footprint, material class, deposition order, estimated age, condition, access history, and lifecycle state. A class-specific time-to-review trigger would mark a layer for inspection, never automatic removal. An age-weighted value-and-risk score would order limited field work, but qualified reviewers would first check whether roots, sediment, adjacent structures, monitoring records, or permits still depend on the material. The responsible manager could continue use, schedule another inspection, retire the layer in place, preserve it as a time-limited exception, or seek separate permission for staged retrieval. Any removal would leave a geospatial “tombstone” recording the former footprint, reason, evidence, and successor layer. The comparator is ordinary project-file review plus current surface inspection without unified identities, expiry triggers, dependency gates, exception states, or tombstones.
The cross-domain transfer¶
The layer-decay archetype maps directly onto successive physical installations that can outlive their original role. Stable identities, review deadlines, risk ordering, dependency checks, differentiated disposition, expiring exceptions, and removal records translate cleanly. The analogy does not make age proof of obsolescence or turn a database state into authority to disturb habitat.
Why it advanced¶
This was a STRICT_SUCCESS because it passed Experiment 5’s strict researched-candidate bar. That status does not show that multiple legacy layers are common at real reaches, that reviewers can classify them reliably, that an adopter will use the register, or that it improves bank stability, wildlife outcomes, or costs.
Prior art and the remaining open claim¶
Existing practice already includes safer material substitution, functional-longevity classes, inspection and maintenance records, removal requirements, permits, monitoring plans, and asset databases. The narrower open claim is that adding installation-level identities, review-only expiry triggers, a mandatory ecological, structural, and regulatory dependency check, time-limited exceptions, and post-removal tombstones will yield more reproducible and dependency-resolved recommendations than ordinary files and surface inspection. It expressly does not claim that older material should be removed or that the register itself reduces erosion, plastic pollution, or wildlife mortality.
Smallest decisive test¶
At a partner-managed reach, sample 20–50 mapped or suspected installations. Two qualified reviewers would independently assess the same nondestructive evidence under randomized workflows: ordinary files plus surface inspection, and the same material augmented by the lifecycle register and dependency checklist. Compare unique layers found, unresolved records, agreement, review time, live dependencies, and actionable recommendations. Stop at observation; do not probe or remove. Falsify the near-term claim if no sequential accumulation exists, the added workflow finds no decision-relevant layers or dependencies, agreement remains below a preregistered threshold such as kappa 0.60, review time more than doubles without better resolution, or either reviewer recommends retrieval despite a live or unresolved dependency.
Deployment and cost¶
Begin with a read-only inventory using existing plans, photographs, permitted observation, and provisional uncertainty labels. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup and operational launch, and \(10,000–\)50,000 annually. These are not quotations; no site partner or data-stewardship arrangement is secured.
Risks and uncertainties¶
- Surface observation may miss buried or visually similar layers and create false confidence in completeness.
- An age-weighted score may disguise subjective judgments as measurement precision.
- Staff may treat a review deadline as an instruction to remove material despite the dependency gate.
- “Retire in place” may become indefinite neglect if its exception and next review date are not enforced.
- Attention to installed mats may divert investigation from hydrologic or geomorphic causes of site failure.
Expert review¶
Useful reviewer backgrounds: Wetland or stream restoration ecologist, Geotechnical or hydraulic engineer, Environmental permitting specialist, Erosion-control contractor, GIS and restoration-records manager.
- Can reviewers identify deposition order and lifecycle state nondestructively and with acceptable agreement?
- Which observed conditions are sufficient to establish that roots, sediment, structures, or permits still depend on a layer?
- How should unknown installation dates and incomplete specifications affect priority without implying obsolescence?
- What authority and evidence are required before staged retrieval at the selected reach?
- Does the register add useful information beyond the site’s existing maintenance, permit, monitoring, and asset records?
Evidence and provenance¶
Selected sources: S1: Use of Sustainable Materials for Erosion and Sediment Control Practices · S2: Make the Change to Wildlife-Friendly Erosion Control Products! · S4: Water Quality General Certification No. 4100REV: Wetland and Stream Restoration and Creation · S5: Survey of wildlife-friendly temporary and permanent erosion control products and approaches · S6: Erosion Control Toolbox: Rolled Erosion Control Product (Blanket) · S7: Erosion prevention practices—erosion control blankets and anchoring devices · S8: Natural Channel Design Review Checklist
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized profiles placed the candidate 35th–37th, band C. This narrow reading-order range uses an affordability proxy and does not alter its strict experimental endpoint or measure field prevalence, ecological benefit, economic value, or deployment readiness.
36. Subtracting Recoater Self-Noise¶
Canonical title: Efference-Copy Residual Sensing for Powder-Layer Recoating
In one sentence: A shadow monitoring system would predict signals caused by a powder recoater’s own commands, route unexplained residuals for review, and revert to complete signals whenever validity or safety checks fail.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-13 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Chemistry Materials |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 64.0/100 · rank range 17–48 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-11, EXP06-PARTNER-12, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
During powder-layer recoating, commanded acceleration and routine powder contact create large, repeatable force, current, vibration, and acoustic signals. These self-generated patterns can conceal smaller signs of an agglomerate, protrusion, foreign object, uneven layer, or changed powder flow. Independent static thresholds can instead produce repeated benign alerts. It is not yet known whether command-caused motion explains enough signal on a target machine to create a useful residual, or whether complete raw monitoring already fits the available data and operator-attention capacity.
What is proposed¶
Each outgoing motion command would be copied into a frozen, versioned forward model that predicts the synchronized sensor response expected for one declared machine, recoater, powder class, speed range, and environment. The system would subtract that prediction from the measured full window, preserve signed and time-aligned residuals, and weight them by sensor reliability, timing, coherence, persistence, consequence, and processing cost. A compatible receiver could reconstruct the full signal from the prediction and residual. Validated residual clusters would route to an operator with a defined inspection action. Random and risk-based raw windows, periodic full-state anchors, checksums, heartbeats, and drift tests would audit the cancellation. Any timing, model, sensor, reconstruction, scope, or protected-safety failure would restore complete processing. Model updates would remain slower, reviewed, and reversible so recurring defects could not immediately be learned away. Comparators are full raw review and existing static thresholds.
The cross-domain transfer¶
Predictive residual processing becomes an efference-copy system: the controller’s own command predicts the machine-generated sensory return, and the unexplained remainder receives scarce transmission and operator attention. The structural mapping is strong, but adjacent disturbance observers, recoater sensors, digital twins, anomaly detectors, and predictive codecs already supply most components separately.
Why it advanced¶
This is an EMPIRICAL_PARTNER_CANDIDATE, not a strict success. It cleared a separate lane for a bounded external laboratory study, but lacks a target-machine dataset, frozen model, controlled challenge results, comparative reviewer evidence, measured capacity savings, validated fallback reliability, and an OEM or laboratory commitment.
Prior art and the remaining open claim¶
Recoater force and vibration sensing, powder-bed imaging, digital-twin recoating control, model-based collision detection, industrial interfaces, and bounded powder-spreading testbeds already exist. The remaining claim is narrower: on one fixed configuration, command-conditioned cancellation with reconstruction, raw audits, version checks, and forced fallback will detect and localize specified external interactions no worse than complete raw review while reducing total data and reviewer burden after computation, audits, fallbacks, inspections, and maintenance are counted. No live recoater experiment establishes that combined comparison, especially for events synchronized with commanded acceleration.
Smallest decisive test¶
An OEM or accredited laboratory would run a preregistered shadow test on one fixed recoater, sensor layout, motion program, medium, and environment while retaining every raw channel. Blinded challenges would include localized resistance, altered layer height, obstacle surrogates, powder-drag changes, sensor bias, timing offsets, dropped residuals, version mismatch, and missing heartbeats. Compare the reconstructed residual path, complete raw review, and static thresholds on protected-event capture, classification, localization, false alerts, reconstruction, fallback behavior, bytes, review time, inspections, and maintenance. Reject advancement if any protected or synchronization challenge fails to trigger full mode, any controlled interaction is over-cancelled, sensitivity or localization breaches the preregistered non-inferiority margin, or total capacity cost is not lower.
Deployment and cost¶
The first authorized step is a non-production shadow test; existing controls, interlocks, raw storage, and operator decisions remain authoritative. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. They are not OEM or supplier quotes.
Risks and uncertainties¶
- A wrong or mistimed model may subtract a real interaction that resembles the expected response to acceleration.
- Powder lot, humidity, wear, mounting, or sensor coupling may invalidate the learned self-signal.
- Sender and receiver can share the same wrong model while still passing a compatibility checksum.
- Rapid model adaptation could normalize recurring defects or wear instead of exposing them.
- Audit sampling may miss rare over-cancelled events, while frequent fallbacks may increase workload and pressure operators to weaken safeguards.
Expert review¶
Useful reviewer backgrounds: Additive-manufacturing process engineer, Machine-controls engineer, Recoater and powder-spreading specialist, Industrial sensing and signal-processing researcher, Machine safety and functional-safety engineer.
- What fraction of each target sensor’s variance is reproducibly explained by the outgoing command under nominal conditions?
- Which controlled interactions are most likely to resemble command-caused acceleration signals and be over-cancelled?
- What non-inferiority margin is acceptable for detection and localization against complete raw review?
- Can clocks, position traces, checksums, heartbeats, and full-state anchors force fallback within the required latency?
- After computation, raw audits, fallbacks, inspections, and maintenance, does the residual path actually reduce total capacity use?
Evidence and provenance¶
Selected sources: S1: In-process monitoring and non-destructive evaluation for metal additive manufacturing processes (NIST IR 8538) · S2: Software Releases · S3: 4040 Development & Demonstration of an Open Layered Protocol for Powder Bed AM · S4: Recoater Force Sensor Array for Spatial and Temporal In-Situ Powder Spreading Quality and Surface Defect Monitoring, US application 20240300024 · S5: Monitoring Laser Powder Bed Fusion Recoater Blade Vibrations for Collision Avoidance · S6: Smart Recoating: A Digital Twin Framework for Optimisation and Control of Powder Spreading in Metal Additive Manufacturing · S7: Practical Aspects of Model-Based Collision Detection · S8: Powder Spreading Testbed for Studying the Powder Spreading Process in Powder Bed Fusion Machines
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: Post-hoc harmonized profiles rank this candidate from 17th to 48th, band C, showing strong sensitivity to review priorities. The score is only a reading-order aid with a cost proxy, not an endpoint, economic-value measure, validation result, or deployment recommendation.
37. A Shared Pause for Audit Independence¶
Canonical title: Absent-User Independence Boundary Cycle
In one sentence: At key engagement changes, an audit team would visibly separate client-service pressures from its duty to absent report users, route uncertain matters through official channels, and make consultation or role substitution socially legitimate.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-27 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Accounting Auditing |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 64.0/100 · rank range 35–37 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-11, EXP06-PARTNER-26, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Audit firms already require independence checks, but completing a form does not ensure that a team confronts the purpose of independence together. During engagement acceptance, renewal, or scope changes, fees, staffing, deadlines, and client relationships may dominate discussion. Team members may struggle to identify the absent financial-statement users whom independence protects, raise uncertain services or relationships early, or step away from a role without appearing disloyal. The prevalence of this collective-meaning problem has not been measured directly.
What is proposed¶
At defined engagement triggers, the firm would hold a structured boundary cycle alongside—not instead of—its official independence process. Cards would represent the paying client relationship and the absent users relying on the audit, while an empty seat would make those users visible in the discussion. After private reflection, participants would sort invented or engagement-level service and role cards into ordinary processing, outside the engagement, or independent review. Personal facts would go only through confidential channels. The engagement sponsor would name commercial pressures, and an independent ethics witness would state limits the sponsor cannot override. Participants could speak, write, remain silent, or skip the symbolic portion. Every unresolved matter would then enter the authorized acceptance, staffing, consultation, or service-approval system with an owner, evidence requirement, decision authority, and deadline.
The cross-domain transfer¶
The source archetype uses a marked, recurring enactment to renew shared meaning and commitments. Here, the enactment represents absent report users, names commercial incentives, and recognizes consultation or role substitution. Its structural mapping is reasonably direct, but the ceremony itself never determines whether a person, service, or engagement satisfies independence requirements.
Why it advanced¶
This candidate cleared the separately calibrated empirical-partner lane because the underlying independence objective is consequential, a synthetic comparison is feasible, and the intervention has explicit authority and privacy safeguards. It did not enter the strict-success lane: no firm has requested it, and no field evidence shows that audit teams need or benefit from the ritual layer.
Prior art and the remaining open claim¶
Independence confirmations, relationship monitoring, service preapproval, acceptance reviews, training, confidential consultation, staffing restrictions, and quality oversight already perform most operational functions. Ethical culture and group rituals provide adjacent support for the underlying rationale, not evidence for this design. The open claim is narrower: with official rules unchanged, does the marked group cycle improve recognition and correct routing of uncertain matters and perceived consultation safety beyond ordinary workflow, without adding coercion, privacy risk, false assurance, or authority confusion?
Smallest decisive test¶
Preregister a synthetic two-arm tabletop with 12–16 participants per arm. Both groups receive the same invented engagement facts and mock confirmation, consultation, and acceptance workflow; only the treatment group receives the boundary cycle. A blinded reviewer would score recognition of absent-user interests and commercial incentives, identification and routing of uncertain matters, and rejection of group sorting as official approval. Also measure pressure, privacy, accessibility, consultation willingness, authority confusion, false assurance, and elapsed time. Reject the incremental claim if recognition or routing does not improve, a simpler comparator performs as well, or safeguards worsen materially.
Deployment and cost¶
The authorized first step is a tabletop using invented facts, not live client or personal information. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for initial setup, \(250,000–\)1 million for operational launch, and \(250,000–\)1 million annually. These are assessment bands, not vendor quotes.
Risks and uncertainties¶
- Silence, passing, or leaving could be interpreted as evidence that a participant has a personal conflict.
- The empty seat could imply that all financial-statement users share one interest.
- Group sorting could be mistaken for an official independence decision or evidence of unanimity.
- A fee-owning partner could control the script while retaining practical power over staffing and revenue.
- Personal financial, family, employment, or client-protected information could enter an unsuitable group setting despite the confidential route policy. Long-term use could become moral theater while incentives and retaliation risks remain unchanged.
Expert review¶
Useful reviewer backgrounds: Audit-independence or professional-ethics specialist, Audit-firm quality-control leader, Employment and privacy counsel, Organizational-behavior researcher, Experimental-design and measurement specialist.
- Do existing engagement processes already produce timely recognition and routing of the uncertain matters this cycle targets?
- Which actions or participation records could create employment, privacy, privilege, or professional-standards exposure?
- Can opting out, remaining silent, or requesting substitution be made credibly nonretaliatory in a partner-led team?
- What prespecified difference would justify the added meeting time over an improved form plus confidential consultation?
- How should reviewers detect whether participants mistake symbolic sorting or witnessed closure for official approval?
Evidence and provenance¶
Selected sources: S1: Revision of the Commission's Auditor Independence Requirements; Final Rule 33-7919 · S2: QC 1000, A Firm's System of Quality Control · S3: Spotlight: Inspection Observations Related to Auditor Independence · S4: Fostering a Healthy “Tone at the Top” at Audit Firms · S5: Ethics, Independence and Objectivity — Transparency Report 2025 · S6: International Code of Ethics for Professional Accountants, Including International Independence Standards · S7: Does Ethical Culture in Audit Firms Support Auditor Objectivity? · S8: Work Group Rituals Enhance the Meaning of Work
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed it in band C, with ranks 35–37 across profiles. That score is only a reading-order aid, uses cost as a pilot-affordability proxy, and does not measure economic value or change its empirical-partner endpoint.
38. Honest Guarantees for Generative Art¶
Canonical title: Guarantee-Bounded Preflight for Interactive Generative Art
In one sentence: An executable-art platform would state exactly which works and interactions it can verify, preserve unknown results outside that boundary, and leave artistic acceptance to curators.
| Field | Record |
|---|---|
| Portfolio ID | EXP04-STRICT-01 |
| Experiment and endpoint | Experiment 4 · Strict success |
| Archetype × domain | Computability Boundary Mapping × Art Aesthetics |
| Proposal position or arm | PROPOSAL_FIRST |
| Post-hoc reading order | Balanced score 63.0/100 · rank range 36–46 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
A platform wants a preflight tool to decide whether every submitted generative artwork will always terminate or reset and will never violate formal display limits under any future interaction or sensor stream. Yet the artwork language, external capabilities, input bounds, and guarantee are not precisely defined. A timeout or unsuccessful search may then be treated as approval or rejection even though it proves neither. The central issue is whether the unrestricted request crosses a computability boundary while narrower program fragments remain analyzable.
What is proposed¶
Before extending the verifier, the platform would define the artwork language, state behavior, permitted plug-ins and network access, input encoding, display constraints, and meaning of termination or reset. It would seek both a total procedure for restricted fragments and an independently checked impossibility argument for the unrestricted class. Works inside an enforceable finite fragment could receive exact or sound finite-state analysis. Explicitly bounded interactions could be checked exhaustively only within those bounds. Other works would receive sound but incomplete analysis that can report a possible violation or UNKNOWN, followed by curator review and runtime containment. Every report would identify its language version, assumptions, bounds, property, and guarantee type. The verifier would assess mechanically testable constraints, not beauty, meaning, or artistic merit.
The cross-domain transfer¶
Computability-boundary mapping separates an open-ended decision demand from regions where exact procedures are possible. In this setting, executable artworks and future audience inputs form the open class; restricted grammars, finite interaction envelopes, sound abstractions, explicit unknowns, and versioned records mark the narrower regions where honest guarantees can be made.
Why it advanced¶
It passed Experiment 4’s strict researched-candidate bar because the boundary problem is clearly specified, close comparators exist, safeguards preserve curator and artist authority, and a falsifiable archived-work study is defined. STRICT_SUCCESS does not mean real-world validation, novelty, deployment authorization, or demonstrated economic impact.
Prior art and the remaining open claim¶
Museums already use source review, risk assessment, emulation, sandboxes, watchdog resets, artist consultation, and cross-functional conservation. Formal model checking and abstract interpretation also exist for restricted programs. The remaining claim is the museum-specific combination: adding an enforceable language-and-property contract plus guarantee-labeled routing will reduce unsupported claims based on timeouts or sampled success, preserve UNKNOWN as non-dispositive, and retain curator-rated fidelity for works admitted to the restricted fragment. No supplied source establishes that joint result.
Smallest decisive test¶
With an authorized museum or platform, freeze one interpreter and examine 12 archived works read-only. Define three machine-testable display constraints, a finite event alphabet, and trace depth 20. Compare the proposed restricted-fragment and routing workflow with sandbox/watchdog review, property-based testing, and manual source review. An independent formal-methods reviewer would check semantics, soundness, coverage, and any impossibility proof. Measure coverage, time, concrete violations, false alarms, unknowns, label retention, scope overstatement, and curator-rated fidelity. Stop or redesign for any false clearance, incomplete declared enumeration, coercion of UNKNOWN, or failure to preserve essential behavior in at least eight works.
Deployment and cost¶
No adopter, corpus, source rights, or interpreter specification is secured. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for initial setup and operational launch, and \(50,000–\)250,000 annually. They exclude institution-specific rights clearance, legacy recovery, procurement, and major exhibition hardware.
Risks and uncertainties¶
- Formal palette, luminance, or exclusion-zone rules could be mistaken for the curator’s richer aesthetic judgment.
- A restricted language could exclude recursion, emergence, duration, or responsiveness essential to the artwork.
- A coarse abstraction could overwhelm reviewers with possible violations, while an unsound one could issue false clearance.
- Finite analysis may still be impractical because the modeled state space grows too large.
- Staff could operationally turn UNKNOWN into rejection despite the stated policy. Archived works may poorly represent future submissions or dependencies.
Expert review¶
Useful reviewer backgrounds: Formal-methods and computability researcher, Software-art conservator, Curator of digital or time-based media, Generative artist or artist-rights representative, Museum technology and data-governance counsel.
- Does the frozen artwork language actually contain the computational features assumed by the boundary argument?
- Can each formal display property be checked without misrepresenting the curator’s intended constraint?
- Which behaviors must the restricted fragment preserve for the selected works to remain faithful?
- Are the abstraction and bounded enumeration sound and complete for exactly the scopes printed on their reports?
- Will downstream staff preserve UNKNOWN and POSSIBLE_VIOLATION labels instead of translating them into acceptance decisions?
Evidence and provenance¶
Selected sources: S1: Handling Digital Assets in Time-Based Media Art · S2: The Conserving Computer-Based Art Initiative · S3: Matters in Media Art · S4: Risk Assessment as a Tool in the Conservation of Software-Based Artworks · S5: Emulation or it Didn’t Happen · S6: Formal Verification for Node-Based Visual Scripts Using Symbolic Model Checking · S7: Abstract Interpretation: A Unified Lattice Model for Static Analysis of Programs by Construction or Approximation of Fixpoints · S8: Software Developers, Quality Assurance Analysts, and Testers
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed this candidate in band C, spanning ranks 36–46 across profiles. It is a reading-order aid using cost as an affordability proxy, not an experimental endpoint, economic-value measure, or qualification of its STRICT_SUCCESS status.
39. Turning Archival Loss into Stewardship¶
Canonical title: Loss-to-Stewardship Assembly for Community-Governed Religious Archives
In one sentence: A community-controlled assembly would acknowledge authorized archival losses, invite bounded offers of help, and convert accepted offers into funded preservation work without demanding access, disclosure, or public grief.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-31 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Religious Studies Theology |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 63.0/100 · rank range 39–42 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
When a community-held religious recording, manuscript, oral history, or access relationship is destroyed, dispersed, withdrawn, or made inaccessible, a consortium may record only a status change or announcement. Later scholars and staff may not know what was lost, what must remain undisclosed, or who may describe it. Related materials can remain exposed, communities can face repeated requests, and proposed responses can lack owners, permissions, resources, or deadlines. The frequency and practical effects of this pattern have not been measured.
What is proposed¶
Only an authorized custodian could place a verified loss before the consortium and decide what may be public, controlled, or unspoken. After an independent review of standing, privacy, cultural protocol, accessibility, emotional pressure, and exit options, an ordinary empty archival sleeve would mark the acknowledged absence without imitating lost or sacred material. A custodian-approved account would state the type of loss and the claims outsiders may not make. Participants could observe a bounded silence, leave, or use another accessible mode. Institutions could then offer specific capacity—staff time, storage assessment, migration, translation, catalog correction, training, or funds—or pass without explanation. Offers would not be accepted during the marked interval. Custodians could later accept, change, or decline them, after which accepted offers would become permission-bound, funded work orders with owners and review dates.
The cross-domain transfer¶
The ritual archetype is instantiated as a recurring, marked occasion that joins shared memory to renewed commitments. Here, an empty sleeve marks absence, witnesses preserve authorized limits, and a circulating marker prompts offers of capacity. The mapping is meaningful, but the ritual cannot preserve materials, settle custodianship, or transfer access rights by itself.
Why it advanced¶
It cleared the empirical-partner lane because documentary losses and community-governed access are consequential, the design separates acknowledgment from consent, and a fictional tabletop comparison is feasible. It did not meet strict success: no consortium or community has requested it, and there is no field evidence that the assembly improves stewardship beyond ordinary governance.
Prior art and the remaining open claim¶
Loss registries, cultural-access protocols, grants, preservation consortia, emergency workflows, working groups, and memorial events already cover the component functions. Community-specific authority cannot be replaced by consortium permission. The open claim concerns only the added assembly: compared with the same layered registry, budget, meeting, and work-order process, does it improve accurate recall, completed custodian-approved work, and avoidance of repeat requests without increasing disclosure, pressure, unequal attention, emotional burden, or administrative cost?
Smallest decisive test¶
Preregister a six-to-eight-week crossover tabletop with 18–30 practitioners, preservation staff, and compensated community-archive advisers, using one fictional low-sensitivity loss. Compare the marked assembly with a permission-layered registry and ordinary preservation meeting; hold facts, restrictions, and capacity budget constant and randomize condition order. Measure immediate and 30-day recall, disclosure and overclaim errors, standing judgments, conversion of offers into funded work orders, task clarity, repeat requests, pressure, accessibility, burden, hours, and cost. Advance only if the assembly improves at least two prespecified stewardship outcomes with no authority or disclosure error and no worse pressure or burden.
Deployment and cost¶
The first authorized step uses no real community or collection. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for initial setup, and \(250,000–\)1 million for both operational launch and annual operation. No staffing model, vendor quote, case volume, or host budget verifies these bands.
Risks and uncertainties¶
- The consortium could appropriate or aestheticize a community’s loss.
- Public acknowledgment could reveal a restricted collection’s existence, location, vulnerability, or former contents.
- Custodians could feel pressured to narrate grief or grant access in exchange for assistance.
- Capacity offers could become donor branding or promises without assigned budgets.
- Visible, dramatic losses could attract resources away from slow deterioration or less prominent communities. The layered memory record could create later surveillance or security risks even when current access controls are followed.
Expert review¶
Useful reviewer backgrounds: Community archive custodian or authorized knowledge holder, Religious-collections archivist, Cultural-protocol and Indigenous data-governance specialist, Preservation-program manager, Privacy, copyright, and records-retention counsel.
- Who has standing to authorize the loss account when custodians or community members disagree?
- What information may be public, controlled, temporarily retained, or omitted entirely?
- Can declining participation, disclosure, or assistance remain genuinely costless under donor and institutional power differences?
- Does the assembly improve completed preservation work beyond an equally funded registry-and-working-group process?
- How should resources be allocated so publicly legible losses do not displace less visible but higher-risk preservation needs?
Evidence and provenance¶
Selected sources: S1: Protecting, preserving and promoting access to the world’s documentary heritage · S2: Religious Archives Group Conference in Association with The National Archives: Conference Report · S3: Protocols for Native American Archival Materials · S4: Understanding Communities and Cultural Protocols · S5: Records at Risk Grants · S6: How APTrust Works · S7: Mourning ritual participation, subjective well-being and prosocial behaviour among the Luhya people of Kenya · S8: About ARCS
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed it in band C, with ranks 39–42 across profiles. This is only a reading-order aid using a cost-band affordability proxy; it does not measure economic value, establish elapsed pilot time, or alter the empirical-partner endpoint.
40. Reviewing Ritual Records by Their Differences¶
Canonical title: Residual-Led Review of Recurrent Ritual Records
In one sentence: Researchers would use a collection-specific prediction model to foreground differences in recurring ritual records while preserving complete sources, independent full-record audits, and automatic fallback whenever the model becomes unreliable.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-06 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Predictive Residual Processing × Religious Studies Theology |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 63.0/100 · rank range 33–48 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-07, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Researchers comparing repeated service transcripts, ceremony descriptions, or ritual-manual editions may spend much of their first review pass rereading recurring structures while locally important changes compete for attention. Full-record reading preserves context but may be slow; simple exception lists can impose one supposedly normal form and hide changes that baseline cannot represent. The actual attention bottleneck has not yet been observed in a selected corpus, and predictability must be assessed separately for each collection rather than across religious traditions.
What is proposed¶
A read-only layer would predict the next descriptive event code—such as a role, action, object, or sequence position—before a held-out record is opened. The model would be versioned, collection-specific, and limited to declared metadata; it could not judge doctrine, authenticity, orthodoxy, or religious value. The interface would show structured insertions, deletions, substitutions, reorderings, and coding disagreements, weighted by uncertainty, source quality, underrepresentation, sensitivity, and interpretive importance. It could collapse only high-confidence expected units during first-pass review. Every packet would reconstruct the complete coded sequence and link to the untouched source. Sensitive, low-confidence, untranslated, permission-ambiguous, harm-related, or poorly represented records would appear in full. Independent random and risk-based full-record audits would detect missed context, while version mismatches, drift, or error-budget breaches would restore full-record review.
The cross-domain transfer¶
Predictive residual processing sends an expected signal plus the differences needed to reconstruct the observation, concentrating attention on prediction errors. Here, the expected signal is a collection-specific ritual-event sequence and the residual is a structured record difference. Checksums, raw-source audits, and fallback preserve reconstruction and expose model failure.
Why it advanced¶
It passed Experiment 6’s strict researched-candidate bar because the transfer is structurally clear, existing corpora and collation tools make a retrospective study plausible, and the comparator, safeguards, and stopping rules are explicit. STRICT_SUCCESS remains a research endpoint only, not field validation, novelty, deployment approval, or evidence of economic value.
Prior art and the remaining open claim¶
TEI apparatuses, CollateX, religious-studies toolkits, synoptic editions, recurrence analysis, anomaly ranking, and manual sampling already expose textual variation and support structured comparison. The remaining claim is their untested combination with pre-observation prediction, reconstructive residuals, consequence-aware routing, independent raw audits, and automatic fallback. For one authorized corpus, that system must save total review time while remaining noninferior to full-record and nonpredictive collation for material-variant recall, reconstruction, context coverage, and protected handling.
Smallest decisive test¶
With a corpus owner and required community steward, preregister a read-only study of 60 authorized, fully coded records. Train and freeze a transparent model on the earliest 30, then counterbalance reviewers across full-record review, nonpredictive collation, and the predictive-residual interface on 30 held-out records. A blinded panel would define material variants and audit every mandatory-bypass record plus random and risk-based collapsed records. Stop if the residual condition misses a bypass item, falls more than five points behind full review on recall or reconstruction, creates a greater than five-point miss disparity for underrepresented contexts, saves less than 20% total person-time, or fails to outperform nonpredictive collation.
Deployment and cost¶
No corpus access, permission determination, adopter commitment, or validated event-code scheme is secured. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for initial setup, \(250,000–\)1 million for operational launch, and \(50,000–\)250,000 annually. Measured engineering, coding, governance, and audit costs remain unavailable.
Risks and uncertainties¶
- A collection-specific expectation could acquire false authority as the normal or canonical ritual form.
- Routine wording, silence, performance, or material context omitted by event codes could carry the important meaning.
- Residual queues could reward novelty and make continuity seem unimportant.
- Underrepresented traditions or periods could be suppressed if consequence safeguards fail.
- Residuals could make sensitive outliers easier to identify. Model expectations could contaminate later coding, while gradual historical change could be absorbed as normal and disappear from attention.
Expert review¶
Useful reviewer backgrounds: Digital humanities or textual-collation specialist, Religious-studies corpus researcher, Community data steward, Machine-learning evaluation specialist, Research-ethics and sensitive-data reviewer.
- Is full-record rereading actually the binding attention cost in the proposed corpus?
- Which descriptive code units preserve performance, silence, translation ambiguity, material practice, and consequential routine wording?
- What material-variant gold standard can a blinded panel apply consistently?
- Do audit and fallback costs leave at least the prespecified 20% net person-time saving?
- Are miss rates, source expansions, and reconstruction errors acceptably balanced across periods, collections, and underrepresented contexts?
Evidence and provenance¶
Selected sources: S1: Arthur Westwell: Digital Techniques for Presenting Liturgical Texts and Building a Database of Carolingian Pontificals · S2: Corpus of Hittite Festive Rituals: Description · S3: WP3 – T-ReS – Toolkit for Religious Studies · S4: TEI Guidelines, Chapter 13: Critical Apparatus · S5: CollateX Documentation · S6: Recurrence Analysis Function, a Dynamic Heatmap for the Visualization of Verse Text and Beyond · S7: Quantifying Text Reuse Across Three Kṛṣṇa Yajurveda Recensions: Using Multi-Algorithm Computational Collation · S8: The CARE Principles for Indigenous Data Governance
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed it in band C, spanning ranks 33–48 across profiles. That wide range is a reading-order aid using cost as an affordability proxy, not an experimental endpoint, an economic-value estimate, or a change to STRICT_SUCCESS.
41. Renewing Ramp Stop-Work Support¶
Canonical title: Shared-Envelope Renewal for Ramp Stop-Work Legitimacy
In one sentence: A monthly, voluntary exercise would let workers from different ramp employers rehearse mutual recognition of approved stop-work calls and expose unsupported response paths without changing operating authority.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-12 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Aviation Aeronautics |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 63.0/100 · rank range 34–46 across three profiles · band C |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-13, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Aircraft turnarounds bring together airline staff and separate fueling, baggage, catering, cleaning, maintenance, towing, and ground-handling employers. Each may teach its own safety rules, yet workers can still disagree about who may pause shared work, how other companies must respond, and whether a junior contractor will be protected for interrupting a senior or time-critical operation. This uncertainty can delay or silence warnings even when every employer has a written stop-work policy.
What is proposed¶
Pilot a voluntary twenty-minute monthly observance at a classroom aircraft outline or closed training stand, never during a live turnaround. A rotating steward explains that the exercise creates no qualification, authority, or procedure. Participants from different employers trace or otherwise follow the aircraft boundary, hear frontline accounts of four cross-company dependencies, and exchange cards listing only approved support such as acknowledgment, interpretation, or escalation routes. Missing or contradictory support becomes a recorded safety-management action rather than an improvised promise. People may speak, write, observe silently, or pass without penalty. A debrief checks pressure, exclusion, ambiguity, and follow-through; independent review can require repair, suspension, or retirement.
The cross-domain transfer¶
The source archetype uses a recurring, marked ritual to make a shared commitment visible and memorable. Here, the ritual becomes a cross-employer aircraft-boundary exercise, reciprocal support-card exchange, plural response, and steward handoff. The mapping is structurally strong, but its symbolic elements have no assumed safety effect; their incremental value must be compared with an ordinary joint briefing.
Why it advanced¶
This candidate passed Experiment 6’s strict researched-candidate bar because it defines a bounded comparator, reversible first test, authority limits, concrete falsifiers, and safeguards against coercion and procedural confusion. STRICT_SUCCESS is only that experimental endpoint: it does not establish real-world effectiveness, novelty, deployment approval, or economic value.
Prior art and the remaining open claim¶
The surrounding parts already exist: joint ramp briefings, written escalation procedures, safety stand-downs, public safety pledges, reporting systems, and operational roles with stop-work authority. The open claim is narrower. Where approved routes already exist, does adding this opt-out monthly enactment improve unaided reconstruction of reciprocal cross-company responses and assignment of unsupported interfaces to authorized owners, compared with an equal-duration conventional briefing, without added pressure, retaliation concern, access loss, or authority confusion?
Smallest decisive test¶
At one willing station, place 8–12 workers from at least three organizations into matched twenty-minute sessions using identical fictional scenarios and route information: the proposed renewal or a conventional joint briefing. Blind-score who may call an approved pause, each organization’s response and escalation path, conflict handling, and what the session did not authorize. Also record whether planted gaps receive an owner, forum, and date, plus anonymous pressure and access measures. Proceed, revise, or stop; do not move to live operations without separate authorization.
Deployment and cost¶
The first-evidence estimate is under $10,000. Initial deployment, operational launch, and annual recurring support are each roughly \(10,000–\)50,000 in 2026 resource-equivalent terms, not vendor quotes. Local policy mapping, translation, shift coverage, accessibility, facilitation, auditing, and correction of discovered communication or staffing gaps could materially change the total.
Risks and uncertainties¶
- Managers could point to visible agreement while leaving anti-retaliation protection weak.
- Contractors or junior workers may feel compelled to affirm support in front of supervisors.
- Participants may mistake ceremonial statements or cards for approved operating procedures.
- Mobility, sensory, language, remote-access, and shift constraints may make participation unequal.
- A dominant airline or handler could control the script, stewardship rotation, or interpretation of mutual support promises.
Expert review¶
Useful reviewer backgrounds: Airport or airline safety-management-system specialist, Ground-handling operations leader, Ramp worker, union, or contractor representative, Aviation human-factors and safety-culture researcher, Accessibility and language-access specialist.
- Do local policies and contracts already define reciprocal stop-work acknowledgment across every participating employer, and where do they conflict?
- Can workers decline speaking, moving, or card exchange without conspicuous refusal or employment consequences?
- Does the matched briefing control contain exactly the same approved route information and facilitator time?
- What blinded scoring rule will distinguish genuine route reconstruction from recall of ceremonial wording?
- Which authority can accept each discovered gap, fund its correction, and verify closure?
Evidence and provenance¶
Selected sources: S1: Ground Operations Safety · S2: Ground Operations — Safety — SIRM 30 · S3: Easy Access Rules for Ground Handling (Regulations (EU) 2025/23 and 2025/20) · S4: Safety Management Systems (SMS) for Airports and Airport Projects · S5: Aircraft Arrival, Turnaround and Departure on Stand, ASGrOps_OSI_093 · S6: Introducing the Aviation Safety Pledge · S7: National Safety Stand-Down to Prevent Falls in Construction · S8: Air Transportation: NAICS 481
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized score is only a post-hoc reading aid: band C, with profile ranks from 34 to 46. Its affordability input proxies pilot speed and does not measure elapsed time, economic value, or the STRICT_SUCCESS endpoint.
42. Expiring Old Training Evidence Safely¶
Canonical title: Cognitive Evidence Lease Manager for Longitudinal Adaptive Training
In one sentence: A shadow system would retire contextually stale training evidence from current predictions while preserving the records and model versions needed to reconstruct earlier recommendations.
| Field | Record |
|---|---|
| Portfolio ID | EXP05-STRICT-02 |
| Experiment and endpoint | Experiment 5 · Strict success |
| Archetype × domain | Layer Decay And Expiration Management × Cognitive Science |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 62.0/100 · rank range 39–44 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
An adaptive cognitive-training system can keep using performance evidence from older sessions after a person’s ability, task design, scoring model, or context has changed. Those records may then distort current difficulty choices. Simply deleting them creates a different problem: researchers may lose the history needed to explain a trajectory, reproduce an earlier recommendation, investigate model behavior, or meet retention duties. The system needs to separate evidence’s authority in current inference from its physical preservation.
What is proposed¶
Assign each trial-derived evidence layer an inference lease when it is created, based on evidence type, task and scoring versions, context, and a review horizon. Expiry removes the layer from ordinary current-state estimation but sends it to a provenance-preserving archive rather than deleting it. New corroboration can renew a lease; task, scoring, or context changes can trigger early review. A multi-factor score may order review work but cannot decide disposition. Before destruction, stewards must check retention authority and downstream dependencies, use reversible quarantine, and leave a tombstone. Sampled restore drills test whether earlier estimates and recommendations remain reproducible. Initial evaluation occurs only through offline shadow replay, leaving source data, the live model, and participant experience unchanged.
The cross-domain transfer¶
The archetype manages accumulated layers whose usefulness decays or expires. In this domain, the layers are timestamped trials, summaries, context annotations, and model bindings. Expiration changes whether evidence influences today’s estimate, while storage tiers, dependency checks, quarantine, tombstones, and restore tests preserve accountable history. The structural mapping is direct, although temporal models may already address the predictive problem.
Why it advanced¶
This candidate passed Experiment 5’s strict researched-candidate bar through a specific comparator set, measurable offline test, reversible implementation, separated decision authority, and explicit failure rules. STRICT_SUCCESS does not show that stale evidence is prevalent, that leases beat strong temporal models, or that production use is warranted, novel, or economical.
Prior art and the remaining open claim¶
Adjacent systems already apply sliding windows, uniform forgetting, state-space or change-point models, time-aware knowledge tracing, immutable event logs, replay, and storage lifecycle policies. The unresolved comparison is whether explicit context-sensitive leases offer a smaller active evidence set that matches or improves held-out current-session prediction against cumulative and fixed-recency estimators—and remains competitive with a strong temporal model—while preserving historical reconstruction and keeping false-staleness, recommendation instability, and steward workload acceptable.
Smallest decisive test¶
With written data-steward approval, replay a completed multi-session working-memory dataset containing at least two documented task or scoring contexts. Freeze outputs from the existing cumulative estimator, then compare cumulative history, a fixed recent window, and context-sensitive leases; add a strong temporal-model sensitivity analysis if feasible. Measure held-out prediction, recommendation quality against a task-owner rubric, calibration, transitions, active-set size, false staleness, dependency recall, review minutes, and exact historical reconstruction. Reject progression after any missed seeded dependency, failed reconstruction, unacceptable instability or burden, unauthorized exposure, or baseline-reproduction failure.
Deployment and cost¶
First evidence is estimated at \(50,000–\)250,000. Initial deployment and operational launch are each roughly \(250,000–\)1 million; annual recurring work is \(50,000–\)250,000, in 2026 resource-equivalent bands rather than quotes. Actual cost depends on data volume, metadata quality, architecture, legal regimes, dependency tracing, expert review, and integration debt.
Risks and uncertainties¶
- Short or poorly chosen leases could overemphasize recent observations and discard stable information from active inference.
- Context rules may encode sensitive attributes or operate unevenly across participants.
- Adaptive task selection can make recent evidence look more informative because the system chose what was observed.
- A composite review score may hide contestable governance judgments behind numerical precision.
- Incomplete dependency mapping could permit removal of evidence needed to explain an earlier recommendation.
Expert review¶
Useful reviewer backgrounds: Cognitive scientist specializing in working-memory measurement, Adaptive-learning or knowledge-tracing modeler, Data steward or research-records officer, Privacy and data-retention counsel, Event-sourcing and archival-reconstruction engineer.
- What observed task, scoring, or context changes make an older trial semantically inapplicable rather than merely old?
- Which primary metric and non-inferiority margins will compare leases with cumulative, fixed-window, and strong temporal models?
- How will experts label false staleness without using the tested lease rules as their own ground truth?
- Can every sampled historical recommendation be recreated with the archived evidence, code, configuration, and model version?
- Which participant permissions, research requirements, holds, or deletion duties govern archive, quarantine, and destruction?
Evidence and provenance¶
Selected sources: s1: Does Working Memory Training Have to Be Adaptive? · s2: Adapting Training in Real Time: An Empirical Test of Adaptive Difficulty Schedules · s3: Rethinking and Improving Student Learning and Forgetting Processes for Attention Based Knowledge Tracing Models · s4: Time-Dependant Bayesian Knowledge Tracing—Robots That Model User Skills Over Time · s5: Event Sourcing Pattern · s6: AI Risk Management Framework Core · s7: Regulation (EU) 2016/679, General Data Protection Regulation · s8: Amazon S3 Pricing
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering places this candidate in band C, with profile ranks from 39 to 44. That ordering is not an experimental endpoint or economic-value estimate; its pilot-speed input is only a cost-band affordability proxy.
43. A Capped Prize for Catalyst Endurance¶
Canonical title: Capped Endurance Prize for a Durable Nanocatalyst
In one sentence: A sponsor would compare coded nanocatalysts through resource-capped, independently audited endurance tests rather than select a demonstration candidate mainly from publications or peak reported performance.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-07 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Nanotechnology |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 62.0/100 · rank range 31–50 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-03, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A sponsor with one follow-on demonstration award may rank nanocatalyst teams using publications, presentations, and peak results obtained under different conditions. That can favor short favorable runs, pure feedstocks, high scarce-metal use, selected batches, or heavy synthesis and computing expenditure. The chosen catalyst may then prove fragile during longer or variable operation. Yet the alleged selection process, incomplete reporting, resource escalation, and connection to later demonstration failure have not been documented for a named sponsor.
What is proposed¶
Replace publication-priority selection with a preregistered endurance prize whose rules are frozen before entrants are known. Teams register preparations, failed attempts, contributors, counted spending and in-kind inputs, reactor and accelerator-compute hours, scarce materials, and waste. Coded samples go to an independent laboratory for public qualification and undisclosed but representative endurance, feed-variation, restart, and regeneration sequences. Safety and containment are non-compensable gates. Passing entries are scored for reproducibility, sustained conversion and selectivity, deactivation and recovery, material intensity, energy, waste, and reporting completeness. Resource caps limit escalation. Two technically distinct milestone awards preserve alternatives before one demonstration award. Audits, appeals, conflict screening, graduated penalties, cleanup guarantees, and a challenger path govern the contest.
The cross-domain transfer¶
Bounded-rivalry governance redirects competition by fixing the arena, limiting inputs, policing interference, preserving alternatives, and reviewing winner lock-in. Here those elements become a capped catalyst prize with coded testing, milestone awards, safety floors, resource and waste accounting, sanctions, and continuation conditions. The mapping is strong, but the prize’s claimed predictive advantage over standardized endurance testing alone remains unmeasured.
Why it advanced¶
This did not enter Experiment 6’s strict-success lane. It cleared the separately calibrated EMPIRICAL_PARTNER_CANDIDATE lane because a bounded retrospective partner study is testable and safety-limited. Progress depends on a named sponsor and proprietary records; neither the field problem nor the full contest design’s incremental benefit has yet been demonstrated.
Prior art and the remaining open claim¶
Catalyst benchmarking, degradation protocols, independent laboratory validation, staged federal prizes, advance scoring rules, safety controls, and appeals already exist separately. The remaining claim concerns their combination: would adding auditable input caps, complete-attempt reporting, lifecycle burdens, independent milestones, hidden representative segments, and a continuation challenge predict later demonstration performance better than both historical peak-result selection and standardized endurance testing alone? Located evidence supports the components, not that comparative or behavioral result.
Smallest decisive test¶
With a named sponsor, preregister a non-awarding shadow study of 5–15 archived projects. Compare the historical decision, an endurance-only ranking, and the full capped and lifecycle-weighted ranking using coded records and held-back time-series segments. Measure rank uncertainty, missingness, reconstruction labor, incumbent effects, audit reversals, appeal burden, and association with later demonstration outcomes. Stop if fewer than 80% of histories are reconstructable, measurement uncertainty exceeds project differences, reasonable weights reverse the rankings, the full design fails to beat endurance alone, or governance cost is disproportionate. No new synthesis, testing, funding change, or award follows.
Deployment and cost¶
The retrospective evidence study is estimated at \(50,000–\)250,000. Initial setup is \(250,000–\)1 million; operational launch and annual operation are each roughly \(1–\)5 million in 2026 resource-equivalent terms. These are not quotes. Reaction-specific protocols, laboratory capacity, prize administration, purse size, environmental review, confidentiality, and demonstration scope remain unpriced.
Risks and uncertainties¶
- Hidden test sequences may reward resistance to surprise conditions rather than representative durability.
- Interlaboratory or segment variability may be larger than genuine differences among catalysts.
- Resource caps could favor teams that already own equipment, datasets, or precursor inventories.
- Broad accounting may expose confidential operations; narrow accounting may shift spending off book.
- Institutional cleanup guarantees could exclude capable teams without wealthy sponsors even when their methods are safe scores can obscure tradeoffs among activity, selectivity, endurance, scarce materials, energy, and waste. A sponsor, catalyst metrologist, independent laboratory, safety specialist, competition designer, and research-accounting expert must determine whether an auditable comparison is feasible.
Expert review¶
Useful reviewer backgrounds: Catalysis scientist with durability-testing expertise, Nanomaterial measurement and interlaboratory-validation specialist, Independent validation-laboratory operator, Chemical safety, exposure, and waste specialist, Federal prize, procurement, or research-competition counsel.
- Does a named sponsor actually make a scarce follow-on decision using publication priority or peak team-reported results?
- Can archived records distinguish every attempted preparation and run from the selected results presented to the sponsor?
- What target reaction, benchmark, operating envelope, and interlaboratory error define a fair endurance comparison?
- Which resource categories can be audited without unfairly favoring incumbents or exposing protected information?
- Does the full ranking predict later demonstration outcomes better than endurance-only scoring across preregistered weights?
Evidence and provenance¶
Selected sources: S1: Towards Benchmarking in Catalysis Science: Best Practices, Challenges, and Opportunities · S2: The rotating disc electrode: measurement protocols and reproducibility in the evaluation of catalysts for the oxygen evolution reaction · S3: Standardized protocols for evaluating platinum group metal-free oxygen reduction reaction electrocatalysts in polymer electrolyte fuel cells · S4: Nanotechnology Measurement Protocols · S5: Electrolysis Catalyst Synthesis, Ex-situ Electrochemical Performance and Durability Characterization, and Standardization · S6: Prize and Challenge Toolkit: A Guide for Federal Innovation Managers · S7: Accelerating Innovation through American-Made Challenges · S8: Guidance and Publications: Nanotechnology
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading aid assigns band C and profile ranks from 31 to 50. It does not upgrade the EMPIRICAL_PARTNER_CANDIDATE endpoint, measure economic value, or establish elapsed pilot speed; affordability only proxies that input.
44. Send Foresight Surprises, Keep Full Records¶
Canonical title: Scenario-Residual Exchange for Distributed Horizon Scanning
In one sentence: Distributed scanners would send structured differences from a shared scenario expectation while retaining full evidence, protected-signal bypasses, independent audits, and automatic return to complete-packet review when the filter becomes unreliable.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-16 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Futurism Foresight |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 62.0/100 · rank range 41–45 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A distributed foresight network may send complete weekly assessments for every monitored driver, forcing central analysts to reread expected material to find a few assumption-breaking observations. A simple report-by-exception rule is unsafe because silence could mean stability, missing reporting, or model blindness. The proposed setting still lacks packet-level evidence that routine material actually exhausts review capacity, that important signals arrive late for this reason, or that downstream decision blockage is not the real bottleneck.
What is proposed¶
Before each weekly intake, scenario stewards publish versioned expectations for each driver’s direction, pace, geography, actors, cross-impacts, uncertainty, scope, and expiry. Local analysts retain full evidence capsules but encode structured differences such as acceleration, reversal, a new actor, altered coupling, or broken assumption. A gate prioritizes these residual cards by reliability, uncertainty, consequence, perspective coverage, and review cost. Heartbeats distinguish expected conditions from missing reports, while protected safety-, rights-, conflict-, and distributional-harm signals travel in full. Central analysts reconstruct the assessment from the matching expectation and residual; validated surprises receive an owner but do not automatically change scenarios. Independent raw sampling, periodic reconciliation, drift monitoring, and version, error, or missingness triggers restore full-packet review.
The cross-domain transfer¶
Predictive-residual processing sends deviations from a synchronized expected state instead of retransmitting the entire state. Here, the expected state is a versioned scenario-conditioned driver assessment, and the residual describes model-relative change. Reconstruction, heartbeats, raw audits, protected bypasses, and fallback make the mapping substantive, though translating ambiguous foresight narratives into reliable residual fields remains an unresolved weakness.
Why it advanced¶
This did not qualify as an Experiment 6 strict success. It entered the EMPIRICAL_PARTNER_CANDIDATE lane because a shadow comparison on an adopter’s historical corpus is bounded and measurable. The essential field evidence—actual backlog prevalence, missed-signal causes, taxonomy reliability, protected-signal recall, and fully loaded labor savings—is still missing.
Prior art and the remaining open claim¶
Scenario-guided scanning, indicator-based early detection, distributed scanning networks, short-list filtering, collaborative platforms, ordinary tags, anomaly alerts, and periodic scenario refreshes are adjacent prior art. The narrower open claim is that a synchronized expectation-plus-residual workflow can reduce total review effort versus complete packets, tagged triage, and existing signpost filtering without materially reducing detection of assumption changes, cross-driver effects, novel actors, or protected-source evidence once preparation, audits, reconciliation, and fallback are counted.
Smallest decisive test¶
Obtain 200–500 authorized historical packets with timestamps. Using only earlier material, freeze expectations for each evaluation window and randomly assign blinded reviewers to complete packets, tagged complete packets, or residual cards with logged fallback. Preregister review-time savings, reconstruction agreement, recall margins, and equal protected-source recall; include raw audits and injected version mismatches and missing heartbeats. Reject the workflow if any safety-class item is suppressed, reconstruction breaches tolerance, underrepresented sources fare worse, fallback fails, systematic residual structure persists, or expectation preparation, auditing, reconciliation, and fallback erase the labor saving.
Deployment and cost¶
First evidence is estimated at \(10,000–\)50,000. Initial deployment is \(50,000–\)250,000, operational launch \(250,000–\)1 million, and annual recurring work \(50,000–\)250,000 in rough 2026 resource-equivalent bands, not quotes. Taxonomy design, expectation maintenance, adjudication, audit independence, access controls, integration, training, and fallback workload may dominate software costs.
Risks and uncertainties¶
- Shared expectations could become a self-confirming filter that suppresses evidence outside current scenarios.
- Precision or consequence weights may discount unfamiliar regions, disciplines, sources, or minority perspectives.
- A heartbeat may be recorded even when the underlying source network has quietly failed.
- Residual cards may remove narrative context needed to interpret ambiguous long-range evidence.
- Analysts may change classifications to attract central attention or avoid scrutiny.
Expert review¶
Useful reviewer backgrounds: Government or corporate foresight program leader, Scenario-planning and horizon-scanning methodologist, Information-retrieval or human-in-the-loop systems researcher, Rights, conflict, and distributional-impact specialist, Records, privacy, and access-control officer.
- Do complete packets currently exceed a declared review budget, and which material findings were delayed specifically by intake saturation?
- Can independent annotators reliably encode and reconstruct the proposed driver-state and residual fields?
- Which evidence classes must bypass filtering regardless of predicted relevance or apparent routine status?
- Does the residual workflow preserve protected and underrepresented-source recall under blinded adjudication?
- After counting preparation, maintenance, audit, reconciliation, and fallback, is total analyst time at least meaningfully lower than complete review?
Evidence and provenance¶
Selected sources: S1: Building capacity in technology horizon scanning: A guide for policymakers · S2: Building Anticipatory Capacity with Strategic Foresight in Government: Lessons from Lithuania, Italy, and Malta · S3: The Futures Toolkit · S4: Enhancing horizon scanning by utilizing pre-developed scenarios: Analysis of current practice and specification of a process improvement to aid the identification of important weak signals · S5: Anticipating surprise: The case of the early warning system of Rijkswaterstaat in the Netherlands · S6: Integrating scenario planning and indicator-based early detection for scenario transfer · S7: FIBRES Pricing · S8: Employer Costs for Employee Compensation — December 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized ordering is a post-hoc reading aid: band C, with profile ranks from 41 to 45. It neither changes the EMPIRICAL_PARTNER_CANDIDATE status nor measures economic value; the pilot-speed input is an affordability proxy, not elapsed time.
45. Residual-First Review of Forensic Timelines¶
Canonical title: Residual-First Triage for Digital Forensic Timelines
In one sentence: A read-only review layer would foreground unexplained timeline changes while preserving the complete forensic record, reconstructible context, independent audits, and automatic return to full review.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-04 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Predictive Residual Processing × Criminology Forensic |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 62.0/100 · rank range 32–51 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-05, EXP06-PARTNER-14, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Digital forensic examiners may confront millions of timestamped system, application, synchronization, and acquisition events. Repeated, predictable background activity can consume attention before an examiner reaches missing, extra, displaced, or altered events that matter to a case. Simply hiding familiar-looking events is unsafe: an apparently routine record may still reveal guilt, innocence, provenance, deletion, clock problems, or evidence-integrity failures, and context-free anomalies can themselves invite overinterpretation.
What is proposed¶
For one declared operating-system, application, and version scope, freeze a reference model that predicts ordinary background event bundles. Compare the complete extraction with those predictions and describe each discrepancy as missing, extra, reordered, time-shifted, or attribute-changed. Show prioritized discrepancies with model identity, provenance, uncertainty, and enough neighboring events to reconstruct the local sequence. Never alter the forensic image or complete extraction. User-authored material, deletion indicators, integrity and acquisition failures, clock discontinuities, required disclosures, and examiner-requested records always appear in full. An independent reviewer examines random and risk-selected raw windows. Unsupported versions, model or parser mismatches, audit disagreement, excessive reconstruction error, and queue overload switch the affected scope back to complete chronological review. Human reviewers, not the model, determine evidentiary meaning.
The cross-domain transfer¶
The predictive-residual archetype maps strongly here: a versioned model predicts routine event sequences, while the analyst receives typed differences between prediction and observation. Reconstruction, version checks, independent raw sampling, protected-event bypasses, drift monitoring, and full-review fallback instantiate the archetype without replacing the underlying evidence.
Why it advanced¶
This candidate passed Experiment 6's strict researched-candidate bar because the burden is documented, a bounded prototype appears feasible, the comparison is falsifiable, and the safeguards directly address context loss and automation risk. That status is not real-world validation, a novelty finding, deployment approval, or evidence of economic impact.
Prior art and the remaining open claim¶
Complete timeline extraction, static filters, analyzers, pattern reconstruction, prioritization, provenance drill-down, clustering, and visualization already exist. The remaining claim is narrower: compared with full chronological review and the strongest static-filter or analyzer workflow, a frozen residual-first layer combining typed discrepancies, reconstructible context, protected bypasses, independent raw-window audits, synchronized versions, and automatic fallback would reduce review effort without increasing material or exculpatory misses, interpretation errors, or total audit and maintenance burden.
Smallest decisive test¶
Pre-register a randomized crossover study with about 8–12 qualified examiners, 4–6 synthetic or reusable closed-case images, and 24–36 matched sections from one software-version scope. Compare complete review, the strongest Plaso or Timesketch static workflow, and residual-first review. Measure examiner minutes, time to inspect material events, inculpatory and exculpatory recall, protected-context recall, false escalation, reconstruction disagreement, disparities, fallback, and total workload. Stop progression for any protected-bypass failure, unreviewed material suppression, unacceptable recall loss, recurrent mismatch, systematic disparity, negligible effort reduction, or audit burden comparable to full review. Passing would not authorize live-case use.
Deployment and cost¶
The first retrospective evidence study is estimated at \(50,000–\)250,000. Initial deployment startup is \(250,000–\)1 million; an operational launch is \(1–\)5 million, with \(250,000–\)1 million recurring annually. These are rough 2026 USD resource-equivalent bands, not vendor quotes, and version maintenance may erase expected savings.
Risks and uncertainties¶
- A mistaken reference model could suppress material inculpatory, exculpatory, provenance, or integrity information.
- Examiners could mistake statistical unusualness for criminal intent, authorship, or evidentiary significance.
- Parser changes, clock ambiguity, missing data, or mismatched model versions could create false discrepancies or false silence.
- Reference data and priority weights could work unevenly across applications, languages, devices, or user contexts.
- Sparse audit samples might miss rare blind spots, while incomplete version records could obstruct disclosure and independent challenge.
Expert review¶
Useful reviewer backgrounds: Digital forensic examiner, Forensic laboratory quality and method-validation specialist, Forensic timeline-tool developer, Defense-side digital-forensics expert, Evidence and disclosure lawyer.
- Can sampled timeline windows be reconstructed to the semantic fidelity required for independent examination and disclosure?
- Which event classes require unconditional full-context presentation under laboratory policy and applicable law?
- Does residual-first review preserve inculpatory and exculpatory recall against both complete review and the strongest static workflow?
- How often would operating-system, application, parser, locale, clock, or acquisition changes force model revalidation?
- Does audit, synchronization, and maintenance work remain below the attention saved during review?
Evidence and provenance¶
Selected sources: S1: Needs Assessment of Forensic Laboratories and Medical Examiner/Coroner Offices: A Report to Congress · S2: Digital Forensics · S3: Method validation in digital forensics (accessible), FSR-G-218 · S4: Using log2timeline.py — Plaso documentation · S5: Create an analyzer — Timesketch · S6: An Automated Timeline Reconstruction Approach for Digital Forensic Investigations · S7: Forensic Science Technicians: Occupational Outlook Handbook · S8: SoK: Timeline-based event reconstruction for digital forensics: Terminology, methodology, and current challenges
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized score is only a post-hoc reading-order aid: this candidate ranked 32–51 across profiles and falls in band C. The ranking is not an experimental endpoint or a measure of novelty, deployment readiness, or economic value.
Band D — lower post-hoc review priority (ranks 46–59)¶
46. A Clinic for Cross-Field Lemma Handoffs¶
Canonical title: Rotating Proof-Interface Clinic for Cross-Subfield Lemma Handoffs
In one sentence: A rotating pair of specialists would turn eligible cross-subfield proof obstructions into precise, review-ready lemma packets without proving the lemma or changing the theorem.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-10 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Catalytic Pathway Enablement × Mathematics |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 61.0/100 · rank range 43–51 across three profiles · band D |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
The problem¶
In a mathematics collaboration spanning subfields, a proof obligation may stall because participants use different definitions, notation, assumptions, or standards of explanation. Contributors repeatedly search for appropriate specialists, reconstruct terminology, and discover lost hypotheses only after work begins. Yet no mathematics-specific audit shows how often this translation and routing work, rather than genuinely missing mathematical insight, causes delay. Informal access to well-connected experts may also determine which obligations receive attention.
What is proposed¶
Create a governed clinic staffed by a rotating pair of specialists covering both sides of an interface. Intake must state the parent theorem, local definitions, requested conclusion, allowed assumptions, known dependencies, attempted approaches, and exact source of confusion. In a capped session, the specialists check compatibility, translate notation, preserve quantifiers and hypotheses, separate established implications from open mathematics, identify side conditions, and produce the smallest faithful lemma packet with an intended recipient or escalation path. An independent reviewer compares the packet with the original request. The clinic must reject disputes about truth, foundations, credit, theorem changes, or genuinely new mathematics. It tracks queue depth, labor, reopens, defects, conflicts, reviewer capacity, and specialist recovery, then rotates or rests staff instead of treating them as unlimited infrastructure.
The cross-domain transfer¶
The catalytic-pathway analogy is plausible but not literal. The clinic acts as a reusable facilitator that lowers recurring translation and routing barriers, releases each packet, and returns to readiness. It cannot change whether a lemma is true or replace the mathematical insight and scrutiny required to prove it.
Why it advanced¶
This candidate did not enter the strict-success lane. It cleared a separately calibrated empirical-partner lane because a small shadow comparison is practical and reversible. Advancement therefore means it merits a bounded external data-partner study, while the prevalence of the problem, available records, specialist willingness, and comparative benefit remain unverified.
Prior art and the remaining open claim¶
Focused mathematics programs, public question intake, shared glossaries, direct consultation, formal dependency blueprints, and large collaborative proof projects already address parts of the problem. The open claim concerns the governed combination: for archived obligations crossing the same two subfields, a capped rotating clinic with a versioned lemma-packet contract would beat an equal-access ad hoc handoff on preparation labor or time without adding statement defects, reopens, bad routing, or hidden downstream burden.
Smallest decisive test¶
Find one willing project and eight authorized, de-identified, resolved obligations from the same subfield pair, including translation successes, substantive proof problems, and an incompatible case. Compare the rotating clinic with an ad hoc team given equal source access and recorded labor; conceal archived resolutions until packets are frozen. Blinded reviewers score statement fidelity, definitions, hypotheses, quantifiers, uncertainty, routing, readiness, reopens, total labor, elapsed time, and downstream burden. Reject progression after a confidentiality breach, undisclosed material alteration, failure to reject incompatibility, mostly substantive escalations, increased defects or burden, or failure to meet the precommitted practical improvement threshold.
Deployment and cost¶
The first evidence study and initial startup are each estimated at \(10,000–\)50,000. Operational launch and annual recurring support are each \(50,000–\)250,000. These rough 2026 USD resource-equivalent bands are not quotes; scarce dual-subfield specialists, independent reviewers, compensation, and confidentiality controls could dominate actual cost.
Risks and uncertainties¶
- Translation could silently remove a hypothesis, change a quantifier, or overstate equivalence between definitions.
- Clinic packets or specialists could acquire informal authority despite having no power to accept mathematical claims.
- Eligibility rules could favor contributors already fluent in preferred notation or connected to clinic staff.
- Rotating specialists could lose continuity, become exhausted, or receive inadequate credit for substantive formulation work.
- Confidential conjectures, correspondence, or attribution information could reach unintended recipients.
Expert review¶
Useful reviewer backgrounds: Mathematician from each participating subfield, Collaborative-project proof architect, Mathematical editor or independent proof reviewer, Research ethics and confidentiality specialist, Academic labor and attribution specialist.
- Can the project identify eight comparable resolved obligations whose reuse is authorized and whose resolutions can be concealed?
- What fraction of archived stalls arose from translation and routing rather than missing mathematical insight?
- Which changes to definitions, hypotheses, or quantifiers count as material fidelity defects?
- Can rotating specialists maintain continuity without exceeding fair workload and compensation limits?
- Does the clinic reduce total submitter, specialist, and downstream-review effort under equal access?
Evidence and provenance¶
Selected sources: S1: Investigating communication hindrance in interdisciplinary collaboration: A grounded theory approach · S2: Cultural barriers to interdisciplinary research collaboration: evidence from Australia · S3: SQuaREs: Structured Quartet Research Ensembles · S4: Collaborate@ICERM · S5: How do I ask a good question? · S6: LeanArchitect: Automating Blueprint Generation for Humans and AI · S7: The Equational Theories Project: Advancing Collaborative Mathematical Research at Scale · S8: National employment and wage data by occupation, May 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized assessment places this candidate 43–51 across profiles, in band D. That post-hoc ordering only guides reading and uses a cost-band affordability proxy; it is not an experimental endpoint, economic-value estimate, or substitute for the missing partner study.
47. Choosing One Proof Route Fairly¶
Canonical title: Bounded Proof-Verification Slot Challenge
In one sentence: A no-stakes challenge would test whether correctness-gated, resource-capped comparison selects a more maintainable route for one scarce formal-verification slot than simpler selection methods.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-02 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Bounded Rivalry Governance × Mathematics |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 61.0/100 · rank range 38–55 across three profiles · band D |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A mathematics consortium may have several proposed routes to a theorem but enough specialist time to formalize only one. If the slot goes to the first route that appears complete, teams may benefit from premature completeness claims, hidden dependencies, elaborate presentation, or silence about fatal flaws. A strategically packaged route can therefore displace a sounder, more maintainable alternative, consuming scarce referee and formalizer time while discouraging useful sharing between teams.
What is proposed¶
Run a voluntary, time-limited challenge under rules frozen before judges see team identities. Every team submits the same proof certificate, listing contributors, axioms, imported results, dependencies, unresolved gaps, and known counterchecks. Independent reproduction is a pass-or-fail gate: presentation quality or expected cost cannot compensate for a failed mathematical step. Only passing routes are ranked on declared criteria such as dependency transparency, modularity, explanatory coverage, and estimated formalization burden. Give teams equal page limits, clarification rounds, and judge access. Permit disclosed collaboration and route merging, while prohibiting tampering, plagiarism, retaliation, private judge contact, sham independence, and concealment of a known fatal audit flaw. Provide a separate procedural appeal and reopen the slot if staged formalization later fails. Preserve attribution and access for every correctness-passing route.
The cross-domain transfer¶
The bounded-rivalry archetype maps directly to competition for one consortium-funded verification slot. A frozen rulebook, noncompensable correctness gate, equal reviewer-facing resource caps, auditable conduct rules, appeal, and reopening constrain strategic escalation while directing comparison toward downstream verification needs. Rivalry remains optional; collaboration may still prove better.
Why it advanced¶
This candidate passed Experiment 6's strict researched-candidate bar because the scarce-resource setting is coherent, adjacent infrastructure exists, authority can be bounded, and a reversible comparison can falsify the claim. It has not demonstrated better selection in practice and is not a novelty, deployment, or economic-impact finding.
Prior art and the remaining open claim¶
Proof certificates, machine checking, contribution rules, dependency graphs, task dashboards, formalization challenges, and large collaborative projects already exist. The unresolved contrast is whether their governance elements work better in this specific allocation decision: can correctness-gated ranking with equal access, frozen secondary criteria, appeal, and staged reopening select a route requiring fewer corrections or formalizer hours than blinded holistic triage or first reproduction, without excess review overhead or suppressed collaboration?
Smallest decisive test¶
With a consenting formalization organization, preregister a no-stakes shadow exercise using eight de-identified packets from a settled theorem: sound routes of differing modularity, known gaps, an undeclared dependency, a polished distraction, and controls. Compare the governed challenge with blinded holistic triage, while logging first independent reproduction as a third comparator. Auditors reproduce critical lemma chains and measure false passage, corrections, observed formalizer hours on a fixed sample, reviewer minutes, agreement, successful planted gaming, appeals, and collaboration effects. Do not progress if an incorrect packet passes, gaming determines selection, agreement misses its threshold, both comparators perform as well or better, or governance exceeds its budget.
Deployment and cost¶
The shadow study is estimated at \(10,000–\)50,000. Startup and operational launch are each \(50,000–\)250,000, with annual recurring costs of \(250,000–\)1 million. These are rough 2026 USD resource-equivalent bands, not budgets or quotes; expert reproduction and governance, rather than software, are the main uncertainties.
Risks and uncertainties¶
- A visible ranking could turn proof development into a status contest and discourage useful collaboration.
- Secondary scoring could reward a fashionable proof style rather than actual maintainability or mathematical value.
- Page and clarification caps could disadvantage routes whose irreducible explanations are longer.
- Dependency disclosure could expose unpublished work, while misconduct controls could stigmatize legitimate collaboration.
- Judges could underestimate formalization effort or apply school-specific preferences inconsistently.
Expert review¶
Useful reviewer backgrounds: Research mathematician familiar with the theorem, Formal-proof engineer or maintainer, Independent mathematical referee, Research-governance and due-process specialist, Mathematical collaboration researcher.
- Does the target consortium actually face several viable routes competing for one indivisible verification slot?
- Can judges reproduce correctness and apply the secondary rubric with preregistered agreement?
- Do page and contact caps equalize reviewer access without shifting effort into hidden channels?
- Does challenge-selected work require fewer corrections and formalizer hours than both comparison methods?
- Would the format measurably discourage route sharing, merging, or disclosure of negative findings?
Evidence and provenance¶
Selected sources: S1: Contributing to mathlib · S2: Completion of the Liquid Tensor Experiment · S3: IMO Grand Challenge · S4: Exponentiating Mathematics (expMath) · S5: Formalising perfectoid spaces · S6: National employment and wage data by occupation, May 2025 · S7: Formalising Fermat · S8: The Equational Theories Project: Advancing Collaborative Mathematical Research at Scale
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized assessment ranks this candidate 38–55 across profiles and assigns band D. This post-hoc reading order is not the Experiment 6 endpoint and does not measure economic value; its affordability input also does not independently estimate elapsed pilot time.
48. Monitoring a Program’s Measurement Footprint¶
Canonical title: Efference-Residual Monitoring for Crime-Prevention Programs
In one sentence: A frozen model would subtract the aggregate records expected from a place-based prevention program so evaluators can inspect unexplained changes without producing individual scores or treating residuals as causal conclusions.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-14 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Criminology Forensic |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 60.0/100 · rank range 40–53 across three profiles · band D |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-04, EXP06-STRICT-05, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A patrol, outreach, reporting, or other place-based prevention program can change both underlying events and how those events are recorded. More officer presence, for example, may predictably alter contacts, detected offenses, calls, complaints, or data completeness. Ordinary dashboards can blur this measurement footprint with external change, displacement, service withdrawal, or harm. Yet unexplained aggregate differences can also be overread as proof of program success, failure, misconduct, or community behavior.
What is proposed¶
Before each monitoring window, copy the authorized program schedule and intensity into a frozen, versioned model that predicts only the aggregate records and observation opportunities the program itself is expected to generate. Compare those expectations with separately retained observations and report typed differences: excess or missing activity, spatial or temporal displacement, source disagreement, complaint or injury changes, and unexplained missingness. Prioritize residuals by reliability, uncertainty, exposure, persistence, rights consequence, and evaluator capacity. Each alert must include its expected value, observed aggregate, uncertainty, provenance, program history, and reconstructible full-window context. Complaints, force, injury, deaths, disparity checks, whistleblower material, outages, integrity failures, and oversight requests always appear in full. Independent reviewers audit random and risk-selected windows. Residuals open evaluation inquiries only; shocks, small cells, mismatch, drift, or audit failures restore complete reporting.
The cross-domain transfer¶
The predictive-residual archetype is instantiated through an efference copy of the agency's own authorized schedule: expected program-generated records are separated from the unexplained remainder. The mapping is structurally strong for monitoring, but residuals cannot identify true crime levels, individual risk, causation, or the correct operational response.
Why it advanced¶
This candidate did not enter the strict-success lane. It cleared the separate empirical-partner lane because a retrospective aggregate comparison is bounded and measurable. Missing field evidence remains central: no agency has supplied schedules, linked sources, reviewers, legal approval, audit access, or evidence that the model separates measurement footprint from omitted context.
Prior art and the remaining open claim¶
Multi-indicator dashboards, pre/post comparisons, matched comparison areas, process and impact evaluations, displacement analysis, and generic anomaly detection already cover much of the substantive work. The narrower open claim is that an action-conditioned residual interface, with reconstructibility, independent raw-window audits, rights-critical bypasses, synchronization, and fallback, would reduce evaluator effort against full and conventional evaluation dashboards without materially missing displacement, reporting divergence, service withdrawal, or harm, or encouraging stronger unsupported causal claims.
Smallest decisive test¶
Pre-register a 12–16-week offline study using 150–250 historical or synthetic windows from one completed program and at least eight blinded evaluators. Compare the existing full dashboard, a conventional process-plus-impact dashboard, and the residual-first interface in balanced crossover order. Include outages, intensity changes, displacement, complaint or injury shifts, service withdrawal, benign shocks, source disagreement, and version mismatch. Require material-change recall within five percentage points of the better comparator and at least 20% lower median review time. Stop for any protected-signal omission or privacy breach, systematic audit misses, excessive reconstruction failure, increased causal overclaiming, failed fallback, or no reduction in total review-plus-maintenance effort.
Deployment and cost¶
The first evidence study is estimated at \(50,000–\)250,000. Startup is \(250,000–\)1 million; operational launch is \(1–\)5 million, with \(250,000–\)1 million recurring annually. These are rough 2026 USD resource-equivalent bands, not procurement quotes, and local data integration, privacy, oversight, and audit requirements remain unpriced.
Risks and uncertainties¶
- The model could normalize repeated over-enforcement, under-service, or rights harm as expected program behavior.
- Police-generated records could dominate independent sources and create a self-confirming baseline.
- Evaluators or leaders could treat aggregate residuals as causal findings despite confounding and spillovers.
- Schedule, boundary, reporting, or threshold changes could be used to make visible residuals disappear.
- Sparse audits, uneven independent-source quality, or minimum cell sizes could hide rare harms or create geographic disparities in uncertainty.
Expert review¶
Useful reviewer backgrounds: Independent crime-program evaluator, Criminologist specializing in place-based interventions, Civil-rights and community-oversight representative, Government privacy and de-identification specialist, Agency data steward.
- Can a frozen schedule-conditioned model distinguish predictable recording changes from external events and omitted context?
- Which force, injury, complaint, disparity, missingness, and oversight signals must bypass filtering in full?
- Are independent data sources sufficiently complete and comparable across all evaluated places?
- Do residual displays increase unsupported causal interpretations relative to conventional dashboards?
- Does total evaluator, audit, calibration, and maintenance time fall while material-change recall remains within tolerance?
Evidence and provenance¶
Selected sources: S1: The Nation’s Two Crime Measures, 2015–2024 · S2: Measurement Error in Calls-For-Service as an Indicator of Crime · S3: Spatial Displacement and Diffusion of Benefits Among Geographically-Focused Policing Initiatives · S4: Assessing Responses to Problems: An Introductory Guide for Police Problem-Solvers · S5: An Ex Post Facto Evaluation Framework for Place-Based Police Interventions · S6: Smart Policing Initiative: Overview · S7: NIST SP 800-188: De-Identifying Government Datasets—Techniques and Governance · S8: Conduct of Law Enforcement Agencies
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized assessment places this candidate 40–53 across profiles, in band D. It is a post-hoc reading-order aid using an affordability proxy, not an experimental endpoint, economic-value measure, field-validation result, or evidence that an agency will adopt it.
49. Making Carbon-Budget Boundaries Visible¶
Canonical title: Open-Boundary Assembly for Regional Carbon-Budget Synthesis
In one sentence: At each regional carbon-budget release freeze, teams would jointly expose mismatched boundaries and unresolved residuals, then carry agreed disclosures into ordinary publication controls.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-15 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Environmental Climate |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 60.0/100 · rank range 45–53 across three profiles · band D |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-28, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Regional carbon budgets combine atmospheric, forest, soil, aquatic, and land-use estimates that may cover different places, periods, or definitions. Although specialists document limitations, those details can remain scattered in appendices and team files. A polished balance table can then look more complete than it is: exclusions disappear from headlines, residual quantities lack an owner or explanation, and contributors give different accounts of what the published total includes and leaves unresolved.
What is proposed¶
After normal technical reconciliation but before release approval, the consortium would hold a voluntary 50-minute Open-Boundary Assembly. Each component team would display a removable layer showing its spatial mask and time window, then state what its estimate includes, excludes, and cannot resolve. The response “heard, not resolved” would acknowledge the statement without implying agreement. A residual marker would pass among interface stewards, who could flag a mismatch, assign it for examination, declare it irreducible for that release, or pass. Contributors could accept, revise, transfer, or decline responsibility for carrying limitations into particular outputs. The editor—not the assembly—would then update the versioned boundary ledger, claims, graphics, metadata, owners, deadlines, and open questions through the existing publication system.
The cross-domain transfer¶
The ritualized-commitment archetype becomes a marked, repeatable scientific checkpoint. Visible layers symbolize that the regional total is assembled from partial views; standardized statements make limitations collectively audible; witnessed handoffs renew responsibility. The mapping is structurally strong, but its claimed benefit depends on ritual features adding something beyond equally careful facilitation and documentation.
Why it advanced¶
This candidate passed Experiment 6’s strict researched-candidate bar because it defines a bounded problem, preserves scientific authority, specifies a close comparator, provides falsifiers and safety controls, and proposes a measurable randomized pilot. STRICT_SUCCESS does not mean the assembly has worked in practice, is novel worldwide, is authorized for deployment, or produces economic value.
Prior art and the remaining open claim¶
The parts have substantial adjacent prior art: carbon accounting already uses QA/QC, completeness and double-counting checks, formal review, uncertainty tables, facilitated workshops, shared records, and version control. Workplace research also suggests rituals can increase perceived meaning, but not carbon-budget accuracy. The remaining claim is narrower: adding this voluntary ritual layer to an otherwise identical 50-minute boundary-review checklist improves detection, delayed recall, and fulfillment of disclosure duties without creating coercion, false assent, confidentiality failures, or authority confusion.
Smallest decisive test¶
Preregister a remote crossover study with 24–40 carbon-cycle or adjacent researchers in four to six balanced teams. Teams would review matched synthetic packets containing hidden boundary, stock-flow, covariance, exclusion, and residual defects, using either the assembly or an equal-duration facilitated checklist with the same facts and ledger. Advance only for at least a 20-percentage-point improvement in detection or delayed recall, or two substantively new correct disclosures per team, with no material loss of mock-output accuracy and no credible safety incident. Equivalent performance or any safety breach falsifies the incremental claim.
Deployment and cost¶
The first study and startup are each estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms. Operational launch is also \(10,000–\)50,000, while recurring annual operation is estimated at \(50,000–\)250,000. No consortium has committed authority, staff time, or funding; live evaluation would also require access to versioned synthesis artifacts.
Risks and uncertainties¶
- Transparent layers may hide nonlinear processes, covariance, scale differences, or nonspatial boundaries.
- “Heard, not resolved” may become rote or be mistaken for scientific agreement.
- Marker circulation may pressure participants to explain quantities outside their expertise.
- Publicly witnessed dissent could expose junior contributors or minority interpretations to retaliation.
- The ceremony could make an unchanged or misleading synthesis appear unusually coherent and legitimate despite unresolved evidence gaps.
Expert review¶
Useful reviewer backgrounds: Regional carbon-budget scientist, Carbon-accounting and uncertainty specialist, Environmental synthesis editor or data steward, Research-team governance and facilitation specialist, Accessibility, consent, and confidentiality reviewer.
- Can existing release records establish that documented boundary limitations actually disappear from headline tables, graphics, or summaries?
- Would the planted defects and scoring rubric represent consequential regional-budget interface failures rather than merely easy checklist items?
- Can the checklist and assembly arms be matched for facts, facilitation quality, documentation, and time?
- What safety threshold should govern reports of pressure, false assent, confidentiality loss, or authority confusion?
- Which live or retrospectively versioned artifacts could measure whether disclosure obligations were ultimately fulfilled?
Evidence and provenance¶
Selected sources: S1: Global Carbon Budget 2024 · S2: RECCAP2—REgional Carbon Cycle Assessment and Processes 2: Overview · S3: 2006 IPCC Guidelines for National Greenhouse Gas Inventories, Volume 1, Chapter 6: Quality Assurance/Quality Control and Verification · S4: IPCC Procedures · S5: Developing Reproducible Workflows Collaboratively · S6: The Science and Practice of Team Science: Reimagining Collaboration in a Changing Research Landscape—Consensus Study Highlights · S7: Work Group Rituals Enhance the Meaning of Work · S8: Resources for Working Groups
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review is only a post-hoc reading aid: scores range from 57 to 63 across profiles, with ranks 45–53 and ordering band D. Its speed input is an affordability proxy, not measured elapsed time or economic value, and it does not alter STRICT_SUCCESS status.
50. Reserve AI Capacity for Harm Inquiry¶
Canonical title: Protected nondeployment reserve in annual AI pilot portfolios
In one sentence: Organizations would protect at least 15% of annual AI-pilot capacity for nondeployment investigations before live-use projects consume the portfolio.
| Field | Record |
|---|---|
| Portfolio ID | EXP04-STRICT-06 |
| Experiment and endpoint | Experiment 4 · Strict success |
| Archetype × domain | Negative Space Design × Tech Ethics Ai Governance |
| Proposal position or arm | RETRIEVAL_FIRST |
| Post-hoc reading order | Balanced score 59.0/100 · rank range 49–50 across three profiles · band D |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
When an organization allocates its annual AI-pilot portfolio, live-use proposals can consume all available capacity before individual approvals occur. Red-teaming, shadow evaluation, and affected-community inquiry must then compete for whatever remains. Material problems may consequently emerge only after people are exposed and projects have accumulated operational dependencies. The underlying prevalence is unknown, however, and existing risk-tiered assurance may already provide adequate predeployment coverage without a fixed reserve.
What is proposed¶
Before approving individual pilots, the portfolio governing body would reserve at least 15% of total annual pilot capacity exclusively for nondeployment harm inquiry. The protected share could fund red-teaming, shadow evaluation, or appropriately safeguarded affected-community work, but not live deployment. Separate accounting would prevent project teams from informally absorbing it. Reallocation would require the governing body to document that eligible inquiry demand was exhausted, obtain concurrence from an independent assurance function, record the destination and reasons, and preserve enough capacity to complete active inquiries. Mandatory privacy, security, accessibility, safety, and incident-response work would remain outside the contest for this reserve. Existing approval and risk-management processes would continue to control deployment decisions.
The cross-domain transfer¶
Negative-space design is instantiated as deliberately unused deployment capacity: a protected portfolio “void” is bounded before surrounding projects are selected. Its positive purpose is to preserve room for inquiry while changes remain feasible. The structural transfer is clear, although percentage ring-fencing and controlled release already exist in adjacent evaluation and innovation portfolios.
Why it advanced¶
This candidate passed Experiment 4’s strict researched-candidate bar by defining the denominator, protected use, release rule, comparator, measurable outcomes, and a records-only first step. STRICT_SUCCESS is limited to that research screen; it is not evidence that 15% is optimal, that harm detection improves, or that an organization will adopt the rule.
Prior art and the remaining open claim¶
Adjacent systems already allocate resources to AI testing, independent assurance, governance boards, sandboxes, recurring evaluations, and risk-tiered review. Other fields also ring-fence evaluation funding or divide portfolios by fixed percentages. The remaining contrastive claim is specifically that a preapproval 15% nondeployment reserve, protected from live-use commitments and released only with documented independent concurrence, increases independently adjudicated material issues found per proposed system without lowering the number of validated pilots by more than 10%.
Smallest decisive test¶
With one willing organization, preregister a records-only reconstruction of a completed annual portfolio. Compare the actual flexible or risk-tiered allocation with a 15% protected shadow allocation applied before proposal selection. Define capacity, eligible inquiry, materiality, and missing-data rules in advance; use two issue reviewers, including one independent of pilot selection. Measure material issues per proposed system, validated-pilot count, inquiry completion, and simulated compliance with release rules. Do not advance if issue detection does not increase, validated pilots fall by more than 10%, or capacity and eligible work cannot be measured consistently.
Deployment and cost¶
First evidence and initial startup are each estimated at \(50,000–\)250,000. Operational launch and annual recurring costs are each estimated at \(250,000–\)1 million, in rough 2026 resource-equivalent bands rather than vendor quotes. Actual cost depends heavily on specialist red teams, community participation, compute, data preparation, and displaced pilot capacity.
Risks and uncertainties¶
- The 15% threshold may be too large, too small, or meaningless for very small portfolios.
- Teams may relabel ordinary development or compliance work as protected inquiry.
- Issue counts may reward numerous trivial findings unless independent reviewers apply a reproducible materiality standard.
- A fixed reserve could sit unused while beneficial pilots wait, or encourage wasteful spending merely to exhaust it.
- Observational results may confuse the reserve’s effect with an organization’s pre-existing safety culture and staffing quality.
Expert review¶
Useful reviewer backgrounds: AI portfolio-governance leader, Independent AI assurance or audit specialist, Causal-inference and program-evaluation researcher, Affected-community research and safeguarding specialist, Organizational finance and capacity-accounting expert.
- Can pilot capacity be expressed in a consistent unit across projects, staff, compute, and external assurance?
- Which work qualifies as nondeployment harm inquiry rather than normal development or mandatory compliance?
- Can historical records establish when an issue was found, whether it was material, and whether it changed design or approval?
- Is flexible risk-tiered assurance a more credible comparator than the organization’s actual historical allocation alone?
- How should continuous deployment, procurement, model updates, and small portfolios be handled outside the annual denominator?
Evidence and provenance¶
Selected sources: S1: AI RMF Core · S2: Industry temperature check: barriers and enablers to AI assurance · S3: OMB Memorandum M-25-21: Accelerating Federal Use of AI through Innovation, Governance, and Public Trust · S4: Costed Evaluation Plan Guidance, Tools and Templates · S5: Innovation portfolios for public sector organizations · S6: Anthropic’s Responsible Scaling Policy · S7: National Employment and Wage Data by Occupation, May 2025 · S8: Guidance to Set Up Your Organization's AI Governance Process
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading aid scores this candidate 58–60 across profiles, ranks it 49–50, and places it in band D. The pilot-speed input reflects cost-band affordability, not observed duration. These figures neither measure economic value nor replace its STRICT_SUCCESS endpoint.
51. Audited Competition for Close-Review Priority¶
Canonical title: Verified Close-Readiness Arena for Scarce Consolidation Review Slots
In one sentence: Business units would compete for scarce early financial-close review slots using independently verified readiness evidence rather than self-declared completion or managerial escalation.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-01 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Accounting Auditing |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 59.0/100 · rank range 47–54 across three profiles · band D |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-02, EXP06-PARTNER-03, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
During a multi-entity financial close, business units may compete for a limited number of early consolidation or technical-accounting reviews. If self-certified readiness or managerial escalation controls the queue, a unit can gain priority by prematurely closing reconciliations, shifting exceptions to another entity, deferring unsupported items, or obtaining privileged reviewer access. Reviewers then reopen supposedly complete work while other units wait. Crucially, no external evidence yet shows that this scarcity or strategic behavior exists in a target organization.
What is proposed¶
The corporate controller would replace the informal queue race with a recurring, bounded readiness contest. Units meeting a minimum control floor could compete for several early review slots. A rulebook fixed before scoring would allow genuine documentation, automation, supported resolution, and timely escalation, while prohibiting concealment, unsupported deferral, exception dumping, shared answers, off-channel influence, and undisclosed borrowed labor. An independent verifier would score hidden, risk-stratified samples for first-pass evidentiary completeness, supported treatment of aged exceptions, intercompany agreement, and absence of later unsupported corrections. Timely disclosure of material issues would have a protected route. Winners would be audited, serious errors appealable, extra contest labor capped, and some awarded capacity retained for remediation. Priority would expire after one close.
The cross-domain transfer¶
Bounded-rivalry governance becomes a formal contest for a scarce operational prize. Eligibility, lawful tactics, fouls, judging, appeals, resource caps, multiple awards, spillover responsibility, and recurring challenger access constrain how units compete. The mapping is detailed, but it remains unproven that queue positions create meaningful strategic interdependence rather than reflecting ordinary systems, complexity, or staffing constraints.
Why it advanced¶
This candidate did not enter Experiment 6’s strict-success lane. It cleared the separately calibrated EMPIRICAL_PARTNER_CANDIDATE lane because a bounded retrospective partner study is feasible and decision-relevant. Its advancement depends on obtaining proprietary field records; there is currently no evidence of local queue scarcity, manipulation, predictive advantage, or safe behavioral response.
Prior art and the remaining open claim¶
Financial-close platforms already provide workflows, dependencies, approvals, dashboards, audit trails, exception monitoring, and risk-based administrative triage. Accounting standards also require evidence, objective verification, and controls over period-end reporting. The narrower open claim, conditional on consequential rivalry being demonstrated, is that hidden-sample, independently verified competitive ranking predicts less first-pass rework, fewer reopened or transferred exceptions, and fewer unsupported corrections than both current queueing and noncompetitive risk-based triage, without delaying protected disclosure of material issues.
Smallest decisive test¶
Preregister a retrospective replay of one completed close across four to eight entities and 40–80 risk-stratified reconciliations. Compare recorded queue order, blinded noncompetitive risk-based triage, and the proposed arena score. Measure verified completeness, rework hours, reopened exceptions, transferred mismatches, unsupported correcting entries, and disclosure timing. Report stability under bootstrap samples and reasonable weight changes. Do not proceed beyond a no-consequence simulation if the arena fails to outperform triage out of sample, Kendall rank stability is below 0.60, size or complexity drives results, data are inconsistent, or issue reporting appears delayed or suppressed.
Deployment and cost¶
A retrospective first study is estimated at \(10,000–\)50,000. Initial startup is \(50,000–\)250,000; operational launch and annual recurring operation are each \(250,000–\)1 million in rough 2026 resource-equivalent terms. Software, data integration, independent verification, appeals, labor-cap auditing, and recurring governance still require organization-specific estimates, not vendor assumptions.
Risks and uncertainties¶
- Teams may game evidence fields that hidden samples do not cover.
- Risk adjustment may embed incumbent complexity assumptions and disadvantage legitimate challengers.
- Labor caps may be evaded through off-books work or restrict units with genuine remediation needs.
- Coordination screens may falsely flag common deadlines or shared system failures as collusion.
- Winning units may accumulate reviewer relationships and procedural knowledge even though formal priority expires.
Expert review¶
Useful reviewer backgrounds: Corporate controller or consolidation leader, Internal-control and ICFR specialist, Internal auditor independent of close operations, Financial-close data and workflow engineer, Employment, legal, and external-audit governance adviser.
- Are early specialist-review slots genuinely scarce, and does their timing materially affect other entities’ outcomes?
- Can queue, reconciliation, exception, rework, escalation, and correction records be joined reliably across entities?
- Does the proposed score outperform administrative risk triage on data not used to construct the ranking?
- How can protected material-issue reporting be separated from competitive scoring and monitored for delay?
- Would labor caps, hidden sampling, and anomaly screening conflict with employment rules, ICFR responsibilities, or external-audit arrangements?
Evidence and provenance¶
Selected sources: S1: Can fine-tuning your financial processes help accelerate your growth? · S2: Stepping into the future of controllership · S3: AS 2201: An Audit of Internal Control Over Financial Reporting That Is Integrated with an Audit of Financial Statements · S4: Financial Close Management Software · S5: Approve Closing Tasks · S6: Executive tournament incentives and audit fees · S7: Commission Guidance Regarding Management’s Report on Internal Control Over Financial Reporting Under Section 13(a) or 15(d) of the Securities Exchange Act of 1934 · S8: Occupational Employment and Wages — May 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading aid gives scores of 56–62, ranks 47–54, and band D. Its speed input is a cost-affordability proxy rather than elapsed time. This ordering is not an experimental endpoint or economic-value estimate and does not upgrade the partner-candidate status.
52. Renewing Climate-Monitoring Responsibilities¶
Canonical title: Season-Turn Stewardship Muster for Long-Term Climate Monitoring
In one sentence: Before each field season, monitoring staff would visibly renew or transfer station-to-archive duties and then record every accepted obligation in the ordinary task system.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-28 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Environmental Climate |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 59.0/100 · rank range 44–58 across three profiles · band D |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-15, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Long-term climate monitoring depends on seasonal sampling, calibration, metadata, custody transfers, and reliable handoffs despite staff turnover. Written protocols may leave particular responsibilities privately understood or attached to departed personnel. A field season can therefore begin without a named primary, backup, needed resource, credential, or escalation route for every station-to-archive link. Missing duties or tacit knowledge may surface only when work is due, creating gaps or ambiguities in a record intended to remain comparable over decades.
What is proposed¶
After the normal technical readiness review and before each field season, the network would hold a voluntary 35-minute Season-Turn Stewardship Muster. A neutral signal would mark the occasion, and participants would trace a fictional or real observation from station to archive by moving plain link cards across a network map. At each step, the current steward could renew, revise, transfer, or decline responsibility without explaining publicly. A backup and missing resources would be named before acceptance. Witnesses could acknowledge completed handoffs, but attendance, speech, silence, or card handling would not create consent or authority. Closure would occur only when ordinary managers enter an owner, backup, resources, due date, and escalation route into the existing task system. Debrief and independent review could modify, pause, or retire the practice.
The cross-domain transfer¶
The ritual archetype becomes a marked seasonal enactment of shared stewardship. Card movement makes the custody chain visible; voluntary renewal prevents stale assignments from persisting silently; witnessed handoffs support memory across turnover. The mapping is plausible, but most operational content duplicates ordinary responsibility maps and structured handoffs, leaving only the ritual features’ incremental social-memory effect to test.
Why it advanced¶
This proposal did not enter the strict-success lane. It cleared the EMPIRICAL_PARTNER_CANDIDATE lane because a small, fictional-record crossover with an external monitoring partner could test the remaining contrast safely. Field evidence is missing on problem prevalence, adopter demand, improved recall, dependency discovery, operational completion, voluntariness, and accessibility.
Prior art and the remaining open claim¶
Climate networks already require sustained operations, calibration, metadata, annual maintenance, anomaly tracking, preseason reviews, assigned roles, and configuration control. Structured handoffs also cover ownership, acknowledgment, next steps, and clarification; workplace ritual studies measured meaning, not monitoring continuity. The remaining claim is that adding a neutral threshold, voluntary card traversal, and witnessed renewal to an equal-duration administrative session improves seven-day recall or surfaces more valid dependencies without increasing pressure, exclusion, religious conflict, or confusion about formal authority.
Smallest decisive test¶
With one willing network, preregister a crossover using four matched fictional station-to-archive records and 12–24 participants. Compare a 35-minute administrative readiness session with the same session plus the ritual features, crossing teams onto a different record. Measure valid actionable dependencies and blinded immediate and seven-day recall of each primary, backup, and escalation route. Falsify incremental promise if the ritual finds no additional valid dependency and improves complete-link recall by less than 15 percentage points, or worsens pressure or access ratings by at least 0.5 on a five-point scale. Any credible coercion, privacy, cultural, or authority incident requires a halt.
Deployment and cost¶
First evidence is estimated below $10,000. Initial startup, operational launch, and annual recurring operation are each estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent bands. Actual costs remain uncertain because network size, travel, participant count, union requirements, accessibility work, facilitation, and task-system integration have not been specified.
Risks and uncertainties¶
- Employment hierarchy may make a formally optional refusal feel professionally costly.
- The event may aestheticize stewardship while staffing, equipment, training, or travel shortages remain unfunded.
- Public handoffs may expose disability, location, performance, or employment information.
- The linear card metaphor may distort parallel, contested, or nonlinear observation pathways.
- Neutral-looking symbolism may still create religious or cultural conflict, especially if local designers add unauthorized ceremonial elements.
Expert review¶
Useful reviewer backgrounds: Climate-monitoring network operator, Field-to-archive data and metadata steward, Human-factors and structured-handoff researcher, Workplace accessibility and religious-accommodation specialist, Program evaluator experienced in crossover trials.
- Do ordinary preseason records actually contain missing owners, backups, resources, credentials, or escalation routes?
- Are the fictional station-to-archive cases realistic enough to test consequential dependencies without exposing operational data?
- Can the administrative comparator match all content, time, facilitation, and task closure except the ritual features?
- How will anonymous measures detect pressure when managers and subordinates participate together?
- What result would justify a prospective no-consequence simulation before any real responsibility transfer is considered?
Evidence and provenance¶
Selected sources: S1: GCOS Surface Reference Network (GSRN): Justification, Requirements, Siting and Instrumentation Options (GCOS-226) · S2: U.S. Climate Reference Network (USCRN) · S3: U.S. Climate Reference Network Data Management Plan · S4: Fire Ecology Monitoring Protocol for the Heartland Inventory and Monitoring Network · S5: Handoff · S6: Work Group Rituals Enhance the Meaning of Work · S7: Section 12: Religious Discrimination · S8: Employer Costs for Employee Compensation—December 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading aid scores this candidate 55–63, with ranks 44–58 and band D. Its speed measure is a cost-affordability proxy, not observed duration. The wide profile range is not economic value evidence and leaves the EMPIRICAL_PARTNER_CANDIDATE endpoint unchanged.
53. Recurring Stewardship for Digital Legacies¶
Canonical title: Digital Memory Stewardship Observance for Algorithmic Afterlives
In one sentence: A governed observance would periodically turn remembrance into authorized, verifiable decisions about access, visibility, retention, resurfacing, correction, and successor responsibility for a digital legacy.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-29 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Human Computer Interaction |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 59.0/100 · rank range 41–57 across three profiles · band D |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
After death or lasting incapacity, a person’s posts, messages, images, inferred attributes, and scheduled reminders may persist under fragmented platform settings. Their recorded wishes may be incomplete, appointed account stewards may leave, and relatives or correspondents may disagree about what should remain visible. A one-time memorialization decision cannot necessarily address later platform changes, automated resurfacing, new audiences, disputed material, or the transfer of practical responsibility to a new steward.
What is proposed¶
At memorial creation, a chosen recurring date, a steward handoff, or a material platform change, authorized participants would hold a Digital Memory Stewardship Observance. A prior review would determine who may participate, submit private input, withhold material involving them, or decline. During the session, recommendations, autoplay, metrics, and nonessential notifications would be paused where technically possible. Participants would review consented, provenance-labeled materials without requiring shared grief or one approved life story. Authorized stewards would then renew, revise, decline, or transfer duties covering access, visibility, retention, resurfacing, moderation, and correction. Closure would create authenticated platform requests, named owners, deadlines, access changes, and an unresolved-matters register, followed by debrief, audit, repair, and retirement options.
The cross-domain transfer¶
The ritualized meaning-and-commitment archetype is transferred through a marked pause in ordinary algorithmic circulation, a recognizable sequence, consented witnessing, and recurring renewal of duties. The mapping is structurally strong at the governance level, but the symbolic elements have not been shown to improve stewardship beyond a well-run checklist using identical platform controls.
Why it advanced¶
This EMPIRICAL_PARTNER_CANDIDATE cleared a separately calibrated lane for a bounded external-partner study, not the strict-success lane. It advanced because the lifecycle problem, platform authorities, adjacent practices, safety controls, and falsifiable comparison are identifiable. Crucially, there is no field evidence that bereaved people would find the observance acceptable, safe, accessible, or useful.
Prior art and the remaining open claim¶
Adjacent prior art includes platform memorialization, legacy contacts, inactivity plans, estate administration, grief gatherings, and collaboratively edited memorials. These already cover many individual components. The narrower open claim is that adding a recurring, consent-aware observance—with an algorithmic pause, plural witnessing, explicit duty renewal, and later audit—to the same controls will produce more correctly authorized and verified actions and reveal more unresolved conflicts than an equal-time static directive-and-action checklist, without causing more privacy breaches, coercion, or distress.
Smallest decisive test¶
A digital-legacy professional body or university HCI laboratory would recruit 12–24 living-volunteer dyads using synthetic or participant-controlled redacted account replicas. Dyads would receive either the complete observance or an equal-time checklist with identical mock controls and authority briefing. Independent reviewers would assess correct authority assignment, verified action completion, detected conflicts, unauthorized access or disclosure, handoff success, distress, coercion, unwanted exposure, accessibility failures, and facilitator halts. Proceed only if the observance materially improves verified correct actions or conflict detection without worse safety or completion outcomes; otherwise adapt, hold, or retire it. This would not demonstrate benefit during bereavement.
Deployment and cost¶
A sandboxed rehearsal and manual action ledger appear feasible; production use depends on platform-specific authentication, recommendation controls, verification, rollback, law, and trained facilitation. Rough 2026 USD resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for initial startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. These are not vendor quotes.
Risks and uncertainties¶
- A powerful family or community faction could turn the memorial into one authorized-looking account of the person’s life.
- Private messages, images, or relationship patterns could be exposed to people who lack standing to see them.
- Survivors could mistake ceremonial agreement for authority over recorded directives or another living person’s data.
- Visible settings may not match the platform’s actual recommendation and resurfacing behavior.
- Recurring dates or notifications could impose remembrance on people who want to disengage permanently or temporarily without explanation.
Expert review¶
Useful reviewer backgrounds: Digital-legacy and estate-planning specialist, Human-computer interaction researcher specializing in death and bereavement, Privacy and fiduciary-access lawyer, Platform trust, safety, and account-support engineer, Grief-informed facilitator or clinical safety reviewer.
- Which proposed decisions belong respectively to the original account holder, a fiduciary, a depicted or corresponding person, and the platform?
- Can the mock platform reliably verify resurfacing behavior and reverse every tested setting or access change?
- Which symbolic elements add information or accountability that the equal-time checklist does not?
- What stopping thresholds for distress, coercion, unwanted exposure, and disputed authority should block continuation?
- Can the study recruit contested or culturally varied dyads without treating participation as consent to memorialization?
Evidence and provenance¶
Selected sources: s1: About Memorialized Accounts · s2: How to add a Legacy Contact for your Apple Account · s3: About Inactive Account Manager · s4: Current Acts—F: Fiduciary Access to Digital Assets Act, Revised · s5: Digital Legacy Association—Home · s6: Engaging with Death Online: An Analysis of Systems that Support Legacy-Making, Bereavement, and Remembrance · s7: Digital Legacy: A Systematic Literature Review · s8: From Personal Data to Digital Legacy: Exploring Conflicts in the Sharing, Security and Privacy of Post-mortem Data
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review is only a post-hoc reading order, not an experimental endpoint or measure of economic value. It placed this candidate between ranks 41 and 57 in band D; the cost input approximated affordability, not elapsed pilot time.
54. Testing Audit Findings Before Awarding Leads¶
Canonical title: Replicable Assurance Challenge for Scarce Audit-Lead Mandates
In one sentence: A separate internal-audit challenge would award temporary lead roles and investigative hours according to independently reproduced risk findings rather than visible finding volume or severity alone.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-02 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Accounting Auditing |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 58.0/100 · rank range 52–54 across three profiles · band D |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-01, EXP06-PARTNER-03, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Internal-audit teams may compete for a small number of prominent engagement-lead roles and discretionary investigative hours. If managers informally reward finding counts, apparent severity, speed, or budget performance, auditors could gain an advantage by splitting one root problem into several findings, favoring easily scored issues, delaying referrals, withholding reusable tests, or pressing auditees toward harsher labels. The result could be a leadership pipeline that rewards persuasive issue production rather than reproducible assurance work.
What is proposed¶
Create a quarterly Replicable Assurance Challenge that remains separate from mandatory reporting and personnel evaluation. Teams would submit evidence from completed work to compete for two temporary lead mandates and capped investigative hours. A frozen rulebook would protect urgent, legal, fraud, whistleblower, and material-misstatement reporting while prohibiting duplicate counting, evidence withholding, retaliation, reciprocal scoring, unsupported severity changes, auditee pressure, and hidden labor. An identity-blinded panel would test leading hypotheses on held-out transactions and score reproducibility, severity support, root-cause coherence, added risk coverage, audit-trail quality, portability, and auditee burden. Awards would expire after one quarter; some hours would remain centrally held for correction. Appeals, coordination screens, later validation, and retirement rules would constrain gaming and entrenchment.
The cross-domain transfer¶
Bounded rivalry is instantiated as a deliberately separate contest for scarce, temporary mandates. Legitimate competition occurs through reproducible evidence under frozen rules; mandatory assurance communication remains outside the arena. Independent re-performance, labor caps, anti-collusion screens, burden scoring, expiring prizes, and later review are intended to keep rivalry tied to assurance quality.
Why it advanced¶
This EMPIRICAL_PARTNER_CANDIDATE entered the bounded data-partner lane, not strict success. It offers a testable shadow study using existing audit records and established quality-review authority. However, no inspected records show that lead roles are truly scarce, that issue production affects selection, or that the proposed strategic behavior occurs in operating audit functions.
Prior art and the remaining open claim¶
Quality-assurance programs, ordinary engagement review, root-cause analysis, consolidated issue tracking, rotations, and risk-adjusted scorecards already address much of the problem. The open claim is conditional: only where genuine scarcity and strategic interdependence exist, blinded re-performance should predict later finding validity and incremental risk coverage better than ordinary quality review or a simpler scorecard, without delaying reporting, increasing defensive documentation, exposing confidential material, or adding auditee burden. This is not a world-novelty claim.
Smallest decisive test¶
Preregister a retrospective shadow replay of six closed engagements, twelve issued or merged findings, and six documented but unissued hypotheses. Three independent quality reviewers would use preserved records and held-out transactions without contacting auditees, publishing ranks, or allocating mandates. Compare the proposed composite with existing QAIP outcomes and a simpler risk-adjusted scorecard. Measure inter-rater reliability, rank stability, reviewer hours, blinding, later finding status, added coverage, reconstructed burden, and strategic markers. Do not advance if reliability is below 0.70, the method adds no discrimination, engagement identity or writing style dominates, confidentiality or blinding fails, or prospective use could inhibit candid reporting.
Deployment and cost¶
Existing audit repositories could support a retrospective replay, but live use would require reliable blinding, risk normalization, protected workpaper access, independent reviewers, and safeguards against informal career use. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(250,000–\)1 million annually; they are not quotes.
Risks and uncertainties¶
- Competition could compromise, or appear to compromise, auditor independence.
- A reproducibility metric could disadvantage emerging or systemic risks that cannot yet be repeated across transactions or locations.
- Teams could optimize documentation for the review panel while moving preparation labor off the recorded budget.
- Audit subjects and methods may reveal team identity despite formal blinding.
- Private standings could still affect promotion or assignment decisions if confidentiality fails.
Expert review¶
Useful reviewer backgrounds: Chief audit executive, Independent internal-audit quality-assurance reviewer, Audit-committee or board governance specialist, Employment, privilege, and workpaper-access counsel, Audit-methodology and measurement researcher.
- Do historical assignment records show scarce lead opportunities and a relationship between visible issue production and selection?
- Can reviewers reconstruct nonissued hypotheses and held-out tests from authorized records without new auditee requests?
- Does the composite remain reliable after controlling for engagement risk, scope, specialization, and workpaper style?
- Would auditors alter urgent reporting or create defensive documentation if a live challenge were introduced?
- Can shadow results be technically and institutionally prevented from entering personnel decisions?
Evidence and provenance¶
Selected sources: S1: Global Internal Audit Standards, 2024 Edition · S2: Quality Assurance and Improvement Program (QAIP) · S3: Risk: The Root of the Matter · S4: Incentives for Dishonesty: An Experimental Study with Internal Auditors · S5: Is the Objectivity of Internal Audit Compromised When the Internal Audit Function Is a Management Training Ground? · S6: TeamMate Audit Management · S7: Connected Risk Quick Start Guide for Internal Audit and Controls Leaders · S8: Accountants and Auditors: Occupational Outlook Handbook
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review is a post-hoc ordering aid, not an endpoint or economic-value estimate. This candidate fell between ranks 52 and 54 in band D. Its cost-based affordability proxy did not independently measure how quickly a pilot could run.
55. Full-Cost Bidding for Preservation Access¶
Canonical title: Externality-Adjusted Reverse Auction for Religious-Collection Preservation
In one sentence: Custodians would submit sealed subsidy bids for comparable preservation work, with lifecycle costs, cultural authority, material safety, and future technical dependence incorporated before scarce laboratory time is awarded.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-08 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Religious Studies Theology |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 58.0/100 · rank range 47–55 across three profiles · band D |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A library consortium may have fewer mobile-laboratory weeks and less subsidy than custodians request for fragile manuscripts, recordings, ritual objects, and community archives. In a narrative competition, applicants could understate preparation, rights, metadata, storage, migration, or remediation costs while exaggerating quantity or urgency. Projects that appear inexpensive may therefore win by shifting work and risk to custodians, communities, or future repositories, while vulnerable materials wait and selected providers gain technical leverage over later rounds.
What is proposed¶
Divide annual laboratory capacity into narrow lots defined by material type, condition, and verified work units. Before bidding, independent reviewers would confirm custodial authority, cultural permissions, conservation safeguards, metadata, access restrictions, storage commitments, and technical feasibility. Eligible custodians would submit sealed reverse bids stating the minimum subsidy required for preparation, treatment, digitization, return copies, metadata, and a defined preservation period. A fixed schedule would add reserves for handling, migration, dependencies, and remediation, without penalizing culturally required restrictions. The lowest adjusted eligible bids would clear, subject to institutional concentration and shared-failure checks. Finalists would be audited; sponsors would secure remediation obligations; outputs would use transferable formats under custodian-controlled access; and later cost, damage, access, and concentration results would recalibrate future lots.
The cross-domain transfer¶
Bounded rivalry becomes a reverse auction for scarce preservation capacity. Price competition is permitted only within pre-authorized, technically comparable lots and above noncontestable cultural and safety floors. Sealed bids, fixed adjustments, audits, guarantees, concentration limits, transferable workflows, rebidding windows, and post-project review seek to prevent cost dumping, coordination, and lock-in.
Why it advanced¶
This EMPIRICAL_PARTNER_CANDIDATE cleared a bounded external data-study lane rather than strict success. Active preservation programs, technical standards, and an authorized retrospective design make the question researchable. The central missing evidence is institutional: no partner has supplied project records, approved even a simulation, or shown that stable comparable lots can be constructed.
Prior art and the remaining open claim¶
Expert grant panels, first-come allocation, collection-size formulas, eligibility lotteries, centralized preservation, and established reverse-auction procedures supply adjacent prior art. The narrower open claim is that, for genuinely comparable and pre-authorized projects, externality-adjusted subsidy bids will predict realized lifecycle cost and allocate fixed laboratory capacity more cost-effectively than those alternatives without increasing cultural exclusion, physical harm, change orders, technical concentration, or stranded outputs. The claim fails if discretionary adjustments overwhelm price or protected collections cannot fit comparable lots.
Smallest decisive test¶
With custodial and records-owner authorization, reconstruct proposed and realized lifecycle costs for at most ten completed projects in no more than three material-and-treatment categories. Two independent conservation-cost reviewers would apply preregistered fields, comparability thresholds, missing-data rules, and harm measures. Test whether adjusted ordering predicts realized cost better than requested budgets and actual narrative rankings, then simulate narrative, first-come, size-formula, and lottery allocations. Permit only a go/no-go decision for a nonbinding sealed-bid simulation. Stop if reviewer agreement is poor, over 25% of projects are incomparable, discretionary adjustments dominate, prediction does not materially improve, protected or under-resourced custodians are disproportionately excluded, or simulated safety or concentration worsens.
Deployment and cost¶
The administrative auction machinery is feasible only for narrow categories with stable units; cultural authority, rights, condition, insurance, storage, and legal characterization remain local. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(250,000–\)1 million annually. They are not vendor quotes.
Risks and uncertainties¶
- Standard work units could conceal collection-specific fragility, urgency, or treatment needs.
- Institutions with already subsidized infrastructure could bid below smaller community archives without being more efficient overall.
- Guarantee requirements could exclude under-resourced custodians even if pooled guarantees are nominally available.
- An adjustment schedule could wrongly treat culturally required access restrictions as costs or inefficiencies.
- Transferability requirements could conflict with community authority over restricted knowledge or materials.
Expert review¶
Useful reviewer backgrounds: Conservator experienced with the selected material categories, Community archive or religious-collection custodian, Cultural authority and Indigenous data-governance specialist, Preservation economist or auction-design researcher, Copyright, privacy, cultural-property, and procurement counsel.
- Which material-and-treatment categories can be normalized without obscuring condition, cultural protocol, or urgency?
- Can proposed and realized preparation, metadata, storage, migration, and remediation costs be reconstructed consistently from closed-project records?
- How should the design prevent infrastructure-rich institutions from appearing artificially inexpensive?
- Which cultural or custodial restrictions must remain outside every price and externality calculation?
- Would the process legally constitute grantmaking, procurement, or another allocation form in the partner’s jurisdiction?
Evidence and provenance¶
Selected sources: S1: Towards Sustainable Preservation and Accessibility of Documentary Heritage · S2: Preserving Endangered Cultural Memory at a Time of Heightened Risk: Evaluating the Recordings at Risk Grant Program · S3: Recordings at Risk · S4: Collections Stewardship · S5: Preservation and Selection for Digitization · S6: Technical Guidelines for Digitizing Cultural Heritage Materials · S7: The CARE Principles for Indigenous Data Governance · S8: Federal Acquisition Regulation Subpart 17.8—Reverse Auctions
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized assessment is only a post-hoc reading order, not an experimental endpoint or economic-value measure. It placed this candidate between ranks 47 and 55 in band D. The pilot input represented cost-band affordability rather than independently assessed elapsed time.
56. Reviewing What Changed in Student Reasoning¶
Canonical title: Predictive-Residual Formative Review for Interpretive Coursework
In one sentence: A shadow review system would predict a student’s next rubric-level performance, foreground the differences from that prediction, and preserve full human access, auditing, and fallback for every consequential interpretation.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-07 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Predictive Residual Processing × Religious Studies Theology |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 58.0/100 · rank range 54–56 across three profiles · band D |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-06, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Instructors reviewing repeated low-stakes religious-studies assignments must keep checking familiar skills while looking for new misconceptions, unsupported comparisons, and unexpectedly strong reasoning. Complete review preserves context but may consume feedback time confirming repeated mastery. Predictive filtering could focus attention, yet it could also lock students into earlier profiles, treat one tradition or writing style as normal, miss defensible alternative readings, or confuse personal belief with academic performance. The challenge is to save attention without weakening plural, contestable human judgment.
What is proposed¶
For one course and one recurring assignment family, a versioned model would use only a consenting student’s earlier work to predict rubric-level features in the next response before it is opened. It would predict taught academic skills—not wording, belief, doctrinal correctness, or grades. Human coding of the untouched submission would then identify signed differences such as improvement, omitted evidence, repeated conflation, changed qualification, unsupported transfer, a defensible alternative, or coding disagreement. A residual-first interface could collapse expected components, but the matching model and residuals must reconstruct the whole rubric profile, with the full submission one action away. Humans would validate every residual. Independent full marking, student challenges, version checks, error budgets, and mandatory bypasses for high-stakes, personal, accessibility-sensitive, culturally disputed, unfamiliar, or low-confidence work would trigger complete review when needed.
The cross-domain transfer¶
Predictive-residual processing is transferred directly: a versioned learner state predicts the next rubric profile, human coding supplies the observed profile, and signed residuals carry informative change to reviewers. Reconstruction, uncertainty thresholds, independent raw-response audits, version handshakes, and fallback preserve access to the full signal rather than making the residual queue authoritative.
Why it advanced¶
This STRICT_SUCCESS passed Experiment 6’s strict researched-candidate bar because it defines a bounded comparator, quantitative success and failure thresholds, human authority, safety bypasses, and a feasible shadow test. That endpoint is not real-world validation, a novelty finding, deployment authorization, evidence of improved learning, or proof that the workflow saves money in practice.
Prior art and the remaining open claim¶
Adjacent systems already include fixed rubrics, essay scoring, mastery dashboards, adaptive quizzes, answer grouping, work sampling, feedback templates, and knowledge tracing. The remaining claim concerns their narrower combination: predicting a learner-specific rubric profile before opening a held-out interpretive response, representing the response as reconstructable signed residuals, and auditing against full marking. Compared with complete manual review, it must reduce total review time by at least 20% while holding reconstruction disagreement to 5% or less and adding no missed defensible alternative or mandatory-bypass case.
Smallest decisive test¶
An authorized course partner would supply 48 de-identified responses from 12 students completing four sequential low-stakes assignments. Assignments 1–2 would create transparent predictions; the rubric, model, thresholds, bypasses, and analysis would be frozen before opening assignments 3–4. Counterbalanced reviewers would compare complete-response and residual-first review, while separate blinded graders would double-mark all 24 held-out responses. Success requires at least 20% lower median total review time, no more than 5% reconstruction disagreement, no additional missed defensible alternative or bypass case, no evident language or interpretive-position error pattern, and lower all-in effort after coding, validation, audit, challenge, maintenance, and fallback. Any failure falsifies or narrows the claim.
Deployment and cost¶
A small transparent shadow prototype is feasible, but live use would require privacy, accessibility, assessment-policy, procurement, and data-governance approval plus domain-literate coding and auditing. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(250,000–\)1 million for startup, \(50,000–\)250,000 for operational launch, and \(50,000–\)250,000 annually; these are not quotes.
Risks and uncertainties¶
- Earlier predictions could anchor reviewers and cause genuine improvement to be discounted.
- The rubric or learner model could systematically misread a less represented language, tradition, or argumentative style.
- A defensible interpretation could appear erroneous merely because it was not predicted from earlier work.
- Residual-focused feedback could omit the educational value of recognizing well-executed continuity.
- Learner-state records could expose distinctive errors or beliefs and create additional education-record privacy risk.
Expert review¶
Useful reviewer backgrounds: Religious-studies or theology instructor, Educational measurement and formative-assessment researcher, Learning-analytics or interpretable-model specialist, Accessibility and student-data governance officer, Tradition- and language-literate independent marker.
- Are two prior responses sufficiently predictive within the frozen task family to justify any collapsed review?
- Can independent graders reconstruct every rubric profile within the 5% disagreement limit?
- Which interpretations, languages, accommodations, or task types must always bypass residual-first review?
- Does total time still fall by 20% after coding, validation, expansion, auditing, challenges, maintenance, and fallback are included?
- Can error patterns by language or interpretive position be examined ethically in a sample this small?
Evidence and provenance¶
Selected sources: S1: An Empirical Analysis Exploring the Impact of Traditional Exams and Multi-Stage Assignments on Academic Workload in a Final Year Engineering Context · S2: EduMark AI: rethinking assessment and feedback with ethical AI · S3: A-level Religious Studies 7062: Scheme of assessment · S4: AI-assisted grading and answer groups · S5: Interpretable Knowledge Tracing: Simple and Efficient Student Modeling with Causal Relations · S6: Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring · S7: Family Educational Rights and Privacy Act regulations · S8: National employment and wage data by occupation, May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review is a post-hoc reading order, not the strict experimental endpoint or a measure of economic value. It placed the candidate between ranks 54 and 56 in band D; its affordability proxy did not independently score elapsed pilot time.
57. Testing Hidden Patterns in Library Exposure¶
Canonical title: Offline modal-control testing for coupled discovery-exposure imbalance
In one sentence: This post-hoc candidate proposes an offline test of whether coupled exposure patterns persist across digital-library updates and whether targeting those patterns would outperform ordinary monitoring and direct category constraints.
| Field | Record |
|---|---|
| Portfolio ID | EXP03-POSTHOC-01 |
| Experiment and endpoint | Experiment 3 · Post-hoc strict innovation-like survivor |
| Archetype × domain | Invariant Mode Decomposition Design × Library Information Science |
| Proposal position or arm | Not recorded |
| Post-hoc reading order | Balanced score 56.0/100 · rank range 49–58 across three profiles · band D |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
Digital libraries usually monitor engagement and exposure one item or category at a time. That can miss a combination of modest subject, format, branch, or creator-group imbalances that persists or grows through repeated ranking updates. Such a pattern could steadily narrow what patrons discover even while average engagement and every individual category measure appear acceptable. It is not yet known whether these cross-cycle patterns occur reproducibly in an actual library system.
What is proposed¶
Using deidentified records from eight ranking-update cycles, analysts would represent each cycle as exposure and engagement deviations across policy-approved resource groups. They would fit a restricted model on six cycles to estimate which combinations persist, decay, or grow, while controlling for catalog availability and query mix. Two sealed cycles would test whether this model predicts exposure better than ordinary category-by-category monitoring. Only then would offline replay compare carefully bounded, mode-targeted ranking adjustments with direct per-category constraints. Nothing would change for patrons. The proposed control would be rejected if it failed utility, privacy, representation, conditioning, residual-error, or stability limits.
The cross-domain transfer¶
The transferred archetype decomposes a changing system into combinations that evolve together. Here, those combinations are recurring mixtures of library-resource exposure deviations, and their gains describe whether they fade or grow between updates. The mapping is plausible but structurally weak because library discovery is nonlinear, partly human-driven, and represented by only a few observed transitions.
Why it advanced¶
This candidate was identified only after Experiment 3 through a stricter combined opportunity screen. It is a post-hoc survivor, not a preregistered experimental success. It advanced because library studies document exposure disparities and recommender feedback effects, while an offline, reversible test offers strong safeguards despite substantial data and identification gaps.
Prior art and the remaining open claim¶
The individual parts already have adjacent prior art: dynamic exposure controls, spectral analysis of recommender feedback, and data-driven mode decomposition with control all exist. The narrower unresolved claim is whether a low-rank operator over approved library-resource groups finds reproducible cross-cycle modes and whether targeting them produces better held-out exposure equity and retrieval utility than both category monitoring and direct category constraints. The search did not establish that library-specific combination or its added value.
Smallest decisive test¶
Preregister a zero-deployment retrospective study using eight consistently defined cycles. First calculate whether five training transitions contain enough information for the proposed state dimension; stop unless a justified low-rank restriction makes estimation credible. Fit category-lag and modal models on cycles one through six, then evaluate prediction, conditioning, residuals, alignment, and spectral separation on cycles seven and eight. Advance to offline control replay only if the modal model wins; reject the intervention if it fails to beat direct constraints without harming utility or represented groups.
Deployment and cost¶
The first evidence study is estimated at a rough 2026 resource-equivalent cost of \(50,000–\)250,000. Initial deployment and operational launch are each estimated at \(250,000–\)1 million, with \(50,000–\)250,000 annually. These are assessment bands, not vendor quotes. Deployment would also require logs, platform access, local metadata mapping, and accountable library review.
Risks and uncertainties¶
- Analysts could misdescribe statistical modes as traits or preferences of patrons or communities.
- The model could conceal missing or poorly cataloged resources by treating their absence as a ranking dynamic.
- With few transitions or a small spectral gap, modes could rotate or exchange identities under minor data changes.
- Improved modal exposure could come at the cost of retrieval relevance or create harm in an omitted group.
- Granular exposure strata could leak private information or lack enough observations for reliable estimation.
Expert review¶
Useful reviewer backgrounds: Library discovery and collection-governance specialist, Recommender-systems fairness researcher, Dynamic-systems or system-identification statistician, Privacy and representation reviewer, Library ranking-platform engineer.
- Can the available logs produce eight consistently defined cycles after catalog, query, seasonal, and policy changes are accounted for?
- What effective state dimension is estimable from five training transitions, and what low-rank restriction would be defensible?
- Do the fitted modes remain aligned under resampling, alternative group definitions, and both sealed cycles?
- Would direct per-category constraints provide equal or better equity and utility with less modeling risk?
- Which minimum-support and aggregation rules prevent privacy leakage and the reification of communities?
Evidence and provenance¶
Selected sources: S1: Feedback Loop and Bias Amplification in Recommender Systems · S2: A Study of Position Bias in Digital Library Recommender Systems · S3: Bias in Book Recommendation: A Case Study on the Danish Public Libraries · S4: Guidance on the Use of Artificial Intelligence in Libraries · S5: Fairness of Exposure in Dynamic Recommendation · S6: Deconvolving Feedback Loops in Recommender Systems · S7: Dynamic Mode Decomposition with Control · S8: On Dynamic Mode Decomposition: Theory and Applications
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed this candidate between ranks 49 and 58, in band D. That score is only a post-hoc reading order using a cost-based affordability proxy; it is neither an experimental endpoint nor evidence of economic value.
58. A Governed Contest Between Ethnographic Explanations¶
Canonical title: Bounded Rival Reanalysis of Ethnographic Explanations
In one sentence: This candidate would test, first with fictional data, whether tightly governed comparison of rival ethnographic explanations exposes cherry-picking and overreach without sacrificing context or unfairly creating a canonical winner.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-09 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Sociology Anthropology |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 56.0/100 · rank range 56–58 across three profiles · band D |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
When several research teams explain why neighborhood mutual-aid organizations survived or dissolved, scarce publication and funding slots can reward selective cases, privileged contextual help, larger teams, rhetorical confidence, or coordinated submissions. Ordinary peer review sees completed manuscripts but may not reveal how teams chose evidence, handled contradictory episodes, or influenced one another. A winning account can then be treated as authoritative, potentially narrowing later inquiry and misrepresenting the people whose records supplied the evidence.
What is proposed¶
Editors and an authorized data steward would run a secure, staged comparison under a rulebook fixed before analysis. Eligible teams would receive the same corpus, clarification archive, funded analyst hours, workspace period, and submission format. Each would preregister its main explanation, expected observations, scope, case-selection rule, and treatment of contrary evidence before viewing reserved cases. Submissions would map claims to episodes, alternatives, negative cases, and uncertainty. Independent reviewers would score several dimensions rather than one proxy, reproduce leading analyses, and audit misconduct indicators. Two complementary explanations could be commissioned, while follow-up funding would seek a discriminating observation rather than declare one cultural truth. Awards would remain nonexclusive and appealable.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a research contest with scarce commissions, equal resources, explicit legal moves, audits, sanctions, appeals, and limits on winner power. The structural mapping is detailed: competition is confined to comparable explanations of one bounded outcome, while participant identity, credibility, lived meaning, and representational authority remain outside the contest.
Why it advanced¶
It qualified only for the separately calibrated empirical-partner lane, not strict success. Existing practices show that secure qualitative-data review, preregistration, claim-to-evidence annotation, registered reports, and adversarial collaboration are feasible. However, no field study, adopter commitment, prevalence estimate, or evidence of the combined package’s comparative benefit exists.
Prior art and the remaining open claim¶
Nearly every epistemic component has adjacent prior art, including qualitative preregistration, transparent claim annotation, registered reports, controlled repositories, and empirical adversarial collaboration. The remaining claim concerns their governed combination: whether equal access, reserved cases, multidimensional scoring, leader audits, resource limits, and nonexclusive portfolio selection detect more planted cherry-picking and unsupported scope claims than ordinary review, without materially reducing contextual adequacy, increasing status sensitivity, or hardening one explanation into a canon.
Smallest decisive test¶
Preregister a synthetic trial using a fictional 18-organization corpus and eight engineered submissions. Randomly assign blinded panels to the bounded arena, ordinary peer review, or a plural symposium, then repeat after changing team names and status cues. Advance only if the arena detects at least 80% of planted cherry-picking or fabrication, keeps false misconduct referrals at or below 10%, maintains rank concordance of at least 0.80, preserves contextual adequacy within 0.30 standard deviations of the symposium, and improves negative-case visibility and scope calibration.
Deployment and cost¶
The synthetic first-evidence study carries a rough 2026 resource-equivalent estimate of \(50,000–\)250,000. Startup and operational launch are each estimated at \(250,000–\)1 million, with annual recurring costs also at \(250,000–\)1 million. These are not quotations and omit uncertain incident, litigation, and follow-up-grant costs.
Risks and uncertainties¶
- A common rubric may falsely treat distinct interpretive traditions as directly comparable.
- Reserved cases may reward simplified prediction and penalize legitimate contextual revision.
- Shared records and evidence maps may expose protected context or detach statements from relationships needed to interpret them.
- Resource limits may still favor teams with existing code, theory, staff, or tacit familiarity.
- Collusion screens may mistake common intellectual lineage or legitimate collaboration for misconduct signals requiring inquiry alone, not guilt findings.
Expert review¶
Useful reviewer backgrounds: Qualitative and ethnographic methods scholar, Research-integrity and journal-governance specialist, Qualitative-data repository steward, Human-subjects, privacy, or IRB specialist, Participant or source-community governance representative.
- Can reviewers reliably distinguish planted evidentiary weakness from legitimate differences between interpretive traditions?
- Do reserved cases measure explanatory adequacy, or do they systematically reward decontextualized prediction?
- Which materials can be shared, mapped, and reproduced without violating consent, cultural-access rules, or confidentiality?
- How should panel-composition sensitivity and false misconduct referrals affect the decision to stop?
- Would a plural symposium, registered report, or collaborative follow-up produce comparable integrity gains with less burden and canonization risk?
Evidence and provenance¶
Selected sources: S1: Fostering Integrity in Research · S2: Measuring the Prevalence of Questionable Research Practices With Incentives for Truth Telling · S3: ReShare Data Review Procedures · S4: Preregistering Qualitative Research · S5: A Guide to Annotation for Transparent Inquiry (ATI), Version 1.0 · S6: Registered Reports · S7: Coded Private Information or Biospecimens Used in Research, Guidance (2018) · S8: Rationale and Guidelines for Empirical Adversarial Collaboration: A Thinking & Reasoning Initiative
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed this candidate between ranks 56 and 58, in band D. This is a post-hoc reading aid, not an experimental endpoint or value estimate; its affordability input substitutes cost bands for an independently measured pilot timeline.
59. Mapping Coupled City Budget Deadlocks¶
Canonical title: Modal Deadlock Map for City Budget Bargaining
In one sentence: This candidate proposes a retrospective test of whether combinations of moderate budget disagreements predict bargaining trouble better than ordinary issue-by-issue tracking and can be linked to reversible procedural responses.
| Field | Record |
|---|---|
| Portfolio ID | EXP04-STRICT-03 |
| Experiment and endpoint | Experiment 4 · Strict success |
| Archetype × domain | Invariant Mode Decomposition Design × Political Science |
| Proposal position or arm | PROPOSAL_FIRST |
| Post-hoc reading order | Balanced score 51.0/100 · rank range 59–59 across three profiles · band D |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
City budget negotiators bargain over packages covering revenue, staffing, capital projects, debt, reserves, and district allocations. Officials usually track each disagreement separately, although movement on one issue can change positions on several others. A combination of moderate gaps may therefore persist, oscillate, or grow while no single issue looks exceptional. The negotiation can drift toward repeated rejection, rushed concessions, or a missed legal deadline before an issue-by-issue dashboard provides a clear warning.
What is proposed¶
For one bounded budget process, analysts would turn formally recorded, caucus-level proposal gaps into a small vector for each bargaining round. A local model would estimate which combinations of gaps decay, persist, oscillate, or grow. A mode would become actionable only if it met preregistered gain, persistence, resampling, reconstruction, and deadline-risk rules. Analysts would then model whether an already authorized procedural step—such as reordering discussion, separating a package, holding a joint factual briefing, or requesting simultaneous clarification—selectively reduces that mode. The facilitator could make a nonbinding recommendation, but elected officials would retain all authority over offers and votes. Drift, residual, or spectral-gap failures would suspend the dashboard.
The cross-domain transfer¶
The mode-decomposition archetype is instantiated as a map of how combined budget gaps change between bargaining rounds. Its modes are weighted packages of disagreement, and their gains indicate decay, growth, or oscillation. The transfer is structurally explicit, but its usefulness depends on an approximately stable local process and enough comparable rounds—both currently unshown.
Why it advanced¶
This candidate passed Experiment 4’s strict researched-candidate bar. That endpoint reflects the experiment’s screening criteria only; it does not establish real-world validation, novelty, deployment authority, or economic impact. It advanced with a falsifiable archival test, clear comparators and stopping rules, identifiable public authorities, and reversible, nonbinding use.
Prior art and the remaining open claim¶
Adjacent systems already track city budget amendments, support multi-issue negotiation, analyze negotiation processes over time, and decompose fitted dynamic operators. The open claim is narrower: within one stable city budget process, can an aggregate proposal-gap model find a resampling-stable coupled mode that improves held-out prediction over separate issue gaps, deadlines, official tracking, and observable shocks, leaves no consequential structured residual, and maps selectively to a procedure the authorized body can reverse? That claim remains unvalidated.
Smallest decisive test¶
Run an eight-week archival feasibility study with one consenting city and no live recommendation. Inventory formal packages, amendments, votes, and dated forecasts; stop if they cannot form reproducible aggregate snapshots. Preregister no more than eight coordinates, rank and sample rules, regularization, stability and spectral thresholds, residual limits, and both comparators. Use forward-chaining held-out tests. Advance only if a stable mode beats separate issue gaps plus deadline and the official tracker plus observable shocks, survives removal of shock rounds, and maps selectively to an authorized reversible procedure.
Deployment and cost¶
First evidence is estimated at a rough 2026 resource-equivalent cost of \(10,000–\)50,000. Initial startup is \(50,000–\)250,000; operational launch is \(250,000–\)1 million; and recurring annual cost is \(50,000–\)250,000. These are not vendor quotes. Live use would additionally require local legal, records, accessibility, governance, and procedural authorization.
Risks and uncertainties¶
- Too few comparable bargaining rounds could produce an overfit or rank-deficient model.
- Modes might reflect how clerks recorded proposals rather than how bargaining actually changed.
- Officials or the public might interpret a descriptive mode as evidence of motive, loyalty, or blame.
- A nominally procedural recommendation could redistribute agenda power or public visibility among constituencies.
- Negotiators could manipulate recorded positions, or conflict could move into an omitted dimension while the retained mode appears to improve.
Expert review¶
Useful reviewer backgrounds: Municipal budget-process and public-law specialist, Negotiation and political-process researcher, Dynamic-mode decomposition or system-identification statistician, Neutral public-sector facilitator, Public-records, privacy, and data-governance counsel.
- Does the archived process contain enough independent full-state snapshots to estimate the preregistered model reliably?
- Do mode shapes and gains remain stable under resampling, minor coordinate changes, and removal of shock rounds?
- Does the modal model beat both comparators on held-out prediction or reconstruction by the declared margin?
- Can any procedural lever be shown to affect the risky mode selectively rather than merely accompany broader political change?
- Who has legal authority to approve each recommended procedure, derived-data retention rule, and publication decision?
Evidence and provenance¶
Selected sources: S1: Coalition governance and municipal stability in South Africa: Institutional challenges and reform imperatives · S2: The Budget Process · S3: Budget Extender, Local Law 102 of 2026 (Int. 0873-2026) · S4: Seattle City Council budget amendment tracker · S5: NegoManage: A System for Supporting Bilateral Negotiations · S6: Analyzing the Multiple Dimensions of Negotiation Processes · S7: On Dynamic Mode Decomposition: Theory and Applications · S8: Math Occupations — Occupational Outlook Handbook
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review ranked this candidate 59th in every profile, in band D. That post-hoc order does not alter its strict-success endpoint and does not measure value; the pilot-speed input was only a cost-band affordability proxy.
Appendix B — Artifact index¶
This index is the audit map for the public report. It emphasizes canonical design, execution, analysis, and raw-evidence locations instead of listing every generated file. Candidate-specific proposal and evaluation links appear in the dossier appendix and candidate index.
Experimental artifacts are authoritative. Public-report files are later harmonizations or editorial explanations.
Program-level materials¶
| Artifact | Purpose |
|---|---|
| Inverse-innovation experiment directory | Program landing page and experiment map. |
| Program roadmap, 2026-08-03 | Decisions and planned work preserved before the literature review and later experiments. |
| Literature review | Search strategy, related work, primary-source synthesis, and contribution boundary. |
| Nearest-systems matrix | Feature-by-feature comparison with close systems and research traditions. |
| Report design | Approved scope, dossier policy, interpretation rules, and release gates. |
| Candidate index | Harmonized 59-row machine-readable evidence record. |
| Report fact table | Cross-experiment counts and canonical analysis pointers. |
| Report validation result | Latest automated invariant, dossier-length, artifact, and local-link audit. |
| Publication QA register | Status and evidence for the eleven approved release gates. |
| Prior-art sentinel audit | Ten-case historical, trade, non-English, and full-composition sensitivity audit. |
| Post-hoc yield decomposition | Archetype and domain concentration diagnostics across compatible Experiment 3–6 strata. |
| Prospective archetype-breadth probe | Probability-sampled generated-archetype productivity under a three-domain light screen. |
| Post-hoc proposal-type analysis | Blinded intervention-substrate and evidence-dependency coding for 422 evaluated proposals. |
| Prospective external context comparison | Uniform external scrutiny of Experiment 2's relevant, archetype-only, and irrelevant-mechanism terminal proposals. |
| Prospective substrate-denial test | Matched intervention on the E9 random stratum, with blinded quality/substrate measurement, public-web screening, and a triggered Max-effort follow-up. |
| Prospective second-slot policy test | New 60-cell comparison of ordinary-diverse and alternative-substrate P2s, with duplicate blinded measurement, opaque public-web screening, and a frozen portfolio decision rule. |
| Prospective Applicability Graph test | Stratified 40-archetype, 120-case benchmark separating candidate retrieval from blinded literal-by-literal DNF route verification. |
Experiment 1 — calibration¶
Purpose: initial sniff test of mechanism context across 60 cells and three conditions.
| Layer | Canonical artifacts |
|---|---|
| Orientation | README · experiment specification · rubric |
| Prompts and schemas | generation prompt · review prompt · cell schema · review schema |
| Cell declarations | calibration cells · full matrix |
| Results | experiment findings · calibration report · summary JSON |
| Audits | protocol QA · prior-art sniff |
| Code | analyze_inverse_innovation.py · validate_inverse_innovation_results.py |
Experiment 2 — relevant mechanisms and iterative revision¶
Purpose: compare relevant mechanisms with archetype-only and equal-length irrelevant-mechanism conditions in 20 paired cells; allow up to five revisions after the original proposal.
| Layer | Canonical artifacts |
|---|---|
| Pilot | pilot README · pilot protocol · completion manifest |
| Frozen design | protocol · analysis plan · design freeze · execution freeze |
| Core prompts | generator · reviser · tester · quality judge · gap classifier |
| Schemas | candidate · test · quality judgment · controller |
| Blinding and keys | quality commitment · revealed arm key · revealed quality key |
| Results | results narrative · results JSON · completion manifest |
| Post-hoc opportunity screen | method · report · ranked results · sources |
| Code | run_inverse_innovation_exp02_controller.py · analyze_inverse_innovation_exp02_main20.py · analyze_inverse_innovation_exp02_opportunities.py · validate_inverse_innovation_exp02_main20.py |
Experiment 3 — 320-cell full matrix and external verification¶
Purpose: cross five archetypes with 64 domains, run the closed-book pipeline, assess opportunities, and externally scrutinize a selected set.
| Layer | Canonical artifacts |
|---|---|
| Frozen design | protocol · analysis plan · design freeze |
| Closed-book prompts and schemas | generator · reviser · tester · candidate schema · test schema |
| Opportunity stage | method · design freeze · evaluator prompt · report |
| External research | method · design freeze · selection manifest · research prompt · assessment schema |
| Results | final report · integrated results · external report · external results · completion manifest |
| Raw candidate evidence | external inputs · external assessments |
| Code | run_inverse_innovation_exp03_full_matrix.py · run_inverse_innovation_exp03_external_research.py · analyze_inverse_innovation_exp03_integrated.py · validate_inverse_innovation_exp03_full_matrix.py |
Experiment 4 — proposal-first versus retrieval-first¶
Purpose: compare a complete-proposal-first workflow with broad retrieval and hypothesis screening in 20 paired cells.
| Layer | Canonical artifacts |
|---|---|
| Design | README · method · design freeze · selection manifest |
| Proposal-first prompts | control generator · evaluator · reviser |
| Retrieval-first prompts | self-scout · treatment ideator · prior-art critic · selector · proposal drafter |
| Schemas | proposal · evaluation · hypothesis pool · prior-art critique |
| Results | report · result JSON · resource accounting · completion manifest |
| Raw evidence | scientific records · run manifest |
| Code | build_inverse_innovation_exp04.py · run_inverse_innovation_exp04_v2.py · analyze_inverse_innovation_exp04.py · validate_inverse_innovation_exp04_complete.py |
Experiment 5 — five-proposal portfolio pilot¶
Purpose: generate and fully evaluate five diverse complete proposals in each of 20 cells, then examine marginal yield by proposal position.
| Layer | Canonical artifacts |
|---|---|
| Design | README · method · design freeze · selection manifest |
| Prompts | proposal 1 generator · portfolio continuation · diversity auditor · opportunity evaluator · reviser |
| Schemas | proposal · evaluation · diversity · controller |
| Results | report · result JSON · proposal-level results · resource accounting · completion manifest |
| Raw evidence | scientific records · run manifest |
| Code | build_inverse_innovation_exp05.py · run_inverse_innovation_exp05.py · analyze_inverse_innovation_exp05.py · validate_inverse_innovation_exp05.py |
Experiment 6 — four-proposal 60-cell generalization¶
Purpose: test the four-proposal rescue rule on five new archetypes and 12 domains; add a calibrated empirical-partner lane.
| Layer | Canonical artifacts |
|---|---|
| Design | README · method · design freeze · selection manifest |
| Prompts | proposal 1 generator · portfolio continuation · diversity auditor · opportunity evaluator · empirical-partner adjudicator |
| Schemas | proposal · evaluation · empirical partner · controller |
| Partner calibration | cases · calibration result |
| Results | report · result JSON · cell-level results · proposal-level results · resource accounting · completion manifest |
| Raw evidence | scientific records · partner adjudications · run manifest |
| Code | build_inverse_innovation_exp06.py · run_inverse_innovation_exp06.py · analyze_inverse_innovation_exp06.py · validate_inverse_innovation_exp06.py |
Experiment 7 — retrospective proposal-only selector¶
Purpose: test whether blinded proposal-only ranking could prefilter Experiment 6 candidates while retaining strict and empirical-partner yield.
| Layer | Canonical artifacts |
|---|---|
| Design | README · method · design freeze · input manifest · blind map |
| Selector | prompt · schema · frozen consensus · policy selections |
| Results | report · result JSON · policy metrics · false negatives · resource accounting · completion manifest |
| Raw evidence | selector outputs · run manifest |
| Code | build_inverse_innovation_exp07.py · run_inverse_innovation_exp07.py · analyze_inverse_innovation_exp07.py · validate_inverse_innovation_exp07.py |
Report-hardening prior-art sentinel¶
Purpose: test whether explicitly searching historical, trade, non-English, and full-composition lanes exposes closer precedents for a hash-selected ten-case sentinel.
| Layer | Canonical artifacts |
|---|---|
| Frozen design | method · design freeze · hash-selected sample |
| Search and cases | query log · case results |
| Result | audit report |
| Code | build_inverse_innovation_prior_art_sentinel.py |
Experiment 8 — post-hoc yield decomposition¶
Purpose: describe how observed outcomes are distributed by tested archetype and domain without pooling unlike protocols.
| Layer | Canonical artifacts |
|---|---|
| Frozen design | method · design freeze |
| Results | report · results JSON · archetype rows · domain rows |
| Code | analyze_inverse_innovation_exp08.py |
Experiment 9 — prospective archetype-breadth probe¶
Purpose: test whether light-screen productivity extends beyond the archetypes selected in earlier experiments, including a probability sample from the previously untested generated-archetype corpus.
| Layer | Canonical artifacts |
|---|---|
| Frozen design | method · design freeze · documented design correction · selection manifest · cell manifest |
| Prompts and schemas | generator · light screen · proposal schema · screen schema |
| Corpus and inputs | archetype source snapshot · 150 cell inputs |
| Results | report · result JSON · archetype rows · cell rows · completion manifest |
| Raw evidence | proposal and screen records · run manifest |
| Code | build_inverse_innovation_exp09.py · seal_inverse_innovation_exp09_design.py · run_inverse_innovation_exp09.py · analyze_inverse_innovation_exp09.py · validate_inverse_innovation_exp09.py |
Experiment 10 — post-hoc proposal type × disposition¶
Purpose: classify all 422 primary externally evaluated proposals on intervention substrate and decisive evidence dependency, with outcome labels withheld from two fresh passes and a third blinded adjudication.
| Layer | Canonical artifacts |
|---|---|
| Frozen design | method · design freeze |
| Blinded record | blinded inputs · withheld outcomes · pass 1 · pass 2 · disagreements · adjudications |
| Results | report · results JSON · joined rows |
| Code | prepare_inverse_innovation_exp10.py · run_inverse_innovation_exp10_classifiers.py · run_inverse_innovation_exp10_adjudicator.py · analyze_inverse_innovation_exp10.py |
Post-hoc archetype substrate analyses¶
Purpose: correct the experimental archetype-score summary, test whether structural/framed character shifts proposal substrate, and distinguish source modulation from a broader governance/computational pipeline constraint.
| Layer | Canonical artifacts |
|---|---|
| Combined synthesis | report · corrected internal note · validation result |
| E10B existing-data join | method · design freeze · report · results · archetype rows |
| E9 blinded follow-up | method · design freeze · classification seal · report · results |
| E9 classification record | blinded inputs · pass 1 · pass 2 · disagreements · adjudications · joined rows |
| Code | analyze_inverse_innovation_e10b.py · prepare_inverse_innovation_e9_substrate_followup.py · run_inverse_innovation_e9_substrate_classifiers.py · run_inverse_innovation_e9_substrate_adjudicator.py · analyze_inverse_innovation_e9_substrate_followup.py · validate_inverse_innovation_substrate_analysis.py |
Experiment 11 — external mechanism-context comparison¶
Purpose: test whether Experiment 2's relevant-mechanism advantage persists when all preserved terminal proposals receive the same external research and blinded comparative judgment.
| Layer | Canonical artifacts |
|---|---|
| Frozen design | protocol · analysis plan · design freeze · input manifest |
| Amendments | transport/schema amendment · its freeze · interrupted-attempt recovery · its freeze |
| Prompts and schemas | external scrutinizer · blinded comparative judge · external-evaluation schema · paired-judgment schema |
| Blinding and adjudication | sealed adjudication record · tiebreak manifest · revealed arm key |
| Results | report · result JSON · completion manifest |
| Raw evidence | external evaluations and comparative judgments · run manifest |
| Code | build_inverse_innovation_exp11.py · seal_inverse_innovation_exp11_design.py · run_inverse_innovation_exp11.py · recover_inverse_innovation_exp11_interrupted.py · analyze_inverse_innovation_exp11.py · validate_inverse_innovation_exp11.py |
Experiment 12 — paired substrate-denial intervention¶
Purpose: test whether the governance/computational concentration observed in E9 is an instruction-sensitive default, measure the quality and light-screen cost of leaving those substrates, and run a frozen Max-effort follow-up only if the High-effort cost trigger fires.
| Layer | Canonical artifacts |
|---|---|
| Frozen design | method · design seal · input manifest · Max subset · README status-file audit correction |
| Prompts and schemas | prompts · schemas |
| Blinding and joins | primary blinded inputs · primary measurement seals · primary revealed joins · Max blinded record |
| Results | complete report · primary result JSON · cell rows · Max result JSON · resource summary · validation result |
| Raw evidence | proposal, screen, judgment, and telemetry records · Max scientific measurement records |
| Code | build_inverse_innovation_exp12.py · run_inverse_innovation_exp12.py · run_inverse_innovation_exp12_screens.py · analyze_inverse_innovation_exp12.py · analyze_inverse_innovation_exp12_max.py · summarize_inverse_innovation_exp12_runtime.py · validate_inverse_innovation_exp12.py |
Experiment 13 — second-slot portfolio policy¶
Purpose: determine prospectively whether one additional proposal should use ordinary conceptual diversity or deliberately seek a non-governance, non-computational causal substrate.
| Layer | Canonical artifacts |
|---|---|
| Frozen design and sample | method · design seal · selection manifest · cell manifest · source snapshot |
| Prompts and schemas | prompts · schemas · batch manifest |
| Blinding and measurement | opaque inputs and pair records · measurement input seal · measurement output seal · adjudication manifest |
| Results | completion status · complete report · result JSON · cell rows · incremental survivors · validation result |
| Raw evidence | 180 proposals, blinded passes, adjudications, screens, and telemetry · scientific run manifest · measurement manifest · adjudication manifest |
| Code | build_inverse_innovation_exp13.py · seal_inverse_innovation_exp13_design.py · run_inverse_innovation_exp13.py · prepare_inverse_innovation_exp13_blinding.py · run_inverse_innovation_exp13_measurement.py · run_inverse_innovation_exp13_adjudication.py · analyze_inverse_innovation_exp13.py · validate_inverse_innovation_exp13.py |
Experiment 14 — Applicability Graph retrieval and DNF verification¶
Purpose: test prospectively whether the new diagnostic Applicability Graph improves problem-to-archetype retrieval over the existing solution-oriented index and whether explicit DNF routes reject one-literal near misses.
| Layer | Canonical artifacts |
|---|---|
| Frozen design and sample | method · design seal · selection manifest · sampling seed |
| Prompts and schemas | prompts · schemas |
| Blinding and measurement | opaque audit and verifier inputs · eligible-case manifest · measurement input seal · verification input seal · audit recovery amendment |
| Results | completion status · complete report · result JSON · case rows · validation result · completion manifest |
| Raw evidence | case bundles, duplicate audits, adjudications, retrieval records, verifier outputs, and telemetry · scientific run manifest · private join keys |
| Code | build_inverse_innovation_exp14.py · seal_inverse_innovation_exp14_design.py · run_inverse_innovation_exp14.py · recover_inverse_innovation_exp14_audit_batch.py · analyze_inverse_innovation_exp14.py · validate_inverse_innovation_exp14.py |
Experiment 15 — route-aware Applicability Graph retrieval¶
Purpose: test prospectively, on a disjoint 40-archetype sample, whether five frozen condition-level DNF aggregators improve problem-to-archetype candidate retrieval over the solution-index and diagnostic maximum-hit baselines.
| Layer | Canonical artifacts |
|---|---|
| Frozen design and sample | method · design seal · selection manifest · sampling seed |
| Prompts and schemas | case writer · case auditor · case adjudicator · schemas |
| Blinding and measurement | opaque audit inputs · eligible-case manifest · measurement input seal · private join keys |
| Results | completion status · complete report · post-hoc failure localization · result JSON · post-hoc diagnostic JSON · case rows · validation result · completion manifest |
| Raw evidence | case bundles, duplicate audits, adjudications, retrieval records, and telemetry · scientific run manifest and approval record · 40 applicability packets |
| Code | build_inverse_innovation_exp15.py · seal_inverse_innovation_exp15_design.py · run_inverse_innovation_exp15.py · inverse_innovation_route_retrieval.py · analyze_inverse_innovation_exp15.py · analyze_inverse_innovation_exp15_diagnostics.py · validate_inverse_innovation_exp15.py |
Public-report derivation¶
| Artifact | Role |
|---|---|
build_inverse_innovation_public_report_data.py |
Extracts and validates the 59 canonical candidate records, report facts, cost bands, and post-hoc sensitivity ordering. |
draft_inverse_innovation_public_report_dossiers.py |
Produces preserved editorial plain-language drafts from the frozen candidate records without web access. These are not experimental outputs. |
build_inverse_innovation_public_report.py |
Deterministically compiles candidate data, editorial drafts, sources, and artifact links into the dossier appendix. |
validate_inverse_innovation_public_report.py |
Audits counts, endpoints, required fields, links, anchors, and headline quantitative claims. |
| Dossier drafting manifest | Records editorial model, settings, batch count, source, and completion status. |
| Dossier drafts | Structured editorial derivatives used by the Markdown compiler. |
Reading provenance correctly¶
- A proposal file records what the generator ultimately proposed.
- An evaluation file records the bounded external evidence and the evaluator's judgment at that time.
- An analysis file aggregates frozen rules across proposals or cells.
- A dossier translates those records for readers and may introduce a short alias; it does not replace them.
- A later expert review should be appended as new evidence, not written into an experimental record.
Appendix C — Expert review template¶
Use this form to evaluate one candidate from the dossier appendix. Reviewers should read the original proposal, external evaluation, and cited sources before reaching a final disposition. A short review based only on the plain-language dossier should be labeled summary-only.
The purpose is not to reward novelty language. It is to determine whether the problem, causal transfer, comparator, and next test survive informed scrutiny.
Review metadata¶
| Field | Response |
|---|---|
| Candidate portfolio ID | |
| Canonical title | |
| Reviewer name or stable pseudonym | |
| Date | |
| Domain(s) of expertise | |
| Years or type of relevant experience | |
| Review depth | SUMMARY_ONLY / FULL_ARTIFACT / FULL_PLUS_ADDITIONAL_SEARCH |
| Conflicts of interest | |
| Information that cannot be made public |
1. Plain-language reconstruction¶
In your own words, state:
- the real-world problem;
- the proposed intervention;
- the strongest existing comparator; and
- the result that would justify taking a next step.
If this cannot be done without guessing, identify the ambiguity and stop the review until it is resolved.
2. Problem reality and importance¶
| Question | Rating | Evidence or explanation |
|---|---|---|
| Does the described problem occur? | NO / RARE / SOMETIMES / COMMON / UNKNOWN |
|
| Is the observable state measured correctly? | 1–5 or UNKNOWN |
|
| Is the consequence causal, or only associated? | CAUSAL / PLAUSIBLE / ASSOCIATION_ONLY / UNSUPPORTED |
|
| Would solving it materially matter? | 1–5 | |
| Is there an actor with incentive and authority to act? | 1–5 |
What evidence would most change these judgments?
3. Prior art and distinctiveness¶
List the closest systems, publications, products, standards, patents, policies, or ordinary practices you know. Include links or citations where possible.
| Question | Rating | Explanation |
|---|---|---|
| Are the dossier's nearest rivals actually the strongest comparators? | YES / PARTLY / NO / UNKNOWN |
|
| Does a close implementation already exist? | YES / PARTLY / NO_CLOSE_MATCH_KNOWN / UNKNOWN |
|
| Is the remaining contrast stated fairly? | 1–5 | |
| Is that contrast consequential rather than cosmetic? | 1–5 |
Recommended prior-art disposition:
-
ESTABLISHED_PRACTICE -
SUBSTANTIAL_COLLISION -
ADJACENT_PRIOR_ART -
NO_CLOSE_MATCH_FOUND_IN_REVIEW -
INDETERMINATE
This label describes your bounded review; it is not a patent or world-novelty opinion.
4. Structural transfer¶
Does the solution archetype do real causal work, or has the proposal merely renamed a domain practice?
| Criterion | Rating (1–5) | Explanation |
|---|---|---|
| Source relations are represented accurately | ||
| Target elements play genuinely corresponding roles | ||
| The mapping yields a non-obvious intervention or test | ||
| Domain constraints are preserved during adaptation | ||
| The proposal would be weaker without this transfer |
Identify any broken correspondence, missing mechanism, or better abstraction.
5. Intervention and implementation¶
| Question | Rating | Explanation |
|---|---|---|
| Is the intervention specified well enough to prototype? | 1–5 | |
| Are the required data or materials obtainable? | 1–5 | |
| Is the proposed authority structure realistic? | 1–5 | |
| Are workflow and incentive effects accounted for? | 1–5 | |
| Are safety and rollback protections adequate for the next test? | 1–5 | |
| Is the first-evidence cost band plausible? | TOO_LOW / PLAUSIBLE / TOO_HIGH / UNKNOWN |
|
| Is the startup cost band plausible? | TOO_LOW / PLAUSIBLE / TOO_HIGH / UNKNOWN |
List any missing technical, legal, ethical, labor, environmental, security, accessibility, or distributional constraint.
6. Decisive test¶
Rewrite the smallest useful test if necessary.
| Element | Review |
|---|---|
| Unit of analysis | |
| Eligible sample or cases | |
| Intervention condition | |
| Strongest comparator | |
| Primary outcome | |
| Safety/noninferiority outcomes | |
| Minimum useful effect or decision threshold | |
| Stop conditions | |
| Required decision-maker or data owner | |
| Estimated duration and cost |
Would a positive result actually distinguish the proposal from its prior art? Would a negative result cause the project to stop or materially change?
7. Recommended disposition¶
Choose one:
-
REJECT_PROBLEM— the problem is absent, trivial, or materially misstated. -
REJECT_COLLISION— existing practice covers the consequential claim. -
REJECT_CAUSAL_CHAIN— the intervention does not plausibly change the outcome. -
REVISE— a bounded conceptual revision could produce a valid candidate. -
MERGE— combine with a named existing or dossier candidate. -
PREVALENCE_STUDY— first determine whether the problem is frequent or consequential. -
PARTNERED_EVIDENCE_STUDY— obtain internal data or operational access before judging it. -
BOUNDED_PILOT— run the reviewed reversible comparison. -
ADOPTION_INQUIRY— the main uncertainty is demand, ownership, or authority.
Overall confidence: LOW / MODERATE / HIGH
What is the single strongest reason for this disposition? What evidence would reverse it?
8. Human revision and inspiration¶
Did the candidate help you formulate something better, even if you rejected it?
- Revised problem:
- Revised intervention:
- Better domain or use case:
- Better archetype or composition of archetypes:
- New measurement or comparator:
- Estimated conceptual distance from the original:
MINOR/MODERATE/MAJOR/ENTIRELY_NEW
This section is important research data. It can reveal value from human–AI co-creation that a binary survivor count misses.
9. Publication permission¶
- The full review may be published with attribution.
- The full review may be published under a pseudonym.
- Only an anonymized structured summary may be published.
- Do not publish; use only for aggregate research.
Optional signature or verification method:
Appendix D — Publication quality-control register¶
Status date: 2026-08-06
Scope: August 2026 public working-paper package
Interpretation: PASS means the named gate has direct evidence. PARTIAL
means that a narrower automated or editorial check passed but the full semantic
claim has not been independently verified. A working manifest is not a final
release seal while any gate remains partial or pending.
| Gate | Status | Evidence and remaining work |
|---|---|---|
| 1. Generate a canonical fact table | PASS — automated | data/report_facts.json is regenerated from preserved artifacts for twelve prospective experiments, one retrospective experiment, and five post-hoc analyses by build_inverse_innovation_public_report_data.py. |
| 2. Reconcile counts and endpoint definitions | PASS — automated | validate_inverse_innovation_public_report.py checks the 59 dossiers, portfolio endpoint counts, experiment scope, ranks, and load-bearing report phrases, including the Experiment 9 and 11–15 frozen results. Protocol-specific denominators remain separate in the report. |
| 3. Verify all 59 candidate lineages and classifications | PASS — automated lineage; PARTIAL — semantic | Every candidate links to existing proposal and evaluation artifacts and has one preserved endpoint class. No independent expert has re-adjudicated all classifications. |
| 4. Check every dossier against proposal and terminal evaluation | PARTIAL | The dossier compiler reads the canonical candidate index and validators check IDs, links, lengths, and required fields. Full sentence-level semantic comparison by an independent reader has not been completed for all 59. |
| 5. Verify external citations and coverage dates | PARTIAL | Source objects require URL, publisher, source class, date fields, access date, and supported claims. The frozen ten-case prior-art sentinel ran 40 targeted historical, trade/standards, non-English, and composition queries: 8/10 cases acquired closer adjacent precedent, 2 did not, and none met its likely-substantial-collision rule. This improves sensitivity evidence but does not verify every citation or prove exhaustive search. |
| 6. Reproduce every table and figure from saved data | PARTIAL | Experiment tables are cross-checked against canonical analysis outputs. Experiments 9 and 11–15 pass their dedicated validators; the Experiment 8 and Experiment 10 tables pass the hardening validator; and the substrate analyses pass freeze-hash, blinding, matrix, clustered-result, and frozen-gate checks. Several main-report prose tables remain editorial rather than generated directly from one script. |
| 7. Audit every numerical claim against the fact table | PARTIAL | Validators check the principal published counts, yields, resources, denominators, hardening results, Experiment 9/11/12/13/14/15 frozen statistics, and E10B/E9 substrate headlines. The 422-record proposal-type analysis used two outcome-blinded passes plus blinded adjudication of all 68 disagreements; the E9 substrate follow-up used two blinded passes plus adjudication of all ten disagreements before key join; Experiment 12 preserves its two-pass classifications, paired judgments, adjudications, web screens, and cell-level join; Experiment 13 preserves 180 proposals, duplicate blinded measurement, all disagreement adjudications, 180 opaque screens, and its paired join; Experiment 14 preserves duplicate case audits, all material-disagreement adjudications, sealed retrieval inputs, and 114 blinded DNF verifications; and Experiment 15 preserves a disjoint archetype sample, duplicate case audits, material-disagreement adjudications, sealed eligibility, and seven deterministic retrieval arms. A complete independent audit of every numeral in narrative prose has not been performed. |
| 8. Run link and anchor checks | PASS — automated | Public-report and publication validators check local targets, dossier anchors, site chapter links, and registered downloads. External reachability can change after release. |
| 9. Review narrative accessibility and technical fidelity | PASS — editorial; PARTIAL — independent | The report received a cold adversarial model review and subsequent correction pass. It has not received independent domain-expert or journal review. |
| 10. Preserve a version, changelog, and correction route | PASS | Machine release metadata, change log, visible revision month, original experiment artifacts, and a correction policy are present. |
| 11. Seal only after all gates pass | PENDING | The current 0.9 manifest identifies a reproducible working publication state. It is not a final QA seal while gates 3–7 and 9 remain partial. |
What the automated checks establish¶
They establish that expected artifacts exist, counts reconcile, candidate IDs and endpoint classes are internally consistent, required source fields are present, local links and anchors resolve, and publication archives match their manifests. They do not establish world novelty, source truth, causal validity, domain-expert agreement, or the absence of older and proprietary precedent.
Release rule¶
The report may be published as a clearly labeled working paper with this status register. A future final seal requires either completing the partial semantic gates or explicitly revising the approved release criteria; silent equivalence between automated conformance and independent verification is not permitted.