Candidate dossiers — Band C¶
Part of Inverse Innovation with the Encyclopedia of Abstractions · Ranks 26–45: intermediate post-hoc review priority · Last revised August 2026
26. Review Materials Results by Model Mismatch¶
Canonical title: Residual-Governed Characterization Queue for Materials Discovery
In one sentence: A blinded partner study would test whether scientists can review model-relative characterization differences instead of every complete package while preserving protected findings and scientific decisions.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-12 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Chemistry Materials |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 10–36 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-11, EXP06-PARTNER-13, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A materials campaign may generate a large characterization package for every composition and processing condition. When scientists review all packages equally, routine confirmations consume the same attention as results that contradict predicted phases, structures, or properties. Fixed features and anomaly labels can shorten the queue, but they may erase unfamiliar peaks, minority phases, or small reliable contradictions. The actual field gap is fundamental: no evidence yet shows that a target laboratory's full-package queue is overloaded or delays important discrepancies.
What is proposed¶
For each candidate, a versioned campaign model would freeze predicted phase, structure, and property outputs before measurements are revealed. The system would compare those predictions with standardized results and preserve signed property errors, categorical disagreements, missing observations, unpredicted peaks or phases, uncertainty, and provenance. Reliable or consequential discrepancies would enter a primary queue with a named reviewer and possible actions such as replication, an orthogonal assay, or a model-scope review. Expected results would remain reconstructable and complete raw files available on demand. Random candidates, protected classes, and risk-selected cases would still receive full-package review. New chemical classes, incompatible versions, invalid measurements, safety flags, stale models, structured residuals, or excessive reconstruction error would suspend triage. Scientists—not the queue—would authorize experiments and model changes.
The cross-domain transfer¶
The residual-processing archetype becomes a scarce-attention system rather than a data link. A model predicts each candidate's standardized characterization; measured-minus-predicted differences become the review message and later learning signal. The mapping is strong for numerical properties but weaker for images, peaks, phases, and unfamiliar features, which cannot always be reconstructed additively and must remain accessible in full.
Why it advanced¶
This candidate cleared only the separately calibrated empirical-partner lane. Autonomous laboratories, active learning, human-guided phase mapping, and large provenance systems make its components plausible. It did not reach strict success because no paired field trial shows that residual presentations reduce total work without changing scientific actions, and the target queue constraint remains unmeasured.
Prior art and the remaining open claim¶
Adjacent systems already use Bayesian optimization, uncertainty-guided measurement selection, automated phase analysis, anomaly detection, and human model updates. The narrower open comparison is an audited review interface: frozen multimodal predictions plus numerical, categorical, missingness, and unfamiliar-feature residuals, with protected and sampled full-package review. On one material family, it claims at least 30% less total reviewer time, at least 95% action concordance, complete capture of protected challenges, no acceptance of missing or incompatible results, and no increase in follow-up demand.
Smallest decisive test¶
A partner laboratory would conduct a paired, blinded shadow study on 60–150 already authorized candidates from one family and fixed assay suite. It would compare complete packages, uncertainty-only ordering, fixed anomaly or feature thresholds, and the frozen residual presentation. Concealed challenges would include small property shifts, phase conflicts, unfamiliar peaks, bad assays, missing data, version mismatches, and out-of-scope chemistry. Reject operational adoption if workload falls less than 30% after audits and fallbacks, concordance is below 95%, any protected case is missed, an incompatible or missing package passes, or one discrepancy class is systematically lost.
Deployment and cost¶
A shadow trial needs laboratory leadership, characterization, model, safety, and data-steward authorization, plus review of trade-secret, export-control, retention, and publication rules. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. No vendor quote supports them.
Risks and uncertainties¶
- Standardization may remove an unfamiliar peak, morphology, or minority phase before the residual is calculated.
- A confident but misspecified campaign model may classify informative chemistry as routine.
- Underestimated assay uncertainty may give noisy results excessive influence over routing and model updates.
- Audits may miss rare discrepancies, especially when risk sampling reflects only anticipated failure modes.
- Routine packages removed from the main queue may contain context needed to recognize a later campaign-wide pattern.
Expert review¶
Useful reviewer backgrounds: Materials discovery scientist, X-ray diffraction or multimodal characterization specialist, Bayesian modeling and active-learning researcher, Scientific data-governance specialist, Experimental-design statistician.
- Is complete-package review currently exceeding a declared capacity or delaying model-relevant discrepancies in the proposed campaign?
- Which raw or standardized observations cannot be represented safely as model-relative residuals?
- What protected challenge set would cover unfamiliar peaks, minority phases, specimen errors, and out-of-scope chemistry?
- How should audit size and sampling be powered for rare but decision-changing discrepancies?
- Can independent reviewers reliably classify each mismatch as chemical evidence, measurement failure, specimen error, or model-scope failure?
Evidence and provenance¶
Selected sources: S1: Achieving AI-Driven Autonomous Laboratories · S2: Autonomous Methods · S3: Workshop Report on Autonomous Methodologies for Accelerating X-ray Measurements · S4: An autonomous laboratory for the accelerated synthesis of inorganic materials · S5: Author Correction: An autonomous laboratory for the accelerated synthesis of inorganic materials · S6: On-the-fly closed-loop materials discovery via Bayesian active learning · S7: Human-in-the-loop for Bayesian autonomous materials phase mapping · S8: NCAL: Data Management
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: Post-hoc harmonized profiles place this candidate from rank 10 to 36, in ordering band C. The spread is a reading-order aid, not a preregistered outcome or estimate of economic value, and it does not remove the need for field evidence.
27. A Common Contract for Synchronized Takes¶
Canonical title: Representation-Independent Synchronized-Take Contract for Dailies
In one sentence: A read-only pilot would test whether different picture-and-sound synchronizers can expose the same take membership, timing, gaps, and conflicts without clients depending on filenames or vendor internals.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-21 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Film Media Production |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 18–32 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-22, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Dailies systems often decide which picture and sound files form a take by using filenames, folder order, embedded timecode, recorder labels, sidecars, or one algorithm's behavior. A camera, recorder, metadata carrier, file split, or synchronization tool can therefore change grouping and frame-to-sample alignment even when the intended take is unchanged. Existing pipelines lack a shared behavioral test for deciding whether a replacement synchronizer truly preserves membership, timing, channel roles, gaps, conflicts, and ambiguity.
What is proposed¶
An opaque SynchronizedTake component would register picture and sound streams by semantic role, accept clock or slate observations, resolve alignment under an explicit policy, map stream positions into common take time, report shared coverage, validate completeness, freeze a resolved version, and create superseding versions. Its contract would require monotonic mapping within continuous segments, one public correspondence for every resolved position, explicit gaps and discontinuities, declared ambiguity and conflict errors, immutability after freezing, and no state change after rejection. Filenames, folders, metadata locations, waveform features, caches, databases, and matching algorithms would remain private. The same black-box fixtures, generated sequences, and representation-change tests would judge manual-anchor, timecode, waveform-assisted, and device-specific implementations. The component could not rewrite, delete, or silently relabel source media.
The cross-domain transfer¶
The interface-contract archetype maps cleanly to dailies: define a synchronized take by observable behavior and invariants, not by its file layout or synchronization algorithm. Each implementation translates private evidence into the same semantic streams and common-time mappings. Substitution depends on passing one oracle, although narrowly authorized diagnostic access may still be needed for physical clock provenance.
Why it advanced¶
This is an empirical-partner candidate, not a strict-success result. Existing synchronization products, timecode and audio standards, interchange schemas, and implementation-neutral timeline models support feasibility. It advanced because a copied-media, read-only comparison is bounded and reversible, while real failure prevalence, operator acceptance, independent implementations, and production authorization remain absent.
Prior art and the remaining open claim¶
Premiere and Resolve already synchronize multicamera material; FCPXML represents synchronized clips; BWF and SMPTE timecode standardize important carriers; and OpenTimelineIO separates timelines from proprietary files. The remaining claim is narrower than a new format: two independent resolvers would return equivalent semantic membership, covered intervals, mappings, gaps, conflicts, and lifecycle outcomes after renaming, relocation, registration reordering, clock-origin shifts, and lossless segmentation—and would outperform a metadata-schema-only approach without exposing private implementation details.
Smallest decisive test¶
With written production approval, duplicate one short, non-live take containing two picture streams, separate sound, aligned coverage, and a known gap or conflict. Freeze expected membership, mappings, errors, positional tolerance, and the no-write rule. Compare the incumbent result, a simple manual-anchor model, one adapter, and a schema-only baseline where possible. Rename and relocate files, reorder registration, shift the clock origin with compensation, and split a continuous stream losslessly. Reject the claim if independent resolvers diverge beyond tolerance, silently resolve conflicts, mutate after rejection, require private fields for ordinary work, or the schema baseline matches them with materially less effort.
Deployment and cost¶
A pilot requires post-production, sound, camera or data, dailies, editorial, rights, and security approval. It must use copied media and leave the active pipeline unchanged. Rough 2026 bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for launch, and \(50,000–\)250,000 annually; no measured effort or quote confirms them.
Risks and uncertainties¶
- The common-time model may omit drift, variable rates, discontinuities, or corrupted-clock behavior found in production media.
- Clock provenance or recorder-specific channel facts may be necessary production information rather than accidental representation.
- A shared test fixture may encode the same mistaken synchronization assumptions as the reference implementation.
- Opacity may make failures difficult to diagnose unless a narrowly governed read-only diagnostic surface is defined.
- Two mappings within a numerical tolerance may still differ in ways assistant editors or sound staff consider operationally unacceptable.
Expert review¶
Useful reviewer backgrounds: Dailies pipeline engineer, Production sound mixer or synchronization specialist, Assistant editor, Media interchange and timecode standards expert, Post-production security and rights specialist.
- Which device-specific facts must remain visible for authorized diagnosis, and which are accidental client dependencies?
- What frame or sample tolerance is acceptable for each downstream viewing, logging, and editorial task?
- How must the abstract model represent drift, discontinuities, variable frame rates, missing coverage, and multichannel sound?
- Can a canonical metadata baseline provide the same invariance and usability with less integration work?
- What production-derived fixtures would adequately represent actual synchronization failures without exposing protected media?
Evidence and provenance¶
Selected sources: S1: The 2030 Vision White Paper Section 3.3 · S2: 2030 Greenlight · S3: Create multi-camera source sequences in Premiere · S4: DaVinci Resolve – Edit · S5: Document Type Definition: Final Cut Pro XML Interchange Format 1.10 · S6: EBU Tech 3285 v2: Specification of the Broadcast Wave Format · S7: SMPTE ST 12-1:2014, Time and Control Code · S8: OpenTimelineIO 0.16.0 Architecture
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized review places this candidate between ranks 18 and 32, in ordering band C. This is only a reading-order aid based partly on an affordability proxy; it is neither a preregistered success nor evidence of production benefit.
28. Portable Lighting Cues With Behavioral Tests¶
Canonical title: Rig-Independent Cinematic Lighting Cue Contract
In one sentence: An isolated trial would test whether different lighting controllers and fixture allocations can reproduce an approved cue trajectory and safe lifecycle without exposing raw console addresses.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-22 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Film Media Production |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 23–30 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-21, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Film lighting cues are often encoded through one console's channels, fixture profiles, addresses, show-file structure, macros, and transition rules. Moving a scene, changing a console, or reallocating fixtures can alter fades, colors, holds, conflicts, or emergency transitions even when the intended look has not changed. File import and transport standards help move data or commands, but the production still lacks a shared test for whether another controller preserves both the photographed cue behavior and its output boundaries.
What is proposed¶
An opaque LightingCueProgram would expose semantic operations: validate a rig, arm without changing output, trigger, sample declared state, hold, resume, cancel, abort to a safe state, report status, freeze, and supersede. Its abstract model would name lighting roles, targets, timed trajectories, interpolation, concurrency and conflict rules, enrolled outputs, a safe state, and version history. Valid event sequences would have deterministic abstract trajectories; rejected or unarmed commands would leave output unchanged; abort would take precedence; and commands could reach only enrolled endpoints. Console syntax, channels, addresses, universes, profiles, show files, and drivers would remain private. A shared black-box suite would compare a reference player and console adapters after address renumbering, registration reordering, and capability-equivalent fixture reallocation, using visual, camera, photometric, timing, and safety tolerances fixed by accountable staff beforehand.
The cross-domain transfer¶
The representation-independent-contract archetype is instantiated as a semantic lighting state machine separated from console and rig details. Adapters translate the same cue operations into different physical implementations, and one conformance oracle governs substitution. The mapping is incomplete unless it captures photographed light: equal console values do not guarantee equal spectra, flicker, optics, placement, or camera response.
Why it advanced¶
This candidate entered only the empirical-partner lane. Show-file import, GDTF/MVR, DMX, ACN, and open control gateways establish adjacent technical pieces, while a simulator test is small and reversible. It lacks measured production prevalence, an independent implementation, approved visual tolerances, vendor or production commitment, and evidence that physical fixture differences will not dominate representation effects.
Prior art and the remaining open claim¶
Current prior art standardizes show-data import, device and scene descriptions, transport, and cross-protocol control. It does not establish a cinematographer-approved cue lifecycle and emitted-light substitution test. The remaining claim is that two independent adapters can hide addressing yet deliver equivalent semantic and photographed trajectories, identical errors and lifecycle behavior, no unenrolled output, and an approved abort transition after representation changes. Existing interchange matching those results at equal or lower effort would defeat the incremental claim.
Smallest decisive test¶
With DP, gaffer, electrical, and safety approval, test one disconnected cue containing arming, a timed fade, concurrent endpoints, hold and resume, a conflict, an unarmed rejection, and abort-to-safe. Freeze expected trajectories, error codes, enrolled outputs, camera and photometric tolerances, timing limits, and abort deadline. Compare a reference player, one console adapter, and the incumbent import-and-rebuild workflow, then repeat after address, registration, and fixture-allocation changes. Reject the claim for any unenrolled command, rejected-command output, lifecycle disagreement, tolerance failure, required console-native escape hatch, cheaper equivalent incumbent performance, or materially different photographed result despite nominal conformance.
Deployment and cost¶
The first test must remain on a simulator or isolated prelight bay with no performers, shooting rig, hazardous effects, machinery, or production automation attached. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for launch, and \(50,000–\)250,000 annually. They are not quotes.
Risks and uncertainties¶
- Normalized semantic values may produce different perceived and photographed light across fixture technologies.
- The contract may omit spectral output, flicker, optics, spatial placement, latency, or resource contention important to the scene.
- Console telemetry may conform even while the emitted light falls outside the approved camera or photometric tolerance.
- A software-defined safe state may conflict with stage electrical procedure or fail over an unreliable lighting transport.
- The contract may accidentally preserve the incumbent console's interpolation or tie-breaking behavior instead of the cinematographer's intent.
Expert review¶
Useful reviewer backgrounds: Director of photography or imaging scientist, Gaffer and lighting console programmer, Stage electrical and entertainment-control safety specialist, Fixture calibration and spectral-measurement specialist, Lighting-control standards and adapter engineer.
- Which semantic endpoint properties and photographed-light measurements define equivalence for the selected cue?
- What tolerances for spectrum, exposure, color, flicker, trajectory timing, and abort completion must be frozen before testing?
- Can the chosen GDTF descriptions represent the relevant fixtures accurately, including unsupported or vendor-specific features?
- Which abort behavior is safe under actual electrical procedure, and what can be guaranteed over a transport that may lose packets?
- Does the incumbent import-and-manual-rebuild workflow already meet the same criteria with equal or lower effort?
Evidence and provenance¶
Selected sources: S1: Importing Show Files · S2: GDTF FAQ · S3: GDTF & MVR Help Pages · S4: ANSI E1.11-2024: USITT DMX512-A Asynchronous Serial Digital Data Transmission Standard for Controlling Lighting Equipment and Accessories · S5: ANSI E1.17-2015 (R2025): Entertainment Technology—Architecture for Control Networks (ACN) · S6: Open Lighting Architecture Developer Documentation · S7: General Device Type Format (GDTF) Fixture Import for Eos Family · S8: Color Reproduction in LED Wall Virtual Production Stages
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: Post-hoc harmonized profiles place this candidate from rank 23 to 30, in ordering band C. The ranking is a reading-order aid using an affordability proxy, not a preregistered endpoint, deployment authorization, novelty determination, or measure of economic value.
29. Common-Wafer Contest for One Fabrication Slot¶
Canonical title: Common-Wafer Trial for Awarding a Scarce Nanofabrication Integration Slot
In one sentence: A nanofabrication facility would compare teams on blinded, equally resourced test wafers before awarding its single process-integration slot.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-03 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Bounded Rivalry Governance × Nanotechnology |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 20–29 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-07, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A shared nanofabrication facility has one integration bay and limited technician time. Teams now compete largely through proposals and results from samples they chose themselves. That can reward unusually favorable devices, extensive private characterization, or incomplete reporting of failed runs. It can also leave the facility and other users bearing contamination, waste, cleanup, and downtime. The facility may therefore select a persuasive process that cannot reproduce safely and reliably on shared equipment.
What is proposed¶
Replace proposal-only selection with a rule-bound common-wafer trial. Before entrants are known, publish eligibility, identical substrates, allowed process steps, equal tool and measurement budgets, safety gates, scoring, tie-breaks, confidentiality, penalties, and appeals. Code the samples and score complete attempt histories, functional yield across wafers, dimensional and electrical reproducibility, equipment compatibility, waste, cleanup, and recovery time. Unsafe performance cannot be offset by technical strength. Independently remeasure the leaders and audit their resource use. Give two teams bounded validation runs before awarding the single integration bay. Afterwards, compare trial scores with actual integration performance, examine whether the winner gained control over shared interfaces or rules, and reopen competition through a scheduled challenger window.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a deliberately governed contest for a genuinely scarce facility slot. Common specimens and resource caps define the arena; blinded scoring, audits, safety floors, spillover accounting, penalties, appeals, post-cycle review, and later challenger access constrain how teams may compete and what winning confers.
Why it advanced¶
It passed Experiment 6's strict researched-candidate bar because the allocation problem, responsible authorities, testing capabilities, safeguards, comparator, and falsifiable shadow study were sufficiently specified. That status concerns the researched candidate only; it does not establish field performance, novelty, deployment authority, adopter demand, or economic benefit.
Prior art and the remaining open claim¶
Proposal review, safety qualification, common-specimen comparisons, metering, contamination controls, equipment standards, and post-selection oversight already exist, so the proposal sits next to substantial prior art. The narrower open claim is that the full score—held-out reproducibility, all attempts, equal counted resources, tool compatibility, and recovery burden—will predict audited performance and avoid ranking reversals better than both proposal-only review and a simpler common-wafer technical score.
Smallest decisive test¶
With facility, safety, and data approval, run a non-awarding shadow study lasting at most 12 weeks and costing at most $50,000. Require at least eight archived entrants from two process families, comparable retained wafers or replicate measurements, and usable operational records. Freeze the rubric and audit split before revealing identities or existing ranks. Advance only if the full score improves held-out Kendall rank correlation over proposal review by at least 0.15, reduces audit reversals by at least 20%, is no worse than technical-only scoring, avoids a leader-changing process-family interaction, keeps essential missingness below 20%, and causes no safety or confidentiality breach.
Deployment and cost¶
The first shadow evidence step is estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms, not a vendor quote. Initial deployment is \(50,000–\)250,000; operational launch and annual operation are each \(250,000–\)1 million. Real use would also require local authority over recipes, records, appeals, reserves, and penalties.
Risks and uncertainties¶
- Common wafers may favor one process family or poorly represent integration conditions.
- A frozen score may encourage teams to optimize measured proxies instead of robust integration performance.
- Equal in-contest resource caps may still favor teams with stronger infrastructure outside the counted arena.
- Recipe submission and access-log audits may expose confidential know-how or identifiable personnel data.
- A remediation reserve could exclude less-capitalized teams or exceed the facility's legal authority to collect it. Managers have not established the needed terms yet, and the effects on those teams require field data. No source establishes how complete the facility's records are or how frequently proposal winners fail shared-condition replication. No facility has committed to host the study, and site-specific technician capacity and opportunity costs remain unknown.
Expert review¶
Useful reviewer backgrounds: Nanofabrication process-integration engineer, Facility operations and metrology manager, Environmental health and contamination-control specialist, Research-allocation and data-governance counsel.
- Can retained wafers from at least eight entrants be compared without introducing process-family-specific measurement bias?
- Which waste, cleanup, downtime, and compatibility measures can be reconstructed reliably from existing records?
- Would the proposed score predict successful integration better than technical yield and reproducibility alone?
- What authority does the facility have to inspect recipes, hear appeals, impose penalties, or require a remediation reserve?
Evidence and provenance¶
Selected sources: S1: User Proposal Review and Evaluation Process · S2: Safety in the NanoFab · S3: Safety and Policies · S4: Nanosensor Manufacturing Workshop: Finding Better Paths to Products · S5: Benchmarking the ACEnano Toolbox for Characterisation of Nanoparticle Size and Concentration by Interlaboratory Comparisons · S6: Fees · S7: FAQ—Individual Standards · S8: Report on the Nanoscience Research Centers
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized score is only a post-hoc reading-order aid: band C, ranking between 20 and 29 across profiles. Its pilot-speed input reflects cost-band affordability, not independently measured elapsed time or economic value, and it does not alter strict-success status.
30. A Stable Contract for Crystal Structures¶
Canonical title: Representation-Independent Contract for Ordered Periodic Material Structures
In one sentence: An opaque software contract would let materials tools exchange the same ordered periodic structure without depending on atom order, file layout, or a particular canonicalization method.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-08 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Representation Independent Interface Contract × Chemistry Materials |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 26–30 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
Materials software passes crystal structures among parsers, databases, simulation tools, and analysis programs. Clients often depend on incidental details such as atom-array order, coordinate wrapping, lattice orientation, cell convention, or serialized field layout. The same physical ordered structure can then receive different identifiers or downstream treatment, while an internal parser or storage change can break clients even when the intended material state has not changed. It also becomes difficult to distinguish a physical change from an encoding change.
What is proposed¶
Define an immutable ordered-periodic-material value by its periodic decorated points in physical space, including species, occupancies, units, tolerance policy, and provenance. Expose validated construction, composition, physical-equivalence tests, invariant summaries, explicit conversions, and serialization, but hide arrays, atom order, coordinate basis, caches, canonical labels, and format-specific fields. Specify preconditions, typed errors, no-mutation-on-failure behavior, and laws for atom permutation, origin shifts, periodic wrapping, unit conversion, and admissible cell changes. Every parser or store must pass the same public-only black-box and metamorphic tests before substitution. Version changes to promised behavior and probe error text, ordering, serialization, debug access, and timing for leaks. Version zero rejects disorder, partial occupancy, trajectories, surfaces, and inferred bonding rather than silently approximating them.
The cross-domain transfer¶
The representation-independent-interface archetype maps strongly here. The abstract component is the physical ordered periodic structure; concrete arrays, graphs, files, and database rows remain hidden. Behavioral laws, typed failures, shared conformance tests, leakage review, and versioning determine whether independently built implementations may substitute for one another.
Why it advanced¶
It passed Experiment 6's strict researched-candidate bar because close comparators, a bounded scope, scientific authority, reversible sandbox, measurable failure conditions, and a live two-adapter benchmark were identified. This endpoint does not show that production repositories suffer widespread coupling or that the contract is novel, performant, deployable, or adopted.
Prior art and the remaining open claim¶
Representation-insensitive structure matching, pymatgen, spglib, CIF, OPTIMADE, and unified materials-data systems already address important parts of the problem. The remaining claim is narrower: two independently structured adapters can obey one public behavioral oracle for equivalence, invariants, errors, immutability, and provenance while leaking fewer internal details and producing fewer semantic disagreements than the existing array baseline, a canonicalization path, or a schema-only round trip.
Smallest decisive test¶
With a platform partner and scientific-method lead, preregister the abstract state, tolerance rules, observables, and exclusions. Test two independently structured adapters on 12 copied fixtures, five seeded representation changes per fixture, fixed assertions, and 200 operation sequences. Compare them with the existing array/serializer behavior, spglib plus pymatgen, and an OPTIMADE/CIF round trip. Reject version zero if any meaning-preserving transformation changes a promised observable, any fixture triple exposes tolerance-driven non-transitivity, provenance is lost, hidden representation must be exposed, or the contract fails to reduce semantic divergences and leakage relative to the best comparator. Also record runtime and memory.
Deployment and cost¶
The read-only sandbox is estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms. Initial deployment is \(50,000–\)250,000; operational launch is \(250,000–\)1 million; annual operation is \(50,000–\)250,000. Production identifier changes, deduplication, parser replacement, or database migration are expressly outside the first test.
Risks and uncertainties¶
- Tolerance-based equivalence may be non-transitive, producing unstable identity groups.
- The initial abstract state may omit meaningful distinctions such as defects, chirality, magnetic order, isotope labels, or provenance.
- An over-specified suite could freeze incidental numerical or serialization behavior.
- An under-specified suite could pass simple structures while missing difficult-cell disagreements.
- Opaque access may obstruct legitimate diagnostics and encourage unsupported escape hatches. Real repository data have not yet shown how frequent or costly representation coupling is. Scientific reviewers have not approved the proposed version-zero semantics, and comparator performance remains unmeasured. Passing finite tests would not establish chemical identity, scientific equivalence, implementation correctness, or acceptable production performance.
Expert review¶
Useful reviewer backgrounds: Computational crystallographer, Materials-data platform architect, Scientific software testing specialist, Materials provenance and standards expert.
- Does the proposed abstract state preserve every scientifically relevant distinction in the selected ordered-periodic scope?
- Can the tolerance policy avoid non-transitive equivalence for realistic near-boundary structures?
- Which current clients depend on atom order, serializer layout, canonical labels, or other hidden details?
- Does the contract reduce semantic divergence and leakage without unacceptable runtime, memory, or diagnostic costs?
Evidence and provenance¶
Selected sources: S1: Identifying duplicate crystal structures: XtalComp, an open-source solution · S2: pymatgen.core package: structure_matcher module · S3: Spglib conventions of standardized unit cell · S4: OPTIMADE API specification v1.3.0 · S5: OPTIMADE, an API for exchanging materials data · S6: Core CIF dictionary · S7: NOMAD — Materials science data, managed and shared · S8: Employer Costs for Employee Compensation — March 2026
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized result is a post-hoc ordering aid: band C, ranks 26–30 across profiles. It is not an experimental endpoint or economic-value measure; its pilot-speed input is an affordability proxy, and the candidate remains a strict success only under Experiment 6's researched bar.
31. Show Reviewers Only Unexpected Proof Effects¶
Canonical title: Residual Impact Maps for Mathematical Theory Revisions
In one sentence: A complete proof-library check would be reconstructed from frozen theorem-level predictions and typed discrepancies, allowing reviewers to focus on surprises while protected changes always remain visible.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-17 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Mathematics |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 66.0/100 · rank range 27–34 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
Revising an axiom, definition, notation rule, or trusted dependency can affect hundreds of machine-checked theorems. A checker can revalidate the whole library, but maintainers may still face a large migration report. Dependency graphs can flag harmless reachability and miss effects from elaboration, automation, or undeclared coupling. Reading every expected result wastes attention, yet showing only predicted failures could hide an unexpected survival, a new dependency, a removed obligation, or a theorem that still passes for the wrong reason.
What is proposed¶
Before migration, freeze a versioned model predicting every theorem's status, changed obligations, and dependency differences. Independently run the trusted checker over the complete authorized corpus. Compare each prediction with the actual result and record typed residuals for unexpected failures or survivals, changed diagnostics, obligations, dependencies, trust assumptions, timeouts, missing results, or confidence disagreements. Reconstruct the full ledger from prediction plus residual, while directing review primarily to consequential mismatches. Statements, axioms, admitted facts, trust changes, removed obligations, checker failures, missing results, and out-of-scope theorems always appear in full. Random predicted-unaffected records and boundary cases receive independent audits. Version, coverage, reconstruction, drift, or audit failures automatically restore the complete theorem-by-theorem report for the affected component.
The cross-domain transfer¶
Predictive residual processing becomes a review codec for mathematical migrations. A frozen impact model supplies the expected theorem ledger; complete checking supplies reality; typed differences carry surprises. Checksums, audits, resynchronization, protected-event bypasses, drift monitoring, and raw-report fallback keep compression from becoming selective proof checking.
Why it advanced¶
It did not enter the strict-success lane. It cleared a separately calibrated empirical-partner-candidate lane because a reversible archived replay and decision thresholds are specified, but the decisive field evidence is missing: no measured report burden, predictable-record share, predictor calibration, reviewer study, audit rate, cost benchmark, maintainer funding, or partner authorization exists.
Prior art and the remaining open claim¶
Static dependency analysis, grouped checker reports, incremental proof checking, iCoq-style regression selection, source diffs, and trust checklists are established neighbors. The open comparison is whether complete independent checking plus frozen predictions, typed residuals, protected full records, random audits, exact reconstruction, and component fallback can reduce report volume and review time without lowering protected-event recall or blinded classification accuracy relative to full and dependency-grouped reports.
Smallest decisive test¶
With maintainer authorization, replay an archived, non-release-blocking migration containing at least 500 declarations. Freeze the model, residual types, protected classes, thresholds, and audit sample before opening the evaluation partition. Randomize blinded reviewers among full reports, dependency-grouped reports, and the residual interface; insert cases covering failures, survivals, dependency and obligation changes, axioms, timeouts, missing outputs, manifest gaps, and version mismatches. Require 100% protected-event recall, zero missing-as-success errors, exact ledger reconstruction, no consequential audit miss, classification no more than five percentage points below the best comparator, and at least 20% lower median review time or report volume. Success permits only another shadow study.
Deployment and cost¶
The archived replay is estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms. Initial deployment is \(50,000–\)250,000; operational launch is \(250,000–\)1 million; annual operation is \(50,000–\)250,000. The model may prioritize review but may never approve revisions, waive checker results, or edit proofs automatically.
Risks and uncertainties¶
- A shared predictor could create correlated blind spots across whole theory components.
- Unexpected theorem survival may conceal a weakened statement or unintended dependency.
- Automation and elaboration may create semantic coupling absent from declared dependency graphs.
- The residual schema may preserve checker status but omit mathematical context needed for judgment.
- Reviewers may anchor on predictions, while thresholds may be tuned to shrink the queue. No evidence yet quantifies present reviewer burden, repeated predictable content, or protected-event frequency. The theorem-level predictor has not been calibrated across the required outcome types, and no human comparison has tested whether residual presentation preserves decisions. Audit rates, repository confidentiality rules, operating costs, partner authorization, and maintainer funding remain unresolved.
Expert review¶
Useful reviewer backgrounds: Formal-mathematics library maintainer, Proof-assistant kernel and elaboration expert, Human-factors researcher for technical review, Statistical audit and anomaly-detection specialist.
- What fraction of a real foundational migration report is predictable repetition, and how much reviewer time does it consume?
- Can the predictor detect unexpected survivals, trust changes, removed obligations, timeouts, and missing outputs with calibrated uncertainty?
- What random and boundary-focused audit rate would detect rare consequential blind spots?
- Does the residual interface preserve blinded reviewer accuracy while reducing total review, model-maintenance, audit, and fallback cost?
Evidence and provenance¶
Selected sources: S1: Maintaining a Library of Formal Mathematics · S2: mathlib4: The Math Library of Lean 4 · S3: RFC: Add left actions and right actions to expression tree elaborator and make ^ be a right action · S4: Pull Request Review Guide · S5: iCoq: Regression Proof Selection for Large-Scale Verification Projects · S6: Practical Machine-Checked Formalization of Change Impact Analysis · S7: The Isabelle System Manual (Isabelle2022) · S8: Axioms — The Lean Language Reference
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized score is a post-hoc reading-order aid only: band C, ranks 27–34 across profiles. The pilot-speed input approximates affordability rather than measured duration. This ordering does not convert the candidate into a strict success or indicate economic value.
32. A Recurring Reckoning for Airport Noise Promises¶
Canonical title: Shared-Sky Accountability Observance for Airport Noise Commitments
In one sentence: A short, consent-based observance would connect remembered airport-noise experiences and past commitments to an authorized public decision and a concrete follow-up ledger.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-13 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Aviation Aeronautics |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 66.0/100 · rank range 23–38 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-12, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Airports can collect noise measurements, complaints, and public comments while losing a shared memory of what communities experienced and what institutions promised. Staff and resident turnover may scatter commitments across minutes and dashboards. Conventional meetings can repeatedly ask people to recount distress without creating a bounded occasion to acknowledge it, renew or revise a promise, record disagreement, and assign follow-through. Participants may then disagree about a commitment's origin, scope, owner, authority, resources, or next review date.
What is proposed¶
Embed a quarterly 30-minute observance in an existing airport-community forum. A rotating community-institution pair first documents its purpose, participant standing, story and recording provenance, access options, obligations, and retirement rule. A restrained threshold introduces 60 seconds of optional listening to consent-cleared low-intensity audio or viewing an acoustic trace; silence, private reflection, remote attendance, or leaving are equivalent choices. Paired community and institutional accounts then reconstruct one commitment, preserving corrections and disagreement. Consenting witnesses acknowledge the experience and exact promise. An authorized representative must renew, revise, or decline it; community members may dissent or remain silent. Closure updates a public ledger with authority limits, owner, next action, resource dependency, forum, review date, and unresolved disagreement, followed by a harm-and-meaning debrief and periodic independent audit.
The cross-domain transfer¶
The ritualized-commitment archetype becomes a marked, recurring accountability occasion rather than an operational aviation decision. Threshold, optional reflection, paired histories, witnessing, explicit institutional recommitment, ledger closure, debrief, rotating stewardship, harm audit, and governed retirement connect symbolic recognition to ordinary authorized follow-through.
Why it advanced¶
It passed Experiment 6's strict researched-candidate bar because the intervention, authority boundary, conventional comparator, consent protections, stopping rules, and matched rehearsal are explicit and testable. This does not show noise reduction, community-wide legitimacy, real commitment fulfillment, adopter willingness, novelty, deployment authorization, or economic impact.
Prior art and the remaining open claim¶
Noise dashboards, complaint systems, community roundtables, facilitated planning, written agreements, action ledgers, and one-time listening sessions already exist. The remaining claim is incremental: adding the governed threshold, optional sensory interval, paired provenance accounts, witnessing, an explicit renew-revise-decline choice, and immediate debrief will improve accurate, retained understanding of one commitment beyond an equally timed facilitated ledger review, without increasing coercion, distress, false representation, or confusion that acknowledgment equals mitigation.
Smallest decisive test¶
With a willing forum, run a preregistered tabletop and randomized matched rehearsal with 12–20 consenting community and institutional participants, using a fictional commitment and synthetic low-intensity media. Compare the full observance with an equally timed agenda-and-ledger review containing identical facts. Score eight commitment facts immediately and after 7–14 days, plus authority understanding, fairness, comfort, manipulation, representation, disclosure pressure, accessibility, and freedom to opt out. Proceed only if median recall improves by at least two facts, at least 80% understand what was not decided, coercion and representation scores are no worse, and no serious distress or retaliation concern occurs. The observer may require revision or stop.
Deployment and cost¶
First evidence, initial deployment, operational launch, and annual recurring operation are each estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms, not external quotes. The observance cannot change routes, schedules, funding, regulation, environmental findings, or airline operations, and it cannot replace complaints, consultation, mediation, or technical analysis.
Risks and uncertainties¶
- The observance could aestheticize residents' distress or turn it into institutional performance.
- Audio or repeated recollection could cause sensory or emotional harm despite opt-out choices.
- Selected recordings, histories, or visible participants could be treated as representing absent communities.
- Ceremonial acknowledgment might be mistaken for mitigation, legal acceptance, or resolution.
- An authorized-looking renewal could conceal missing funding, regulatory power, or operational feasibility. No site has shown that commitment-memory failures occur at the assumed rate, and no airport or community group has agreed to host the pilot. Live evidence is absent on nonretaliatory refusal, adverse events, comparative benefit over a disciplined ledger meeting, and whether real representatives can decide without bypassing other authorities. Site-specific costs are also unknown.
Expert review¶
Useful reviewer backgrounds: Airport community-engagement and noise-management lead, Affected-community representative with independent standing, Accessibility, trauma-informed facilitation, and emotional-safety specialist, Aviation governance and authority-boundary expert.
- Do participants currently fail to reconstruct commitment history, scope, authority, ownership, dependencies, and unresolved disagreement from ordinary records?
- Does the ritual sequence improve delayed factual recall beyond an equally timed facilitated ledger review?
- Can audio-free participation, silence, dissent, and exit remain genuinely nonretaliatory under the forum's power relationships?
- Which representative can renew, revise, or decline a real commitment, and which decisions must remain with airports, airlines, regulators, boards, or air-traffic authorities?
Evidence and provenance¶
Selected sources: S1: Neighborhood Environmental Survey · S2: Aircraft Noise: Military Helicopter Operators Should Improve Outreach to Affected Communities in the D.C. Area (GAO-26-107758) · S3: Advisory Circular 150/5050-4A: Community Involvement in Airport Planning · S4: What is a community roundtable? · S5: SFO Airport/Community Roundtable · S6: Community Engagement on Aircraft Noise · S7: CAP3041: Guidance for Airport Engagement and Complaints Handling Around Environmental Sustainability · S8: Being a Fair Neighbor—Development and Validation of the Aircraft Noise-Related Fairness Inventory (fAIR-In)
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized result is a post-hoc reading-order aid: band C, ranks 23–38 across profiles. It neither measures economic value nor changes the strict-success endpoint; the pilot-speed input is an affordability proxy, and elapsed pilot time was not independently scored.
33. Retiring Outdated Criminal-Record Labels¶
Canonical title: Label Sunset: Criminal-Record Authority Retirement Cycle
In one sentence: A voluntary quarterly exercise would help record custodians distinguish an ended authorization from deletion, preserve lawful history, and assign unresolved downstream corrections to named owners.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-14 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Criminology Forensic |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 66.0/100 · rank range 31–33 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A court or other competent authority may restrict or end the permitted use of a criminal-justice label, yet copies, vendor feeds, access permissions, derived flags, and staff assumptions can persist across disconnected systems. Correcting the source record does not necessarily tell every downstream custodian what changed. The result can be continued reliance on a superseded status; however, careless cleanup can also destroy records that must lawfully remain available for provenance, oversight, or defined exceptions.
What is proposed¶
Authorized custodians would hold a quarterly, voluntary “Label Sunset” cycle after counsel confirms which disposition categories are eligible. Using only invented records, participants would move a fictional label through a map of source, repository, vendor, and user systems. A records witness would verify a mock retirement instruction, remove the label from the active layer, and place it in an archive sleeve marked “retired—provenance preserved.” Participants could decline any part of the exercise. The group would recognize only verified batch-level reconciliation and route every unresolved exception category into protected workflows with an owner, resources, and deadline. An immediate debrief and independent audit could revise, pause, or end the cycle. Unlike an ordinary reconciliation briefing, the proposal adds a witnessed enactment, recommitment, recognition, and explicit exception handoff; it changes no legal status or production record itself.
The cross-domain transfer¶
The ritualized meaning-and-commitment archetype becomes a governed enactment of how a label spreads and loses operational authority. Its marked object, witnesses, voluntary commitment, closure, debrief, and retirement path map clearly to the domain, while real legal powers, system privileges, and individual remedies remain entirely outside the ritual.
Why it advanced¶
This was a STRICT_SUCCESS because it passed Experiment 6’s strict researched-candidate bar. That endpoint reflects the quality and testability of the researched proposal, not field validation, novelty, legal authorization, adoption, or demonstrated effects on record accuracy, employment, stigma, recidivism, or other real-world outcomes.
Prior art and the remaining open claim¶
Legal relief workflows, automated access and retention rules, reconciliation audits, privacy training, and record-clearing programs already perform much of the practical work. The narrower open claim is that adding a voluntary fictional propagation exercise, witnessed retirement, recommitment, aggregate recognition, exception routing, and debrief to a matched reconciliation briefing will improve accurate legal distinctions, delayed recall, and assignment of downstream exceptions without increasing coercion, stigma, privacy risk, or beliefs that the ceremony itself completed legal relief. There is no field evidence for that comparison.
Smallest decisive test¶
A court or repository unit would recruit 24–40 volunteers from at least four custodial functions and randomize intact teams to the 45-minute cycle or a content-, facilitator-, and time-matched briefing. Blind scoring would cover six invented scenarios before, immediately after, and two weeks later. Advance only if the cycle gains at least 0.4 standard deviations on the delayed composite or 15 percentage points in fully correct exception routing, with no material safety loss. Retire the mechanism if neither threshold is met, real records enter the exercise, more than 10% feel dissent is unsafe, or false-relief beliefs exceed the limit.
Deployment and cost¶
The authorized first step is a synthetic pilot with invented cases and no production queries or changes. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup and operational launch, and \(10,000–\)50,000 annually. These are assessment bands, not vendor quotes; no institution has committed funding or adoption.
Risks and uncertainties¶
- Participants may mistake symbolic retirement for sealing, deletion, expungement, or complete downstream correction.
- Small aggregate exception categories could expose protected case information.
- Staff may experience participation, affirmation, or dissent as an employment loyalty test.
- Recognition could create ceremonial closure while vendor copies, derived flags, or informal assumptions persist.
- The archive metaphor could encourage destruction or concealment of records that must lawfully remain preserved or accessible under an exception.
Expert review¶
Useful reviewer backgrounds: Criminal-records counsel, Court or repository data-governance lead, Privacy and information-governance specialist, Background-screening systems expert, Lived-experience or record-subject advocate.
- Can the six scenarios reliably distinguish correction, sealing, restricted use, deletion, retention, and lawful exceptions?
- Which custodian has authority to own each downstream exception without exposing protected case data?
- Would staff reasonably perceive passing, dissenting, or challenging the script as safe and nonretaliatory?
- Can the matched briefing contain exactly the same legal and system information so the enactment is the only material difference?
- What evidence would show that the ceremony is hiding incomplete vendor or repository reconciliation?
Evidence and provenance¶
Selected sources: S1: Criminal History Records: Additional Actions Could Enhance the Completeness of Records Used for Employment-Related Background Checks · S2: Fair Credit Reporting; Background Screening · S3: Technical and Operational Challenges of Implementing Clean Slate: Research Findings · S4: Making the Promise of Expungement a Reality: A Guide to Record Relief in the State Courts · S5: NIST Privacy Framework: A Tool for Improving Privacy through Enterprise Risk Management, Version 1.0 · S6: New Jersey P.L. 2019, c.269: Automated Clean Slate Process · S8: Occupational Employment and Wages — May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed it 31st–33rd, band C, across post-hoc profiles. That score only sets reading order and uses an affordability proxy; it is neither an experimental endpoint nor a measure of economic value, deployment readiness, or impact.
34. Stress-Testing Flight-Control Downselects¶
Canonical title: Envelope-Robustness Downselect Arena for Flight-Control Candidates
In one sentence: A shadow contest would compare whether hidden scenarios, safety gates, resource caps, and two-candidate advancement select flight controllers more consistently than public benchmarks or expert panels.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-04 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Aviation Aeronautics |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 65.0/100 · rank range 22–40 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
Aircraft programs must decide which competing flight-control candidates receive scarce hardware-in-the-loop and flight-test resources. If teams know every scored simulation case, they may tune to those cases, simulator artifacts, or inferred test structure, while escalating compute and engineering effort. A visible-score winner may therefore be less reliable under new conditions than its rank suggests. The actual frequency of this problem in flight-control programs is unknown, and general competition research supplies meaningful counterevidence against assuming severe leaderboard overfitting is routine.
What is proposed¶
The program would freeze eligibility, configuration deadlines, allowed resources, safety floors, scoring, audit rules, fouls, appeals, and judge authority before submissions. Teams could develop against a disclosed suite, but an independent custodian would rank frozen builds using undisclosed seeds and disturbance combinations. Every controller would first face noncompensable safety gates. Passing candidates would then be scored for hidden-condition robustness, simulated pilot-intervention demand, control activity, reproducibility, and declared integration burden under resource caps. Two materially different candidates, rather than one leaderboard winner, would receive hypothetical advancement in the first shadow study. Independent reproduction and later hardware-in-the-loop, pilot, and maintainer evidence would govern any real progression. The comparator includes a fully public single-winner benchmark and an unranked expert-panel choice using the same builds, scenario families, safety constraints, and nominal compute limits. No contest score authorizes flight.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a controlled engineering contest with a scarce prize, fixed legal moves, safety floors, resource limits, audits, appeals, multiple winners, and a later challenger window. The mapping is strong, but most individual governance features already appear in aviation challenges, technical competitions, or staged verification programs.
Why it advanced¶
This is an EMPIRICAL_PARTNER_CANDIDATE, not a strict success. It cleared a separately calibrated lane for a bounded external data-partner study because the comparison is testable, but it lacks proprietary frozen controllers, program-specific scenarios, validated resource accounting, downstream hardware evidence, and any partner commitment.
Prior art and the remaining open claim¶
Robust flight-control challenges, protected leaderboards, aviation contest rules, safety stop authority, formal protests, and staged simulation-to-hardware validation already exist. The remaining claim concerns their combination: holding builds and approved conditions constant, a hidden, safety-gated, resource-capped arena that advances two complementary candidates will produce more stable selections on independently seeded validation and later authorized hardware or human evaluation than either a public-suite winner or an expert-panel choice. That comparative effect has not been measured, and portfolio discretion could simply advance a favored lower scorer.
Smallest decisive test¶
With an aircraft-program partner, preregister an offline shadow study using three to six frozen controllers. Apply the proposed arena, the disclosed single-winner benchmark, and an unranked panel to identical approved materials, then test their selections on a second inaccessible scenario set. Measure selection agreement, rank stability, safety failures, reproducibility, intervention demand, actuator activity, resources, disputes, and sensitivity to lawful seeds and weights. Do not seek hardware authority unless the arena is more stable or predictive. Falsify it if innocuous seeds reverse choices, results do not beat both comparators, portfolio selection adds an inferior redundant candidate, or resource accounting systematically favors incumbents. Void results after leakage, unverifiable builds, compensable safety violations, or disputed simulator validity.
Deployment and cost¶
The first step is offline shadow evaluation with no contract, integration, certification, personnel, or flight consequences. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. They are not facility or supplier quotes.
Risks and uncertainties¶
- Secret scenarios may contain the same modeling errors as the disclosed simulator while making those errors harder to challenge.
- Existing tools and pretrained models could let incumbents exceed the practical resource cap without recorded spending.
- A discretionary definition of “complementary” could justify advancing a favored lower-scoring controller.
- Teams may tune to the inferred scenario generator rather than to operationally relevant robustness.
- Simulation rank may fail to predict hardware behavior, pilot workload, maintainability, integration effort, or certification findings.
Expert review¶
Useful reviewer backgrounds: Flight-control engineer, Aircraft verification and validation lead, Test pilot or pilot-in-the-loop specialist, Airworthiness and certification specialist, Simulation-model validation expert.
- Which approved scenario variations are independent enough to test robustness without leaving the simulator’s validity envelope?
- Can legacy tools, reusable models, supplier labor, compute, and sponsor support be counted fairly across entrants?
- How should complementarity be defined before scores are known so it cannot become discretionary favoritism?
- What rank-stability or downstream-prediction improvement would justify the added contest administration?
- Which results may inform a later hardware study without being mistaken for certification evidence or flight authorization?
Evidence and provenance¶
Selected sources: S1: AC 25.1309-1B—System Design and Analysis · S2: AC 25.1329-1B—Approval of Flight Guidance Systems · S3: NASA Systems Engineering Handbook, Section 5.0—Product Realization · S5: Robust Flight Control Design Challenge: Problem Formulation and Manual—The Research Civil Aircraft Model · S6: The Ladder: A Reliable Leaderboard for Machine Learning Competitions · S7: A Meta-Analysis of Overfitting in Machine Learning · S8: DARPA Lift Challenge—Competitors, Rules, Scoring, Safety, and Protest Procedure
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized profiles rank this candidate from 22nd to 40th, band C. The wide range reflects different review weights. It is a reading-order aid using a cost proxy, not an experimental endpoint, economic-value estimate, or partner commitment.
35. Tracking Aging Wetland Erosion Mats¶
Canonical title: Lifecycle-Governed Retirement of Wetland Erosion-Control Mats
In one sentence: A site register would identify successive erosion-control layers, trigger review as they age, block unsafe removal, and preserve a record after authorized retrieval.
| Field | Record |
|---|---|
| Portfolio ID | EXP05-STRICT-03 |
| Experiment and endpoint | Experiment 5 · Strict success |
| Archetype × domain | Layer Decay And Expiration Management × Environmental Climate |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 64.0/100 · rank range 35–37 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
Wetland and streambank repairs can leave several generations of blankets, coir mats, synthetic netting, and anchors in the same location. Older layers may be exposed, buried, torn, rooted through, or missing from project records. Crews may cover an unidentified layer again, leaving fragmenting mesh in habitat, or remove material that still holds vegetation and sediment. The field prevalence of such sequential legacy layers is not yet established for a candidate site, and hydrology or geomorphology may matter more than installed material.
What is proposed¶
At one bounded restoration reach, managers would give each mapped installation a stable identity, footprint, material class, deposition order, estimated age, condition, access history, and lifecycle state. A class-specific time-to-review trigger would mark a layer for inspection, never automatic removal. An age-weighted value-and-risk score would order limited field work, but qualified reviewers would first check whether roots, sediment, adjacent structures, monitoring records, or permits still depend on the material. The responsible manager could continue use, schedule another inspection, retire the layer in place, preserve it as a time-limited exception, or seek separate permission for staged retrieval. Any removal would leave a geospatial “tombstone” recording the former footprint, reason, evidence, and successor layer. The comparator is ordinary project-file review plus current surface inspection without unified identities, expiry triggers, dependency gates, exception states, or tombstones.
The cross-domain transfer¶
The layer-decay archetype maps directly onto successive physical installations that can outlive their original role. Stable identities, review deadlines, risk ordering, dependency checks, differentiated disposition, expiring exceptions, and removal records translate cleanly. The analogy does not make age proof of obsolescence or turn a database state into authority to disturb habitat.
Why it advanced¶
This was a STRICT_SUCCESS because it passed Experiment 5’s strict researched-candidate bar. That status does not show that multiple legacy layers are common at real reaches, that reviewers can classify them reliably, that an adopter will use the register, or that it improves bank stability, wildlife outcomes, or costs.
Prior art and the remaining open claim¶
Existing practice already includes safer material substitution, functional-longevity classes, inspection and maintenance records, removal requirements, permits, monitoring plans, and asset databases. The narrower open claim is that adding installation-level identities, review-only expiry triggers, a mandatory ecological, structural, and regulatory dependency check, time-limited exceptions, and post-removal tombstones will yield more reproducible and dependency-resolved recommendations than ordinary files and surface inspection. It expressly does not claim that older material should be removed or that the register itself reduces erosion, plastic pollution, or wildlife mortality.
Smallest decisive test¶
At a partner-managed reach, sample 20–50 mapped or suspected installations. Two qualified reviewers would independently assess the same nondestructive evidence under randomized workflows: ordinary files plus surface inspection, and the same material augmented by the lifecycle register and dependency checklist. Compare unique layers found, unresolved records, agreement, review time, live dependencies, and actionable recommendations. Stop at observation; do not probe or remove. Falsify the near-term claim if no sequential accumulation exists, the added workflow finds no decision-relevant layers or dependencies, agreement remains below a preregistered threshold such as kappa 0.60, review time more than doubles without better resolution, or either reviewer recommends retrieval despite a live or unresolved dependency.
Deployment and cost¶
Begin with a read-only inventory using existing plans, photographs, permitted observation, and provisional uncertainty labels. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup and operational launch, and \(10,000–\)50,000 annually. These are not quotations; no site partner or data-stewardship arrangement is secured.
Risks and uncertainties¶
- Surface observation may miss buried or visually similar layers and create false confidence in completeness.
- An age-weighted score may disguise subjective judgments as measurement precision.
- Staff may treat a review deadline as an instruction to remove material despite the dependency gate.
- “Retire in place” may become indefinite neglect if its exception and next review date are not enforced.
- Attention to installed mats may divert investigation from hydrologic or geomorphic causes of site failure.
Expert review¶
Useful reviewer backgrounds: Wetland or stream restoration ecologist, Geotechnical or hydraulic engineer, Environmental permitting specialist, Erosion-control contractor, GIS and restoration-records manager.
- Can reviewers identify deposition order and lifecycle state nondestructively and with acceptable agreement?
- Which observed conditions are sufficient to establish that roots, sediment, structures, or permits still depend on a layer?
- How should unknown installation dates and incomplete specifications affect priority without implying obsolescence?
- What authority and evidence are required before staged retrieval at the selected reach?
- Does the register add useful information beyond the site’s existing maintenance, permit, monitoring, and asset records?
Evidence and provenance¶
Selected sources: S1: Use of Sustainable Materials for Erosion and Sediment Control Practices · S2: Make the Change to Wildlife-Friendly Erosion Control Products! · S4: Water Quality General Certification No. 4100REV: Wetland and Stream Restoration and Creation · S5: Survey of wildlife-friendly temporary and permanent erosion control products and approaches · S6: Erosion Control Toolbox: Rolled Erosion Control Product (Blanket) · S7: Erosion prevention practices—erosion control blankets and anchoring devices · S8: Natural Channel Design Review Checklist
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized profiles placed the candidate 35th–37th, band C. This narrow reading-order range uses an affordability proxy and does not alter its strict experimental endpoint or measure field prevalence, ecological benefit, economic value, or deployment readiness.
36. Subtracting Recoater Self-Noise¶
Canonical title: Efference-Copy Residual Sensing for Powder-Layer Recoating
In one sentence: A shadow monitoring system would predict signals caused by a powder recoater’s own commands, route unexplained residuals for review, and revert to complete signals whenever validity or safety checks fail.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-13 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Chemistry Materials |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 64.0/100 · rank range 17–48 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-11, EXP06-PARTNER-12, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
During powder-layer recoating, commanded acceleration and routine powder contact create large, repeatable force, current, vibration, and acoustic signals. These self-generated patterns can conceal smaller signs of an agglomerate, protrusion, foreign object, uneven layer, or changed powder flow. Independent static thresholds can instead produce repeated benign alerts. It is not yet known whether command-caused motion explains enough signal on a target machine to create a useful residual, or whether complete raw monitoring already fits the available data and operator-attention capacity.
What is proposed¶
Each outgoing motion command would be copied into a frozen, versioned forward model that predicts the synchronized sensor response expected for one declared machine, recoater, powder class, speed range, and environment. The system would subtract that prediction from the measured full window, preserve signed and time-aligned residuals, and weight them by sensor reliability, timing, coherence, persistence, consequence, and processing cost. A compatible receiver could reconstruct the full signal from the prediction and residual. Validated residual clusters would route to an operator with a defined inspection action. Random and risk-based raw windows, periodic full-state anchors, checksums, heartbeats, and drift tests would audit the cancellation. Any timing, model, sensor, reconstruction, scope, or protected-safety failure would restore complete processing. Model updates would remain slower, reviewed, and reversible so recurring defects could not immediately be learned away. Comparators are full raw review and existing static thresholds.
The cross-domain transfer¶
Predictive residual processing becomes an efference-copy system: the controller’s own command predicts the machine-generated sensory return, and the unexplained remainder receives scarce transmission and operator attention. The structural mapping is strong, but adjacent disturbance observers, recoater sensors, digital twins, anomaly detectors, and predictive codecs already supply most components separately.
Why it advanced¶
This is an EMPIRICAL_PARTNER_CANDIDATE, not a strict success. It cleared a separate lane for a bounded external laboratory study, but lacks a target-machine dataset, frozen model, controlled challenge results, comparative reviewer evidence, measured capacity savings, validated fallback reliability, and an OEM or laboratory commitment.
Prior art and the remaining open claim¶
Recoater force and vibration sensing, powder-bed imaging, digital-twin recoating control, model-based collision detection, industrial interfaces, and bounded powder-spreading testbeds already exist. The remaining claim is narrower: on one fixed configuration, command-conditioned cancellation with reconstruction, raw audits, version checks, and forced fallback will detect and localize specified external interactions no worse than complete raw review while reducing total data and reviewer burden after computation, audits, fallbacks, inspections, and maintenance are counted. No live recoater experiment establishes that combined comparison, especially for events synchronized with commanded acceleration.
Smallest decisive test¶
An OEM or accredited laboratory would run a preregistered shadow test on one fixed recoater, sensor layout, motion program, medium, and environment while retaining every raw channel. Blinded challenges would include localized resistance, altered layer height, obstacle surrogates, powder-drag changes, sensor bias, timing offsets, dropped residuals, version mismatch, and missing heartbeats. Compare the reconstructed residual path, complete raw review, and static thresholds on protected-event capture, classification, localization, false alerts, reconstruction, fallback behavior, bytes, review time, inspections, and maintenance. Reject advancement if any protected or synchronization challenge fails to trigger full mode, any controlled interaction is over-cancelled, sensitivity or localization breaches the preregistered non-inferiority margin, or total capacity cost is not lower.
Deployment and cost¶
The first authorized step is a non-production shadow test; existing controls, interlocks, raw storage, and operator decisions remain authoritative. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. They are not OEM or supplier quotes.
Risks and uncertainties¶
- A wrong or mistimed model may subtract a real interaction that resembles the expected response to acceleration.
- Powder lot, humidity, wear, mounting, or sensor coupling may invalidate the learned self-signal.
- Sender and receiver can share the same wrong model while still passing a compatibility checksum.
- Rapid model adaptation could normalize recurring defects or wear instead of exposing them.
- Audit sampling may miss rare over-cancelled events, while frequent fallbacks may increase workload and pressure operators to weaken safeguards.
Expert review¶
Useful reviewer backgrounds: Additive-manufacturing process engineer, Machine-controls engineer, Recoater and powder-spreading specialist, Industrial sensing and signal-processing researcher, Machine safety and functional-safety engineer.
- What fraction of each target sensor’s variance is reproducibly explained by the outgoing command under nominal conditions?
- Which controlled interactions are most likely to resemble command-caused acceleration signals and be over-cancelled?
- What non-inferiority margin is acceptable for detection and localization against complete raw review?
- Can clocks, position traces, checksums, heartbeats, and full-state anchors force fallback within the required latency?
- After computation, raw audits, fallbacks, inspections, and maintenance, does the residual path actually reduce total capacity use?
Evidence and provenance¶
Selected sources: S1: In-process monitoring and non-destructive evaluation for metal additive manufacturing processes (NIST IR 8538) · S2: Software Releases · S3: 4040 Development & Demonstration of an Open Layered Protocol for Powder Bed AM · S4: Recoater Force Sensor Array for Spatial and Temporal In-Situ Powder Spreading Quality and Surface Defect Monitoring, US application 20240300024 · S5: Monitoring Laser Powder Bed Fusion Recoater Blade Vibrations for Collision Avoidance · S6: Smart Recoating: A Digital Twin Framework for Optimisation and Control of Powder Spreading in Metal Additive Manufacturing · S7: Practical Aspects of Model-Based Collision Detection · S8: Powder Spreading Testbed for Studying the Powder Spreading Process in Powder Bed Fusion Machines
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: Post-hoc harmonized profiles rank this candidate from 17th to 48th, band C, showing strong sensitivity to review priorities. The score is only a reading-order aid with a cost proxy, not an endpoint, economic-value measure, validation result, or deployment recommendation.
37. A Shared Pause for Audit Independence¶
Canonical title: Absent-User Independence Boundary Cycle
In one sentence: At key engagement changes, an audit team would visibly separate client-service pressures from its duty to absent report users, route uncertain matters through official channels, and make consultation or role substitution socially legitimate.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-27 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Accounting Auditing |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 64.0/100 · rank range 35–37 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-11, EXP06-PARTNER-26, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Audit firms already require independence checks, but completing a form does not ensure that a team confronts the purpose of independence together. During engagement acceptance, renewal, or scope changes, fees, staffing, deadlines, and client relationships may dominate discussion. Team members may struggle to identify the absent financial-statement users whom independence protects, raise uncertain services or relationships early, or step away from a role without appearing disloyal. The prevalence of this collective-meaning problem has not been measured directly.
What is proposed¶
At defined engagement triggers, the firm would hold a structured boundary cycle alongside—not instead of—its official independence process. Cards would represent the paying client relationship and the absent users relying on the audit, while an empty seat would make those users visible in the discussion. After private reflection, participants would sort invented or engagement-level service and role cards into ordinary processing, outside the engagement, or independent review. Personal facts would go only through confidential channels. The engagement sponsor would name commercial pressures, and an independent ethics witness would state limits the sponsor cannot override. Participants could speak, write, remain silent, or skip the symbolic portion. Every unresolved matter would then enter the authorized acceptance, staffing, consultation, or service-approval system with an owner, evidence requirement, decision authority, and deadline.
The cross-domain transfer¶
The source archetype uses a marked, recurring enactment to renew shared meaning and commitments. Here, the enactment represents absent report users, names commercial incentives, and recognizes consultation or role substitution. Its structural mapping is reasonably direct, but the ceremony itself never determines whether a person, service, or engagement satisfies independence requirements.
Why it advanced¶
This candidate cleared the separately calibrated empirical-partner lane because the underlying independence objective is consequential, a synthetic comparison is feasible, and the intervention has explicit authority and privacy safeguards. It did not enter the strict-success lane: no firm has requested it, and no field evidence shows that audit teams need or benefit from the ritual layer.
Prior art and the remaining open claim¶
Independence confirmations, relationship monitoring, service preapproval, acceptance reviews, training, confidential consultation, staffing restrictions, and quality oversight already perform most operational functions. Ethical culture and group rituals provide adjacent support for the underlying rationale, not evidence for this design. The open claim is narrower: with official rules unchanged, does the marked group cycle improve recognition and correct routing of uncertain matters and perceived consultation safety beyond ordinary workflow, without adding coercion, privacy risk, false assurance, or authority confusion?
Smallest decisive test¶
Preregister a synthetic two-arm tabletop with 12–16 participants per arm. Both groups receive the same invented engagement facts and mock confirmation, consultation, and acceptance workflow; only the treatment group receives the boundary cycle. A blinded reviewer would score recognition of absent-user interests and commercial incentives, identification and routing of uncertain matters, and rejection of group sorting as official approval. Also measure pressure, privacy, accessibility, consultation willingness, authority confusion, false assurance, and elapsed time. Reject the incremental claim if recognition or routing does not improve, a simpler comparator performs as well, or safeguards worsen materially.
Deployment and cost¶
The authorized first step is a tabletop using invented facts, not live client or personal information. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for initial setup, \(250,000–\)1 million for operational launch, and \(250,000–\)1 million annually. These are assessment bands, not vendor quotes.
Risks and uncertainties¶
- Silence, passing, or leaving could be interpreted as evidence that a participant has a personal conflict.
- The empty seat could imply that all financial-statement users share one interest.
- Group sorting could be mistaken for an official independence decision or evidence of unanimity.
- A fee-owning partner could control the script while retaining practical power over staffing and revenue.
- Personal financial, family, employment, or client-protected information could enter an unsuitable group setting despite the confidential route policy. Long-term use could become moral theater while incentives and retaliation risks remain unchanged.
Expert review¶
Useful reviewer backgrounds: Audit-independence or professional-ethics specialist, Audit-firm quality-control leader, Employment and privacy counsel, Organizational-behavior researcher, Experimental-design and measurement specialist.
- Do existing engagement processes already produce timely recognition and routing of the uncertain matters this cycle targets?
- Which actions or participation records could create employment, privacy, privilege, or professional-standards exposure?
- Can opting out, remaining silent, or requesting substitution be made credibly nonretaliatory in a partner-led team?
- What prespecified difference would justify the added meeting time over an improved form plus confidential consultation?
- How should reviewers detect whether participants mistake symbolic sorting or witnessed closure for official approval?
Evidence and provenance¶
Selected sources: S1: Revision of the Commission's Auditor Independence Requirements; Final Rule 33-7919 · S2: QC 1000, A Firm's System of Quality Control · S3: Spotlight: Inspection Observations Related to Auditor Independence · S4: Fostering a Healthy “Tone at the Top” at Audit Firms · S5: Ethics, Independence and Objectivity — Transparency Report 2025 · S6: International Code of Ethics for Professional Accountants, Including International Independence Standards · S7: Does Ethical Culture in Audit Firms Support Auditor Objectivity? · S8: Work Group Rituals Enhance the Meaning of Work
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed it in band C, with ranks 35–37 across profiles. That score is only a reading-order aid, uses cost as a pilot-affordability proxy, and does not measure economic value or change its empirical-partner endpoint.
38. Honest Guarantees for Generative Art¶
Canonical title: Guarantee-Bounded Preflight for Interactive Generative Art
In one sentence: An executable-art platform would state exactly which works and interactions it can verify, preserve unknown results outside that boundary, and leave artistic acceptance to curators.
| Field | Record |
|---|---|
| Portfolio ID | EXP04-STRICT-01 |
| Experiment and endpoint | Experiment 4 · Strict success |
| Archetype × domain | Computability Boundary Mapping × Art Aesthetics |
| Proposal position or arm | PROPOSAL_FIRST |
| Post-hoc reading order | Balanced score 63.0/100 · rank range 36–46 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
A platform wants a preflight tool to decide whether every submitted generative artwork will always terminate or reset and will never violate formal display limits under any future interaction or sensor stream. Yet the artwork language, external capabilities, input bounds, and guarantee are not precisely defined. A timeout or unsuccessful search may then be treated as approval or rejection even though it proves neither. The central issue is whether the unrestricted request crosses a computability boundary while narrower program fragments remain analyzable.
What is proposed¶
Before extending the verifier, the platform would define the artwork language, state behavior, permitted plug-ins and network access, input encoding, display constraints, and meaning of termination or reset. It would seek both a total procedure for restricted fragments and an independently checked impossibility argument for the unrestricted class. Works inside an enforceable finite fragment could receive exact or sound finite-state analysis. Explicitly bounded interactions could be checked exhaustively only within those bounds. Other works would receive sound but incomplete analysis that can report a possible violation or UNKNOWN, followed by curator review and runtime containment. Every report would identify its language version, assumptions, bounds, property, and guarantee type. The verifier would assess mechanically testable constraints, not beauty, meaning, or artistic merit.
The cross-domain transfer¶
Computability-boundary mapping separates an open-ended decision demand from regions where exact procedures are possible. In this setting, executable artworks and future audience inputs form the open class; restricted grammars, finite interaction envelopes, sound abstractions, explicit unknowns, and versioned records mark the narrower regions where honest guarantees can be made.
Why it advanced¶
It passed Experiment 4’s strict researched-candidate bar because the boundary problem is clearly specified, close comparators exist, safeguards preserve curator and artist authority, and a falsifiable archived-work study is defined. STRICT_SUCCESS does not mean real-world validation, novelty, deployment authorization, or demonstrated economic impact.
Prior art and the remaining open claim¶
Museums already use source review, risk assessment, emulation, sandboxes, watchdog resets, artist consultation, and cross-functional conservation. Formal model checking and abstract interpretation also exist for restricted programs. The remaining claim is the museum-specific combination: adding an enforceable language-and-property contract plus guarantee-labeled routing will reduce unsupported claims based on timeouts or sampled success, preserve UNKNOWN as non-dispositive, and retain curator-rated fidelity for works admitted to the restricted fragment. No supplied source establishes that joint result.
Smallest decisive test¶
With an authorized museum or platform, freeze one interpreter and examine 12 archived works read-only. Define three machine-testable display constraints, a finite event alphabet, and trace depth 20. Compare the proposed restricted-fragment and routing workflow with sandbox/watchdog review, property-based testing, and manual source review. An independent formal-methods reviewer would check semantics, soundness, coverage, and any impossibility proof. Measure coverage, time, concrete violations, false alarms, unknowns, label retention, scope overstatement, and curator-rated fidelity. Stop or redesign for any false clearance, incomplete declared enumeration, coercion of UNKNOWN, or failure to preserve essential behavior in at least eight works.
Deployment and cost¶
No adopter, corpus, source rights, or interpreter specification is secured. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for initial setup and operational launch, and \(50,000–\)250,000 annually. They exclude institution-specific rights clearance, legacy recovery, procurement, and major exhibition hardware.
Risks and uncertainties¶
- Formal palette, luminance, or exclusion-zone rules could be mistaken for the curator’s richer aesthetic judgment.
- A restricted language could exclude recursion, emergence, duration, or responsiveness essential to the artwork.
- A coarse abstraction could overwhelm reviewers with possible violations, while an unsound one could issue false clearance.
- Finite analysis may still be impractical because the modeled state space grows too large.
- Staff could operationally turn UNKNOWN into rejection despite the stated policy. Archived works may poorly represent future submissions or dependencies.
Expert review¶
Useful reviewer backgrounds: Formal-methods and computability researcher, Software-art conservator, Curator of digital or time-based media, Generative artist or artist-rights representative, Museum technology and data-governance counsel.
- Does the frozen artwork language actually contain the computational features assumed by the boundary argument?
- Can each formal display property be checked without misrepresenting the curator’s intended constraint?
- Which behaviors must the restricted fragment preserve for the selected works to remain faithful?
- Are the abstraction and bounded enumeration sound and complete for exactly the scopes printed on their reports?
- Will downstream staff preserve UNKNOWN and POSSIBLE_VIOLATION labels instead of translating them into acceptance decisions?
Evidence and provenance¶
Selected sources: S1: Handling Digital Assets in Time-Based Media Art · S2: The Conserving Computer-Based Art Initiative · S3: Matters in Media Art · S4: Risk Assessment as a Tool in the Conservation of Software-Based Artworks · S5: Emulation or it Didn’t Happen · S6: Formal Verification for Node-Based Visual Scripts Using Symbolic Model Checking · S7: Abstract Interpretation: A Unified Lattice Model for Static Analysis of Programs by Construction or Approximation of Fixpoints · S8: Software Developers, Quality Assurance Analysts, and Testers
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed this candidate in band C, spanning ranks 36–46 across profiles. It is a reading-order aid using cost as an affordability proxy, not an experimental endpoint, economic-value measure, or qualification of its STRICT_SUCCESS status.
39. Turning Archival Loss into Stewardship¶
Canonical title: Loss-to-Stewardship Assembly for Community-Governed Religious Archives
In one sentence: A community-controlled assembly would acknowledge authorized archival losses, invite bounded offers of help, and convert accepted offers into funded preservation work without demanding access, disclosure, or public grief.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-31 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Religious Studies Theology |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 63.0/100 · rank range 39–42 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
When a community-held religious recording, manuscript, oral history, or access relationship is destroyed, dispersed, withdrawn, or made inaccessible, a consortium may record only a status change or announcement. Later scholars and staff may not know what was lost, what must remain undisclosed, or who may describe it. Related materials can remain exposed, communities can face repeated requests, and proposed responses can lack owners, permissions, resources, or deadlines. The frequency and practical effects of this pattern have not been measured.
What is proposed¶
Only an authorized custodian could place a verified loss before the consortium and decide what may be public, controlled, or unspoken. After an independent review of standing, privacy, cultural protocol, accessibility, emotional pressure, and exit options, an ordinary empty archival sleeve would mark the acknowledged absence without imitating lost or sacred material. A custodian-approved account would state the type of loss and the claims outsiders may not make. Participants could observe a bounded silence, leave, or use another accessible mode. Institutions could then offer specific capacity—staff time, storage assessment, migration, translation, catalog correction, training, or funds—or pass without explanation. Offers would not be accepted during the marked interval. Custodians could later accept, change, or decline them, after which accepted offers would become permission-bound, funded work orders with owners and review dates.
The cross-domain transfer¶
The ritual archetype is instantiated as a recurring, marked occasion that joins shared memory to renewed commitments. Here, an empty sleeve marks absence, witnesses preserve authorized limits, and a circulating marker prompts offers of capacity. The mapping is meaningful, but the ritual cannot preserve materials, settle custodianship, or transfer access rights by itself.
Why it advanced¶
It cleared the empirical-partner lane because documentary losses and community-governed access are consequential, the design separates acknowledgment from consent, and a fictional tabletop comparison is feasible. It did not meet strict success: no consortium or community has requested it, and there is no field evidence that the assembly improves stewardship beyond ordinary governance.
Prior art and the remaining open claim¶
Loss registries, cultural-access protocols, grants, preservation consortia, emergency workflows, working groups, and memorial events already cover the component functions. Community-specific authority cannot be replaced by consortium permission. The open claim concerns only the added assembly: compared with the same layered registry, budget, meeting, and work-order process, does it improve accurate recall, completed custodian-approved work, and avoidance of repeat requests without increasing disclosure, pressure, unequal attention, emotional burden, or administrative cost?
Smallest decisive test¶
Preregister a six-to-eight-week crossover tabletop with 18–30 practitioners, preservation staff, and compensated community-archive advisers, using one fictional low-sensitivity loss. Compare the marked assembly with a permission-layered registry and ordinary preservation meeting; hold facts, restrictions, and capacity budget constant and randomize condition order. Measure immediate and 30-day recall, disclosure and overclaim errors, standing judgments, conversion of offers into funded work orders, task clarity, repeat requests, pressure, accessibility, burden, hours, and cost. Advance only if the assembly improves at least two prespecified stewardship outcomes with no authority or disclosure error and no worse pressure or burden.
Deployment and cost¶
The first authorized step uses no real community or collection. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for initial setup, and \(250,000–\)1 million for both operational launch and annual operation. No staffing model, vendor quote, case volume, or host budget verifies these bands.
Risks and uncertainties¶
- The consortium could appropriate or aestheticize a community’s loss.
- Public acknowledgment could reveal a restricted collection’s existence, location, vulnerability, or former contents.
- Custodians could feel pressured to narrate grief or grant access in exchange for assistance.
- Capacity offers could become donor branding or promises without assigned budgets.
- Visible, dramatic losses could attract resources away from slow deterioration or less prominent communities. The layered memory record could create later surveillance or security risks even when current access controls are followed.
Expert review¶
Useful reviewer backgrounds: Community archive custodian or authorized knowledge holder, Religious-collections archivist, Cultural-protocol and Indigenous data-governance specialist, Preservation-program manager, Privacy, copyright, and records-retention counsel.
- Who has standing to authorize the loss account when custodians or community members disagree?
- What information may be public, controlled, temporarily retained, or omitted entirely?
- Can declining participation, disclosure, or assistance remain genuinely costless under donor and institutional power differences?
- Does the assembly improve completed preservation work beyond an equally funded registry-and-working-group process?
- How should resources be allocated so publicly legible losses do not displace less visible but higher-risk preservation needs?
Evidence and provenance¶
Selected sources: S1: Protecting, preserving and promoting access to the world’s documentary heritage · S2: Religious Archives Group Conference in Association with The National Archives: Conference Report · S3: Protocols for Native American Archival Materials · S4: Understanding Communities and Cultural Protocols · S5: Records at Risk Grants · S6: How APTrust Works · S7: Mourning ritual participation, subjective well-being and prosocial behaviour among the Luhya people of Kenya · S8: About ARCS
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed it in band C, with ranks 39–42 across profiles. This is only a reading-order aid using a cost-band affordability proxy; it does not measure economic value, establish elapsed pilot time, or alter the empirical-partner endpoint.
40. Reviewing Ritual Records by Their Differences¶
Canonical title: Residual-Led Review of Recurrent Ritual Records
In one sentence: Researchers would use a collection-specific prediction model to foreground differences in recurring ritual records while preserving complete sources, independent full-record audits, and automatic fallback whenever the model becomes unreliable.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-06 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Predictive Residual Processing × Religious Studies Theology |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 63.0/100 · rank range 33–48 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-07, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Researchers comparing repeated service transcripts, ceremony descriptions, or ritual-manual editions may spend much of their first review pass rereading recurring structures while locally important changes compete for attention. Full-record reading preserves context but may be slow; simple exception lists can impose one supposedly normal form and hide changes that baseline cannot represent. The actual attention bottleneck has not yet been observed in a selected corpus, and predictability must be assessed separately for each collection rather than across religious traditions.
What is proposed¶
A read-only layer would predict the next descriptive event code—such as a role, action, object, or sequence position—before a held-out record is opened. The model would be versioned, collection-specific, and limited to declared metadata; it could not judge doctrine, authenticity, orthodoxy, or religious value. The interface would show structured insertions, deletions, substitutions, reorderings, and coding disagreements, weighted by uncertainty, source quality, underrepresentation, sensitivity, and interpretive importance. It could collapse only high-confidence expected units during first-pass review. Every packet would reconstruct the complete coded sequence and link to the untouched source. Sensitive, low-confidence, untranslated, permission-ambiguous, harm-related, or poorly represented records would appear in full. Independent random and risk-based full-record audits would detect missed context, while version mismatches, drift, or error-budget breaches would restore full-record review.
The cross-domain transfer¶
Predictive residual processing sends an expected signal plus the differences needed to reconstruct the observation, concentrating attention on prediction errors. Here, the expected signal is a collection-specific ritual-event sequence and the residual is a structured record difference. Checksums, raw-source audits, and fallback preserve reconstruction and expose model failure.
Why it advanced¶
It passed Experiment 6’s strict researched-candidate bar because the transfer is structurally clear, existing corpora and collation tools make a retrospective study plausible, and the comparator, safeguards, and stopping rules are explicit. STRICT_SUCCESS remains a research endpoint only, not field validation, novelty, deployment approval, or evidence of economic value.
Prior art and the remaining open claim¶
TEI apparatuses, CollateX, religious-studies toolkits, synoptic editions, recurrence analysis, anomaly ranking, and manual sampling already expose textual variation and support structured comparison. The remaining claim is their untested combination with pre-observation prediction, reconstructive residuals, consequence-aware routing, independent raw audits, and automatic fallback. For one authorized corpus, that system must save total review time while remaining noninferior to full-record and nonpredictive collation for material-variant recall, reconstruction, context coverage, and protected handling.
Smallest decisive test¶
With a corpus owner and required community steward, preregister a read-only study of 60 authorized, fully coded records. Train and freeze a transparent model on the earliest 30, then counterbalance reviewers across full-record review, nonpredictive collation, and the predictive-residual interface on 30 held-out records. A blinded panel would define material variants and audit every mandatory-bypass record plus random and risk-based collapsed records. Stop if the residual condition misses a bypass item, falls more than five points behind full review on recall or reconstruction, creates a greater than five-point miss disparity for underrepresented contexts, saves less than 20% total person-time, or fails to outperform nonpredictive collation.
Deployment and cost¶
No corpus access, permission determination, adopter commitment, or validated event-code scheme is secured. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for initial setup, \(250,000–\)1 million for operational launch, and \(50,000–\)250,000 annually. Measured engineering, coding, governance, and audit costs remain unavailable.
Risks and uncertainties¶
- A collection-specific expectation could acquire false authority as the normal or canonical ritual form.
- Routine wording, silence, performance, or material context omitted by event codes could carry the important meaning.
- Residual queues could reward novelty and make continuity seem unimportant.
- Underrepresented traditions or periods could be suppressed if consequence safeguards fail.
- Residuals could make sensitive outliers easier to identify. Model expectations could contaminate later coding, while gradual historical change could be absorbed as normal and disappear from attention.
Expert review¶
Useful reviewer backgrounds: Digital humanities or textual-collation specialist, Religious-studies corpus researcher, Community data steward, Machine-learning evaluation specialist, Research-ethics and sensitive-data reviewer.
- Is full-record rereading actually the binding attention cost in the proposed corpus?
- Which descriptive code units preserve performance, silence, translation ambiguity, material practice, and consequential routine wording?
- What material-variant gold standard can a blinded panel apply consistently?
- Do audit and fallback costs leave at least the prespecified 20% net person-time saving?
- Are miss rates, source expansions, and reconstruction errors acceptably balanced across periods, collections, and underrepresented contexts?
Evidence and provenance¶
Selected sources: S1: Arthur Westwell: Digital Techniques for Presenting Liturgical Texts and Building a Database of Carolingian Pontificals · S2: Corpus of Hittite Festive Rituals: Description · S3: WP3 – T-ReS – Toolkit for Religious Studies · S4: TEI Guidelines, Chapter 13: Critical Apparatus · S5: CollateX Documentation · S6: Recurrence Analysis Function, a Dynamic Heatmap for the Visualization of Verse Text and Beyond · S7: Quantifying Text Reuse Across Three Kṛṣṇa Yajurveda Recensions: Using Multi-Algorithm Computational Collation · S8: The CARE Principles for Indigenous Data Governance
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed it in band C, spanning ranks 33–48 across profiles. That wide range is a reading-order aid using cost as an affordability proxy, not an experimental endpoint, an economic-value estimate, or a change to STRICT_SUCCESS.
41. Renewing Ramp Stop-Work Support¶
Canonical title: Shared-Envelope Renewal for Ramp Stop-Work Legitimacy
In one sentence: A monthly, voluntary exercise would let workers from different ramp employers rehearse mutual recognition of approved stop-work calls and expose unsupported response paths without changing operating authority.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-12 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Aviation Aeronautics |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 63.0/100 · rank range 34–46 across three profiles · band C |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-13, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Aircraft turnarounds bring together airline staff and separate fueling, baggage, catering, cleaning, maintenance, towing, and ground-handling employers. Each may teach its own safety rules, yet workers can still disagree about who may pause shared work, how other companies must respond, and whether a junior contractor will be protected for interrupting a senior or time-critical operation. This uncertainty can delay or silence warnings even when every employer has a written stop-work policy.
What is proposed¶
Pilot a voluntary twenty-minute monthly observance at a classroom aircraft outline or closed training stand, never during a live turnaround. A rotating steward explains that the exercise creates no qualification, authority, or procedure. Participants from different employers trace or otherwise follow the aircraft boundary, hear frontline accounts of four cross-company dependencies, and exchange cards listing only approved support such as acknowledgment, interpretation, or escalation routes. Missing or contradictory support becomes a recorded safety-management action rather than an improvised promise. People may speak, write, observe silently, or pass without penalty. A debrief checks pressure, exclusion, ambiguity, and follow-through; independent review can require repair, suspension, or retirement.
The cross-domain transfer¶
The source archetype uses a recurring, marked ritual to make a shared commitment visible and memorable. Here, the ritual becomes a cross-employer aircraft-boundary exercise, reciprocal support-card exchange, plural response, and steward handoff. The mapping is structurally strong, but its symbolic elements have no assumed safety effect; their incremental value must be compared with an ordinary joint briefing.
Why it advanced¶
This candidate passed Experiment 6’s strict researched-candidate bar because it defines a bounded comparator, reversible first test, authority limits, concrete falsifiers, and safeguards against coercion and procedural confusion. STRICT_SUCCESS is only that experimental endpoint: it does not establish real-world effectiveness, novelty, deployment approval, or economic value.
Prior art and the remaining open claim¶
The surrounding parts already exist: joint ramp briefings, written escalation procedures, safety stand-downs, public safety pledges, reporting systems, and operational roles with stop-work authority. The open claim is narrower. Where approved routes already exist, does adding this opt-out monthly enactment improve unaided reconstruction of reciprocal cross-company responses and assignment of unsupported interfaces to authorized owners, compared with an equal-duration conventional briefing, without added pressure, retaliation concern, access loss, or authority confusion?
Smallest decisive test¶
At one willing station, place 8–12 workers from at least three organizations into matched twenty-minute sessions using identical fictional scenarios and route information: the proposed renewal or a conventional joint briefing. Blind-score who may call an approved pause, each organization’s response and escalation path, conflict handling, and what the session did not authorize. Also record whether planted gaps receive an owner, forum, and date, plus anonymous pressure and access measures. Proceed, revise, or stop; do not move to live operations without separate authorization.
Deployment and cost¶
The first-evidence estimate is under $10,000. Initial deployment, operational launch, and annual recurring support are each roughly \(10,000–\)50,000 in 2026 resource-equivalent terms, not vendor quotes. Local policy mapping, translation, shift coverage, accessibility, facilitation, auditing, and correction of discovered communication or staffing gaps could materially change the total.
Risks and uncertainties¶
- Managers could point to visible agreement while leaving anti-retaliation protection weak.
- Contractors or junior workers may feel compelled to affirm support in front of supervisors.
- Participants may mistake ceremonial statements or cards for approved operating procedures.
- Mobility, sensory, language, remote-access, and shift constraints may make participation unequal.
- A dominant airline or handler could control the script, stewardship rotation, or interpretation of mutual support promises.
Expert review¶
Useful reviewer backgrounds: Airport or airline safety-management-system specialist, Ground-handling operations leader, Ramp worker, union, or contractor representative, Aviation human-factors and safety-culture researcher, Accessibility and language-access specialist.
- Do local policies and contracts already define reciprocal stop-work acknowledgment across every participating employer, and where do they conflict?
- Can workers decline speaking, moving, or card exchange without conspicuous refusal or employment consequences?
- Does the matched briefing control contain exactly the same approved route information and facilitator time?
- What blinded scoring rule will distinguish genuine route reconstruction from recall of ceremonial wording?
- Which authority can accept each discovered gap, fund its correction, and verify closure?
Evidence and provenance¶
Selected sources: S1: Ground Operations Safety · S2: Ground Operations — Safety — SIRM 30 · S3: Easy Access Rules for Ground Handling (Regulations (EU) 2025/23 and 2025/20) · S4: Safety Management Systems (SMS) for Airports and Airport Projects · S5: Aircraft Arrival, Turnaround and Departure on Stand, ASGrOps_OSI_093 · S6: Introducing the Aviation Safety Pledge · S7: National Safety Stand-Down to Prevent Falls in Construction · S8: Air Transportation: NAICS 481
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized score is only a post-hoc reading aid: band C, with profile ranks from 34 to 46. Its affordability input proxies pilot speed and does not measure elapsed time, economic value, or the STRICT_SUCCESS endpoint.
42. Expiring Old Training Evidence Safely¶
Canonical title: Cognitive Evidence Lease Manager for Longitudinal Adaptive Training
In one sentence: A shadow system would retire contextually stale training evidence from current predictions while preserving the records and model versions needed to reconstruct earlier recommendations.
| Field | Record |
|---|---|
| Portfolio ID | EXP05-STRICT-02 |
| Experiment and endpoint | Experiment 5 · Strict success |
| Archetype × domain | Layer Decay And Expiration Management × Cognitive Science |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 62.0/100 · rank range 39–44 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
An adaptive cognitive-training system can keep using performance evidence from older sessions after a person’s ability, task design, scoring model, or context has changed. Those records may then distort current difficulty choices. Simply deleting them creates a different problem: researchers may lose the history needed to explain a trajectory, reproduce an earlier recommendation, investigate model behavior, or meet retention duties. The system needs to separate evidence’s authority in current inference from its physical preservation.
What is proposed¶
Assign each trial-derived evidence layer an inference lease when it is created, based on evidence type, task and scoring versions, context, and a review horizon. Expiry removes the layer from ordinary current-state estimation but sends it to a provenance-preserving archive rather than deleting it. New corroboration can renew a lease; task, scoring, or context changes can trigger early review. A multi-factor score may order review work but cannot decide disposition. Before destruction, stewards must check retention authority and downstream dependencies, use reversible quarantine, and leave a tombstone. Sampled restore drills test whether earlier estimates and recommendations remain reproducible. Initial evaluation occurs only through offline shadow replay, leaving source data, the live model, and participant experience unchanged.
The cross-domain transfer¶
The archetype manages accumulated layers whose usefulness decays or expires. In this domain, the layers are timestamped trials, summaries, context annotations, and model bindings. Expiration changes whether evidence influences today’s estimate, while storage tiers, dependency checks, quarantine, tombstones, and restore tests preserve accountable history. The structural mapping is direct, although temporal models may already address the predictive problem.
Why it advanced¶
This candidate passed Experiment 5’s strict researched-candidate bar through a specific comparator set, measurable offline test, reversible implementation, separated decision authority, and explicit failure rules. STRICT_SUCCESS does not show that stale evidence is prevalent, that leases beat strong temporal models, or that production use is warranted, novel, or economical.
Prior art and the remaining open claim¶
Adjacent systems already apply sliding windows, uniform forgetting, state-space or change-point models, time-aware knowledge tracing, immutable event logs, replay, and storage lifecycle policies. The unresolved comparison is whether explicit context-sensitive leases offer a smaller active evidence set that matches or improves held-out current-session prediction against cumulative and fixed-recency estimators—and remains competitive with a strong temporal model—while preserving historical reconstruction and keeping false-staleness, recommendation instability, and steward workload acceptable.
Smallest decisive test¶
With written data-steward approval, replay a completed multi-session working-memory dataset containing at least two documented task or scoring contexts. Freeze outputs from the existing cumulative estimator, then compare cumulative history, a fixed recent window, and context-sensitive leases; add a strong temporal-model sensitivity analysis if feasible. Measure held-out prediction, recommendation quality against a task-owner rubric, calibration, transitions, active-set size, false staleness, dependency recall, review minutes, and exact historical reconstruction. Reject progression after any missed seeded dependency, failed reconstruction, unacceptable instability or burden, unauthorized exposure, or baseline-reproduction failure.
Deployment and cost¶
First evidence is estimated at \(50,000–\)250,000. Initial deployment and operational launch are each roughly \(250,000–\)1 million; annual recurring work is \(50,000–\)250,000, in 2026 resource-equivalent bands rather than quotes. Actual cost depends on data volume, metadata quality, architecture, legal regimes, dependency tracing, expert review, and integration debt.
Risks and uncertainties¶
- Short or poorly chosen leases could overemphasize recent observations and discard stable information from active inference.
- Context rules may encode sensitive attributes or operate unevenly across participants.
- Adaptive task selection can make recent evidence look more informative because the system chose what was observed.
- A composite review score may hide contestable governance judgments behind numerical precision.
- Incomplete dependency mapping could permit removal of evidence needed to explain an earlier recommendation.
Expert review¶
Useful reviewer backgrounds: Cognitive scientist specializing in working-memory measurement, Adaptive-learning or knowledge-tracing modeler, Data steward or research-records officer, Privacy and data-retention counsel, Event-sourcing and archival-reconstruction engineer.
- What observed task, scoring, or context changes make an older trial semantically inapplicable rather than merely old?
- Which primary metric and non-inferiority margins will compare leases with cumulative, fixed-window, and strong temporal models?
- How will experts label false staleness without using the tested lease rules as their own ground truth?
- Can every sampled historical recommendation be recreated with the archived evidence, code, configuration, and model version?
- Which participant permissions, research requirements, holds, or deletion duties govern archive, quarantine, and destruction?
Evidence and provenance¶
Selected sources: s1: Does Working Memory Training Have to Be Adaptive? · s2: Adapting Training in Real Time: An Empirical Test of Adaptive Difficulty Schedules · s3: Rethinking and Improving Student Learning and Forgetting Processes for Attention Based Knowledge Tracing Models · s4: Time-Dependant Bayesian Knowledge Tracing—Robots That Model User Skills Over Time · s5: Event Sourcing Pattern · s6: AI Risk Management Framework Core · s7: Regulation (EU) 2016/679, General Data Protection Regulation · s8: Amazon S3 Pricing
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering places this candidate in band C, with profile ranks from 39 to 44. That ordering is not an experimental endpoint or economic-value estimate; its pilot-speed input is only a cost-band affordability proxy.
43. A Capped Prize for Catalyst Endurance¶
Canonical title: Capped Endurance Prize for a Durable Nanocatalyst
In one sentence: A sponsor would compare coded nanocatalysts through resource-capped, independently audited endurance tests rather than select a demonstration candidate mainly from publications or peak reported performance.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-07 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Nanotechnology |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 62.0/100 · rank range 31–50 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-03, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A sponsor with one follow-on demonstration award may rank nanocatalyst teams using publications, presentations, and peak results obtained under different conditions. That can favor short favorable runs, pure feedstocks, high scarce-metal use, selected batches, or heavy synthesis and computing expenditure. The chosen catalyst may then prove fragile during longer or variable operation. Yet the alleged selection process, incomplete reporting, resource escalation, and connection to later demonstration failure have not been documented for a named sponsor.
What is proposed¶
Replace publication-priority selection with a preregistered endurance prize whose rules are frozen before entrants are known. Teams register preparations, failed attempts, contributors, counted spending and in-kind inputs, reactor and accelerator-compute hours, scarce materials, and waste. Coded samples go to an independent laboratory for public qualification and undisclosed but representative endurance, feed-variation, restart, and regeneration sequences. Safety and containment are non-compensable gates. Passing entries are scored for reproducibility, sustained conversion and selectivity, deactivation and recovery, material intensity, energy, waste, and reporting completeness. Resource caps limit escalation. Two technically distinct milestone awards preserve alternatives before one demonstration award. Audits, appeals, conflict screening, graduated penalties, cleanup guarantees, and a challenger path govern the contest.
The cross-domain transfer¶
Bounded-rivalry governance redirects competition by fixing the arena, limiting inputs, policing interference, preserving alternatives, and reviewing winner lock-in. Here those elements become a capped catalyst prize with coded testing, milestone awards, safety floors, resource and waste accounting, sanctions, and continuation conditions. The mapping is strong, but the prize’s claimed predictive advantage over standardized endurance testing alone remains unmeasured.
Why it advanced¶
This did not enter Experiment 6’s strict-success lane. It cleared the separately calibrated EMPIRICAL_PARTNER_CANDIDATE lane because a bounded retrospective partner study is testable and safety-limited. Progress depends on a named sponsor and proprietary records; neither the field problem nor the full contest design’s incremental benefit has yet been demonstrated.
Prior art and the remaining open claim¶
Catalyst benchmarking, degradation protocols, independent laboratory validation, staged federal prizes, advance scoring rules, safety controls, and appeals already exist separately. The remaining claim concerns their combination: would adding auditable input caps, complete-attempt reporting, lifecycle burdens, independent milestones, hidden representative segments, and a continuation challenge predict later demonstration performance better than both historical peak-result selection and standardized endurance testing alone? Located evidence supports the components, not that comparative or behavioral result.
Smallest decisive test¶
With a named sponsor, preregister a non-awarding shadow study of 5–15 archived projects. Compare the historical decision, an endurance-only ranking, and the full capped and lifecycle-weighted ranking using coded records and held-back time-series segments. Measure rank uncertainty, missingness, reconstruction labor, incumbent effects, audit reversals, appeal burden, and association with later demonstration outcomes. Stop if fewer than 80% of histories are reconstructable, measurement uncertainty exceeds project differences, reasonable weights reverse the rankings, the full design fails to beat endurance alone, or governance cost is disproportionate. No new synthesis, testing, funding change, or award follows.
Deployment and cost¶
The retrospective evidence study is estimated at \(50,000–\)250,000. Initial setup is \(250,000–\)1 million; operational launch and annual operation are each roughly \(1–\)5 million in 2026 resource-equivalent terms. These are not quotes. Reaction-specific protocols, laboratory capacity, prize administration, purse size, environmental review, confidentiality, and demonstration scope remain unpriced.
Risks and uncertainties¶
- Hidden test sequences may reward resistance to surprise conditions rather than representative durability.
- Interlaboratory or segment variability may be larger than genuine differences among catalysts.
- Resource caps could favor teams that already own equipment, datasets, or precursor inventories.
- Broad accounting may expose confidential operations; narrow accounting may shift spending off book.
- Institutional cleanup guarantees could exclude capable teams without wealthy sponsors even when their methods are safe scores can obscure tradeoffs among activity, selectivity, endurance, scarce materials, energy, and waste. A sponsor, catalyst metrologist, independent laboratory, safety specialist, competition designer, and research-accounting expert must determine whether an auditable comparison is feasible.
Expert review¶
Useful reviewer backgrounds: Catalysis scientist with durability-testing expertise, Nanomaterial measurement and interlaboratory-validation specialist, Independent validation-laboratory operator, Chemical safety, exposure, and waste specialist, Federal prize, procurement, or research-competition counsel.
- Does a named sponsor actually make a scarce follow-on decision using publication priority or peak team-reported results?
- Can archived records distinguish every attempted preparation and run from the selected results presented to the sponsor?
- What target reaction, benchmark, operating envelope, and interlaboratory error define a fair endurance comparison?
- Which resource categories can be audited without unfairly favoring incumbents or exposing protected information?
- Does the full ranking predict later demonstration outcomes better than endurance-only scoring across preregistered weights?
Evidence and provenance¶
Selected sources: S1: Towards Benchmarking in Catalysis Science: Best Practices, Challenges, and Opportunities · S2: The rotating disc electrode: measurement protocols and reproducibility in the evaluation of catalysts for the oxygen evolution reaction · S3: Standardized protocols for evaluating platinum group metal-free oxygen reduction reaction electrocatalysts in polymer electrolyte fuel cells · S4: Nanotechnology Measurement Protocols · S5: Electrolysis Catalyst Synthesis, Ex-situ Electrochemical Performance and Durability Characterization, and Standardization · S6: Prize and Challenge Toolkit: A Guide for Federal Innovation Managers · S7: Accelerating Innovation through American-Made Challenges · S8: Guidance and Publications: Nanotechnology
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading aid assigns band C and profile ranks from 31 to 50. It does not upgrade the EMPIRICAL_PARTNER_CANDIDATE endpoint, measure economic value, or establish elapsed pilot speed; affordability only proxies that input.
44. Send Foresight Surprises, Keep Full Records¶
Canonical title: Scenario-Residual Exchange for Distributed Horizon Scanning
In one sentence: Distributed scanners would send structured differences from a shared scenario expectation while retaining full evidence, protected-signal bypasses, independent audits, and automatic return to complete-packet review when the filter becomes unreliable.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-16 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Futurism Foresight |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 62.0/100 · rank range 41–45 across three profiles · band C |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A distributed foresight network may send complete weekly assessments for every monitored driver, forcing central analysts to reread expected material to find a few assumption-breaking observations. A simple report-by-exception rule is unsafe because silence could mean stability, missing reporting, or model blindness. The proposed setting still lacks packet-level evidence that routine material actually exhausts review capacity, that important signals arrive late for this reason, or that downstream decision blockage is not the real bottleneck.
What is proposed¶
Before each weekly intake, scenario stewards publish versioned expectations for each driver’s direction, pace, geography, actors, cross-impacts, uncertainty, scope, and expiry. Local analysts retain full evidence capsules but encode structured differences such as acceleration, reversal, a new actor, altered coupling, or broken assumption. A gate prioritizes these residual cards by reliability, uncertainty, consequence, perspective coverage, and review cost. Heartbeats distinguish expected conditions from missing reports, while protected safety-, rights-, conflict-, and distributional-harm signals travel in full. Central analysts reconstruct the assessment from the matching expectation and residual; validated surprises receive an owner but do not automatically change scenarios. Independent raw sampling, periodic reconciliation, drift monitoring, and version, error, or missingness triggers restore full-packet review.
The cross-domain transfer¶
Predictive-residual processing sends deviations from a synchronized expected state instead of retransmitting the entire state. Here, the expected state is a versioned scenario-conditioned driver assessment, and the residual describes model-relative change. Reconstruction, heartbeats, raw audits, protected bypasses, and fallback make the mapping substantive, though translating ambiguous foresight narratives into reliable residual fields remains an unresolved weakness.
Why it advanced¶
This did not qualify as an Experiment 6 strict success. It entered the EMPIRICAL_PARTNER_CANDIDATE lane because a shadow comparison on an adopter’s historical corpus is bounded and measurable. The essential field evidence—actual backlog prevalence, missed-signal causes, taxonomy reliability, protected-signal recall, and fully loaded labor savings—is still missing.
Prior art and the remaining open claim¶
Scenario-guided scanning, indicator-based early detection, distributed scanning networks, short-list filtering, collaborative platforms, ordinary tags, anomaly alerts, and periodic scenario refreshes are adjacent prior art. The narrower open claim is that a synchronized expectation-plus-residual workflow can reduce total review effort versus complete packets, tagged triage, and existing signpost filtering without materially reducing detection of assumption changes, cross-driver effects, novel actors, or protected-source evidence once preparation, audits, reconciliation, and fallback are counted.
Smallest decisive test¶
Obtain 200–500 authorized historical packets with timestamps. Using only earlier material, freeze expectations for each evaluation window and randomly assign blinded reviewers to complete packets, tagged complete packets, or residual cards with logged fallback. Preregister review-time savings, reconstruction agreement, recall margins, and equal protected-source recall; include raw audits and injected version mismatches and missing heartbeats. Reject the workflow if any safety-class item is suppressed, reconstruction breaches tolerance, underrepresented sources fare worse, fallback fails, systematic residual structure persists, or expectation preparation, auditing, reconciliation, and fallback erase the labor saving.
Deployment and cost¶
First evidence is estimated at \(10,000–\)50,000. Initial deployment is \(50,000–\)250,000, operational launch \(250,000–\)1 million, and annual recurring work \(50,000–\)250,000 in rough 2026 resource-equivalent bands, not quotes. Taxonomy design, expectation maintenance, adjudication, audit independence, access controls, integration, training, and fallback workload may dominate software costs.
Risks and uncertainties¶
- Shared expectations could become a self-confirming filter that suppresses evidence outside current scenarios.
- Precision or consequence weights may discount unfamiliar regions, disciplines, sources, or minority perspectives.
- A heartbeat may be recorded even when the underlying source network has quietly failed.
- Residual cards may remove narrative context needed to interpret ambiguous long-range evidence.
- Analysts may change classifications to attract central attention or avoid scrutiny.
Expert review¶
Useful reviewer backgrounds: Government or corporate foresight program leader, Scenario-planning and horizon-scanning methodologist, Information-retrieval or human-in-the-loop systems researcher, Rights, conflict, and distributional-impact specialist, Records, privacy, and access-control officer.
- Do complete packets currently exceed a declared review budget, and which material findings were delayed specifically by intake saturation?
- Can independent annotators reliably encode and reconstruct the proposed driver-state and residual fields?
- Which evidence classes must bypass filtering regardless of predicted relevance or apparent routine status?
- Does the residual workflow preserve protected and underrepresented-source recall under blinded adjudication?
- After counting preparation, maintenance, audit, reconciliation, and fallback, is total analyst time at least meaningfully lower than complete review?
Evidence and provenance¶
Selected sources: S1: Building capacity in technology horizon scanning: A guide for policymakers · S2: Building Anticipatory Capacity with Strategic Foresight in Government: Lessons from Lithuania, Italy, and Malta · S3: The Futures Toolkit · S4: Enhancing horizon scanning by utilizing pre-developed scenarios: Analysis of current practice and specification of a process improvement to aid the identification of important weak signals · S5: Anticipating surprise: The case of the early warning system of Rijkswaterstaat in the Netherlands · S6: Integrating scenario planning and indicator-based early detection for scenario transfer · S7: FIBRES Pricing · S8: Employer Costs for Employee Compensation — December 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized ordering is a post-hoc reading aid: band C, with profile ranks from 41 to 45. It neither changes the EMPIRICAL_PARTNER_CANDIDATE status nor measures economic value; the pilot-speed input is an affordability proxy, not elapsed time.
45. Residual-First Review of Forensic Timelines¶
Canonical title: Residual-First Triage for Digital Forensic Timelines
In one sentence: A read-only review layer would foreground unexplained timeline changes while preserving the complete forensic record, reconstructible context, independent audits, and automatic return to full review.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-04 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Predictive Residual Processing × Criminology Forensic |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 62.0/100 · rank range 32–51 across three profiles · band C |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-05, EXP06-PARTNER-14, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Digital forensic examiners may confront millions of timestamped system, application, synchronization, and acquisition events. Repeated, predictable background activity can consume attention before an examiner reaches missing, extra, displaced, or altered events that matter to a case. Simply hiding familiar-looking events is unsafe: an apparently routine record may still reveal guilt, innocence, provenance, deletion, clock problems, or evidence-integrity failures, and context-free anomalies can themselves invite overinterpretation.
What is proposed¶
For one declared operating-system, application, and version scope, freeze a reference model that predicts ordinary background event bundles. Compare the complete extraction with those predictions and describe each discrepancy as missing, extra, reordered, time-shifted, or attribute-changed. Show prioritized discrepancies with model identity, provenance, uncertainty, and enough neighboring events to reconstruct the local sequence. Never alter the forensic image or complete extraction. User-authored material, deletion indicators, integrity and acquisition failures, clock discontinuities, required disclosures, and examiner-requested records always appear in full. An independent reviewer examines random and risk-selected raw windows. Unsupported versions, model or parser mismatches, audit disagreement, excessive reconstruction error, and queue overload switch the affected scope back to complete chronological review. Human reviewers, not the model, determine evidentiary meaning.
The cross-domain transfer¶
The predictive-residual archetype maps strongly here: a versioned model predicts routine event sequences, while the analyst receives typed differences between prediction and observation. Reconstruction, version checks, independent raw sampling, protected-event bypasses, drift monitoring, and full-review fallback instantiate the archetype without replacing the underlying evidence.
Why it advanced¶
This candidate passed Experiment 6's strict researched-candidate bar because the burden is documented, a bounded prototype appears feasible, the comparison is falsifiable, and the safeguards directly address context loss and automation risk. That status is not real-world validation, a novelty finding, deployment approval, or evidence of economic impact.
Prior art and the remaining open claim¶
Complete timeline extraction, static filters, analyzers, pattern reconstruction, prioritization, provenance drill-down, clustering, and visualization already exist. The remaining claim is narrower: compared with full chronological review and the strongest static-filter or analyzer workflow, a frozen residual-first layer combining typed discrepancies, reconstructible context, protected bypasses, independent raw-window audits, synchronized versions, and automatic fallback would reduce review effort without increasing material or exculpatory misses, interpretation errors, or total audit and maintenance burden.
Smallest decisive test¶
Pre-register a randomized crossover study with about 8–12 qualified examiners, 4–6 synthetic or reusable closed-case images, and 24–36 matched sections from one software-version scope. Compare complete review, the strongest Plaso or Timesketch static workflow, and residual-first review. Measure examiner minutes, time to inspect material events, inculpatory and exculpatory recall, protected-context recall, false escalation, reconstruction disagreement, disparities, fallback, and total workload. Stop progression for any protected-bypass failure, unreviewed material suppression, unacceptable recall loss, recurrent mismatch, systematic disparity, negligible effort reduction, or audit burden comparable to full review. Passing would not authorize live-case use.
Deployment and cost¶
The first retrospective evidence study is estimated at \(50,000–\)250,000. Initial deployment startup is \(250,000–\)1 million; an operational launch is \(1–\)5 million, with \(250,000–\)1 million recurring annually. These are rough 2026 USD resource-equivalent bands, not vendor quotes, and version maintenance may erase expected savings.
Risks and uncertainties¶
- A mistaken reference model could suppress material inculpatory, exculpatory, provenance, or integrity information.
- Examiners could mistake statistical unusualness for criminal intent, authorship, or evidentiary significance.
- Parser changes, clock ambiguity, missing data, or mismatched model versions could create false discrepancies or false silence.
- Reference data and priority weights could work unevenly across applications, languages, devices, or user contexts.
- Sparse audit samples might miss rare blind spots, while incomplete version records could obstruct disclosure and independent challenge.
Expert review¶
Useful reviewer backgrounds: Digital forensic examiner, Forensic laboratory quality and method-validation specialist, Forensic timeline-tool developer, Defense-side digital-forensics expert, Evidence and disclosure lawyer.
- Can sampled timeline windows be reconstructed to the semantic fidelity required for independent examination and disclosure?
- Which event classes require unconditional full-context presentation under laboratory policy and applicable law?
- Does residual-first review preserve inculpatory and exculpatory recall against both complete review and the strongest static workflow?
- How often would operating-system, application, parser, locale, clock, or acquisition changes force model revalidation?
- Does audit, synchronization, and maintenance work remain below the attention saved during review?
Evidence and provenance¶
Selected sources: S1: Needs Assessment of Forensic Laboratories and Medical Examiner/Coroner Offices: A Report to Congress · S2: Digital Forensics · S3: Method validation in digital forensics (accessible), FSR-G-218 · S4: Using log2timeline.py — Plaso documentation · S5: Create an analyzer — Timesketch · S6: An Automated Timeline Reconstruction Approach for Digital Forensic Investigations · S7: Forensic Science Technicians: Occupational Outlook Handbook · S8: SoK: Timeline-based event reconstruction for digital forensics: Terminology, methodology, and current challenges
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized score is only a post-hoc reading-order aid: this candidate ranked 32–51 across profiles and falls in band C. The ranking is not an experimental endpoint or a measure of novelty, deployment readiness, or economic value.