Candidate dossiers — Band B¶
Part of Inverse Innovation with the Encyclopedia of Abstractions · Ranks 11–25: high post-hoc review priority · Last revised August 2026
11. A Credible End to Financial Close¶
Canonical title: Close-Mode Release and Recovery Observance
In one sentence: An optional post-close observance would mark the end of exceptional work, preserve every residual task, and turn gratitude into authorized, reviewable recovery commitments.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-11 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Accounting Auditing |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 70.0/100 · rank range 8–26 across three profiles · band B |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-26, EXP06-PARTNER-27, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A financial close may be officially complete while staff still face vague cleanup, monitoring, and audit-support duties. Completion emails and celebrations can praise visible overtime without clarifying who owns remaining work or what recovery will occur. Fatigue and ambiguous responsibility may then carry into ordinary operations and the next close. Symbolic gratitude can also conceal unresolved staffing, pay, leave, or process needs, while employees who are remote, junior, quiet, or unwilling to celebrate remain unseen.
What is proposed¶
After an authorized reporting milestone, hold a governed and optional release observance that has no power to close books, waive controls, or cancel work. A steward states that boundary, then switches off a nonoperational close-mode marker. Participants may use a bounded reflection period, remain off-camera, join asynchronously, or decline. With prior consent, witnesses recognize specific work such as error prevention, documentation, coordination, boundary setting, and asking for help, without ranking hours or glorifying exhaustion. A shared board separates completed work from residual obligations, which enter the official tracker with owners, limits, evidence needs, and escalation paths. Only authorized recovery, backfill, compensation-review, meeting-relief, staffing, or process commitments are published. An immediate debrief tests coercion and credibility; a later independent review compares promises with schedules, recovery, and task completion.
The cross-domain transfer¶
The ritualized-commitment archetype becomes a visible transition out of exceptional close mode. Its marker, silence, witnessing, commitment renewal, and later accountability review enact shared meaning and reciprocal obligations. The structural mapping is plausible, but the ceremony cannot itself reduce workload, provide recovery, satisfy wage rules, or correct accounting processes.
Why it advanced¶
The candidate passed Experiment 6’s strict researched-candidate bar because the underlying fatigue and close-management problems are consequential, a low-cost synthetic comparison is feasible, and explicit consent and authority safeguards make the claim testable. STRICT_SUCCESS does not establish workplace benefit, adopter demand, novelty, or permission for live use.
Prior art and the remaining open claim¶
Workplace rituals, project-completion ceremonies, close-management software, task trackers, retrospectives, overtime and leave policies, fatigue programs, and recognition systems already exist. The narrower open claim is that, with these controls held constant, an optional governed release sequence can improve retained understanding of official completion versus residual work and the credibility of authorized recovery commitments without increasing coercion, privacy exposure, status confusion, overwork glorification, or unpaid extra-role effort.
Smallest decisive test¶
Run a preregistered synthetic comparison with 8–12 volunteer accounting and workforce participants. Counterbalance two equivalent close scenarios: ordinary status communication, task tracking, recovery policy, and retrospective; and those same controls plus the observance. Seed residual tasks, unauthorized recovery proposals, an accessibility need, and a status-confusion cue. Measure immediate and 72-hour understanding of official status, ownership, escalation, and authorized commitments, alongside anonymous safety ratings. Revise or reject the observance after any lost task, increased authority error, noncredible opt-out, pressured positivity, unsafe disclosure, or failure to improve retained boundary clarity.
Deployment and cost¶
The first synthetic study is estimated below $10,000 in rough 2026 resource-equivalent terms. Startup and operational launch are each estimated at \(10,000–\)50,000, with \(50,000–\)250,000 annually. These bands exclude potentially dominant costs such as paid recovery, overtime, backfill, staffing, compensation changes, and automation, and are not vendor quotes.
Risks and uncertainties¶
- Staff may mistake switching off the marker for official completion of books, controls, or audit obligations.
- The expectation of celebration or silence may become coercive despite a formal opt-out.
- Recognition may reward visible long hours while overlooking remote, junior, contingent, or upstream contributors.
- Managers may promise recovery, leave, pay, staffing, or backfill without authority or resources.
- Private health, family, grievance, or workload information could be exposed during reflection or recognition and then mishandled by others involved in the session or organization afterward.
Expert review¶
Useful reviewer backgrounds: Corporate controller, Human-resources or labor-relations specialist, Accounting close-process owner, Occupational fatigue and workplace-safety specialist, Accessibility, privacy, and organizational-behavior specialist.
- Can every participant decline the symbolic portion without practical or perceived retaliation?
- Which statements and visual cues reliably distinguish the observance from official close status?
- Who has authority and funding to approve each proposed recovery or workload commitment?
- Does recognition capture preventive and boundary-setting work without rewarding excessive hours?
- What evidence before the next close would show that promised recovery and residual-task ownership actually occurred?
Evidence and provenance¶
Selected sources: S1: The Agentic Close: From Month-End Sprint to Always-Ready Finance · S2: How organizations can streamline the month-end close · S3: 2025 Corporate Finance & Accounting Talent Study · S4: Work group rituals enhance the meaning of work · S5: Financial period close workspace · S6: Fatigue and Work · S7: Wages and the Fair Labor Standards Act · S8: National employment and wage data by occupation, May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed this candidate between ranks 8 and 26 across profiles, in band B. Its relatively inexpensive test improved deployment-heavy ordering, but the score is only a reading aid—not an endpoint, elapsed-time measure, or estimate of value.
12. Blind Testing for Shared-Service Cost Models¶
Canonical title: Holdout Cost-Driver Tournament for Shared-Service Allocations
In one sentence: A controller would compare competing shared-service allocation models on withheld data under fixed rules, without posting the results or using them in live decisions.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-03 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Accounting Auditing |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 11–21 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-01, EXP06-PARTNER-02, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Business units often favor shared-service cost models that reduce their own allocated share, because one enterprise rule shifts costs among all units. Sponsors may choose favorable historical periods, redefine usage, exclude transactions, add opaque complexity, or influence judges. Yet competing proposals can reveal better cost drivers and weak data. The challenge is to distinguish legitimate measurement improvement from burden shifting when no organization-specific evidence yet shows that sponsorship incentives actually distort the selection process.
What is proposed¶
For one reconciled shared-service pool, run an identity-blinded tournament whose sole prize is a one-year designation as the managerial allocation basis. Freeze eligibility, data sources, transformations, holdout periods, scoring weights, conflicts, appeals, and prohibited conduct before outputs are inspected. Each model must disclose beneficiaries, use governed data, reproduce its logic, and reconcile the full pool. Blinded tests score usage linkage, out-of-period stability, auditability, data burden, and sensitivity to discretionary assumptions; sponsor savings are disclosed to validators but not scored. Independent reviewers reperform leading models. Development effort and complexity are capped, suspicious cross-model exclusions prompt investigation, and common data remains available to later challengers. A central holdback supports correction of trial-year defects. The designation expires, and post-contest review may revise or retire the process.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a controlled competition among allocation models. The scarce prize, fixed arena, eligibility rules, foul schedule, resource caps, independent judging, challenger access, and winner review map clearly. The empirical premise is weak, however: no partner data yet show systematic self-favoring sponsorship, selection influence, or manipulation beyond ordinary technical disagreement.
Why it advanced¶
This candidate did not enter the strict-success lane. It qualified only as an EMPIRICAL_PARTNER_CANDIDATE because a bounded, nonposting partner study could test the premise safely, while essential field evidence remains missing. Its status is an invitation to investigate actual sponsor behavior and model performance, not evidence that the proposed problem or remedy exists in practice.
Prior art and the remaining open claim¶
Managerial-costing standards, standard allocation drivers, commercial allocation software, model-risk governance, holdout testing, independent validation, and controller judgment are adjacent prior art. The open comparison is narrower and conditional: when direct tracing is infeasible and sponsors have distributive exposure, a blinded, preregistered tournament may select a more stable, reproducible, usage-linked model than controller selection or a non-blinded panel while weakening the relationship between sponsor savings and rank.
Smallest decisive test¶
With a controller and data owners, preregister a closed-year shadow study comparing the incumbent basis, a conventional non-blinded panel choice, and the blinded tournament using identical candidate models. Freeze the holdout quarter, weights, and stress tests first. Measure reconciliation, reproducibility, stability, usage linkage, assumption sensitivity, burden, transfers to nonparticipants, direct-tracing feasibility, and correlation between sponsor savings and rank. Do not post allocations or use them for budgets, pay, tax, transfer pricing, or reporting. Stop if data access, common-pool reconciliation, validator independence, or target validity fails; require material comparative improvement before considering another study.
Deployment and cost¶
The first partner study is estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms. Startup is estimated at \(50,000–\)250,000, operational launch at \(250,000–\)1 million, and annual operation at \(50,000–\)250,000. Actual costs, authority, data availability, and recurring validation burden have not been observed and these bands are not vendor quotes.
Risks and uncertainties¶
- Sponsor identity may remain obvious from a model’s structure, defeating blinding.
- Withheld historical periods may not represent future service consumption or behavior.
- A composite score may conceal disputed judgments about causal linkage, stability, and administrative burden.
- Complexity caps may reject a justified model for genuinely heterogeneous services.
- Historical driver data may already reflect earlier allocation choices, missing usage, or incumbent control and therefore bias comparisons in ways difficult to detect without detailed organization-specific investigation.
Expert review¶
Useful reviewer backgrounds: Corporate controller or managerial-accounting policy owner, Shared-service cost-modeling specialist, Independent model validator or internal auditor, Operational data owner and data-governance specialist, Business-unit finance representative without judging authority.
- Do sponsored models favor their sponsors after legitimate service differences are controlled?
- Can sponsor identity be hidden well enough for blinded judging to be meaningful?
- What operational measure can serve as a defensible usage-linkage target?
- Is direct metering or decomposition feasible at an acceptable authorized burden?
- Are model rankings stable across preregistered holdouts, stress scenarios, and reasonable scoring weights?
Evidence and provenance¶
Selected sources: S1: Why do so many organisations struggle with cost allocation? · S2: Who should pay for support functions? · S3: Distorted Cost Allocation: An Encouragement or Discouragement? · S4: The Conceptual Framework for Managerial Costing · S5: TBM Modeling Allocation Methods · S6: Supervisory Guidance on Model Risk Management · S7: The TBM Taxonomy · S8: HHS Policy for Information Technology Portfolio Management (PfM)
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed this candidate between ranks 11 and 21 across profiles, in band B. This reading-order score uses an affordability proxy and is not an experimental endpoint, economic-value estimate, or upgrade from its EMPIRICAL_PARTNER_CANDIDATE status.
13. Separate Release Effects from River Disturbances¶
Canonical title: Command-Conditioned Residual Watch for Managed River Releases
In one sentence: Test whether a model of verified reservoir operations can filter predictable downstream changes from operators’ attention without hiding protected signals, coincident environmental events, or model failures.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-15 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Environmental Climate |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 13–19 across three profiles · band B |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
Reservoir releases can predictably alter downstream water level, temperature, dissolved oxygen, conductivity, and turbidity. These large, expected changes can flood ordinary dashboards and fixed alarms, leaving operators to decide whether each warning reflects the release or a separate problem such as contamination, bank failure, or unexpected sediment movement. Widening thresholds during releases reduces nuisance alarms but can also create a blind spot precisely when an external event may coincide with operations.
What is proposed¶
Before downstream readings arrive, copy each authorized gate or turbine command and compare it with measured actuator behavior. A frozen, versioned model would predict when and how the release should affect each monitoring station. The system would preserve every raw reading, calculate signed observed-minus-predicted residuals, and direct reliable or consequential unexplained changes to an operator queue. A residual would always retain enough command, model, station, timing, and sensor information to reconstruct the full observation. Critical ecological limits, dam-safety variables, compliance records, sensor faults, and public-warning conditions would bypass filtering. Unverified actuation, excessive uncertainty, model drift, missing heartbeats, incompatible versions, or failed reconstruction would restore full-signal review. Residuals could prompt investigation or separately reviewed recalibration, but could not alter gates or emergency actions.
The cross-domain transfer¶
The predictive-residual archetype is mapped strongly here: a copy of the reservoir’s own command predicts its delayed sensory consequences, and the difference between prediction and observation highlights what the command does not explain. The mapping adds actuator verification, uncertainty weighting, reconstruction, raw-data audits, and a fallback because an environmental predictor must never become grounds for erasing consequential observations.
Why it advanced¶
This entered the empirical-partner lane because release modeling, gate-operation modeling, continuous monitoring, quality control, and alarm-management practices make a bounded replay plausible. It did not clear the strict-success lane: no site has supplied synchronized records, and there is no measured evidence yet for event recall, protected-signal routing, reconstruction, fallback reliability, workload reduction, or lower total burden.
Prior art and the remaining open claim¶
Adjacent systems already model reservoir operations and downstream water quality, detect sensor anomalies, and manage alarms. The narrower open comparison is whether conditioning a frozen model on both the issued command and measured actuation can outperform an unchanged raw-threshold dashboard: at least 30% less routine review, 100% routing of protected signals, at least 95% detection and acknowledgement of scripted coincident departures, reconstruction within declared tolerances, and correct fallback for every injected fault. No exact implementation or comparative result was established.
Smallest decisive test¶
With one reservoir partner, preregister a historical release replay with a contiguous untouched holdout. Freeze the model, stations, uncertainty assumptions, protected classes, tolerances, and fallback rules. Blindly add conductivity, turbidity, and dissolved-oxygen departures; inject actuator mismatch, dropout, checksum conflict, travel-time shift, and persistent bias; and compare with the unchanged dashboard. Reject the claim if review falls by under 30%, any protected signal is suppressed, scripted-event detection is below 95%, any injected fault misses fallback, reconstruction exceeds tolerance, or total labor is not lower. Only a complete pass permits a read-only shadow release.
Deployment and cost¶
The first step is retrospective and then shadow-only; existing dashboards, staffing, operating procedures, and control authority remain unchanged. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for initial startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. These are assessment bands, not vendor quotes or site estimates.
Risks and uncertainties¶
- A contaminant or sediment pulse arriving during a release could be predicted away as an operational effect.
- A logged command may not match physical gate or turbine movement, corrupting the downstream prediction.
- Wrong travel time could turn one event into misleading positive and negative residuals at different times.
- Season, tributary inflow, channel change, stratification, or shared upstream-data errors could invalidate the model.
- Expected release effects may still breach ecological or compliance limits and therefore cannot disappear from review context or records, even if they are accurately predicted.
Expert review¶
Useful reviewer backgrounds: Reservoir operations engineer, River water-quality scientist, Hydrologic or hydraulic modeler, Environmental incident-response lead, Continuous-sensor quality specialist subject to relevant independence safeguards.
- Are command, measured-actuator, upstream-condition, tributary, weather, sensor, alarm, and acknowledgement records synchronized well enough for the preregistered replay?
- Which stage, dissolved-oxygen, conductivity, and turbidity conditions must always bypass residual filtering?
- What variable-specific reconstruction errors and travel-time errors are acceptable before full-signal fallback?
- Can blinded coincident events be designed so they represent consequential external disturbances without being trivially detectable?
- Does total operator, modeling, audit, and investigation labor remain below the raw-dashboard baseline at equal completeness?
Evidence and provenance¶
Selected sources: S1: Water Quality · S2: The Effects of Hydropower Releases from Lake Texoma on Downstream Water Quality · S3: HEC-ResSim Version 4.0, New Water Quality Feature · S4: HOW TO Add Real-Time Gate Operation Capability to a ResSim Model · S5: Guidelines and Standard Procedures for Continuous Water-Quality Monitors: Station Operation, Record Computation, and Data Reporting · S6: ContDataQC: An R Package and Shiny App for Quality Control of Continuous Water Quality Sensor Data · S7: Alarm Management for Hydropower Plants · S8: Guidelines for Collecting Data to Support Riverine Hydrodynamic and Water Quality Simulation Models
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized score placed this candidate between ranks 13 and 19, in band B. That score is only a post-hoc reading order using a cost-band affordability proxy; it is not an experimental endpoint, elapsed-time estimate, deployment finding, or measure of economic value.
14. A Storage-Independent Aircraft Maintenance Ledger¶
Canonical title: Opaque Maintenance-Obligation Ledger for Aircraft Configuration and Usage Evidence
In one sentence: Define and test maintenance-obligation behavior independently of database layout so backend changes cannot silently alter configuration, usage, credit, conflict, or due-status answers.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-18 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Aviation Aeronautics |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 7–28 across three profiles · band B |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
Aircraft maintenance clients may read shared tables and calculated columns directly, each making its own assumptions about record order, nulls, unit conversions, corrections, installations, counter resets, and maintenance credit. A database migration, cache redesign, or replay into a new backend can therefore change whether an obligation appears satisfied, remaining, overdue, or indeterminate even when the underlying evidence is logically unchanged. Different clients may also produce conflicting answers with no common behavioral standard for deciding which is correct.
What is proposed¶
Create an opaque aircraft maintenance-obligation ledger whose public operations establish a baseline, submit typed evidence, supersede rather than erase an erroneous event, produce an as-of snapshot, query an obligation, compare revisions, and explain the evidence behind a status. Its abstract state would include configuration, serialized components, usage, requirements, credited accomplishments, conflicts, supersession links, and immutable history. Preconditions would define identifiers, units, timing, authority, applicability, and correction links. Invariants would prohibit deleted history, double installation, unsupported credit, and definite answers from contradictory evidence. Relational, event-log, graph, or cached implementations would remain hidden and would have to pass the same generated operation-sequence tests. The ledger would report evidence and conflict states; authorized personnel would still decide maintenance credit, deferrals, aircraft status, and operational release.
The cross-domain transfer¶
The representation-independent contract archetype maps directly to the ledger. It replaces database rows as the effective interface with an abstract state, explicit transitions, stable errors, evidence explanations, and a common black-box oracle. Opacity alone is insufficient: each backend must map to the same ledger meaning, and client code must stop reconstructing maintenance status from hidden storage details.
Why it advanced¶
This qualified for a bounded partner study because regulated record integrity, digital interoperability, mature maintenance systems, and established migration audits make isolated replay feasible. It did not enter strict success: no dependency audit, controlled representation perturbation, independent model, mutant suite, leakage audit, or comparative result exists, and no operator has approved the complete semantics for even one obligation type.
Prior art and the remaining open claim¶
Maintenance standards already cover information exchange, electronic logbooks, allowable configuration, transfer records, and due-status business rules; commercial systems already manage configurations, utilization, compliance, and component status. Append-only logs, typed APIs, and migration audits also address parts of the problem. The narrower open claim is that an opaque evidence-state contract plus an independent reference model and shared sequence oracle will catch semantic defects and preserve identical observable behavior across two differently represented backends more reliably than schema checks, record counts, and sampled reports.
Smallest decisive test¶
With an operator or MRO and an authorized records reviewer, preregister one obligation type and at least 100 curated histories plus generated sequences. Compare an incumbent wrapper and independent immutable model with schema, count, and sampled-report checks. Seed defects for duplicated utilization, erased superseded evidence, insertion-order dependence, double installation, and silent conflict resolution. Require identical revisions, statuses, errors, and provenance explanations between nominal implementations, and rejection of every seeded defect. Reject the claim if the baseline catches the same defects, any mutant passes, valid implementations need schema exposure to agree, contradictory evidence must become definite, or representation-only perturbations reveal no actual client dependency.
Deployment and cost¶
Begin with de-identified or synthetic histories in a read-only isolated environment; do not write official records or connect results to dispatch or release workflows. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. They are not quotations or organization-specific estimates.
Risks and uncertainties¶
- The abstract state may omit maintenance-program, jurisdiction, configuration, or evidence-authority context required to interpret an obligation.
- The independent model may repeat the incumbent’s mistaken rule or be too simple to adjudicate complex histories.
- Generated cases may miss irregular but valid corrections, counter changes, or applicability sequences.
- Events labeled independent may actually have order-dependent maintenance meaning.
- Opaque storage could hinder investigation if evidence explanations and sanctioned audit access are inadequate, while deterministic outputs could be mistaken for release authority.
Expert review¶
Useful reviewer backgrounds: Authorized aircraft maintenance-records specialist, Maintenance-planning or continuing-airworthiness engineer, MRO systems architect, Aviation software assurance and test specialist, Regulatory compliance or airworthiness counsel.
- Which single obligation type has semantics complete enough to specify units, applicability, cutoffs, corrections, conflicts, and authority?
- Which clients currently read tables, calculated columns, row order, nulls, or local credit rules?
- What contract-level evidence explanation is necessary for an authorized reviewer to resolve every seeded divergence?
- Which event pairs truly commute, and which only appear independent until configuration applicability is considered?
- What regulatory acceptance, retention, privacy, cybersecurity, and controlled-data conditions would govern any move beyond shadow replay?
Evidence and provenance¶
Selected sources: S1: 14 CFR §121.380 Maintenance Recording Requirements and §121.380a Transfer of Maintenance Records · S2: AC 120-78B: Electronic Signatures, Electronic Recordkeeping, and Electronic Manuals · S3: Easy Access Rules for Continuing Airworthiness: ML.A.305 Aircraft Continuing-Airworthiness Record System · S4: Adopting Aircraft Electronic Records, First Edition · S5: ATA e-Business Program Standards · S6: Regulatory Compliance Module · S7: Case Study: Endeavor Air—Managing Legacy MRO Systems · S8: Copa Airlines Selects GE Aviation for Digital Records Management
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed this candidate from rank 7 to 28, in band B. This wide, post-hoc reading range is not an experimental endpoint or evidence of value; its pilot-speed input was only a cost-band affordability proxy, with elapsed time unscored.
15. Portable Rules for Drought-Stage Advice¶
Canonical title: Representation-Independent Contract for Drought-Stage Recommendations
In one sentence: Specify drought-stage recommendation behavior independently of any spreadsheet or dashboard, then test whether a second engine can reproduce it without relying on hidden formulas, fields, colors, or evaluation order.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-20 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Environmental Climate |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 14–20 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-19, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A drought plan’s stage calculation may live in one spreadsheet, dashboard, indicator service, or rules engine. Staff and connected procedures can start depending on cell locations, colors, proprietary scores, rounding, formula order, or undocumented treatment of missing data. Replacing or refactoring that implementation can then change an escalation, recovery, or indeterminate result even though the approved policy and hydrologic evidence were supposed to remain the same. The reason for the change may be impossible to reconstruct cleanly.
What is proposed¶
Define an opaque drought-stage recommender whose public state includes the current recommended stage, assessment time, evidence sufficiency, pending transition, policy and parameter versions, and stable reason categories. Operations would initialize a version, assess or advance with typed evidence, invalidate withdrawn inputs, query the recommendation, and export an immutable receipt. Inputs must meet declared rules for units, time windows, location, eligibility, freshness, and coverage. The same abstract history must produce the same result; missing required evidence cannot silently become a normal stage; hysteresis and persistence must follow the approved policy; and rejected calls cannot change state. Formulas, database structures, vendor fields, caches, colors, and evaluation order remain hidden. A shared black-box suite would test boundaries, histories, corrections, missingness, receipts, errors, and leakage. Official declarations, restrictions, allocations, and public messages remain with the designated authority.
The cross-domain transfer¶
The representation-independent contract archetype is a strong structural match. The proposal separates the policy-facing behavior—stages, sufficiency, transitions, reasons, errors, and receipts—from any spreadsheet or rules engine. A replacement is acceptable only under the same approved policy version and after common conformance tests; changing a threshold or hysteresis rule is a policy change, not an implementation substitution.
Why it advanced¶
This reached the empirical-partner lane because real programs use multi-indicator stages, formal trigger governance exists, and environmental-data standards demonstrate typed exchange and black-box conformance. It did not reach strict success: no real workflow dependency has been documented, no independent engine has been compared, and no authority has approved the proposed semantics for missingness, corrections, invalidation, or hysteresis.
Prior art and the remaining open claim¶
Drought programs already define indicators, stages, triggers, governance, and history-sensitive assessment; WaterML standardizes observation exchange, and environmental-data APIs have executable conformance tests. These are adjacent rather than empty territory. The remaining claim is narrower: for one frozen policy version, a stateful contract covering insufficiency, time order, hysteresis, corrections, invalidation, stable reasons and errors, and receipts will let an independently built engine substitute without any contracted-output divergence or client dependence on internal representation.
Smallest decisive test¶
In a six-to-eight-week non-live partner study, freeze one approved policy and completed assessment packet. Preregister evidence quantities, units, freshness, coverage, stages, transitions, missingness, corrections, reasons, receipts, boundary outcomes, and tolerances. Compare the incumbent through an adapter, an independent decision-table engine, and current manual or schema-only validation across threshold, unit, field-order, stale-data, repeated-snapshot, escalation, recovery, correction, and invalidation cases. Reject the problem premise if no hidden dependency appears and substitution requires no client change or output correction. Reject the intervention if any mutant passes, two passing engines disagree, or expressing the policy requires exposing the incumbent representation.
Deployment and cost¶
The first study is sandboxed, non-authoritative, and disconnected from declarations, restrictions, allocations, controls, and public communications. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(50,000–\)250,000 annually. These bands are not vendor quotes and lack external comparable-cost evidence.
Risks and uncertainties¶
- The contract could accidentally turn an undocumented spreadsheet quirk into approved drought policy.
- A finite suite may miss a threshold boundary or assessment history that changes a consequential recommendation.
- Equivalent software behavior would not show that the indicators, thresholds, or policy are scientifically, legally, or equitably appropriate.
- Stable reason categories may omit nuance required by hydrologists or decision-makers.
- Generated cases may miss correlated missing data or unusual corrections, and shadow outputs may still encourage unauthorized automation.
Expert review¶
Useful reviewer backgrounds: Drought-program policy owner, Operational hydrologist, Water-utility drought planner, Environmental software and conformance-test engineer, Legal, equity, and public-communications reviewer.
- Which document and authority determine the approved policy when spreadsheet behavior and written rules disagree?
- What exact evidence, freshness, coverage, hysteresis, and missing-data rules must the contract express?
- Do any current clients depend on cells, colors, vendor scores, rounding, reason text, or evaluation order?
- Which deliberately mutated engines would represent the most decision-relevant implementation errors?
- What outputs and labels are necessary to prevent a shadow recommendation from being treated as an official declaration?
Evidence and provenance¶
Selected sources: S1: Monitoring Drought · S2: Drought Response and Recovery: A Basic Guide for Water Utilities · S3: Building a Drought Planning Platform · S4: Drought Indicators & Assessment · S5: Water company drought plan guideline, 2025 · S6: What is the USDM? · S7: WaterML · S8: OGC API - Environmental Data Retrieval 1.0 Conformance Test Suite
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed this candidate between ranks 14 and 20, in band B. This is a post-hoc reading order, not an experimental endpoint, economic-value score, or pilot-duration estimate; affordability was proxied from the rough cost band.
16. Portable Rules for Delphi Rounds¶
Canonical title: Representation-Independent Contract for Delphi Elicitation Rounds
In one sentence: Test whether survey, spreadsheet, and analysis implementations can preserve the same Delphi-round lifecycle and anonymity boundary under one storage-independent behavioral contract.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-23 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Futurism Foresight |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 13–24 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-24, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A multiround public-health Delphi exercise may spread enrollment, anonymous responses, closure, aggregation, feedback, revision, and withdrawal across a survey platform, spreadsheets, email, and analysis scripts. Each tool can treat eligibility, missing answers, late submissions, replacements, withdrawal, and participant identity differently. When data moves or a tool changes, matching questionnaires and columns does not show that the same responses entered the aggregate, the same feedback population was used, or the same anonymity promises survived.
What is proposed¶
Define a Blind Iterative Elicitation State with operations to enroll an eligible participant through a separate identity service, issue an unlinkable study handle, open a round, accept or replace a response before closure, record withdrawal, close the round, calculate a declared aggregate, publish controlled feedback, begin the next round, and export an audit view. The contract would allow one active response per eligible handle, question, and round; forbid post-closure mutation; derive feedback only from the declared closed population; apply a sponsor-approved withdrawal policy; and block ordinary analysis from resolving identities. Preconditions, errors, side effects, versioning, and sanctioned identity recovery would be explicit while vendor fields, spreadsheet rows, email routing, and storage remain hidden. Every implementation would have to pass the same sequence and leakage tests before holding authoritative round state.
The cross-domain transfer¶
The representation-independent contract archetype maps well to the procedural state of Delphi rounds. It defines observable enrollment, response, withdrawal, closure, aggregate, feedback, and identity-access behavior without choosing a survey or storage format. The transfer is limited, however: technical equivalence cannot preserve participant experience, facilitator judgment, or the social context that may influence elicitation quality.
Why it advanced¶
This entered the empirical-partner lane because Delphi is established in government and health research, mature tools implement most lifecycle functions, and a synthetic two-implementation test is readily bounded. It was not a strict success: no migration incident evidence, committed sponsor, executable contract, adapter comparison, mutant result, or leakage result exists, and the governing legal and withdrawal rules remain unspecified.
Prior art and the remaining open claim¶
Existing Delphi products already support anonymous participation, multiple rounds, response revision, feedback, stopping rules, and study identifiers. Reporting guidance makes methodological choices more explicit, while vendor-neutral study-data standards support portable records. The narrower unresolved claim is that two independent implementations of one sponsor-approved synthetic protocol will agree on every contracted lifecycle and identity-access observation, reject every single-rule semantic mutant, and reveal no predeclared identity-linkage channel through allowed outputs. This is compositional proximity, not a novelty finding.
Smallest decisive test¶
Within six weeks, have a method owner preregister a synthetic two-round protocol covering eligibility changes, duplicates, replacement, missing and late answers, pre- and post-closure withdrawal, failed closure, two aggregation rules, feedback, and identity-resolution attempts. Build independent in-memory and spreadsheet-backed implementations, run identical scripted and generated histories, and add at least eight single-rule mutants. Reject the claim if an approved workflow cannot be expressed, implementations pass while disagreeing on any contracted observation, any mutant survives, or handles, timestamps, ordering, filenames, errors, or exports expose a predeclared linkage channel. Compare with reconstruction through the ordinary spreadsheet/export procedure; use no real experts or active records.
Deployment and cost¶
Start entirely offline with synthetic participants, non-live questions, isolated identity mappings, and no authoritative study migration. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(50,000–\)250,000 for operational launch, and \(10,000–\)50,000 annually. These labor-equivalent bands are not vendor quotes and exclude unverified licensing, security, procurement, and integration costs.
Risks and uncertainties¶
- The contract may encode one contested Delphi method as if it were a universal technical invariant.
- Generated histories may omit the unusual event orderings most likely to expose disagreement.
- Study handles could remain linkable through timing, ordering, filenames, metadata, or error messages.
- A simple reference implementation could acquire unwarranted authority over sponsor-approved methodological choices.
- Passing lifecycle tests would not show that the expert panel is representative, its judgments are accurate, or the participant experience is equivalent across tools.
Expert review¶
Useful reviewer backgrounds: Delphi methodologist, Public-health workforce foresight program owner, Research data-protection or privacy officer, Survey-platform and research-software engineer, Study sponsor or research-governance reviewer.
- Which withdrawal, replacement, eligibility, aggregation, feedback, and identity-recovery rules has the sponsor actually approved?
- Which ordinary outputs could link pseudonymous handles to people through timing, ordering, filenames, metadata, or errors?
- Do the generated sequences cover every meaningful pre- and post-closure failure path?
- Can two implementations represent the approved workflow without importing vendor-specific identifiers or spreadsheet behavior?
- Does the ordinary spreadsheet/export comparator detect the same seeded mutants, and at what review effort?
Evidence and provenance¶
Selected sources: s1: The Futures Toolkit HTML · s2: Guidance on Conducting and REporting DElphi Studies (CREDES) in palliative care: Recommendations based on a methodological systematic review · s3: ACCORD (ACcurate COnsensus Reporting Document): A reporting guideline for consensus methods in biomedicine developed via a modified Delphi · s4: eDelphi 2026 · s5: DelphiManager Frequently Asked Questions · s6: ODM v2.0 · s7: Regulation (EU) 2016/679 (General Data Protection Regulation) · s8: Software Developers, Quality Assurance Analysts, and Testers
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed this candidate between ranks 13 and 24, in band B. This post-hoc ordering is not a preregistered endpoint, validation, or economic-value measure; its pilot-speed input is only an affordability proxy derived from the rough cost band.
17. Stable Sampling Across Changing Data Systems¶
Canonical title: Opaque Probability-Sampling Frame Contract for Multi-Wave Social Research
In one sentence: This external-partner study candidate would test whether an opaque sampling contract can preserve eligible units, seeded selections, and inclusion probabilities when the same approved household frame is stored differently; no field study or adopter commitment yet supports it.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-25 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Sociology Anthropology |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 15–21 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A household study may store its recruitment frame as spreadsheets, rosters, or databases that use different ordering, identifiers, and blank-value conventions. Sampling scripts can quietly depend on those details. Reordering or migrating an otherwise unchanged frame may then alter eligibility, give duplicate records extra selection chances, change a seeded draw, or assign different inclusion probabilities. Researchers could mistake those technical effects for genuine field outcomes or an authorized change in sampling design.
What is proposed¶
Define the sampling frame through public behavior rather than a particular table. Custodians would map source records to unique abstract sampling units, keep direct identifiers behind a protected boundary, and freeze an immutable snapshot with explicit eligible, ineligible, unresolved, unknown, and withdrawn states. Authorized clients could draw a sample, query probabilities, verify a draw, supersede a snapshot, or export opaque recruitment handles. The contract would require identical results for the same snapshot, design version, and seed despite row permutation, lossless reserialization, or replacement of internal keys. Duplicate physical records could not create extra chances, invalid requests could not return partial samples, and drawing could never initiate contact. Shared black-box, generated-case, transformation, and leakage tests would gate backend replacement.
The cross-domain transfer¶
The transferred archetype is a representation-independent interface contract: multiple storage systems must expose the same abstract units and observable sampling behavior while hiding their internal representation. The structural mapping is strong, although the difficult mapping from real records to unique people, households, or addresses remains a substantive governance decision rather than a software detail.
Why it advanced¶
It cleared the separately calibrated external-partner lane because the operations, comparator, failure conditions, and synthetic test are unusually specific, and mature sampling and metadata tools make a prototype feasible. It did not enter strict success: representation-caused frame divergence has not been quantified, and no study or field office has offered data or committed to evaluation.
Prior art and the remaining open claim¶
Sampling standards already require accurate, unduplicated frames, retained design information, testing, documentation, and protection of restricted data. Metadata standards and sampling products already represent frames, selection probabilities, seeds, and reproducible draws. The narrower open claim is that independently built backends passing one opaque behavioral oracle will preserve abstract snapshots, selections, probabilities, errors, and non-contact side effects across representation-only changes. This is adjacent prior art, not a novelty finding.
Smallest decisive test¶
Run a two-week trial with 200 fictional units encoded independently in a shuffled flat file and normalized database. Compare the current file-coupled scripts with both adapters used only through the contract across at least 1,000 prespecified seeds. Apply row permutations, reserialization, key replacement, ineligible-record insertion, and duplicate consolidation; also insert deliberately defective adapters. Advance only if every legitimate transformation preserves all specified outputs, every mutant is caught, no source information leaks, and reviewers can map every valid state without hiding an identity judgment. Any changed draw or probability defeats the claim.
Deployment and cost¶
The first synthetic evidence step is assessed at roughly \(10,000-\)50,000. Startup, operational launch, and annual recurring work are each roughly \(50,000-\)250,000 in 2026 resource-equivalent terms, not vendor quotes. Live deployment would additionally require approved identity rules, privacy and ethics review, security controls, version governance, training, and a certified rollback path.
Risks and uncertainties¶
- The abstract unit model could erase meaningful distinctions among households, addresses, or memberships.
- Incorrect deduplication could merge distinct eligible units; insufficient deduplication could create extra selection chances.
- Predictable seeds, handles, output order, or error details could expose identities or reveal selection.
- A technically conforming frame could still omit important parts of the intended population.
- Immutable snapshots could preserve known mistakes if supersession is too difficult to use promptly.
Expert review¶
Useful reviewer backgrounds: Survey sampling statistician, Social-research data engineer, Frame custodian or field-operations lead, Research ethics and privacy specialist, Community or participant-governance representative.
- Can every valid source state map to a unique abstract unit without concealing a disputed identity or eligibility judgment?
- Which draw outputs must be bit-for-bit identical across backends, and which implementation differences may legitimately vary?
- Do current migrations or reorderings measurably change units, selections, or weights after authorized design changes are excluded?
- Could opaque handles, deterministic seeds, ordering, errors, or timing permit reidentification or prediction?
- Who may approve eligibility, deduplication, design-version, and live-backend changes?
Evidence and provenance¶
Selected sources: S1: Statistical Quality Standard A3: Developing and Implementing a Sample Design · S2: Statistics Canada Quality Guidelines: Coverage and frames · S3: Methodology: 2022–23 Survey of Asian Americans · S4: Survey Development — DDI Lifecycle 3.3 Technical Guide · S5: sampling: Survey Sampling, version 2.11 · S6: PROC SURVEYSELECT Statement, SAS/STAT 13.1 User's Guide · S7: Pseudonymisation
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed it between ranks 15 and 21, in band B. That score only sets reading order, uses affordability as a pilot-speed proxy, and measures neither experimental success nor economic value; its endpoint remains EMPIRICAL_PARTNER_CANDIDATE.
18. Renewing Nanomaterial Stewardship Commitments¶
Canonical title: The Visible Boundary: Nanomaterial Stewardship Renewal
In one sentence: This external-partner study candidate adds a brief, consent-governed stewardship observance to ordinary nanomaterial custody controls, but no participant study yet shows that it improves ownership accuracy beyond a strong checklist and briefing.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-30 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Nanotechnology |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 1–43 across three profiles · band B |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
The problem¶
Nanomaterials and related duties pass among synthesis, fabrication, measurement, safety, storage, and waste teams. Records may exist while people disagree about who owns an unresolved containment, labeling, transport, or disposal task, when it is due, or when to escalate it. Routine signatures can conceal that gap, especially after turnover. The result may be stranded anomalies, unfunded controls, ambiguous custody, poor traceability, or avoidable exposure and contamination risks.
What is proposed¶
At project launch, each custody transfer, and quarterly thereafter, hold a 12-minute “Visible Boundary” observance outside controlled work areas. An inert token moves along a printed lifecycle map from synthesis through disposal; no real material enters the session. Relevant participants may speak, write privately, observe, or pass without explanation. They identify the current stage and unresolved uncertainties, then propose, revise, or decline symbolic commitments. An accepted commitment counts only after the normal action system records an owner, resources, deadline, and escalation condition. A safety witness reads those entries back and states that the event does not certify safety or replace procedures, training, stop-work rights, custody records, or engineering controls. Anonymous feedback and independent reviews can revise, pause, or retire the practice.
The cross-domain transfer¶
The archetype transfers recurring ritualized meaning and commitment into a technical custody setting. The marked time, lifecycle symbol, witnessing, renewal, debrief, and retirement paths map clearly. The proposed effect, however, depends on a social mechanism—shared enactment improving accurate recall and ownership—that has not been demonstrated in nanomaterial work.
Why it advanced¶
It cleared the external-partner lane because a low-cost tabletop can compare the symbolic layer directly with a strengthened checklist and briefing, using delayed individual measurements and explicit safety failures. It did not meet strict success: no facility has committed, the proposed ownership problem lacks prevalence evidence, and the central comparative effect remains untested.
Prior art and the remaining open claim¶
Lifecycle controls, training, labeling, waste procedures, worker participation, action tracking, project rituals, and safety-culture programs already exist. Research suggests group rituals can increase perceived meaning, but it does not establish better safety-relevant recall or voluntary participation under laboratory hierarchies. The remaining claim is narrower: adding this governed enactment to the same strong checklist and briefing improves agreement about owners, deadlines, uncertainties, downstream parties, and stop conditions without pressure, added errors, or false safety assurance.
Smallest decisive test¶
With facility and independent safety approval, counter-order two fictitious custody-transfer scenarios for one volunteer cross-functional group. Condition A uses a strengthened checklist and conventional briefing; condition B adds the observance. Measure each participant’s answers about ownership, deadlines, uncertainty, downstream parties, and stop conditions immediately and 48-72 hours later. A blinded safety reviewer should code omissions, invalid commitments, and false closure, while anonymous responses assess accessibility, pressure, and certification confusion. Do not advance if B lacks the prespecified accuracy gain, creates more errors, makes refusal consequential, or is mistaken for safety certification.
Deployment and cost¶
The tabletop is assessed below $10,000. Startup, launch, and annual recurring effort are each roughly \(10,000-\)50,000 in 2026 resource-equivalent terms, not vendor quotes. Any operational use would require local labor, accessibility, privacy, records, and safety review, trained independent facilitation, protected opt-outs, and continued funding for the actual controls and commitments.
Risks and uncertainties¶
- Employees or junior researchers could experience visible participation as a loyalty test.
- Ceremonial completion could be mistaken for evidence that a material or process is safe.
- Token passing could become rote and suppress uncertainty instead of surfacing it.
- The session could substitute recognition for money, staff, engineering controls, or corrective action.
- Remote, disabled, multilingual, contract, night-shift, or waste-handling workers could be excluded from the shared account.
Expert review¶
Useful reviewer backgrounds: Nanomaterial laboratory scientist, Environment, health, and safety professional, Human-factors or organizational-behavior researcher, Labor and research-ethics specialist, Waste-management or downstream custody representative.
- Do actual transfers produce disagreement about owners, deadlines, escalation conditions, or downstream effects after current records are reviewed?
- What accuracy improvement over a strengthened checklist would justify the additional social and time burden?
- Can people decline every symbolic act without their refusal becoming visible or consequential?
- Which wording and facilitation practices best prevent ceremonial closure from being interpreted as safety certification?
- Who can independently halt or retire the practice if commitments remain unfunded or participants report pressure?
Evidence and provenance¶
Selected sources: S1: General Safe Practices for Working with Engineered Nanomaterials in Research Laboratories · S2: Nanomaterials · S3: Results of the 2019 Survey of Engineered Nanomaterial Occupational Health and Safety Practices · S4: Safety Management—Worker Participation · S5: Safe Science: Promoting a Culture of Safety in Academic Chemical Research · S6: Work Group Rituals Enhance the Meaning of Work · S7: The Point of No Return: Ritual Performance and Strategy Making in Project Organizations · S8: Occupational Health and Safety Specialists and Technicians
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc ordering ranged from rank 1 under deployment-heavy weights to 43 under impact-heavy weights, placing it in band B. This volatility reflects weighting and a cost proxy, not experimental success or economic value; the endpoint remains EMPIRICAL_PARTNER_CANDIDATE.
19. Ending Mutual-Aid Mandates Truthfully¶
Canonical title: Mutual-Aid Mandate Renewal and Release Assembly
In one sentence: This external-partner study candidate would add a witnessed renewal-and-release assembly to lawful closure and transition procedures, while prominently requiring evidence that the ritual layer improves shared understanding without coercing volunteers or stranding residents.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-32 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Sociology Anthropology |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 69.0/100 · rank range 14–25 across three profiles · band B |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
The problem¶
A neighborhood mutual-aid network may outlive the emergency that created it. Public channels, volunteer titles, pooled funds, routes, and phone trees can still signal readiness even when capacity and authority are unclear. Residents may rely on services that no longer operate; volunteers may feel unable to leave; and founders may retain accounts after nominal handoffs. Ordinary meetings can cancel tasks but may not create a legitimate, witnessed separation between ending volunteer roles and completing remaining duties.
What is proposed¶
Hold a quarterly and trigger-based Mandate Renewal and Release Assembly while any function remains open. Beforehand, stewards inventory advertised services, capacity, dependencies, funds, data, decision authority, reliance, and possible recipients. A status board marks each function as under review, and function cards enter provisional lanes: renew, transfer, suspend, repair, or release. Card placement is not a decision. Willing volunteers, lawful resource holders, and safeguards for affected residents must support the operational disposition. Closure then updates notices, rosters, escalation paths, funds, credentials, data, referrals, and deadlines. Participants may speak, submit privately, remain silent, delegate, or leave. Retiring a role never extinguishes debts, safeguarding duties, contracts, or promised transitions, and later audits compare public claims with actual capacity and custody.
The cross-domain transfer¶
The ritualized commitment archetype becomes a recurring, witnessed review of collective mandates, using marked time, function cards, reflection, renewed consent, release, handoff, debrief, and later audit. The mapping is coherent, but the distinct contribution of ceremony is uncertain because sunset rules, transition plans, demobilization, closure checklists, and volunteer recognition already perform much of the work.
Why it advanced¶
It entered the external-partner lane because it separates voluntary labor from surviving legal and service duties, names the necessary authorities, and offers a bounded comparator study. It did not enter strict success: no network has committed, stale-capacity prevalence is unknown, and there is no evidence that the assembly outperforms ordinary closure and transition practice.
Prior art and the remaining open claim¶
Emergency demobilization, nonprofit dissolution, service transfer, volunteer disengagement, debriefing, recognition, and care-oriented organizational closure are established. The open comparison begins only after a valid sunset rule, authorized closure checklist, and safe transition plan exist. The question is whether the witnessed assembly further improves volunteers’, residents’, and custodians’ accuracy about service status, authority, refusal rights, and unfinished duties without added coercion, disclosure, delay, or stranded work. It is a combination-level claim amid adjacent prior art.
Smallest decisive test¶
With a named network and legal or fiscal sponsor, run a compensated, preregistered tabletop with 24-40 volunteers, resident representatives, custodians, and potential partners. Counterbalance matched, noncritical mock mandates between ordinary sunset, checklist, and transition procedures and the same package plus the assembly. Measure service status, authority, refusal paths, notices, custody, and unfinished duties immediately and seven days later, alongside pressure, disclosure, distress, time, and false-closure beliefs. Reject incremental advantage if accuracy does not materially improve, burdens rise without compensating benefit, or any simulated duty is stranded. Halt for pressure, distress, privacy failure, or disputed authority.
Deployment and cost¶
The tabletop is assessed below $10,000; startup, launch, and annual recurring effort are each roughly \(10,000-\)50,000 in 2026 resource-equivalent terms, not vendor quotes. Live costs could rise with legal complexity, critical services, debts, accessibility needs, data cleanup, and funded transitions. Ordinary legal and operational authority remains mandatory.
Risks and uncertainties¶
- Residents could lose needed support before a capable alternative and truthful notice are in place.
- Public discussion could expose a resident’s need or a volunteer’s exhaustion.
- Recognition, gratitude, or lament could pressure volunteers to renew their labor.
- Founders could curate the evidence or retain credentials despite a witnessed handoff.
- Symbolic release could be misused to avoid debts, delete data improperly, or leave contracts and safeguarding duties incomplete.
Expert review¶
Useful reviewer backgrounds: Mutual-aid organizer or former volunteer, Affected-resident representative, Nonprofit or fiscal-sponsor lawyer, Data-protection and safeguarding specialist, Service-transition or emergency-demobilization practitioner.
- Which functions are legally controlled, informally controlled, safety-critical, or dependent on outside partners?
- Can volunteers withdraw immediately while residents still receive a safe and truthful transition?
- What evidence would show that the assembly adds understanding beyond a sunset clause, closure checklist, and facilitated transition meeting?
- How will private reliance statements and participation choices be protected from founders, employers, or public records?
- Who verifies that credentials, funds, notices, data, property, and surviving duties actually moved after the witnessed decision?
Evidence and provenance¶
Selected sources: S1: More Than a COVID-19 Response: Sustaining Mutual Aid Groups During and Beyond the Pandemic · S2: Covid Aid ceasing as a charity, with Support Community to continue under independent member ownership · S3: Demobilizing the VRC (2 of 2) · S4: Dissolution · S5: The Wind Down—For Better Organizational and Project Endings · S6: Group Termination: Completing the Study of Group Development · S7: Disengagement · S8: Occupational Employment and Wages—May 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading order placed it between ranks 14 and 25, in band B. The calculation uses affordability as a pilot-speed proxy and does not establish experimental success, deployment merit, or economic value; its endpoint remains EMPIRICAL_PARTNER_CANDIDATE.
20. Safer Cleanup of Post-Production Assets¶
Canonical title: Dependency-aware lifecycle governance for post-production assets
In one sentence: This post-hoc survivor proposes dependency-aware lifecycle rules for media-production files, but every major component has adjacent prior art and the complete workflow has not been compared with a manual sweep or configured asset-management system.
| Field | Record |
|---|---|
| Portfolio ID | EXP03-POSTHOC-02 |
| Experiment and endpoint | Experiment 3 · Post-hoc strict innovation-like survivor |
| Archetype × domain | Layer Decay And Expiration Management × Film Media Production |
| Proposal position or arm | Not recorded |
| Post-hoc reading order | Balanced score 68.0/100 · rank range 15–20 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A production accumulates edit versions, visual-effects renders, proxies, sound stems, exports, and project files whose usefulness changes over time. Old artifacts can remain visible as if current, consume active storage, and slow retrieval. Yet deleting them by age, name, or staff intuition can break project relinking, reconstruction, rights evidence, or future remastering. The core tension is distinguishing safely demotable clutter from creative or evidentiary material whose future dependencies are incomplete or hidden.
What is proposed¶
Assign each artifact an explicit lifecycle and creative-authority state, then assess dependencies, rights, legal holds, ownership, reconstruction needs, and restoration context before moving or deleting anything. Approved artifacts could move to cheaper storage or reversible quarantine; persistent markers would prevent restored or superseded files from silently appearing current again. Restore drills would test checksums, relinking, codecs, plug-ins, and software context, while recurring reviews would revisit earlier classifications. Owners would approve exceptions, archive or IT staff would verify recoverability, and legal or rights staff would control relevant holds and irreversible destruction. The proposal aims to bound active storage and search noise without silently sacrificing usable production history, but automatic deletion based only on age, access, score, or filename is excluded.
The cross-domain transfer¶
The layer-decay archetype maps to media files whose authority and usefulness expire at different rates. Lifecycle states nominate refresh, demotion, quarantine, or deletion, while dependency and restoration checks prevent unsafe transitions. The analogy is imperfect because dormant creative work can regain value years later and may depend on proprietary tools, external vendors, or undocumented practice.
Why it advanced¶
After Experiment 3, the candidate survived a stricter combined opportunity screen because the problem is concrete, existing technical primitives permit a bounded replay, and safety outcomes are measurable. This was a post-hoc result, not a preregistered success. Prevalence, net benefit, complete dependency capture, and adopter acceptance remain unestablished.
Prior art and the remaining open claim¶
Media-asset systems already archive and restore files with permissions; media ontologies represent versions and provenance; preservation standards cover rights, events, relationships, and technical dependencies; workflow tools track stale data; and cloud systems automate storage tiers. The only remaining contrast is the operating-policy combination: creative-authority state, dependency and reconstruction coverage, rights and owner gates, reversible quarantine, anti-staleness markers, restore drills, and recurring revalidation. The search did not establish novelty or superiority over a configured media-asset and archive workflow.
Smallest decisive test¶
On one completed, non-litigated production, conduct a read-only replay using a documented random sample of 200 artifacts from two classes. Independent reviewers compare a manual sweep, verified off-the-shelf asset-management rules, and the full proposal. Using isolated copies, simulate tiering and quarantine and test checksums, restoration, relinking, and software context. Compare labor, retrieval time, authoritative-version errors, safely demotable bytes, disputed labels, dependency misses, and restore success. Do not advance unless the full workflow materially improves a user outcome and active-tier burden after labor without worsening safety. Any missed live dependency, hold conflict, failed integrity check, relink, or restore defeats advancement.
Deployment and cost¶
First evidence is assessed at roughly \(10,000-\)50,000. Startup is roughly \(50,000-\)250,000, operational launch \(250,000-\)1 million, and annual recurring work \(50,000-\)250,000 in 2026 resource-equivalent terms, not vendor quotes. Actual costs depend on artifact volume, integrations, storage, licensing, legal review, vendor access, and recurring classification labor.
Risks and uncertainties¶
- The workflow could destroy unique creative work, contractual evidence, or material needed for a later reconstruction.
- Dependency graphs may omit external vendors, proprietary project formats, plug-ins, codecs, fonts, licenses, or tacit creative steps.
- Lifecycle labels could falsely imply that authority or disposal eligibility is settled.
- Archive latency and retrieval charges could disrupt an unexpected reuse request.
- Quarantine records and anti-staleness markers could themselves become another unmanaged layer.
Expert review¶
Useful reviewer backgrounds: Post-production supervisor or assistant editor, Media archivist or preservation engineer, VFX and sound pipeline engineer, Production IT or media-asset-management administrator, Entertainment rights, contracts, or records counsel.
- Can the production identify a canonical artifact and its authoritative status across edit, VFX, sound, and delivery systems?
- How complete are dependency records for external vendors, software versions, plug-ins, codecs, fonts, and licenses?
- Which outcomes would justify the added review labor compared with a configured media-asset system?
- Who has authority to approve demotion, quarantine, exceptions, holds, and irreversible destruction for each artifact class?
- What restore and relink failures are acceptable, if any, before all destructive transitions must stop?
Evidence and provenance¶
Selected sources: S1: Learning to Love End-to-End Cloud Production in Love Hurts · S2: Archive and Restore · S3: Ontology for Media Creation: Versions Ontology v2.7 · S4: Managing Audiovisual Records · S5: Depends: Workflow Management for Research and Visual Effects · S6: PREMIS Data Dictionary, Version 3.0, Now Available · S7: Archive a Blob · S8: National Employment and Wage Data by Occupation, May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading order placed it between ranks 15 and 20, in band B. This ordering is not an experimental endpoint or economic-value measure. Its endpoint remains POST_HOC_STRICT_INNOVATION_LIKE_SURVIVOR, not a preregistered success.
21. Show What Actually Solved the Model¶
Canonical title: Oracle-Relative Guarantee Passports for Macroeconomic Policy Models
In one sentence: A passport would distinguish what a macroeconomic policy platform computed itself from results supplied by solvers, data services, or people, while preserving bounded and unresolved outcomes.
| Field | Record |
|---|---|
| Portfolio ID | EXP05-STRICT-01 |
| Experiment and endpoint | Experiment 5 · Strict success |
| Archetype × domain | Computability Boundary Mapping × Economics Finance |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 68.0/100 · rank range 16–30 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
Policy models can contain unrestricted agent programs, continuous calculations, external equilibrium solvers, live data, and human choices among possible equilibria. Yet a final policy ranking may simply say that the model converged. That label can hide who or what supplied the answer, whether the supplier could abstain, whether other trajectories remained unresolved, and whether a finite timeout was mistaken for universal stability. The result may therefore claim more computational certainty than the platform actually established.
What is proposed¶
Before a model enters formal policy comparison, give it a capability-relative guarantee passport. Freeze the model language, economic target, equilibrium and stability definitions, time domain, and quantifiers. Define what the base platform can compute, then register each solver, data source, real-number routine, interactive environment, and human selection step as a named external capability with explicit promises, errors, latency, abstention, accountability, and version behavior. Test whether conclusions survive withdrawal, invalid responses, and promise violations. Search fairly, within declared bounds, for trajectories that refute convergence or target attainment. Route every result as BASE_CERTIFIED, ORACLE_RELATIVE, COUNTEREXAMPLE_FOUND, BOUNDED_ONLY, UNKNOWN, OUT_OF_MODEL, or SYSTEM_FAILURE. A timeout remains UNKNOWN, and a human-selected equilibrium remains assisted. Any claimed unrestricted impossibility reduction must be independently checked before use.
The cross-domain transfer¶
The computability-boundary archetype maps strongly here. External solvers, live information, and adaptive human choices act like added computational capabilities. The passport records which question the base program answers, which answer depends on a named capability and its promises, and which questions remain unresolved. Checked countertrajectories can refute precise universal claims without pretending to decide every possible model.
Why it advanced¶
This candidate passed Experiment 5's strict researched-candidate bar because the problem is consequential, the classifications and offline pilot are technically feasible, and existing model-governance functions could own them. That endpoint is only a researched-candidate result: it does not establish real-world effectiveness, novelty, economic value, adopter commitment, or authorization for policy use.
Prior art and the remaining open claim¶
Model inventories, validation, lifecycle governance, model cards, and documentation of human or third-party limits already exist. Numerical platforms also disclose dependence on initial guesses, iteration limits, and equilibrium selection. The narrower untested claim is that named capability contracts, promise checks, withdrawal tests, and mechanically preserved result labels will catch more capability-laundering errors than an ordinary model inventory and validation package, without changing the economic question being asked. No world-novelty finding was made.
Smallest decisive test¶
Preregister an offline experiment using six synthetic models and stubbed external services. Randomize cases between the ordinary inventory and validation template with raw logs, and the passport workflow. Blinded reviewers classify each result's computational dependency and status. Compare classification accuracy, false BASE_CERTIFIED results, preservation of UNKNOWN, promise-violation detection, review time, and semantic fidelity. Independently check any reduction and countertrajectory. Reject the incremental claim if accuracy does not improve, any false BASE_CERTIFIED label appears, UNKNOWN is lost, withdrawal cannot localize dependency, or at least two independent macroeconomists find that formalization materially changes the intended question.
Deployment and cost¶
The authorized first step is an offline synthetic pilot with no live policy instruments, confidential feeds, forecasts, or production materials. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(250,000–\)1 million annually. These are assessment bands, not vendor quotes.
Risks and uncertainties¶
- The formal equilibrium or target definition may omit what policy staff actually mean to evaluate.
- Staff may treat capability labels as paperwork and still collapse assisted, bounded, and unknown results into one ranking.
- Analysts may make adaptive equilibrium choices outside the registered procedure.
- A solver or data service may change behavior without a visible version change.
- Some solver promises may depend on semantic or empirical conditions that cannot be checked mechanically or conservatively screened with confidence. A base-certified calculation may still use an empirically poor economic model and produce a misleading policy conclusion. A valid impossibility result for unrestricted models may be wrongly extended to finite or promised subclasses.
Expert review¶
Useful reviewer backgrounds: Macroeconomist who develops or validates policy models, Computability or formal-methods researcher, Numerical equilibrium-solver specialist, Central-bank model-governance lead, Independent policy-model reviewer.
- Can the proposed formal convergence property preserve the economic question used by policy staff without silently narrowing it?
- For each external solver or human step, which promise conditions can be checked before its output is accepted?
- Can the unrestricted model class and policy query support a correct, independently checkable undecidability reduction?
- Do downstream policy materials preserve ORACLE_RELATIVE and UNKNOWN labels, or collapse them into a single ranking?
- Against ordinary validation documentation, how many dependency-classification errors does the passport prevent, and at what review cost?
Evidence and provenance¶
Selected sources: S1: Annual Report 2025, Box 7: Macroeconomic modelling in times of uncertainty · S2: The Dynare Reference Manual, version 7.1: The model file · S3: No-Regret Learning in Games is Turing Complete · S4: Supervisory Guidance on Model Risk Management, SR 26-2 Attachment · S5: Model Cards for Model Reporting · S6: Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 · S7: ECB macroeconometric models for forecasting and policy analysis: Development, current practices and prospective challenges · S8: Employer Costs for Employee Compensation—March 2026
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed this candidate in Band B, ranking 16–30 across three weighting profiles. That score is only a post-hoc reading order; it is not an experimental endpoint or economic-value measure, and its pilot-speed input is a cost-band proxy.
22. Keep Wetland Records Stable Across Systems¶
Canonical title: Representation-Independent Evidence Ledger for Wetland Greenhouse-Gas Monitoring
In one sentence: A behavioral evidence ledger would test whether different storage systems produce the same wetland monitoring history, coverage, provenance, and greenhouse-gas aggregates without exposing their internal layouts.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-19 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Environmental Climate |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 68.0/100 · rank range 11–29 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-20, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Wetland observations, laboratory results, corrections, and quality decisions may live in workbooks, database tables, raster layers, or files. Analysis scripts can quietly depend on row order, worksheet names, sentinel values, grid resolution, or a database's duplicate rules. A storage migration or regridding can then change which observations count, how revisions apply, or how values are aggregated even though the nominal fields remain. That makes it hard to tell a scientific change from an implementation-induced discontinuity.
What is proposed¶
Define the monitoring record by its behavior, not by its files or tables. The ledger would accept observation batches, retain superseded versions and reasons, apply or revoke quality decisions, answer version-specific queries, compute only declared variable-appropriate aggregates, report coverage and provenance, and create repeatable snapshots. Submissions must include stable identity, units, spatial and temporal support, method, and authorized provenance. Rejected operations must leave state unchanged and return stable errors for malformed data, incompatible units, duplicate identities, unsupported aggregation, unauthorized revision, or unavailable coverage. Workbooks, relational stores, rasters, indexes, and caches remain hidden behind the interface. Every backend must pass the same examples, generated operation sequences, and representation-leakage audit. Breaking semantic changes require a new contract version rather than reinterpretation of an existing snapshot.
The cross-domain transfer¶
The representation-independent interface archetype maps directly to storage substitution: different physical systems should denote the same versioned evidence state and obey the same operations, invariants, errors, and side-effect rules. The mapping is weaker for scientific validity. Backend agreement cannot show that field measurements, quality policies, ecological assumptions, or aggregation conventions are themselves correct.
Why it advanced¶
This candidate did not enter the strict-success lane. It cleared a separately calibrated lane for a bounded external data-partner study because the components are implementable and the claim is testable. Crucially, there is no real dependency inventory, authorized replay, measured discrepancy rate, comparative result, committed wetland partner, or project-specific cost evidence.
Prior art and the remaining open claim¶
Observation schemas, provenance standards, greenhouse-gas repositories, automated quality checks, versioning, and ecological workflow tools already provide close constituent parts. The remaining claim is narrower: for one frozen wetland workflow, a contract covering identity, supersession, quality decisions, units, coverage, provenance, errors, snapshots, and variable-aware aggregation will make an incumbent adapter and an independent backend agree on every management-relevant public result without clients inspecting storage internals. That comparison has not been run.
Smallest decisive test¶
With written owner approval, copy one completed project's data into a read-only sandbox and freeze 80–200 real or redacted operations spanning measurement, correction, quality review, snapshot, and aggregation. Compare the unchanged workflow, a schema-normalized export, and the behavioral contract implemented by both an incumbent adapter and an independent in-memory model. Add generated duplicates, reordered calls, incompatible units, revoked decisions, incomplete coverage, and authorization failures. Pass only with zero unexplained differences across declared results and no internal access. Reject the problem locally if no internal dependence exists; reject the intervention if conforming implementations still produce a management-relevant divergence or require exposing the incumbent layout or algorithm.
Deployment and cost¶
The first study must remain read-only and cannot alter official balances, records, eligibility, crediting, compliance, or management decisions. Rough 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for launch, and \(50,000–\)250,000 annually; they are not vendor quotes.
Risks and uncertainties¶
- The contract may preserve an incorrect scientific convention because incumbent behavior is mistaken for intended meaning.
- Unit or spatial-support normalization may combine observations that should instead be rejected as incompatible.
- An adapter defect may be misdiagnosed as a difference between the underlying storage systems.
- Finite conformance tests may miss operation sequences that alter a management-relevant result.
- Opaque storage may impede legitimate scientific inspection unless provenance and approved diagnostic views are adequate. Stable identifiers and detailed provenance may disclose sensitive site information if access controls are weak. A broad contract may become expensive to govern, while a narrow one may omit decision-critical behavior.
Expert review¶
Useful reviewer backgrounds: Wetland greenhouse-gas monitoring scientist, Environmental data architect, Scientific provenance and standards specialist, Carbon-program monitoring or methodology lead, Data-governance and confidential-location reviewer.
- Which observation, correction, quality, coverage, and aggregation behaviors can change a management-relevant greenhouse-gas result?
- Do any current client scripts inspect worksheet coordinates, schemas, file paths, raster cells, or private status codes?
- Which variables are extensive or intensive, and what aggregation rules preserve their scientific meaning?
- Can two independently implemented backends pass the contract yet still disagree on an official or management-relevant output?
- What access, retention, and disclosure rules apply to site identities and provenance in the proposed sandbox?
Evidence and provenance¶
Selected sources: S1: 2013 Supplement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories: Wetlands · S2: VM0033 Methodology for Tidal Wetland and Seagrass Restoration, v2.1 · S3: Global Greenhouse Gas Watch: Data Management · S4: World’s Largest Carbon Program Pilots Digital Measuring of Forest Carbon · S5: Observations, Measurements, and Samples · S6: PROV-N: The Provenance Notation · S7: OpenGHG 0.19.0 Developer API · S8: Developing a modern data workflow for regularly updated data
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized review placed this candidate in Band B, ranking 11–29 depending on weighting. This is a reading-order aid, not its endpoint or a value estimate; the pilot-speed input reflects cost-band affordability rather than measured elapsed time.
23. Retire Outdated Course Guidance Safely¶
Canonical title: Lifecycle Registry for Accumulated Course Scaffolds
In one sentence: An artifact registry would help course teams find stale guidance across copied course offerings while protecting accessibility supports, assessment dependencies, and records that must remain recoverable.
| Field | Record |
|---|---|
| Portfolio ID | EXP04-STRICT-04 |
| Experiment and endpoint | Experiment 4 · Strict success |
| Archetype × domain | Layer Decay And Expiration Management × Education Pedagogy |
| Proposal position or arm | PROPOSAL_FIRST |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 18–27 across three profiles · band B |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
Repeated course copies accumulate hints, examples, rubrics, refreshers, prompts, accommodations, and corrective notes. Older items may remain visible after the syllabus, assessment, learner group, or linked resource changes, leaving students with contradictory or obsolete guidance. Staff cannot safely clean everything up: an old-looking item may still support an accommodation, assessment, grade review, accreditation record, or reconstruction of a prior offering. Current course tools expose some clutter and broken links but do not resolve these lifecycle decisions.
What is proposed¶
Give each instructional-support artifact an identity, owner, source course version, lifecycle state, validation date, dependencies, retention class, and provisional expiry trigger. A dashboard would flag suspected stale items and rank review priority using age, present access, contradiction, ownership, and reconstruction value, but a person would decide what happens. Authorized reviewers could refresh, retain, unpublish, archive, compact, hold, or place an item in reversible quarantine. Accessibility, assessment, grade-review, accreditation, legal, and records dependencies must be checked before removal. Retired items leave a marker pointing to a successor, archive, or explicit unavailable state. Temporary supports receive review dates when created. Periodic revalidation revisits active items and exception holds, while restore tests confirm that archived materials remain readable, attributable, and connected to the correct historical course.
The cross-domain transfer¶
The layer-decay archetype has a clear structural match: each course offering deposits another support layer, and old layers can look current. Lifecycle states, expiry reviews, dependency checks, quarantine, retirement markers, and restore tests bound the active layer without erasing protected history. Age is only a review signal, however; it cannot determine whether instructional content remains valid.
Why it advanced¶
This candidate passed Experiment 4's strict researched-candidate bar because institutions already perform course cleanup and retention work, relevant owners exist, and a read-only comparison is feasible. The endpoint does not mean the registry works in practice, is novel, saves money, improves learning, or has institutional authorization beyond a bounded study.
Prior art and the remaining open claim¶
Manual course audits, link validation, cleanup tools, content templates, version distribution, retention schedules, and archives are established. The open contrast is whether an artifact-level registry spanning course generations improves identification of current versus stale supports and surfaces assessment, accessibility, and records dependencies better than a careful manual audit aided by Canvas Link Validator and TidyUP. It must also preserve recovery and add no more than 20% reviewer time. Only that integrated, measured comparison remains open; world novelty was not assessed.
Smallest decisive test¶
With instructor, LMS, accessibility, privacy, and records approval, use four sandboxed snapshots of one course, excluding submissions, grades, and identifiable accommodation records. Sample at most 80 instructor-created supports and compare the existing manual audit plus Link Validator and TidyUP with the same evidence plus the registry. Two authorized reviewers classify currentness, visibility, dependencies, exceptions, and proposed state; seed up to eight synthetic defects. Proceed only with sensitivity of at least 0.85, false flags at most 0.10, Cohen's kappa at least 0.70, no protected-data exposure, and median review time within 20% of baseline. Reject incremental advantage if classification or dependency recall does not improve.
Deployment and cost¶
Begin with a read-only inventory; do not change visibility, content, grades, assessments, accommodations, or retention. Rough 2026 resource-equivalent bands are under $10,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(50,000–\)250,000 annually. These are assessment bands, not vendor prices.
Risks and uncertainties¶
- Reviewers may mistake age for staleness even when an old explanation remains correct.
- Low access counts may unfairly flag essential but rarely used accessibility or safety materials.
- Incomplete links to assessments, external tools, grade reviews, or accommodations may make a load-bearing artifact appear safe to retire.
- A composite score may create false precision or encode local pedagogical preferences.
- Quarantine could conflict with mandatory destruction rules or retain sensitive material too long. Lifecycle labels may confuse students if exposed without careful wording. Registry metadata and temporary holds may themselves become stale. Successful restore tests on sampled formats may conceal unreadable materials elsewhere in the archive.
Expert review¶
Useful reviewer backgrounds: Instructional designer, Instructor experienced with repeated LMS course copies, LMS administrator or integration specialist, Accessibility and disability-services representative, Academic records, privacy, or retention officer.
- Can artifacts be matched reliably across copied course shells without confusing distinct items or missing descendants?
- Which dependencies can the LMS and external tools expose, and which still require human review?
- What local rules distinguish ordinary instructional supports from protected educational or grade-evaluation records?
- Do reviewers agree on currentness, contradiction, dependencies, and lifecycle state at the required thresholds?
- Does the registry improve detection over Link Validator, TidyUP, and manual review without exceeding the 20% time limit?
Evidence and provenance¶
Selected sources: S1: Declutter your Learn.UQ course checklist · S2: Course Data Purge Process · S3: Postgraduate Students’ Experience of Using a Learning Management System to Support Their Learning: A Qualitative Descriptive Study · S4: How do I validate links in a course? · S5: TidyUP · S6: Course Content Distribution Comparison · S7: Grades and Records of Student Performance — Academic Policy 1480.10 · S8: Occupational Employment and Wage Statistics: Instructional Coordinators
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The harmonized review placed the candidate in Band B, with ranks from 18 to 27 across weighting profiles. This post-hoc ordering does not alter its strict endpoint or measure economic value; pilot speed was represented only by a cost-band proxy.
24. Test Cold-Case Theories on Equal Terms¶
Canonical title: Parallel-Hypothesis Gate for Cold-Case Resources
In one sentence: A bounded hypothesis gate would compare rival cold-case explanations using equal records, preregistered predictions, capped resources, and rights-aware scoring without turning advancement into a finding of guilt.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-05 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Criminology Forensic |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 24–28 across three profiles · band B |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
Cold-case units may have several plausible explanations but limited analyst time, specialist review, and laboratory capacity. The incumbent team can control the chronology, evidence requests, and briefing order, allowing one theory to advance through ownership or presentation rather than discriminating evidence. Rival teams can also duplicate work, hoard information, or escalate intrusive requests. A resource-allocation ranking may then be mistaken for proof about a named person even though it only selects what to examine next.
What is proposed¶
For at least two distinct, minimally supported hypotheses, create separated teams with equal access to a read-only, citation-addressable record. Before new results arrive, each team registers its theory, contrary evidence, falsifiers, discriminating predictions, uncertainty, proposed tests, and expected privacy, evidence-consumption, and third-party burdens. Freeze eligibility, scoring, conflicts, appeals, and prohibited conduct in advance. Equal analyst-hour and submission caps limit escalation. An independent panel scores evidence coverage, falsifiability, distinctiveness, provenance, treatment of contradictions, feasibility, and rights-adjusted information value. At most two complementary hypotheses receive a bounded validation slot, not investigative authority. Advancement establishes neither guilt, probable cause, admissibility, nor permission for contact, search, surveillance, testing, charging, or public accusation. New evidence or failed predictions can reopen entry, and later review compares preregistered predictions with results.
The cross-domain transfer¶
The bounded-rivalry archetype maps to several hypotheses competing for scarce analysis and testing. The gate defines who may compete, equalizes inputs and budgets, penalizes harmful off-process behavior, permits more than one bounded award, and reopens competition after new evidence. The structural mapping is plausible, but rivalry could intensify team commitment or strategic withholding instead of improving reasoning.
Why it advanced¶
This candidate did not reach the strict-success lane. It cleared a separate empirical-partner lane because a masked retrospective comparison is measurable and can be kept from affecting cases. The central field evidence is missing: no agency partner, prevalence estimate, comparative trial, validated scoring rubric, measured local cost, or evidence that team separation is safer than collaboration exists.
Prior art and the remaining open claim¶
Multidisciplinary cold-case review, Analysis of Competing Hypotheses, explicit alternatives, evidence for and against each theory, ordered information exposure, provenance, and auditable forensic reasoning already exist. Research also warns that formal hypothesis layouts do not reliably reduce bias. The remaining claim is that separated teams, identical cutoff records, registered predictions, equal budgets, rights-adjusted scoring, and no more than two validation slots outperform both case conferences and ordinary ACH without increasing unsupported allegations, intrusion, evidence use, entrenchment, withholding, or panel inconsistency.
Smallest decisive test¶
With an authorized partner, preregister a three-condition retrospective crossover using 8–12 masked, time-split closed or synthetic cases: the proposed gate, a multidisciplinary conference, and a noncompetitive ACH worksheet. Give each condition identical cutoff records and analyst hours, rotate qualified participants, and prevent case recognition or outcome leakage. Before revealing later evidence, capture citations, contradictions, probabilities, predictions, actions, expected information value, privacy burden, and evidence consumption. Blinded reviewers assess calibration, discrimination, citation completeness, unsupported allegations, redundancy, reliability, time, and intrusion. Reject the claim if the gate fails to beat the better comparator or worsens calibration, agreement, withholding, entrenchment, allegations, privacy burdens, or evidence-consumption proposals.
Deployment and cost¶
Only a retrospective, no-case-impact simulation using masked copies is authorized initially; it cannot reopen a case or affect any person or evidence. Rough 2026 resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for startup and operational launch, and \(1–\)5 million annually. They are not vendor quotes.
Risks and uncertainties¶
- Assigned teams may become more committed to their hypotheses and resist contrary evidence.
- Teams may withhold urgent exculpatory information until scoring or relabel similar explanations to qualify separately.
- Scorers may reward narrative confidence or specificity rather than evidentiary value and calibration.
- Historical cutoff packets may omit context legitimately available to investigators at the time.
- Case recognition or later-outcome knowledge may leak to participants or reviewers. Resource caps may disadvantage a genuinely complex explanation. Even confidential rankings could stigmatize investigators or people named in hypotheses. Officials or the public may misrepresent advancement as evidence of guilt or authority for coercive action.
Expert review¶
Useful reviewer backgrounds: Cold-case or major-case review supervisor, Forensic scientist or laboratory evidence custodian, Investigator trained in competing-hypothesis analysis, Prosecutor, defense-disclosure, or criminal-procedure specialist, Privacy, research-ethics, and experimental-design reviewer.
- Can masked cutoff packets provide equal and sufficiently complete records without revealing case identities or later outcomes?
- Does the scoring rubric produce reliable rankings across independent panels and resist gaming through confidence or narrative style?
- Compared with case conferences and ACH, does the gate improve held-out calibration and rights-adjusted information value?
- Does team separation increase withholding, hypothesis entrenchment, unsupported allegations, privacy burden, or duplicative evidence requests?
- Which local authorities must approve record access, disclosure handling, laboratory recommendations, retention, and participant involvement?
Evidence and provenance¶
Selected sources: S1: National Best Practices for Implementing and Sustaining a Cold Case Investigation Unit · S2: NamUs Cold Case Advisory Process · S3: Conducting Effective Investigations: Practice Evidence · S4: Tunnel Vision and Confirmation Bias Among Police Investigators and Laypeople in Hypothetical Criminal Contexts · S5: Psychology of Intelligence Analysis · S6: Effects of Task Structure and Confirmation Bias in Alternative Hypotheses Evaluation · S7: OSAC 2022-S-0030 Standard Methodology in Bloodstain Pattern Analysis, Version 2.1 · S8: Disclosure Manual: Chapter 5 — Reasonable Lines of Enquiry and Third Parties
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized review placed this candidate in Band B, ranking 24–28 across weighting profiles. That ordering is neither its empirical-partner endpoint nor an economic-value measure, and the pilot-speed input is an affordability proxy rather than observed study duration.
25. Send Raman Surprises, Preserve Full Evidence¶
Canonical title: Versioned Residual Raman Monitoring for Remote Reaction Experiments
In one sentence: A shadow study would test whether transmitting reconstructable differences from predicted Raman spectra can reduce data and review burdens without hiding consequential chemical or instrument changes.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-11 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Predictive Residual Processing × Chemistry Materials |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 67.0/100 · rank range 12–39 across three profiles · band B |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-12, EXP06-PARTNER-13, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Remote reaction monitoring can produce a dense stream of Raman spectra. Most of each spectrum may reflect predictable solvent bands, baseline behavior, or gradual temperature effects, yet every spectrum still consumes transmission, storage, and chemist attention. A small unexpected peak, shifted band, or changed intensity may therefore be reviewed late. Simply sending selected peaks or anomaly labels is also risky because it can hide unfamiliar changes and deny the chemist enough context to reconstruct or challenge the result.
What is proposed¶
Before each acquisition, matched, versioned models at the instrument and review station would predict the next wavelength-aligned spectrum from accepted prior data and declared reaction conditions. The instrument computer would subtract that frozen prediction from the complete observation and transmit the signed correction, uncertainty, context, heartbeat, and model checksum. The receiver could reconstruct the spectrum by adding the correction to its synchronized prediction. Reliable, persistent, or consequential differences would bring the corresponding full spectral window to a named chemist. Random complete spectra and scheduled full-state anchors would provide independent checks. Missing heartbeats, incompatible versions, poor detector quality, stale calibration, structured errors, excessive reconstruction error, or protected safety conditions would restore full-spectrum transmission. Independent safety instruments and the original reaction controls would remain outside this system.
The cross-domain transfer¶
The predictive-residual archetype maps directly onto the spectral stream: predict one complete spectrum, compare it with the actual detector output, communicate the signed mismatch, and reconstruct the observation at the receiver. Version checks, raw audits, full anchors, and automatic decompression address the central danger that both endpoints could share the same wrong expectation.
Why it advanced¶
This is an empirical-partner candidate, not a strict success or real-world validation. Real-time Raman acquisition and spectral compression already exist, and the proposed shadow comparison is technically bounded. It advanced because its measurements and rejection rules are unusually concrete, despite having no field evidence that the target workflow suffers a binding bandwidth or attention constraint.
Prior art and the remaining open claim¶
Adjacent prior art includes commercial real-time Raman systems with multivariate predictors, compressive Raman acquisition, and standardized lossless or near-lossless spectral compression. The open claim is narrower: for one instrument, probe, reaction family, and one-scan horizon, a checksum-gated reconstructive residual stream with raw audits, heartbeats, full anchors, protected bypasses, and forced fallback would use fewer total resources than both full-spectrum streaming and ordinary lossless compression while preserving blinded decisions, protected perturbations, and regional error limits.
Smallest decisive test¶
A laboratory partner would run twelve non-hazard-escalating reactions: eight ordinary runs and four containing approved spectral, calibration, detector, heartbeat, and version challenges. Every raw spectrum would remain authoritative. Blinded chemists would compare full streaming, lossless compression, and the frozen residual path on total bytes, storage, computation, audit and fallback traffic, review minutes, decisions, reconstruction error, and protected-event capture. Reject the claim if any protected challenge is missed, reconstruction changes a material decision, a heartbeat or checksum fault fails to force raw mode, an error limit is exceeded, or net savings are absent.
Deployment and cost¶
The first shadow study requires instrument and API access plus bench-chemist, instrumentation-owner, safety, and information-security approval. Its rough first-evidence band is \(10,000–\)50,000 in 2026 resource equivalents. Estimated startup is \(50,000–\)250,000; operational launch \(250,000–\)1 million; and annual recurring work \(50,000–\)250,000. These are assessment bands, not vendor quotes.
Risks and uncertainties¶
- A shared wrong predictor could suppress an unexpected low-amplitude peak at both endpoints while producing a plausible reconstruction.
- Instrument fouling or changing chemistry could gradually be treated as normal rather than as evidence of a new regime.
- Alignment error, quantization, or lost corrections could distort later reconstructed spectral shapes.
- A chemist-approved consequence table could still privilege anticipated bands and discount unfamiliar features.
- Random raw audits may miss rare failures; a clean small sample cannot demonstrate completeness or justify later raw-data deletion.
Expert review¶
Useful reviewer backgrounds: Raman spectroscopy and chemometrics specialist, Reaction or process chemist, Laboratory instrumentation engineer, Laboratory safety and data-integrity specialist, Statistical audit-sampling expert.
- Which low-amplitude or unfamiliar spectral changes must always bypass suppression for the selected reaction family?
- What regional and cumulative reconstruction limits would preserve the chemist's actual reaction-state decisions?
- Does the intended workflow have measured transmission, storage, queue-delay, or review burdens that ordinary lossless compression cannot address?
- What random-audit size and stratification would have adequate power to reveal rare suppressed changes?
- Which instrument faults and calibration conditions can be injected safely without changing the controlling record?
Evidence and provenance¶
Selected sources: S1: PAT — A Framework for Innovative Pharmaceutical Development, Manufacturing, and Quality Assurance · S2: Raman spectrometry as a tool to study minimization of batch age effects and make product quality decisions for biotherapeutic antibody production · S3: Relative Intensity Correction Standards for Fluorescence and Raman Spectroscopy · S4: Questions and Answers on Current Good Manufacturing Practice Requirements—Records and Reports · S5: Raman Rxn2 analyzer powered by Kaiser Raman technology · S6: Integration of a Raman spectroscopic platform based on online sampling to monitor chemical reaction processes · S7: Recent Trends in Compressive Raman Spectroscopy Using DMD-Based Binary Detection · S8: CCSDS 123.0-B-2: Low-Complexity Lossless and Near-Lossless Multispectral and Hyperspectral Image Compression
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The harmonized scores place this candidate between ranks 12 and 39, in ordering band B. That range is only a post-hoc reading aid using an affordability proxy; it is not an experimental endpoint, economic-value estimate, or change to partner-candidate status.