Candidate dossiers — Band A¶
Part of Inverse Innovation with the Encyclopedia of Abstractions · Ranks 1–10: highest post-hoc review priority · Last revised August 2026
1. Fairer Selection for a Festival Slate¶
Canonical title: Campaign-Bounded Selection for a Film Festival Slate
In one sentence: A festival would limit lobbying and unequal access, standardize review, and choose qualifying films for their contribution to a declared slate, then compare that process with its ordinary workflow.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-01 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Bounded Rivalry Governance × Film Media Production |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 73.0/100 · rank range 1–6 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
The problem¶
A festival has fewer competition slots than eligible films, yet entrants may receive unequal access to programmers, deadlines, screening attention, exceptions, or opportunities to supply extra material. Private pitches, gifts, sponsor pressure, and intermediary relationships can influence visibility. Screeners may also carry uneven workloads or undisclosed conflicts. The process can therefore reward campaign resources and access rather than each film’s contribution to the festival’s stated purpose, while leaving entrants unable to challenge clear procedural errors.
What is proposed¶
For one competition section, publish the purpose, available slots, eligibility rules, accepted materials, review stages, conflicts, individual quality threshold, slate-level criteria, tie-breaks, and procedural appeal route before submissions open. Ban gifts, private selection pitches, undisclosed referrals, and entrant-initiated lobbying. Randomly assign each eligible film to two screeners with capped caseloads, record recusals, and obtain a third review when assessments diverge materially. Only films clearing the individual floor reach slate selection. The panel then chooses a feasible, nonredundant portfolio using disclosed considerations such as runtime balance, programmatic perspectives, and audience pathways. An independent reviewer audits sampled files. Appeals may correct process errors but may not replace curatorial judgment. Selection applies to one edition only.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a clearly delimited competition for scarce screenings. A published rulebook defines the permitted arena; material and contact limits dampen campaign escalation; workload controls, recusals, sanctions, and audits constrain manipulation; and portfolio selection directs rivalry toward contribution to the whole slate. Post-edition review supplies a revise-or-retire decision.
Why it advanced¶
Experiment 6 classified this candidate as STRICT_SUCCESS because the problem, plausible adopters, component feasibility, reversible shadow test, and contrast with ordinary selection were sufficiently supported for its strict researched-candidate bar. That status does not establish real-world effectiveness, distinctiveness, deployment authority, or economic impact; the proposal remains a partnered research program.
Prior art and the remaining open claim¶
Transparent rules, controlled submission materials, assigned viewers, second reviews, conflict recusals, authorized exceptions, and intentional slate composition already exist in festival practice. The open claim is narrower: compared with a festival’s current workflow, combining standardized access, random double review, workload ceilings, an individual floor, disclosed portfolio constraints, sampled audit, and process-only appeals will reduce treatment discrepancies and access-linked variation without unacceptable losses in curatorial fit, reviewer attention, or timeliness. That advantage has not been demonstrated.
Smallest decisive test¶
With one festival, preregister a shadow comparison on 60–120 opt-in short-film submissions. Apply ordinary selection and the proposed workflow to the same files without changing official outcomes. Compare missing reviews, time, agreement, recusals, expertise reassignments, threshold stability, portfolio reasons, audit errors, appeal corrections, slate overlap, concentration, and blind coherence ratings. Reject the claim if discrepancies do not fall, agreement worsens, expertise corrections exceed 20%, portfolio choices breach the floor, median labor rises over 50% without matching benefit, appeals reopen taste, or auditors cannot separate rule breaches from lawful curation.
Deployment and cost¶
The authorized first step is a consented shadow pilot, not a live selection change. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence and startup, and \(50,000–\)250,000 for operational launch and annual recurring work. These are bottom-up assessment bands, not vendor quotes; no festival partner has committed.
Risks and uncertainties¶
- Standard material limits may omit cultural, production, safety, or accessibility context needed to interpret a film responsibly.
- Portfolio criteria could become vague cover for favoritism or token selection, including admission below the declared quality floor.
- Double screening and audits could overload programmers, delay decisions, or force a smaller review pool.
- A communications firewall may still favor films already legible through credits, press coverage, or established intermediaries.
- Audit samples may miss selective misconduct or wrongly treat innocent intermediary concentration as evidence of wrongdoing.
Expert review¶
Useful reviewer backgrounds: Festival artistic director or programming lead, Festival operations and submissions manager, Independent procedural auditor, Film-sector conflict, privacy, and labor counsel, Curatorial evaluation or portfolio-design researcher.
- Can the festival define portfolio criteria precisely enough to guide decisions without disguising unrestricted discretion?
- What caseload ceiling permits substantive double review within the existing calendar and staffing budget?
- Can auditors reliably distinguish a correctable process violation from a lawful curatorial judgment?
- Which contextual materials must remain available so standardization does not systematically disadvantage particular films?
Evidence and provenance¶
Selected sources: S1: Film Festival Best Practices · S2: 2025 Sundance Film Festival Unveils Short Film Program Presented by Vimeo · S3: Submitting Your Project to the 2026 Sundance Film Festival: FAQ · S4: BFI London Film Festival Assistant Programmer Job Pack · S5: Festival Entry Regulations 2026 · S6: Terms of Submission—Official Selection 2026 · S7: Pre-Jury Transparency in Short Film Festivals Organized in Turkey within the Framework of Gatekeeping Theory · S8: Socioeconomic Factors of National Representation in the Global Film Festival Circuit: Skewed Toward the Large and Wealthy, but Small Countries Can Beat the Odds
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: In the post-hoc harmonized reading order, this candidate scored 73–74 across profiles, ranked between 1 and 6, and fell in band A. The ordering used cost-band affordability as a pilot proxy; it is neither an experimental endpoint nor an estimate of economic value.
2. A Common Meaning for Evidence Records¶
Canonical title: Representation-Independent Evidence Continuity Record
In one sentence: A forensic organization would define custody records by their permitted actions and observable meaning, then require any replacement system to pass the same behavior-based tests rather than merely matching fields.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-09 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Representation Independent Interface Contract × Criminology Forensic |
| Proposal position or arm | P1 |
| Post-hoc reading order | Balanced score 73.0/100 · rank range 2–7 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-10, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Forensic evidence systems can encode custody, seals, sampling, consumption, and corrections in different tables, status codes, display orders, or editable notes. During migration or system integration, two implementations may contain similar fields yet report different current custodians, unresolved transfers, seal histories, sample ancestry, remaining quantities, or corrections. Without a representation-independent rule, an organization cannot determine whether a new implementation preserves the procedural meaning of the original record rather than merely copying its visible data.
What is proposed¶
Define an Evidence Continuity Record as an abstract state machine, not a database layout. Its public operations would register items, offer and accept transfers, record seal opening and resealing, derive child samples, record consumption, append linked corrections, query current state, render authorized history, and check named continuity conditions. Each operation would specify permissions, preconditions, results, errors, and side-effect limits. Invariants would prohibit unmatched completed transfers, negative quantities, unlinked samples, silent deletion, and more than one current accountable custodian. Storage tables, vendor codes, indexes, internal identifiers, and row order would remain non-contractual. A shared black-box suite would run examples, generated action sequences, authorization checks, and representation-leakage probes before an implementation could qualify as a substitute.
The cross-domain transfer¶
The representation-independent interface archetype becomes a behavioral contract for one evidence record. Abstract states and operations define what custody history means; an opaque boundary prevents clients from depending on vendor internals; and a common conformance oracle tests whether different implementations preserve the same authorized observations, errors, transitions, and invariants.
Why it advanced¶
Experiment 6 classified this as STRICT_SUCCESS because heterogeneous forensic systems, migration difficulty, relevant institutional authority, established software techniques, and a safe synthetic test were sufficiently supported for the strict researched-candidate bar. This does not validate the contract, authorize production migration, establish legal sufficiency, demonstrate distinctiveness, or measure economic impact.
Prior art and the remaining open claim¶
Custody standards, common schemas, ontologies, audit trails, access controls, two-sided transfers, and migration tools already cover much of the surrounding territory. None of that alone establishes behavioral equivalence between implementations. The remaining claim is that an opaque, vendor-neutral state-machine contract plus one behavior-only conformance and leakage suite will identify semantically acceptable substitutes more reliably than schema matching, audit-trail presence, or vendor-specific acceptance tests. It remains open until the suite detects seeded defects and survives review of material legal and scientific distinctions.
Smallest decisive test¶
With a laboratory quality authority and authorized reviewer, preregister 12–20 synthetic scenarios covering transfers, seals, samples, consumption, corrections, repeated calls, unauthorized queries, and exports. Run them against a transparent reference model, an incumbent-behavior adapter, and at least five mutated adapters containing specified semantic defects. Advance only if the first two agree on every mandatory assertion, expose no prohibited information, and the suite rejects every mutant. Revise or reject the contract if an incumbent sequence is unrepresentable, a mutant passes, or reviewers find a material contract-observable difference after both implementations pass.
Deployment and cost¶
Begin only with synthetic data and sandbox adapters; do not alter live evidence or infer admissibility. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup and annual work, and \(250,000–\)1 million for operational launch. They are assessment estimates, not vendor quotes.
Risks and uncertainties¶
- The abstract state may omit jurisdiction-specific signatures, native documents, instrument metadata, or another legally or scientifically material distinction.
- Tests may copy incumbent quirks and mistakenly preserve them as required behavior.
- A passing suite could be misrepresented as proof that recorded events are true or that evidence is admissible.
- Incomplete authorized export operations could make the opaque boundary obstruct legitimate audit or disclosure.
- Generated tests may miss failures that appear only in long, concurrent, or malformed action sequences.
Expert review¶
Useful reviewer backgrounds: Forensic laboratory quality manager, Evidence custodian or property-unit supervisor, Forensic LIMS architect or integration engineer, Evidence and disclosure counsel, Software conformance and property-based testing specialist.
- Which representation-specific artifacts are legally or scientifically material and therefore must be contract-observable?
- Can an adapter reproduce incumbent behavior without relying on undocumented or mutable vendor features?
- Do the proposed invariants cover every authorized transfer, correction, sampling, and consumption sequence in the bounded scope?
- What mutant set would provide a credible test that the suite distinguishes field similarity from semantic equivalence?
Evidence and provenance¶
Selected sources: S1: Evidence Management Steering Committee Report: Opportunities to Strengthen Evidence Management Processes · S2: Laboratory Information Management Systems in Forensic Science Service Provider Laboratories: Current State and Next Generation · S3: Landscape Study of Software-Based Evidence Management Systems for Law Enforcement · S4: OSAC 2021-N-0018: Standard for On-Scene Collection and Preservation of Physical Evidence, Version 2.0 · S5: ISO 22095:2020 — Chain of custody — General terminology and models · S6: CASE Ontology Design and Specification · S7: Privacy Impact Assessment for the Laboratory Information Management System · S8: Best Practices for Authenticating Digital Evidence
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading order scored this candidate 73–74, placed it between ranks 2 and 7, and assigned band A. That order used a cost-band affordability proxy and is not an experimental endpoint, legal finding, or measure of economic value.
3. Truthful Limits for Safety Verification¶
Canonical title: Assurance-Boundary Gate for Extensible Safety-Control Designs
In one sentence: A design organization would classify what a safety verifier can honestly decide before routing each controller to exact checking, bounded or approximate analysis, or accountable human review.
| Field | Record |
|---|---|
| Portfolio ID | EXP04-STRICT-02 |
| Experiment and endpoint | Experiment 4 · Strict success |
| Archetype × domain | Computability Boundary Mapping × Engineering Design |
| Proposal position or arm | PROPOSAL_FIRST |
| Post-hoc reading order | Balanced score 72.0/100 · rank range 1–4 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(250,000–\)1 million |
The problem¶
An organization asks one verifier to return a correct, terminating SAFE or UNSAFE answer for every programmable safety controller, even when submissions include unrestricted scripts, unbounded variables, and supplier extensions. The actual computation model and guarantee are left unclear. Engineers may then treat timeouts as failures, keep adding computing power to a requirement that may be undecidable, or quietly restrict accepted designs while continuing to make a universal assurance claim. Release records may not reveal which bounds or assumptions support a verdict.
What is proposed¶
Place an assurance-boundary gate before verification and release. The gate would formalize the controller language, state representation, safety property, environmental assumptions, quantifiers, computation model, and requested guarantee. It would admit exact terminating verification only for mechanically enforceable finite or otherwise decidable fragments. Richer designs would go to sound over-approximation, explicitly bounded exploration, or accountable expert escalation. UNKNOWN, OUT_OF_SCOPE, TIMEOUT, and TOOL_FAILURE would remain separate from SAFE and UNSAFE. Claims of impossibility would require a checked reduction or equivalent proof, not repeated timeouts. Every result would record its fragment, bounds, abstraction, assumptions, guarantee, residual uncertainty, and triggers for reclassification when the model, tool, supplier interface, or environment changes.
The cross-domain transfer¶
Computability-boundary mapping becomes an intake and routing decision for controller assurance. The gate separates exact solvability, weaker machine-checkable evidence, and unresolved cases relative to an explicit model. Enforced language fragments and typed results prevent bounded searches, abstractions, tool failures, or nonanswers from inheriting a stronger universal safety claim.
Why it advanced¶
Experiment 4 classified this candidate as STRICT_SUCCESS because the boundary is technically grounded, component practices are mature, relevant authorities exist, and a reversible non-release comparison is clearly testable. Passing that strict researched-candidate bar does not show field prevalence, better safety outcomes, adopter acceptance, deployment authorization, economic impact, or broad distinctiveness.
Prior art and the remaining open claim¶
Finite-state model checking, abstraction, bounded exploration, UNKNOWN results, witnesses, and formal-assurance records are established practices. The narrower open claim concerns their mandatory ordering and governance: placing a model-relative solvability and scope classification plus an enforced status schema before an otherwise standards-conformant workflow will reduce over-strong or mislabeled verdicts and improve reviewer routing agreement, while losing no more than one conclusive decision in a 12-model pilot. The individual mechanisms are adjacent prior art, and comparative advantage has not been measured.
Smallest decisive test¶
Freeze 12 previously adjudicated controller models spanning finite, bounded, abstractable, unrestricted, unsafe, and timeout cases. Blind and randomize reviewer pairs to the strongest existing workflow or the proposed gate, using equal tools, evidence, and time. Compare adjudicated mislabeled verdicts, routing agreement, hours, nonanswer rates, and conclusive decisions. Advance only with zero mislabeled gate verdicts, improvement in mislabeling or agreement, at most one lost conclusive result, at most one fragment-membership disagreement, and cost below $50,000. Reject if weak or incomplete evidence becomes unqualified SAFE/UNSAFE, agreement fails to improve, or physical semantics cannot be reproduced.
Deployment and cost¶
The first step is a frozen, non-release pilot that cannot alter deployed logic or release status. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence and \(250,000–\)1 million for startup, operational launch, and annual recurring work. These are assessment estimates rather than vendor quotes or certification budgets.
Risks and uncertainties¶
- A correct proof may concern a model that omits important physical-controller or environmental behavior.
- An unsound abstraction could create false confidence; a coarse but sound abstraction could generate too many unresolved alarms.
- Designers may move excluded but safety-relevant behavior into informal escape hatches outside the decidable fragment.
- A finite analysis may exhaust resources, inviting staff to confuse infeasibility or timeout with a semantic verdict.
- The recorded boundary can become stale after language, verifier, environment, or supplier-interface changes.
Expert review¶
Useful reviewer backgrounds: Safety-control chief engineer, Independent formal-verification specialist, Design assurance or certification reviewer, Controller-language and toolchain engineer, Domain regulator or licensing specialist.
- Does the declared controller language actually include unrestricted computation, or is the accepted class already enforceably decidable?
- Can fragment membership and environmental assumptions be reproduced independently for every pilot model?
- Are the approximate-analysis modes sound, and how are spurious counterexamples labeled and escalated?
- Would the gate add information beyond the strongest existing standards-conformant workflow under equal time and tools?
Evidence and provenance¶
Selected sources: S1: IEC 61508-3:2010 — Functional safety of electrical/electronic/programmable electronic safety-related systems, Part 3: Software requirements · S2: Digital Instrumentation and Controls Research · S3: DOT/FAA/TC-19/22: Use of Virtual Machines in Avionics Systems and Assurance Concerns · S4: Explainable Verification for Rapid Certification · S5: What is Formal Methods? · S6: CBMC: What is loop unwinding? · S7: SV-COMP 2013 Competition Procedure and Verification Result Definitions · S8: National employment and wage data by occupation, May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: In the post-hoc harmonized reading order, this candidate scored 72–74, ranked between 1 and 4, and fell in band A. The ordering used affordability as a pilot proxy; it is not a preregistered endpoint, safety finding, or economic-value measure.
4. Keeping Test Norms Current and Traceable¶
Canonical title: Psychometric Norm Lifecycle Registry and Scoring Gate
In one sentence: A registry and scoring gate would keep context-mismatched norm packages out of new psychological scoring while preserving exact historical packages for authorized reconstruction.
| Field | Record |
|---|---|
| Portfolio ID | EXP05-STRICT-04 |
| Experiment and endpoint | Experiment 5 · Strict success |
| Archetype × domain | Layer Decay And Expiration Management × Psychology |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 72.0/100 · rank range 2–5 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
A psychological test may accumulate several norm tables, scoring transformations, cutoff rules, and population-specific calibration files. Older packages can remain in manuals, spreadsheets, scripts, or software without a clear current, restricted, superseded, or archived status. Users may then score against an unintended population or edition, producing inconsistent standardized results. Simply deleting old packages is also unsafe because prior reports, longitudinal studies, audits, or research reproductions may depend on the exact historical transformation.
What is proposed¶
Create a lifecycle registry linked to a scoring gate. Give every norm package an immutable identifier and record its instrument version, reference population, collection period, intended uses, transformations, cutoffs, validation evidence, software dependencies, latest review, successor, and lifecycle state. A review lease or relevant context change would move a package to review-due, not declare it invalid. New scoring would require an active, context-matched package or a documented expert override. Superseded packages would disappear from ordinary prospective selection but remain retrievable by identifier for authorized reconstruction. Dependency checks would block removal while reports, studies, longitudinal series, or audits still need a package. Quarantine, archival tiers, disposition markers, and fixed-vector restore drills would make retirement reversible and test whether historical scoring remains executable.
The cross-domain transfer¶
Layer-decay management becomes governance for successive norm and scoring packages. The registry identifies each deposited layer, detects expired reviews and changed contexts, and separates active authority from historical retention. Dependency checks, quarantine, archives, deletion markers, and restoration tests prevent cleanup from destroying the scoring basis of earlier work.
Why it advanced¶
Experiment 5 classified this candidate as STRICT_SUCCESS because norm changes can matter, relevant professional guidance and adopter classes exist, the software pattern is feasible, and a bounded read-only comparison is testable. This strict researched-candidate result does not establish field effectiveness, cross-publisher portability, distinctiveness, deployment authorization, or economic value; partnered research remains necessary.
Prior art and the remaining open claim¶
Professional guidance already calls for current, population-relevant, identifiable norms; publishers already renorm tests, label legacy products, retire scoring software, and maintain successor platforms. The open claim is the added effect of integrating immutable package identity, expiration of prospective authority, context-gated selection, dependency-constrained retirement, and tested restoration. Compared with ordinary files, manuals, or platform presentation, that package should reduce context-mismatched selection and improve exact historical reconstruction without excessive valid-use blocks or expert overrides. This comparative claim remains untested.
Smallest decisive test¶
Use one licensed or synthetic test family with at least three packages and randomly assign 20–30 qualified evaluators to the current presentation or a read-only registry across 24–40 scenarios. Compare mismatched selections and exact reconstruction, plus time, blocks, overrides, identity errors, and preservation choices; restore one archived package against a fixed vector. Advance only with at least a 10-point mismatch reduction, 15-point reconstruction gain, no more than a 5-point increase in valid-use blocks, at most 10% unnecessary overrides, zero identity errors, and exact restoration. Redesign if either primary outcome fails or users confuse lifecycle status with validity.
Deployment and cost¶
Start with a read-only inventory and mock gate; do not change production scores, reports, or source files. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup and annual work, and \(250,000–\)1 million for operational launch. They are not vendor quotes.
Risks and uncertainties¶
- Users may mistake an active lifecycle state for evidence that a norm is psychometrically valid for the individual case.
- Incomplete population, instrument, or intended-use metadata could falsely authorize a mismatched package.
- Missing references from reports, spreadsheets, printed tables, or local scripts could allow destructive retirement.
- Archived transformations may become non-executable as software environments and formats age.
- Registry access or preservation policies could expose restricted test content or retain sensitive norm data longer than authorized.
Expert review¶
Useful reviewer backgrounds: Psychometrician responsible for test norms, Test publisher or assessment-program owner, Practicing psychologist or qualified assessment user, Scoring-platform and archival systems engineer, Test-security, privacy, and records counsel.
- Can package metadata express intended population and use precisely enough to gate scoring without implying validity?
- What events should trigger review-due status, and who may renew, restrict, supersede, or override a package?
- How completely can inbound dependencies from reports, studies, scripts, and longitudinal datasets be discovered?
- Can an archived package reproduce the fixed historical score in a controlled environment without exposing restricted materials?
Evidence and provenance¶
Selected sources: S1: Guidelines for Practitioner Use of Test Revisions, Obsolete Tests, and Test Disposal · S2: ITC Guidelines on Test Use, Final Version 1.2 · S3: Model for the Review, Description and Evaluation of Psychological and Educational Tests, Version 5.0 · S4: The Flynn Effect: A Meta-analysis · S5: The Woodcock Reading Mastery Test: Impact of Normative Changes · S6: NEO Personality Inventory-3 (Normative Update) · S7: Pearson Scoring Software · S8: Occupational Employment and Wages—May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading order scored this candidate 72–74, ranked it between 2 and 5, and placed it in band A. That order used cost-band affordability as a proxy and is neither an experimental endpoint nor a measure of economic value.
5. Protected Quiet Periods Between Neural Stimuli¶
Canonical title: Protected Null Epochs for Closed-Loop Neural Stimulation
In one sentence: Explicitly reserved, carefully labeled periods without stimulation could reveal whether apparent neural responses reflect the current stimulus or lingering effects from earlier ones.
| Field | Record |
|---|---|
| Portfolio ID | EXP04-STRICT-05 |
| Experiment and endpoint | Experiment 4 · Strict success |
| Archetype × domain | Negative Space Design × Neuroscience |
| Proposal position or arm | PROPOSAL_FIRST |
| Post-hoc reading order | Balanced score 71.0/100 · rank range 2–6 across three profiles · band A |
| First-evidence resource band | under $10,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
The problem¶
A closed-loop neural experiment may deliver another stimulus whenever timing and safety rules allow, treating unused time as wasted. If the brain has not returned to a stable reference state, responses can overlap with residual activation, adaptation, or changing neural state. Researchers may then attribute a response to the current stimulus when part of it came from preceding trials. Sparse or ambiguously recorded quiet periods also make scheduling histories harder to compare and reproduce.
What is proposed¶
The experiment generator would omit selected low-priority stimuli at genuine block or state boundaries and protect the resulting null epochs from automatic backfilling. Recording, synchronization, contextual logging, and safety monitoring would continue throughout each pause. Codes would distinguish an intentional pause, a recovery-triggered pause, an operator decision, and equipment or command failure. Stimulation would resume at a predeclared time or when an approved recovery measure reached a bounded criterion, with a maximum pause and manual override. The comparator is a dense schedule paired with the best history-dependent response model; additional comparisons include a uniformly longer interval and randomized omission. The aim is to separate stimulus effects from sequence-history effects without sacrificing unacceptable condition coverage.
The cross-domain transfer¶
The negative-space archetype becomes deliberately protected time in the experiment schedule. The omitted event creates a measured reference interval connected to the next response, while explicit boundaries and state codes give the absence a clear meaning. The mapping is strong: the empty space is an active experimental condition, not merely slower scheduling or missing data.
Why it advanced¶
This candidate passed Experiment 4's strict researched-candidate bar because the problem is measurable, offline testing is feasible, the intervention has explicit safety and rollback boundaries, and its strongest rivals are directly comparable. STRICT_SUCCESS denotes passage of that research screen only; it does not establish real-world effectiveness, novelty, deployment permission, or economic value.
Prior art and the remaining open claim¶
Null events, washout periods, adaptive stimulation, history-dependent models, and audit logs already exist, so this is adjacent prior art. The narrower open claim is that recovery-bounded, explicitly coded null epochs outperform dense scheduling with model correction, uniformly longer intervals, and random omission. Any advantage must remain after matching stimulated-trial count and elapsed time and must be large enough to justify reduced condition coverage.
Smallest decisive test¶
Using one synchronized historical session, preregister a response model containing stimulus identity, recent history, elapsed time, and prestimulus state. Estimate recovery without held-out trials, then replay fixed-gap, recovery-threshold, and matched-random-omission policies. Compare held-out response error, stimulus-versus-history identifiability, baseline stability, coverage, and simulated duration against the dense schedule with its best history model and a uniformly longer interval. Reject the proposal if recovery-based spacing does not improve the preregistered identifiability measure, random omission or modeling matches it, coverage falls below its floor, or recovery estimates cannot support bounded resumption.
Deployment and cost¶
The authorized first step is offline analysis; live stimulation requires separate ethics, institutional-safety, and possibly device-regulatory approval. Estimated 2026 resource-equivalent bands are under $10,000 for first evidence, \(10,000–\)50,000 for initial deployment, \(50,000–\)250,000 for operational launch, and \(10,000–\)50,000 annually. These are assessment bands, not vendor quotes.
Risks and uncertainties¶
- Fewer stimulated trials may reduce statistical precision or leave required conditions underrepresented.
- Replacing omitted trials could lengthen sessions, increasing participant fatigue or animal burden.
- A recovery-triggered rule could preferentially sample particular neural states and introduce selection bias.
- Faulty state codes could make equipment failure look like intentional silence.
- An unreliable recovery signal could repeatedly suppress conditions or destabilize the closed loop.
Expert review¶
Useful reviewer backgrounds: Closed-loop neural-stimulation researcher, Neural time-series statistician, Research ethics or animal-care specialist, Neurotechnology safety engineer, Experimental-design methodologist.
- Can the available event-level data distinguish recovery dynamics from ordinary time drift and recent stimulus history?
- What recovery measure and maximum pause could be specified without using held-out outcomes?
- Which conditions must remain above a prespecified coverage floor after omissions?
- Does model-based correction match the proposed schedule after trial count and elapsed time are equalized?
- Which approvals would a live pilot require for the specific device, population, and protocol?
Evidence and provenance¶
Selected sources: S01: Learning to Control the Brain through Adaptive Closed-Loop Patterned Stimulation · S02: Influence of Inter-Stimulus Interval on Auditory Evoked Potentials · S03: Stochastic Designs in Event-Related fMRI · S04: A Deep Brain Stimulation Trial Period for Treating Chronic Pain · S05: Subcortical Short-Term Plasticity Elicited by Deep Brain Stimulation · S06: Neural Recording and Modulation · S07: IDE Approval Process · S08: Center for Research Informatics Price Guide, FY26
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: A post-hoc harmonized score placed the candidate between ranks 2 and 6 across three reading profiles, in band A. This ordering is not an experimental endpoint or economic-value measure; its pilot-speed input is only a cost-affordability proxy.
6. A Safer Accessibility Testing Challenge¶
Canonical title: Accessibility Barrier Discovery Portfolio Challenge
In one sentence: A sealed, batch-scored challenge would reward a complementary set of reproducible accessibility barriers instead of whichever reports arrive first.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-06 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Bounded Rivalry Governance × Human Computer Interaction |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 71.0/100 · rank range 6–9 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
The problem¶
An accessibility bounty with a fixed reward pool can reward speed or report volume rather than useful discovery. Testers may submit scanner output, reserve findings before documenting them, split one barrier into several reports, or withhold reproduction details. Reviewers then spend time resolving duplicates while difficult, task-blocking barriers remain poorly described. Unequal automation, unsafe testing, and access to live or personal data can also shift costs onto users, maintainers, and less-resourced testers.
What is proposed¶
Run bounded rounds in a synthetic-data environment with equal access windows, scoped accounts, submission and request caps, sealed reports, and published rules. Reports must describe the affected task, starting state, interaction sequence, observed barrier, expected behavior, environment, reproducible evidence, and remediation-relevant trace. Authorization, privacy, fixture integrity, and reproducibility are mandatory gates. Verified reports are scored for task obstruction, clarity, reproducibility, distinctness, and repair usefulness. Awards are selected as a portfolio covering complementary journeys, interaction modes, and failure classes, rather than paid in filing order. Independent verification, separate appeals, predefined foul rules, recurring challenger entry, and post-round review address favoritism, sabotage, collusion, lock-in, and rubric gaming.
The cross-domain transfer¶
The bounded-rivalry archetype becomes a controlled contest for a finite accessibility reward pool. Rivalry remains, but a rulebook, equal resource ceilings, sealed batches, safety gates, portfolio awards, appeals, and recurring access limit how contestants can gain advantage. The structural mapping is direct, although whether rivalry adds value over collaboration remains untested.
Why it advanced¶
This candidate cleared the separate EMPIRICAL_PARTNER_CANDIDATE lane because a controlled mock study can compare allocation systems without touching production. It did not pass the strict-success lane. Crucially, there is no field evidence that accessibility-specific rivalry improves coverage over a well-run paid panel, and no organization has committed authority, staff, funding, or an environment.
Prior art and the remaining open claim¶
Accessibility audits, disability-led participatory testing, automated scanning, usability studies, and first-valid bug bounties already cover much of the work. The open claim concerns the combined allocation system: with expertise, scope, safety rules, and reward resources held constant, sealed batches, complementarity-based awards, and equal caps should yield more distinct, independently reproducible, repair-usable barriers per reviewer hour than either first-valid allocation or a noncompetitive paid panel, without added harm or exclusion.
Smallest decisive test¶
Run a preregistered two-round crossover on an isolated prototype with synthetic accounts, four essential journeys, hidden seeded barriers, and about 12 compensated testers spanning assistive technologies and input modes. Compare the proposed challenge with a flat-fee panel, then rescore all reports under first-valid allocation. Measure distinct reproduced barriers, seed recall, coverage, repair usefulness, reviewer time, duplicates, fragmentation, burden, urgent-report delay, appeals, and safety events. Advance only if the portfolio condition beats both comparators by the prespecified material margin—suggested as 20% per reviewer hour—without worse safety, delay, or loss of manual work.
Deployment and cost¶
Begin with a nonmonetary mock using fictitious points and equal base compensation. Estimated 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup and operational launch, and \(250,000–\)1 million annually. No direct pricing or internal labor study confirms these bands; they are not vendor quotes.
Risks and uncertainties¶
- Submission caps may suppress slow, uncertain, manual, or assistive-technology-intensive findings.
- Batch sealing may delay urgent accessibility or security disclosure.
- Portfolio scoring may encode the sponsor's incomplete view of important journeys and interaction modes.
- Identity or collusion screens may wrongly combine independent testers or treat suspicious patterns as guilt.
- Synthetic journeys may omit contextual, longitudinal, or socially mediated barriers found in real use.
Expert review¶
Useful reviewer backgrounds: Disabled accessibility tester, Accessibility evaluation methodologist, Bug-bounty program designer, Security and privacy reviewer, Experimental economist or contest-design researcher.
- Can independent scorers apply the complementarity rubric consistently across assistive-technology configurations?
- Do equal caps disproportionately exclude manual or assistive-technology-intensive investigation?
- What urgent-disclosure path preserves safety without compromising sealed allocation?
- Does the competitive condition outperform an equally funded noncompetitive panel after reviewer time is counted?
- Which identity and collusion signals can support investigation without becoming automatic penalties?
Evidence and provenance¶
Selected sources: S1: Guidance on Web Accessibility and the ADA · S2: Selecting Web Accessibility Evaluation Tools · S3: Tips for Usability Testing with People with Disabilities · S4: WCAG Evaluation Methodology (WCAG-EM) 2.0 · S5: Detailed Platform Standards · S6: Fable: Accessibility Research, Powered by People with Disabilities · S7: Crowdsourced Security Vulnerability Discovery: Modeling and Organizing Bug-Bounty Programs · S8: Vulnerability Disclosure Policy
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized reading aid placed this candidate between ranks 6 and 9 across profiles, in band A. It does not change the partner-candidate endpoint or measure economic value; its pilot-speed input is a cost-affordability proxy.
7. Consistent Rules for Climate Signpost Alerts¶
Canonical title: Behavioral Contract for Adaptive-Policy Signposts and Alerts
In one sentence: A representation-neutral behavioral contract would require spreadsheets, dashboards, and monitoring services to interpret the same climate signpost history in the same way.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-24 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Representation Independent Interface Contract × Futurism Foresight |
| Proposal position or arm | P3 |
| Post-hoc reading order | Balanced score 71.0/100 · rank range 7–12 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-23, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Coastal adaptation signposts may be copied among policy documents, spreadsheets, dashboards, and provider pipelines. Each implementation can handle dates, revisions, missing readings, sustained thresholds, units, acknowledgments, and retirement differently. Two systems can therefore produce different review alerts from the same observations even when their indicator names and displayed thresholds match. A software or data-source change may silently advance, delay, repeat, or suppress a policy review while officials believe the governing commitment is unchanged.
What is proposed¶
Create an Adaptive Signpost Register with defined operations for registering a versioned signpost, ingesting observations, marking missing periods, revising or retracting data, evaluating status, issuing and acknowledging alerts, suspending evaluation, superseding a signpost, and replaying history. The contract specifies units, time rules, persistence, missingness, revision behavior, typed errors, side effects, and lifecycle states. It requires provenance, an immutable event history, one active version, deterministic replay, and a firm separation between an advisory alert and authority to act. Storage layouts, formulas, provider payloads, caches, and visual presentation stay hidden. A replacement is accepted only if a shared black-box test suite produces the same governed states, alerts, errors, and audit records.
The cross-domain transfer¶
The representation-independent-interface archetype becomes a behavioral boundary around climate signposts. Different technical implementations may store and display data differently, but must expose the same operations, states, errors, replay behavior, and alerts. The mapping is strong at the software layer, though it cannot settle whether an indicator or threshold is scientifically or politically appropriate.
Why it advanced¶
This proposal cleared the EMPIRICAL_PARTNER_CANDIDATE lane because two offline implementations and a synthetic event corpus can test behavioral equivalence safely. It did not enter the strict-success lane. There is no field measurement of signpost divergence, no integrated prototype comparison, and no coastal authority has committed specifications, staff, historical data, or adoption authority.
Prior art and the remaining open claim¶
Adaptive pathways already use signposts and triggers; observation schemas, event sourcing, conformance tests, and migration reviews also exist. The remaining claim is narrower: independently built implementations following one behavioral contract should agree on governed outputs and expose more seeded semantic mistakes than schema validation plus manual spot checks. The contract fails if important divergences survive, approved semantics require provider-specific internals, or the existing comparator finds the same defects with materially less effort.
Smallest decisive test¶
With a named authority, freeze two approved synthetic signpost specifications and 30 histories covering missingness, late and duplicate data, corrections, retractions, persistence, acknowledgments, suspensions, supersession, and replay. Seed at least 12 semantic mutations, including missing-as-zero, arrival-time ordering, one-sample breaches, destructive revisions, duplicate alerts, and hidden rounding. Separate people build an event-ledger model and spreadsheet adapter. Compare contract tests with schema validation plus current manual checks. Reject the proposal if any decision-relevant mutant is missed, conforming systems disagree, approved rules require implementation leakage, or the comparator achieves equivalent detection at materially lower staff effort.
Deployment and cost¶
The first step keeps synthetic observations and copied specifications offline; it must not issue live alerts or change policy thresholds. Estimated 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(50,000–\)250,000 annually. These are not vendor quotes.
Risks and uncertainties¶
- Formal consistency may create false confidence in a weak indicator or poorly chosen threshold.
- The contract may hide contested judgments about time, missingness, or revisions inside technical rules.
- An incomplete test generator may omit consequential event sequences.
- Overly strict tests may freeze harmless choices such as display rounding or notification format.
- A conforming replacement may still have unacceptable latency or resource use outside the initial test.
Expert review¶
Useful reviewer backgrounds: Climate adaptation-pathways specialist, Municipal resilience policy owner, Data-contract or API architect, Geospatial observation-data steward, Software conformance-testing specialist.
- Can each approved signpost be expressed without provider-specific storage or payload fields?
- Which states and alerts are advisory, and which authority separately initiates a policy review?
- Does replay preserve the interpretation of histories created under earlier specification versions?
- Which semantic mutations would materially alter a real review decision?
- How much staff effort does the contract suite require compared with current schema and migration checks?
Evidence and provenance¶
Selected sources: S1: Dynamic adaptive policy pathways: A method for crafting robust decisions for a deeply uncertain world · S2: Designing a monitoring system to detect signals to adapt to uncertain climate change · S3: Taking an adaptive approach: Thames Estuary 2100 · S4: Coastal hazards and climate change guidance · S5: A Summary of Adaptation Pathways Approaches · S6: OGC SensorThings API Part 1: Sensing · S7: Event sourcing pattern · S8: National employment and wage data by occupation, May 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized aid ranked the candidate from 7 to 12 across reading profiles, in band A. This is neither an experimental endpoint nor an economic-value estimate; the pilot-speed input reflects cost-band affordability, not independently measured elapsed time.
8. Remembering Why Financial Controls Exist¶
Canonical title: Why This Control Exists: Control-Memory and Repair Cycle
In one sentence: A carefully bounded reenactment would help control owners remember a control's verified origin, dependencies, and reciprocal duties, then route accepted changes through ordinary governance.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-PARTNER-26 |
| Experiment and endpoint | Experiment 6 · Empirical-partner candidate |
| Archetype × domain | Ritualized Meaning And Commitment Enactment × Accounting Auditing |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 71.0/100 · rank range 3–19 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(10,000–\)50,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-11, EXP06-PARTNER-27, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
A financial-reporting control may survive staff turnover as a checklist while its rationale and dependencies disappear. Preparers and reviewers can still perform the documented step without knowing which past failure it addresses, whose information it protects, what upstream inputs it assumes, or what conduct must continue. Generic certifications and refresher training may reinforce procedure without restoring that shared understanding. The result can be mechanical compliance, repeated handoff failures, obsolete controls, and false reassurance that remediation is complete.
What is proposed¶
Once each quarter, a cross-functional group would examine one cleared control family through a synthetic or redacted transaction path. A steward first verifies the history, separates fact from interpretation, removes protected information, and identifies affected stakeholders. Participants use functional roles and accessible cards to trace where information was lost, duplicated, misclassified, or repaired. Upstream and accounting roles exchange reciprocal commitments about reliable inputs, validation, feedback, and escalation. Recognition is optional and consent-based; participants may revise or decline proposed commitments without escaping formal job duties. Accepted changes move into authorized narratives, responsibility maps, remediation plans, budgets, or escalation records with owners and dates. Debrief and independent review may revise, pause, repair, or retire the practice.
The cross-domain transfer¶
The ritualized-meaning archetype becomes a marked, recurring examination of why a control exists and what people owe one another to keep it workable. Reenactment, symbolic framing, witnessing, reciprocal exchange, renewal, and retirement are present. The mapping is structurally clear, but evidence that these ritual features outperform simpler instruction in this domain is absent.
Why it advanced¶
This candidate cleared the EMPIRICAL_PARTNER_CANDIDATE lane because a synthetic, content-matched trial can measure learning and harm without changing live controls. It did not pass the strict-success lane. No accounting-specific evidence shows added benefit over a walkthrough, and no controller, internal-control leader, audit committee, or funder has committed to a pilot.
Prior art and the remaining open claim¶
Root-cause reports, training, simulations, certifications, governance software, audit testing, process mining, and ordinary knowledge transfer are established comparators. Ritual features have adjacent workplace and educational evidence, but not for durable financial-control operation. The open claim is that consent-governed symbolic reenactment, reciprocal obligation exchange, and revisable renewal improve seven-day recall, responsibility mapping, and workflow-ready follow-through beyond a content-equivalent walkthrough or ordinary training, without added coercion, blame, religious conflict, or confusion about control evidence.
Smallest decisive test¶
Recruit 36–60 volunteers, stratified by tenure, into three equal-content and equal-time arms: the full cycle, a facilitated walkthrough without ritual features, and ordinary case-based training. Use only synthetic materials. Blind-score immediate and seven-day recall of facts, causal sequence, assertions, dependencies, duties, authority limits, escalation, and the boundary between learning artifacts and control evidence. Reject the proposal if the full cycle trails the best comparator by the prespecified 10-point advantage, adds no workflow-readiness benefit, produces blame or pressure above 10%, exposes opt-outs to supervisors, fails an accommodation, or causes anyone to treat ceremonial closure as proof of effectiveness.
Deployment and cost¶
The authorized first step is one synthetic 75-minute session with 8–12 volunteers; it cannot change or evaluate a live control. Estimated 2026 resource-equivalent bands are \(10,000–\)50,000 for first evidence and startup, and \(50,000–\)250,000 for operational launch and annual recurrence. No time study, wage benchmark, or quote validates these bands.
Risks and uncertainties¶
- A selective origin story may legitimize leadership or suppress disputed facts.
- Functional reenactment may still expose or symbolically blame identifiable employees.
- Visible passing, silence, or dissent may become legible to supervisors despite formal consent rules.
- Participants may mistake recognition or ceremonial closure for evidence that a control works or remediation is finished.
- The practice may become costly training theater while underlying resources, authority, and incentives remain unchanged.
Expert review¶
Useful reviewer backgrounds: Corporate controller or internal-control leader, Internal-audit specialist, Organizational learning researcher, Employment, privacy, and religious-accommodation adviser, Accessible facilitation and workplace-safety specialist.
- Can the selected control history be verified and presented without exposing confidential or identifiable information?
- Does the ritual arm improve seven-day recall beyond an otherwise identical walkthrough?
- Can employees decline symbolic participation without supervisors learning or inferring that choice?
- Do proposed commitments enter authorized remediation workflows with owners, resources, and dates?
- Can participants reliably distinguish the learning exercise from control evidence, approval, or certification?
Evidence and provenance¶
Selected sources: S1: Importance of Audits of Internal Controls · S2: Improving Organizational Performance and Governance: How the COSO Frameworks Can Help · S3: The Importance of a Comprehensive Risk Assessment by Auditors and Management · S4: IIA Response to IAASB Proposed Standard on Audits of Financial Statements of Less Complex Entities · S5: Resolving Internal Control Deficiencies and Restatements · S6: Work Group Rituals Enhance the Meaning of Work · S7: Experiential Learning: A Game Changer for Accountants · S8: Section 12: Religious Discrimination
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized aid placed this candidate between ranks 3 and 19 across profiles, in band A. The broad range reflects different reading priorities. It is not an experimental endpoint or economic-value measure, and affordability only proxies pilot speed.
9. Predictive Checks for Evidence Custody¶
Canonical title: Residual Assurance for Physical-Evidence Custody Transitions
In one sentence: A frozen workflow model would flag meaningful differences between expected and recorded evidence-custody transitions while preserving the complete ledger and requiring human review.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-05 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Predictive Residual Processing × Criminology Forensic |
| Proposal position or arm | P2 |
| Post-hoc reading order | Balanced score 71.0/100 · rank range 5–17 across three profiles · band A |
| First-evidence resource band | \(50,000–\)250,000 |
| Initial deployment startup band | \(250,000–\)1 million |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-04, EXP06-PARTNER-14, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Forensic laboratories must preserve every handoff, storage move, checkout, return, and seal change for physical evidence. Routine entries can overwhelm reviewers, allowing a missing scan, unauthorized custodian, location mismatch, or broken seal to remain unnoticed. Static rules catch known violations but may miss problems whose meaning depends on the item’s prior state. Yet treating every unusual entry as suspected tampering can create needless quarantines and unsupported conclusions about evidence or personnel.
What is proposed¶
Keep the append-only custody ledger as the authoritative record, then add a separate assurance layer. For one evidence class and workflow, a frozen, versioned model predicts the next authorized custodian, location, seal state, and timing window from the last fully reconciled state. Each actual event is compared with that prediction and labeled as missing, extra, duplicated, late, reordered, unauthorized, or otherwise inconsistent. A gate weighs source reliability, uncertainty, integrity consequences, and reviewer capacity before routing the discrepancy, with its full context, to a named human reviewer. Missing evidence, seal changes, unauthorized access, recorder failure, legal holds, and reviewer requests bypass filtering. Heartbeats verify that silence is genuine, complete-ledger audits test quiet cases, and drift or reconstruction failures return affected items to manual reconciliation.
The cross-domain transfer¶
The predictive-residual archetype becomes a custody-state monitor: expected transitions provide the prediction, recorded scans provide the observation, and their directional difference becomes the residual. The mapping is structurally strong because the workflow has observable states and ordered transitions, but the model remains subordinate to the immutable ledger, physical reconciliation, and human judgment.
Why it advanced¶
This candidate passed Experiment 6’s strict researched-candidate bar because the problem is consequential, most technical components already exist, responsible laboratory authorities and a bounded retrospective test are identifiable, and strong safeguards are specified. STRICT_SUCCESS does not mean real-world validation, novelty, deployment authorization, or demonstrated economic impact.
Prior art and the remaining open claim¶
Barcode and RFID tracking, forensic laboratory systems, deterministic alerts, complete-ledger review, physical inventory, and process-conformance checking already provide adjacent parts. The remaining claim is narrower: combining a frozen custody predictor, structured directional residuals, verified heartbeats, mandatory full-context bypasses, audits of quiet ledgers, and automatic fallback can improve protected-discrepancy recall, acknowledgement speed, reconstruction fidelity, false escalation, and total review workload against optimized rules and full-ledger review.
Smallest decisive test¶
Run a preregistered, read-only three-arm crossover using synthetic ledgers or approved closed-case records from one laboratory workflow. Qualified blinded reviewers compare complete-ledger review, optimized deterministic rules, and the residual interface on seeded custody errors, outages, version changes, and legitimate exceptions. Measure protected-discrepancy recall, acknowledgement time, false escalation, unsupported tampering interpretations, reconstruction error, audit misses, fallback success, reviewer time, and maintenance effort. Reject the approach after any unexplained protected-class miss, irrecoverable reconstruction, failed silence-versus-outage distinction, or failure to improve the joint recall-latency-workload criterion over both comparators.
Deployment and cost¶
The first evidence study is estimated at \(50,000–\)250,000 in rough 2026 resource-equivalent terms. Initial deployment and operational launch are each estimated at \(250,000–\)1 million, with \(50,000–\)250,000 annually. These are planning bands, not quotes; workflow integration, validation, security, training, discovery rules, and accreditation review remain substantial.
Risks and uncertainties¶
- A repeatedly used but unauthorized workflow could become the model’s expected pattern.
- Failed heartbeats could make missing observations look like legitimate quiet periods.
- Reviewers could interpret an unusual transition as evidence of tampering rather than a discrepancy requiring reconciliation.
- Threshold changes intended to quiet the queue could conceal integrity-relevant events.
- One missing or reordered event could corrupt later state reconstruction until full resynchronization occurs unless fallback works correctly and promptly.
Expert review¶
Useful reviewer backgrounds: Forensic laboratory quality manager, Evidence custodian, Forensic LIMS and data-integration engineer, Process-conformance or state-estimation specialist, Forensic accreditation and legal-discovery specialist.
- Which custody discrepancies must have zero unexplained misses in the retrospective test?
- Can approved records or synthetic sequences represent realistic missing, delayed, duplicated, reordered, and emergency-exception events?
- How accurately can the complete custody state be reconstructed after recorder outages or version changes?
- Do quiet-ledger audits reveal material discrepancies that the predictor suppresses or never represents?
- What jurisdiction-specific retention, discovery, privacy, labor, cybersecurity, and accreditation rules govern residual metadata?
Evidence and provenance¶
Selected sources: S1: Evidence Management Survey Results · S3: Evidence Management Steering Committee Report: Opportunities to Strengthen Evidence Management Processes · S4: A Landscape Study of Laboratory Information Management Systems (LIMS) for Forensic Crime Laboratories · S5: RFID Technology in Forensic Evidence Management: An Assessment of Barriers, Benefits, and Costs · S6: Conformance Checking: Foundations, Milestones and Challenges · S7: State of Michigan Contract 071B4300089: Laboratory Information Management System · S8: Laboratory Information Management Systems in Forensic Science Service Provider Laboratories: Current State and Next Generation
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed this candidate between ranks 5 and 17 across three profiles, in band A. That score only sets reading order, uses an affordability proxy for pilot speed, and is neither an experimental endpoint nor an estimate of economic value.
10. A Stable Contract for Justice Histories¶
Canonical title: Observation-Bounded Justice Event Trajectory Contract
In one sentence: An opaque data contract would make differently structured justice datasets return the same study-defined event histories, corrections, observation limits, and uncertainty states.
| Field | Record |
|---|---|
| Portfolio ID | EXP06-STRICT-10 |
| Experiment and endpoint | Experiment 6 · Strict success |
| Archetype × domain | Representation Independent Interface Contract × Criminology Forensic |
| Proposal position or arm | P4 |
| Post-hoc reading order | Balanced score 70.0/100 · rank range 8–22 across three profiles · band A |
| First-evidence resource band | \(10,000–\)50,000 |
| Initial deployment startup band | \(50,000–\)250,000 |
Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-09, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.
The problem¶
Longitudinal justice studies often inherit meaning from a warehouse’s rows, null values, keys, joins, and default ordering. A booking row may be treated as an event, a missing field as proof of absence, or a later correction as if it had always been known. Rebuilding the warehouse or analysis layer can therefore change cohorts, event counts, or timing without changing the intended study definition, while inadequate observation may be silently converted into a negative finding.
What is proposed¶
Define each pseudonymous subject’s trajectory through an opaque contract rather than a table layout. The contract stores source-attributed event assertions, occurrence and recording times, equivalence links, non-erasing corrections, explicit periods of source observation, and versioned classification policies. Authorized operations add assertions, relate duplicates, append corrections, register observation periods, reconstruct what was known at a cutoff, and classify study windows. A window returns one of three answers: a qualifying event was recorded, none was recorded during adequate observation, or observation was insufficient. Invalid operations leave the abstract state unchanged. Warehouse keys, joins, normalization, nesting, caches, and ordering stay hidden. Relational, document, or graph implementations are substitutable only when shared black-box, generated-sequence, metamorphic, and representation-leakage tests produce equivalent histories, classifications, uncertainty states, and permitted errors.
The cross-domain transfer¶
The representation-independent interface archetype becomes a justice-trajectory contract. Its abstract state contains events, corrections, observation coverage, and policy versions; its public operations expose study meanings rather than storage details. The mapping is strong for bounded research datasets, but it cannot erase source-specific legal meaning or establish that underlying records and linkages are accurate.
Why it advanced¶
The candidate passed Experiment 6’s strict researched-candidate bar because documented longitudinal-data errors could alter study classifications, the required software primitives exist, and a synthetic cross-representation test is feasible. STRICT_SUCCESS remains a researched-candidate result, not proof of adoption, factual accuracy, universal applicability, novelty, or real-data performance.
Prior art and the remaining open claim¶
Common justice schemas, harmonized research tables, temporal databases, observation-period models, statistical plans, linkage systems, and immutable snapshots are adjacent prior art. The open claim concerns their narrower combination: one opaque contract with source attribution, non-erasing corrections, two kinds of time, explicit observation intervals, versioned policies, three-valued window answers, and a cross-representation oracle can preserve approved-study classifications better than direct warehouse queries, schema checks, and aggregate comparisons.
Smallest decisive test¶
Preregister 15–25 synthetic trajectories containing duplicates, split and consolidated episodes, late dispositions, corrections, observation gaps, overlapping coverage, timing conflicts, and policy changes. Two independent teams encode them in relational and document stores. Compare direct queries, schema and aggregate checks, a harmonized-table baseline, and the contract against fixed expected outputs and at least 12 seeded semantic mutants. Advance only with zero material classification or uncertainty divergences across valid implementations and at least 90% mutant detection, outperforming baseline checks. Redesign or reject it if inadequate observation becomes absence or passing implementations disagree materially.
Deployment and cost¶
A synthetic first study is estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms. Startup is estimated at \(50,000–\)250,000, operational launch at \(250,000–\)1 million, and annual operation at \(50,000–\)250,000. These are not vendor quotes; real-data use also requires study governance, privacy review, and authorization.
Risks and uncertainties¶
- The abstract event categories could erase legally or analytically important source distinctions.
- An equivalence rule could merge separate events or count one event more than once.
- Weak observation criteria could convert incomplete coverage into an apparent absence.
- Synthetic cases may omit irregular corrections, delayed entry, and overlapping-source behavior found in actual data.
- A passing conformance suite could be mistaken for evidence that source assertions or record linkages are correct and complete across all cases.
Expert review¶
Useful reviewer backgrounds: Criminology research methodologist, Administrative-data steward, Temporal data-modeling specialist, Software conformance-testing engineer, Privacy, ethics, and institutional-review specialist.
- Which source-specific distinctions must remain visible to preserve the approved study’s interpretation?
- What evidence is sufficient to call an observation interval adequate for a negative finding?
- Can independent implementers agree on event-equivalence and correction behavior before seeing test results?
- Which seeded mutants represent material study errors rather than harmless implementation differences?
- Would real pseudonymous records remain identifiable or require additional institutional and legal authorization?
Evidence and provenance¶
Selected sources: s1: Overview · s2: Criminal Justice Administrative Records System (CJARS): Data Documentation, 2023 Q3 · s3: Longitudinal linkage of administrative data: design principles and the total error framework · s4: NIEM 5.0 Justice Domain: j:Arrest · s5: OMOP Common Data Model v5.3: Observation Period · s6: Temporal tables · s7: Coded Private Information or Biospecimens Used in Research, Guidance (2018) · s8: Software Developers, Quality Assurance Analysts, and Testers
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index
Ordering note: The post-hoc harmonized ordering placed this candidate between ranks 8 and 22 across profiles, in band A. It is a reading-order aid using a cost-based pilot proxy, not an experimental endpoint, economic-value measure, or change to its STRICT_SUCCESS status.