Skip to content

Candidate dossiers — Band D

Part of Inverse Innovation with the Encyclopedia of Abstractions · Ranks 46–59: lower post-hoc review priority · Last revised August 2026

46. A Clinic for Cross-Field Lemma Handoffs

Canonical title: Rotating Proof-Interface Clinic for Cross-Subfield Lemma Handoffs
In one sentence: A rotating pair of specialists would turn eligible cross-subfield proof obstructions into precise, review-ready lemma packets without proving the lemma or changing the theorem.

Field Record
Portfolio ID EXP06-PARTNER-10
Experiment and endpoint Experiment 6 · Empirical-partner candidate
Archetype × domain Catalytic Pathway Enablement × Mathematics
Proposal position or arm P2
Post-hoc reading order Balanced score 61.0/100 · rank range 43–51 across three profiles · band D
First-evidence resource band \(10,000–\)50,000
Initial deployment startup band \(10,000–\)50,000

The problem

In a mathematics collaboration spanning subfields, a proof obligation may stall because participants use different definitions, notation, assumptions, or standards of explanation. Contributors repeatedly search for appropriate specialists, reconstruct terminology, and discover lost hypotheses only after work begins. Yet no mathematics-specific audit shows how often this translation and routing work, rather than genuinely missing mathematical insight, causes delay. Informal access to well-connected experts may also determine which obligations receive attention.

What is proposed

Create a governed clinic staffed by a rotating pair of specialists covering both sides of an interface. Intake must state the parent theorem, local definitions, requested conclusion, allowed assumptions, known dependencies, attempted approaches, and exact source of confusion. In a capped session, the specialists check compatibility, translate notation, preserve quantifiers and hypotheses, separate established implications from open mathematics, identify side conditions, and produce the smallest faithful lemma packet with an intended recipient or escalation path. An independent reviewer compares the packet with the original request. The clinic must reject disputes about truth, foundations, credit, theorem changes, or genuinely new mathematics. It tracks queue depth, labor, reopens, defects, conflicts, reviewer capacity, and specialist recovery, then rotates or rests staff instead of treating them as unlimited infrastructure.

The cross-domain transfer

The catalytic-pathway analogy is plausible but not literal. The clinic acts as a reusable facilitator that lowers recurring translation and routing barriers, releases each packet, and returns to readiness. It cannot change whether a lemma is true or replace the mathematical insight and scrutiny required to prove it.

Why it advanced

This candidate did not enter the strict-success lane. It cleared a separately calibrated empirical-partner lane because a small shadow comparison is practical and reversible. Advancement therefore means it merits a bounded external data-partner study, while the prevalence of the problem, available records, specialist willingness, and comparative benefit remain unverified.

Prior art and the remaining open claim

Focused mathematics programs, public question intake, shared glossaries, direct consultation, formal dependency blueprints, and large collaborative proof projects already address parts of the problem. The open claim concerns the governed combination: for archived obligations crossing the same two subfields, a capped rotating clinic with a versioned lemma-packet contract would beat an equal-access ad hoc handoff on preparation labor or time without adding statement defects, reopens, bad routing, or hidden downstream burden.

Smallest decisive test

Find one willing project and eight authorized, de-identified, resolved obligations from the same subfield pair, including translation successes, substantive proof problems, and an incompatible case. Compare the rotating clinic with an ad hoc team given equal source access and recorded labor; conceal archived resolutions until packets are frozen. Blinded reviewers score statement fidelity, definitions, hypotheses, quantifiers, uncertainty, routing, readiness, reopens, total labor, elapsed time, and downstream burden. Reject progression after a confidentiality breach, undisclosed material alteration, failure to reject incompatibility, mostly substantive escalations, increased defects or burden, or failure to meet the precommitted practical improvement threshold.

Deployment and cost

The first evidence study and initial startup are each estimated at \(10,000–\)50,000. Operational launch and annual recurring support are each \(50,000–\)250,000. These rough 2026 USD resource-equivalent bands are not quotes; scarce dual-subfield specialists, independent reviewers, compensation, and confidentiality controls could dominate actual cost.

Risks and uncertainties

  • Translation could silently remove a hypothesis, change a quantifier, or overstate equivalence between definitions.
  • Clinic packets or specialists could acquire informal authority despite having no power to accept mathematical claims.
  • Eligibility rules could favor contributors already fluent in preferred notation or connected to clinic staff.
  • Rotating specialists could lose continuity, become exhausted, or receive inadequate credit for substantive formulation work.
  • Confidential conjectures, correspondence, or attribution information could reach unintended recipients.

Expert review

Useful reviewer backgrounds: Mathematician from each participating subfield, Collaborative-project proof architect, Mathematical editor or independent proof reviewer, Research ethics and confidentiality specialist, Academic labor and attribution specialist.

  1. Can the project identify eight comparable resolved obligations whose reuse is authorized and whose resolutions can be concealed?
  2. What fraction of archived stalls arose from translation and routing rather than missing mathematical insight?
  3. Which changes to definitions, hypotheses, or quantifiers count as material fidelity defects?
  4. Can rotating specialists maintain continuity without exceeding fair workload and compensation limits?
  5. Does the clinic reduce total submitter, specialist, and downstream-review effort under equal access?

Evidence and provenance

Selected sources: S1: Investigating communication hindrance in interdisciplinary collaboration: A grounded theory approach · S2: Cultural barriers to interdisciplinary research collaboration: evidence from Australia · S3: SQuaREs: Structured Quartet Research Ensembles · S4: Collaborate@ICERM · S5: How do I ask a good question? · S6: LeanArchitect: Automating Blueprint Generation for Humans and AI · S7: The Equational Theories Project: Advancing Collaborative Mathematical Research at Scale · S8: National employment and wage data by occupation, May 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index

Ordering note: The harmonized assessment places this candidate 43–51 across profiles, in band D. That post-hoc ordering only guides reading and uses a cost-band affordability proxy; it is not an experimental endpoint, economic-value estimate, or substitute for the missing partner study.


47. Choosing One Proof Route Fairly

Canonical title: Bounded Proof-Verification Slot Challenge
In one sentence: A no-stakes challenge would test whether correctness-gated, resource-capped comparison selects a more maintainable route for one scarce formal-verification slot than simpler selection methods.

Field Record
Portfolio ID EXP06-STRICT-02
Experiment and endpoint Experiment 6 · Strict success
Archetype × domain Bounded Rivalry Governance × Mathematics
Proposal position or arm P1
Post-hoc reading order Balanced score 61.0/100 · rank range 38–55 across three profiles · band D
First-evidence resource band \(10,000–\)50,000
Initial deployment startup band \(50,000–\)250,000

The problem

A mathematics consortium may have several proposed routes to a theorem but enough specialist time to formalize only one. If the slot goes to the first route that appears complete, teams may benefit from premature completeness claims, hidden dependencies, elaborate presentation, or silence about fatal flaws. A strategically packaged route can therefore displace a sounder, more maintainable alternative, consuming scarce referee and formalizer time while discouraging useful sharing between teams.

What is proposed

Run a voluntary, time-limited challenge under rules frozen before judges see team identities. Every team submits the same proof certificate, listing contributors, axioms, imported results, dependencies, unresolved gaps, and known counterchecks. Independent reproduction is a pass-or-fail gate: presentation quality or expected cost cannot compensate for a failed mathematical step. Only passing routes are ranked on declared criteria such as dependency transparency, modularity, explanatory coverage, and estimated formalization burden. Give teams equal page limits, clarification rounds, and judge access. Permit disclosed collaboration and route merging, while prohibiting tampering, plagiarism, retaliation, private judge contact, sham independence, and concealment of a known fatal audit flaw. Provide a separate procedural appeal and reopen the slot if staged formalization later fails. Preserve attribution and access for every correctness-passing route.

The cross-domain transfer

The bounded-rivalry archetype maps directly to competition for one consortium-funded verification slot. A frozen rulebook, noncompensable correctness gate, equal reviewer-facing resource caps, auditable conduct rules, appeal, and reopening constrain strategic escalation while directing comparison toward downstream verification needs. Rivalry remains optional; collaboration may still prove better.

Why it advanced

This candidate passed Experiment 6's strict researched-candidate bar because the scarce-resource setting is coherent, adjacent infrastructure exists, authority can be bounded, and a reversible comparison can falsify the claim. It has not demonstrated better selection in practice and is not a novelty, deployment, or economic-impact finding.

Prior art and the remaining open claim

Proof certificates, machine checking, contribution rules, dependency graphs, task dashboards, formalization challenges, and large collaborative projects already exist. The unresolved contrast is whether their governance elements work better in this specific allocation decision: can correctness-gated ranking with equal access, frozen secondary criteria, appeal, and staged reopening select a route requiring fewer corrections or formalizer hours than blinded holistic triage or first reproduction, without excess review overhead or suppressed collaboration?

Smallest decisive test

With a consenting formalization organization, preregister a no-stakes shadow exercise using eight de-identified packets from a settled theorem: sound routes of differing modularity, known gaps, an undeclared dependency, a polished distraction, and controls. Compare the governed challenge with blinded holistic triage, while logging first independent reproduction as a third comparator. Auditors reproduce critical lemma chains and measure false passage, corrections, observed formalizer hours on a fixed sample, reviewer minutes, agreement, successful planted gaming, appeals, and collaboration effects. Do not progress if an incorrect packet passes, gaming determines selection, agreement misses its threshold, both comparators perform as well or better, or governance exceeds its budget.

Deployment and cost

The shadow study is estimated at \(10,000–\)50,000. Startup and operational launch are each \(50,000–\)250,000, with annual recurring costs of \(250,000–\)1 million. These are rough 2026 USD resource-equivalent bands, not budgets or quotes; expert reproduction and governance, rather than software, are the main uncertainties.

Risks and uncertainties

  • A visible ranking could turn proof development into a status contest and discourage useful collaboration.
  • Secondary scoring could reward a fashionable proof style rather than actual maintainability or mathematical value.
  • Page and clarification caps could disadvantage routes whose irreducible explanations are longer.
  • Dependency disclosure could expose unpublished work, while misconduct controls could stigmatize legitimate collaboration.
  • Judges could underestimate formalization effort or apply school-specific preferences inconsistently.

Expert review

Useful reviewer backgrounds: Research mathematician familiar with the theorem, Formal-proof engineer or maintainer, Independent mathematical referee, Research-governance and due-process specialist, Mathematical collaboration researcher.

  1. Does the target consortium actually face several viable routes competing for one indivisible verification slot?
  2. Can judges reproduce correctness and apply the secondary rubric with preregistered agreement?
  3. Do page and contact caps equalize reviewer access without shifting effort into hidden channels?
  4. Does challenge-selected work require fewer corrections and formalizer hours than both comparison methods?
  5. Would the format measurably discourage route sharing, merging, or disclosure of negative findings?

Evidence and provenance

Selected sources: S1: Contributing to mathlib · S2: Completion of the Liquid Tensor Experiment · S3: IMO Grand Challenge · S4: Exponentiating Mathematics (expMath) · S5: Formalising perfectoid spaces · S6: National employment and wage data by occupation, May 2025 · S7: Formalising Fermat · S8: The Equational Theories Project: Advancing Collaborative Mathematical Research at Scale
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index

Ordering note: The harmonized assessment ranks this candidate 38–55 across profiles and assigns band D. This post-hoc reading order is not the Experiment 6 endpoint and does not measure economic value; its affordability input also does not independently estimate elapsed pilot time.


48. Monitoring a Program’s Measurement Footprint

Canonical title: Efference-Residual Monitoring for Crime-Prevention Programs
In one sentence: A frozen model would subtract the aggregate records expected from a place-based prevention program so evaluators can inspect unexplained changes without producing individual scores or treating residuals as causal conclusions.

Field Record
Portfolio ID EXP06-PARTNER-14
Experiment and endpoint Experiment 6 · Empirical-partner candidate
Archetype × domain Predictive Residual Processing × Criminology Forensic
Proposal position or arm P4
Post-hoc reading order Balanced score 60.0/100 · rank range 40–53 across three profiles · band D
First-evidence resource band \(50,000–\)250,000
Initial deployment startup band \(250,000–\)1 million

Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-04, EXP06-STRICT-05, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.

The problem

A patrol, outreach, reporting, or other place-based prevention program can change both underlying events and how those events are recorded. More officer presence, for example, may predictably alter contacts, detected offenses, calls, complaints, or data completeness. Ordinary dashboards can blur this measurement footprint with external change, displacement, service withdrawal, or harm. Yet unexplained aggregate differences can also be overread as proof of program success, failure, misconduct, or community behavior.

What is proposed

Before each monitoring window, copy the authorized program schedule and intensity into a frozen, versioned model that predicts only the aggregate records and observation opportunities the program itself is expected to generate. Compare those expectations with separately retained observations and report typed differences: excess or missing activity, spatial or temporal displacement, source disagreement, complaint or injury changes, and unexplained missingness. Prioritize residuals by reliability, uncertainty, exposure, persistence, rights consequence, and evaluator capacity. Each alert must include its expected value, observed aggregate, uncertainty, provenance, program history, and reconstructible full-window context. Complaints, force, injury, deaths, disparity checks, whistleblower material, outages, integrity failures, and oversight requests always appear in full. Independent reviewers audit random and risk-selected windows. Residuals open evaluation inquiries only; shocks, small cells, mismatch, drift, or audit failures restore complete reporting.

The cross-domain transfer

The predictive-residual archetype is instantiated through an efference copy of the agency's own authorized schedule: expected program-generated records are separated from the unexplained remainder. The mapping is structurally strong for monitoring, but residuals cannot identify true crime levels, individual risk, causation, or the correct operational response.

Why it advanced

This candidate did not enter the strict-success lane. It cleared the separate empirical-partner lane because a retrospective aggregate comparison is bounded and measurable. Missing field evidence remains central: no agency has supplied schedules, linked sources, reviewers, legal approval, audit access, or evidence that the model separates measurement footprint from omitted context.

Prior art and the remaining open claim

Multi-indicator dashboards, pre/post comparisons, matched comparison areas, process and impact evaluations, displacement analysis, and generic anomaly detection already cover much of the substantive work. The narrower open claim is that an action-conditioned residual interface, with reconstructibility, independent raw-window audits, rights-critical bypasses, synchronization, and fallback, would reduce evaluator effort against full and conventional evaluation dashboards without materially missing displacement, reporting divergence, service withdrawal, or harm, or encouraging stronger unsupported causal claims.

Smallest decisive test

Pre-register a 12–16-week offline study using 150–250 historical or synthetic windows from one completed program and at least eight blinded evaluators. Compare the existing full dashboard, a conventional process-plus-impact dashboard, and the residual-first interface in balanced crossover order. Include outages, intensity changes, displacement, complaint or injury shifts, service withdrawal, benign shocks, source disagreement, and version mismatch. Require material-change recall within five percentage points of the better comparator and at least 20% lower median review time. Stop for any protected-signal omission or privacy breach, systematic audit misses, excessive reconstruction failure, increased causal overclaiming, failed fallback, or no reduction in total review-plus-maintenance effort.

Deployment and cost

The first evidence study is estimated at \(50,000–\)250,000. Startup is \(250,000–\)1 million; operational launch is \(1–\)5 million, with \(250,000–\)1 million recurring annually. These are rough 2026 USD resource-equivalent bands, not procurement quotes, and local data integration, privacy, oversight, and audit requirements remain unpriced.

Risks and uncertainties

  • The model could normalize repeated over-enforcement, under-service, or rights harm as expected program behavior.
  • Police-generated records could dominate independent sources and create a self-confirming baseline.
  • Evaluators or leaders could treat aggregate residuals as causal findings despite confounding and spillovers.
  • Schedule, boundary, reporting, or threshold changes could be used to make visible residuals disappear.
  • Sparse audits, uneven independent-source quality, or minimum cell sizes could hide rare harms or create geographic disparities in uncertainty.

Expert review

Useful reviewer backgrounds: Independent crime-program evaluator, Criminologist specializing in place-based interventions, Civil-rights and community-oversight representative, Government privacy and de-identification specialist, Agency data steward.

  1. Can a frozen schedule-conditioned model distinguish predictable recording changes from external events and omitted context?
  2. Which force, injury, complaint, disparity, missingness, and oversight signals must bypass filtering in full?
  3. Are independent data sources sufficiently complete and comparable across all evaluated places?
  4. Do residual displays increase unsupported causal interpretations relative to conventional dashboards?
  5. Does total evaluator, audit, calibration, and maintenance time fall while material-change recall remains within tolerance?

Evidence and provenance

Selected sources: S1: The Nation’s Two Crime Measures, 2015–2024 · S2: Measurement Error in Calls-For-Service as an Indicator of Crime · S3: Spatial Displacement and Diffusion of Benefits Among Geographically-Focused Policing Initiatives · S4: Assessing Responses to Problems: An Introductory Guide for Police Problem-Solvers · S5: An Ex Post Facto Evaluation Framework for Place-Based Police Interventions · S6: Smart Policing Initiative: Overview · S7: NIST SP 800-188: De-Identifying Government Datasets—Techniques and Governance · S8: Conduct of Law Enforcement Agencies
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index

Ordering note: The harmonized assessment places this candidate 40–53 across profiles, in band D. It is a post-hoc reading-order aid using an affordability proxy, not an experimental endpoint, economic-value measure, field-validation result, or evidence that an agency will adopt it.


49. Making Carbon-Budget Boundaries Visible

Canonical title: Open-Boundary Assembly for Regional Carbon-Budget Synthesis
In one sentence: At each regional carbon-budget release freeze, teams would jointly expose mismatched boundaries and unresolved residuals, then carry agreed disclosures into ordinary publication controls.

Field Record
Portfolio ID EXP06-STRICT-15
Experiment and endpoint Experiment 6 · Strict success
Archetype × domain Ritualized Meaning And Commitment Enactment × Environmental Climate
Proposal position or arm P4
Post-hoc reading order Balanced score 60.0/100 · rank range 45–53 across three profiles · band D
First-evidence resource band \(10,000–\)50,000
Initial deployment startup band \(10,000–\)50,000

Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-28, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.

The problem

Regional carbon budgets combine atmospheric, forest, soil, aquatic, and land-use estimates that may cover different places, periods, or definitions. Although specialists document limitations, those details can remain scattered in appendices and team files. A polished balance table can then look more complete than it is: exclusions disappear from headlines, residual quantities lack an owner or explanation, and contributors give different accounts of what the published total includes and leaves unresolved.

What is proposed

After normal technical reconciliation but before release approval, the consortium would hold a voluntary 50-minute Open-Boundary Assembly. Each component team would display a removable layer showing its spatial mask and time window, then state what its estimate includes, excludes, and cannot resolve. The response “heard, not resolved” would acknowledge the statement without implying agreement. A residual marker would pass among interface stewards, who could flag a mismatch, assign it for examination, declare it irreducible for that release, or pass. Contributors could accept, revise, transfer, or decline responsibility for carrying limitations into particular outputs. The editor—not the assembly—would then update the versioned boundary ledger, claims, graphics, metadata, owners, deadlines, and open questions through the existing publication system.

The cross-domain transfer

The ritualized-commitment archetype becomes a marked, repeatable scientific checkpoint. Visible layers symbolize that the regional total is assembled from partial views; standardized statements make limitations collectively audible; witnessed handoffs renew responsibility. The mapping is structurally strong, but its claimed benefit depends on ritual features adding something beyond equally careful facilitation and documentation.

Why it advanced

This candidate passed Experiment 6’s strict researched-candidate bar because it defines a bounded problem, preserves scientific authority, specifies a close comparator, provides falsifiers and safety controls, and proposes a measurable randomized pilot. STRICT_SUCCESS does not mean the assembly has worked in practice, is novel worldwide, is authorized for deployment, or produces economic value.

Prior art and the remaining open claim

The parts have substantial adjacent prior art: carbon accounting already uses QA/QC, completeness and double-counting checks, formal review, uncertainty tables, facilitated workshops, shared records, and version control. Workplace research also suggests rituals can increase perceived meaning, but not carbon-budget accuracy. The remaining claim is narrower: adding this voluntary ritual layer to an otherwise identical 50-minute boundary-review checklist improves detection, delayed recall, and fulfillment of disclosure duties without creating coercion, false assent, confidentiality failures, or authority confusion.

Smallest decisive test

Preregister a remote crossover study with 24–40 carbon-cycle or adjacent researchers in four to six balanced teams. Teams would review matched synthetic packets containing hidden boundary, stock-flow, covariance, exclusion, and residual defects, using either the assembly or an equal-duration facilitated checklist with the same facts and ledger. Advance only for at least a 20-percentage-point improvement in detection or delayed recall, or two substantively new correct disclosures per team, with no material loss of mock-output accuracy and no credible safety incident. Equivalent performance or any safety breach falsifies the incremental claim.

Deployment and cost

The first study and startup are each estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent terms. Operational launch is also \(10,000–\)50,000, while recurring annual operation is estimated at \(50,000–\)250,000. No consortium has committed authority, staff time, or funding; live evaluation would also require access to versioned synthesis artifacts.

Risks and uncertainties

  • Transparent layers may hide nonlinear processes, covariance, scale differences, or nonspatial boundaries.
  • “Heard, not resolved” may become rote or be mistaken for scientific agreement.
  • Marker circulation may pressure participants to explain quantities outside their expertise.
  • Publicly witnessed dissent could expose junior contributors or minority interpretations to retaliation.
  • The ceremony could make an unchanged or misleading synthesis appear unusually coherent and legitimate despite unresolved evidence gaps.

Expert review

Useful reviewer backgrounds: Regional carbon-budget scientist, Carbon-accounting and uncertainty specialist, Environmental synthesis editor or data steward, Research-team governance and facilitation specialist, Accessibility, consent, and confidentiality reviewer.

  1. Can existing release records establish that documented boundary limitations actually disappear from headline tables, graphics, or summaries?
  2. Would the planted defects and scoring rubric represent consequential regional-budget interface failures rather than merely easy checklist items?
  3. Can the checklist and assembly arms be matched for facts, facilitation quality, documentation, and time?
  4. What safety threshold should govern reports of pressure, false assent, confidentiality loss, or authority confusion?
  5. Which live or retrospectively versioned artifacts could measure whether disclosure obligations were ultimately fulfilled?

Evidence and provenance

Selected sources: S1: Global Carbon Budget 2024 · S2: RECCAP2—REgional Carbon Cycle Assessment and Processes 2: Overview · S3: 2006 IPCC Guidelines for National Greenhouse Gas Inventories, Volume 1, Chapter 6: Quality Assurance/Quality Control and Verification · S4: IPCC Procedures · S5: Developing Reproducible Workflows Collaboratively · S6: The Science and Practice of Team Science: Reimagining Collaboration in a Changing Research Landscape—Consensus Study Highlights · S7: Work Group Rituals Enhance the Meaning of Work · S8: Resources for Working Groups
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index

Ordering note: The harmonized review is only a post-hoc reading aid: scores range from 57 to 63 across profiles, with ranks 45–53 and ordering band D. Its speed input is an affordability proxy, not measured elapsed time or economic value, and it does not alter STRICT_SUCCESS status.


50. Reserve AI Capacity for Harm Inquiry

Canonical title: Protected nondeployment reserve in annual AI pilot portfolios
In one sentence: Organizations would protect at least 15% of annual AI-pilot capacity for nondeployment investigations before live-use projects consume the portfolio.

Field Record
Portfolio ID EXP04-STRICT-06
Experiment and endpoint Experiment 4 · Strict success
Archetype × domain Negative Space Design × Tech Ethics Ai Governance
Proposal position or arm RETRIEVAL_FIRST
Post-hoc reading order Balanced score 59.0/100 · rank range 49–50 across three profiles · band D
First-evidence resource band \(50,000–\)250,000
Initial deployment startup band \(50,000–\)250,000

The problem

When an organization allocates its annual AI-pilot portfolio, live-use proposals can consume all available capacity before individual approvals occur. Red-teaming, shadow evaluation, and affected-community inquiry must then compete for whatever remains. Material problems may consequently emerge only after people are exposed and projects have accumulated operational dependencies. The underlying prevalence is unknown, however, and existing risk-tiered assurance may already provide adequate predeployment coverage without a fixed reserve.

What is proposed

Before approving individual pilots, the portfolio governing body would reserve at least 15% of total annual pilot capacity exclusively for nondeployment harm inquiry. The protected share could fund red-teaming, shadow evaluation, or appropriately safeguarded affected-community work, but not live deployment. Separate accounting would prevent project teams from informally absorbing it. Reallocation would require the governing body to document that eligible inquiry demand was exhausted, obtain concurrence from an independent assurance function, record the destination and reasons, and preserve enough capacity to complete active inquiries. Mandatory privacy, security, accessibility, safety, and incident-response work would remain outside the contest for this reserve. Existing approval and risk-management processes would continue to control deployment decisions.

The cross-domain transfer

Negative-space design is instantiated as deliberately unused deployment capacity: a protected portfolio “void” is bounded before surrounding projects are selected. Its positive purpose is to preserve room for inquiry while changes remain feasible. The structural transfer is clear, although percentage ring-fencing and controlled release already exist in adjacent evaluation and innovation portfolios.

Why it advanced

This candidate passed Experiment 4’s strict researched-candidate bar by defining the denominator, protected use, release rule, comparator, measurable outcomes, and a records-only first step. STRICT_SUCCESS is limited to that research screen; it is not evidence that 15% is optimal, that harm detection improves, or that an organization will adopt the rule.

Prior art and the remaining open claim

Adjacent systems already allocate resources to AI testing, independent assurance, governance boards, sandboxes, recurring evaluations, and risk-tiered review. Other fields also ring-fence evaluation funding or divide portfolios by fixed percentages. The remaining contrastive claim is specifically that a preapproval 15% nondeployment reserve, protected from live-use commitments and released only with documented independent concurrence, increases independently adjudicated material issues found per proposed system without lowering the number of validated pilots by more than 10%.

Smallest decisive test

With one willing organization, preregister a records-only reconstruction of a completed annual portfolio. Compare the actual flexible or risk-tiered allocation with a 15% protected shadow allocation applied before proposal selection. Define capacity, eligible inquiry, materiality, and missing-data rules in advance; use two issue reviewers, including one independent of pilot selection. Measure material issues per proposed system, validated-pilot count, inquiry completion, and simulated compliance with release rules. Do not advance if issue detection does not increase, validated pilots fall by more than 10%, or capacity and eligible work cannot be measured consistently.

Deployment and cost

First evidence and initial startup are each estimated at \(50,000–\)250,000. Operational launch and annual recurring costs are each estimated at \(250,000–\)1 million, in rough 2026 resource-equivalent bands rather than vendor quotes. Actual cost depends heavily on specialist red teams, community participation, compute, data preparation, and displaced pilot capacity.

Risks and uncertainties

  • The 15% threshold may be too large, too small, or meaningless for very small portfolios.
  • Teams may relabel ordinary development or compliance work as protected inquiry.
  • Issue counts may reward numerous trivial findings unless independent reviewers apply a reproducible materiality standard.
  • A fixed reserve could sit unused while beneficial pilots wait, or encourage wasteful spending merely to exhaust it.
  • Observational results may confuse the reserve’s effect with an organization’s pre-existing safety culture and staffing quality.

Expert review

Useful reviewer backgrounds: AI portfolio-governance leader, Independent AI assurance or audit specialist, Causal-inference and program-evaluation researcher, Affected-community research and safeguarding specialist, Organizational finance and capacity-accounting expert.

  1. Can pilot capacity be expressed in a consistent unit across projects, staff, compute, and external assurance?
  2. Which work qualifies as nondeployment harm inquiry rather than normal development or mandatory compliance?
  3. Can historical records establish when an issue was found, whether it was material, and whether it changed design or approval?
  4. Is flexible risk-tiered assurance a more credible comparator than the organization’s actual historical allocation alone?
  5. How should continuous deployment, procurement, model updates, and small portfolios be handled outside the annual denominator?

Evidence and provenance

Selected sources: S1: AI RMF Core · S2: Industry temperature check: barriers and enablers to AI assurance · S3: OMB Memorandum M-25-21: Accelerating Federal Use of AI through Innovation, Governance, and Public Trust · S4: Costed Evaluation Plan Guidance, Tools and Templates · S5: Innovation portfolios for public sector organizations · S6: Anthropic’s Responsible Scaling Policy · S7: National Employment and Wage Data by Occupation, May 2025 · S8: Guidance to Set Up Your Organization's AI Governance Process
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index

Ordering note: The post-hoc harmonized reading aid scores this candidate 58–60 across profiles, ranks it 49–50, and places it in band D. The pilot-speed input reflects cost-band affordability, not observed duration. These figures neither measure economic value nor replace its STRICT_SUCCESS endpoint.


51. Audited Competition for Close-Review Priority

Canonical title: Verified Close-Readiness Arena for Scarce Consolidation Review Slots
In one sentence: Business units would compete for scarce early financial-close review slots using independently verified readiness evidence rather than self-declared completion or managerial escalation.

Field Record
Portfolio ID EXP06-PARTNER-01
Experiment and endpoint Experiment 6 · Empirical-partner candidate
Archetype × domain Bounded Rivalry Governance × Accounting Auditing
Proposal position or arm P1
Post-hoc reading order Balanced score 59.0/100 · rank range 47–54 across three profiles · band D
First-evidence resource band \(10,000–\)50,000
Initial deployment startup band \(50,000–\)250,000

Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-02, EXP06-PARTNER-03, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.

The problem

During a multi-entity financial close, business units may compete for a limited number of early consolidation or technical-accounting reviews. If self-certified readiness or managerial escalation controls the queue, a unit can gain priority by prematurely closing reconciliations, shifting exceptions to another entity, deferring unsupported items, or obtaining privileged reviewer access. Reviewers then reopen supposedly complete work while other units wait. Crucially, no external evidence yet shows that this scarcity or strategic behavior exists in a target organization.

What is proposed

The corporate controller would replace the informal queue race with a recurring, bounded readiness contest. Units meeting a minimum control floor could compete for several early review slots. A rulebook fixed before scoring would allow genuine documentation, automation, supported resolution, and timely escalation, while prohibiting concealment, unsupported deferral, exception dumping, shared answers, off-channel influence, and undisclosed borrowed labor. An independent verifier would score hidden, risk-stratified samples for first-pass evidentiary completeness, supported treatment of aged exceptions, intercompany agreement, and absence of later unsupported corrections. Timely disclosure of material issues would have a protected route. Winners would be audited, serious errors appealable, extra contest labor capped, and some awarded capacity retained for remediation. Priority would expire after one close.

The cross-domain transfer

Bounded-rivalry governance becomes a formal contest for a scarce operational prize. Eligibility, lawful tactics, fouls, judging, appeals, resource caps, multiple awards, spillover responsibility, and recurring challenger access constrain how units compete. The mapping is detailed, but it remains unproven that queue positions create meaningful strategic interdependence rather than reflecting ordinary systems, complexity, or staffing constraints.

Why it advanced

This candidate did not enter Experiment 6’s strict-success lane. It cleared the separately calibrated EMPIRICAL_PARTNER_CANDIDATE lane because a bounded retrospective partner study is feasible and decision-relevant. Its advancement depends on obtaining proprietary field records; there is currently no evidence of local queue scarcity, manipulation, predictive advantage, or safe behavioral response.

Prior art and the remaining open claim

Financial-close platforms already provide workflows, dependencies, approvals, dashboards, audit trails, exception monitoring, and risk-based administrative triage. Accounting standards also require evidence, objective verification, and controls over period-end reporting. The narrower open claim, conditional on consequential rivalry being demonstrated, is that hidden-sample, independently verified competitive ranking predicts less first-pass rework, fewer reopened or transferred exceptions, and fewer unsupported corrections than both current queueing and noncompetitive risk-based triage, without delaying protected disclosure of material issues.

Smallest decisive test

Preregister a retrospective replay of one completed close across four to eight entities and 40–80 risk-stratified reconciliations. Compare recorded queue order, blinded noncompetitive risk-based triage, and the proposed arena score. Measure verified completeness, rework hours, reopened exceptions, transferred mismatches, unsupported correcting entries, and disclosure timing. Report stability under bootstrap samples and reasonable weight changes. Do not proceed beyond a no-consequence simulation if the arena fails to outperform triage out of sample, Kendall rank stability is below 0.60, size or complexity drives results, data are inconsistent, or issue reporting appears delayed or suppressed.

Deployment and cost

A retrospective first study is estimated at \(10,000–\)50,000. Initial startup is \(50,000–\)250,000; operational launch and annual recurring operation are each \(250,000–\)1 million in rough 2026 resource-equivalent terms. Software, data integration, independent verification, appeals, labor-cap auditing, and recurring governance still require organization-specific estimates, not vendor assumptions.

Risks and uncertainties

  • Teams may game evidence fields that hidden samples do not cover.
  • Risk adjustment may embed incumbent complexity assumptions and disadvantage legitimate challengers.
  • Labor caps may be evaded through off-books work or restrict units with genuine remediation needs.
  • Coordination screens may falsely flag common deadlines or shared system failures as collusion.
  • Winning units may accumulate reviewer relationships and procedural knowledge even though formal priority expires.

Expert review

Useful reviewer backgrounds: Corporate controller or consolidation leader, Internal-control and ICFR specialist, Internal auditor independent of close operations, Financial-close data and workflow engineer, Employment, legal, and external-audit governance adviser.

  1. Are early specialist-review slots genuinely scarce, and does their timing materially affect other entities’ outcomes?
  2. Can queue, reconciliation, exception, rework, escalation, and correction records be joined reliably across entities?
  3. Does the proposed score outperform administrative risk triage on data not used to construct the ranking?
  4. How can protected material-issue reporting be separated from competitive scoring and monitored for delay?
  5. Would labor caps, hidden sampling, and anomaly screening conflict with employment rules, ICFR responsibilities, or external-audit arrangements?

Evidence and provenance

Selected sources: S1: Can fine-tuning your financial processes help accelerate your growth? · S2: Stepping into the future of controllership · S3: AS 2201: An Audit of Internal Control Over Financial Reporting That Is Integrated with an Audit of Financial Statements · S4: Financial Close Management Software · S5: Approve Closing Tasks · S6: Executive tournament incentives and audit fees · S7: Commission Guidance Regarding Management’s Report on Internal Control Over Financial Reporting Under Section 13(a) or 15(d) of the Securities Exchange Act of 1934 · S8: Occupational Employment and Wages — May 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index

Ordering note: The post-hoc harmonized reading aid gives scores of 56–62, ranks 47–54, and band D. Its speed input is a cost-affordability proxy rather than elapsed time. This ordering is not an experimental endpoint or economic-value estimate and does not upgrade the partner-candidate status.


52. Renewing Climate-Monitoring Responsibilities

Canonical title: Season-Turn Stewardship Muster for Long-Term Climate Monitoring
In one sentence: Before each field season, monitoring staff would visibly renew or transfer station-to-archive duties and then record every accepted obligation in the ordinary task system.

Field Record
Portfolio ID EXP06-PARTNER-28
Experiment and endpoint Experiment 6 · Empirical-partner candidate
Archetype × domain Ritualized Meaning And Commitment Enactment × Environmental Climate
Proposal position or arm P1
Post-hoc reading order Balanced score 59.0/100 · rank range 44–58 across three profiles · band D
First-evidence resource band under $10,000
Initial deployment startup band \(10,000–\)50,000

Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-15, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.

The problem

Long-term climate monitoring depends on seasonal sampling, calibration, metadata, custody transfers, and reliable handoffs despite staff turnover. Written protocols may leave particular responsibilities privately understood or attached to departed personnel. A field season can therefore begin without a named primary, backup, needed resource, credential, or escalation route for every station-to-archive link. Missing duties or tacit knowledge may surface only when work is due, creating gaps or ambiguities in a record intended to remain comparable over decades.

What is proposed

After the normal technical readiness review and before each field season, the network would hold a voluntary 35-minute Season-Turn Stewardship Muster. A neutral signal would mark the occasion, and participants would trace a fictional or real observation from station to archive by moving plain link cards across a network map. At each step, the current steward could renew, revise, transfer, or decline responsibility without explaining publicly. A backup and missing resources would be named before acceptance. Witnesses could acknowledge completed handoffs, but attendance, speech, silence, or card handling would not create consent or authority. Closure would occur only when ordinary managers enter an owner, backup, resources, due date, and escalation route into the existing task system. Debrief and independent review could modify, pause, or retire the practice.

The cross-domain transfer

The ritual archetype becomes a marked seasonal enactment of shared stewardship. Card movement makes the custody chain visible; voluntary renewal prevents stale assignments from persisting silently; witnessed handoffs support memory across turnover. The mapping is plausible, but most operational content duplicates ordinary responsibility maps and structured handoffs, leaving only the ritual features’ incremental social-memory effect to test.

Why it advanced

This proposal did not enter the strict-success lane. It cleared the EMPIRICAL_PARTNER_CANDIDATE lane because a small, fictional-record crossover with an external monitoring partner could test the remaining contrast safely. Field evidence is missing on problem prevalence, adopter demand, improved recall, dependency discovery, operational completion, voluntariness, and accessibility.

Prior art and the remaining open claim

Climate networks already require sustained operations, calibration, metadata, annual maintenance, anomaly tracking, preseason reviews, assigned roles, and configuration control. Structured handoffs also cover ownership, acknowledgment, next steps, and clarification; workplace ritual studies measured meaning, not monitoring continuity. The remaining claim is that adding a neutral threshold, voluntary card traversal, and witnessed renewal to an equal-duration administrative session improves seven-day recall or surfaces more valid dependencies without increasing pressure, exclusion, religious conflict, or confusion about formal authority.

Smallest decisive test

With one willing network, preregister a crossover using four matched fictional station-to-archive records and 12–24 participants. Compare a 35-minute administrative readiness session with the same session plus the ritual features, crossing teams onto a different record. Measure valid actionable dependencies and blinded immediate and seven-day recall of each primary, backup, and escalation route. Falsify incremental promise if the ritual finds no additional valid dependency and improves complete-link recall by less than 15 percentage points, or worsens pressure or access ratings by at least 0.5 on a five-point scale. Any credible coercion, privacy, cultural, or authority incident requires a halt.

Deployment and cost

First evidence is estimated below $10,000. Initial startup, operational launch, and annual recurring operation are each estimated at \(10,000–\)50,000 in rough 2026 resource-equivalent bands. Actual costs remain uncertain because network size, travel, participant count, union requirements, accessibility work, facilitation, and task-system integration have not been specified.

Risks and uncertainties

  • Employment hierarchy may make a formally optional refusal feel professionally costly.
  • The event may aestheticize stewardship while staffing, equipment, training, or travel shortages remain unfunded.
  • Public handoffs may expose disability, location, performance, or employment information.
  • The linear card metaphor may distort parallel, contested, or nonlinear observation pathways.
  • Neutral-looking symbolism may still create religious or cultural conflict, especially if local designers add unauthorized ceremonial elements.

Expert review

Useful reviewer backgrounds: Climate-monitoring network operator, Field-to-archive data and metadata steward, Human-factors and structured-handoff researcher, Workplace accessibility and religious-accommodation specialist, Program evaluator experienced in crossover trials.

  1. Do ordinary preseason records actually contain missing owners, backups, resources, credentials, or escalation routes?
  2. Are the fictional station-to-archive cases realistic enough to test consequential dependencies without exposing operational data?
  3. Can the administrative comparator match all content, time, facilitation, and task closure except the ritual features?
  4. How will anonymous measures detect pressure when managers and subordinates participate together?
  5. What result would justify a prospective no-consequence simulation before any real responsibility transfer is considered?

Evidence and provenance

Selected sources: S1: GCOS Surface Reference Network (GSRN): Justification, Requirements, Siting and Instrumentation Options (GCOS-226) · S2: U.S. Climate Reference Network (USCRN) · S3: U.S. Climate Reference Network Data Management Plan · S4: Fire Ecology Monitoring Protocol for the Heartland Inventory and Monitoring Network · S5: Handoff · S6: Work Group Rituals Enhance the Meaning of Work · S7: Section 12: Religious Discrimination · S8: Employer Costs for Employee Compensation—December 2025
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index

Ordering note: The post-hoc harmonized reading aid scores this candidate 55–63, with ranks 44–58 and band D. Its speed measure is a cost-affordability proxy, not observed duration. The wide profile range is not economic value evidence and leaves the EMPIRICAL_PARTNER_CANDIDATE endpoint unchanged.


53. Recurring Stewardship for Digital Legacies

Canonical title: Digital Memory Stewardship Observance for Algorithmic Afterlives
In one sentence: A governed observance would periodically turn remembrance into authorized, verifiable decisions about access, visibility, retention, resurfacing, correction, and successor responsibility for a digital legacy.

Field Record
Portfolio ID EXP06-PARTNER-29
Experiment and endpoint Experiment 6 · Empirical-partner candidate
Archetype × domain Ritualized Meaning And Commitment Enactment × Human Computer Interaction
Proposal position or arm P3
Post-hoc reading order Balanced score 59.0/100 · rank range 41–57 across three profiles · band D
First-evidence resource band \(50,000–\)250,000
Initial deployment startup band \(250,000–\)1 million

The problem

After death or lasting incapacity, a person’s posts, messages, images, inferred attributes, and scheduled reminders may persist under fragmented platform settings. Their recorded wishes may be incomplete, appointed account stewards may leave, and relatives or correspondents may disagree about what should remain visible. A one-time memorialization decision cannot necessarily address later platform changes, automated resurfacing, new audiences, disputed material, or the transfer of practical responsibility to a new steward.

What is proposed

At memorial creation, a chosen recurring date, a steward handoff, or a material platform change, authorized participants would hold a Digital Memory Stewardship Observance. A prior review would determine who may participate, submit private input, withhold material involving them, or decline. During the session, recommendations, autoplay, metrics, and nonessential notifications would be paused where technically possible. Participants would review consented, provenance-labeled materials without requiring shared grief or one approved life story. Authorized stewards would then renew, revise, decline, or transfer duties covering access, visibility, retention, resurfacing, moderation, and correction. Closure would create authenticated platform requests, named owners, deadlines, access changes, and an unresolved-matters register, followed by debrief, audit, repair, and retirement options.

The cross-domain transfer

The ritualized meaning-and-commitment archetype is transferred through a marked pause in ordinary algorithmic circulation, a recognizable sequence, consented witnessing, and recurring renewal of duties. The mapping is structurally strong at the governance level, but the symbolic elements have not been shown to improve stewardship beyond a well-run checklist using identical platform controls.

Why it advanced

This EMPIRICAL_PARTNER_CANDIDATE cleared a separately calibrated lane for a bounded external-partner study, not the strict-success lane. It advanced because the lifecycle problem, platform authorities, adjacent practices, safety controls, and falsifiable comparison are identifiable. Crucially, there is no field evidence that bereaved people would find the observance acceptable, safe, accessible, or useful.

Prior art and the remaining open claim

Adjacent prior art includes platform memorialization, legacy contacts, inactivity plans, estate administration, grief gatherings, and collaboratively edited memorials. These already cover many individual components. The narrower open claim is that adding a recurring, consent-aware observance—with an algorithmic pause, plural witnessing, explicit duty renewal, and later audit—to the same controls will produce more correctly authorized and verified actions and reveal more unresolved conflicts than an equal-time static directive-and-action checklist, without causing more privacy breaches, coercion, or distress.

Smallest decisive test

A digital-legacy professional body or university HCI laboratory would recruit 12–24 living-volunteer dyads using synthetic or participant-controlled redacted account replicas. Dyads would receive either the complete observance or an equal-time checklist with identical mock controls and authority briefing. Independent reviewers would assess correct authority assignment, verified action completion, detected conflicts, unauthorized access or disclosure, handoff success, distress, coercion, unwanted exposure, accessibility failures, and facilitator halts. Proceed only if the observance materially improves verified correct actions or conflict detection without worse safety or completion outcomes; otherwise adapt, hold, or retire it. This would not demonstrate benefit during bereavement.

Deployment and cost

A sandboxed rehearsal and manual action ledger appear feasible; production use depends on platform-specific authentication, recommendation controls, verification, rollback, law, and trained facilitation. Rough 2026 USD resource-equivalent bands are \(50,000–\)250,000 for first evidence, \(250,000–\)1 million for initial startup, \(1–\)5 million for operational launch, and \(250,000–\)1 million annually. These are not vendor quotes.

Risks and uncertainties

  • A powerful family or community faction could turn the memorial into one authorized-looking account of the person’s life.
  • Private messages, images, or relationship patterns could be exposed to people who lack standing to see them.
  • Survivors could mistake ceremonial agreement for authority over recorded directives or another living person’s data.
  • Visible settings may not match the platform’s actual recommendation and resurfacing behavior.
  • Recurring dates or notifications could impose remembrance on people who want to disengage permanently or temporarily without explanation.

Expert review

Useful reviewer backgrounds: Digital-legacy and estate-planning specialist, Human-computer interaction researcher specializing in death and bereavement, Privacy and fiduciary-access lawyer, Platform trust, safety, and account-support engineer, Grief-informed facilitator or clinical safety reviewer.

  1. Which proposed decisions belong respectively to the original account holder, a fiduciary, a depicted or corresponding person, and the platform?
  2. Can the mock platform reliably verify resurfacing behavior and reverse every tested setting or access change?
  3. Which symbolic elements add information or accountability that the equal-time checklist does not?
  4. What stopping thresholds for distress, coercion, unwanted exposure, and disputed authority should block continuation?
  5. Can the study recruit contested or culturally varied dyads without treating participation as consent to memorialization?

Evidence and provenance

Selected sources: s1: About Memorialized Accounts · s2: How to add a Legacy Contact for your Apple Account · s3: About Inactive Account Manager · s4: Current Acts—F: Fiduciary Access to Digital Assets Act, Revised · s5: Digital Legacy Association—Home · s6: Engaging with Death Online: An Analysis of Systems that Support Legacy-Making, Bereavement, and Remembrance · s7: Digital Legacy: A Systematic Literature Review · s8: From Personal Data to Digital Legacy: Exploring Conflicts in the Sharing, Security and Privacy of Post-mortem Data
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index

Ordering note: The harmonized review is only a post-hoc reading order, not an experimental endpoint or measure of economic value. It placed this candidate between ranks 41 and 57 in band D; the cost input approximated affordability, not elapsed pilot time.


54. Testing Audit Findings Before Awarding Leads

Canonical title: Replicable Assurance Challenge for Scarce Audit-Lead Mandates
In one sentence: A separate internal-audit challenge would award temporary lead roles and investigative hours according to independently reproduced risk findings rather than visible finding volume or severity alone.

Field Record
Portfolio ID EXP06-PARTNER-02
Experiment and endpoint Experiment 6 · Empirical-partner candidate
Archetype × domain Bounded Rivalry Governance × Accounting Auditing
Proposal position or arm P2
Post-hoc reading order Balanced score 58.0/100 · rank range 52–54 across three profiles · band D
First-evidence resource band \(10,000–\)50,000
Initial deployment startup band \(50,000–\)250,000

Lineage note: This record shares its archetype–domain cell with EXP06-PARTNER-01, EXP06-PARTNER-03, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.

The problem

Internal-audit teams may compete for a small number of prominent engagement-lead roles and discretionary investigative hours. If managers informally reward finding counts, apparent severity, speed, or budget performance, auditors could gain an advantage by splitting one root problem into several findings, favoring easily scored issues, delaying referrals, withholding reusable tests, or pressing auditees toward harsher labels. The result could be a leadership pipeline that rewards persuasive issue production rather than reproducible assurance work.

What is proposed

Create a quarterly Replicable Assurance Challenge that remains separate from mandatory reporting and personnel evaluation. Teams would submit evidence from completed work to compete for two temporary lead mandates and capped investigative hours. A frozen rulebook would protect urgent, legal, fraud, whistleblower, and material-misstatement reporting while prohibiting duplicate counting, evidence withholding, retaliation, reciprocal scoring, unsupported severity changes, auditee pressure, and hidden labor. An identity-blinded panel would test leading hypotheses on held-out transactions and score reproducibility, severity support, root-cause coherence, added risk coverage, audit-trail quality, portability, and auditee burden. Awards would expire after one quarter; some hours would remain centrally held for correction. Appeals, coordination screens, later validation, and retirement rules would constrain gaming and entrenchment.

The cross-domain transfer

Bounded rivalry is instantiated as a deliberately separate contest for scarce, temporary mandates. Legitimate competition occurs through reproducible evidence under frozen rules; mandatory assurance communication remains outside the arena. Independent re-performance, labor caps, anti-collusion screens, burden scoring, expiring prizes, and later review are intended to keep rivalry tied to assurance quality.

Why it advanced

This EMPIRICAL_PARTNER_CANDIDATE entered the bounded data-partner lane, not strict success. It offers a testable shadow study using existing audit records and established quality-review authority. However, no inspected records show that lead roles are truly scarce, that issue production affects selection, or that the proposed strategic behavior occurs in operating audit functions.

Prior art and the remaining open claim

Quality-assurance programs, ordinary engagement review, root-cause analysis, consolidated issue tracking, rotations, and risk-adjusted scorecards already address much of the problem. The open claim is conditional: only where genuine scarcity and strategic interdependence exist, blinded re-performance should predict later finding validity and incremental risk coverage better than ordinary quality review or a simpler scorecard, without delaying reporting, increasing defensive documentation, exposing confidential material, or adding auditee burden. This is not a world-novelty claim.

Smallest decisive test

Preregister a retrospective shadow replay of six closed engagements, twelve issued or merged findings, and six documented but unissued hypotheses. Three independent quality reviewers would use preserved records and held-out transactions without contacting auditees, publishing ranks, or allocating mandates. Compare the proposed composite with existing QAIP outcomes and a simpler risk-adjusted scorecard. Measure inter-rater reliability, rank stability, reviewer hours, blinding, later finding status, added coverage, reconstructed burden, and strategic markers. Do not advance if reliability is below 0.70, the method adds no discrimination, engagement identity or writing style dominates, confidentiality or blinding fails, or prospective use could inhibit candid reporting.

Deployment and cost

Existing audit repositories could support a retrospective replay, but live use would require reliable blinding, risk normalization, protected workpaper access, independent reviewers, and safeguards against informal career use. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(250,000–\)1 million annually; they are not quotes.

Risks and uncertainties

  • Competition could compromise, or appear to compromise, auditor independence.
  • A reproducibility metric could disadvantage emerging or systemic risks that cannot yet be repeated across transactions or locations.
  • Teams could optimize documentation for the review panel while moving preparation labor off the recorded budget.
  • Audit subjects and methods may reveal team identity despite formal blinding.
  • Private standings could still affect promotion or assignment decisions if confidentiality fails.

Expert review

Useful reviewer backgrounds: Chief audit executive, Independent internal-audit quality-assurance reviewer, Audit-committee or board governance specialist, Employment, privilege, and workpaper-access counsel, Audit-methodology and measurement researcher.

  1. Do historical assignment records show scarce lead opportunities and a relationship between visible issue production and selection?
  2. Can reviewers reconstruct nonissued hypotheses and held-out tests from authorized records without new auditee requests?
  3. Does the composite remain reliable after controlling for engagement risk, scope, specialization, and workpaper style?
  4. Would auditors alter urgent reporting or create defensive documentation if a live challenge were introduced?
  5. Can shadow results be technically and institutionally prevented from entering personnel decisions?

Evidence and provenance

Selected sources: S1: Global Internal Audit Standards, 2024 Edition · S2: Quality Assurance and Improvement Program (QAIP) · S3: Risk: The Root of the Matter · S4: Incentives for Dishonesty: An Experimental Study with Internal Auditors · S5: Is the Objectivity of Internal Audit Compromised When the Internal Audit Function Is a Management Training Ground? · S6: TeamMate Audit Management · S7: Connected Risk Quick Start Guide for Internal Audit and Controls Leaders · S8: Accountants and Auditors: Occupational Outlook Handbook
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index

Ordering note: The harmonized review is a post-hoc ordering aid, not an endpoint or economic-value estimate. This candidate fell between ranks 52 and 54 in band D. Its cost-based affordability proxy did not independently measure how quickly a pilot could run.


55. Full-Cost Bidding for Preservation Access

Canonical title: Externality-Adjusted Reverse Auction for Religious-Collection Preservation
In one sentence: Custodians would submit sealed subsidy bids for comparable preservation work, with lifecycle costs, cultural authority, material safety, and future technical dependence incorporated before scarce laboratory time is awarded.

Field Record
Portfolio ID EXP06-PARTNER-08
Experiment and endpoint Experiment 6 · Empirical-partner candidate
Archetype × domain Bounded Rivalry Governance × Religious Studies Theology
Proposal position or arm P4
Post-hoc reading order Balanced score 58.0/100 · rank range 47–55 across three profiles · band D
First-evidence resource band \(10,000–\)50,000
Initial deployment startup band \(50,000–\)250,000

The problem

A library consortium may have fewer mobile-laboratory weeks and less subsidy than custodians request for fragile manuscripts, recordings, ritual objects, and community archives. In a narrative competition, applicants could understate preparation, rights, metadata, storage, migration, or remediation costs while exaggerating quantity or urgency. Projects that appear inexpensive may therefore win by shifting work and risk to custodians, communities, or future repositories, while vulnerable materials wait and selected providers gain technical leverage over later rounds.

What is proposed

Divide annual laboratory capacity into narrow lots defined by material type, condition, and verified work units. Before bidding, independent reviewers would confirm custodial authority, cultural permissions, conservation safeguards, metadata, access restrictions, storage commitments, and technical feasibility. Eligible custodians would submit sealed reverse bids stating the minimum subsidy required for preparation, treatment, digitization, return copies, metadata, and a defined preservation period. A fixed schedule would add reserves for handling, migration, dependencies, and remediation, without penalizing culturally required restrictions. The lowest adjusted eligible bids would clear, subject to institutional concentration and shared-failure checks. Finalists would be audited; sponsors would secure remediation obligations; outputs would use transferable formats under custodian-controlled access; and later cost, damage, access, and concentration results would recalibrate future lots.

The cross-domain transfer

Bounded rivalry becomes a reverse auction for scarce preservation capacity. Price competition is permitted only within pre-authorized, technically comparable lots and above noncontestable cultural and safety floors. Sealed bids, fixed adjustments, audits, guarantees, concentration limits, transferable workflows, rebidding windows, and post-project review seek to prevent cost dumping, coordination, and lock-in.

Why it advanced

This EMPIRICAL_PARTNER_CANDIDATE cleared a bounded external data-study lane rather than strict success. Active preservation programs, technical standards, and an authorized retrospective design make the question researchable. The central missing evidence is institutional: no partner has supplied project records, approved even a simulation, or shown that stable comparable lots can be constructed.

Prior art and the remaining open claim

Expert grant panels, first-come allocation, collection-size formulas, eligibility lotteries, centralized preservation, and established reverse-auction procedures supply adjacent prior art. The narrower open claim is that, for genuinely comparable and pre-authorized projects, externality-adjusted subsidy bids will predict realized lifecycle cost and allocate fixed laboratory capacity more cost-effectively than those alternatives without increasing cultural exclusion, physical harm, change orders, technical concentration, or stranded outputs. The claim fails if discretionary adjustments overwhelm price or protected collections cannot fit comparable lots.

Smallest decisive test

With custodial and records-owner authorization, reconstruct proposed and realized lifecycle costs for at most ten completed projects in no more than three material-and-treatment categories. Two independent conservation-cost reviewers would apply preregistered fields, comparability thresholds, missing-data rules, and harm measures. Test whether adjusted ordering predicts realized cost better than requested budgets and actual narrative rankings, then simulate narrative, first-come, size-formula, and lottery allocations. Permit only a go/no-go decision for a nonbinding sealed-bid simulation. Stop if reviewer agreement is poor, over 25% of projects are incomparable, discretionary adjustments dominate, prediction does not materially improve, protected or under-resourced custodians are disproportionately excluded, or simulated safety or concentration worsens.

Deployment and cost

The administrative auction machinery is feasible only for narrow categories with stable units; cultural authority, rights, condition, insurance, storage, and legal characterization remain local. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(50,000–\)250,000 for startup, \(250,000–\)1 million for operational launch, and \(250,000–\)1 million annually. They are not vendor quotes.

Risks and uncertainties

  • Standard work units could conceal collection-specific fragility, urgency, or treatment needs.
  • Institutions with already subsidized infrastructure could bid below smaller community archives without being more efficient overall.
  • Guarantee requirements could exclude under-resourced custodians even if pooled guarantees are nominally available.
  • An adjustment schedule could wrongly treat culturally required access restrictions as costs or inefficiencies.
  • Transferability requirements could conflict with community authority over restricted knowledge or materials.

Expert review

Useful reviewer backgrounds: Conservator experienced with the selected material categories, Community archive or religious-collection custodian, Cultural authority and Indigenous data-governance specialist, Preservation economist or auction-design researcher, Copyright, privacy, cultural-property, and procurement counsel.

  1. Which material-and-treatment categories can be normalized without obscuring condition, cultural protocol, or urgency?
  2. Can proposed and realized preparation, metadata, storage, migration, and remediation costs be reconstructed consistently from closed-project records?
  3. How should the design prevent infrastructure-rich institutions from appearing artificially inexpensive?
  4. Which cultural or custodial restrictions must remain outside every price and externality calculation?
  5. Would the process legally constitute grantmaking, procurement, or another allocation form in the partner’s jurisdiction?

Evidence and provenance

Selected sources: S1: Towards Sustainable Preservation and Accessibility of Documentary Heritage · S2: Preserving Endangered Cultural Memory at a Time of Heightened Risk: Evaluating the Recordings at Risk Grant Program · S3: Recordings at Risk · S4: Collections Stewardship · S5: Preservation and Selection for Digitization · S6: Technical Guidelines for Digitizing Cultural Heritage Materials · S7: The CARE Principles for Indigenous Data Governance · S8: Federal Acquisition Regulation Subpart 17.8—Reverse Auctions
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index

Ordering note: The harmonized assessment is only a post-hoc reading order, not an experimental endpoint or economic-value measure. It placed this candidate between ranks 47 and 55 in band D. The pilot input represented cost-band affordability rather than independently assessed elapsed time.


56. Reviewing What Changed in Student Reasoning

Canonical title: Predictive-Residual Formative Review for Interpretive Coursework
In one sentence: A shadow review system would predict a student’s next rubric-level performance, foreground the differences from that prediction, and preserve full human access, auditing, and fallback for every consequential interpretation.

Field Record
Portfolio ID EXP06-STRICT-07
Experiment and endpoint Experiment 6 · Strict success
Archetype × domain Predictive Residual Processing × Religious Studies Theology
Proposal position or arm P4
Post-hoc reading order Balanced score 58.0/100 · rank range 54–56 across three profiles · band D
First-evidence resource band \(10,000–\)50,000
Initial deployment startup band \(250,000–\)1 million

Lineage note: This record shares its archetype–domain cell with EXP06-STRICT-06, but preserves a distinct proposal or experimental arm. Treat them as related candidate instances, not independent cell-level evidence.

The problem

Instructors reviewing repeated low-stakes religious-studies assignments must keep checking familiar skills while looking for new misconceptions, unsupported comparisons, and unexpectedly strong reasoning. Complete review preserves context but may consume feedback time confirming repeated mastery. Predictive filtering could focus attention, yet it could also lock students into earlier profiles, treat one tradition or writing style as normal, miss defensible alternative readings, or confuse personal belief with academic performance. The challenge is to save attention without weakening plural, contestable human judgment.

What is proposed

For one course and one recurring assignment family, a versioned model would use only a consenting student’s earlier work to predict rubric-level features in the next response before it is opened. It would predict taught academic skills—not wording, belief, doctrinal correctness, or grades. Human coding of the untouched submission would then identify signed differences such as improvement, omitted evidence, repeated conflation, changed qualification, unsupported transfer, a defensible alternative, or coding disagreement. A residual-first interface could collapse expected components, but the matching model and residuals must reconstruct the whole rubric profile, with the full submission one action away. Humans would validate every residual. Independent full marking, student challenges, version checks, error budgets, and mandatory bypasses for high-stakes, personal, accessibility-sensitive, culturally disputed, unfamiliar, or low-confidence work would trigger complete review when needed.

The cross-domain transfer

Predictive-residual processing is transferred directly: a versioned learner state predicts the next rubric profile, human coding supplies the observed profile, and signed residuals carry informative change to reviewers. Reconstruction, uncertainty thresholds, independent raw-response audits, version handshakes, and fallback preserve access to the full signal rather than making the residual queue authoritative.

Why it advanced

This STRICT_SUCCESS passed Experiment 6’s strict researched-candidate bar because it defines a bounded comparator, quantitative success and failure thresholds, human authority, safety bypasses, and a feasible shadow test. That endpoint is not real-world validation, a novelty finding, deployment authorization, evidence of improved learning, or proof that the workflow saves money in practice.

Prior art and the remaining open claim

Adjacent systems already include fixed rubrics, essay scoring, mastery dashboards, adaptive quizzes, answer grouping, work sampling, feedback templates, and knowledge tracing. The remaining claim concerns their narrower combination: predicting a learner-specific rubric profile before opening a held-out interpretive response, representing the response as reconstructable signed residuals, and auditing against full marking. Compared with complete manual review, it must reduce total review time by at least 20% while holding reconstruction disagreement to 5% or less and adding no missed defensible alternative or mandatory-bypass case.

Smallest decisive test

An authorized course partner would supply 48 de-identified responses from 12 students completing four sequential low-stakes assignments. Assignments 1–2 would create transparent predictions; the rubric, model, thresholds, bypasses, and analysis would be frozen before opening assignments 3–4. Counterbalanced reviewers would compare complete-response and residual-first review, while separate blinded graders would double-mark all 24 held-out responses. Success requires at least 20% lower median total review time, no more than 5% reconstruction disagreement, no additional missed defensible alternative or bypass case, no evident language or interpretive-position error pattern, and lower all-in effort after coding, validation, audit, challenge, maintenance, and fallback. Any failure falsifies or narrows the claim.

Deployment and cost

A small transparent shadow prototype is feasible, but live use would require privacy, accessibility, assessment-policy, procurement, and data-governance approval plus domain-literate coding and auditing. Rough 2026 USD resource-equivalent bands are \(10,000–\)50,000 for first evidence, \(250,000–\)1 million for startup, \(50,000–\)250,000 for operational launch, and \(50,000–\)250,000 annually; these are not quotes.

Risks and uncertainties

  • Earlier predictions could anchor reviewers and cause genuine improvement to be discounted.
  • The rubric or learner model could systematically misread a less represented language, tradition, or argumentative style.
  • A defensible interpretation could appear erroneous merely because it was not predicted from earlier work.
  • Residual-focused feedback could omit the educational value of recognizing well-executed continuity.
  • Learner-state records could expose distinctive errors or beliefs and create additional education-record privacy risk.

Expert review

Useful reviewer backgrounds: Religious-studies or theology instructor, Educational measurement and formative-assessment researcher, Learning-analytics or interpretable-model specialist, Accessibility and student-data governance officer, Tradition- and language-literate independent marker.

  1. Are two prior responses sufficiently predictive within the frozen task family to justify any collapsed review?
  2. Can independent graders reconstruct every rubric profile within the 5% disagreement limit?
  3. Which interpretations, languages, accommodations, or task types must always bypass residual-first review?
  4. Does total time still fall by 20% after coding, validation, expansion, auditing, challenges, maintenance, and fallback are included?
  5. Can error patterns by language or interpretive position be examined ethically in a sample this small?

Evidence and provenance

Selected sources: S1: An Empirical Analysis Exploring the Impact of Traditional Exams and Multi-Stage Assignments on Academic Workload in a Final Year Engineering Context · S2: EduMark AI: rethinking assessment and feedback with ethical AI · S3: A-level Religious Studies 7062: Scheme of assessment · S4: AI-assisted grading and answer groups · S5: Interpretable Knowledge Tracing: Simple and Efficient Student Modeling with Causal Relations · S6: Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring · S7: Family Educational Rights and Privacy Act regulations · S8: National employment and wage data by occupation, May 2025
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index

Ordering note: The harmonized review is a post-hoc reading order, not the strict experimental endpoint or a measure of economic value. It placed the candidate between ranks 54 and 56 in band D; its affordability proxy did not independently score elapsed pilot time.


57. Testing Hidden Patterns in Library Exposure

Canonical title: Offline modal-control testing for coupled discovery-exposure imbalance
In one sentence: This post-hoc candidate proposes an offline test of whether coupled exposure patterns persist across digital-library updates and whether targeting those patterns would outperform ordinary monitoring and direct category constraints.

Field Record
Portfolio ID EXP03-POSTHOC-01
Experiment and endpoint Experiment 3 · Post-hoc strict innovation-like survivor
Archetype × domain Invariant Mode Decomposition Design × Library Information Science
Proposal position or arm Not recorded
Post-hoc reading order Balanced score 56.0/100 · rank range 49–58 across three profiles · band D
First-evidence resource band \(50,000–\)250,000
Initial deployment startup band \(250,000–\)1 million

The problem

Digital libraries usually monitor engagement and exposure one item or category at a time. That can miss a combination of modest subject, format, branch, or creator-group imbalances that persists or grows through repeated ranking updates. Such a pattern could steadily narrow what patrons discover even while average engagement and every individual category measure appear acceptable. It is not yet known whether these cross-cycle patterns occur reproducibly in an actual library system.

What is proposed

Using deidentified records from eight ranking-update cycles, analysts would represent each cycle as exposure and engagement deviations across policy-approved resource groups. They would fit a restricted model on six cycles to estimate which combinations persist, decay, or grow, while controlling for catalog availability and query mix. Two sealed cycles would test whether this model predicts exposure better than ordinary category-by-category monitoring. Only then would offline replay compare carefully bounded, mode-targeted ranking adjustments with direct per-category constraints. Nothing would change for patrons. The proposed control would be rejected if it failed utility, privacy, representation, conditioning, residual-error, or stability limits.

The cross-domain transfer

The transferred archetype decomposes a changing system into combinations that evolve together. Here, those combinations are recurring mixtures of library-resource exposure deviations, and their gains describe whether they fade or grow between updates. The mapping is plausible but structurally weak because library discovery is nonlinear, partly human-driven, and represented by only a few observed transitions.

Why it advanced

This candidate was identified only after Experiment 3 through a stricter combined opportunity screen. It is a post-hoc survivor, not a preregistered experimental success. It advanced because library studies document exposure disparities and recommender feedback effects, while an offline, reversible test offers strong safeguards despite substantial data and identification gaps.

Prior art and the remaining open claim

The individual parts already have adjacent prior art: dynamic exposure controls, spectral analysis of recommender feedback, and data-driven mode decomposition with control all exist. The narrower unresolved claim is whether a low-rank operator over approved library-resource groups finds reproducible cross-cycle modes and whether targeting them produces better held-out exposure equity and retrieval utility than both category monitoring and direct category constraints. The search did not establish that library-specific combination or its added value.

Smallest decisive test

Preregister a zero-deployment retrospective study using eight consistently defined cycles. First calculate whether five training transitions contain enough information for the proposed state dimension; stop unless a justified low-rank restriction makes estimation credible. Fit category-lag and modal models on cycles one through six, then evaluate prediction, conditioning, residuals, alignment, and spectral separation on cycles seven and eight. Advance to offline control replay only if the modal model wins; reject the intervention if it fails to beat direct constraints without harming utility or represented groups.

Deployment and cost

The first evidence study is estimated at a rough 2026 resource-equivalent cost of \(50,000–\)250,000. Initial deployment and operational launch are each estimated at \(250,000–\)1 million, with \(50,000–\)250,000 annually. These are assessment bands, not vendor quotes. Deployment would also require logs, platform access, local metadata mapping, and accountable library review.

Risks and uncertainties

  • Analysts could misdescribe statistical modes as traits or preferences of patrons or communities.
  • The model could conceal missing or poorly cataloged resources by treating their absence as a ranking dynamic.
  • With few transitions or a small spectral gap, modes could rotate or exchange identities under minor data changes.
  • Improved modal exposure could come at the cost of retrieval relevance or create harm in an omitted group.
  • Granular exposure strata could leak private information or lack enough observations for reliable estimation.

Expert review

Useful reviewer backgrounds: Library discovery and collection-governance specialist, Recommender-systems fairness researcher, Dynamic-systems or system-identification statistician, Privacy and representation reviewer, Library ranking-platform engineer.

  1. Can the available logs produce eight consistently defined cycles after catalog, query, seasonal, and policy changes are accounted for?
  2. What effective state dimension is estimable from five training transitions, and what low-rank restriction would be defensible?
  3. Do the fitted modes remain aligned under resampling, alternative group definitions, and both sealed cycles?
  4. Would direct per-category constraints provide equal or better equity and utility with less modeling risk?
  5. Which minimum-support and aggregation rules prevent privacy leakage and the reification of communities?

Evidence and provenance

Selected sources: S1: Feedback Loop and Bias Amplification in Recommender Systems · S2: A Study of Position Bias in Digital Library Recommender Systems · S3: Bias in Book Recommendation: A Case Study on the Danish Public Libraries · S4: Guidance on the Use of Artificial Intelligence in Libraries · S5: Fairness of Exposure in Dynamic Recommendation · S6: Deconvolving Feedback Loops in Recommender Systems · S7: Dynamic Mode Decomposition with Control · S8: On Dynamic Mode Decomposition: Theory and Applications
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index

Ordering note: The harmonized review placed this candidate between ranks 49 and 58, in band D. That score is only a post-hoc reading order using a cost-based affordability proxy; it is neither an experimental endpoint nor evidence of economic value.


58. A Governed Contest Between Ethnographic Explanations

Canonical title: Bounded Rival Reanalysis of Ethnographic Explanations
In one sentence: This candidate would test, first with fictional data, whether tightly governed comparison of rival ethnographic explanations exposes cherry-picking and overreach without sacrificing context or unfairly creating a canonical winner.

Field Record
Portfolio ID EXP06-PARTNER-09
Experiment and endpoint Experiment 6 · Empirical-partner candidate
Archetype × domain Bounded Rivalry Governance × Sociology Anthropology
Proposal position or arm P4
Post-hoc reading order Balanced score 56.0/100 · rank range 56–58 across three profiles · band D
First-evidence resource band \(50,000–\)250,000
Initial deployment startup band \(250,000–\)1 million

The problem

When several research teams explain why neighborhood mutual-aid organizations survived or dissolved, scarce publication and funding slots can reward selective cases, privileged contextual help, larger teams, rhetorical confidence, or coordinated submissions. Ordinary peer review sees completed manuscripts but may not reveal how teams chose evidence, handled contradictory episodes, or influenced one another. A winning account can then be treated as authoritative, potentially narrowing later inquiry and misrepresenting the people whose records supplied the evidence.

What is proposed

Editors and an authorized data steward would run a secure, staged comparison under a rulebook fixed before analysis. Eligible teams would receive the same corpus, clarification archive, funded analyst hours, workspace period, and submission format. Each would preregister its main explanation, expected observations, scope, case-selection rule, and treatment of contrary evidence before viewing reserved cases. Submissions would map claims to episodes, alternatives, negative cases, and uncertainty. Independent reviewers would score several dimensions rather than one proxy, reproduce leading analyses, and audit misconduct indicators. Two complementary explanations could be commissioned, while follow-up funding would seek a discriminating observation rather than declare one cultural truth. Awards would remain nonexclusive and appealable.

The cross-domain transfer

The bounded-rivalry archetype becomes a research contest with scarce commissions, equal resources, explicit legal moves, audits, sanctions, appeals, and limits on winner power. The structural mapping is detailed: competition is confined to comparable explanations of one bounded outcome, while participant identity, credibility, lived meaning, and representational authority remain outside the contest.

Why it advanced

It qualified only for the separately calibrated empirical-partner lane, not strict success. Existing practices show that secure qualitative-data review, preregistration, claim-to-evidence annotation, registered reports, and adversarial collaboration are feasible. However, no field study, adopter commitment, prevalence estimate, or evidence of the combined package’s comparative benefit exists.

Prior art and the remaining open claim

Nearly every epistemic component has adjacent prior art, including qualitative preregistration, transparent claim annotation, registered reports, controlled repositories, and empirical adversarial collaboration. The remaining claim concerns their governed combination: whether equal access, reserved cases, multidimensional scoring, leader audits, resource limits, and nonexclusive portfolio selection detect more planted cherry-picking and unsupported scope claims than ordinary review, without materially reducing contextual adequacy, increasing status sensitivity, or hardening one explanation into a canon.

Smallest decisive test

Preregister a synthetic trial using a fictional 18-organization corpus and eight engineered submissions. Randomly assign blinded panels to the bounded arena, ordinary peer review, or a plural symposium, then repeat after changing team names and status cues. Advance only if the arena detects at least 80% of planted cherry-picking or fabrication, keeps false misconduct referrals at or below 10%, maintains rank concordance of at least 0.80, preserves contextual adequacy within 0.30 standard deviations of the symposium, and improves negative-case visibility and scope calibration.

Deployment and cost

The synthetic first-evidence study carries a rough 2026 resource-equivalent estimate of \(50,000–\)250,000. Startup and operational launch are each estimated at \(250,000–\)1 million, with annual recurring costs also at \(250,000–\)1 million. These are not quotations and omit uncertain incident, litigation, and follow-up-grant costs.

Risks and uncertainties

  • A common rubric may falsely treat distinct interpretive traditions as directly comparable.
  • Reserved cases may reward simplified prediction and penalize legitimate contextual revision.
  • Shared records and evidence maps may expose protected context or detach statements from relationships needed to interpret them.
  • Resource limits may still favor teams with existing code, theory, staff, or tacit familiarity.
  • Collusion screens may mistake common intellectual lineage or legitimate collaboration for misconduct signals requiring inquiry alone, not guilt findings.

Expert review

Useful reviewer backgrounds: Qualitative and ethnographic methods scholar, Research-integrity and journal-governance specialist, Qualitative-data repository steward, Human-subjects, privacy, or IRB specialist, Participant or source-community governance representative.

  1. Can reviewers reliably distinguish planted evidentiary weakness from legitimate differences between interpretive traditions?
  2. Do reserved cases measure explanatory adequacy, or do they systematically reward decontextualized prediction?
  3. Which materials can be shared, mapped, and reproduced without violating consent, cultural-access rules, or confidentiality?
  4. How should panel-composition sensitivity and false misconduct referrals affect the decision to stop?
  5. Would a plural symposium, registered report, or collaborative follow-up produce comparable integrity gains with less burden and canonization risk?

Evidence and provenance

Selected sources: S1: Fostering Integrity in Research · S2: Measuring the Prevalence of Questionable Research Practices With Incentives for Truth Telling · S3: ReShare Data Review Procedures · S4: Preregistering Qualitative Research · S5: A Guide to Annotation for Transparent Inquiry (ATI), Version 1.0 · S6: Registered Reports · S7: Coded Private Information or Biospecimens Used in Research, Guidance (2018) · S8: Rationale and Guidelines for Empirical Adversarial Collaboration: A Thinking & Reasoning Initiative
Original records: Original proposal · External evaluation · Partner-lane adjudication · Experiment report · Machine-readable candidate index

Ordering note: The harmonized review placed this candidate between ranks 56 and 58, in band D. This is a post-hoc reading aid, not an experimental endpoint or value estimate; its affordability input substitutes cost bands for an independently measured pilot timeline.


59. Mapping Coupled City Budget Deadlocks

Canonical title: Modal Deadlock Map for City Budget Bargaining
In one sentence: This candidate proposes a retrospective test of whether combinations of moderate budget disagreements predict bargaining trouble better than ordinary issue-by-issue tracking and can be linked to reversible procedural responses.

Field Record
Portfolio ID EXP04-STRICT-03
Experiment and endpoint Experiment 4 · Strict success
Archetype × domain Invariant Mode Decomposition Design × Political Science
Proposal position or arm PROPOSAL_FIRST
Post-hoc reading order Balanced score 51.0/100 · rank range 59–59 across three profiles · band D
First-evidence resource band \(10,000–\)50,000
Initial deployment startup band \(50,000–\)250,000

The problem

City budget negotiators bargain over packages covering revenue, staffing, capital projects, debt, reserves, and district allocations. Officials usually track each disagreement separately, although movement on one issue can change positions on several others. A combination of moderate gaps may therefore persist, oscillate, or grow while no single issue looks exceptional. The negotiation can drift toward repeated rejection, rushed concessions, or a missed legal deadline before an issue-by-issue dashboard provides a clear warning.

What is proposed

For one bounded budget process, analysts would turn formally recorded, caucus-level proposal gaps into a small vector for each bargaining round. A local model would estimate which combinations of gaps decay, persist, oscillate, or grow. A mode would become actionable only if it met preregistered gain, persistence, resampling, reconstruction, and deadline-risk rules. Analysts would then model whether an already authorized procedural step—such as reordering discussion, separating a package, holding a joint factual briefing, or requesting simultaneous clarification—selectively reduces that mode. The facilitator could make a nonbinding recommendation, but elected officials would retain all authority over offers and votes. Drift, residual, or spectral-gap failures would suspend the dashboard.

The cross-domain transfer

The mode-decomposition archetype is instantiated as a map of how combined budget gaps change between bargaining rounds. Its modes are weighted packages of disagreement, and their gains indicate decay, growth, or oscillation. The transfer is structurally explicit, but its usefulness depends on an approximately stable local process and enough comparable rounds—both currently unshown.

Why it advanced

This candidate passed Experiment 4’s strict researched-candidate bar. That endpoint reflects the experiment’s screening criteria only; it does not establish real-world validation, novelty, deployment authority, or economic impact. It advanced with a falsifiable archival test, clear comparators and stopping rules, identifiable public authorities, and reversible, nonbinding use.

Prior art and the remaining open claim

Adjacent systems already track city budget amendments, support multi-issue negotiation, analyze negotiation processes over time, and decompose fitted dynamic operators. The open claim is narrower: within one stable city budget process, can an aggregate proposal-gap model find a resampling-stable coupled mode that improves held-out prediction over separate issue gaps, deadlines, official tracking, and observable shocks, leaves no consequential structured residual, and maps selectively to a procedure the authorized body can reverse? That claim remains unvalidated.

Smallest decisive test

Run an eight-week archival feasibility study with one consenting city and no live recommendation. Inventory formal packages, amendments, votes, and dated forecasts; stop if they cannot form reproducible aggregate snapshots. Preregister no more than eight coordinates, rank and sample rules, regularization, stability and spectral thresholds, residual limits, and both comparators. Use forward-chaining held-out tests. Advance only if a stable mode beats separate issue gaps plus deadline and the official tracker plus observable shocks, survives removal of shock rounds, and maps selectively to an authorized reversible procedure.

Deployment and cost

First evidence is estimated at a rough 2026 resource-equivalent cost of \(10,000–\)50,000. Initial startup is \(50,000–\)250,000; operational launch is \(250,000–\)1 million; and recurring annual cost is \(50,000–\)250,000. These are not vendor quotes. Live use would additionally require local legal, records, accessibility, governance, and procedural authorization.

Risks and uncertainties

  • Too few comparable bargaining rounds could produce an overfit or rank-deficient model.
  • Modes might reflect how clerks recorded proposals rather than how bargaining actually changed.
  • Officials or the public might interpret a descriptive mode as evidence of motive, loyalty, or blame.
  • A nominally procedural recommendation could redistribute agenda power or public visibility among constituencies.
  • Negotiators could manipulate recorded positions, or conflict could move into an omitted dimension while the retained mode appears to improve.

Expert review

Useful reviewer backgrounds: Municipal budget-process and public-law specialist, Negotiation and political-process researcher, Dynamic-mode decomposition or system-identification statistician, Neutral public-sector facilitator, Public-records, privacy, and data-governance counsel.

  1. Does the archived process contain enough independent full-state snapshots to estimate the preregistered model reliably?
  2. Do mode shapes and gains remain stable under resampling, minor coordinate changes, and removal of shock rounds?
  3. Does the modal model beat both comparators on held-out prediction or reconstruction by the declared margin?
  4. Can any procedural lever be shown to affect the risky mode selectively rather than merely accompany broader political change?
  5. Who has legal authority to approve each recommended procedure, derived-data retention rule, and publication decision?

Evidence and provenance

Selected sources: S1: Coalition governance and municipal stability in South Africa: Institutional challenges and reform imperatives · S2: The Budget Process · S3: Budget Extender, Local Law 102 of 2026 (Int. 0873-2026) · S4: Seattle City Council budget amendment tracker · S5: NegoManage: A System for Supporting Bilateral Negotiations · S6: Analyzing the Multiple Dimensions of Negotiation Processes · S7: On Dynamic Mode Decomposition: Theory and Applications · S8: Math Occupations — Occupational Outlook Handbook
Original records: Original proposal · External evaluation · Experiment report · Machine-readable candidate index

Ordering note: The harmonized review ranked this candidate 59th in every profile, in band D. That post-hoc order does not alter its strict-success endpoint and does not measure value; the pilot-speed input was only a cost-band affordability proxy.