{"dossiers":[{"portfolio_id":"EXP06-STRICT-15","plain_language_title":"Making Carbon-Budget Boundaries Visible","one_sentence_summary":"At each regional carbon-budget release freeze, teams would jointly expose mismatched boundaries and unresolved residuals, then carry agreed disclosures into ordinary publication controls.","problem_plain":"Regional carbon budgets combine atmospheric, forest, soil, aquatic, and land-use estimates that may cover different places, periods, or definitions. Although specialists document limitations, those details can remain scattered in appendices and team files. A polished balance table can then look more complete than it is: exclusions disappear from headlines, residual quantities lack an owner or explanation, and contributors give different accounts of what the published total includes and leaves unresolved.","proposal_plain":"After normal technical reconciliation but before release approval, the consortium would hold a voluntary 50-minute Open-Boundary Assembly. Each component team would display a removable layer showing its spatial mask and time window, then state what its estimate includes, excludes, and cannot resolve. The response “heard, not resolved” would acknowledge the statement without implying agreement. A residual marker would pass among interface stewards, who could flag a mismatch, assign it for examination, declare it irreducible for that release, or pass. Contributors could accept, revise, transfer, or decline responsibility for carrying limitations into particular outputs. The editor—not the assembly—would then update the versioned boundary ledger, claims, graphics, metadata, owners, deadlines, and open questions through the existing publication system.","transfer_plain":"The ritualized-commitment archetype becomes a marked, repeatable scientific checkpoint. Visible layers symbolize that the regional total is assembled from partial views; standardized statements make limitations collectively audible; witnessed handoffs renew responsibility. The mapping is structurally strong, but its claimed benefit depends on ritual features adding something beyond equally careful facilitation and documentation.","why_it_advanced":"This candidate passed Experiment 6’s strict researched-candidate bar because it defines a bounded problem, preserves scientific authority, specifies a close comparator, provides falsifiers and safety controls, and proposes a measurable randomized pilot. STRICT_SUCCESS does not mean the assembly has worked in practice, is novel worldwide, is authorized for deployment, or produces economic value.","prior_art_and_open_claim":"The parts have substantial adjacent prior art: carbon accounting already uses QA/QC, completeness and double-counting checks, formal review, uncertainty tables, facilitated workshops, shared records, and version control. Workplace research also suggests rituals can increase perceived meaning, but not carbon-budget accuracy. The remaining claim is narrower: adding this voluntary ritual layer to an otherwise identical 50-minute boundary-review checklist improves detection, delayed recall, and fulfillment of disclosure duties without creating coercion, false assent, confidentiality failures, or authority confusion.","test_and_decision":"Preregister a remote crossover study with 24–40 carbon-cycle or adjacent researchers in four to six balanced teams. Teams would review matched synthetic packets containing hidden boundary, stock-flow, covariance, exclusion, and residual defects, using either the assembly or an equal-duration facilitated checklist with the same facts and ledger. Advance only for at least a 20-percentage-point improvement in detection or delayed recall, or two substantively new correct disclosures per team, with no material loss of mock-output accuracy and no credible safety incident. Equivalent performance or any safety breach falsifies the incremental claim.","deployment_and_cost":"The first study and startup are each estimated at $10,000–$50,000 in rough 2026 resource-equivalent terms. Operational launch is also $10,000–$50,000, while recurring annual operation is estimated at $50,000–$250,000. No consortium has committed authority, staff time, or funding; live evaluation would also require access to versioned synthesis artifacts.","risks_and_uncertainties":["Transparent layers may hide nonlinear processes, covariance, scale differences, or nonspatial boundaries.","“Heard, not resolved” may become rote or be mistaken for scientific agreement.","Marker circulation may pressure participants to explain quantities outside their expertise.","Publicly witnessed dissent could expose junior contributors or minority interpretations to retaliation.","The ceremony could make an unchanged or misleading synthesis appear unusually coherent and legitimate despite unresolved evidence gaps."],"expert_types":["Regional carbon-budget scientist","Carbon-accounting and uncertainty specialist","Environmental synthesis editor or data steward","Research-team governance and facilitation specialist","Accessibility, consent, and confidentiality reviewer"],"expert_questions":["Can existing release records establish that documented boundary limitations actually disappear from headline tables, graphics, or summaries?","Would the planted defects and scoring rubric represent consequential regional-budget interface failures rather than merely easy checklist items?","Can the checklist and assembly arms be matched for facts, facilitation quality, documentation, and time?","What safety threshold should govern reports of pressure, false assent, confidentiality loss, or authority confusion?","Which live or retrospectively versioned artifacts could measure whether disclosure obligations were ultimately fulfilled?"],"ranking_note":"The harmonized review is only a post-hoc reading aid: scores range from 57 to 63 across profiles, with ranks 45–53 and ordering band D. Its speed input is an affordability proxy, not measured elapsed time or economic value, and it does not alter STRICT_SUCCESS status.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP04-STRICT-06","plain_language_title":"Reserve AI Capacity for Harm Inquiry","one_sentence_summary":"Organizations would protect at least 15% of annual AI-pilot capacity for nondeployment investigations before live-use projects consume the portfolio.","problem_plain":"When an organization allocates its annual AI-pilot portfolio, live-use proposals can consume all available capacity before individual approvals occur. Red-teaming, shadow evaluation, and affected-community inquiry must then compete for whatever remains. Material problems may consequently emerge only after people are exposed and projects have accumulated operational dependencies. The underlying prevalence is unknown, however, and existing risk-tiered assurance may already provide adequate predeployment coverage without a fixed reserve.","proposal_plain":"Before approving individual pilots, the portfolio governing body would reserve at least 15% of total annual pilot capacity exclusively for nondeployment harm inquiry. The protected share could fund red-teaming, shadow evaluation, or appropriately safeguarded affected-community work, but not live deployment. Separate accounting would prevent project teams from informally absorbing it. Reallocation would require the governing body to document that eligible inquiry demand was exhausted, obtain concurrence from an independent assurance function, record the destination and reasons, and preserve enough capacity to complete active inquiries. Mandatory privacy, security, accessibility, safety, and incident-response work would remain outside the contest for this reserve. Existing approval and risk-management processes would continue to control deployment decisions.","transfer_plain":"Negative-space design is instantiated as deliberately unused deployment capacity: a protected portfolio “void” is bounded before surrounding projects are selected. Its positive purpose is to preserve room for inquiry while changes remain feasible. The structural transfer is clear, although percentage ring-fencing and controlled release already exist in adjacent evaluation and innovation portfolios.","why_it_advanced":"This candidate passed Experiment 4’s strict researched-candidate bar by defining the denominator, protected use, release rule, comparator, measurable outcomes, and a records-only first step. STRICT_SUCCESS is limited to that research screen; it is not evidence that 15% is optimal, that harm detection improves, or that an organization will adopt the rule.","prior_art_and_open_claim":"Adjacent systems already allocate resources to AI testing, independent assurance, governance boards, sandboxes, recurring evaluations, and risk-tiered review. Other fields also ring-fence evaluation funding or divide portfolios by fixed percentages. The remaining contrastive claim is specifically that a preapproval 15% nondeployment reserve, protected from live-use commitments and released only with documented independent concurrence, increases independently adjudicated material issues found per proposed system without lowering the number of validated pilots by more than 10%.","test_and_decision":"With one willing organization, preregister a records-only reconstruction of a completed annual portfolio. Compare the actual flexible or risk-tiered allocation with a 15% protected shadow allocation applied before proposal selection. Define capacity, eligible inquiry, materiality, and missing-data rules in advance; use two issue reviewers, including one independent of pilot selection. Measure material issues per proposed system, validated-pilot count, inquiry completion, and simulated compliance with release rules. Do not advance if issue detection does not increase, validated pilots fall by more than 10%, or capacity and eligible work cannot be measured consistently.","deployment_and_cost":"First evidence and initial startup are each estimated at $50,000–$250,000. Operational launch and annual recurring costs are each estimated at $250,000–$1 million, in rough 2026 resource-equivalent bands rather than vendor quotes. Actual cost depends heavily on specialist red teams, community participation, compute, data preparation, and displaced pilot capacity.","risks_and_uncertainties":["The 15% threshold may be too large, too small, or meaningless for very small portfolios.","Teams may relabel ordinary development or compliance work as protected inquiry.","Issue counts may reward numerous trivial findings unless independent reviewers apply a reproducible materiality standard.","A fixed reserve could sit unused while beneficial pilots wait, or encourage wasteful spending merely to exhaust it.","Observational results may confuse the reserve’s effect with an organization’s pre-existing safety culture and staffing quality."],"expert_types":["AI portfolio-governance leader","Independent AI assurance or audit specialist","Causal-inference and program-evaluation researcher","Affected-community research and safeguarding specialist","Organizational finance and capacity-accounting expert"],"expert_questions":["Can pilot capacity be expressed in a consistent unit across projects, staff, compute, and external assurance?","Which work qualifies as nondeployment harm inquiry rather than normal development or mandatory compliance?","Can historical records establish when an issue was found, whether it was material, and whether it changed design or approval?","Is flexible risk-tiered assurance a more credible comparator than the organization’s actual historical allocation alone?","How should continuous deployment, procurement, model updates, and small portfolios be handled outside the annual denominator?"],"ranking_note":"The post-hoc harmonized reading aid scores this candidate 58–60 across profiles, ranks it 49–50, and places it in band D. The pilot-speed input reflects cost-band affordability, not observed duration. These figures neither measure economic value nor replace its STRICT_SUCCESS endpoint.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-PARTNER-01","plain_language_title":"Audited Competition for Close-Review Priority","one_sentence_summary":"Business units would compete for scarce early financial-close review slots using independently verified readiness evidence rather than self-declared completion or managerial escalation.","problem_plain":"During a multi-entity financial close, business units may compete for a limited number of early consolidation or technical-accounting reviews. If self-certified readiness or managerial escalation controls the queue, a unit can gain priority by prematurely closing reconciliations, shifting exceptions to another entity, deferring unsupported items, or obtaining privileged reviewer access. Reviewers then reopen supposedly complete work while other units wait. Crucially, no external evidence yet shows that this scarcity or strategic behavior exists in a target organization.","proposal_plain":"The corporate controller would replace the informal queue race with a recurring, bounded readiness contest. Units meeting a minimum control floor could compete for several early review slots. A rulebook fixed before scoring would allow genuine documentation, automation, supported resolution, and timely escalation, while prohibiting concealment, unsupported deferral, exception dumping, shared answers, off-channel influence, and undisclosed borrowed labor. An independent verifier would score hidden, risk-stratified samples for first-pass evidentiary completeness, supported treatment of aged exceptions, intercompany agreement, and absence of later unsupported corrections. Timely disclosure of material issues would have a protected route. Winners would be audited, serious errors appealable, extra contest labor capped, and some awarded capacity retained for remediation. Priority would expire after one close.","transfer_plain":"Bounded-rivalry governance becomes a formal contest for a scarce operational prize. Eligibility, lawful tactics, fouls, judging, appeals, resource caps, multiple awards, spillover responsibility, and recurring challenger access constrain how units compete. The mapping is detailed, but it remains unproven that queue positions create meaningful strategic interdependence rather than reflecting ordinary systems, complexity, or staffing constraints.","why_it_advanced":"This candidate did not enter Experiment 6’s strict-success lane. It cleared the separately calibrated EMPIRICAL_PARTNER_CANDIDATE lane because a bounded retrospective partner study is feasible and decision-relevant. Its advancement depends on obtaining proprietary field records; there is currently no evidence of local queue scarcity, manipulation, predictive advantage, or safe behavioral response.","prior_art_and_open_claim":"Financial-close platforms already provide workflows, dependencies, approvals, dashboards, audit trails, exception monitoring, and risk-based administrative triage. Accounting standards also require evidence, objective verification, and controls over period-end reporting. The narrower open claim, conditional on consequential rivalry being demonstrated, is that hidden-sample, independently verified competitive ranking predicts less first-pass rework, fewer reopened or transferred exceptions, and fewer unsupported corrections than both current queueing and noncompetitive risk-based triage, without delaying protected disclosure of material issues.","test_and_decision":"Preregister a retrospective replay of one completed close across four to eight entities and 40–80 risk-stratified reconciliations. Compare recorded queue order, blinded noncompetitive risk-based triage, and the proposed arena score. Measure verified completeness, rework hours, reopened exceptions, transferred mismatches, unsupported correcting entries, and disclosure timing. Report stability under bootstrap samples and reasonable weight changes. Do not proceed beyond a no-consequence simulation if the arena fails to outperform triage out of sample, Kendall rank stability is below 0.60, size or complexity drives results, data are inconsistent, or issue reporting appears delayed or suppressed.","deployment_and_cost":"A retrospective first study is estimated at $10,000–$50,000. Initial startup is $50,000–$250,000; operational launch and annual recurring operation are each $250,000–$1 million in rough 2026 resource-equivalent terms. Software, data integration, independent verification, appeals, labor-cap auditing, and recurring governance still require organization-specific estimates, not vendor assumptions.","risks_and_uncertainties":["Teams may game evidence fields that hidden samples do not cover.","Risk adjustment may embed incumbent complexity assumptions and disadvantage legitimate challengers.","Labor caps may be evaded through off-books work or restrict units with genuine remediation needs.","Coordination screens may falsely flag common deadlines or shared system failures as collusion.","Winning units may accumulate reviewer relationships and procedural knowledge even though formal priority expires."],"expert_types":["Corporate controller or consolidation leader","Internal-control and ICFR specialist","Internal auditor independent of close operations","Financial-close data and workflow engineer","Employment, legal, and external-audit governance adviser"],"expert_questions":["Are early specialist-review slots genuinely scarce, and does their timing materially affect other entities’ outcomes?","Can queue, reconciliation, exception, rework, escalation, and correction records be joined reliably across entities?","Does the proposed score outperform administrative risk triage on data not used to construct the ranking?","How can protected material-issue reporting be separated from competitive scoring and monitored for delay?","Would labor caps, hidden sampling, and anomaly screening conflict with employment rules, ICFR responsibilities, or external-audit arrangements?"],"ranking_note":"The post-hoc harmonized reading aid gives scores of 56–62, ranks 47–54, and band D. Its speed input is a cost-affordability proxy rather than elapsed time. This ordering is not an experimental endpoint or economic-value estimate and does not upgrade the partner-candidate status.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-PARTNER-28","plain_language_title":"Renewing Climate-Monitoring Responsibilities","one_sentence_summary":"Before each field season, monitoring staff would visibly renew or transfer station-to-archive duties and then record every accepted obligation in the ordinary task system.","problem_plain":"Long-term climate monitoring depends on seasonal sampling, calibration, metadata, custody transfers, and reliable handoffs despite staff turnover. Written protocols may leave particular responsibilities privately understood or attached to departed personnel. A field season can therefore begin without a named primary, backup, needed resource, credential, or escalation route for every station-to-archive link. Missing duties or tacit knowledge may surface only when work is due, creating gaps or ambiguities in a record intended to remain comparable over decades.","proposal_plain":"After the normal technical readiness review and before each field season, the network would hold a voluntary 35-minute Season-Turn Stewardship Muster. A neutral signal would mark the occasion, and participants would trace a fictional or real observation from station to archive by moving plain link cards across a network map. At each step, the current steward could renew, revise, transfer, or decline responsibility without explaining publicly. A backup and missing resources would be named before acceptance. Witnesses could acknowledge completed handoffs, but attendance, speech, silence, or card handling would not create consent or authority. Closure would occur only when ordinary managers enter an owner, backup, resources, due date, and escalation route into the existing task system. Debrief and independent review could modify, pause, or retire the practice.","transfer_plain":"The ritual archetype becomes a marked seasonal enactment of shared stewardship. Card movement makes the custody chain visible; voluntary renewal prevents stale assignments from persisting silently; witnessed handoffs support memory across turnover. The mapping is plausible, but most operational content duplicates ordinary responsibility maps and structured handoffs, leaving only the ritual features’ incremental social-memory effect to test.","why_it_advanced":"This proposal did not enter the strict-success lane. It cleared the EMPIRICAL_PARTNER_CANDIDATE lane because a small, fictional-record crossover with an external monitoring partner could test the remaining contrast safely. Field evidence is missing on problem prevalence, adopter demand, improved recall, dependency discovery, operational completion, voluntariness, and accessibility.","prior_art_and_open_claim":"Climate networks already require sustained operations, calibration, metadata, annual maintenance, anomaly tracking, preseason reviews, assigned roles, and configuration control. Structured handoffs also cover ownership, acknowledgment, next steps, and clarification; workplace ritual studies measured meaning, not monitoring continuity. The remaining claim is that adding a neutral threshold, voluntary card traversal, and witnessed renewal to an equal-duration administrative session improves seven-day recall or surfaces more valid dependencies without increasing pressure, exclusion, religious conflict, or confusion about formal authority.","test_and_decision":"With one willing network, preregister a crossover using four matched fictional station-to-archive records and 12–24 participants. Compare a 35-minute administrative readiness session with the same session plus the ritual features, crossing teams onto a different record. Measure valid actionable dependencies and blinded immediate and seven-day recall of each primary, backup, and escalation route. Falsify incremental promise if the ritual finds no additional valid dependency and improves complete-link recall by less than 15 percentage points, or worsens pressure or access ratings by at least 0.5 on a five-point scale. Any credible coercion, privacy, cultural, or authority incident requires a halt.","deployment_and_cost":"First evidence is estimated below $10,000. Initial startup, operational launch, and annual recurring operation are each estimated at $10,000–$50,000 in rough 2026 resource-equivalent bands. Actual costs remain uncertain because network size, travel, participant count, union requirements, accessibility work, facilitation, and task-system integration have not been specified.","risks_and_uncertainties":["Employment hierarchy may make a formally optional refusal feel professionally costly.","The event may aestheticize stewardship while staffing, equipment, training, or travel shortages remain unfunded.","Public handoffs may expose disability, location, performance, or employment information.","The linear card metaphor may distort parallel, contested, or nonlinear observation pathways.","Neutral-looking symbolism may still create religious or cultural conflict, especially if local designers add unauthorized ceremonial elements."],"expert_types":["Climate-monitoring network operator","Field-to-archive data and metadata steward","Human-factors and structured-handoff researcher","Workplace accessibility and religious-accommodation specialist","Program evaluator experienced in crossover trials"],"expert_questions":["Do ordinary preseason records actually contain missing owners, backups, resources, credentials, or escalation routes?","Are the fictional station-to-archive cases realistic enough to test consequential dependencies without exposing operational data?","Can the administrative comparator match all content, time, facilitation, and task closure except the ritual features?","How will anonymous measures detect pressure when managers and subordinates participate together?","What result would justify a prospective no-consequence simulation before any real responsibility transfer is considered?"],"ranking_note":"The post-hoc harmonized reading aid scores this candidate 55–63, with ranks 44–58 and band D. Its speed measure is a cost-affordability proxy, not observed duration. The wide profile range is not economic value evidence and leaves the EMPIRICAL_PARTNER_CANDIDATE endpoint unchanged.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]}]}