{"dossiers":[{"portfolio_id":"EXP06-STRICT-12","plain_language_title":"Renewing Ramp Stop-Work Support","one_sentence_summary":"A monthly, voluntary exercise would let workers from different ramp employers rehearse mutual recognition of approved stop-work calls and expose unsupported response paths without changing operating authority.","problem_plain":"Aircraft turnarounds bring together airline staff and separate fueling, baggage, catering, cleaning, maintenance, towing, and ground-handling employers. Each may teach its own safety rules, yet workers can still disagree about who may pause shared work, how other companies must respond, and whether a junior contractor will be protected for interrupting a senior or time-critical operation. This uncertainty can delay or silence warnings even when every employer has a written stop-work policy.","proposal_plain":"Pilot a voluntary twenty-minute monthly observance at a classroom aircraft outline or closed training stand, never during a live turnaround. A rotating steward explains that the exercise creates no qualification, authority, or procedure. Participants from different employers trace or otherwise follow the aircraft boundary, hear frontline accounts of four cross-company dependencies, and exchange cards listing only approved support such as acknowledgment, interpretation, or escalation routes. Missing or contradictory support becomes a recorded safety-management action rather than an improvised promise. People may speak, write, observe silently, or pass without penalty. A debrief checks pressure, exclusion, ambiguity, and follow-through; independent review can require repair, suspension, or retirement.","transfer_plain":"The source archetype uses a recurring, marked ritual to make a shared commitment visible and memorable. Here, the ritual becomes a cross-employer aircraft-boundary exercise, reciprocal support-card exchange, plural response, and steward handoff. The mapping is structurally strong, but its symbolic elements have no assumed safety effect; their incremental value must be compared with an ordinary joint briefing.","why_it_advanced":"This candidate passed Experiment 6’s strict researched-candidate bar because it defines a bounded comparator, reversible first test, authority limits, concrete falsifiers, and safeguards against coercion and procedural confusion. STRICT_SUCCESS is only that experimental endpoint: it does not establish real-world effectiveness, novelty, deployment approval, or economic value.","prior_art_and_open_claim":"The surrounding parts already exist: joint ramp briefings, written escalation procedures, safety stand-downs, public safety pledges, reporting systems, and operational roles with stop-work authority. The open claim is narrower. Where approved routes already exist, does adding this opt-out monthly enactment improve unaided reconstruction of reciprocal cross-company responses and assignment of unsupported interfaces to authorized owners, compared with an equal-duration conventional briefing, without added pressure, retaliation concern, access loss, or authority confusion?","test_and_decision":"At one willing station, place 8–12 workers from at least three organizations into matched twenty-minute sessions using identical fictional scenarios and route information: the proposed renewal or a conventional joint briefing. Blind-score who may call an approved pause, each organization’s response and escalation path, conflict handling, and what the session did not authorize. Also record whether planted gaps receive an owner, forum, and date, plus anonymous pressure and access measures. Proceed, revise, or stop; do not move to live operations without separate authorization.","deployment_and_cost":"The first-evidence estimate is under $10,000. Initial deployment, operational launch, and annual recurring support are each roughly $10,000–$50,000 in 2026 resource-equivalent terms, not vendor quotes. Local policy mapping, translation, shift coverage, accessibility, facilitation, auditing, and correction of discovered communication or staffing gaps could materially change the total.","risks_and_uncertainties":["Managers could point to visible agreement while leaving anti-retaliation protection weak.","Contractors or junior workers may feel compelled to affirm support in front of supervisors.","Participants may mistake ceremonial statements or cards for approved operating procedures.","Mobility, sensory, language, remote-access, and shift constraints may make participation unequal.","A dominant airline or handler could control the script, stewardship rotation, or interpretation of mutual support promises pressure, retaliation concern, accessibility loss, or confusion about operational authority increase? The proposal needs aviation SMS specialists, ground-operations practitioners, worker representatives, accessibility experts, and organizational-behavior researchers to validate local authority and measurement design."],"expert_types":["Airport or airline safety-management-system specialist","Ground-handling operations leader","Ramp worker, union, or contractor representative","Aviation human-factors and safety-culture researcher","Accessibility and language-access specialist"],"expert_questions":["Do local policies and contracts already define reciprocal stop-work acknowledgment across every participating employer, and where do they conflict?","Can workers decline speaking, moving, or card exchange without conspicuous refusal or employment consequences?","Does the matched briefing control contain exactly the same approved route information and facilitator time?","What blinded scoring rule will distinguish genuine route reconstruction from recall of ceremonial wording?","Which authority can accept each discovered gap, fund its correction, and verify closure?"] ,"ranking_note":"The harmonized score is only a post-hoc reading aid: band C, with profile ranks from 34 to 46. Its affordability input proxies pilot speed and does not measure elapsed time, economic value, or the STRICT_SUCCESS endpoint.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP05-STRICT-02","plain_language_title":"Expiring Old Training Evidence Safely","one_sentence_summary":"A shadow system would retire contextually stale training evidence from current predictions while preserving the records and model versions needed to reconstruct earlier recommendations.","problem_plain":"An adaptive cognitive-training system can keep using performance evidence from older sessions after a person’s ability, task design, scoring model, or context has changed. Those records may then distort current difficulty choices. Simply deleting them creates a different problem: researchers may lose the history needed to explain a trajectory, reproduce an earlier recommendation, investigate model behavior, or meet retention duties. The system needs to separate evidence’s authority in current inference from its physical preservation.","proposal_plain":"Assign each trial-derived evidence layer an inference lease when it is created, based on evidence type, task and scoring versions, context, and a review horizon. Expiry removes the layer from ordinary current-state estimation but sends it to a provenance-preserving archive rather than deleting it. New corroboration can renew a lease; task, scoring, or context changes can trigger early review. A multi-factor score may order review work but cannot decide disposition. Before destruction, stewards must check retention authority and downstream dependencies, use reversible quarantine, and leave a tombstone. Sampled restore drills test whether earlier estimates and recommendations remain reproducible. Initial evaluation occurs only through offline shadow replay, leaving source data, the live model, and participant experience unchanged.","transfer_plain":"The archetype manages accumulated layers whose usefulness decays or expires. In this domain, the layers are timestamped trials, summaries, context annotations, and model bindings. Expiration changes whether evidence influences today’s estimate, while storage tiers, dependency checks, quarantine, tombstones, and restore tests preserve accountable history. The structural mapping is direct, although temporal models may already address the predictive problem.","why_it_advanced":"This candidate passed Experiment 5’s strict researched-candidate bar through a specific comparator set, measurable offline test, reversible implementation, separated decision authority, and explicit failure rules. STRICT_SUCCESS does not show that stale evidence is prevalent, that leases beat strong temporal models, or that production use is warranted, novel, or economical.","prior_art_and_open_claim":"Adjacent systems already apply sliding windows, uniform forgetting, state-space or change-point models, time-aware knowledge tracing, immutable event logs, replay, and storage lifecycle policies. The unresolved comparison is whether explicit context-sensitive leases offer a smaller active evidence set that matches or improves held-out current-session prediction against cumulative and fixed-recency estimators—and remains competitive with a strong temporal model—while preserving historical reconstruction and keeping false-staleness, recommendation instability, and steward workload acceptable.","test_and_decision":"With written data-steward approval, replay a completed multi-session working-memory dataset containing at least two documented task or scoring contexts. Freeze outputs from the existing cumulative estimator, then compare cumulative history, a fixed recent window, and context-sensitive leases; add a strong temporal-model sensitivity analysis if feasible. Measure held-out prediction, recommendation quality against a task-owner rubric, calibration, transitions, active-set size, false staleness, dependency recall, review minutes, and exact historical reconstruction. Reject progression after any missed seeded dependency, failed reconstruction, unacceptable instability or burden, unauthorized exposure, or baseline-reproduction failure.","deployment_and_cost":"First evidence is estimated at $50,000–$250,000. Initial deployment and operational launch are each roughly $250,000–$1 million; annual recurring work is $50,000–$250,000, in 2026 resource-equivalent bands rather than quotes. Actual cost depends on data volume, metadata quality, architecture, legal regimes, dependency tracing, expert review, and integration debt.","risks_and_uncertainties":["Short or poorly chosen leases could overemphasize recent observations and discard stable information from active inference.","Context rules may encode sensitive attributes or operate unevenly across participants.","Adaptive task selection can make recent evidence look more informative because the system chose what was observed.","A composite review score may hide contestable governance judgments behind numerical precision.","Incomplete dependency mapping could permit removal of evidence needed to explain an earlier recommendation or report error rates versus cumulative, fixed-window, and strong temporal-model comparators? Data governance, cognitive science, adaptive-modeling, privacy law, and archival systems expertise are required before production consideration."],"expert_types":["Cognitive scientist specializing in working-memory measurement","Adaptive-learning or knowledge-tracing modeler","Data steward or research-records officer","Privacy and data-retention counsel","Event-sourcing and archival-reconstruction engineer"],"expert_questions":["What observed task, scoring, or context changes make an older trial semantically inapplicable rather than merely old?","Which primary metric and non-inferiority margins will compare leases with cumulative, fixed-window, and strong temporal models?","How will experts label false staleness without using the tested lease rules as their own ground truth?","Can every sampled historical recommendation be recreated with the archived evidence, code, configuration, and model version?","Which participant permissions, research requirements, holds, or deletion duties govern archive, quarantine, and destruction?"] ,"ranking_note":"The post-hoc harmonized ordering places this candidate in band C, with profile ranks from 39 to 44. That ordering is not an experimental endpoint or economic-value estimate; its pilot-speed input is only a cost-band affordability proxy.","source_ids_used":["s1","s2","s3","s4","s5","s6","s7","s8"]},{"portfolio_id":"EXP06-PARTNER-07","plain_language_title":"A Capped Prize for Catalyst Endurance","one_sentence_summary":"A sponsor would compare coded nanocatalysts through resource-capped, independently audited endurance tests rather than select a demonstration candidate mainly from publications or peak reported performance.","problem_plain":"A sponsor with one follow-on demonstration award may rank nanocatalyst teams using publications, presentations, and peak results obtained under different conditions. That can favor short favorable runs, pure feedstocks, high scarce-metal use, selected batches, or heavy synthesis and computing expenditure. The chosen catalyst may then prove fragile during longer or variable operation. Yet the alleged selection process, incomplete reporting, resource escalation, and connection to later demonstration failure have not been documented for a named sponsor.","proposal_plain":"Replace publication-priority selection with a preregistered endurance prize whose rules are frozen before entrants are known. Teams register preparations, failed attempts, contributors, counted spending and in-kind inputs, reactor and accelerator-compute hours, scarce materials, and waste. Coded samples go to an independent laboratory for public qualification and undisclosed but representative endurance, feed-variation, restart, and regeneration sequences. Safety and containment are non-compensable gates. Passing entries are scored for reproducibility, sustained conversion and selectivity, deactivation and recovery, material intensity, energy, waste, and reporting completeness. Resource caps limit escalation. Two technically distinct milestone awards preserve alternatives before one demonstration award. Audits, appeals, conflict screening, graduated penalties, cleanup guarantees, and a challenger path govern the contest.","transfer_plain":"Bounded-rivalry governance redirects competition by fixing the arena, limiting inputs, policing interference, preserving alternatives, and reviewing winner lock-in. Here those elements become a capped catalyst prize with coded testing, milestone awards, safety floors, resource and waste accounting, sanctions, and continuation conditions. The mapping is strong, but the prize’s claimed predictive advantage over standardized endurance testing alone remains unmeasured.","why_it_advanced":"This did not enter Experiment 6’s strict-success lane. It cleared the separately calibrated EMPIRICAL_PARTNER_CANDIDATE lane because a bounded retrospective partner study is testable and safety-limited. Progress depends on a named sponsor and proprietary records; neither the field problem nor the full contest design’s incremental benefit has yet been demonstrated.","prior_art_and_open_claim":"Catalyst benchmarking, degradation protocols, independent laboratory validation, staged federal prizes, advance scoring rules, safety controls, and appeals already exist separately. The remaining claim concerns their combination: would adding auditable input caps, complete-attempt reporting, lifecycle burdens, independent milestones, hidden representative segments, and a continuation challenge predict later demonstration performance better than both historical peak-result selection and standardized endurance testing alone? Located evidence supports the components, not that comparative or behavioral result.","test_and_decision":"With a named sponsor, preregister a non-awarding shadow study of 5–15 archived projects. Compare the historical decision, an endurance-only ranking, and the full capped and lifecycle-weighted ranking using coded records and held-back time-series segments. Measure rank uncertainty, missingness, reconstruction labor, incumbent effects, audit reversals, appeal burden, and association with later demonstration outcomes. Stop if fewer than 80% of histories are reconstructable, measurement uncertainty exceeds project differences, reasonable weights reverse the rankings, the full design fails to beat endurance alone, or governance cost is disproportionate. No new synthesis, testing, funding change, or award follows.","deployment_and_cost":"The retrospective evidence study is estimated at $50,000–$250,000. Initial setup is $250,000–$1 million; operational launch and annual operation are each roughly $1–$5 million in 2026 resource-equivalent terms. These are not quotes. Reaction-specific protocols, laboratory capacity, prize administration, purse size, environmental review, confidentiality, and demonstration scope remain unpriced.","risks_and_uncertainties":["Hidden test sequences may reward resistance to surprise conditions rather than representative durability.","Interlaboratory or segment variability may be larger than genuine differences among catalysts.","Resource caps could favor teams that already own equipment, datasets, or precursor inventories.","Broad accounting may expose confidential operations; narrow accounting may shift spending off book.","Institutional cleanup guarantees could exclude capable teams without wealthy sponsors even when their methods are safe scores can obscure tradeoffs among activity, selectivity, endurance, scarce materials, energy, and waste. A sponsor, catalyst metrologist, independent laboratory, safety specialist, competition designer, and research-accounting expert must determine whether an auditable comparison is feasible."],"expert_types":["Catalysis scientist with durability-testing expertise","Nanomaterial measurement and interlaboratory-validation specialist","Independent validation-laboratory operator","Chemical safety, exposure, and waste specialist","Federal prize, procurement, or research-competition counsel"],"expert_questions":["Does a named sponsor actually make a scarce follow-on decision using publication priority or peak team-reported results?","Can archived records distinguish every attempted preparation and run from the selected results presented to the sponsor?","What target reaction, benchmark, operating envelope, and interlaboratory error define a fair endurance comparison?","Which resource categories can be audited without unfairly favoring incumbents or exposing protected information?","Does the full ranking predict later demonstration outcomes better than endurance-only scoring across preregistered weights?"] ,"ranking_note":"The post-hoc harmonized reading aid assigns band C and profile ranks from 31 to 50. It does not upgrade the EMPIRICAL_PARTNER_CANDIDATE endpoint, measure economic value, or establish elapsed pilot speed; affordability only proxies that input.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-PARTNER-16","plain_language_title":"Send Foresight Surprises, Keep Full Records","one_sentence_summary":"Distributed scanners would send structured differences from a shared scenario expectation while retaining full evidence, protected-signal bypasses, independent audits, and automatic return to complete-packet review when the filter becomes unreliable.","problem_plain":"A distributed foresight network may send complete weekly assessments for every monitored driver, forcing central analysts to reread expected material to find a few assumption-breaking observations. A simple report-by-exception rule is unsafe because silence could mean stability, missing reporting, or model blindness. The proposed setting still lacks packet-level evidence that routine material actually exhausts review capacity, that important signals arrive late for this reason, or that downstream decision blockage is not the real bottleneck.","proposal_plain":"Before each weekly intake, scenario stewards publish versioned expectations for each driver’s direction, pace, geography, actors, cross-impacts, uncertainty, scope, and expiry. Local analysts retain full evidence capsules but encode structured differences such as acceleration, reversal, a new actor, altered coupling, or broken assumption. A gate prioritizes these residual cards by reliability, uncertainty, consequence, perspective coverage, and review cost. Heartbeats distinguish expected conditions from missing reports, while protected safety-, rights-, conflict-, and distributional-harm signals travel in full. Central analysts reconstruct the assessment from the matching expectation and residual; validated surprises receive an owner but do not automatically change scenarios. Independent raw sampling, periodic reconciliation, drift monitoring, and version, error, or missingness triggers restore full-packet review.","transfer_plain":"Predictive-residual processing sends deviations from a synchronized expected state instead of retransmitting the entire state. Here, the expected state is a versioned scenario-conditioned driver assessment, and the residual describes model-relative change. Reconstruction, heartbeats, raw audits, protected bypasses, and fallback make the mapping substantive, though translating ambiguous foresight narratives into reliable residual fields remains an unresolved weakness.","why_it_advanced":"This did not qualify as an Experiment 6 strict success. It entered the EMPIRICAL_PARTNER_CANDIDATE lane because a shadow comparison on an adopter’s historical corpus is bounded and measurable. The essential field evidence—actual backlog prevalence, missed-signal causes, taxonomy reliability, protected-signal recall, and fully loaded labor savings—is still missing.","prior_art_and_open_claim":"Scenario-guided scanning, indicator-based early detection, distributed scanning networks, short-list filtering, collaborative platforms, ordinary tags, anomaly alerts, and periodic scenario refreshes are adjacent prior art. The narrower open claim is that a synchronized expectation-plus-residual workflow can reduce total review effort versus complete packets, tagged triage, and existing signpost filtering without materially reducing detection of assumption changes, cross-driver effects, novel actors, or protected-source evidence once preparation, audits, reconciliation, and fallback are counted.","test_and_decision":"Obtain 200–500 authorized historical packets with timestamps. Using only earlier material, freeze expectations for each evaluation window and randomly assign blinded reviewers to complete packets, tagged complete packets, or residual cards with logged fallback. Preregister review-time savings, reconstruction agreement, recall margins, and equal protected-source recall; include raw audits and injected version mismatches and missing heartbeats. Reject the workflow if any safety-class item is suppressed, reconstruction breaches tolerance, underrepresented sources fare worse, fallback fails, systematic residual structure persists, or expectation preparation, auditing, reconciliation, and fallback erase the labor saving.","deployment_and_cost":"First evidence is estimated at $10,000–$50,000. Initial deployment is $50,000–$250,000, operational launch $250,000–$1 million, and annual recurring work $50,000–$250,000 in rough 2026 resource-equivalent bands, not quotes. Taxonomy design, expectation maintenance, adjudication, audit independence, access controls, integration, training, and fallback workload may dominate software costs.","risks_and_uncertainties":["Shared expectations could become a self-confirming filter that suppresses evidence outside current scenarios.","Precision or consequence weights may discount unfamiliar regions, disciplines, sources, or minority perspectives.","A heartbeat may be recorded even when the underlying source network has quietly failed.","Residual cards may remove narrative context needed to interpret ambiguous long-range evidence.","Analysts may change classifications to attract central attention or avoid scrutiny expected-state versions, and protected bypasses remain independent of the scenario stewards? Foresight-method, information-retrieval, human-factors, rights-impact, records-governance, and adopter workflow expertise are needed."],"expert_types":["Government or corporate foresight program leader","Scenario-planning and horizon-scanning methodologist","Information-retrieval or human-in-the-loop systems researcher","Rights, conflict, and distributional-impact specialist","Records, privacy, and access-control officer"],"expert_questions":["Do complete packets currently exceed a declared review budget, and which material findings were delayed specifically by intake saturation?","Can independent annotators reliably encode and reconstruct the proposed driver-state and residual fields?","Which evidence classes must bypass filtering regardless of predicted relevance or apparent routine status?","Does the residual workflow preserve protected and underrepresented-source recall under blinded adjudication?","After counting preparation, maintenance, audit, reconciliation, and fallback, is total analyst time at least meaningfully lower than complete review?"] ,"ranking_note":"The harmonized ordering is a post-hoc reading aid: band C, with profile ranks from 41 to 45. It neither changes the EMPIRICAL_PARTNER_CANDIDATE status nor measures economic value; the pilot-speed input is an affordability proxy, not elapsed time.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]}]}