{"dossiers":[{"portfolio_id":"EXP06-STRICT-14","plain_language_title":"Retiring Outdated Criminal-Record Labels","one_sentence_summary":"A voluntary quarterly exercise would help record custodians distinguish an ended authorization from deletion, preserve lawful history, and assign unresolved downstream corrections to named owners.","problem_plain":"A court or other competent authority may restrict or end the permitted use of a criminal-justice label, yet copies, vendor feeds, access permissions, derived flags, and staff assumptions can persist across disconnected systems. Correcting the source record does not necessarily tell every downstream custodian what changed. The result can be continued reliance on a superseded status; however, careless cleanup can also destroy records that must lawfully remain available for provenance, oversight, or defined exceptions.","proposal_plain":"Authorized custodians would hold a quarterly, voluntary “Label Sunset” cycle after counsel confirms which disposition categories are eligible. Using only invented records, participants would move a fictional label through a map of source, repository, vendor, and user systems. A records witness would verify a mock retirement instruction, remove the label from the active layer, and place it in an archive sleeve marked “retired—provenance preserved.” Participants could decline any part of the exercise. The group would recognize only verified batch-level reconciliation and route every unresolved exception category into protected workflows with an owner, resources, and deadline. An immediate debrief and independent audit could revise, pause, or end the cycle. Unlike an ordinary reconciliation briefing, the proposal adds a witnessed enactment, recommitment, recognition, and explicit exception handoff; it changes no legal status or production record itself.","transfer_plain":"The ritualized meaning-and-commitment archetype becomes a governed enactment of how a label spreads and loses operational authority. Its marked object, witnesses, voluntary commitment, closure, debrief, and retirement path map clearly to the domain, while real legal powers, system privileges, and individual remedies remain entirely outside the ritual.","why_it_advanced":"This was a STRICT_SUCCESS because it passed Experiment 6’s strict researched-candidate bar. That endpoint reflects the quality and testability of the researched proposal, not field validation, novelty, legal authorization, adoption, or demonstrated effects on record accuracy, employment, stigma, recidivism, or other real-world outcomes.","prior_art_and_open_claim":"Legal relief workflows, automated access and retention rules, reconciliation audits, privacy training, and record-clearing programs already perform much of the practical work. The narrower open claim is that adding a voluntary fictional propagation exercise, witnessed retirement, recommitment, aggregate recognition, exception routing, and debrief to a matched reconciliation briefing will improve accurate legal distinctions, delayed recall, and assignment of downstream exceptions without increasing coercion, stigma, privacy risk, or beliefs that the ceremony itself completed legal relief. There is no field evidence for that comparison.","test_and_decision":"A court or repository unit would recruit 24–40 volunteers from at least four custodial functions and randomize intact teams to the 45-minute cycle or a content-, facilitator-, and time-matched briefing. Blind scoring would cover six invented scenarios before, immediately after, and two weeks later. Advance only if the cycle gains at least 0.4 standard deviations on the delayed composite or 15 percentage points in fully correct exception routing, with no material safety loss. Retire the mechanism if neither threshold is met, real records enter the exercise, more than 10% feel dissent is unsafe, or false-relief beliefs exceed the limit.","deployment_and_cost":"The authorized first step is a synthetic pilot with invented cases and no production queries or changes. Rough 2026 resource-equivalent bands are $10,000–$50,000 for first evidence, $50,000–$250,000 for startup and operational launch, and $10,000–$50,000 annually. These are assessment bands, not vendor quotes; no institution has committed funding or adoption.","risks_and_uncertainties":["Participants may mistake symbolic retirement for sealing, deletion, expungement, or complete downstream correction.","Small aggregate exception categories could expose protected case information.","Staff may experience participation, affirmation, or dissent as an employment loyalty test.","Recognition could create ceremonial closure while vendor copies, derived flags, or informal assumptions persist.","The archive metaphor could encourage destruction or concealment of records that must lawfully remain preserved or accessible under an exception."],"expert_types":["Criminal-records counsel","Court or repository data-governance lead","Privacy and information-governance specialist","Background-screening systems expert","Lived-experience or record-subject advocate"],"expert_questions":["Can the six scenarios reliably distinguish correction, sealing, restricted use, deletion, retention, and lawful exceptions?","Which custodian has authority to own each downstream exception without exposing protected case data?","Would staff reasonably perceive passing, dissenting, or challenging the script as safe and nonretaliatory?","Can the matched briefing contain exactly the same legal and system information so the enactment is the only material difference?","What evidence would show that the ceremony is hiding incomplete vendor or repository reconciliation?"],"ranking_note":"The harmonized review placed it 31st–33rd, band C, across post-hoc profiles. That score only sets reading order and uses an affordability proxy; it is neither an experimental endpoint nor a measure of economic value, deployment readiness, or impact.","source_ids_used":["S1","S2","S3","S4","S5","S6","S8"]},{"portfolio_id":"EXP06-PARTNER-04","plain_language_title":"Stress-Testing Flight-Control Downselects","one_sentence_summary":"A shadow contest would compare whether hidden scenarios, safety gates, resource caps, and two-candidate advancement select flight controllers more consistently than public benchmarks or expert panels.","problem_plain":"Aircraft programs must decide which competing flight-control candidates receive scarce hardware-in-the-loop and flight-test resources. If teams know every scored simulation case, they may tune to those cases, simulator artifacts, or inferred test structure, while escalating compute and engineering effort. A visible-score winner may therefore be less reliable under new conditions than its rank suggests. The actual frequency of this problem in flight-control programs is unknown, and general competition research supplies meaningful counterevidence against assuming severe leaderboard overfitting is routine.","proposal_plain":"The program would freeze eligibility, configuration deadlines, allowed resources, safety floors, scoring, audit rules, fouls, appeals, and judge authority before submissions. Teams could develop against a disclosed suite, but an independent custodian would rank frozen builds using undisclosed seeds and disturbance combinations. Every controller would first face noncompensable safety gates. Passing candidates would then be scored for hidden-condition robustness, simulated pilot-intervention demand, control activity, reproducibility, and declared integration burden under resource caps. Two materially different candidates, rather than one leaderboard winner, would receive hypothetical advancement in the first shadow study. Independent reproduction and later hardware-in-the-loop, pilot, and maintainer evidence would govern any real progression. The comparator includes a fully public single-winner benchmark and an unranked expert-panel choice using the same builds, scenario families, safety constraints, and nominal compute limits. No contest score authorizes flight.","transfer_plain":"The bounded-rivalry archetype becomes a controlled engineering contest with a scarce prize, fixed legal moves, safety floors, resource limits, audits, appeals, multiple winners, and a later challenger window. The mapping is strong, but most individual governance features already appear in aviation challenges, technical competitions, or staged verification programs.","why_it_advanced":"This is an EMPIRICAL_PARTNER_CANDIDATE, not a strict success. It cleared a separately calibrated lane for a bounded external data-partner study because the comparison is testable, but it lacks proprietary frozen controllers, program-specific scenarios, validated resource accounting, downstream hardware evidence, and any partner commitment.","prior_art_and_open_claim":"Robust flight-control challenges, protected leaderboards, aviation contest rules, safety stop authority, formal protests, and staged simulation-to-hardware validation already exist. The remaining claim concerns their combination: holding builds and approved conditions constant, a hidden, safety-gated, resource-capped arena that advances two complementary candidates will produce more stable selections on independently seeded validation and later authorized hardware or human evaluation than either a public-suite winner or an expert-panel choice. That comparative effect has not been measured, and portfolio discretion could simply advance a favored lower scorer.","test_and_decision":"With an aircraft-program partner, preregister an offline shadow study using three to six frozen controllers. Apply the proposed arena, the disclosed single-winner benchmark, and an unranked panel to identical approved materials, then test their selections on a second inaccessible scenario set. Measure selection agreement, rank stability, safety failures, reproducibility, intervention demand, actuator activity, resources, disputes, and sensitivity to lawful seeds and weights. Do not seek hardware authority unless the arena is more stable or predictive. Falsify it if innocuous seeds reverse choices, results do not beat both comparators, portfolio selection adds an inferior redundant candidate, or resource accounting systematically favors incumbents. Void results after leakage, unverifiable builds, compensable safety violations, or disputed simulator validity.","deployment_and_cost":"The first step is offline shadow evaluation with no contract, integration, certification, personnel, or flight consequences. Rough 2026 resource-equivalent bands are $50,000–$250,000 for first evidence, $250,000–$1 million for startup, $1–$5 million for operational launch, and $250,000–$1 million annually. They are not facility or supplier quotes.","risks_and_uncertainties":["Secret scenarios may contain the same modeling errors as the disclosed simulator while making those errors harder to challenge.","Existing tools and pretrained models could let incumbents exceed the practical resource cap without recorded spending.","A discretionary definition of “complementary” could justify advancing a favored lower-scoring controller.","Teams may tune to the inferred scenario generator rather than to operationally relevant robustness.","Simulation rank may fail to predict hardware behavior, pilot workload, maintainability, integration effort, or certification findings."],"expert_types":["Flight-control engineer","Aircraft verification and validation lead","Test pilot or pilot-in-the-loop specialist","Airworthiness and certification specialist","Simulation-model validation expert"],"expert_questions":["Which approved scenario variations are independent enough to test robustness without leaving the simulator’s validity envelope?","Can legacy tools, reusable models, supplier labor, compute, and sponsor support be counted fairly across entrants?","How should complementarity be defined before scores are known so it cannot become discretionary favoritism?","What rank-stability or downstream-prediction improvement would justify the added contest administration?","Which results may inform a later hardware study without being mistaken for certification evidence or flight authorization?"],"ranking_note":"The post-hoc harmonized profiles rank this candidate from 22nd to 40th, band C. The wide range reflects different review weights. It is a reading-order aid using a cost proxy, not an experimental endpoint, economic-value estimate, or partner commitment.","source_ids_used":["S1","S2","S3","S5","S6","S7","S8"]},{"portfolio_id":"EXP05-STRICT-03","plain_language_title":"Tracking Aging Wetland Erosion Mats","one_sentence_summary":"A site register would identify successive erosion-control layers, trigger review as they age, block unsafe removal, and preserve a record after authorized retrieval.","problem_plain":"Wetland and streambank repairs can leave several generations of blankets, coir mats, synthetic netting, and anchors in the same location. Older layers may be exposed, buried, torn, rooted through, or missing from project records. Crews may cover an unidentified layer again, leaving fragmenting mesh in habitat, or remove material that still holds vegetation and sediment. The field prevalence of such sequential legacy layers is not yet established for a candidate site, and hydrology or geomorphology may matter more than installed material.","proposal_plain":"At one bounded restoration reach, managers would give each mapped installation a stable identity, footprint, material class, deposition order, estimated age, condition, access history, and lifecycle state. A class-specific time-to-review trigger would mark a layer for inspection, never automatic removal. An age-weighted value-and-risk score would order limited field work, but qualified reviewers would first check whether roots, sediment, adjacent structures, monitoring records, or permits still depend on the material. The responsible manager could continue use, schedule another inspection, retire the layer in place, preserve it as a time-limited exception, or seek separate permission for staged retrieval. Any removal would leave a geospatial “tombstone” recording the former footprint, reason, evidence, and successor layer. The comparator is ordinary project-file review plus current surface inspection without unified identities, expiry triggers, dependency gates, exception states, or tombstones.","transfer_plain":"The layer-decay archetype maps directly onto successive physical installations that can outlive their original role. Stable identities, review deadlines, risk ordering, dependency checks, differentiated disposition, expiring exceptions, and removal records translate cleanly. The analogy does not make age proof of obsolescence or turn a database state into authority to disturb habitat.","why_it_advanced":"This was a STRICT_SUCCESS because it passed Experiment 5’s strict researched-candidate bar. That status does not show that multiple legacy layers are common at real reaches, that reviewers can classify them reliably, that an adopter will use the register, or that it improves bank stability, wildlife outcomes, or costs.","prior_art_and_open_claim":"Existing practice already includes safer material substitution, functional-longevity classes, inspection and maintenance records, removal requirements, permits, monitoring plans, and asset databases. The narrower open claim is that adding installation-level identities, review-only expiry triggers, a mandatory ecological, structural, and regulatory dependency check, time-limited exceptions, and post-removal tombstones will yield more reproducible and dependency-resolved recommendations than ordinary files and surface inspection. It expressly does not claim that older material should be removed or that the register itself reduces erosion, plastic pollution, or wildlife mortality.","test_and_decision":"At a partner-managed reach, sample 20–50 mapped or suspected installations. Two qualified reviewers would independently assess the same nondestructive evidence under randomized workflows: ordinary files plus surface inspection, and the same material augmented by the lifecycle register and dependency checklist. Compare unique layers found, unresolved records, agreement, review time, live dependencies, and actionable recommendations. Stop at observation; do not probe or remove. Falsify the near-term claim if no sequential accumulation exists, the added workflow finds no decision-relevant layers or dependencies, agreement remains below a preregistered threshold such as kappa 0.60, review time more than doubles without better resolution, or either reviewer recommends retrieval despite a live or unresolved dependency.","deployment_and_cost":"Begin with a read-only inventory using existing plans, photographs, permitted observation, and provisional uncertainty labels. Rough 2026 resource-equivalent bands are $10,000–$50,000 for first evidence, $50,000–$250,000 for startup and operational launch, and $10,000–$50,000 annually. These are not quotations; no site partner or data-stewardship arrangement is secured.","risks_and_uncertainties":["Surface observation may miss buried or visually similar layers and create false confidence in completeness.","An age-weighted score may disguise subjective judgments as measurement precision.","Staff may treat a review deadline as an instruction to remove material despite the dependency gate.","“Retire in place” may become indefinite neglect if its exception and next review date are not enforced.","Attention to installed mats may divert investigation from hydrologic or geomorphic causes of site failure."],"expert_types":["Wetland or stream restoration ecologist","Geotechnical or hydraulic engineer","Environmental permitting specialist","Erosion-control contractor","GIS and restoration-records manager"],"expert_questions":["Can reviewers identify deposition order and lifecycle state nondestructively and with acceptable agreement?","Which observed conditions are sufficient to establish that roots, sediment, structures, or permits still depend on a layer?","How should unknown installation dates and incomplete specifications affect priority without implying obsolescence?","What authority and evidence are required before staged retrieval at the selected reach?","Does the register add useful information beyond the site’s existing maintenance, permit, monitoring, and asset records?"],"ranking_note":"The post-hoc harmonized profiles placed the candidate 35th–37th, band C. This narrow reading-order range uses an affordability proxy and does not alter its strict experimental endpoint or measure field prevalence, ecological benefit, economic value, or deployment readiness.","source_ids_used":["S1","S2","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-PARTNER-13","plain_language_title":"Subtracting Recoater Self-Noise","one_sentence_summary":"A shadow monitoring system would predict signals caused by a powder recoater’s own commands, route unexplained residuals for review, and revert to complete signals whenever validity or safety checks fail.","problem_plain":"During powder-layer recoating, commanded acceleration and routine powder contact create large, repeatable force, current, vibration, and acoustic signals. These self-generated patterns can conceal smaller signs of an agglomerate, protrusion, foreign object, uneven layer, or changed powder flow. Independent static thresholds can instead produce repeated benign alerts. It is not yet known whether command-caused motion explains enough signal on a target machine to create a useful residual, or whether complete raw monitoring already fits the available data and operator-attention capacity.","proposal_plain":"Each outgoing motion command would be copied into a frozen, versioned forward model that predicts the synchronized sensor response expected for one declared machine, recoater, powder class, speed range, and environment. The system would subtract that prediction from the measured full window, preserve signed and time-aligned residuals, and weight them by sensor reliability, timing, coherence, persistence, consequence, and processing cost. A compatible receiver could reconstruct the full signal from the prediction and residual. Validated residual clusters would route to an operator with a defined inspection action. Random and risk-based raw windows, periodic full-state anchors, checksums, heartbeats, and drift tests would audit the cancellation. Any timing, model, sensor, reconstruction, scope, or protected-safety failure would restore complete processing. Model updates would remain slower, reviewed, and reversible so recurring defects could not immediately be learned away. Comparators are full raw review and existing static thresholds.","transfer_plain":"Predictive residual processing becomes an efference-copy system: the controller’s own command predicts the machine-generated sensory return, and the unexplained remainder receives scarce transmission and operator attention. The structural mapping is strong, but adjacent disturbance observers, recoater sensors, digital twins, anomaly detectors, and predictive codecs already supply most components separately.","why_it_advanced":"This is an EMPIRICAL_PARTNER_CANDIDATE, not a strict success. It cleared a separate lane for a bounded external laboratory study, but lacks a target-machine dataset, frozen model, controlled challenge results, comparative reviewer evidence, measured capacity savings, validated fallback reliability, and an OEM or laboratory commitment.","prior_art_and_open_claim":"Recoater force and vibration sensing, powder-bed imaging, digital-twin recoating control, model-based collision detection, industrial interfaces, and bounded powder-spreading testbeds already exist. The remaining claim is narrower: on one fixed configuration, command-conditioned cancellation with reconstruction, raw audits, version checks, and forced fallback will detect and localize specified external interactions no worse than complete raw review while reducing total data and reviewer burden after computation, audits, fallbacks, inspections, and maintenance are counted. No live recoater experiment establishes that combined comparison, especially for events synchronized with commanded acceleration.","test_and_decision":"An OEM or accredited laboratory would run a preregistered shadow test on one fixed recoater, sensor layout, motion program, medium, and environment while retaining every raw channel. Blinded challenges would include localized resistance, altered layer height, obstacle surrogates, powder-drag changes, sensor bias, timing offsets, dropped residuals, version mismatch, and missing heartbeats. Compare the reconstructed residual path, complete raw review, and static thresholds on protected-event capture, classification, localization, false alerts, reconstruction, fallback behavior, bytes, review time, inspections, and maintenance. Reject advancement if any protected or synchronization challenge fails to trigger full mode, any controlled interaction is over-cancelled, sensitivity or localization breaches the preregistered non-inferiority margin, or total capacity cost is not lower.","deployment_and_cost":"The first authorized step is a non-production shadow test; existing controls, interlocks, raw storage, and operator decisions remain authoritative. Rough 2026 resource-equivalent bands are $50,000–$250,000 for first evidence, $250,000–$1 million for startup, $1–$5 million for operational launch, and $250,000–$1 million annually. They are not OEM or supplier quotes.","risks_and_uncertainties":["A wrong or mistimed model may subtract a real interaction that resembles the expected response to acceleration.","Powder lot, humidity, wear, mounting, or sensor coupling may invalidate the learned self-signal.","Sender and receiver can share the same wrong model while still passing a compatibility checksum.","Rapid model adaptation could normalize recurring defects or wear instead of exposing them.","Audit sampling may miss rare over-cancelled events, while frequent fallbacks may increase workload and pressure operators to weaken safeguards."],"expert_types":["Additive-manufacturing process engineer","Machine-controls engineer","Recoater and powder-spreading specialist","Industrial sensing and signal-processing researcher","Machine safety and functional-safety engineer"],"expert_questions":["What fraction of each target sensor’s variance is reproducibly explained by the outgoing command under nominal conditions?","Which controlled interactions are most likely to resemble command-caused acceleration signals and be over-cancelled?","What non-inferiority margin is acceptable for detection and localization against complete raw review?","Can clocks, position traces, checksums, heartbeats, and full-state anchors force fallback within the required latency?","After computation, raw audits, fallbacks, inspections, and maintenance, does the residual path actually reduce total capacity use?"],"ranking_note":"Post-hoc harmonized profiles rank this candidate from 17th to 48th, band C, showing strong sensitivity to review priorities. The score is only a reading-order aid with a cost proxy, not an endpoint, economic-value measure, validation result, or deployment recommendation.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]}]}