{"dossiers":[{"portfolio_id":"EXP06-STRICT-03","plain_language_title":"Common-Wafer Contest for One Fabrication Slot","one_sentence_summary":"A nanofabrication facility would compare teams on blinded, equally resourced test wafers before awarding its single process-integration slot.","problem_plain":"A shared nanofabrication facility has one integration bay and limited technician time. Teams now compete largely through proposals and results from samples they chose themselves. That can reward unusually favorable devices, extensive private characterization, or incomplete reporting of failed runs. It can also leave the facility and other users bearing contamination, waste, cleanup, and downtime. The facility may therefore select a persuasive process that cannot reproduce safely and reliably on shared equipment.","proposal_plain":"Replace proposal-only selection with a rule-bound common-wafer trial. Before entrants are known, publish eligibility, identical substrates, allowed process steps, equal tool and measurement budgets, safety gates, scoring, tie-breaks, confidentiality, penalties, and appeals. Code the samples and score complete attempt histories, functional yield across wafers, dimensional and electrical reproducibility, equipment compatibility, waste, cleanup, and recovery time. Unsafe performance cannot be offset by technical strength. Independently remeasure the leaders and audit their resource use. Give two teams bounded validation runs before awarding the single integration bay. Afterwards, compare trial scores with actual integration performance, examine whether the winner gained control over shared interfaces or rules, and reopen competition through a scheduled challenger window.","transfer_plain":"The bounded-rivalry archetype becomes a deliberately governed contest for a genuinely scarce facility slot. Common specimens and resource caps define the arena; blinded scoring, audits, safety floors, spillover accounting, penalties, appeals, post-cycle review, and later challenger access constrain how teams may compete and what winning confers.","why_it_advanced":"It passed Experiment 6's strict researched-candidate bar because the allocation problem, responsible authorities, testing capabilities, safeguards, comparator, and falsifiable shadow study were sufficiently specified. That status concerns the researched candidate only; it does not establish field performance, novelty, deployment authority, adopter demand, or economic benefit.","prior_art_and_open_claim":"Proposal review, safety qualification, common-specimen comparisons, metering, contamination controls, equipment standards, and post-selection oversight already exist, so the proposal sits next to substantial prior art. The narrower open claim is that the full score—held-out reproducibility, all attempts, equal counted resources, tool compatibility, and recovery burden—will predict audited performance and avoid ranking reversals better than both proposal-only review and a simpler common-wafer technical score.","test_and_decision":"With facility, safety, and data approval, run a non-awarding shadow study lasting at most 12 weeks and costing at most $50,000. Require at least eight archived entrants from two process families, comparable retained wafers or replicate measurements, and usable operational records. Freeze the rubric and audit split before revealing identities or existing ranks. Advance only if the full score improves held-out Kendall rank correlation over proposal review by at least 0.15, reduces audit reversals by at least 20%, is no worse than technical-only scoring, avoids a leader-changing process-family interaction, keeps essential missingness below 20%, and causes no safety or confidentiality breach.","deployment_and_cost":"The first shadow evidence step is estimated at $10,000–$50,000 in rough 2026 resource-equivalent terms, not a vendor quote. Initial deployment is $50,000–$250,000; operational launch and annual operation are each $250,000–$1 million. Real use would also require local authority over recipes, records, appeals, reserves, and penalties.","risks_and_uncertainties":["Common wafers may favor one process family or poorly represent integration conditions.","A frozen score may encourage teams to optimize measured proxies instead of robust integration performance.","Equal in-contest resource caps may still favor teams with stronger infrastructure outside the counted arena.","Recipe submission and access-log audits may expose confidential know-how or identifiable personnel data.","A remediation reserve could exclude less-capitalized teams or exceed the facility's legal authority to collect it. Managers have not established the needed terms yet, and the effects on those teams require field data. No source establishes how complete the facility's records are or how frequently proposal winners fail shared-condition replication. No facility has committed to host the study, and site-specific technician capacity and opportunity costs remain unknown."],"expert_types":["Nanofabrication process-integration engineer","Facility operations and metrology manager","Environmental health and contamination-control specialist","Research-allocation and data-governance counsel"],"expert_questions":["Can retained wafers from at least eight entrants be compared without introducing process-family-specific measurement bias?","Which waste, cleanup, downtime, and compatibility measures can be reconstructed reliably from existing records?","Would the proposed score predict successful integration better than technical yield and reproducibility alone?","What authority does the facility have to inspect recipes, hear appeals, impose penalties, or require a remediation reserve?"],"ranking_note":"The harmonized score is only a post-hoc reading-order aid: band C, ranking between 20 and 29 across profiles. Its pilot-speed input reflects cost-band affordability, not independently measured elapsed time or economic value, and it does not alter strict-success status.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-STRICT-08","plain_language_title":"A Stable Contract for Crystal Structures","one_sentence_summary":"An opaque software contract would let materials tools exchange the same ordered periodic structure without depending on atom order, file layout, or a particular canonicalization method.","problem_plain":"Materials software passes crystal structures among parsers, databases, simulation tools, and analysis programs. Clients often depend on incidental details such as atom-array order, coordinate wrapping, lattice orientation, cell convention, or serialized field layout. The same physical ordered structure can then receive different identifiers or downstream treatment, while an internal parser or storage change can break clients even when the intended material state has not changed. It also becomes difficult to distinguish a physical change from an encoding change.","proposal_plain":"Define an immutable ordered-periodic-material value by its periodic decorated points in physical space, including species, occupancies, units, tolerance policy, and provenance. Expose validated construction, composition, physical-equivalence tests, invariant summaries, explicit conversions, and serialization, but hide arrays, atom order, coordinate basis, caches, canonical labels, and format-specific fields. Specify preconditions, typed errors, no-mutation-on-failure behavior, and laws for atom permutation, origin shifts, periodic wrapping, unit conversion, and admissible cell changes. Every parser or store must pass the same public-only black-box and metamorphic tests before substitution. Version changes to promised behavior and probe error text, ordering, serialization, debug access, and timing for leaks. Version zero rejects disorder, partial occupancy, trajectories, surfaces, and inferred bonding rather than silently approximating them.","transfer_plain":"The representation-independent-interface archetype maps strongly here. The abstract component is the physical ordered periodic structure; concrete arrays, graphs, files, and database rows remain hidden. Behavioral laws, typed failures, shared conformance tests, leakage review, and versioning determine whether independently built implementations may substitute for one another.","why_it_advanced":"It passed Experiment 6's strict researched-candidate bar because close comparators, a bounded scope, scientific authority, reversible sandbox, measurable failure conditions, and a live two-adapter benchmark were identified. This endpoint does not show that production repositories suffer widespread coupling or that the contract is novel, performant, deployable, or adopted.","prior_art_and_open_claim":"Representation-insensitive structure matching, pymatgen, spglib, CIF, OPTIMADE, and unified materials-data systems already address important parts of the problem. The remaining claim is narrower: two independently structured adapters can obey one public behavioral oracle for equivalence, invariants, errors, immutability, and provenance while leaking fewer internal details and producing fewer semantic disagreements than the existing array baseline, a canonicalization path, or a schema-only round trip.","test_and_decision":"With a platform partner and scientific-method lead, preregister the abstract state, tolerance rules, observables, and exclusions. Test two independently structured adapters on 12 copied fixtures, five seeded representation changes per fixture, fixed assertions, and 200 operation sequences. Compare them with the existing array/serializer behavior, spglib plus pymatgen, and an OPTIMADE/CIF round trip. Reject version zero if any meaning-preserving transformation changes a promised observable, any fixture triple exposes tolerance-driven non-transitivity, provenance is lost, hidden representation must be exposed, or the contract fails to reduce semantic divergences and leakage relative to the best comparator. Also record runtime and memory.","deployment_and_cost":"The read-only sandbox is estimated at $10,000–$50,000 in rough 2026 resource-equivalent terms. Initial deployment is $50,000–$250,000; operational launch is $250,000–$1 million; annual operation is $50,000–$250,000. Production identifier changes, deduplication, parser replacement, or database migration are expressly outside the first test.","risks_and_uncertainties":["Tolerance-based equivalence may be non-transitive, producing unstable identity groups.","The initial abstract state may omit meaningful distinctions such as defects, chirality, magnetic order, isotope labels, or provenance.","An over-specified suite could freeze incidental numerical or serialization behavior.","An under-specified suite could pass simple structures while missing difficult-cell disagreements.","Opaque access may obstruct legitimate diagnostics and encourage unsupported escape hatches. Real repository data have not yet shown how frequent or costly representation coupling is. Scientific reviewers have not approved the proposed version-zero semantics, and comparator performance remains unmeasured. Passing finite tests would not establish chemical identity, scientific equivalence, implementation correctness, or acceptable production performance."],"expert_types":["Computational crystallographer","Materials-data platform architect","Scientific software testing specialist","Materials provenance and standards expert"],"expert_questions":["Does the proposed abstract state preserve every scientifically relevant distinction in the selected ordered-periodic scope?","Can the tolerance policy avoid non-transitive equivalence for realistic near-boundary structures?","Which current clients depend on atom order, serializer layout, canonical labels, or other hidden details?","Does the contract reduce semantic divergence and leakage without unacceptable runtime, memory, or diagnostic costs?"],"ranking_note":"The harmonized result is a post-hoc ordering aid: band C, ranks 26–30 across profiles. It is not an experimental endpoint or economic-value measure; its pilot-speed input is an affordability proxy, and the candidate remains a strict success only under Experiment 6's researched bar.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-PARTNER-17","plain_language_title":"Show Reviewers Only Unexpected Proof Effects","one_sentence_summary":"A complete proof-library check would be reconstructed from frozen theorem-level predictions and typed discrepancies, allowing reviewers to focus on surprises while protected changes always remain visible.","problem_plain":"Revising an axiom, definition, notation rule, or trusted dependency can affect hundreds of machine-checked theorems. A checker can revalidate the whole library, but maintainers may still face a large migration report. Dependency graphs can flag harmless reachability and miss effects from elaboration, automation, or undeclared coupling. Reading every expected result wastes attention, yet showing only predicted failures could hide an unexpected survival, a new dependency, a removed obligation, or a theorem that still passes for the wrong reason.","proposal_plain":"Before migration, freeze a versioned model predicting every theorem's status, changed obligations, and dependency differences. Independently run the trusted checker over the complete authorized corpus. Compare each prediction with the actual result and record typed residuals for unexpected failures or survivals, changed diagnostics, obligations, dependencies, trust assumptions, timeouts, missing results, or confidence disagreements. Reconstruct the full ledger from prediction plus residual, while directing review primarily to consequential mismatches. Statements, axioms, admitted facts, trust changes, removed obligations, checker failures, missing results, and out-of-scope theorems always appear in full. Random predicted-unaffected records and boundary cases receive independent audits. Version, coverage, reconstruction, drift, or audit failures automatically restore the complete theorem-by-theorem report for the affected component.","transfer_plain":"Predictive residual processing becomes a review codec for mathematical migrations. A frozen impact model supplies the expected theorem ledger; complete checking supplies reality; typed differences carry surprises. Checksums, audits, resynchronization, protected-event bypasses, drift monitoring, and raw-report fallback keep compression from becoming selective proof checking.","why_it_advanced":"It did not enter the strict-success lane. It cleared a separately calibrated empirical-partner-candidate lane because a reversible archived replay and decision thresholds are specified, but the decisive field evidence is missing: no measured report burden, predictable-record share, predictor calibration, reviewer study, audit rate, cost benchmark, maintainer funding, or partner authorization exists.","prior_art_and_open_claim":"Static dependency analysis, grouped checker reports, incremental proof checking, iCoq-style regression selection, source diffs, and trust checklists are established neighbors. The open comparison is whether complete independent checking plus frozen predictions, typed residuals, protected full records, random audits, exact reconstruction, and component fallback can reduce report volume and review time without lowering protected-event recall or blinded classification accuracy relative to full and dependency-grouped reports.","test_and_decision":"With maintainer authorization, replay an archived, non-release-blocking migration containing at least 500 declarations. Freeze the model, residual types, protected classes, thresholds, and audit sample before opening the evaluation partition. Randomize blinded reviewers among full reports, dependency-grouped reports, and the residual interface; insert cases covering failures, survivals, dependency and obligation changes, axioms, timeouts, missing outputs, manifest gaps, and version mismatches. Require 100% protected-event recall, zero missing-as-success errors, exact ledger reconstruction, no consequential audit miss, classification no more than five percentage points below the best comparator, and at least 20% lower median review time or report volume. Success permits only another shadow study.","deployment_and_cost":"The archived replay is estimated at $10,000–$50,000 in rough 2026 resource-equivalent terms. Initial deployment is $50,000–$250,000; operational launch is $250,000–$1 million; annual operation is $50,000–$250,000. The model may prioritize review but may never approve revisions, waive checker results, or edit proofs automatically.","risks_and_uncertainties":["A shared predictor could create correlated blind spots across whole theory components.","Unexpected theorem survival may conceal a weakened statement or unintended dependency.","Automation and elaboration may create semantic coupling absent from declared dependency graphs.","The residual schema may preserve checker status but omit mathematical context needed for judgment.","Reviewers may anchor on predictions, while thresholds may be tuned to shrink the queue. No evidence yet quantifies present reviewer burden, repeated predictable content, or protected-event frequency. The theorem-level predictor has not been calibrated across the required outcome types, and no human comparison has tested whether residual presentation preserves decisions. Audit rates, repository confidentiality rules, operating costs, partner authorization, and maintainer funding remain unresolved."],"expert_types":["Formal-mathematics library maintainer","Proof-assistant kernel and elaboration expert","Human-factors researcher for technical review","Statistical audit and anomaly-detection specialist"],"expert_questions":["What fraction of a real foundational migration report is predictable repetition, and how much reviewer time does it consume?","Can the predictor detect unexpected survivals, trust changes, removed obligations, timeouts, and missing outputs with calibrated uncertainty?","What random and boundary-focused audit rate would detect rare consequential blind spots?","Does the residual interface preserve blinded reviewer accuracy while reducing total review, model-maintenance, audit, and fallback cost?"],"ranking_note":"The harmonized score is a post-hoc reading-order aid only: band C, ranks 27–34 across profiles. The pilot-speed input approximates affordability rather than measured duration. This ordering does not convert the candidate into a strict success or indicate economic value.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-STRICT-13","plain_language_title":"A Recurring Reckoning for Airport Noise Promises","one_sentence_summary":"A short, consent-based observance would connect remembered airport-noise experiences and past commitments to an authorized public decision and a concrete follow-up ledger.","problem_plain":"Airports can collect noise measurements, complaints, and public comments while losing a shared memory of what communities experienced and what institutions promised. Staff and resident turnover may scatter commitments across minutes and dashboards. Conventional meetings can repeatedly ask people to recount distress without creating a bounded occasion to acknowledge it, renew or revise a promise, record disagreement, and assign follow-through. Participants may then disagree about a commitment's origin, scope, owner, authority, resources, or next review date.","proposal_plain":"Embed a quarterly 30-minute observance in an existing airport-community forum. A rotating community-institution pair first documents its purpose, participant standing, story and recording provenance, access options, obligations, and retirement rule. A restrained threshold introduces 60 seconds of optional listening to consent-cleared low-intensity audio or viewing an acoustic trace; silence, private reflection, remote attendance, or leaving are equivalent choices. Paired community and institutional accounts then reconstruct one commitment, preserving corrections and disagreement. Consenting witnesses acknowledge the experience and exact promise. An authorized representative must renew, revise, or decline it; community members may dissent or remain silent. Closure updates a public ledger with authority limits, owner, next action, resource dependency, forum, review date, and unresolved disagreement, followed by a harm-and-meaning debrief and periodic independent audit.","transfer_plain":"The ritualized-commitment archetype becomes a marked, recurring accountability occasion rather than an operational aviation decision. Threshold, optional reflection, paired histories, witnessing, explicit institutional recommitment, ledger closure, debrief, rotating stewardship, harm audit, and governed retirement connect symbolic recognition to ordinary authorized follow-through.","why_it_advanced":"It passed Experiment 6's strict researched-candidate bar because the intervention, authority boundary, conventional comparator, consent protections, stopping rules, and matched rehearsal are explicit and testable. This does not show noise reduction, community-wide legitimacy, real commitment fulfillment, adopter willingness, novelty, deployment authorization, or economic impact.","prior_art_and_open_claim":"Noise dashboards, complaint systems, community roundtables, facilitated planning, written agreements, action ledgers, and one-time listening sessions already exist. The remaining claim is incremental: adding the governed threshold, optional sensory interval, paired provenance accounts, witnessing, an explicit renew-revise-decline choice, and immediate debrief will improve accurate, retained understanding of one commitment beyond an equally timed facilitated ledger review, without increasing coercion, distress, false representation, or confusion that acknowledgment equals mitigation.","test_and_decision":"With a willing forum, run a preregistered tabletop and randomized matched rehearsal with 12–20 consenting community and institutional participants, using a fictional commitment and synthetic low-intensity media. Compare the full observance with an equally timed agenda-and-ledger review containing identical facts. Score eight commitment facts immediately and after 7–14 days, plus authority understanding, fairness, comfort, manipulation, representation, disclosure pressure, accessibility, and freedom to opt out. Proceed only if median recall improves by at least two facts, at least 80% understand what was not decided, coercion and representation scores are no worse, and no serious distress or retaliation concern occurs. The observer may require revision or stop.","deployment_and_cost":"First evidence, initial deployment, operational launch, and annual recurring operation are each estimated at $10,000–$50,000 in rough 2026 resource-equivalent terms, not external quotes. The observance cannot change routes, schedules, funding, regulation, environmental findings, or airline operations, and it cannot replace complaints, consultation, mediation, or technical analysis.","risks_and_uncertainties":["The observance could aestheticize residents' distress or turn it into institutional performance.","Audio or repeated recollection could cause sensory or emotional harm despite opt-out choices.","Selected recordings, histories, or visible participants could be treated as representing absent communities.","Ceremonial acknowledgment might be mistaken for mitigation, legal acceptance, or resolution.","An authorized-looking renewal could conceal missing funding, regulatory power, or operational feasibility. No site has shown that commitment-memory failures occur at the assumed rate, and no airport or community group has agreed to host the pilot. Live evidence is absent on nonretaliatory refusal, adverse events, comparative benefit over a disciplined ledger meeting, and whether real representatives can decide without bypassing other authorities. Site-specific costs are also unknown."],"expert_types":["Airport community-engagement and noise-management lead","Affected-community representative with independent standing","Accessibility, trauma-informed facilitation, and emotional-safety specialist","Aviation governance and authority-boundary expert"],"expert_questions":["Do participants currently fail to reconstruct commitment history, scope, authority, ownership, dependencies, and unresolved disagreement from ordinary records?","Does the ritual sequence improve delayed factual recall beyond an equally timed facilitated ledger review?","Can audio-free participation, silence, dissent, and exit remain genuinely nonretaliatory under the forum's power relationships?","Which representative can renew, revise, or decline a real commitment, and which decisions must remain with airports, airlines, regulators, boards, or air-traffic authorities?"],"ranking_note":"The harmonized result is a post-hoc reading-order aid: band C, ranks 23–38 across profiles. It neither measures economic value nor changes the strict-success endpoint; the pilot-speed input is an affordability proxy, and elapsed pilot time was not independently scored.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]}]}