{"dossiers":[{"portfolio_id":"EXP05-STRICT-01","plain_language_title":"Show What Actually Solved the Model","one_sentence_summary":"A passport would distinguish what a macroeconomic policy platform computed itself from results supplied by solvers, data services, or people, while preserving bounded and unresolved outcomes.","problem_plain":"Policy models can contain unrestricted agent programs, continuous calculations, external equilibrium solvers, live data, and human choices among possible equilibria. Yet a final policy ranking may simply say that the model converged. That label can hide who or what supplied the answer, whether the supplier could abstain, whether other trajectories remained unresolved, and whether a finite timeout was mistaken for universal stability. The result may therefore claim more computational certainty than the platform actually established.","proposal_plain":"Before a model enters formal policy comparison, give it a capability-relative guarantee passport. Freeze the model language, economic target, equilibrium and stability definitions, time domain, and quantifiers. Define what the base platform can compute, then register each solver, data source, real-number routine, interactive environment, and human selection step as a named external capability with explicit promises, errors, latency, abstention, accountability, and version behavior. Test whether conclusions survive withdrawal, invalid responses, and promise violations. Search fairly, within declared bounds, for trajectories that refute convergence or target attainment. Route every result as BASE_CERTIFIED, ORACLE_RELATIVE, COUNTEREXAMPLE_FOUND, BOUNDED_ONLY, UNKNOWN, OUT_OF_MODEL, or SYSTEM_FAILURE. A timeout remains UNKNOWN, and a human-selected equilibrium remains assisted. Any claimed unrestricted impossibility reduction must be independently checked before use.","transfer_plain":"The computability-boundary archetype maps strongly here. External solvers, live information, and adaptive human choices act like added computational capabilities. The passport records which question the base program answers, which answer depends on a named capability and its promises, and which questions remain unresolved. Checked countertrajectories can refute precise universal claims without pretending to decide every possible model.","why_it_advanced":"This candidate passed Experiment 5's strict researched-candidate bar because the problem is consequential, the classifications and offline pilot are technically feasible, and existing model-governance functions could own them. That endpoint is only a researched-candidate result: it does not establish real-world effectiveness, novelty, economic value, adopter commitment, or authorization for policy use.","prior_art_and_open_claim":"Model inventories, validation, lifecycle governance, model cards, and documentation of human or third-party limits already exist. Numerical platforms also disclose dependence on initial guesses, iteration limits, and equilibrium selection. The narrower untested claim is that named capability contracts, promise checks, withdrawal tests, and mechanically preserved result labels will catch more capability-laundering errors than an ordinary model inventory and validation package, without changing the economic question being asked. No world-novelty finding was made.","test_and_decision":"Preregister an offline experiment using six synthetic models and stubbed external services. Randomize cases between the ordinary inventory and validation template with raw logs, and the passport workflow. Blinded reviewers classify each result's computational dependency and status. Compare classification accuracy, false BASE_CERTIFIED results, preservation of UNKNOWN, promise-violation detection, review time, and semantic fidelity. Independently check any reduction and countertrajectory. Reject the incremental claim if accuracy does not improve, any false BASE_CERTIFIED label appears, UNKNOWN is lost, withdrawal cannot localize dependency, or at least two independent macroeconomists find that formalization materially changes the intended question.","deployment_and_cost":"The authorized first step is an offline synthetic pilot with no live policy instruments, confidential feeds, forecasts, or production materials. Rough 2026 resource-equivalent bands are $10,000–$50,000 for first evidence, $50,000–$250,000 for startup, $250,000–$1 million for operational launch, and $250,000–$1 million annually. These are assessment bands, not vendor quotes.","risks_and_uncertainties":["The formal equilibrium or target definition may omit what policy staff actually mean to evaluate.","Staff may treat capability labels as paperwork and still collapse assisted, bounded, and unknown results into one ranking.","Analysts may make adaptive equilibrium choices outside the registered procedure.","A solver or data service may change behavior without a visible version change.","Some solver promises may depend on semantic or empirical conditions that cannot be checked mechanically or conservatively screened with confidence. A base-certified calculation may still use an empirically poor economic model and produce a misleading policy conclusion. A valid impossibility result for unrestricted models may be wrongly extended to finite or promised subclasses."],"expert_types":["Macroeconomist who develops or validates policy models","Computability or formal-methods researcher","Numerical equilibrium-solver specialist","Central-bank model-governance lead","Independent policy-model reviewer"],"expert_questions":["Can the proposed formal convergence property preserve the economic question used by policy staff without silently narrowing it?","For each external solver or human step, which promise conditions can be checked before its output is accepted?","Can the unrestricted model class and policy query support a correct, independently checkable undecidability reduction?","Do downstream policy materials preserve ORACLE_RELATIVE and UNKNOWN labels, or collapse them into a single ranking?","Against ordinary validation documentation, how many dependency-classification errors does the passport prevent, and at what review cost?"],"ranking_note":"The harmonized review placed this candidate in Band B, ranking 16–30 across three weighting profiles. That score is only a post-hoc reading order; it is not an experimental endpoint or economic-value measure, and its pilot-speed input is a cost-band proxy.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-PARTNER-19","plain_language_title":"Keep Wetland Records Stable Across Systems","one_sentence_summary":"A behavioral evidence ledger would test whether different storage systems produce the same wetland monitoring history, coverage, provenance, and greenhouse-gas aggregates without exposing their internal layouts.","problem_plain":"Wetland observations, laboratory results, corrections, and quality decisions may live in workbooks, database tables, raster layers, or files. Analysis scripts can quietly depend on row order, worksheet names, sentinel values, grid resolution, or a database's duplicate rules. A storage migration or regridding can then change which observations count, how revisions apply, or how values are aggregated even though the nominal fields remain. That makes it hard to tell a scientific change from an implementation-induced discontinuity.","proposal_plain":"Define the monitoring record by its behavior, not by its files or tables. The ledger would accept observation batches, retain superseded versions and reasons, apply or revoke quality decisions, answer version-specific queries, compute only declared variable-appropriate aggregates, report coverage and provenance, and create repeatable snapshots. Submissions must include stable identity, units, spatial and temporal support, method, and authorized provenance. Rejected operations must leave state unchanged and return stable errors for malformed data, incompatible units, duplicate identities, unsupported aggregation, unauthorized revision, or unavailable coverage. Workbooks, relational stores, rasters, indexes, and caches remain hidden behind the interface. Every backend must pass the same examples, generated operation sequences, and representation-leakage audit. Breaking semantic changes require a new contract version rather than reinterpretation of an existing snapshot.","transfer_plain":"The representation-independent interface archetype maps directly to storage substitution: different physical systems should denote the same versioned evidence state and obey the same operations, invariants, errors, and side-effect rules. The mapping is weaker for scientific validity. Backend agreement cannot show that field measurements, quality policies, ecological assumptions, or aggregation conventions are themselves correct.","why_it_advanced":"This candidate did not enter the strict-success lane. It cleared a separately calibrated lane for a bounded external data-partner study because the components are implementable and the claim is testable. Crucially, there is no real dependency inventory, authorized replay, measured discrepancy rate, comparative result, committed wetland partner, or project-specific cost evidence.","prior_art_and_open_claim":"Observation schemas, provenance standards, greenhouse-gas repositories, automated quality checks, versioning, and ecological workflow tools already provide close constituent parts. The remaining claim is narrower: for one frozen wetland workflow, a contract covering identity, supersession, quality decisions, units, coverage, provenance, errors, snapshots, and variable-aware aggregation will make an incumbent adapter and an independent backend agree on every management-relevant public result without clients inspecting storage internals. That comparison has not been run.","test_and_decision":"With written owner approval, copy one completed project's data into a read-only sandbox and freeze 80–200 real or redacted operations spanning measurement, correction, quality review, snapshot, and aggregation. Compare the unchanged workflow, a schema-normalized export, and the behavioral contract implemented by both an incumbent adapter and an independent in-memory model. Add generated duplicates, reordered calls, incompatible units, revoked decisions, incomplete coverage, and authorization failures. Pass only with zero unexplained differences across declared results and no internal access. Reject the problem locally if no internal dependence exists; reject the intervention if conforming implementations still produce a management-relevant divergence or require exposing the incumbent layout or algorithm.","deployment_and_cost":"The first study must remain read-only and cannot alter official balances, records, eligibility, crediting, compliance, or management decisions. Rough 2026 resource-equivalent bands are $10,000–$50,000 for first evidence, $50,000–$250,000 for startup, $250,000–$1 million for launch, and $50,000–$250,000 annually; they are not vendor quotes.","risks_and_uncertainties":["The contract may preserve an incorrect scientific convention because incumbent behavior is mistaken for intended meaning.","Unit or spatial-support normalization may combine observations that should instead be rejected as incompatible.","An adapter defect may be misdiagnosed as a difference between the underlying storage systems.","Finite conformance tests may miss operation sequences that alter a management-relevant result.","Opaque storage may impede legitimate scientific inspection unless provenance and approved diagnostic views are adequate. Stable identifiers and detailed provenance may disclose sensitive site information if access controls are weak. A broad contract may become expensive to govern, while a narrow one may omit decision-critical behavior."],"expert_types":["Wetland greenhouse-gas monitoring scientist","Environmental data architect","Scientific provenance and standards specialist","Carbon-program monitoring or methodology lead","Data-governance and confidential-location reviewer"],"expert_questions":["Which observation, correction, quality, coverage, and aggregation behaviors can change a management-relevant greenhouse-gas result?","Do any current client scripts inspect worksheet coordinates, schemas, file paths, raster cells, or private status codes?","Which variables are extensive or intensive, and what aggregation rules preserve their scientific meaning?","Can two independently implemented backends pass the contract yet still disagree on an official or management-relevant output?","What access, retention, and disclosure rules apply to site identities and provenance in the proposed sandbox?"],"ranking_note":"The post-hoc harmonized review placed this candidate in Band B, ranking 11–29 depending on weighting. This is a reading-order aid, not its endpoint or a value estimate; the pilot-speed input reflects cost-band affordability rather than measured elapsed time.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP04-STRICT-04","plain_language_title":"Retire Outdated Course Guidance Safely","one_sentence_summary":"An artifact registry would help course teams find stale guidance across copied course offerings while protecting accessibility supports, assessment dependencies, and records that must remain recoverable.","problem_plain":"Repeated course copies accumulate hints, examples, rubrics, refreshers, prompts, accommodations, and corrective notes. Older items may remain visible after the syllabus, assessment, learner group, or linked resource changes, leaving students with contradictory or obsolete guidance. Staff cannot safely clean everything up: an old-looking item may still support an accommodation, assessment, grade review, accreditation record, or reconstruction of a prior offering. Current course tools expose some clutter and broken links but do not resolve these lifecycle decisions.","proposal_plain":"Give each instructional-support artifact an identity, owner, source course version, lifecycle state, validation date, dependencies, retention class, and provisional expiry trigger. A dashboard would flag suspected stale items and rank review priority using age, present access, contradiction, ownership, and reconstruction value, but a person would decide what happens. Authorized reviewers could refresh, retain, unpublish, archive, compact, hold, or place an item in reversible quarantine. Accessibility, assessment, grade-review, accreditation, legal, and records dependencies must be checked before removal. Retired items leave a marker pointing to a successor, archive, or explicit unavailable state. Temporary supports receive review dates when created. Periodic revalidation revisits active items and exception holds, while restore tests confirm that archived materials remain readable, attributable, and connected to the correct historical course.","transfer_plain":"The layer-decay archetype has a clear structural match: each course offering deposits another support layer, and old layers can look current. Lifecycle states, expiry reviews, dependency checks, quarantine, retirement markers, and restore tests bound the active layer without erasing protected history. Age is only a review signal, however; it cannot determine whether instructional content remains valid.","why_it_advanced":"This candidate passed Experiment 4's strict researched-candidate bar because institutions already perform course cleanup and retention work, relevant owners exist, and a read-only comparison is feasible. The endpoint does not mean the registry works in practice, is novel, saves money, improves learning, or has institutional authorization beyond a bounded study.","prior_art_and_open_claim":"Manual course audits, link validation, cleanup tools, content templates, version distribution, retention schedules, and archives are established. The open contrast is whether an artifact-level registry spanning course generations improves identification of current versus stale supports and surfaces assessment, accessibility, and records dependencies better than a careful manual audit aided by Canvas Link Validator and TidyUP. It must also preserve recovery and add no more than 20% reviewer time. Only that integrated, measured comparison remains open; world novelty was not assessed.","test_and_decision":"With instructor, LMS, accessibility, privacy, and records approval, use four sandboxed snapshots of one course, excluding submissions, grades, and identifiable accommodation records. Sample at most 80 instructor-created supports and compare the existing manual audit plus Link Validator and TidyUP with the same evidence plus the registry. Two authorized reviewers classify currentness, visibility, dependencies, exceptions, and proposed state; seed up to eight synthetic defects. Proceed only with sensitivity of at least 0.85, false flags at most 0.10, Cohen's kappa at least 0.70, no protected-data exposure, and median review time within 20% of baseline. Reject incremental advantage if classification or dependency recall does not improve.","deployment_and_cost":"Begin with a read-only inventory; do not change visibility, content, grades, assessments, accommodations, or retention. Rough 2026 resource-equivalent bands are under $10,000 for first evidence, $50,000–$250,000 for startup, $250,000–$1 million for operational launch, and $50,000–$250,000 annually. These are assessment bands, not vendor prices.","risks_and_uncertainties":["Reviewers may mistake age for staleness even when an old explanation remains correct.","Low access counts may unfairly flag essential but rarely used accessibility or safety materials.","Incomplete links to assessments, external tools, grade reviews, or accommodations may make a load-bearing artifact appear safe to retire.","A composite score may create false precision or encode local pedagogical preferences.","Quarantine could conflict with mandatory destruction rules or retain sensitive material too long. Lifecycle labels may confuse students if exposed without careful wording. Registry metadata and temporary holds may themselves become stale. Successful restore tests on sampled formats may conceal unreadable materials elsewhere in the archive."],"expert_types":["Instructional designer","Instructor experienced with repeated LMS course copies","LMS administrator or integration specialist","Accessibility and disability-services representative","Academic records, privacy, or retention officer"],"expert_questions":["Can artifacts be matched reliably across copied course shells without confusing distinct items or missing descendants?","Which dependencies can the LMS and external tools expose, and which still require human review?","What local rules distinguish ordinary instructional supports from protected educational or grade-evaluation records?","Do reviewers agree on currentness, contradiction, dependencies, and lifecycle state at the required thresholds?","Does the registry improve detection over Link Validator, TidyUP, and manual review without exceeding the 20% time limit?"],"ranking_note":"The harmonized review placed the candidate in Band B, with ranks from 18 to 27 across weighting profiles. This post-hoc ordering does not alter its strict endpoint or measure economic value; pilot speed was represented only by a cost-band proxy.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-PARTNER-05","plain_language_title":"Test Cold-Case Theories on Equal Terms","one_sentence_summary":"A bounded hypothesis gate would compare rival cold-case explanations using equal records, preregistered predictions, capped resources, and rights-aware scoring without turning advancement into a finding of guilt.","problem_plain":"Cold-case units may have several plausible explanations but limited analyst time, specialist review, and laboratory capacity. The incumbent team can control the chronology, evidence requests, and briefing order, allowing one theory to advance through ownership or presentation rather than discriminating evidence. Rival teams can also duplicate work, hoard information, or escalate intrusive requests. A resource-allocation ranking may then be mistaken for proof about a named person even though it only selects what to examine next.","proposal_plain":"For at least two distinct, minimally supported hypotheses, create separated teams with equal access to a read-only, citation-addressable record. Before new results arrive, each team registers its theory, contrary evidence, falsifiers, discriminating predictions, uncertainty, proposed tests, and expected privacy, evidence-consumption, and third-party burdens. Freeze eligibility, scoring, conflicts, appeals, and prohibited conduct in advance. Equal analyst-hour and submission caps limit escalation. An independent panel scores evidence coverage, falsifiability, distinctiveness, provenance, treatment of contradictions, feasibility, and rights-adjusted information value. At most two complementary hypotheses receive a bounded validation slot, not investigative authority. Advancement establishes neither guilt, probable cause, admissibility, nor permission for contact, search, surveillance, testing, charging, or public accusation. New evidence or failed predictions can reopen entry, and later review compares preregistered predictions with results.","transfer_plain":"The bounded-rivalry archetype maps to several hypotheses competing for scarce analysis and testing. The gate defines who may compete, equalizes inputs and budgets, penalizes harmful off-process behavior, permits more than one bounded award, and reopens competition after new evidence. The structural mapping is plausible, but rivalry could intensify team commitment or strategic withholding instead of improving reasoning.","why_it_advanced":"This candidate did not reach the strict-success lane. It cleared a separate empirical-partner lane because a masked retrospective comparison is measurable and can be kept from affecting cases. The central field evidence is missing: no agency partner, prevalence estimate, comparative trial, validated scoring rubric, measured local cost, or evidence that team separation is safer than collaboration exists.","prior_art_and_open_claim":"Multidisciplinary cold-case review, Analysis of Competing Hypotheses, explicit alternatives, evidence for and against each theory, ordered information exposure, provenance, and auditable forensic reasoning already exist. Research also warns that formal hypothesis layouts do not reliably reduce bias. The remaining claim is that separated teams, identical cutoff records, registered predictions, equal budgets, rights-adjusted scoring, and no more than two validation slots outperform both case conferences and ordinary ACH without increasing unsupported allegations, intrusion, evidence use, entrenchment, withholding, or panel inconsistency.","test_and_decision":"With an authorized partner, preregister a three-condition retrospective crossover using 8–12 masked, time-split closed or synthetic cases: the proposed gate, a multidisciplinary conference, and a noncompetitive ACH worksheet. Give each condition identical cutoff records and analyst hours, rotate qualified participants, and prevent case recognition or outcome leakage. Before revealing later evidence, capture citations, contradictions, probabilities, predictions, actions, expected information value, privacy burden, and evidence consumption. Blinded reviewers assess calibration, discrimination, citation completeness, unsupported allegations, redundancy, reliability, time, and intrusion. Reject the claim if the gate fails to beat the better comparator or worsens calibration, agreement, withholding, entrenchment, allegations, privacy burdens, or evidence-consumption proposals.","deployment_and_cost":"Only a retrospective, no-case-impact simulation using masked copies is authorized initially; it cannot reopen a case or affect any person or evidence. Rough 2026 resource-equivalent bands are $50,000–$250,000 for first evidence, $250,000–$1 million for startup and operational launch, and $1–$5 million annually. They are not vendor quotes.","risks_and_uncertainties":["Assigned teams may become more committed to their hypotheses and resist contrary evidence.","Teams may withhold urgent exculpatory information until scoring or relabel similar explanations to qualify separately.","Scorers may reward narrative confidence or specificity rather than evidentiary value and calibration.","Historical cutoff packets may omit context legitimately available to investigators at the time.","Case recognition or later-outcome knowledge may leak to participants or reviewers. Resource caps may disadvantage a genuinely complex explanation. Even confidential rankings could stigmatize investigators or people named in hypotheses. Officials or the public may misrepresent advancement as evidence of guilt or authority for coercive action."],"expert_types":["Cold-case or major-case review supervisor","Forensic scientist or laboratory evidence custodian","Investigator trained in competing-hypothesis analysis","Prosecutor, defense-disclosure, or criminal-procedure specialist","Privacy, research-ethics, and experimental-design reviewer"],"expert_questions":["Can masked cutoff packets provide equal and sufficiently complete records without revealing case identities or later outcomes?","Does the scoring rubric produce reliable rankings across independent panels and resist gaming through confidence or narrative style?","Compared with case conferences and ACH, does the gate improve held-out calibration and rights-adjusted information value?","Does team separation increase withholding, hypothesis entrenchment, unsupported allegations, privacy burden, or duplicative evidence requests?","Which local authorities must approve record access, disclosure handling, laboratory recommendations, retention, and participant involvement?"],"ranking_note":"The post-hoc harmonized review placed this candidate in Band B, ranking 24–28 across weighting profiles. That ordering is neither its empirical-partner endpoint nor an economic-value measure, and the pilot-speed input is an affordability proxy rather than observed study duration.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]}]}