{"dossiers":[{"portfolio_id":"EXP06-STRICT-04","plain_language_title":"Residual-First Review of Forensic Timelines","one_sentence_summary":"A read-only review layer would foreground unexplained timeline changes while preserving the complete forensic record, reconstructible context, independent audits, and automatic return to full review.","problem_plain":"Digital forensic examiners may confront millions of timestamped system, application, synchronization, and acquisition events. Repeated, predictable background activity can consume attention before an examiner reaches missing, extra, displaced, or altered events that matter to a case. Simply hiding familiar-looking events is unsafe: an apparently routine record may still reveal guilt, innocence, provenance, deletion, clock problems, or evidence-integrity failures, and context-free anomalies can themselves invite overinterpretation.","proposal_plain":"For one declared operating-system, application, and version scope, freeze a reference model that predicts ordinary background event bundles. Compare the complete extraction with those predictions and describe each discrepancy as missing, extra, reordered, time-shifted, or attribute-changed. Show prioritized discrepancies with model identity, provenance, uncertainty, and enough neighboring events to reconstruct the local sequence. Never alter the forensic image or complete extraction. User-authored material, deletion indicators, integrity and acquisition failures, clock discontinuities, required disclosures, and examiner-requested records always appear in full. An independent reviewer examines random and risk-selected raw windows. Unsupported versions, model or parser mismatches, audit disagreement, excessive reconstruction error, and queue overload switch the affected scope back to complete chronological review. Human reviewers, not the model, determine evidentiary meaning.","transfer_plain":"The predictive-residual archetype maps strongly here: a versioned model predicts routine event sequences, while the analyst receives typed differences between prediction and observation. Reconstruction, version checks, independent raw sampling, protected-event bypasses, drift monitoring, and full-review fallback instantiate the archetype without replacing the underlying evidence.","why_it_advanced":"This candidate passed Experiment 6's strict researched-candidate bar because the burden is documented, a bounded prototype appears feasible, the comparison is falsifiable, and the safeguards directly address context loss and automation risk. That status is not real-world validation, a novelty finding, deployment approval, or evidence of economic impact.","prior_art_and_open_claim":"Complete timeline extraction, static filters, analyzers, pattern reconstruction, prioritization, provenance drill-down, clustering, and visualization already exist. The remaining claim is narrower: compared with full chronological review and the strongest static-filter or analyzer workflow, a frozen residual-first layer combining typed discrepancies, reconstructible context, protected bypasses, independent raw-window audits, synchronized versions, and automatic fallback would reduce review effort without increasing material or exculpatory misses, interpretation errors, or total audit and maintenance burden.","test_and_decision":"Pre-register a randomized crossover study with about 8–12 qualified examiners, 4–6 synthetic or reusable closed-case images, and 24–36 matched sections from one software-version scope. Compare complete review, the strongest Plaso or Timesketch static workflow, and residual-first review. Measure examiner minutes, time to inspect material events, inculpatory and exculpatory recall, protected-context recall, false escalation, reconstruction disagreement, disparities, fallback, and total workload. Stop progression for any protected-bypass failure, unreviewed material suppression, unacceptable recall loss, recurrent mismatch, systematic disparity, negligible effort reduction, or audit burden comparable to full review. Passing would not authorize live-case use.","deployment_and_cost":"The first retrospective evidence study is estimated at $50,000–$250,000. Initial deployment startup is $250,000–$1 million; an operational launch is $1–$5 million, with $250,000–$1 million recurring annually. These are rough 2026 USD resource-equivalent bands, not vendor quotes, and version maintenance may erase expected savings.","risks_and_uncertainties":["A mistaken reference model could suppress material inculpatory, exculpatory, provenance, or integrity information.","Examiners could mistake statistical unusualness for criminal intent, authorship, or evidentiary significance.","Parser changes, clock ambiguity, missing data, or mismatched model versions could create false discrepancies or false silence.","Reference data and priority weights could work unevenly across applications, languages, devices, or user contexts.","Sparse audit samples might miss rare blind spots, while incomplete version records could obstruct disclosure and independent challenge."],"expert_types":["Digital forensic examiner","Forensic laboratory quality and method-validation specialist","Forensic timeline-tool developer","Defense-side digital-forensics expert","Evidence and disclosure lawyer"],"expert_questions":["Can sampled timeline windows be reconstructed to the semantic fidelity required for independent examination and disclosure?","Which event classes require unconditional full-context presentation under laboratory policy and applicable law?","Does residual-first review preserve inculpatory and exculpatory recall against both complete review and the strongest static workflow?","How often would operating-system, application, parser, locale, clock, or acquisition changes force model revalidation?","Does audit, synchronization, and maintenance work remain below the attention saved during review?"],"ranking_note":"The harmonized score is only a post-hoc reading-order aid: this candidate ranked 32–51 across profiles and falls in band C. The ranking is not an experimental endpoint or a measure of novelty, deployment readiness, or economic value.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-PARTNER-10","plain_language_title":"A Clinic for Cross-Field Lemma Handoffs","one_sentence_summary":"A rotating pair of specialists would turn eligible cross-subfield proof obstructions into precise, review-ready lemma packets without proving the lemma or changing the theorem.","problem_plain":"In a mathematics collaboration spanning subfields, a proof obligation may stall because participants use different definitions, notation, assumptions, or standards of explanation. Contributors repeatedly search for appropriate specialists, reconstruct terminology, and discover lost hypotheses only after work begins. Yet no mathematics-specific audit shows how often this translation and routing work, rather than genuinely missing mathematical insight, causes delay. Informal access to well-connected experts may also determine which obligations receive attention.","proposal_plain":"Create a governed clinic staffed by a rotating pair of specialists covering both sides of an interface. Intake must state the parent theorem, local definitions, requested conclusion, allowed assumptions, known dependencies, attempted approaches, and exact source of confusion. In a capped session, the specialists check compatibility, translate notation, preserve quantifiers and hypotheses, separate established implications from open mathematics, identify side conditions, and produce the smallest faithful lemma packet with an intended recipient or escalation path. An independent reviewer compares the packet with the original request. The clinic must reject disputes about truth, foundations, credit, theorem changes, or genuinely new mathematics. It tracks queue depth, labor, reopens, defects, conflicts, reviewer capacity, and specialist recovery, then rotates or rests staff instead of treating them as unlimited infrastructure.","transfer_plain":"The catalytic-pathway analogy is plausible but not literal. The clinic acts as a reusable facilitator that lowers recurring translation and routing barriers, releases each packet, and returns to readiness. It cannot change whether a lemma is true or replace the mathematical insight and scrutiny required to prove it.","why_it_advanced":"This candidate did not enter the strict-success lane. It cleared a separately calibrated empirical-partner lane because a small shadow comparison is practical and reversible. Advancement therefore means it merits a bounded external data-partner study, while the prevalence of the problem, available records, specialist willingness, and comparative benefit remain unverified.","prior_art_and_open_claim":"Focused mathematics programs, public question intake, shared glossaries, direct consultation, formal dependency blueprints, and large collaborative proof projects already address parts of the problem. The open claim concerns the governed combination: for archived obligations crossing the same two subfields, a capped rotating clinic with a versioned lemma-packet contract would beat an equal-access ad hoc handoff on preparation labor or time without adding statement defects, reopens, bad routing, or hidden downstream burden.","test_and_decision":"Find one willing project and eight authorized, de-identified, resolved obligations from the same subfield pair, including translation successes, substantive proof problems, and an incompatible case. Compare the rotating clinic with an ad hoc team given equal source access and recorded labor; conceal archived resolutions until packets are frozen. Blinded reviewers score statement fidelity, definitions, hypotheses, quantifiers, uncertainty, routing, readiness, reopens, total labor, elapsed time, and downstream burden. Reject progression after a confidentiality breach, undisclosed material alteration, failure to reject incompatibility, mostly substantive escalations, increased defects or burden, or failure to meet the precommitted practical improvement threshold.","deployment_and_cost":"The first evidence study and initial startup are each estimated at $10,000–$50,000. Operational launch and annual recurring support are each $50,000–$250,000. These rough 2026 USD resource-equivalent bands are not quotes; scarce dual-subfield specialists, independent reviewers, compensation, and confidentiality controls could dominate actual cost.","risks_and_uncertainties":["Translation could silently remove a hypothesis, change a quantifier, or overstate equivalence between definitions.","Clinic packets or specialists could acquire informal authority despite having no power to accept mathematical claims.","Eligibility rules could favor contributors already fluent in preferred notation or connected to clinic staff.","Rotating specialists could lose continuity, become exhausted, or receive inadequate credit for substantive formulation work.","Confidential conjectures, correspondence, or attribution information could reach unintended recipients."],"expert_types":["Mathematician from each participating subfield","Collaborative-project proof architect","Mathematical editor or independent proof reviewer","Research ethics and confidentiality specialist","Academic labor and attribution specialist"],"expert_questions":["Can the project identify eight comparable resolved obligations whose reuse is authorized and whose resolutions can be concealed?","What fraction of archived stalls arose from translation and routing rather than missing mathematical insight?","Which changes to definitions, hypotheses, or quantifiers count as material fidelity defects?","Can rotating specialists maintain continuity without exceeding fair workload and compensation limits?","Does the clinic reduce total submitter, specialist, and downstream-review effort under equal access?"],"ranking_note":"The harmonized assessment places this candidate 43–51 across profiles, in band D. That post-hoc ordering only guides reading and uses a cost-band affordability proxy; it is not an experimental endpoint, economic-value estimate, or substitute for the missing partner study.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-STRICT-02","plain_language_title":"Choosing One Proof Route Fairly","one_sentence_summary":"A no-stakes challenge would test whether correctness-gated, resource-capped comparison selects a more maintainable route for one scarce formal-verification slot than simpler selection methods.","problem_plain":"A mathematics consortium may have several proposed routes to a theorem but enough specialist time to formalize only one. If the slot goes to the first route that appears complete, teams may benefit from premature completeness claims, hidden dependencies, elaborate presentation, or silence about fatal flaws. A strategically packaged route can therefore displace a sounder, more maintainable alternative, consuming scarce referee and formalizer time while discouraging useful sharing between teams.","proposal_plain":"Run a voluntary, time-limited challenge under rules frozen before judges see team identities. Every team submits the same proof certificate, listing contributors, axioms, imported results, dependencies, unresolved gaps, and known counterchecks. Independent reproduction is a pass-or-fail gate: presentation quality or expected cost cannot compensate for a failed mathematical step. Only passing routes are ranked on declared criteria such as dependency transparency, modularity, explanatory coverage, and estimated formalization burden. Give teams equal page limits, clarification rounds, and judge access. Permit disclosed collaboration and route merging, while prohibiting tampering, plagiarism, retaliation, private judge contact, sham independence, and concealment of a known fatal audit flaw. Provide a separate procedural appeal and reopen the slot if staged formalization later fails. Preserve attribution and access for every correctness-passing route.","transfer_plain":"The bounded-rivalry archetype maps directly to competition for one consortium-funded verification slot. A frozen rulebook, noncompensable correctness gate, equal reviewer-facing resource caps, auditable conduct rules, appeal, and reopening constrain strategic escalation while directing comparison toward downstream verification needs. Rivalry remains optional; collaboration may still prove better.","why_it_advanced":"This candidate passed Experiment 6's strict researched-candidate bar because the scarce-resource setting is coherent, adjacent infrastructure exists, authority can be bounded, and a reversible comparison can falsify the claim. It has not demonstrated better selection in practice and is not a novelty, deployment, or economic-impact finding.","prior_art_and_open_claim":"Proof certificates, machine checking, contribution rules, dependency graphs, task dashboards, formalization challenges, and large collaborative projects already exist. The unresolved contrast is whether their governance elements work better in this specific allocation decision: can correctness-gated ranking with equal access, frozen secondary criteria, appeal, and staged reopening select a route requiring fewer corrections or formalizer hours than blinded holistic triage or first reproduction, without excess review overhead or suppressed collaboration?","test_and_decision":"With a consenting formalization organization, preregister a no-stakes shadow exercise using eight de-identified packets from a settled theorem: sound routes of differing modularity, known gaps, an undeclared dependency, a polished distraction, and controls. Compare the governed challenge with blinded holistic triage, while logging first independent reproduction as a third comparator. Auditors reproduce critical lemma chains and measure false passage, corrections, observed formalizer hours on a fixed sample, reviewer minutes, agreement, successful planted gaming, appeals, and collaboration effects. Do not progress if an incorrect packet passes, gaming determines selection, agreement misses its threshold, both comparators perform as well or better, or governance exceeds its budget.","deployment_and_cost":"The shadow study is estimated at $10,000–$50,000. Startup and operational launch are each $50,000–$250,000, with annual recurring costs of $250,000–$1 million. These are rough 2026 USD resource-equivalent bands, not budgets or quotes; expert reproduction and governance, rather than software, are the main uncertainties.","risks_and_uncertainties":["A visible ranking could turn proof development into a status contest and discourage useful collaboration.","Secondary scoring could reward a fashionable proof style rather than actual maintainability or mathematical value.","Page and clarification caps could disadvantage routes whose irreducible explanations are longer.","Dependency disclosure could expose unpublished work, while misconduct controls could stigmatize legitimate collaboration.","Judges could underestimate formalization effort or apply school-specific preferences inconsistently."],"expert_types":["Research mathematician familiar with the theorem","Formal-proof engineer or maintainer","Independent mathematical referee","Research-governance and due-process specialist","Mathematical collaboration researcher"],"expert_questions":["Does the target consortium actually face several viable routes competing for one indivisible verification slot?","Can judges reproduce correctness and apply the secondary rubric with preregistered agreement?","Do page and contact caps equalize reviewer access without shifting effort into hidden channels?","Does challenge-selected work require fewer corrections and formalizer hours than both comparison methods?","Would the format measurably discourage route sharing, merging, or disclosure of negative findings?"],"ranking_note":"The harmonized assessment ranks this candidate 38–55 across profiles and assigns band D. This post-hoc reading order is not the Experiment 6 endpoint and does not measure economic value; its affordability input also does not independently estimate elapsed pilot time.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-PARTNER-14","plain_language_title":"Monitoring a Program’s Measurement Footprint","one_sentence_summary":"A frozen model would subtract the aggregate records expected from a place-based prevention program so evaluators can inspect unexplained changes without producing individual scores or treating residuals as causal conclusions.","problem_plain":"A patrol, outreach, reporting, or other place-based prevention program can change both underlying events and how those events are recorded. More officer presence, for example, may predictably alter contacts, detected offenses, calls, complaints, or data completeness. Ordinary dashboards can blur this measurement footprint with external change, displacement, service withdrawal, or harm. Yet unexplained aggregate differences can also be overread as proof of program success, failure, misconduct, or community behavior.","proposal_plain":"Before each monitoring window, copy the authorized program schedule and intensity into a frozen, versioned model that predicts only the aggregate records and observation opportunities the program itself is expected to generate. Compare those expectations with separately retained observations and report typed differences: excess or missing activity, spatial or temporal displacement, source disagreement, complaint or injury changes, and unexplained missingness. Prioritize residuals by reliability, uncertainty, exposure, persistence, rights consequence, and evaluator capacity. Each alert must include its expected value, observed aggregate, uncertainty, provenance, program history, and reconstructible full-window context. Complaints, force, injury, deaths, disparity checks, whistleblower material, outages, integrity failures, and oversight requests always appear in full. Independent reviewers audit random and risk-selected windows. Residuals open evaluation inquiries only; shocks, small cells, mismatch, drift, or audit failures restore complete reporting.","transfer_plain":"The predictive-residual archetype is instantiated through an efference copy of the agency's own authorized schedule: expected program-generated records are separated from the unexplained remainder. The mapping is structurally strong for monitoring, but residuals cannot identify true crime levels, individual risk, causation, or the correct operational response.","why_it_advanced":"This candidate did not enter the strict-success lane. It cleared the separate empirical-partner lane because a retrospective aggregate comparison is bounded and measurable. Missing field evidence remains central: no agency has supplied schedules, linked sources, reviewers, legal approval, audit access, or evidence that the model separates measurement footprint from omitted context.","prior_art_and_open_claim":"Multi-indicator dashboards, pre/post comparisons, matched comparison areas, process and impact evaluations, displacement analysis, and generic anomaly detection already cover much of the substantive work. The narrower open claim is that an action-conditioned residual interface, with reconstructibility, independent raw-window audits, rights-critical bypasses, synchronization, and fallback, would reduce evaluator effort against full and conventional evaluation dashboards without materially missing displacement, reporting divergence, service withdrawal, or harm, or encouraging stronger unsupported causal claims.","test_and_decision":"Pre-register a 12–16-week offline study using 150–250 historical or synthetic windows from one completed program and at least eight blinded evaluators. Compare the existing full dashboard, a conventional process-plus-impact dashboard, and the residual-first interface in balanced crossover order. Include outages, intensity changes, displacement, complaint or injury shifts, service withdrawal, benign shocks, source disagreement, and version mismatch. Require material-change recall within five percentage points of the better comparator and at least 20% lower median review time. Stop for any protected-signal omission or privacy breach, systematic audit misses, excessive reconstruction failure, increased causal overclaiming, failed fallback, or no reduction in total review-plus-maintenance effort.","deployment_and_cost":"The first evidence study is estimated at $50,000–$250,000. Startup is $250,000–$1 million; operational launch is $1–$5 million, with $250,000–$1 million recurring annually. These are rough 2026 USD resource-equivalent bands, not procurement quotes, and local data integration, privacy, oversight, and audit requirements remain unpriced.","risks_and_uncertainties":["The model could normalize repeated over-enforcement, under-service, or rights harm as expected program behavior.","Police-generated records could dominate independent sources and create a self-confirming baseline.","Evaluators or leaders could treat aggregate residuals as causal findings despite confounding and spillovers.","Schedule, boundary, reporting, or threshold changes could be used to make visible residuals disappear.","Sparse audits, uneven independent-source quality, or minimum cell sizes could hide rare harms or create geographic disparities in uncertainty."],"expert_types":["Independent crime-program evaluator","Criminologist specializing in place-based interventions","Civil-rights and community-oversight representative","Government privacy and de-identification specialist","Agency data steward"],"expert_questions":["Can a frozen schedule-conditioned model distinguish predictable recording changes from external events and omitted context?","Which force, injury, complaint, disparity, missingness, and oversight signals must bypass filtering in full?","Are independent data sources sufficiently complete and comparable across all evaluated places?","Do residual displays increase unsupported causal interpretations relative to conventional dashboards?","Does total evaluator, audit, calibration, and maintenance time fall while material-change recall remains within tolerance?"],"ranking_note":"The harmonized assessment places this candidate 40–53 across profiles, in band D. It is a post-hoc reading-order aid using an affordability proxy, not an experimental endpoint, economic-value measure, field-validation result, or evidence that an agency will adopt it.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]}]}