{"dossiers":[{"portfolio_id":"EXP06-STRICT-01","plain_language_title":"Fairer Selection for a Festival Slate","one_sentence_summary":"A festival would limit lobbying and unequal access, standardize review, and choose qualifying films for their contribution to a declared slate, then compare that process with its ordinary workflow.","problem_plain":"A festival has fewer competition slots than eligible films, yet entrants may receive unequal access to programmers, deadlines, screening attention, exceptions, or opportunities to supply extra material. Private pitches, gifts, sponsor pressure, and intermediary relationships can influence visibility. Screeners may also carry uneven workloads or undisclosed conflicts. The process can therefore reward campaign resources and access rather than each film’s contribution to the festival’s stated purpose, while leaving entrants unable to challenge clear procedural errors.","proposal_plain":"For one competition section, publish the purpose, available slots, eligibility rules, accepted materials, review stages, conflicts, individual quality threshold, slate-level criteria, tie-breaks, and procedural appeal route before submissions open. Ban gifts, private selection pitches, undisclosed referrals, and entrant-initiated lobbying. Randomly assign each eligible film to two screeners with capped caseloads, record recusals, and obtain a third review when assessments diverge materially. Only films clearing the individual floor reach slate selection. The panel then chooses a feasible, nonredundant portfolio using disclosed considerations such as runtime balance, programmatic perspectives, and audience pathways. An independent reviewer audits sampled files. Appeals may correct process errors but may not replace curatorial judgment. Selection applies to one edition only.","transfer_plain":"The bounded-rivalry archetype becomes a clearly delimited competition for scarce screenings. A published rulebook defines the permitted arena; material and contact limits dampen campaign escalation; workload controls, recusals, sanctions, and audits constrain manipulation; and portfolio selection directs rivalry toward contribution to the whole slate. Post-edition review supplies a revise-or-retire decision.","why_it_advanced":"Experiment 6 classified this candidate as STRICT_SUCCESS because the problem, plausible adopters, component feasibility, reversible shadow test, and contrast with ordinary selection were sufficiently supported for its strict researched-candidate bar. That status does not establish real-world effectiveness, distinctiveness, deployment authority, or economic impact; the proposal remains a partnered research program.","prior_art_and_open_claim":"Transparent rules, controlled submission materials, assigned viewers, second reviews, conflict recusals, authorized exceptions, and intentional slate composition already exist in festival practice. The open claim is narrower: compared with a festival’s current workflow, combining standardized access, random double review, workload ceilings, an individual floor, disclosed portfolio constraints, sampled audit, and process-only appeals will reduce treatment discrepancies and access-linked variation without unacceptable losses in curatorial fit, reviewer attention, or timeliness. That advantage has not been demonstrated.","test_and_decision":"With one festival, preregister a shadow comparison on 60–120 opt-in short-film submissions. Apply ordinary selection and the proposed workflow to the same files without changing official outcomes. Compare missing reviews, time, agreement, recusals, expertise reassignments, threshold stability, portfolio reasons, audit errors, appeal corrections, slate overlap, concentration, and blind coherence ratings. Reject the claim if discrepancies do not fall, agreement worsens, expertise corrections exceed 20%, portfolio choices breach the floor, median labor rises over 50% without matching benefit, appeals reopen taste, or auditors cannot separate rule breaches from lawful curation.","deployment_and_cost":"The authorized first step is a consented shadow pilot, not a live selection change. Rough 2026 USD resource-equivalent bands are $10,000–$50,000 for first evidence and startup, and $50,000–$250,000 for operational launch and annual recurring work. These are bottom-up assessment bands, not vendor quotes; no festival partner has committed.","risks_and_uncertainties":["Standard material limits may omit cultural, production, safety, or accessibility context needed to interpret a film responsibly.","Portfolio criteria could become vague cover for favoritism or token selection, including admission below the declared quality floor.","Double screening and audits could overload programmers, delay decisions, or force a smaller review pool.","A communications firewall may still favor films already legible through credits, press coverage, or established intermediaries.","Audit samples may miss selective misconduct or wrongly treat innocent intermediary concentration as evidence of wrongdoing."],"expert_types":["Festival artistic director or programming lead","Festival operations and submissions manager","Independent procedural auditor","Film-sector conflict, privacy, and labor counsel","Curatorial evaluation or portfolio-design researcher"],"expert_questions":["Can the festival define portfolio criteria precisely enough to guide decisions without disguising unrestricted discretion?","What caseload ceiling permits substantive double review within the existing calendar and staffing budget?","Can auditors reliably distinguish a correctable process violation from a lawful curatorial judgment?","Which contextual materials must remain available so standardization does not systematically disadvantage particular films?"],"ranking_note":"In the post-hoc harmonized reading order, this candidate scored 73–74 across profiles, ranked between 1 and 6, and fell in band A. The ordering used cost-band affordability as a pilot proxy; it is neither an experimental endpoint nor an estimate of economic value.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP06-STRICT-09","plain_language_title":"A Common Meaning for Evidence Records","one_sentence_summary":"A forensic organization would define custody records by their permitted actions and observable meaning, then require any replacement system to pass the same behavior-based tests rather than merely matching fields.","problem_plain":"Forensic evidence systems can encode custody, seals, sampling, consumption, and corrections in different tables, status codes, display orders, or editable notes. During migration or system integration, two implementations may contain similar fields yet report different current custodians, unresolved transfers, seal histories, sample ancestry, remaining quantities, or corrections. Without a representation-independent rule, an organization cannot determine whether a new implementation preserves the procedural meaning of the original record rather than merely copying its visible data.","proposal_plain":"Define an Evidence Continuity Record as an abstract state machine, not a database layout. Its public operations would register items, offer and accept transfers, record seal opening and resealing, derive child samples, record consumption, append linked corrections, query current state, render authorized history, and check named continuity conditions. Each operation would specify permissions, preconditions, results, errors, and side-effect limits. Invariants would prohibit unmatched completed transfers, negative quantities, unlinked samples, silent deletion, and more than one current accountable custodian. Storage tables, vendor codes, indexes, internal identifiers, and row order would remain non-contractual. A shared black-box suite would run examples, generated action sequences, authorization checks, and representation-leakage probes before an implementation could qualify as a substitute.","transfer_plain":"The representation-independent interface archetype becomes a behavioral contract for one evidence record. Abstract states and operations define what custody history means; an opaque boundary prevents clients from depending on vendor internals; and a common conformance oracle tests whether different implementations preserve the same authorized observations, errors, transitions, and invariants.","why_it_advanced":"Experiment 6 classified this as STRICT_SUCCESS because heterogeneous forensic systems, migration difficulty, relevant institutional authority, established software techniques, and a safe synthetic test were sufficiently supported for the strict researched-candidate bar. This does not validate the contract, authorize production migration, establish legal sufficiency, demonstrate distinctiveness, or measure economic impact.","prior_art_and_open_claim":"Custody standards, common schemas, ontologies, audit trails, access controls, two-sided transfers, and migration tools already cover much of the surrounding territory. None of that alone establishes behavioral equivalence between implementations. The remaining claim is that an opaque, vendor-neutral state-machine contract plus one behavior-only conformance and leakage suite will identify semantically acceptable substitutes more reliably than schema matching, audit-trail presence, or vendor-specific acceptance tests. It remains open until the suite detects seeded defects and survives review of material legal and scientific distinctions.","test_and_decision":"With a laboratory quality authority and authorized reviewer, preregister 12–20 synthetic scenarios covering transfers, seals, samples, consumption, corrections, repeated calls, unauthorized queries, and exports. Run them against a transparent reference model, an incumbent-behavior adapter, and at least five mutated adapters containing specified semantic defects. Advance only if the first two agree on every mandatory assertion, expose no prohibited information, and the suite rejects every mutant. Revise or reject the contract if an incumbent sequence is unrepresentable, a mutant passes, or reviewers find a material contract-observable difference after both implementations pass.","deployment_and_cost":"Begin only with synthetic data and sandbox adapters; do not alter live evidence or infer admissibility. Rough 2026 USD resource-equivalent bands are $10,000–$50,000 for first evidence, $50,000–$250,000 for startup and annual work, and $250,000–$1 million for operational launch. They are assessment estimates, not vendor quotes.","risks_and_uncertainties":["The abstract state may omit jurisdiction-specific signatures, native documents, instrument metadata, or another legally or scientifically material distinction.","Tests may copy incumbent quirks and mistakenly preserve them as required behavior.","A passing suite could be misrepresented as proof that recorded events are true or that evidence is admissible.","Incomplete authorized export operations could make the opaque boundary obstruct legitimate audit or disclosure.","Generated tests may miss failures that appear only in long, concurrent, or malformed action sequences."],"expert_types":["Forensic laboratory quality manager","Evidence custodian or property-unit supervisor","Forensic LIMS architect or integration engineer","Evidence and disclosure counsel","Software conformance and property-based testing specialist"],"expert_questions":["Which representation-specific artifacts are legally or scientifically material and therefore must be contract-observable?","Can an adapter reproduce incumbent behavior without relying on undocumented or mutable vendor features?","Do the proposed invariants cover every authorized transfer, correction, sampling, and consumption sequence in the bounded scope?","What mutant set would provide a credible test that the suite distinguishes field similarity from semantic equivalence?"],"ranking_note":"The post-hoc harmonized reading order scored this candidate 73–74, placed it between ranks 2 and 7, and assigned band A. That order used a cost-band affordability proxy and is not an experimental endpoint, legal finding, or measure of economic value.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP04-STRICT-02","plain_language_title":"Truthful Limits for Safety Verification","one_sentence_summary":"A design organization would classify what a safety verifier can honestly decide before routing each controller to exact checking, bounded or approximate analysis, or accountable human review.","problem_plain":"An organization asks one verifier to return a correct, terminating SAFE or UNSAFE answer for every programmable safety controller, even when submissions include unrestricted scripts, unbounded variables, and supplier extensions. The actual computation model and guarantee are left unclear. Engineers may then treat timeouts as failures, keep adding computing power to a requirement that may be undecidable, or quietly restrict accepted designs while continuing to make a universal assurance claim. Release records may not reveal which bounds or assumptions support a verdict.","proposal_plain":"Place an assurance-boundary gate before verification and release. The gate would formalize the controller language, state representation, safety property, environmental assumptions, quantifiers, computation model, and requested guarantee. It would admit exact terminating verification only for mechanically enforceable finite or otherwise decidable fragments. Richer designs would go to sound over-approximation, explicitly bounded exploration, or accountable expert escalation. UNKNOWN, OUT_OF_SCOPE, TIMEOUT, and TOOL_FAILURE would remain separate from SAFE and UNSAFE. Claims of impossibility would require a checked reduction or equivalent proof, not repeated timeouts. Every result would record its fragment, bounds, abstraction, assumptions, guarantee, residual uncertainty, and triggers for reclassification when the model, tool, supplier interface, or environment changes.","transfer_plain":"Computability-boundary mapping becomes an intake and routing decision for controller assurance. The gate separates exact solvability, weaker machine-checkable evidence, and unresolved cases relative to an explicit model. Enforced language fragments and typed results prevent bounded searches, abstractions, tool failures, or nonanswers from inheriting a stronger universal safety claim.","why_it_advanced":"Experiment 4 classified this candidate as STRICT_SUCCESS because the boundary is technically grounded, component practices are mature, relevant authorities exist, and a reversible non-release comparison is clearly testable. Passing that strict researched-candidate bar does not show field prevalence, better safety outcomes, adopter acceptance, deployment authorization, economic impact, or broad distinctiveness.","prior_art_and_open_claim":"Finite-state model checking, abstraction, bounded exploration, UNKNOWN results, witnesses, and formal-assurance records are established practices. The narrower open claim concerns their mandatory ordering and governance: placing a model-relative solvability and scope classification plus an enforced status schema before an otherwise standards-conformant workflow will reduce over-strong or mislabeled verdicts and improve reviewer routing agreement, while losing no more than one conclusive decision in a 12-model pilot. The individual mechanisms are adjacent prior art, and comparative advantage has not been measured.","test_and_decision":"Freeze 12 previously adjudicated controller models spanning finite, bounded, abstractable, unrestricted, unsafe, and timeout cases. Blind and randomize reviewer pairs to the strongest existing workflow or the proposed gate, using equal tools, evidence, and time. Compare adjudicated mislabeled verdicts, routing agreement, hours, nonanswer rates, and conclusive decisions. Advance only with zero mislabeled gate verdicts, improvement in mislabeling or agreement, at most one lost conclusive result, at most one fragment-membership disagreement, and cost below $50,000. Reject if weak or incomplete evidence becomes unqualified SAFE/UNSAFE, agreement fails to improve, or physical semantics cannot be reproduced.","deployment_and_cost":"The first step is a frozen, non-release pilot that cannot alter deployed logic or release status. Rough 2026 USD resource-equivalent bands are $10,000–$50,000 for first evidence and $250,000–$1 million for startup, operational launch, and annual recurring work. These are assessment estimates rather than vendor quotes or certification budgets.","risks_and_uncertainties":["A correct proof may concern a model that omits important physical-controller or environmental behavior.","An unsound abstraction could create false confidence; a coarse but sound abstraction could generate too many unresolved alarms.","Designers may move excluded but safety-relevant behavior into informal escape hatches outside the decidable fragment.","A finite analysis may exhaust resources, inviting staff to confuse infeasibility or timeout with a semantic verdict.","The recorded boundary can become stale after language, verifier, environment, or supplier-interface changes."],"expert_types":["Safety-control chief engineer","Independent formal-verification specialist","Design assurance or certification reviewer","Controller-language and toolchain engineer","Domain regulator or licensing specialist"],"expert_questions":["Does the declared controller language actually include unrestricted computation, or is the accepted class already enforceably decidable?","Can fragment membership and environmental assumptions be reproduced independently for every pilot model?","Are the approximate-analysis modes sound, and how are spurious counterexamples labeled and escalated?","Would the gate add information beyond the strongest existing standards-conformant workflow under equal time and tools?"],"ranking_note":"In the post-hoc harmonized reading order, this candidate scored 72–74, ranked between 1 and 4, and fell in band A. The ordering used affordability as a pilot proxy; it is not a preregistered endpoint, safety finding, or economic-value measure.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]},{"portfolio_id":"EXP05-STRICT-04","plain_language_title":"Keeping Test Norms Current and Traceable","one_sentence_summary":"A registry and scoring gate would keep context-mismatched norm packages out of new psychological scoring while preserving exact historical packages for authorized reconstruction.","problem_plain":"A psychological test may accumulate several norm tables, scoring transformations, cutoff rules, and population-specific calibration files. Older packages can remain in manuals, spreadsheets, scripts, or software without a clear current, restricted, superseded, or archived status. Users may then score against an unintended population or edition, producing inconsistent standardized results. Simply deleting old packages is also unsafe because prior reports, longitudinal studies, audits, or research reproductions may depend on the exact historical transformation.","proposal_plain":"Create a lifecycle registry linked to a scoring gate. Give every norm package an immutable identifier and record its instrument version, reference population, collection period, intended uses, transformations, cutoffs, validation evidence, software dependencies, latest review, successor, and lifecycle state. A review lease or relevant context change would move a package to review-due, not declare it invalid. New scoring would require an active, context-matched package or a documented expert override. Superseded packages would disappear from ordinary prospective selection but remain retrievable by identifier for authorized reconstruction. Dependency checks would block removal while reports, studies, longitudinal series, or audits still need a package. Quarantine, archival tiers, disposition markers, and fixed-vector restore drills would make retirement reversible and test whether historical scoring remains executable.","transfer_plain":"Layer-decay management becomes governance for successive norm and scoring packages. The registry identifies each deposited layer, detects expired reviews and changed contexts, and separates active authority from historical retention. Dependency checks, quarantine, archives, deletion markers, and restoration tests prevent cleanup from destroying the scoring basis of earlier work.","why_it_advanced":"Experiment 5 classified this candidate as STRICT_SUCCESS because norm changes can matter, relevant professional guidance and adopter classes exist, the software pattern is feasible, and a bounded read-only comparison is testable. This strict researched-candidate result does not establish field effectiveness, cross-publisher portability, distinctiveness, deployment authorization, or economic value; partnered research remains necessary.","prior_art_and_open_claim":"Professional guidance already calls for current, population-relevant, identifiable norms; publishers already renorm tests, label legacy products, retire scoring software, and maintain successor platforms. The open claim is the added effect of integrating immutable package identity, expiration of prospective authority, context-gated selection, dependency-constrained retirement, and tested restoration. Compared with ordinary files, manuals, or platform presentation, that package should reduce context-mismatched selection and improve exact historical reconstruction without excessive valid-use blocks or expert overrides. This comparative claim remains untested.","test_and_decision":"Use one licensed or synthetic test family with at least three packages and randomly assign 20–30 qualified evaluators to the current presentation or a read-only registry across 24–40 scenarios. Compare mismatched selections and exact reconstruction, plus time, blocks, overrides, identity errors, and preservation choices; restore one archived package against a fixed vector. Advance only with at least a 10-point mismatch reduction, 15-point reconstruction gain, no more than a 5-point increase in valid-use blocks, at most 10% unnecessary overrides, zero identity errors, and exact restoration. Redesign if either primary outcome fails or users confuse lifecycle status with validity.","deployment_and_cost":"Start with a read-only inventory and mock gate; do not change production scores, reports, or source files. Rough 2026 USD resource-equivalent bands are $10,000–$50,000 for first evidence, $50,000–$250,000 for startup and annual work, and $250,000–$1 million for operational launch. They are not vendor quotes.","risks_and_uncertainties":["Users may mistake an active lifecycle state for evidence that a norm is psychometrically valid for the individual case.","Incomplete population, instrument, or intended-use metadata could falsely authorize a mismatched package.","Missing references from reports, spreadsheets, printed tables, or local scripts could allow destructive retirement.","Archived transformations may become non-executable as software environments and formats age.","Registry access or preservation policies could expose restricted test content or retain sensitive norm data longer than authorized."],"expert_types":["Psychometrician responsible for test norms","Test publisher or assessment-program owner","Practicing psychologist or qualified assessment user","Scoring-platform and archival systems engineer","Test-security, privacy, and records counsel"],"expert_questions":["Can package metadata express intended population and use precisely enough to gate scoring without implying validity?","What events should trigger review-due status, and who may renew, restrict, supersede, or override a package?","How completely can inbound dependencies from reports, studies, scripts, and longitudinal datasets be discovered?","Can an archived package reproduce the fixed historical score in a controlled environment without exposing restricted materials?"],"ranking_note":"The post-hoc harmonized reading order scored this candidate 72–74, ranked it between 2 and 5, and placed it in band A. That order used cost-band affordability as a proxy and is neither an experimental endpoint nor a measure of economic value.","source_ids_used":["S1","S2","S3","S4","S5","S6","S7","S8"]}]}