{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","source_assessment_id":"predictive_residual_processing__mathematics:P4:v0","cell_id":"predictive_residual_processing__mathematics","proposal_index":4,"qualification":"EMPIRICAL_PARTNER_CANDIDATE","criteria":{"specific_differentiated_claim":{"status":"YES","reason":"The proposal makes a specific, falsifiable contrast: complete independent checking plus frozen theorem-level predictions, typed residual review, protected-event bypasses, audits, reconstruction, and fallback will reduce review time or report volume without materially degrading recall, reconstruction, or reviewer accuracy versus full and dependency-grouped reports."},"credible_problem_signal":{"status":"YES","reason":"External evidence supports formal-library review burden, costly regression discovery, and migration effects arising from elaboration and timeouts, although the prevalence and magnitude of redundant theorem-record review remain unmeasured."},"identifiable_partner_or_adopter":{"status":"YES","reason":"Mathlib maintainers and reviewers are a concrete adopter and authorizer class with corpus stewardship, merge authority, and access to relevant migration workflows."},"partner_access_is_necessary":{"status":"YES","reason":"The remaining questions require an authorized real migration corpus with retained checker records, ecosystem-specific dependency and trust metadata, and blinded participation by qualified reviewers; public research cannot establish predictor performance, workflow effects, rare-event handling, or operating cost."},"safe_authorized_first_step":{"status":"YES","reason":"The first step is a read-only replay of an archived, non-release-blocking migration that preserves complete checking and the canonical report, freezes the model and thresholds, leaves approval authority with maintainers, and restores full reporting on any validity failure."},"bounded_decisive_empirical_design":{"status":"YES","reason":"One preregistered migration of at least 500 declarations compares full, dependency-grouped, and residual interfaces using an untouched partition, blinded reviewers, injected failure modes, exact reconstruction checks, protected-event recall, accuracy, time, volume, fallback, audit, and cost thresholds with explicit falsifiers."},"no_material_negative_gate":{"status":"YES","reason":"All substantive pipeline gates are YES. Cost scope is uncertain rather than materially negative, and the test directly measures model, audit, fallback, and operating effort alongside reviewer savings."},"not_merely_more_research":{"status":"YES","reason":"The next action is a named, operational shadow replay with specified partner authorization, data, comparators, assignments, measurements, success thresholds, and stop conditions—not open-ended investigation or additional web research."}},"uncertainty_types":["PROBLEM_PREVALENCE","ADOPTER_PULL","INCREMENTAL_EFFECT","WORKFLOW_FIT","DATA_ACCESS","COST_SCOPE"],"partner_profile":"A large machine-checked mathematics library—most concretely mathlib—with maintainers authorized to approve access to an archived migration, proof-infrastructure staff able to export complete checker and dependency records, and qualified reviewers able to participate in a blinded workflow study.","required_access":"Maintainer authorization; one archived, non-release-blocking migration with at least 500 eligible declarations; complete pre- and post-revision checker records, corpus manifests, dependency and trust metadata, diagnostics, and version provenance; an untouched evaluation partition; and qualified reviewers for randomized blinded comparison.","bounded_empirical_test":"Pre-register and run one read-only archived replay. Freeze the predictor, residual taxonomy, protected classes, thresholds, and audit sample before revealing the untouched outcomes. Randomly assign blinded reviewers to full theorem-level reports, dependency-grouped full reports, or the residual interface. Inject unexpected failures and survivals, dependency changes, removed obligations, axiom or admitted-fact changes, timeouts, missing outputs, manifest gaps, and version mismatches; measure reconstruction, protected recall, classification accuracy, review time, report volume, audit disagreement, fallback, and total resource-equivalent cost.","success_condition":"The residual interface achieves exact canonical ledger reconstruction, 100% protected-event recall, zero missing-as-success errors, no consequential independently audited mismatch left unrouted or without fallback, reviewer classification accuracy whose lower confidence bound is no more than five percentage points below the best comparator, at least 20% improvement in median review time or report volume, correct fallback on structured drift, and lower total model, audit, and operating cost than baseline review at equal fidelity.","falsification_condition":"Reject the intervention if any protected case is suppressed, any ledger cannot be reconstructed, any missing output appears as success, an audit finds a consequential mismatch that was neither routed nor covered by fallback, reviewer accuracy breaches the preregistered noninferiority margin, time and volume improve by less than 20%, structured drift fails to trigger fallback, or total operating cost is not lower than baseline reviewer cost at equal fidelity.","rationale":"This is an appropriate empirical-partner candidate because adjacent prior art leaves a narrow architectural claim intact, credible evidence identifies both the maintenance problem and concrete authorizers, and the decisive uncertainties concern performance inside a real corpus and human review workflow. Those uncertainties cannot be resolved through ordinary public research, while the proposed shadow replay is authorized, reversible, comparator-based, tightly bounded, and capable of stopping the line on clear evidence."}