{"cases":[{"alternative_route_blocks":[{"blocked_by_condition_id":"b038_a06_c1","condition_set_id":"B1","how_blocked":"The research extract contains application-time and contemporaneous adjudication fields, but nothing created after the disposition event."},{"blocked_by_condition_id":"b038_a06_c3","condition_set_id":"B3","how_blocked":"Transformations were fitted within the training partition and did not inspect validation records."},{"blocked_by_condition_id":"b038_a06_c4","condition_set_id":"B4","how_blocked":"Applicants and authorization requests are disjoint across partitions."},{"blocked_by_condition_id":"b038_a06_c5","condition_set_id":"B5","how_blocked":"Development did not use a public score, answer key, or repeatedly viewed evaluation result."},{"blocked_by_condition_id":"b038_a06_c6a","condition_set_id":"B6","how_blocked":"The offending status existed at the forecast timestamp inside the insurer; it was not a later correction or retrospective enrichment."},{"blocked_by_condition_id":"b038_a06_c7","condition_set_id":"B7","how_blocked":"Archived API payloads provide a complete record of what the hospital actually received at each forecast time."}],"case_id":"E15A002__DIRECT","case_type":"DIRECT_POSITIVE","domain":"Hospital prior-authorization forecasting","intended_route_id":"B2","omitted_condition_id":null,"remedy_leakage_audit":"The statement reports symptoms, provenance, and controls without naming a solution pattern or recommending any corrective action.","route_evidence":[{"condition_id":"b038_a06_leak_core","intended_status":"SATISFIED","scenario_evidence":"An insurer-only status absent from the hospital's operational request is present during model development and evaluation."},{"condition_id":"b038_a06_c2","intended_status":"SATISFIED","scenario_evidence":"The internal status token mechanically determines the authorization label and is used as a predictor."}],"scenario_text":"A hospital network built a classifier to estimate whether an insurer would approve a prior-authorization request. Offline validation reported 96% accuracy, yet a six-week shadow deployment was barely better than the hospital’s existing rules. The historical research extract included an insurer field called “adjudication status token.” That token was generated by the insurer’s rules service and mechanically determined the approval label. It already existed inside the insurer when the hospital submitted its forecast, but the hospital-facing API never returned it; only the retrospective research feed exposed it. The model ranked the token as its strongest input. No fields were added after the disposition was recorded, applicants and requests were separated across validation partitions, and transformations were estimated only from the training partition. The team also retained the exact API payload delivered for every shadow prediction. It had not consulted a public leaderboard or received interim results from the sealed evaluation. The unexplained issue is why the historical score remains exceptionally high while performance on otherwise similar live requests collapses.","scenario_title":"Prior-Authorization Model Fails in Shadow Use","vocabulary_separation_audit":"This case uses hospital, insurer, authorization, adjudication, and API terminology; it shares no industrial inspection vocabulary with the transfer case."},{"alternative_route_blocks":[{"blocked_by_condition_id":"b038_a06_c1","condition_set_id":"B1","how_blocked":"The joined study table has pre-decision measurements and the contemporaneous buyer disposition flag, but no fields created after the acceptance event."},{"blocked_by_condition_id":"b038_a06_c3","condition_set_id":"B3","how_blocked":"Feature transformations were fitted on supplier training lots without access to evaluation lots."},{"blocked_by_condition_id":"b038_a06_c4","condition_set_id":"B4","how_blocked":"Heat, blade family, production batch, and customer order are disjoint across partitions."},{"blocked_by_condition_id":"b038_a06_c5","condition_set_id":"B5","how_blocked":"There was one sealed evaluation and no recurring benchmark feedback."},{"blocked_by_condition_id":"b038_a06_c6a","condition_set_id":"B6","how_blocked":"The buyer's flag existed at the forecast timestamp and was not produced by later correction or retrospective reconstruction."},{"blocked_by_condition_id":"b038_a06_c7","condition_set_id":"B7","how_blocked":"Timestamped supplier message packets preserve the complete operational input for each lot."}],"case_id":"E15A002__TRANSFER","case_type":"TRANSFER_POSITIVE","domain":"Aerospace component acceptance forecasting","intended_route_id":"B2","omitted_condition_id":null,"remedy_leakage_audit":"The description identifies the discrepancy and factual data path without suggesting an intervention or naming the underlying design pattern.","route_evidence":[{"condition_id":"b038_a06_leak_core","intended_status":"SATISFIED","scenario_evidence":"A buyer-confidential flag absent from supplier messages enters the development and evaluation table."},{"condition_id":"b038_a06_c2","intended_status":"SATISFIED","scenario_evidence":"The buyer's release flag is the direct source of the acceptance outcome and is included among the estimator's inputs."}],"scenario_text":"An aerospace forging supplier developed a scoring system to estimate whether a customer would accept each turbine-blade lot. The retrospective study showed near-perfect discrimination, but estimates issued from the supplier’s production console were unreliable. A joined study table contained a buyer-side “release flag.” The customer’s laboratory system generated this flag from its internal disposition logic, and the reported acceptance outcome was copied directly from it. Although the flag already existed inside the customer’s system when the supplier issued its estimate, it was confidential and never appeared in the supplier’s incoming message packet. The historical data exchange nevertheless included it among the candidate measurements. All other columns came from pre-decision metallurgical records; nothing was appended after acceptance. Validation lots were separated by heat, blade family, production batch, and customer order. Transformations used only training lots, the evaluation was scored once, and no public comparison table was consulted. Timestamped operational packets are complete. Engineers are trying to explain why the study appears decisive while the same scoring code struggles when fed the fields actually delivered to the supplier.","scenario_title":"Turbine-Lot Acceptance Scores Collapse in Production","vocabulary_separation_audit":"The transfer case replaces medical-administrative language with forging, metallurgy, turbine lots, laboratory release, and supplier-customer data exchange terminology."},{"alternative_route_blocks":[{"blocked_by_condition_id":"b038_a06_c1","condition_set_id":"B1","how_blocked":"Predictor measurements occupy a separate pre-test file; outcome labels remain in a label vault rather than a mixed-lifecycle dataset."},{"blocked_by_condition_id":"b038_a06_c3","condition_set_id":"B3","how_blocked":"Model fitting and preprocessing finished before the evaluator opened the sealed labels, so neither could observe evaluation material."},{"blocked_by_condition_id":"b038_a06_c4","condition_set_id":"B4","how_blocked":"Evaluation lots are disjoint from training by heat, batch, blade design, and customer order."},{"blocked_by_condition_id":"b038_a06_c5","condition_set_id":"B5","how_blocked":"The labels were opened once after predictions were frozen, with no development feedback or repeated benchmark consultation."},{"blocked_by_condition_id":"b038_a06_c6a","condition_set_id":"B6","how_blocked":"Predictions were recorded prospectively from contemporaneous packets rather than reconstructed later from corrected records."},{"blocked_by_condition_id":"b038_a06_c7","condition_set_id":"B7","how_blocked":"Signed message archives preserve the exact inputs available for every operational estimate."}],"case_id":"E15A002__NEAR_MISS","case_type":"ONE_LITERAL_NEAR_MISS","domain":"Aerospace component acceptance forecasting","intended_route_id":"B2","omitted_condition_id":"b038_a06_c2","remedy_leakage_audit":"The statement presents the scoring anomaly and explicit exclusions only; it does not propose a remedy, intervention, or named archetype.","route_evidence":[{"condition_id":"b038_a06_leak_core","intended_status":"SATISFIED","scenario_evidence":"Sealed outcomes influence the validation score through label-dependent stratum weights chosen after predictions are frozen."},{"condition_id":"b038_a06_c2","intended_status":"CONTRADICTED","scenario_evidence":"Every predictor is recorded before testing, and none is calculated from, copied from, or updated after the acceptance outcome."}],"scenario_text":"An aerospace supplier ran a prospective evaluation of a difficult estimator for turbine-blade lot acceptance. Every prediction was frozen before destructive testing. The input packet contained only chemistry, furnace telemetry, ultrasonic measurements, and pre-test dimensions. Audit timestamps confirm that every value existed before the test; none was calculated from the pass/fail result, copied from a disposition field, or revised afterward. Nevertheless, the reported evaluation score is implausibly strong. After opening the sealed outcome file, the evaluator selected separate weights for alloy, customer, and defect-severity strata, trying several combinations internally and retaining the one that made the frozen predictions look best. This weighting affected only the reported score, not model fitting, feature processing, or the submitted predictions. Training and evaluation lots share no heat, batch, blade design, or customer order. The outcome file was opened only for this final scoring exercise, so developers received no interim results. Predictor files and labels remained separate, and signed message archives preserve the exact operational packet for every estimate. The concern is confined to how unavailable outcome knowledge entered the final assessment despite the estimator itself using only timely measurements.","scenario_title":"Sealed Blade Trial Produces an Implausible Score","vocabulary_separation_audit":"The near miss preserves the transfer domain and technical difficulty but replaces the buyer-status input mechanism with outcome-informed final scoring, while explicitly excluding label-derived and post-event predictors."}],"experiment_id":"eoa_inverse_innovation_exp15_route_aware_retrieval40_20260814","sample_id":"E15A002","schema_version":1}