{"cases":[{"alternative_route_blocks":[],"case_id":"E14A031__DIRECT","case_type":"DIRECT_POSITIVE","domain":"Industrial ceramic glaze formulation","intended_route_id":"B2","omitted_condition_id":null,"remedy_leakage_audit":"The statement reports formulation behavior and comparative trial evidence without recommending any intervention or naming a solution pattern.","route_evidence":[{"condition_id":"b037_a07_c2","intended_status":"SATISFIED","scenario_evidence":"Matched trials beginning from different recipes finish with a large, reproducible difference in defect rate despite identical materials, equipment, evaluation rules, and trial budgets."},{"condition_id":"b037_a07_c3","intended_status":"SATISFIED","scenario_evidence":"Single-ingredient adjustments cease to help, while a superior recipe differs in several ingredients and is separated by individually worse intermediate batches."}],"scenario_text":"A ceramic tile plant uses an automated formulation program to reduce pinholes in a new matte glaze. Each trial changes one ingredient slightly, fires a test batch, and retains the change only when the measured defect rate falls without violating color and viscosity limits. Engineers ran two matched campaigns with identical raw materials, kiln schedules, measurement rules, and forty-batch budgets. The campaign initialized from the plant’s standard glossy-glaze recipe finished at 11.8% defective tiles. The one initialized from an unusual university recipe finished at 1.6%, a difference repeated across three production weeks.\n\nLogs from the poorer campaign show that it eventually reached a recipe for which every permitted one-ingredient adjustment either increased defects or breached a limit. Nevertheless, archived experiments contain a 2.1% recipe using the same ingredients and limits. That recipe differs in four ingredient levels. Each recorded path from the stalled recipe toward it includes at least one intermediate batch with a higher defect rate, so the program rejects the path before reaching the better recipe.","scenario_title":"Glaze Trials Depend on the Initial Recipe","vocabulary_separation_audit":"The case uses plant-specific language about glaze ingredients, firing, viscosity, and tile defects rather than abstract spatial or cognitive terminology."},{"alternative_route_blocks":[],"case_id":"E14A031__TRANSFER","case_type":"TRANSFER_POSITIVE","domain":"Hospital nursing roster construction","intended_route_id":"B2","omitted_condition_id":null,"remedy_leakage_audit":"The statement contains observed roster outcomes and constraints only; it neither proposes a corrective action nor describes a named solution archetype.","route_evidence":[{"condition_id":"b037_a07_c2","intended_status":"SATISFIED","scenario_evidence":"The same scheduling procedure, data, limits, and runtime produces sharply different penalty totals when begun from the previous roster versus a manually drafted roster."},{"condition_id":"b037_a07_c3","intended_status":"SATISFIED","scenario_evidence":"No single swap improves the poorer roster, although a much better roster exists and differs through a coordinated reassignment whose partial stages score worse."}],"scenario_text":"A regional hospital generates monthly nursing rosters by repeatedly swapping one nurse’s shift or exchanging two assignments whenever the change lowers a penalty score. The score combines uncovered specialties, overtime, consecutive nights, and denied leave under fixed contractual limits. For September, analysts replayed the same staffing data and two-hour runtime from two initial rosters. Beginning with August’s roster produced a final score of 184. Beginning with a manually drafted roster produced 61. The gap persisted when the replay was repeated on an isolated server with the same tie-breaking rules.\n\nThe 184-point roster was not merely unfinished: every legal single reassignment or pairwise exchange available from it raised the score. A roster scoring 73 had already been verified by payroll and ward managers, but it differed through six linked reassignments. Applying any proper subset of those reassignments either temporarily left an intensive-care shift uncovered or increased consecutive-night penalties. Consequently, the accepted sequence of small improvements stopped at the poorer roster even though a substantially better feasible roster was present among the evaluated records.","scenario_title":"September Roster Quality Hinges on the Draft Used","vocabulary_separation_audit":"The transfer case replaces laboratory formulation language with staffing, shifts, specialties, overtime, leave, and contractual scheduling constraints."},{"alternative_route_blocks":[{"blocked_by_condition_id":"b037_a07_c1","condition_set_id":"B1","how_blocked":"The scenario explicitly states that improvement never stalls: before the minimum-score roster, every roster has an admissible one-swap improvement, and no plateau or dead end appears."},{"blocked_by_condition_id":"b037_a07_c5","condition_set_id":"B3","how_blocked":"The department evaluates rosters with one authoritative scalar measure under fixed eligibility rules, with no competing objectives or proxy measures."}],"case_id":"E14A031__NEAR_MISS","case_type":"ONE_LITERAL_NEAR_MISS","domain":"University examination invigilator scheduling","intended_route_id":"B2","omitted_condition_id":"b037_a07_c3","remedy_leakage_audit":"The statement diagnoses a timing-dependent outcome through recorded behavior but offers no recommendation, intervention, or archetype description.","route_evidence":[{"condition_id":"b037_a07_c2","intended_status":"SATISFIED","scenario_evidence":"Under an identical thirty-minute limit, two starting timetables produce final violation counts of 38 and 7, demonstrating a strong initialization effect on the reached result."},{"condition_id":"b037_a07_c3","intended_status":"CONTRADICTED","scenario_evidence":"Audit and exhaustive checks establish that every nonminimal timetable has an improving admissible single swap; runs improve continuously and stop only when time expires, with no trap or plateau."}],"scenario_text":"A university assigns invigilators to examination rooms with software that makes one admissible assignment swap at a time whenever the swap reduces the number of rule violations. Eligibility, availability, room coverage, and rest requirements are fixed before each run. The registrar uses one authoritative count of violations; there are no secondary targets or surrogate scores.\n\nFor the spring session, two runs used identical records, deterministic tie-breaking, and a thirty-minute limit. Starting from last year’s timetable left 38 violations when time expired, while starting from a recent hand-built timetable left 7. Thus, the timetable supplied at launch strongly affects the result delivered to staff.\n\nHowever, the audit trail shows uninterrupted progress in both runs until the cutoff. Exhaustive checking of every timetable encountered found at least one admissible single swap that reduced the violation count whenever the count was above the verified minimum. There were no flat stretches, dead ends, or occasions when a worsening swap was needed to continue. Longer replays from both starting timetables eventually reached the same minimum-count timetable; the poorer scheduled output arose solely because it began farther away and the clock expired.","scenario_title":"Invigilator Runs Differ at the Deadline but Never Get Stuck","vocabulary_separation_audit":"The near miss retains the transfer case’s constrained personnel-scheduling difficulty while using examination rooms, invigilator eligibility, rest rules, and deadline-limited processing; it avoids glaze terminology and abstract labels."}],"experiment_id":"eoa_inverse_innovation_exp14_applicability_retrieval40_20260813","sample_id":"E14A031","schema_version":1}