{"cases":[{"alternative_route_blocks":[],"case_id":"E15A005__DIRECT","case_type":"DIRECT_POSITIVE","domain":"Diagnostic radiology imaging","intended_route_id":"B1","omitted_condition_id":null,"remedy_leakage_audit":"The problem statement reports observations, operational consequences, and missing interpretive context without proposing any corrective action.","route_evidence":[{"condition_id":"b039_a10_c1","intended_status":"SATISFIED","scenario_evidence":"Radiologists use the reconstructed images and displayed measurements to make follow-up and biopsy decisions."},{"condition_id":"b039_a10_c2","intended_status":"SATISFIED","scenario_evidence":"The workstation's reconstruction process predictably sharpens some boundaries and broadens others according to density, size, and chest region."},{"condition_id":"b039_a10_c3","intended_status":"SATISFIED","scenario_evidence":"Repeat scans and phantom studies reproduce the error, and its magnitude can be charted by nodule size, density, and location."},{"condition_id":"b039_a10_c6","intended_status":"SATISFIED","scenario_evidence":"The interface presents derived measurements as ordinary millimeter values without indicating reconstruction-dependent limits, so clinicians treat them as anatomical facts."}],"scenario_text":"A regional hospital recently consolidated chest CT studies from several scanner models onto one reading workstation. Radiologists use its reconstructed images and automatically displayed nodule diameters to decide whether patients remain under surveillance or proceed to biopsy. Since the change, small low-density nodules near the lung edge appear slightly larger, while dense central nodules appear slightly smaller than they do in the scanners’ native reconstructions. Repeat scans of calibration phantoms reproduce the same direction and magnitude of error. The difference grows below six millimeters, changes with tissue density, and is strongest near the pleural boundary, allowing physicists to chart a stable profile by size, density, and location. The workstation applies one reconstruction process to all incoming studies, sharpening some boundaries and broadening others. Its viewer presents the resulting contours and measurements as ordinary millimeter readings, with no indication that they depend on this processing. Clinicians reviewing longitudinal studies have therefore interpreted apparent growth or stability as a change in the patient rather than a property of the workstation’s rendering.","scenario_title":"Nodule Measurements Shift After Workstation Consolidation","vocabulary_separation_audit":"This case uses radiology-specific language concerning CT reconstruction, nodules, tissue density, phantoms, and biopsy decisions; it does not reuse the procurement vocabulary of the transfer cases."},{"alternative_route_blocks":[],"case_id":"E15A005__TRANSFER","case_type":"TRANSFER_POSITIVE","domain":"Municipal procurement review","intended_route_id":"B1","omitted_condition_id":null,"remedy_leakage_audit":"The statement contains no recommendation, intervention, or proposed redesign; it only describes the procurement failure and its evidentiary pattern.","route_evidence":[{"condition_id":"b039_a10_c1","intended_status":"SATISFIED","scenario_evidence":"Evaluation panels rely on generated comparison cards when scoring bids and selecting finalists."},{"condition_id":"b039_a10_c2","intended_status":"SATISFIED","scenario_evidence":"The summarization process consistently converts conditional commitments into categorical entries and compresses long, cross-referenced qualifications."},{"condition_id":"b039_a10_c3","intended_status":"SATISFIED","scenario_evidence":"Repeated processing gives the same results, with a measurable pattern concentrated in long proposals, cross-referenced exceptions, and certain service categories."},{"condition_id":"b039_a10_c6","intended_status":"SATISFIED","scenario_evidence":"Panelists treat polished card entries as vendor commitments even though the cards do not disclose their loss of qualifications and exceptions."}],"scenario_text":"A city purchasing office receives hundred-page proposals for outsourced services and feeds them into software that produces one-page comparison cards. Evaluation panels score vendors from those cards and use the rankings to choose finalists, consulting the original proposals only when a card contains an obvious gap. In several completed procurements, promises that were conditional on staffing levels, contract length, or customer-supplied data appeared on the cards as unconditional commitments. Tests using the same proposal produce the same card each time. The effect is strongest in long submissions, in sections containing cross-referenced exceptions, and in bids for services whose performance terms are spread across technical and commercial volumes. Short, self-contained qualifications are usually retained. Analysts can quantify how often conditions disappear by proposal length, clause location, and service category. The software’s compression rules favor concise affirmative sentences and merge related passages into a single entry. Because every entry uses polished declarative language and sits beside directly quoted prices and dates, panelists commonly treat the entire card as a factual account of what each vendor offered. The cards do not distinguish quoted commitments from condensed interpretations.","scenario_title":"Bid Conditions Disappear from Evaluation Cards","vocabulary_separation_audit":"The transfer case shifts from medical imaging to public purchasing and uses bids, clauses, comparison cards, vendor commitments, and panel scoring rather than spatial or measurement terminology."},{"alternative_route_blocks":[],"case_id":"E15A005__NEAR_MISS","case_type":"ONE_LITERAL_NEAR_MISS","domain":"Municipal procurement review","intended_route_id":"B1","omitted_condition_id":"b039_a10_c3","remedy_leakage_audit":"The scenario supplies diagnostic facts and explicit negative evidence about reproducibility without recommending any response.","route_evidence":[{"condition_id":"b039_a10_c1","intended_status":"SATISFIED","scenario_evidence":"A review board uses generated briefing cards to score submissions and determine which bidders advance."},{"condition_id":"b039_a10_c2","intended_status":"SATISFIED","scenario_evidence":"The mandatory short-card conversion is non-neutral because it condenses, reorders, and paraphrases source material rather than reproducing it."},{"condition_id":"b039_a10_c3","intended_status":"CONTRADICTED","scenario_evidence":"Repeated runs yield different discrepancies, and an audit finds no stable relationship with length, clause type, vendor class, category, or document location."},{"condition_id":"b039_a10_c6","intended_status":"SATISFIED","scenario_evidence":"Board members treat fluent card statements as source-backed commitments although the display does not mark paraphrases or disclose instability."}],"scenario_text":"A county review board uses automatically generated briefing cards to score complex facilities-management bids and determine which firms reach oral presentations. The card generator always converts each submission into fixed sections with short declarative entries, necessarily condensing, reordering, and paraphrasing the source material. During one procurement, a card stated that a bidder included weekend staffing, although the proposal described weekend coverage only as an optional add-on. Board members scored the entry as an included service because the card displayed it in the same style as copied prices and dates and gave no indication that it was a generated paraphrase. However, investigators cannot reproduce the discrepancy. Running the unchanged proposal again alternately preserves the qualification, omits the entire staffing entry, or produces a different wording that remains accurate. Other proposals show similarly shifting outcomes. Across repeated runs, the errors have no stable direction and no measurable association with document length, clause type, vendor class, service category, or passage location. Identical passages placed in identical contexts also yield different results. Thus, while conversion into a compressed card consistently alters the form presented to reviewers, the observed factual deviations are neither repeatable nor profileable.","scenario_title":"Unstable Briefing Cards Influence Bid Scores","vocabulary_separation_audit":"The near miss retains the transfer domain and comparable procurement complexity while using fresh wording; its explicit instability evidence is confined to the omitted repeatability-and-profile condition."}],"experiment_id":"eoa_inverse_innovation_exp15_route_aware_retrieval40_20260814","sample_id":"E15A005","schema_version":1}