{"judgments":[{"pair_id":"E12MQ042","scores_a":{"structural_fidelity":5,"domain_fidelity":5,"causal_coherence":5,"operational_specificity":5,"testability":5,"practicality":4,"contrivance_risk":2},"scores_b":{"structural_fidelity":2,"domain_fidelity":3,"causal_coherence":4,"operational_specificity":5,"testability":4,"practicality":2,"contrivance_risk":5},"winner":"A","reason":"A directly establishes bounded, competent, accountable authority for a genuine shared-facility disposition decision. B is technically articulated but translates legitimacy components into physical interlock analogies too decoratively."},{"pair_id":"E12MQ002","scores_a":{"structural_fidelity":5,"domain_fidelity":4,"causal_coherence":5,"operational_specificity":5,"testability":5,"practicality":2,"contrivance_risk":4},"scores_b":{"structural_fidelity":5,"domain_fidelity":5,"causal_coherence":5,"operational_specificity":5,"testability":5,"practicality":4,"contrivance_risk":2},"winner":"B","reason":"Both genuinely decompose an entangled accounting object with interfaces and reintegration checks, but B targets a common accounting change-control problem with a proportionate shadow test. A's nested mechanical weighing system is unusually elaborate versus simpler inventory-control rivals."},{"pair_id":"E12MQ025","scores_a":{"structural_fidelity":5,"domain_fidelity":5,"causal_coherence":5,"operational_specificity":5,"testability":5,"practicality":4,"contrivance_risk":2},"scores_b":{"structural_fidelity":2,"domain_fidelity":4,"causal_coherence":4,"operational_specificity":5,"testability":4,"practicality":2,"contrivance_risk":5},"winner":"A","reason":"A faithfully implements safe, plural, autonomy-preserving mentor-anchored enculturation for tacit incident judgment and tests the relational contribution against matched rotating experts. B's device-calibration mechanism is coherent but its mentorship mapping is substantially metaphorical."},{"pair_id":"E12MQ012","scores_a":{"structural_fidelity":5,"domain_fidelity":4,"causal_coherence":5,"operational_specificity":5,"testability":5,"practicality":3,"contrivance_risk":3},"scores_b":{"structural_fidelity":5,"domain_fidelity":5,"causal_coherence":5,"operational_specificity":5,"testability":5,"practicality":4,"contrivance_risk":2},"winner":"B","reason":"B directly diagnoses a trial-balance release wave, bounds admission by capacity, coalesces valid retrieval work, and preserves audit independence and urgent exceptions. A is a strong physical realization but is less practically proportionate for the accounting setting."},{"pair_id":"E12MQ022","scores_a":{"structural_fidelity":5,"domain_fidelity":4,"causal_coherence":5,"operational_specificity":5,"testability":5,"practicality":3,"contrivance_risk":3},"scores_b":{"structural_fidelity":5,"domain_fidelity":4,"causal_coherence":4,"operational_specificity":5,"testability":4,"practicality":3,"contrivance_risk":4},"winner":"A","reason":"A has a directly observable single-carton measurement choke, mechanically enforced service-window clearance, and a well-bounded comparative test. B is plausible, but its completion-by-cradle-weight proxy and sole-path cash-office redesign leave more custody and operational assumptions unresolved."},{"pair_id":"E12MQ006","scores_a":{"structural_fidelity":5,"domain_fidelity":5,"causal_coherence":5,"operational_specificity":5,"testability":5,"practicality":4,"contrivance_risk":2},"scores_b":{"structural_fidelity":5,"domain_fidelity":5,"causal_coherence":5,"operational_specificity":4,"testability":4,"practicality":5,"contrivance_risk":2},"winner":"A","reason":"Both are natural herd-dampening designs, but A more completely models the recovery signal, authorization-safe coalescing, capacity feedback, retry behavior, fairness, and bounded comparative replay. B is simpler and practical but tests fewer conditions."}]}