{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp11_mechanism_context_external20_20260804","research_id":"eoa_inverse_innovation_exp11_external_scrutiny_20260804","cell_id":"negative_space_design__computer_science","opaque_id":"negative_space_design__computer_science__B","search_lanes":{"direct_problem":{"queries":["software verification results interface UNKNOWN status developer comprehension study","static analysis warning interface cognitive load developer study","formal verification user interface verdict unknown results"],"source_ids":["SRC1","SRC2"],"no_result_note":"No retained source directly tested whether crowding alone causes misinterpretation when guarantee wording, evidence, ordering, and controls are held constant."},"closest_prior_art":{"queries":["\"verification\" \"white space\" dashboard results","\"formal verification\" \"visual hierarchy\" interface","\"static analysis\" \"whitespace\" results interface","\"status summary\" diagnostics interface spacing"],"source_ids":["SRC1","SRC3","SRC4","SRC6"],"no_result_note":"No close implementation or controlled study was found combining verification-status isolation, duplicate-chrome omission, preserved evidence, and an explicit UNKNOWN-safety endpoint."},"historical_terminology":{"queries":["older terminology screen crowding display density human factors visual search","\"visual clutter\" user interface task performance study whitespace","\"white space\" interface usability experiment grouping proximity"],"source_ids":["SRC2","SRC3","SRC4"],"no_result_note":null},"products_practices_standards":{"queries":["site:docs.oasis-open.org/sarif SARIF result level kind pass fail notApplicable open standard","site:mathworks.com Polyspace verification result colors unproven unknown dashboard","site:w3.org/WAI/WCAG22 spacing whitespace visual presentation"],"source_ids":["SRC5","SRC6","SRC7"],"no_result_note":null},"non_english_regional":{"queries":["Ergebnisse formale Verifikation Benutzeroberfläche unbekannt Status Übersicht","interface résultats vérification formelle statut inconnu tableau de bord","検証結果 未知 ステータス ダッシュボード 形式検証 UI","interfaz resultados verificación estado desconocido espacio en blanco usabilidad"],"source_ids":["SRC7","SRC8"],"no_result_note":"The Japanese testing syllabus establishes regional terminology and dashboard/reporting practice, but no non-English retained source tested the proposed verification-specific whitespace mechanism."},"composition_subproblems":{"queries":["static analysis warnings visual salience grouping developer decisions experiment","security dashboard alert fatigue visual clutter developer findings missed","formal methods results explanation user study counterexample raw output","unknown verdict user interface unsafe inference software analysis"],"source_ids":["SRC1","SRC2","SRC3","SRC5","SRC6"],"no_result_note":"The component literatures cover result comprehension, perceptual grouping, clutter, explicit indeterminate semantics, and product dashboards separately; no retained source evaluated the full composition."}},"sources":[{"source_id":"SRC1","title":"A user study for evaluation of formal verification results and their explanation at Bosch","url":"https://link.springer.com/article/10.1007/s10664-023-10353-4","publisher":"Springer Nature, Empirical Software Engineering","date_or_year":"2023","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Automotive engineers are identifiable users of formal-verification results and safety analysis.","Only 6 of 13 participants regarded highlighted raw model-checker output as easy to understand, versus 12 of 13 for a structured counterexample explanation.","Task performance and participant reports indicate that presentation can affect understanding of verification output.","The study changed explanation content and highlighting and used a one-group pretest-posttest design, so it does not establish a spacing-only effect."]},{"source_id":"SRC2","title":"Display clutter: a review of definitions and measurement techniques","url":"https://pubmed.ncbi.nlm.nih.gov/25790571/","publisher":"Human Factors and Ergonomics Society / SAGE","date_or_year":"2015","source_type":"SECONDARY_RESEARCH","language":"English","claims_supported":["Display clutter can impose attentional and visual-search performance costs.","Clutter is multifaceted and must be defined through task-performance consequences rather than appearance alone.","The literature identifies a tradeoff between excessive data and insufficient information, supporting preservation of evidence and direct measurement."]},{"source_id":"SRC3","title":"An intuitive model of perceptual grouping for HCI design","url":"https://research.ibm.com/publications/an-intuitive-model-of-perceptual-grouping-for-hci-design","publisher":"IBM Research / ACM","date_or_year":"2009","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Proximity and feature similarity cause visual elements to be perceived as coherent groups.","Perceptual organization is relevant to quick and accurate recognition of key interface components.","This is mechanism-level HCI prior art for using spacing to isolate a status, but it is not verification-specific outcome evidence."]},{"source_id":"SRC4","title":"Reading Online Text: A Comparison of Four White Space Layouts","url":"https://researchinuserexperience.wordpress.com/2004/07/12/reading-online-text-a-comparison-of-four-white-space-layouts/","publisher":"Software Usability Research Laboratory, Wichita State University; RUX archive","date_or_year":"2004","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Manipulating margin whitespace affected comprehension and reading speed: larger margins improved comprehension but slowed reading.","Line leading did not improve performance in that experiment, illustrating that not every spacing manipulation works.","The source reports mixed prior recommendations, including limiting whitespace for scanning and search, providing counterevidence to a universal sparse-is-better claim."]},{"source_id":"SRC5","title":"Static Analysis Results Interchange Format (SARIF) Version 2.1.0","url":"https://docs.oasis-open.org/sarif/sarif/v2.1.0/os/sarif-v2.1.0-os.html","publisher":"OASIS Open","date_or_year":"2020","source_type":"OFFICIAL_STANDARD","language":"English","claims_supported":["Static-analysis results already have standardized semantic distinctions including pass, fail, open, review, informational, and notApplicable.","The standard defines open as insufficient information to decide whether a problem exists, directly supporting the rule that UNKNOWN-like results must not imply safety.","Result severity and priority are distinct from result kind, so a display must preserve the authoritative status contract rather than rely on salience alone."]},{"source_id":"SRC6","title":"Interpret Results — Polyspace Access","url":"https://www.mathworks.com/help/polyspace_access/interpret-results.html","publisher":"MathWorks","date_or_year":"Accessed 2026-08-04","source_type":"FIRST_PARTY_PRODUCT","language":"English","claims_supported":["A deployed verification product distinguishes proven error, unproven possible error, and proven absence of error.","Polyspace combines overview dashboards with access to result details and navigation aids.","Product practice shows identifiable verification-service owners and developer reviewers, but the documentation supplies no controlled evidence for protected whitespace."]},{"source_id":"SRC7","title":"Web Content Accessibility Guidelines (WCAG) 2.2","url":"https://www.w3.org/TR/WCAG22/","publisher":"World Wide Web Consortium (W3C)","date_or_year":"2023","source_type":"OFFICIAL_STANDARD","language":"English","claims_supported":["Color must not be the sole visual means of conveying information or state.","Information, structure, relationships, and meaningful sequence must remain available independently of presentation.","Interfaces must tolerate specified text-spacing changes without losing content or functionality, supporting the proposal's text-label and evidence-preservation guardrails."]},{"source_id":"SRC8","title":"JSTQB Foundation Level Syllabus Version 4.0 J02","url":"https://www.jstqb.jp/dl/JSTQB-SyllabusFoundation_VersionV40.J02.pdf","publisher":"Japan Software Testing Qualifications Board (JSTQB)","date_or_year":"2024","source_type":"OFFICIAL_GUIDANCE","language":"Japanese","claims_supported":["Japanese testing guidance identifies dashboards, including CI/CD dashboards, as established ways to communicate test status.","It states that reporting should be tailored to stakeholder information needs.","It references ISO/IEC/IEEE 29119-3 status and completion reporting, showing that result communication has identifiable organizational stakeholders and established regional practice."]}],"problem_evidence":{"status":"PARTLY_SUPPORTED","finding":"The broader problem is real: formal-verification output can be difficult for engineers to interpret, and display clutter can impair attention and visual search. However, no retained source isolates the proposal's exact claim that uninterrupted density causes guarantee and UNKNOWN misinterpretation when wording, evidence, order, and controls are identical.","source_ids":["SRC1","SRC2"],"uncertainty":"Bosch studied changed explanations and highlighting with only 13 experimental participants; the clutter review spans other data-rich domains. Baseline error prevalence for the four proposed verification strata remains unknown."},"adopter_evidence":{"status":"SUPPORTED","finding":"Identifiable adopters and authorizers exist: automotive engineers and safety analysts consume formal-verification output; Polyspace serves developers reviewing proven, unproven, and error results; test-management guidance identifies teams and stakeholders receiving status reports. A verification-system owner and software-safety authority are plausible accountable authorizers for an offline study.","source_ids":["SRC1","SRC6","SRC8"],"uncertainty":"The sources establish adopter classes and organizational roles, not commitment by a named organization to run this particular study."},"implementation_evidence":{"status":"PARTLY_SUPPORTED","finding":"The components are feasible and individually grounded: proximity can visually group status information, whitespace can affect comprehension, deployed products separate semantic result classes and details, and accessibility standards require preserved labels, structure, and functionality. Evidence does not establish the proposed eight-point dual-endpoint effect or superiority to typography-only emphasis.","source_ids":["SRC3","SRC4","SRC5","SRC6","SRC7"],"uncertainty":"Whitespace can slow reading or impair scanning, experience may moderate benefit, and additional vertical space may push evidence below the fold. The specific reserved margins, pacing gap, and chrome omissions require prototyping and accessibility testing."},"prior_art":{"disposition":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Bosch counterexample-explanation interface study","source_ids":["SRC1"],"same_problem":true,"same_causal_lever":false,"overlap":"Targets difficulty understanding formal-verification output and measures engineers' interpretation of model-checker results.","remaining_difference":"It changes explanatory content and highlighting rather than holding all evidence and wording constant while manipulating only whitespace and duplicate chrome."},{"name":"Polyspace Access result interpretation and dashboard","source_ids":["SRC6"],"same_problem":true,"same_causal_lever":false,"overlap":"Separates proven error, unproven possible error, and proven absence, while providing overview and drill-down access to diagnostic details.","remaining_difference":"The documentation does not specify protected status margins, a pacing gap, duplicate-chrome omission, or comparative comprehension testing."},{"name":"Perceptual grouping by proximity","source_ids":["SRC3"],"same_problem":false,"same_causal_lever":true,"overlap":"Uses spatial proximity and separation to make coherent visual groups and key interface components perceptible.","remaining_difference":"It is generic HCI mechanism prior art, not an evaluation of verification guarantees, UNKNOWN safety, or release decisions."},{"name":"Online-text whitespace experiment","source_ids":["SRC4"],"same_problem":false,"same_causal_lever":true,"overlap":"Directly manipulates margins and leading and measures comprehension, speed, and preference.","remaining_difference":"It concerns prose reading rather than dense verification statuses and diagnostics; it also identifies speed and scanning tradeoffs."},{"name":"SARIF explicit indeterminate-result semantics","source_ids":["SRC5"],"same_problem":true,"same_causal_lever":false,"overlap":"Standardizes pass, fail, open, and review semantics so insufficient information is not represented as absence of a problem.","remaining_difference":"It specifies data semantics, not visual isolation or evidence-preserving negative-space design."}],"contrastive_claim_remaining":"For displays with identical status wording, evidence, ordering, and recovery controls, reserving whitespace around the authoritative status, removing only duplicate non-evidentiary chrome, and inserting a gap before diagnostics will improve both guarantee-interpretation accuracy and release/remediation decision accuracy by at least eight percentage points versus a dense baseline, outperform typography-only emphasis, keep evidence-recovery accuracy within three points, and not increase unsafe approval on UNKNOWN by more than two points.","contrastive_claim_falsifier":"The claim is falsified if the preregistered controlled study misses either eight-point primary-endpoint threshold, either confidence interval includes zero, performs no better than typography-only emphasis, breaches the evidence-recovery margin, raises unsafe UNKNOWN approvals beyond two points, or fails accessibility checks. Discovery of an earlier controlled verification-interface study or routine product practice implementing the same package with equivalent outcomes would also overturn the adjacent-art disposition.","confidence":"MODERATE","search_limitations":"Bounded public-web search covered six lanes and retained exactly eight opened sources. It did not systematically inspect patents, proprietary interfaces, conference-demo archives, internal design systems, or every regional language. Product documentation was largely textual rather than a complete versioned screenshot inventory. The search cannot establish world novelty, patentability, freedom to operate, market size, or realized impact."},"researchability_gates":{"externally_supported_problem":{"status":"PASS","rationale":"Independent evidence supports both difficulty interpreting formal-verification results and performance costs from display clutter, although their exact proposed interaction remains untested.","source_ids":["SRC1","SRC2"]},"identifiable_adopter_or_authorizer":{"status":"PASS","rationale":"Developers, automotive safety engineers, verification-product owners, test managers, and safety authorities are identifiable users or accountable decision roles.","source_ids":["SRC1","SRC6","SRC8"]},"distinct_testable_incremental_claim":{"status":"PASS","rationale":"Adjacent art does not resolve the spacing-and-chrome-only comparison with fixed semantics, a typography-only rival, dual accuracy endpoints, evidence-recovery noninferiority, and an UNKNOWN-safety bound.","source_ids":["SRC1","SRC3","SRC4","SRC5","SRC6"]},"bounded_next_evidence_step":{"status":"PASS","rationale":"The proposed counterbalanced study is bounded to 60 adjudicated queries, four declared semantic strata, fixed content, preregistered thresholds, and measurable rollback criteria. Existing work demonstrates that engineer-facing verification-result studies are practicable.","source_ids":["SRC1","SRC5"]},"no_unresolved_safety_or_authority_stop":{"status":"PASS","rationale":"The next step is an offline display study; it preserves status contracts, evidence, assumptions, obligations, recovery controls, and text labels. Joint owner/safety approval, accessibility checks, UNKNOWN monitoring, and halt thresholds bound the authority and safety risks.","source_ids":["SRC5","SRC7"]},"adequate_search_evidence":{"status":"PASS","rationale":"The adversarial search covered direct evidence, closest art, historical terminology, standards/products, non-English regional material, and component combinations, using eight direct sources from multiple independent publishers with primary research, standards, official guidance, and first-party product documentation.","source_ids":["SRC1","SRC2","SRC3","SRC4","SRC5","SRC6","SRC7","SRC8"]}},"strict_success":true,"screen_survival":true,"remaining_research_value":"HIGH","recommended_next_step":"Preregister and run the proposed 60-case counterbalanced within-participant study, adding the typography-only rival as an explicit third condition or preregistered pairwise comparison. Pilot viewport and zoom behavior first to ensure the sparse layout does not move evidence below the fold; log guarantee interpretation, release/remediation decisions, evidence recovery, response time, unsafe UNKNOWN approvals, participant experience, and accessibility failures by query stratum.","world_novelty_boundary":"This review supports only a bounded conclusion: no retained source closely matched the full problem–intervention–measurement package, while substantial adjacent HCI, standards, product, and verification-explanation art exists. It does not establish world novelty, patentability, freedom to operate, market size, routine deployability, or realized impact."}