{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp09_archetype_breadth150_20260804","research_id":"eoa_inverse_innovation_exp09_light_prior_art_20260804","cell_id":"progressive_disclosure__computer_science","search_lanes":{"direct_problem_and_intervention":{"queries":["continuous integration failure diagnosis logs first error downstream noise root cause study","CI failure root cause diagnosis layered evidence logs","CI failure root cause analysis product summary relevant log lines raw log evidence GitLab Duo"],"source_ids":["SRC1","SRC3","SRC4"],"no_result_note":null},"synonyms_and_historical_terms":{"queries":["build failure triage first error downstream errors log folding","CI pipeline troubleshooting root cause log drill down","continuous integration failure diagnosis product root cause analysis evidence log links"],"source_ids":["SRC1","SRC2","SRC4"],"no_result_note":null},"products_practices_and_standards":{"queries":["site:docs.github.com actions workflow run logs failed job annotations raw logs","site:docs.gitlab.com CI job logs collapsible sections raw log","GitLab Duo AI-powered CI/CD pipeline root cause analysis"],"source_ids":["SRC1","SRC2","SRC3"],"no_result_note":null},"component_combination":{"queries":["CI logs AI root cause analysis supporting evidence raw log links","pipeline failure summary relevant log lines contextual assistance","CI build failure log progressive disclosure interface root cause summary drilldown"],"source_ids":["SRC1","SRC3","SRC4"],"no_result_note":null}},"sources":[{"source_id":"SRC1","title":"Using workflow run logs","publisher":"GitHub Docs","url":"https://docs.github.com/en/actions/how-tos/monitor-workflows/use-workflow-run-logs","source_type":"OFFICIAL_GUIDANCE","claims_supported":["GitHub Actions organizes run records by jobs and steps, automatically expands failed steps, and lets investigators search build logs.","Investigators can permalink individual log lines and download log archives, providing source-level traceability and access to fuller records."]},{"source_id":"SRC2","title":"CI/CD job logs","publisher":"GitLab Docs","url":"https://docs.gitlab.com/ci/jobs/job_logs/","source_type":"OFFICIAL_GUIDANCE","claims_supported":["GitLab describes a job log as the full execution history of a CI/CD job and directs investigators to scroll through it for details.","GitLab supports collapsible command or custom sections, default-collapsed sections, timestamps, and a complete raw-log view."]},{"source_id":"SRC3","title":"Developing GitLab Duo: Blending AI and Root Cause Analysis to fix CI/CD pipelines","publisher":"GitLab","url":"https://about.gitlab.com/blog/developing-gitlab-duo-blending-ai-and-root-cause-analysis-to-fix-ci-cd/","source_type":"FIRST_PARTY_PRODUCT","claims_supported":["GitLab Duo analyzes a portion of a failed job log, summarizes a proposed root cause, and suggests a fix within the GitLab interface.","GitLab reports that pipeline failures span code, dependency, test, infrastructure-as-code, deployment, timeout, and environment problems, while manual review of mixed application and system messages can be challenging and time-consuming.","Users can continue with follow-up questions, but the described feature does not require contrastive inspection or automatically widen disclosure when signals conflict."]},{"source_id":"SRC4","title":"LLM-Based Automated Diagnosis Of Integration Test Failures At Google","publisher":"arXiv","url":"https://arxiv.org/abs/2604.12108","source_type":"PRIMARY_RESEARCH","claims_supported":["The study reports that integration-test failures produce massive, heterogeneous, unstructured logs with high cognitive load and low signal-to-noise ratio, making diagnosis difficult and time-consuming.","Auto-Diagnose generates concise diagnoses with the most relevant log lines, converts cited lines into links for further investigation, and presents results contextually in Google's code-review workflow.","The authors evaluated 71 real failures with expert review and reported 90.14% diagnosis accuracy; the paper also documents failures caused by missing component or driver logs."]}],"problem_evidence":{"status":"SUPPORTED","finding":"The problem is visible. Official product documentation requires investigators to locate failed jobs or steps and search, scroll, expand, or download logs. GitLab describes mixed application and system messages, varied failure domains, repeated iterations, and context switching; Google's primary study directly reports large heterogeneous logs, high cognitive load, low signal-to-noise ratio, and time-consuming diagnosis. The sources do not separately quantify confusion between initiating errors and downstream symptoms in general CI populations.","source_ids":["SRC1","SRC2","SRC3","SRC4"]},"closest_prior_art":[{"name":"Google Auto-Diagnose","source_ids":["SRC4"],"overlap":"Produces an in-workflow failure diagnosis, concise summary, selected relevant evidence, uncertainty when information is insufficient, and links from extracted evidence back into logs.","remaining_difference":"The reported interface presents an automated finding rather than a user-navigated ladder whose layers change with an explicit diagnostic hypothesis and that obligatorily widens when competing evidence, risk conditions, or ambiguity appear."},{"name":"GitLab Duo Root Cause Analysis","source_ids":["SRC3"],"overlap":"Analyzes failed CI job logs, summarizes a root cause, proposes remediation, supports follow-up questions, and addresses several code, dependency, configuration, and infrastructure failure classes.","remaining_difference":"The documented product forwards only a log portion and centers on an AI answer and suggested fix; it does not document raw-offset breadcrumbs, a persistent complete-record escape hatch, mandatory competing-hypothesis validation, or risk-triggered disclosure."},{"name":"GitHub and GitLab hierarchical CI log viewers","source_ids":["SRC1","SRC2"],"overlap":"Provide progressive expansion, failed-step emphasis, search, timestamps or line links, and access to downloaded or raw logs.","remaining_difference":"Disclosure follows static workflow or authored section structure rather than the investigator's changing hypothesis, and no documented trigger forces contradictory or critical evidence into view before a diagnostic decision."}],"prior_art_disposition":"ADJACENT_PRIOR_ART","contrastive_claim_remaining":"Compared with automatic failed-step expansion, static collapsible sections, and AI-generated diagnoses with selected log lines, a read-only console that changes evidence depth according to the investigator's explicit failure hypothesis and obligatorily widens to source-linked contradictory or risk evidence will increase correct retrieval of an adjudicated initiating cause without reducing discovery of critical exceptions, orientation, or access to the complete record.","contrastive_claim_falsifier":"The claim is falsified if, on adjudicated failed runs, the ladder does not outperform static job-and-step folding on correct cause identification and decisive-evidence retrieval, or if it increases anchoring on a wrong first-layer candidate, missed conflicts or warnings, orientation loss, or inability to reach the complete raw record.","gates":{"adequate_source_search":{"status":"PASS","rationale":"The bounded search covered the proposal directly, synonymous failure-triage and root-cause terminology, major CI log interfaces, AI diagnostic products, and combinations of summaries, relevant-line extraction, source links, folding, and raw-log access. Four opened sources from GitHub, GitLab, and arXiv were retained, including official, first-party, and primary sources.","source_ids":["SRC1","SRC2","SRC3","SRC4"]},"supported_problem":{"status":"PASS","rationale":"Official and primary sources directly show large or mixed CI/test logs, manual searching and scrolling, varied failure domains, low signal-to-noise ratio, cognitive load, and time-consuming diagnosis.","source_ids":["SRC1","SRC2","SRC3","SRC4"]},"distinct_testable_claim":{"status":"PASS","rationale":"The surviving contrast is narrower than summary, folding, or evidence linking alone: hypothesis-conditioned disclosure with mandatory widening on conflict, ambiguity, or risk. Its outcomes and failure modes are measurable against the documented adjacent interfaces.","source_ids":["SRC1","SRC2","SRC3","SRC4"]},"bounded_next_test":{"status":"PASS","rationale":"A read-only prototype can be compared with a static chronological or folded-log view using a small sanitized set of historical failures with independently adjudicated causes. The test can record cause accuracy, decisive-line retrieval, conflict and warning detection, orientation, and complete-log access without deployment or effect-size claims.","source_ids":["SRC1","SRC2","SRC4"]},"no_obvious_safety_or_authority_stop":{"status":"PASS","rationale":"The proposed first test is non-operational, access-controlled, and limited to sanitized copied records; it grants no merge, rerun, remediation, or pipeline authority. Existing products demonstrate read-restricted log inspection and fuller-record access. Anchoring, extraction of sensitive values, missing warnings, and broken source traceability remain explicit halt criteria rather than unavoidable stops.","source_ids":["SRC1","SRC2"]}},"screen_survival":true,"world_novelty_boundary":"This four-source public-web screen found close components and adjacent systems, especially Google Auto-Diagnose and GitLab Duo, but not the full claimed combination of hypothesis-conditioned progressive disclosure, mandatory conflict/risk escalation, source-oriented breadcrumbs, and a persistent complete-record escape hatch. The bounded search cannot establish world novelty, patentability, market size, expert acceptance, or realized value; patent databases, observability and AIOps products, internal enterprise tooling, and broader HCI/debugging literature could contain closer prior art."}