{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp09_archetype_breadth150_20260804","research_id":"eoa_inverse_innovation_exp09_light_prior_art_20260804","cell_id":"agentic_control_loop_design__computer_science","search_lanes":{"direct_problem_and_intervention":{"queries":["CI build failure recovery agent hypotheses safe actions feedback loop","automated CI failure diagnosis remediation agent rerun logs"],"source_ids":["SRC3","SRC4"],"no_result_note":null},"synonyms_and_historical_terms":{"queries":["self healing CI pipeline failure recovery MAPE-K","autonomic computing monitor analyze plan execute CI failure recovery"],"source_ids":["SRC3","SRC4"],"no_result_note":null},"products_practices_and_standards":{"queries":["CI flaky test automatic retry official documentation","GitLab CI automatic retry runner system failure official","self healing CI/CD agent policy allowlist retry escalate"],"source_ids":["SRC2","SRC3"],"no_result_note":null},"component_combination":{"queries":["CI failure root cause analysis safe action feedback autonomous remediation","CI recovery agent allowlist safe actions hypotheses rerun escalate","closed loop agentic AI self healing CI CD automation"],"source_ids":["SRC1","SRC3","SRC4"],"no_result_note":null}},"sources":[{"source_id":"SRC1","title":"Empirical Study of Restarted and Flaky Builds on Travis CI","publisher":"arXiv / software-engineering researchers","url":"https://arxiv.org/abs/2003.11772","source_type":"PRIMARY_RESEARCH","claims_supported":["An empirical analysis found at least 56,522 manually restarted builds, with initial failures commonly involving tests, networks, or CI-service limitations.","Restarted builds interrupted developer workflow and were associated with substantially longer pull-request merge times, supporting the visibility and consequence of ambiguous CI failures."]},{"source_id":"SRC2","title":"The Custom executor","publisher":"GitLab Docs","url":"https://docs.gitlab.com/runner/executors/custom/","source_type":"OFFICIAL_GUIDANCE","claims_supported":["GitLab Runner distinguishes build failures from system failures and automatically retries specified stages for system failures.","The documented behavior exemplifies fixed, failure-code-driven retry policy rather than hypothesis-based selection and updating."]},{"source_id":"SRC3","title":"Helix: Self-Healing CI/CD Agent","publisher":"Neo Research Inc.","url":"https://docs.heyneo.com/projects/helix-self-healing-ci-cd","source_type":"FIRST_PARTY_PRODUCT","claims_supported":["Helix ingests CI failures, classifies them into test, infrastructure, and dependency families, proposes patches or retries with risk scores, and can rerun, fix, or escalate.","Execution is constrained by an allowlist, creating close overlap with the proposal's diagnosis, bounded action, execution, and escalation elements.","The opened description does not document competing causal hypotheses, information-value-based action selection, or explicit action-attributed hypothesis revision."]},{"source_id":"SRC4","title":"A Closed-Loop Agentic AI Framework for Self-Configuring, Self-Healing, and Optimized CI/CD Automation","publisher":"IJRASET","url":"https://www.ijraset.com/research-paper/closed-loop-agentic-ai-framework-for-self-configuring-self-healing","source_type":"PRIMARY_RESEARCH","claims_supported":["DevOps-Pilot is a CI/CD-specific closed-loop framework integrating repository analysis, planning, execution monitoring, recovery actions, execution history, and policy updates from observed rewards.","Its prototype recommends dependency reinstallation, test reruns, and deployment rollback and adjusts caching, retries, timeouts, and parallelization using execution outcomes.","Its nine-run simulation is preliminary and optimizes pipeline policy broadly; the opened paper does not specify the proposal's competing-hypothesis register, trustworthy-build criterion, reversible diagnostic-action menu, or divided repository/platform authority."]}],"problem_evidence":{"status":"PARTLY_SUPPORTED","finding":"The problem is visible at a coarse level. Primary research documents numerous restarted builds caused by heterogeneous test, network, and CI-platform conditions, developer interruption, and delayed merging. Official GitLab documentation separately distinguishes build and system failures and applies fixed retries, while newer systems assemble diagnosis and remediation capabilities. The retained evidence does not directly establish the proposal's claims about premature escalation, cross-tool evidence fragmentation, or blame between service and platform teams.","source_ids":["SRC1","SRC2","SRC3"]},"closest_prior_art":[{"name":"Helix Self-Healing CI/CD Agent","source_ids":["SRC3"],"overlap":"Directly targets CI/CD failures with log ingestion, cause-family classification, risk-scored retry or patch proposals, allowlisted automatic execution, and escalation.","remaining_difference":"The available description does not show an inspectable register of competing hypotheses, selection for expected hypothesis discrimination, prohibition of non-informative repetition, or explicit revision of hypotheses after each action."},{"name":"DevOps-Pilot closed-loop agentic CI/CD framework","source_ids":["SRC4"],"overlap":"Combines CI/CD context analysis, LLM-assisted planning, execution, monitoring, self-healing actions, outcome history, and feedback-based policy improvement in one closed loop.","remaining_difference":"It learns broad configuration policy from rewards and includes persistent changes such as pipeline generation and rollback; it does not document the candidate's narrowly governed per-failure diagnostic loop, trustworthy-pass objective, reversible action boundary, competing causal assumptions, or repository/platform authority split."},{"name":"GitLab Runner system-failure retry behavior","source_ids":["SRC2"],"overlap":"Classifies a runner condition separately from an ordinary build failure and performs bounded automatic retries of selected stages.","remaining_difference":"Retry selection is predetermined by failure codes and stage settings, without cross-source causal modeling, discriminating action choice, or action-effect model updates."}],"prior_art_disposition":"ADJACENT_PRIOR_ART","contrastive_claim_remaining":"For ambiguous CI failures, a per-incident mechanism that jointly maintains inspectable competing causal hypotheses, chooses one preauthorized reversible action for expected hypothesis discrimination, attributes the next observation to that action, updates the hypotheses, blocks unsupported repetition, and produces authority-specific escalation will reduce redundant actions and improve cause discrimination and escalation completeness relative to fixed retries and existing classify-propose-execute systems, without weakening required checks.","contrastive_claim_falsifier":"The claim is falsified if closer evidence shows an existing system already implements all of those CI-specific couplings, or if fixed-budget blinded replay finds no improvement over the baseline in correct cause discrimination, redundant-action rate, prohibited recommendations, or completeness and routing of escalation packages.","gates":{"adequate_source_search":{"status":"PASS","rationale":"The bounded search covered the proposal directly, self-healing and autonomic terminology, official retry practice, CI-specific agent products, empirical failure research, and closed-loop component combinations. Four opened sources from four publisher contexts include primary research, official guidance, and first-party documentation.","source_ids":["SRC1","SRC2","SRC3","SRC4"]},"supported_problem":{"status":"PASS","rationale":"Empirical evidence supports heterogeneous CI failure causes, widespread manual restarts, developer interruption, and delivery delay. Official and product documentation confirms that retry, classification, diagnosis, execution, and escalation capabilities exist but are implemented with materially different degrees of coupling. Some organizational particulars remain unverified, so support is partial.","source_ids":["SRC1","SRC2","SRC3"]},"distinct_testable_claim":{"status":"PASS","rationale":"After accounting for close CI-specific closed-loop and self-healing systems, the remaining claim is narrowed to measurable coupling of competing hypotheses, information-seeking reversible actions, attributed observations, model revision, repetition control, and authority-aware escalation.","source_ids":["SRC2","SRC3","SRC4"]},"bounded_next_test":{"status":"PASS","rationale":"A read-only blinded replay of 30 resolved failures in one consenting repository is fixed in scope and can compare hypothesis discrimination, action redundancy, prohibited advice, and escalation quality against archived baseline behavior without triggering CI jobs or modifying repositories.","source_ids":["SRC1","SRC3","SRC4"]},"no_obvious_safety_or_authority_stop":{"status":"PASS","rationale":"The proposed first test is shadow-only and uses no execution or repository-write authority. Consent, access minimization, artifact redaction, secret scanning, retention limits, and human review are necessary controls, but no obvious safety or authority issue prevents the bounded test.","source_ids":["SRC2","SRC3"]}},"screen_survival":true,"world_novelty_boundary":"This four-source public-web screen found close CI-specific self-healing agents and a closed-loop CI/CD research prototype, but no retained source documented the complete narrowly governed hypothesis-action-observation-update loop claimed here. This bounded result cannot establish world novelty, patentability, market size, expert acceptance, or realized value; patents, internal systems, other products, and unindexed publications may contain a closer match."}