{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp11_mechanism_context_external20_20260804","research_id":"eoa_inverse_innovation_exp11_external_scrutiny_20260804","cell_id":"invariant_mode_decomposition_design__computer_science","opaque_id":"invariant_mode_decomposition_design__computer_science__C","search_lanes":{"direct_problem":{"queries":["microservice autoscaling oscillation coupled services retries queues instability","microservice autoscaling control theory eigenvalues stability coupled services","dynamic mode decomposition cloud autoscaling microservices","\"Kubernetes Microservices Instability\" Jacobian eigenvalue modal participation"],"source_ids":["SRC1","SRC2","SRC3","SRC8"],"no_result_note":null},"closest_prior_art":{"queries":["\"autoscaling\" \"eigenvalues\" cloud service stability","\"microservice\" \"Jacobian\" eigenvalues autoscaling","dependency-aware coordinated autoscaling microservices","system identification coordinated autoscaling microservices queues retries replicas"],"source_ids":["SRC1","SRC3","SRC4","SRC5","SRC8"],"no_result_note":null},"historical_terminology":{"queries":["autoscaling thrashing oscillation coupled tiers cloud historical","MIMO control autoscaling multi tier web application eigenvalue","\"Automated Control of Multiple Virtualized Resources\" Padala AutoControl","eigenspace anomaly detection service dependency graph computer systems"],"source_ids":["SRC5","SRC7"],"no_result_note":null},"products_practices_standards":{"queries":["site:kubernetes.io horizontal pod autoscaler stabilization window oscillation","Kubernetes HorizontalPodAutoscaler prevent flapping stabilization tolerance","production collaborative overload control microservices load shedding dependencies","dependency-aware autoscaling microservice product queue scaling"],"source_ids":["SRC2","SRC4"],"no_result_note":null},"non_english_regional":{"queries":["微服务 自动伸缩 振荡 依赖 协同 控制","微服务 弹性伸缩 特征值 稳定性 控制","escalonamento automático microsserviços oscilação dependências controle","autoscaling microservicios oscilación dependencias control estabilidad"],"source_ids":["SRC4","SRC8"],"no_result_note":null},"composition_subproblems":{"queries":["\"dynamic mode decomposition\" microservices telemetry","service dependency graph eigenmode anomaly detection autoscaling","dynamic mode decomposition with control state transition operator eigenmodes","online model estimator MIMO controller multi-tier resource allocation"],"source_ids":["SRC5","SRC6","SRC7","SRC8"],"no_result_note":null}},"sources":[{"source_id":"SRC1","title":"On the Stability of the Kubernetes Horizontal Autoscaler Control Loop","url":"https://upcommons.upc.edu/server/api/core/bitstreams/34caab31-d883-4b17-a72d-43ed4e49346a/content","publisher":"IEEE Access / Universitat Politècnica de Catalunya repository","date_or_year":"2025","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Kubernetes HPA independently scales interconnected services from local metrics.","The authors explicitly investigate whether individual scaling decisions can collectively produce large-scale oscillations.","A control-theoretic model, simulation, and three-service Kubernetes testbed establish that stability analysis of service graphs is technically feasible.","Under the paper's restricted CPU-intensive, independent-service assumptions, HPA converged after a transient; this limits any universal claim that coupled instability is inevitable."]},{"source_id":"SRC2","title":"Horizontal Pod Autoscaling","url":"https://kubernetes.io/docs/concepts/workloads/autoscaling/horizontal-pod-autoscale/","publisher":"Kubernetes","date_or_year":"2026","source_type":"OFFICIAL_GUIDANCE","language":"English","claims_supported":["HPA calculates desired replicas from workload metrics and supports multiple metrics.","Configurable stabilization windows, scaling-rate policies, and tolerance are official mechanisms for smoothing fluctuating recommendations and preventing overly eager scaling.","Platform operators have concrete authority and configuration surfaces for bounded, reversible scaling changes."]},{"source_id":"SRC3","title":"PBScaler: A Bottleneck-aware Autoscaling Framework for Microservice-based Applications","url":"https://arxiv.org/abs/2303.14620","publisher":"arXiv; authors from academic research institutions","date_or_year":"2023","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Complex microservice interactions propagate performance anomalies and complicate bottleneck localization and scaling decisions.","Repeated online optimization can cause replica and latency oscillations.","PBScaler combines dependency-aware TopoRank bottleneck localization with offline performance-aware scaling optimization and experimentally compares resource and performance outcomes."]},{"source_id":"SRC4","title":"Overload Control for Scaling WeChat Microservices","url":"https://arxiv.org/abs/1806.04075","publisher":"ACM Symposium on Cloud Computing / arXiv","date_or_year":"2018","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Service-specific overload controls can harm the overall system because of intricate service dependencies.","DAGOR performs collaborative load shedding among related microservices rather than relying only on isolated local action.","The approach had an identifiable production adopter: the WeChat backend, where it had reportedly been used for five years."]},{"source_id":"SRC5","title":"Automated Control of Multiple Virtualized Resources","url":"https://shiftleft.com/mirrors/www.hpl.hp.com/techreports/2008/HPL-2008-123R1.html","publisher":"HP Laboratories","date_or_year":"2008","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["AutoControl combined an online model estimator with a multi-input, multi-output resource controller.","The estimator captured relationships between application performance and multiple resource allocations across nodes.","Experiments with multi-tier benchmarks and production-trace-driven workloads showed coordinated control of CPU and disk resources to meet service objectives."]},{"source_id":"SRC6","title":"Dynamic Mode Decomposition with Control","url":"https://arxiv.org/abs/1409.6358","publisher":"arXiv; Proctor, Brunton, and Kutz","date_or_year":"2014","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["DMD with control estimates low-order state and input mappings from state and actuation snapshots.","The method extracts dynamic modes and eigenvalues while separating intrinsic dynamics from applied control.","The paper warns that ordinary DMD modes can be corrupted by external forcing, directly supporting the need to model autoscaler and operator actions as inputs."]},{"source_id":"SRC7","title":"Eigenspace-based Anomaly Detection in Computer Systems","url":"https://research.ibm.com/publications/eigenspace-based-anomaly-detection-in-computer-systems","publisher":"IBM Research / KDD","date_or_year":"2004","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["A multi-node web system can be represented as a time-varying weighted graph of services and dependencies.","Principal-eigenvector analysis was used for online anomaly detection and faulty-service identification.","Spectral analysis of service-dependency telemetry therefore predates modern microservice autoscaling terminology."]},{"source_id":"SRC8","title":"Kubernetes Microservices Instability: Theory, Worked Example, and Defensive Pipeline","url":"https://www.ajms.in/index.php/ajms/article/view/636","publisher":"Asian Journal of Mathematical Sciences","date_or_year":"2026","source_type":"OTHER","language":"English","claims_supported":["The article directly proposes estimating a Jacobian from Kubernetes microservice telemetry, computing eigenvalues, tracking spectral drift, and using modal participation to identify implicated services.","Its worked example uses frontend, API, and database queue-depth or latency proxies and demonstrates eigenvalue migration toward an instability threshold.","It is close diagnostic prior art, but presents a small numerical worked example rather than a comparative production canary of automated mode-targeted control."]}],"problem_evidence":{"status":"SUPPORTED","finding":"The bounded problem exists: independent autoscaling and service-specific overload responses operate inside interconnected service graphs; published experiments report propagated anomalies, replica oscillation, transient traffic loss, and cases where local overload control can be detrimental system-wide. The evidence does not show that every incident has persistent modal structure, and one restricted HPA model instead converged under independent-service assumptions.","source_ids":["SRC1","SRC2","SRC3","SRC4","SRC8"],"uncertainty":"Observed instability may instead arise from workload shocks, discrete defects, delayed startup, topology changes, or optimizer behavior. The frequency and operational importance of reproducible cross-service modes are not established."},"adopter_evidence":{"status":"SUPPORTED","finding":"Platform/SRE teams and affected service owners are identifiable adopters and authorizers because Kubernetes exposes operator-controlled scaling policies, while PBScaler and the production deployment of DAGOR demonstrate that dependency-aware scaling and coordinated overload control are decisions made at the platform/application boundary.","source_ids":["SRC2","SRC3","SRC4"],"uncertainty":"Joint authority across independently owned services may be organizationally difficult, and no evidence establishes willingness to authorize experimental perturbations in a particular production environment."},"implementation_evidence":{"status":"PARTLY_SUPPORTED","finding":"The components are technically grounded: online multivariable model estimation and MIMO control were demonstrated in multi-tier systems; DMDc estimates state and actuation operators and their modes; spectral service-graph anomaly detection is longstanding; and a recent article applies Jacobian eigenanalysis and drift tracking directly to Kubernetes microservices. Dependency-aware scaling and collaborative load shedding have also been implemented. No retained source validates the complete package of locally estimated queue/retry/replica modes, residual and drift checks, and reversible coordinated modal intervention against a dependency-aware rival in a live canary.","source_ids":["SRC3","SRC4","SRC5","SRC6","SRC7","SRC8"],"uncertainty":"Sparse, confounded telemetry, controller delay, nonlinear regime changes, changing topology, and actuation-induced model invalidation may prevent stable mode estimation or reliable mapping from modes to safe controls."},"prior_art":{"disposition":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Kubernetes microservice Jacobian, eigenvalue, modal-participation, and spectral-drift pipeline","source_ids":["SRC8"],"same_problem":true,"same_causal_lever":true,"overlap":"It closely matches the diagnostic core: telemetry-derived local linearization, eigenvalues, spectral drift, instability thresholds, and mapping modes to participating services.","remaining_difference":"It does not establish an automated state-and-input model over queues, retries, throughput, replicas, and controls, nor a guarded comparative canary showing that coordinated mode-targeted actuation improves recovery and protected outcomes."},{"name":"On the Stability of the Kubernetes Horizontal Autoscaler Control Loop","source_ids":["SRC1"],"same_problem":true,"same_causal_lever":false,"overlap":"It models HPA as a control loop, extends analysis to a service graph, tests a three-service Kubernetes deployment, and explicitly considers emergent large-scale oscillation.","remaining_difference":"It assumes independently operating services, analyzes CPU-target HPA convergence rather than estimating empirical interaction modes, and does not design coordinated modal interventions."},{"name":"PBScaler","source_ids":["SRC3"],"same_problem":true,"same_causal_lever":false,"overlap":"It addresses propagated anomalies and replica oscillation using dependency-aware bottleneck localization followed by coordinated scaling optimization.","remaining_difference":"It ranks bottlenecks using topology and anomaly potential rather than estimating a local transition operator and selecting weakly damped or amplified eigenmodes."},{"name":"DAGOR collaborative overload control","source_ids":["SRC4"],"same_problem":true,"same_causal_lever":false,"overlap":"It rejects purely service-local overload handling and coordinates load shedding across relevant dependent microservices in production.","remaining_difference":"It uses overload status and request-accounting rules rather than learned dynamic modes, modal gain, reconstruction residuals, or spectral drift."},{"name":"AutoControl online estimation and MIMO resource control","source_ids":["SRC5"],"same_problem":false,"same_causal_lever":false,"overlap":"It demonstrates online identification of coupled performance/resource relationships and coordinated multivariable resource control across application tiers.","remaining_difference":"Its target is resource allocation for multi-tier SLOs, not hidden microservice autoscaler cascades, and it does not use invariant-mode selection as the intervention coordinate system."},{"name":"Dynamic Mode Decomposition with Control","source_ids":["SRC6"],"same_problem":false,"same_causal_lever":true,"overlap":"It supplies the generic data-driven state-transition, input-map, eigenmode, and reduced-order control machinery required by the proposal.","remaining_difference":"It neither applies the method to microservice telemetry nor validates mode-targeted autoscaling, retry limiting, or load shedding."},{"name":"Eigenspace-based anomaly detection in computer systems","source_ids":["SRC7"],"same_problem":false,"same_causal_lever":false,"overlap":"It uses eigenstructure of a time-varying service-dependency graph to detect anomalies and identify services.","remaining_difference":"It analyzes graph-activity eigenvectors for detection, not eigenmodes of a dynamical transition operator for coordinated control."}],"contrastive_claim_remaining":"Within a preregistered topology/load window, a control-aware local transition model can reveal a reproducible, separated risky interaction mode whose targeted coordinated scaling, retry limiting, or load shedding reduces oscillation energy and recovery time relative to dependency-aware anomaly detection plus coordinated runbooks, without unacceptable error rate, tail latency, saturation, or cost.","contrastive_claim_falsifier":"The claim is falsified if no reproducible separated risky mode appears; if per-service variables and exogenous-event labels predict propagation equally well; if modal participation fails to identify controllable services; or if a guarded modal intervention does not outperform the dependency-aware rival on recovery, oscillation, and protected outcomes. Discovery of an earlier routinely deployed system implementing this entire diagnostic-and-intervention package would also overturn the disposition.","confidence":"MODERATE","search_limitations":"The bounded search covered direct formulations, adjacent prior art, older MIMO/eigenspace terminology, official Kubernetes practice, Chinese/Portuguese/Spanish terminology, and component combinations. It retained exactly eight direct sources. It was not an exhaustive patent, proprietary-product, dissertation, or citation-network search; one unusually close 2026 source is a review article with only a small worked example. The search cannot establish world novelty, patentability, freedom to operate, market size, or realized impact."},"researchability_gates":{"externally_supported_problem":{"status":"PASS","rationale":"Primary studies and official platform guidance support propagated performance anomalies, oscillation or flapping risk, transient loss during scaling, and the inadequacy of purely local decisions in some interconnected services.","source_ids":["SRC1","SRC2","SRC3","SRC4"]},"identifiable_adopter_or_authorizer":{"status":"PASS","rationale":"Platform/SRE teams and service owners control HPA policy and cross-service mitigations; production use of DAGOR demonstrates an identifiable organizational adopter for coordinated microservice control.","source_ids":["SRC2","SRC4"]},"distinct_testable_incremental_claim":{"status":"PASS","rationale":"Although the diagnostic eigenmode concept has close prior art, the comparative causal claim that guarded mode-targeted coordinated intervention outperforms dependency-aware detection and runbooks remains distinct and measurable.","source_ids":["SRC3","SRC4","SRC6","SRC8"]},"bounded_next_evidence_step":{"status":"PASS","rationale":"A shadow estimator followed by one low-traffic, reversible dependency-neighborhood canary can be preregistered with explicit model-validity, residual, drift, service-level, and cost thresholds.","source_ids":["SRC2","SRC3","SRC6"]},"no_unresolved_safety_or_authority_stop":{"status":"PASS","rationale":"The proposed first step is observational and then canary-scoped, uses existing reversible scaling or retry/load-shed controls, requires platform and service-owner authorization, and excludes production-wide unsupervised actuation and durability or security changes.","source_ids":["SRC2","SRC4","SRC6"]},"adequate_search_evidence":{"status":"PASS","rationale":"All six required lanes were searched adversarially using direct and synonymous terminology, including historical MIMO/eigenspace work, official product behavior, regional-language queries, and combinations of identification, spectral analysis, dependency graphs, and coordinated control. Eight direct sources from independent publishers were retained, seven of them primary, official, or first-party research/guidance.","source_ids":["SRC1","SRC2","SRC3","SRC4","SRC5","SRC6","SRC7","SRC8"]}},"strict_success":true,"screen_survival":true,"remaining_research_value":"MODERATE","recommended_next_step":"Run a shadow-only DMDc-style estimate on one stable dependency neighborhood using synchronized queues, latency, retries, throughput, replica counts, exogenous load, and recorded control actions. Preregister minimum spectral separation, mode repeatability, reconstruction-error, and drift criteria. If those pass, replay matched incidents or execute a low-traffic canary comparing a dependency-aware runbook with a reversible mode-targeted coordinated adjustment. Measure oscillation energy, time to recovery, error rate, tail latency, saturation, and cost; automatically restore prior configurations at any protected-outcome, residual, or drift threshold.","world_novelty_boundary":"This result identifies close public prior art and a remaining falsifiable incremental claim only. The bounded search does not establish world novelty, patentability, freedom to operate, market size, routine deployability, or realized impact."}