{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp11_mechanism_context_external20_20260804","research_id":"eoa_inverse_innovation_exp11_external_scrutiny_20260804","cell_id":"invariant_mode_decomposition_design__computer_science","opaque_id":"invariant_mode_decomposition_design__computer_science__C","search_lanes":{"direct_problem":{"queries":["microservice autoscaling oscillation coupled services retry queues instability","microservice autoscaling cascading failure control loop oscillation","microservice eigenvalue modal analysis autoscaling state transition operator"],"source_ids":["SRC1","SRC2","SRC4","SRC6"],"no_result_note":null},"closest_prior_art":{"queries":["dependency-aware coordinated autoscaling microservices control theory","microservices coordinated autoscaling queueing network control multiple services paper","microservice autoscaling system identification state space model eigenvalue stability","modal analysis microservices autoscaling"],"source_ids":["SRC1","SRC2","SRC3","SRC4"],"no_result_note":null},"historical_terminology":{"queries":["microservice autoscaling MIMO control eigenvalues stability queueing network older terminology","dynamic mode decomposition cloud autoscaling distributed system","dynamic mode decomposition with control state snapshots actuation"],"source_ids":["SRC2","SRC7"],"no_result_note":null},"products_practices_standards":{"queries":["site:kubernetes.io horizontal pod autoscaler stabilization window flapping scaling behavior","site:aws.amazon.com autoscaling oscillation cooldown target tracking microservices","Alibaba adaptive horizontal pod autoscaling observer dry run"],"source_ids":["SRC5","SRC8"],"no_result_note":null},"non_english_regional":{"queries":["Kubernetes 微服务 自动扩缩容 振荡 级联故障 耦合 稳定性 特征值","Kubernetes microservicios autoescalado oscilación fallos en cascada estabilidad","Kubernetes Microservices Autoskalierung Oszillation gekoppelte Regelkreise Stabilität Eigenwerte","Kubernetes マイクロサービス オートスケーリング 振動 連鎖障害 安定性 固有値"],"source_ids":["SRC5","SRC8"],"no_result_note":"Chinese, Spanish, German, and Japanese terminology searches recovered regional autoscaling material and localized Kubernetes documentation, but no additional distinct non-English implementation of mode-directed coordinated control."},"composition_subproblems":{"queries":["microservice coupled control loops autoscaling stability eigenmode eigenvalue","microservices autoscaling spectral analysis dependency graph control","microservice telemetry Jacobian eigenvalues spectral drift canary mitigation","service dependency graph retries queues replica oscillation coordinated scaling"],"source_ids":["SRC1","SRC3","SRC4","SRC6","SRC7"],"no_result_note":null}},"sources":[{"source_id":"SRC1","title":"Kubernetes Microservices Instability: Theory, Worked Example, and Defensive Pipeline","url":"https://www.ajms.in/index.php/ajms/article/download/636/303","publisher":"Asian Journal of Mathematical Sciences / B.R. Nahata Smriti Sansthan","date_or_year":"2025 issue; published 2026","source_type":"SECONDARY_RESEARCH","language":"English","claims_supported":["Describes estimating a telemetry-derived Jacobian for Kubernetes microservices, computing eigenvalues, monitoring spectral drift, and using modal participation to prioritize services.","Prescribes reversible mitigations including throttling, circuit breakers, autoscaler damping, capacity changes, and staging or canary validation.","Provides only a three-service numerical worked example rather than a comparative operational control trial."]},{"source_id":"SRC2","title":"On the Stability of the Kubernetes Horizontal Autoscaler Control Loop","url":"https://doaj.org/article/5550fbb44f13438687e6958ecc076db4","publisher":"IEEE Access (record via DOAJ)","date_or_year":"2025","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Models Kubernetes HPA with control theory and analyzes autoscaling stability using simulation and a real testbed.","Extends analysis from individual services to a whole graph of interconnected services.","Confirms that independently controlled services and graph-level interactions are recognized research objects."]},{"source_id":"SRC3","title":"DeepScaler: Holistic Autoscaling for Microservices Based on Spatiotemporal GNN with Adaptive Graph Learning","url":"https://arxiv.org/abs/2309.00859","publisher":"arXiv / paper authors","date_or_year":"2023","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Treats time-varying service dependencies as causes of cascading provisioning effects.","Learns latent dependency affinities and spatiotemporal features, then simultaneously reconfigures interacting services.","Reports experimental SLA and cost improvements, but does not select or suppress eigenmodes of a local transition operator."]},{"source_id":"SRC4","title":"PBScaler: A Bottleneck-aware Autoscaling Framework for Microservice-based Applications","url":"https://arxiv.org/abs/2303.14620","publisher":"arXiv / paper authors","date_or_year":"2023","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Documents anomaly propagation across interacting microservices and replica and latency oscillations caused by repeated online scaling attempts.","Uses a dependency graph, topology-based bottleneck ranking, and offline optimization to avoid unnecessary scaling.","Implements and evaluates the controller on Kubernetes microservice applications, but targets bottleneck services rather than coupled dynamical modes."]},{"source_id":"SRC5","title":"Horizontal Pod Autoscaling","url":"https://kubernetes.io/docs/concepts/workloads/autoscaling/horizontal-pod-autoscale/","publisher":"Kubernetes","date_or_year":"2026 documentation","source_type":"OFFICIAL_GUIDANCE","language":"English","claims_supported":["Establishes the deployed baseline: an HPA periodically changes a target workload's replicas from observed resource or custom metrics.","Explicitly recognizes frequent replica fluctuation as thrashing or flapping.","Provides stabilization windows, tolerance, and rate policies as established safeguards, while each HPA remains bound to a single workload."]},{"source_id":"SRC6","title":"Addressing Cascading Failures","url":"https://sre.google/sre-book/addressing-cascading-failures/","publisher":"Google Site Reliability Engineering","date_or_year":"2016","source_type":"OFFICIAL_GUIDANCE","language":"English","claims_supported":["Documents retries amplifying overload across frontend and backend layers and contributing to cascading failures.","Identifies SREs and service owners as operational actors and recommends load shedding, retry budgets, backoff, capacity changes, and testing.","Shows that local retry and overload behavior can reinforce system-wide failure and that interventions can themselves affect users."]},{"source_id":"SRC7","title":"Dynamic Mode Decomposition with Control","url":"https://arxiv.org/abs/1409.6358","publisher":"arXiv / paper authors","date_or_year":"2014 preprint; 2016 journal publication","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Provides the older general system-identification method for estimating low-order dynamics and control effects from synchronized state and actuation snapshots.","Extracts unstable growth modes and spectral properties from complex systems.","Does not apply the method to microservice autoscaling, leaving domain-specific validity and intervention evidence unresolved."]},{"source_id":"SRC8","title":"AHPA: Adaptive Horizontal Pod Autoscaling Systems on Alibaba Cloud Container Service for Kubernetes","url":"https://arxiv.org/abs/2303.03640","publisher":"Alibaba Cloud researchers / arXiv","date_or_year":"2023","source_type":"PRIMARY_RESEARCH","language":"English","claims_supported":["Reports a production-deployed autoscaler across Alibaba Cloud customer scenarios, identifying a concrete platform adopter.","Uses workload decomposition, queue-based performance modeling, action-frequency limits, monitoring, and an Observer dry-run mode.","Demonstrates that shadow observation and bounded autoscaling tests are operationally implementable, but does not model cross-service eigenmodes."]}],"problem_evidence":{"status":"SUPPORTED","finding":"Independent per-workload autoscaling is established practice, frequent replica fluctuation is explicitly recognized as thrashing or flapping, and research and operational guidance document anomaly, overload, and retry propagation across service dependencies. PBScaler directly observes replica/latency oscillation, while Google SRE documents self-amplifying retry cascades.","source_ids":["SRC2","SRC4","SRC5","SRC6"],"uncertainty":"The sources establish oscillation and cascading interactions, but do not show that most incidents contain reproducible, separated invariant modes; discrete bugs, traffic shocks, topology changes, and delays remain competing explanations."},"adopter_evidence":{"status":"SUPPORTED","finding":"Platform/SRE teams and service owners are identifiable authorizers because Kubernetes exposes scaling controls to cluster operators, Google describes SRE intervention during cascading overload, and Alibaba reports production deployment of AHPA across multiple customer scenarios.","source_ids":["SRC5","SRC6","SRC8"],"uncertainty":"No evidence establishes willingness by a named organization to authorize this exact modal-control canary; joint approval and local change-management requirements remain deployment-specific."},"implementation_evidence":{"status":"PARTLY_SUPPORTED","finding":"Necessary components exist separately: synchronized metric and actuation snapshots can identify low-order controlled dynamics; Kubernetes exposes telemetry-driven replica control and stabilization safeguards; dependency-aware controllers have been implemented; and AHPA supports dry-run observation. The closest spectral microservice source supplies a defensive pipeline and canary recommendation, but only a toy numerical example rather than an operational comparative trial.","source_ids":["SRC1","SRC3","SRC4","SRC5","SRC7","SRC8"],"uncertainty":"Sparse/confounded telemetry, changing topology, delays, nonlinear saturation, intervention-induced model change, and lack of a demonstrated spectral gap could invalidate the local linear model or its control map."},"prior_art":{"disposition":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Kubernetes Microservices Instability defensive spectral pipeline","source_ids":["SRC1"],"same_problem":true,"same_causal_lever":true,"overlap":"Nearly the full conceptual mechanism overlaps: telemetry-derived Jacobian, eigenvalues, spectral drift, modal participation, mode-prioritized mitigation, autoscaler damping, and staging/canary validation.","remaining_difference":"It presents an illustrative three-service calculation and defensive recommendations, not a reproducible operational comparison showing that separated modes persist and mode-targeted coordinated control outperforms dependency-aware detection/runbooks on recovery and protected outcomes."},{"name":"Whole-service-graph control-theoretic HPA stability analysis","source_ids":["SRC2"],"same_problem":true,"same_causal_lever":false,"overlap":"Analyzes HPA stability for interconnected services using control theory, simulation, and a Kubernetes testbed.","remaining_difference":"It studies and predicts stability rather than estimating dominant empirical service-interaction modes and mapping those modes to coordinated corrective actuation."},{"name":"DeepScaler","source_ids":["SRC3"],"same_problem":true,"same_causal_lever":false,"overlap":"Learns latent, time-varying dependencies and simultaneously scales interacting services to avoid cascading provisioning effects.","remaining_difference":"Its lever is graph-neural demand estimation and holistic allocation, not spectral identification and suppression of weakly damped or growing transition modes."},{"name":"PBScaler","source_ids":["SRC4"],"same_problem":true,"same_causal_lever":false,"overlap":"Targets propagated anomalies and observed autoscaling oscillations with dependency-aware bottleneck localization and safer offline scaling optimization.","remaining_difference":"It ranks individual bottleneck services rather than coupled perturbation directions and does not use modal gain, residual, spectral-gap, or mode-drift criteria."},{"name":"Dynamic Mode Decomposition with Control","source_ids":["SRC7"],"same_problem":false,"same_causal_lever":true,"overlap":"Supplies the general data-driven method for recovering a state-transition operator, control map, unstable modes, and low-order dynamics from state and actuation snapshots.","remaining_difference":"It supplies no microservice state definition, dependency-neighborhood validity contract, safe autoscaling intervention map, or evidence on cloud incidents."}],"contrastive_claim_remaining":"Within a pre-registered, locally stable microservice dependency neighborhood, synchronized telemetry and known actuation can reveal a reproducible, spectrally separated risky interaction mode, and a reversible coordinated intervention chosen from that mode will reduce oscillation or recovery time versus dependency-aware anomaly detection plus coordinated runbooks without violating latency, availability, cost, residual-error, or drift limits.","contrastive_claim_falsifier":"The claim fails if no reproducible separated mode survives controls for workload, releases, topology, and known incidents, or if a mode is reproducible but mode-directed intervention does not improve recovery or protected outcomes relative to the nearest rival in the bounded canary.","confidence":"MODERATE","search_limitations":"This was a bounded public-web search across six lanes, not a systematic review. Exactly eight direct sources were retained and their pages or indexed full-text/PDF records were examined; the AJMS publisher PDF intermittently returned a cache miss, although its indexed direct PDF text exposed the method and mitigation sections. Paywalled literature, patents beyond surfaced results, proprietary postmortems, source code behavior, and unpublished industry practice may contain closer art."},"researchability_gates":{"externally_supported_problem":{"status":"PASS","rationale":"Official guidance and primary studies independently support autoscaling flapping, propagated performance anomalies, and retry-driven cascading instability.","source_ids":["SRC2","SRC4","SRC5","SRC6"]},"identifiable_adopter_or_authorizer":{"status":"PASS","rationale":"Platform/SRE teams and affected service owners control the relevant scaling, retry, and load-shedding mechanisms; Alibaba provides a concrete production platform precedent for adopting advanced autoscaling.","source_ids":["SRC5","SRC6","SRC8"]},"distinct_testable_incremental_claim":{"status":"PASS","rationale":"Although the spectral concept substantially overlaps SRC1, a distinct empirical claim remains: reproducible separated modes and superior protected-outcome performance of mode-directed coordinated control versus the specified dependency-aware rival.","source_ids":["SRC1","SRC2","SRC3","SRC4"]},"bounded_next_evidence_step":{"status":"PASS","rationale":"A bounded sequence is feasible: offline incident replay and shadow estimation, followed only after pre-registered identification criteria by a low-traffic reversible canary. SRC1 recommends staging/canary validation and SRC8 demonstrates an Observer dry-run mode.","source_ids":["SRC1","SRC5","SRC8"]},"no_unresolved_safety_or_authority_stop":{"status":"PASS","rationale":"No inherent safety or authority stop appears if production-wide autonomous control is excluded, owners approve the canary, prior configurations are retained, and error, latency, saturation, cost, residual, and drift limits trigger immediate rollback. Established guidance supports reversible load shedding, retry control, damping, and staged validation.","source_ids":["SRC1","SRC5","SRC6","SRC8"]},"adequate_search_evidence":{"status":"PASS","rationale":"The bounded search covered all six required lanes, including historical system-identification terminology, official products and practices, four non-English/regional formulations, and combinations of the component subproblems. Eight direct sources from more than two publishers were retained, including multiple primary and official sources.","source_ids":["SRC1","SRC2","SRC3","SRC4","SRC5","SRC6","SRC7","SRC8"]}},"strict_success":true,"screen_survival":true,"remaining_research_value":"MODERATE","recommended_next_step":"Pre-register a shadow study on synchronized queue, latency, retry, throughput, replica, workload, deployment, and control-event data from one dependency neighborhood. Compare modal prediction against per-service/exogenous baselines on held-out windows; require stable mode identity, gain separation, and residual bounds. Only if those criteria pass, run an owner-approved low-traffic canary comparing a reversible mode-directed scaling or retry-limit action with the dependency-aware runbook rival, with automatic rollback on protected-outcome limits.","world_novelty_boundary":"The search found a very close published conceptual spectral pipeline, so no broad novelty claim is warranted. The remaining boundary is the narrower comparative empirical claim about reproducible modes and safe mode-directed control in an operational microservice canary. This bounded review cannot establish world novelty, patentability, freedom to operate, market size, or realized impact."}