Live Kubernetes Cluster Anomaly Detection & Chaos Platform
AIOps dashboard ingesting container metrics to predict OOMKills and disk saturation before outages occur, with automated node cordoning.
Project Overview
Designed for DevOps and Site Reliability Engineering (SRE) teams. Scrapes Prometheus metrics across worker nodes and applies Holt-Winters time-series forecasting to anticipate node exhaustion, automatically alerting Slack and scheduling graceful container migration.
Designed for DevOps and Site Reliability Engineering (SRE) teams. Scrapes Prometheus metrics across worker nodes and applies Holt-Winters time-series forecasting to anticipate node exhaustion, automatically alerting Slack and scheduling graceful container migration.
AIOps dashboard ingesting container metrics to predict OOMKills and disk saturation before outages occur, with automated node cordoning.
Core Project Objectives
Capture continuous analog/digital sensor readings with robust noise filtering and hardware calibration.