Cloud computing has become a cornerstone of modern IT infrastructure, providing scalable, on-demand services that power a wide range of applications. However, ensuring high availability remains a significant challenge due to factors such as hardware failures, software bugs, and security threats. Traditional fault management approaches often require manual intervention, leading to increased downtime, inefficiencies, and operational costs. To address these challenges, AI-driven self-healing cloud infrastructure offers a proactive solution for fault detection, diagnosis, and autonomous recovery. This paper explores machine learning (ML)-based anomaly detection, predictive failure analytics, and automated remediation in cloud environments. AI techniques such as deep learning, reinforcement learning, and graph-based analysis enable real-time anomaly detection and intelligent fault recovery. Additionally, cloud-native orchestration tools like Kubernetes and Terraform play a crucial role in automating fault resolution and infrastructure management. Through case studies and empirical analysis, we evaluate the effectiveness of AI-driven self-healing mechanisms in reducing Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR), thereby enhancing overall system uptime. Finally, we discuss key challenges and future research directions, including the development of adaptive AI models, security-aware self-healing, and multi-cloud resilience strategies to achieve fully autonomous and reliable cloud computing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-Healing Cloud Infrastructure: Leveraging AI for Fault Detection and Recovery

  • Vaibhav Sawalkar,
  • Nitin More,
  • Sachin Jagadale,
  • Bhagyashree Shendkar,
  • Pankaj Chandre,
  • Chhaya Mhaske

摘要

Cloud computing has become a cornerstone of modern IT infrastructure, providing scalable, on-demand services that power a wide range of applications. However, ensuring high availability remains a significant challenge due to factors such as hardware failures, software bugs, and security threats. Traditional fault management approaches often require manual intervention, leading to increased downtime, inefficiencies, and operational costs. To address these challenges, AI-driven self-healing cloud infrastructure offers a proactive solution for fault detection, diagnosis, and autonomous recovery. This paper explores machine learning (ML)-based anomaly detection, predictive failure analytics, and automated remediation in cloud environments. AI techniques such as deep learning, reinforcement learning, and graph-based analysis enable real-time anomaly detection and intelligent fault recovery. Additionally, cloud-native orchestration tools like Kubernetes and Terraform play a crucial role in automating fault resolution and infrastructure management. Through case studies and empirical analysis, we evaluate the effectiveness of AI-driven self-healing mechanisms in reducing Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR), thereby enhancing overall system uptime. Finally, we discuss key challenges and future research directions, including the development of adaptive AI models, security-aware self-healing, and multi-cloud resilience strategies to achieve fully autonomous and reliable cloud computing.