Job Failure Prediction and Recommendation System to Prevent Job Failure in Cloud Computing Environment
摘要
Cloud computing has emerged as a critical technology, providing unparalleled scalability and flexibility for enterprises. However, ensuring the reliability and efficiency of cloud services remains a significant challenge for providers. This paper presents an in-depth analysis of task failure characteristics in cloud environments, utilizing publicly available datasets such as the Google Cluster Trace, Mustang Trace, and LANL Trinity Trace. Through comprehensive investigation, correlations between task termination statuses and various factors including CPU demand, memory demand, priority, number of nodes, tasks requested, class level, and scheduling level are uncovered. Additionally, a novel framework leveraging advanced machine learning techniques, including artificial neural networks (ANN) and convolutional neural networks (CNN), is proposed for predicting task failure. By extracting features from the datasets, the prediction model aims to identify tasks at risk of failure preemptively, thereby enhancing cloud service reliability and performance. Furthermore, a recommendation system is developed to optimize resource allocation and scheduling decisions, mitigating the impact of task failures on overall system throughput and efficiency. Integrating advanced machine learning methodologies with real-world cloud data, this study contributes significantly to the advancement of fault tolerance and quality of service in cloud computing environments.