Augmenting Cloud Resource Management with the Necessary Amount of Machine Intelligence
摘要
Cloud computing and data center environments often suffer from low resource efficiency due to overprovisioning and suboptimal management decisions. Improving resource efficiency requires accurate forecasting of metrics like resource consumption and user utilization patterns. However, forecasting resource usage is challenging due to the diverse and dynamic patterns across different levels and types of hardware resources. This work aims to contribute a comprehensive approach to designing an end-to-end cloud resource management system aimed at enhancing resource, energy, and cost efficiency. The research explores a range of machine learning (ML) and non-ML-based methods for predicting future resource usage and evaluates them based on model accuracy and practicality, allowing for minimal learning overheads and high generalizability. A key contribution is the development of a novel cloud resource forecasting model that combines ML and non-ML strategies to achieve a balance between resource efficiency and operational overheads. This model will be integrated into an end-to-end cloud resource management system, leveraging techniques such as overcommitment and autoscaling to optimize resource utilization and application performance. Overall, this work aims to advance the state-of-the-art in cloud resource management and contribute to more efficient and sustainable cloud computing infrastructures, leveraging only the necessary amount of machine intelligence.