Exploring cloud instance options for optimal performance-cost efficiency
摘要
The increasing reliance on data-intensive applications in cloud computing necessitates a cost-effective approach to utilizing cloud servers. However, the wide range of instance types and configurations can lead to suboptimal choices, resulting in unnecessary expenses. This highlights the need for scalable and economically efficient strategies that can enhance performance while minimizing costs. In this manuscript, we explore the question: How can users effectively select cloud instances–across CPU and GPU configurations from various providers–to achieve optimal performance-cost efficiency for parallel and AI workloads, without resorting to exhaustive benchmarking? To address this, we conducted an extensive evaluation of performance, cost, and their trade-offs by executing eighteen parallel workloads on fifty-three CPU-based instances and four AI inference models on twelve GPU-based instances from four major cloud providers. Our findings reveal that the most cost-effective instances do not necessarily offer the highest performance, and the cheapest options often fail to deliver ideal efficiency. For AI workloads, NVIDIA H200-based instances demonstrated the best performance, while the NVIDIA L40S-based instances provided the lowest cost, albeit at the expense of significantly lower throughput. Furthermore, while we explored AI models for instance selection, their recommendations did not consistently align with the optimal choices. We also found that by identifying the most suitable instance for each workload, users could achieve a 29.6% increase in performance and a 6.5% reduction in costs compared to using a fixed, best-performing instance.