Selecting the Performance Metrics to Control the CPU Oversubscription Ratio in a Cloud Server
摘要
Virtual machines (VMs) in IaaS public clouds often request more resources than will actually be used, in particular the CPU resource expressed in the number of cores. To improve CPU utilization, the cloud server performs CPU oversubscription, i.e. offers more cores to VMs than it actually has. To avoid degradation of user applications performance (e.g. significant increase of application latency) due to oversubscription, it is necessary to monitor server performance metrics that correlate with application latency. In this paper, we address the problem of selecting such metrics. Selection is performed on the base of experiments with synthetic workload on a developed testbed, which includes VMs with preconfigured applications, VMs with request generators, and servers to run these two sets of VMs separately. Over 120 metrics are collected on the server running VMs with applications, and in every VM with application. Several series of experiments were performed with homogeneous and heterogeneous workloads. For every series of experiments, a metric best correlating with application latency was selected. Finally, a candidate for “universal” metric was proposed, based on the host runqueue length. The experimentally selected metrics can be used in a group of methods that control CPU oversubscription ratio based on monitoring of performance metrics.