With the widespread application of pre-trained large models in natural language processing and multimodal inference, their high computational power demands, dynamic load characteristics, and strict low-latency requirements have posed serious challenges to computing infrastructure. Traditional physical GPU resource allocation models are difficult to meet demands due to low utilization and poor multi-tenant isolation. In contrast, vGPU technology can improve efficiency through fine-grained resource partitioning and dynamic scheduling, but its performance evaluation methods still face critical bottlenecks: existing testing frameworks are primarily designed for physical GPUs, lack support for large model inference scenarios, and are inadequate in dynamic load simulation, multidimensional performance monitoring, and performance loss analysis of the virtualization layer. In this paper, we proposed a vGPU performance testing framework for large model inference, integrating dynamic computation resource provisioning, automated load generation, and multidimensional fine-grained monitoring through modular design. This framework enables second-level synchronization and collection of monitoring data for both physical and virtual GPU resources, as well as inference workloads. It built a parallel inference performance dataset covering various heterogeneous large models under multiple resource partitioning schemes for the first time. Through quantitative analysis of experiment, it is found that the parallel inference mode for multiple large models results in an average vGPU performance loss of more than 100% compared to the exclusive inference mode. The results of this paper provide theoretical support and practical tools for vGPU resource scheduling and memory optimization for large model inference services in cloud environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

vGPU Performance Testing Framework for Large Model Inference

  • Hu Zhang,
  • Youli Zhang,
  • Hongjun Dai,
  • Ming Gao,
  • Huifeng Liu

摘要

With the widespread application of pre-trained large models in natural language processing and multimodal inference, their high computational power demands, dynamic load characteristics, and strict low-latency requirements have posed serious challenges to computing infrastructure. Traditional physical GPU resource allocation models are difficult to meet demands due to low utilization and poor multi-tenant isolation. In contrast, vGPU technology can improve efficiency through fine-grained resource partitioning and dynamic scheduling, but its performance evaluation methods still face critical bottlenecks: existing testing frameworks are primarily designed for physical GPUs, lack support for large model inference scenarios, and are inadequate in dynamic load simulation, multidimensional performance monitoring, and performance loss analysis of the virtualization layer. In this paper, we proposed a vGPU performance testing framework for large model inference, integrating dynamic computation resource provisioning, automated load generation, and multidimensional fine-grained monitoring through modular design. This framework enables second-level synchronization and collection of monitoring data for both physical and virtual GPU resources, as well as inference workloads. It built a parallel inference performance dataset covering various heterogeneous large models under multiple resource partitioning schemes for the first time. Through quantitative analysis of experiment, it is found that the parallel inference mode for multiple large models results in an average vGPU performance loss of more than 100% compared to the exclusive inference mode. The results of this paper provide theoretical support and practical tools for vGPU resource scheduling and memory optimization for large model inference services in cloud environments.