The widespread adoption of large language models (LLMs), such as the Generative Pre-trained Transformer (GPT), on cloud computing platforms (e.g., Azure) has resulted in a substantial increase in resource demand. This increase poses significant challenges for resource management within cloud environments. This paper aims to highlight these challenges by initially delineating the unique characteristics of resource management for GPT-based models. Subsequently, we analyze the specific challenges faced by resource management when applied to GPT-based models deployed on cloud platforms, and we propose algorithms for resource profiling and prediction concerning inference requests. To facilitate effective resource management, we present a comprehensive resource management framework that includes resource profiling and forecasting methodologies specifically designed for GPT-based models. Additionally, we discuss the future directions for resource management in the context of GPT-based models, emphasizing potential areas for further exploration and improvement. Through this analysis, we aim to provide valuable insights into resource management for GPT-based models deployed in cloud environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Resource Management for GPT-Based Model Deployed on Clouds: Challenges, Solutions, and Future Directions

  • Yongkang Dang,
  • Yiyuan He,
  • Minxian Xu,
  • Kejiang Ye

摘要

The widespread adoption of large language models (LLMs), such as the Generative Pre-trained Transformer (GPT), on cloud computing platforms (e.g., Azure) has resulted in a substantial increase in resource demand. This increase poses significant challenges for resource management within cloud environments. This paper aims to highlight these challenges by initially delineating the unique characteristics of resource management for GPT-based models. Subsequently, we analyze the specific challenges faced by resource management when applied to GPT-based models deployed on cloud platforms, and we propose algorithms for resource profiling and prediction concerning inference requests. To facilitate effective resource management, we present a comprehensive resource management framework that includes resource profiling and forecasting methodologies specifically designed for GPT-based models. Additionally, we discuss the future directions for resource management in the context of GPT-based models, emphasizing potential areas for further exploration and improvement. Through this analysis, we aim to provide valuable insights into resource management for GPT-based models deployed in cloud environments.