<p>The huge and growing demand for AI-native services requires more and more computing resources, especially GPU resources. Therefore, it is important to allocate resources to AI-native services efficiently to improve resource utilization and reduce infrastructure costs while still ensuring the Service Level Objectives (SLOs). To solve these issues, we proposes DRAFAS: an Dynamic Resource Allocation for AI-native Services framework with support for multi-tenants. To ensure isolation between services, we containerized services and managed them with a container orchestration framework. The GPU resources are spatially partitioned and allocated to each container. We then dynamically scaled the number of containers for each service based on the current load. We developed and evaluated two resource allocation algorithms for DRAFAS: a parameter-optimized rule-based algorithm and a Deep Reinforcement Learning (DRL)-based algorithm. The evaluation results showed that the DRL-based algorithm performed better in most cases when compared to the optimized rule-based algorithm and also generalized better when applied to different environment settings. In DRAFAS, we utilized network slicing to reserve network resources for the AI-native services. The evaluation results showed that using network slicing helped maintain SLOs when there was competing traffic, especially for the AI services that have high bandwidth requirements.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DRAFAS: Dynamic Resource Allocation for AI-Native Services

  • Nguyen Van Tu,
  • Sukhyun Nam,
  • Lizhuang Tan,
  • James Won-ki Hong

摘要

The huge and growing demand for AI-native services requires more and more computing resources, especially GPU resources. Therefore, it is important to allocate resources to AI-native services efficiently to improve resource utilization and reduce infrastructure costs while still ensuring the Service Level Objectives (SLOs). To solve these issues, we proposes DRAFAS: an Dynamic Resource Allocation for AI-native Services framework with support for multi-tenants. To ensure isolation between services, we containerized services and managed them with a container orchestration framework. The GPU resources are spatially partitioned and allocated to each container. We then dynamically scaled the number of containers for each service based on the current load. We developed and evaluated two resource allocation algorithms for DRAFAS: a parameter-optimized rule-based algorithm and a Deep Reinforcement Learning (DRL)-based algorithm. The evaluation results showed that the DRL-based algorithm performed better in most cases when compared to the optimized rule-based algorithm and also generalized better when applied to different environment settings. In DRAFAS, we utilized network slicing to reserve network resources for the AI-native services. The evaluation results showed that using network slicing helped maintain SLOs when there was competing traffic, especially for the AI services that have high bandwidth requirements.