AI computing resource pooling leverages software-defined technology to virtualize and containerize AI resources, separating hardware from software. It allows centralized management and on-demand allocation at the cluster level. Users can harness cluster computing power without worrying about the specifics of physical devices, bypassing the complexity of device selection and adaptation. This enhances overall resource utilization. AI services employ CPU and heterogeneous accelerator architecture. By decoupling CPU and accelerator computing, it broadens computational reach and facilitates cross-node execution. This is achieved through application layer API interception and forwarding, directing instructions and data from local accelerators to other nodes for processing, then returning results to the local task, enabling node-collaborative computation. This flexibility enables CPUs and GPUs in AI containers to be dynamically scheduled across the cluster, minimizing resource fragmentation. Research aims to understand AI computing and rendering load characteristics and to develop a novel mechanism for software-defined computing power. It aims to decouple diverse computing capabilities from business applications, enabling granular segmentation and intelligent scheduling. The objective is to create a unified software mechanism for pooling AI heterogeneous computing power, allowing for on-demand allocation and scheduling. This simplifies computing power invocation, optimizes resource utilization in large clusters, and standardizes the provision of computing resources, independent of device types. It substantially boosts the utilization rate and management flexibility of AI computing resource pools, promising extensive market applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Fine-Grained Management and Scheduling System for Artificial Intelligence Computing Resources

  • Chao Wang,
  • Xi Chen,
  • Qingshan Chen,
  • Shaohua Wu

摘要

AI computing resource pooling leverages software-defined technology to virtualize and containerize AI resources, separating hardware from software. It allows centralized management and on-demand allocation at the cluster level. Users can harness cluster computing power without worrying about the specifics of physical devices, bypassing the complexity of device selection and adaptation. This enhances overall resource utilization. AI services employ CPU and heterogeneous accelerator architecture. By decoupling CPU and accelerator computing, it broadens computational reach and facilitates cross-node execution. This is achieved through application layer API interception and forwarding, directing instructions and data from local accelerators to other nodes for processing, then returning results to the local task, enabling node-collaborative computation. This flexibility enables CPUs and GPUs in AI containers to be dynamically scheduled across the cluster, minimizing resource fragmentation. Research aims to understand AI computing and rendering load characteristics and to develop a novel mechanism for software-defined computing power. It aims to decouple diverse computing capabilities from business applications, enabling granular segmentation and intelligent scheduling. The objective is to create a unified software mechanism for pooling AI heterogeneous computing power, allowing for on-demand allocation and scheduling. This simplifies computing power invocation, optimizes resource utilization in large clusters, and standardizes the provision of computing resources, independent of device types. It substantially boosts the utilization rate and management flexibility of AI computing resource pools, promising extensive market applications.