As artificial intelligence (AI) continues to drive industrial upgrading, the demands for efficient and real-time AI services have escalated. However, meeting these demands poses challenges, particularly within the framework of 5G networks, where the infrastructure for providing AI services lacks real-time collaborative control over the required resources, such as the limited communication, computing and model resources. Anticipating the future of 6G networks, this paper advocates for the design with native AI integration from inception, leveraging wireless network infrastructure for ubiquitous real-time AI services. In this case, we address the issue of limited computing, model resources by proposing a distributed optimal resource utilization model migration algorithm. Additionally, a joint communication and computing resources allocation algorithm, based on a greedy strategy, is introduced to enhance resource utilization and support a higher number of AI inference tasks. Simulation results demonstrate that the performance of the proposed algorithm can approach the exhaustive search scheme at lower computational complexity.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Model Migration and Joint Communication and Computing Resource Allocation in Native AI Wireless Networks

  • Zhimin He,
  • Juan Deng,
  • Yaru Li,
  • Xiaoyun Wang

摘要

As artificial intelligence (AI) continues to drive industrial upgrading, the demands for efficient and real-time AI services have escalated. However, meeting these demands poses challenges, particularly within the framework of 5G networks, where the infrastructure for providing AI services lacks real-time collaborative control over the required resources, such as the limited communication, computing and model resources. Anticipating the future of 6G networks, this paper advocates for the design with native AI integration from inception, leveraging wireless network infrastructure for ubiquitous real-time AI services. In this case, we address the issue of limited computing, model resources by proposing a distributed optimal resource utilization model migration algorithm. Additionally, a joint communication and computing resources allocation algorithm, based on a greedy strategy, is introduced to enhance resource utilization and support a higher number of AI inference tasks. Simulation results demonstrate that the performance of the proposed algorithm can approach the exhaustive search scheme at lower computational complexity.