<p>As Internet of Things (IoT) technologies advance and computational demands grow, Unmanned Aerial Vehicle (UAV)-assisted mobile edge computing (MEC) is gradually growing into a feasible technology. However, limited communication and computing resources pose challenges to data processing for compute-intensive tasks in emergency or temporary scenarios. For this purpose, this paper explores a UAV-assisted MEC system model for data processing, aiming to minimize system latency. To address this problem, we propose a joint optimization framework for the NOMA-MEC system, which jointly optimizes the UAV’s trajectory, power and computational resource allocation. The joint optimization of UAV trajectory and resources allocation entails complex spatio-temporal dependencies, rendering it intractable for conventional optimization methods. Motivated by the use of deep reinforcement learning (DRL) in addressing high-dimensional decision-making problems under dynamic environments, we propose a solution based on the Proximal Policy Optimization (PPO) algorithm to address this problem. Finally, our method significantly outperforms baseline approaches in reducing system latency, as demonstrated by simulations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Joint Trajectory and Resources Allocation in a UAV-Assisted NOMA-MEC System Using Deep Reinforcement Learning

  • Junteng Liu,
  • Jie Ding,
  • Jinlong Shi,
  • Xu Zhao

摘要

As Internet of Things (IoT) technologies advance and computational demands grow, Unmanned Aerial Vehicle (UAV)-assisted mobile edge computing (MEC) is gradually growing into a feasible technology. However, limited communication and computing resources pose challenges to data processing for compute-intensive tasks in emergency or temporary scenarios. For this purpose, this paper explores a UAV-assisted MEC system model for data processing, aiming to minimize system latency. To address this problem, we propose a joint optimization framework for the NOMA-MEC system, which jointly optimizes the UAV’s trajectory, power and computational resource allocation. The joint optimization of UAV trajectory and resources allocation entails complex spatio-temporal dependencies, rendering it intractable for conventional optimization methods. Motivated by the use of deep reinforcement learning (DRL) in addressing high-dimensional decision-making problems under dynamic environments, we propose a solution based on the Proximal Policy Optimization (PPO) algorithm to address this problem. Finally, our method significantly outperforms baseline approaches in reducing system latency, as demonstrated by simulations.