In the scenario of multiple terminal devices and multiple edge servers, a computation offloading algorithm based on proximal policy optimization (PPO) is proposed to enhance the success rate of task offloading for terminal devices and improve the resource utilization of edge servers. In this paper, advantage normalization and reward normalization are incorporated into the PPO algorithm to enhance its stability and performance. Additionally, the Softmax function is introduced into the PPO algorithm to enable its application in decision problems with discrete action spaces. A comparison is made between three algorithms in discrete action spaces: Deep deterministic policy gradient discrete (DDPG-D), deep Q-network (DQN), and PPO-discrete (PPO-D). Experimental results demonstrate that the PPO-D algorithm can effectively reduce decision error rates and improve overall resource utilization in computation offloading scenarios. Furthermore, experimental results show that incorporating advantage normalization and reward normalization can effectively enhance the stability of the PPO-D algorithm (ARN-PPO-D), thereby further improving its performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MEC Computation Offloading Decision Based on ARN-PPO-D

  • Xinhao Mao,
  • Xinyu Zhang,
  • Xiaoyu Wang

摘要

In the scenario of multiple terminal devices and multiple edge servers, a computation offloading algorithm based on proximal policy optimization (PPO) is proposed to enhance the success rate of task offloading for terminal devices and improve the resource utilization of edge servers. In this paper, advantage normalization and reward normalization are incorporated into the PPO algorithm to enhance its stability and performance. Additionally, the Softmax function is introduced into the PPO algorithm to enable its application in decision problems with discrete action spaces. A comparison is made between three algorithms in discrete action spaces: Deep deterministic policy gradient discrete (DDPG-D), deep Q-network (DQN), and PPO-discrete (PPO-D). Experimental results demonstrate that the PPO-D algorithm can effectively reduce decision error rates and improve overall resource utilization in computation offloading scenarios. Furthermore, experimental results show that incorporating advantage normalization and reward normalization can effectively enhance the stability of the PPO-D algorithm (ARN-PPO-D), thereby further improving its performance.