In this paper, a heuristic function-improved deep reinforcement learning TD3-based robot navigation method is proposed for local navigation in a simulated substation environment. First, a 3D map of the nonstatic simulated substation environment is constructed. Second, 2D OU noise is added to the output action space to improve exploration efficiency. Then, a heuristic function based on state information is designed to adaptively adjust the hyperparameters of the OU noise action space. Finally, migration learning is used to train the robot for obstacle avoidance navigation in a simulated substation environment with dynamic obstacles added. The experimental results obtained by comparing the H-OU, H-Gauss and original Gaussian noise methods reveal that our proposed method achieves the highest average reward function among the three methods, with a 16.46% increase in the average reward value compared with that of the original method and a navigation success rate of 92%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Deep Reinforcement Learning Method with Adaptive Heuristic Function Improvement for Mobile Robot Navigation in Substation Environment

  • Fuzhi Jian,
  • En Li,
  • Lirong Xie,
  • Zhuang Yan,
  • Yifan Bian,
  • Guzhalayi Bali

摘要

In this paper, a heuristic function-improved deep reinforcement learning TD3-based robot navigation method is proposed for local navigation in a simulated substation environment. First, a 3D map of the nonstatic simulated substation environment is constructed. Second, 2D OU noise is added to the output action space to improve exploration efficiency. Then, a heuristic function based on state information is designed to adaptively adjust the hyperparameters of the OU noise action space. Finally, migration learning is used to train the robot for obstacle avoidance navigation in a simulated substation environment with dynamic obstacles added. The experimental results obtained by comparing the H-OU, H-Gauss and original Gaussian noise methods reveal that our proposed method achieves the highest average reward function among the three methods, with a 16.46% increase in the average reward value compared with that of the original method and a navigation success rate of 92%.