Motion planning is an indispensable part of unmanned aerial vehicle (UAV) technology, especially in unknown environments. Many studies have employed deep reinforcement learning to address the challenges of motion planning in unknown environments, making significant progress. However, the generalizability of these algorithms in new environments still requires improvement. Furthermore, current research lacks the capability to effectively adapt to varying scale of different environments. Therefore, this paper focuses on enhancing the generalizability of algorithms in new environments and their adaptability to different environments scales. To address this issue, this paper first establishes a scale-independent state framework and then employs data augmentation and a pair-actor exploration mechanism to improve the algorithm's generalizability. The test results demonstrate that the proposed algorithm outperforms both the PPO and TD3 algorithm, which are currently considered as the most popular reinforcement learning algorithm in this field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Motion Planning of UAV in Large-Scale Unknown Environments Using a Deep Reinforcement Learning Approach

  • Shiqi Huang,
  • Zhou Fang

摘要

Motion planning is an indispensable part of unmanned aerial vehicle (UAV) technology, especially in unknown environments. Many studies have employed deep reinforcement learning to address the challenges of motion planning in unknown environments, making significant progress. However, the generalizability of these algorithms in new environments still requires improvement. Furthermore, current research lacks the capability to effectively adapt to varying scale of different environments. Therefore, this paper focuses on enhancing the generalizability of algorithms in new environments and their adaptability to different environments scales. To address this issue, this paper first establishes a scale-independent state framework and then employs data augmentation and a pair-actor exploration mechanism to improve the algorithm's generalizability. The test results demonstrate that the proposed algorithm outperforms both the PPO and TD3 algorithm, which are currently considered as the most popular reinforcement learning algorithm in this field.