Helicopter control is considered a highly challenging problem because of its complex nonlinear dynamic characteristics. Deep reinforcement learning (DRL), a machine learning approach that combines the nonlinear representation power of deep learning with the sequential decision-making capabilities of reinforcement learning, effectively addresses complex dynamic decision-making problems. This paper presents a deep reinforcement learning-based approach for autonomous helicopter flight strategy learning. The proposed method integrates the Soft Actor-Critic (SAC) algorithm with the curiosity-driven exploration mechanism of Random Network Distillation, which generates intrinsic reward signals to encourage the agent to explore novel environments, thereby enhancing exploration efficiency. Additionally, the model incorporates a self-attention mechanism, enabling the actor network to capture dependencies between state information. Experimental results demonstrate that the proposed method substantially enhances training efficiency and control performance in random target point hovering tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improved Soft Actor-Critic Algorithm for Autonomous Helicopter Target Hovering

  • Ting Cheng,
  • Xiangping Bryce Zhai,
  • Jing Zhu,
  • Chenkai Cao,
  • Qi Zhu

摘要

Helicopter control is considered a highly challenging problem because of its complex nonlinear dynamic characteristics. Deep reinforcement learning (DRL), a machine learning approach that combines the nonlinear representation power of deep learning with the sequential decision-making capabilities of reinforcement learning, effectively addresses complex dynamic decision-making problems. This paper presents a deep reinforcement learning-based approach for autonomous helicopter flight strategy learning. The proposed method integrates the Soft Actor-Critic (SAC) algorithm with the curiosity-driven exploration mechanism of Random Network Distillation, which generates intrinsic reward signals to encourage the agent to explore novel environments, thereby enhancing exploration efficiency. Additionally, the model incorporates a self-attention mechanism, enabling the actor network to capture dependencies between state information. Experimental results demonstrate that the proposed method substantially enhances training efficiency and control performance in random target point hovering tasks.