<p>Controlling a two-arm robot presents a multidimensional and intricate challenge, further compounded by the overlapping workspace of the arms, which significantly complicates trajectory planning. Therefore, in this paper, we propose a trajectory planning approach based on the neural network Soft Actor-Critic (SAC) algorithm tailored for a 7-degree-of-freedom dual-arm robot. To improve the efficiency of SAC-based methods for trajectory planning in dual-arm cooperation, we introduce a novel hybrid reward function. This function integrates the artificial potential field method with the posture reward function. Additionally, we incorporate the hindsight experience replay (HER) algorithm to optimize the utilization of neural network training data. Moreover, experiments were carried out using both the Baxter robot and the Coppeliasim simulation model to validate the effectiveness of our proposed method. Our results demonstrate that the introduced hybrid reward function significantly enhances the convergence rate of the SAC algorithm, thereby improving the overall performance of trajectory planning for dual-arm robots.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep reinforcement learning with hindsight experience replay for dual-arm robot trajectory planning

  • Shoutao Li,
  • Zhanggen Chu,
  • Zhidong Hu,
  • Zhenze Liu

摘要

Controlling a two-arm robot presents a multidimensional and intricate challenge, further compounded by the overlapping workspace of the arms, which significantly complicates trajectory planning. Therefore, in this paper, we propose a trajectory planning approach based on the neural network Soft Actor-Critic (SAC) algorithm tailored for a 7-degree-of-freedom dual-arm robot. To improve the efficiency of SAC-based methods for trajectory planning in dual-arm cooperation, we introduce a novel hybrid reward function. This function integrates the artificial potential field method with the posture reward function. Additionally, we incorporate the hindsight experience replay (HER) algorithm to optimize the utilization of neural network training data. Moreover, experiments were carried out using both the Baxter robot and the Coppeliasim simulation model to validate the effectiveness of our proposed method. Our results demonstrate that the introduced hybrid reward function significantly enhances the convergence rate of the SAC algorithm, thereby improving the overall performance of trajectory planning for dual-arm robots.