To achieve the great safety requirements for space collaborative robots and astronauts working in remote, unknown and complex space environment, the robots should be clever enough to avoid moving human or obstacles in a quite short time to keep both safe. In this paper, Deep Reinforcement Learning (DRL) techniques were applied to robot manipulators with a grasping workspace invaded by unpredictable moving obstacles. By introducing a hierarchical curriculum learning framework, a multidimensional agent was trained in a high-precision virtual environment. The dynamic obstacle avoidance success rate of strategy trained by Soft Actor-Critic (SAC) achieves 94.6%, higher than 68.4% by Deep Deterministic Policy Gradient (DDPG). Then the Actor network’s weight and bias matrices were copied and successfully implanted in a real RM63-6F anthropomorphic robot manipulator. The robot manipulator’s end clamp can reach about 0.02 m near the target site, which verifies that using DRL to obtain the motion planning strategy of space robot manipulator with limited computing resources is feasible.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on Intelligent Dynamic Obstacle Avoidance Strategy for Space Collaborative Robots

  • Yaobing Wang,
  • Fanglin Xie,
  • Yahang Zhang,
  • Lingxin Wang

摘要

To achieve the great safety requirements for space collaborative robots and astronauts working in remote, unknown and complex space environment, the robots should be clever enough to avoid moving human or obstacles in a quite short time to keep both safe. In this paper, Deep Reinforcement Learning (DRL) techniques were applied to robot manipulators with a grasping workspace invaded by unpredictable moving obstacles. By introducing a hierarchical curriculum learning framework, a multidimensional agent was trained in a high-precision virtual environment. The dynamic obstacle avoidance success rate of strategy trained by Soft Actor-Critic (SAC) achieves 94.6%, higher than 68.4% by Deep Deterministic Policy Gradient (DDPG). Then the Actor network’s weight and bias matrices were copied and successfully implanted in a real RM63-6F anthropomorphic robot manipulator. The robot manipulator’s end clamp can reach about 0.02 m near the target site, which verifies that using DRL to obtain the motion planning strategy of space robot manipulator with limited computing resources is feasible.