错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Modular Reinforcement Learning In Long-Horizon Manipulation Tasks

  • Michal Vavrecka,
  • Jonas Kriz,
  • Nikita Sokovnin,
  • Gabriela Sejnova

摘要

Recently, a number of reinforcement learning (RL) algorithms have been proposed in the area of robotic manipulation. As most of the current robotic benchmarks are focused on simple, non-diverse tasks such as the translation of objects within the scene, various single-policy algorithms are able to solve them with a high success rate. However, when a sequence of diverse subgoals is required (translation, rotation, 6DOF manipulation, trajectory following), the single-policy networks are shown to fail. In this work, we propose two modular multi-policy algorithms (MultiPPO2 and MultiACKTR) that improve diverse long-horizon tasks by adopting a separate policy for each skill that follows its own subgoal. We tested our algorithm in a virtual robotic simulator both on single and multi-step tasks requiring non-diverse (translation) skills and also diverse (translation, rotation and path following) skills. Both algorithms (MultiPPO2 and MultiACKTR) achieved similar performance as single-policy algorithms in the single-step and multi-step translational tasks and outperformed the single-policy algorithms in tasks that require object translation, rotation, and path following. The result proves that a multi-policy algorithm offers better scalability for complex long-horizon tasks and can thus be a promising tool in combination with hierarchical reinforcement learning.