<p>There is a growing expectation that deep reinforcement learning will enable multi-degree-of-freedom robots to acquire policies suitable for real-world applications. However, a robot system with a variety of components requires many learning trials for each different combination of robot modules. In this study, we propose a hierarchical policy design to segment tasks according to different robot components. The tasks of the multi-module robot are performed by skill sets trained on a component-by-component basis. In our learning approach, each module learns reusable skills, which are then integrated to control the whole robotic system. By adopting component-based learning and reusing previously acquired policies, we transform the action space from continuous to discrete. This transformation reduces the complexity of exploration across the entire robotic system. We validated our proposed method by applying it to a valve rotation task using a combination of a robotic arm and a robotic gripper. Evaluation based on physical simulations showed that hierarchical policy construction improved sample efficiency, achieving performance comparable to the baseline with 46.3% fewer samples.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An empirical evaluation of a hierarchical reinforcement learning method towards modular robot control

  • Sho Takeda,
  • Satoshi Yamamori,
  • Satoshi Yagi,
  • Jun Morimoto

摘要

There is a growing expectation that deep reinforcement learning will enable multi-degree-of-freedom robots to acquire policies suitable for real-world applications. However, a robot system with a variety of components requires many learning trials for each different combination of robot modules. In this study, we propose a hierarchical policy design to segment tasks according to different robot components. The tasks of the multi-module robot are performed by skill sets trained on a component-by-component basis. In our learning approach, each module learns reusable skills, which are then integrated to control the whole robotic system. By adopting component-based learning and reusing previously acquired policies, we transform the action space from continuous to discrete. This transformation reduces the complexity of exploration across the entire robotic system. We validated our proposed method by applying it to a valve rotation task using a combination of a robotic arm and a robotic gripper. Evaluation based on physical simulations showed that hierarchical policy construction improved sample efficiency, achieving performance comparable to the baseline with 46.3% fewer samples.