This chapter presents a Model-driven Deep Deterministic Policy Gradient (MDDPG) algorithm that enables the robot to learn a high-level assembly policy in realistic scenarios. The training process of the proposed MDDPG method is driven by a traditional force control strategy, which can improve the sample efficiency and avoid risky actions. A fuzzy reward system is designed, which utilizes prior knowledge to evaluate the assembly process. The reward system not only improves the learning efficiency but also prevents the agent from becoming trapped in a local minimum. A feedback exploration strategy incorporating expert knowledge is also presented. The effectiveness of the algorithm is validated by experiments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LFE: Model-Free RL Method for PiH Assembly

  • Jing Xu,
  • Hao Su,
  • Rui Chen,
  • Zhimin Hou

摘要

This chapter presents a Model-driven Deep Deterministic Policy Gradient (MDDPG) algorithm that enables the robot to learn a high-level assembly policy in realistic scenarios. The training process of the proposed MDDPG method is driven by a traditional force control strategy, which can improve the sample efficiency and avoid risky actions. A fuzzy reward system is designed, which utilizes prior knowledge to evaluate the assembly process. The reward system not only improves the learning efficiency but also prevents the agent from becoming trapped in a local minimum. A feedback exploration strategy incorporating expert knowledge is also presented. The effectiveness of the algorithm is validated by experiments.