LFE: Model-Free RL Method for PiH Assembly
摘要
This chapter presents a Model-driven Deep Deterministic Policy Gradient (MDDPG) algorithm that enables the robot to learn a high-level assembly policy in realistic scenarios. The training process of the proposed MDDPG method is driven by a traditional force control strategy, which can improve the sample efficiency and avoid risky actions. A fuzzy reward system is designed, which utilizes prior knowledge to evaluate the assembly process. The reward system not only improves the learning efficiency but also prevents the agent from becoming trapped in a local minimum. A feedback exploration strategy incorporating expert knowledge is also presented. The effectiveness of the algorithm is validated by experiments.