E-GAIL: efficient GAIL through including negative corruption and long-term rewards for robotic manipulations
摘要
Learning an effective manipulation policy with high efficiency in robotics continues to be a significant challenge. In this paper, we propose E-GAIL, which aims to learn manipulation policies efficiently from a limited set of demonstrations with negative corruption and long-term rewards under the framework of GAIL. Specifically, we propose two techniques: 1) Utilizing both short-term and long-term observations to offer additional rewards for training, accelerating convergence. 2) Incorporating negative actions into generated trajectories for corruption to improve data effectiveness and increase success rates. E-GAIL achieves a 25% improvement in success rates across multiple manipulation tasks, requiring 70% fewer episodes for policy convergence, highlighting its efficiency with limited demonstrations. Our video is available at https://youtu.be/bIDfOjYcY54.