Data Aggregation (DAgger) Algorithm Using Adversarial Agent Policy for Dynamic Situations
摘要
To handle dynamic and static situations in robots using deep learning, the behavior of dynamic obstacles is manually modeled. This is limited in generating diverse situations and does not account for improvements in the ego agent’s policy during training. To address these limitations, a method has been proposed that involves training by modeling the dynamic obstacle’s behavior as the adversarial agent policy. However, this method has been applied only to reinforcement learning, not to the data aggregation (DAgger) algorithm that repeats imitation learning. Thus, this paper proposes a novel DAgger training framework that adopts an adversarial agent policy having a competitive relationship with the ego agent policy. The adversarial agent policy is trained to aggregate data of various and non-redundant dynamic situations, taking into account the gradually improving performance of the ego agent’s policy. The proposed method was applied to autonomous driving and obtained a safer ego agent policy than a policy trained without the adversarial agent policy.