Human-Guided Deep Reinforcement Learning for UAV Flocking Through Rule Extraction and Policy Transfer
摘要
Deep reinforcement learning (DRL) has proved effective for developing flocking behavior in a swarm of UAVs. However, training the UAV swarm is typically time-consuming. This paper proposes a human-guidable DRL-based framework for the leader-follower flocking scenario that allows a human to advise a specific follower with corrections in the action space. We extract useful rules to replace human intervention for more efficient DRL while reducing the workload of the human supervisor. Besides, we propose a transfer learning method to speed up the learning of the flocking policy. In simulation, we demonstrate the effectiveness of the proposed method by comparing DQN with and without human guidance. It is verified that the integration of human guidance into DQN can improve learning efficiency and strategy performance.