Attack–defense strategy of UAV swarm based on DEP-SIQ in the active target defense scenario
摘要
In this paper, a UAV swarm engagement scenario is considered, where the defender swarm tries to intercept the attacker swarm cooperatively to prevent it from entering into the target area. Different from the previous attack–defense strategy on deep reinforcement learning, we remove the assumption that the swarms have perfect knowledge of situation information of both sides and the environment. On this basis, a double experience pool strategic interaction Q-learning (DEP-SIQ) swarm attack and defense algorithm is proposed, which makes the dimension of the network input reduced and the UAV with the same task use the same network. The algorithm sets up different experience pools for the attacker and the defender, respectively. During training, both sides take samples from their own experience pool to train their own network. The estimation reward of the UAV is decomposed into the sum of the interaction reward values with other friendly UAVs, which is effectively suitable for large-scale swarm. Simulation experiments shows the feasibility of the proposed algorithm. Compare with other algorithms, the proposed DEP-SIQ algorithm has a faster learning efficiency, higher win rate and better applicability.