An Information-Guided Cooperative Hunting Control Strategy for Multi-UAV Based on Reinforcement Learning
摘要
In recent years, reinforcement learning has been widely applied in multi-UAV cooperative hunting tasks. In the dynamic environment of cooperative hunting, existing reinforcement learning methods suffer from unstable convergence under rapidly changing conditions and demonstrate suboptimal training efficiency owing to inadequate information guidance during the process of forming the encirclement formation. Therefore, this paper proposes an information-guided cooperative hunting control strategy for multi-UAV to pursuit an evader UAV based on reinforcement learning (HA-MAPPO). Firstly, considering the complexity of the hunting task, we divided it into two phases: “approaching” and “encircling”, allowing UAV to separately approach the target and form an encirclement formation, significantly improving training efficiency. Secondly, based on the virtual structure method, we generate multiple encirclement points, and use the Hungarian Algorithm (HA) to assign one encirclement point to each UAV, to introduce the information of encirclement points for efficient training. In the experimental results, HA-MAPPO algorithm attains stable convergence within approximately 50% fewer training steps compared to Multi-Agent Proximal Policy Optimization (MAPPO), accompanied by markedly enhanced stability in reward curve progression.