A multi-UAV rapid post-disaster search and rescue method based on deep reinforcement learning
摘要
Deep reinforcement learning shows broad prospects in multi-unmanned aerial vehicle(UAV) collaborative search and rescue tasks. However, in the face of high-dimensional collaborative decision-making spaces and limited computing resources, its performance is vulnerable to limitations. This paper proposes a deep deterministic policy gradient method based on linear attention. By introducing the linear attention mechanism based on random feature mapping, while effectively modeling the interaction among UAVs, the computational and storage overcosts caused by the increase in the number of UAVs have been significantly reduced. Furthermore, by combining smooth experience replay and adaptive importance sampling mechanism, the training efficiency and strategy stability have been further improved. The simulation experiments on both post-disaster response search and dynamic containment tasks demonstrate that the proposed algorithm consistently outperforms existing methods. In small-scale scenarios, it maintains nearly perfect success rates, while in medium- and large-scale settings it achieves up to 90.6% and 85.2% success rates in the post-disaster response search task and up to 90.1% and 80.2% in the containment task, corresponding to relative improvements of 15–21% over baselines. These results highlight both the robustness of the method in simple cases and its clear advantage under more challenging multi-UAV conditions.