A weighted Stackelberg strategy approach for three pursuers against a superior evader using reinforcement learning
摘要
Multi-player pursuit-evasion games under measurement noise are of great importance in aerospace and robotics. This paper investigates a differential game involving three cooperative pursuers and a single superior evader under system uncertainties. The payoff function is the terminal distance of the evader and its nearest pursuer. To tackle this problem, a deep reinforcement learning framework is introduced. The main contributions are threefold: First, a weighted pursuit strategy structure is proposed, where the weighting coefficient is optimized through the Deep Deterministic Policy Gradient. Second, a novel reward function, defined by the change in Stackelberg game values between consecutive time steps, is designed to mitigate sparse rewards and enhance training stability. Third, comparative simulations against a reachable set strategy demonstrate the effectiveness of the proposed approach, achieving a smaller terminal miss and confirming its potential for real-world applications.