<p>Multi-player pursuit-evasion games under measurement noise are of great importance in aerospace and robotics. This paper investigates a differential game involving three cooperative pursuers and a single superior evader under system uncertainties. The payoff function is the terminal distance of the evader and its nearest pursuer. To tackle this problem, a deep reinforcement learning framework is introduced. The main contributions are threefold: First, a weighted pursuit strategy structure is proposed, where the weighting coefficient is optimized through the Deep Deterministic Policy Gradient. Second, a novel reward function, defined by the change in Stackelberg game values between consecutive time steps, is designed to mitigate sparse rewards and enhance training stability. Third, comparative simulations against a reachable set strategy demonstrate the effectiveness of the proposed approach, achieving a smaller terminal miss and confirming its potential for real-world applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A weighted Stackelberg strategy approach for three pursuers against a superior evader using reinforcement learning

  • Ziyi Zhan,
  • Yiqun Zhang,
  • Pengfei Zhang

摘要

Multi-player pursuit-evasion games under measurement noise are of great importance in aerospace and robotics. This paper investigates a differential game involving three cooperative pursuers and a single superior evader under system uncertainties. The payoff function is the terminal distance of the evader and its nearest pursuer. To tackle this problem, a deep reinforcement learning framework is introduced. The main contributions are threefold: First, a weighted pursuit strategy structure is proposed, where the weighting coefficient is optimized through the Deep Deterministic Policy Gradient. Second, a novel reward function, defined by the change in Stackelberg game values between consecutive time steps, is designed to mitigate sparse rewards and enhance training stability. Third, comparative simulations against a reachable set strategy demonstrate the effectiveness of the proposed approach, achieving a smaller terminal miss and confirming its potential for real-world applications.