<p>Aiming at the problem of efficient collaborative decision-making among multiple satellites in multi-satellite orbital interception mission, a multi-satellite cooperative orbital interception strategy based on attention mechanism and deep reinforcement learning is proposed. Firstly, considering the orbital dynamics, maneuverability, mission duration, and collision avoidance constraints faced in orbital interception mission, a Markov decision process is designed, and a multi-satellite game strategy solution framework based on the actor–critic network is built; then, a guided reward function is designed to effectively guide the pursuit satellite to approach and intercept the escaping satellite to accelerate the convergence speed of the algorithm; finally, the attention mechanism is used to capture the potential relationship between satellites and generate coded information with biased attention effect, which helps satellites form an efficient collaborative interception strategy. The simulation experiments show that the trained satellites can conduct autonomous learning and decision-making in a dynamic and uncertain environment. In orbital interception mission, the pursuit satellite can adopt an effective collaborative strategy to use the advantage of quantity to make up for the disadvantage of speed, maintain a high mission success rate, and a series of intelligent game behaviors emerge.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention-driven reinforcement learning for multi-satellite collaborative orbital interception strategy solution

  • Wenxiu Zhang,
  • Yamin Wang,
  • Yonghe Zhang

摘要

Aiming at the problem of efficient collaborative decision-making among multiple satellites in multi-satellite orbital interception mission, a multi-satellite cooperative orbital interception strategy based on attention mechanism and deep reinforcement learning is proposed. Firstly, considering the orbital dynamics, maneuverability, mission duration, and collision avoidance constraints faced in orbital interception mission, a Markov decision process is designed, and a multi-satellite game strategy solution framework based on the actor–critic network is built; then, a guided reward function is designed to effectively guide the pursuit satellite to approach and intercept the escaping satellite to accelerate the convergence speed of the algorithm; finally, the attention mechanism is used to capture the potential relationship between satellites and generate coded information with biased attention effect, which helps satellites form an efficient collaborative interception strategy. The simulation experiments show that the trained satellites can conduct autonomous learning and decision-making in a dynamic and uncertain environment. In orbital interception mission, the pursuit satellite can adopt an effective collaborative strategy to use the advantage of quantity to make up for the disadvantage of speed, maintain a high mission success rate, and a series of intelligent game behaviors emerge.