In the orbit cluster game problem, using multi-agent reinforcement learning (MARL) algorithms for policy solving has become a common method. However, MARL methods typically train in a single environment, which fails to capture the underlying causality of the environment [5], resulting in poor performance in real-world environments different from the training environment, especially in scenarios with multiple highly variable disturbances. To address the issue, we propose a Causality Diversity Maximal Marginal Relevance (CDMMR) environment selection algorithm, aimed at selecting suitable training environments for multi-disturbance scenarios in orbit games. By analyzing the causality generated by disturbance factors on the game policy, we filter the environment subsets that can maximize diversity for parallel interaction through multi-agent proximal policy optimization (MAPPO), thereby constructing a robust orbit game policy suitable for multiple disturbances environments. Without loss of generality, the satellite cluster cooperation pursuit of space debris is selected as the scenario. Atmospheric drag, solar radiation pressure, satellite mass, and solar illumination are used as multiple disturbances to numerically simulate and analyze the proposed method. Simulation experiments show that the proposed CDMMR method significantly improves robustness compared to the original MAPPO in multi-disturbance environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust Orbital Game Policy in Multiple Disturbed Environments: An Approach Based on Causality Diversity Maximal Marginal Relevance Algorithm

  • Xiao Wang,
  • Yuying Han,
  • Min Tang,
  • Fei Zhang

摘要

In the orbit cluster game problem, using multi-agent reinforcement learning (MARL) algorithms for policy solving has become a common method. However, MARL methods typically train in a single environment, which fails to capture the underlying causality of the environment [5], resulting in poor performance in real-world environments different from the training environment, especially in scenarios with multiple highly variable disturbances. To address the issue, we propose a Causality Diversity Maximal Marginal Relevance (CDMMR) environment selection algorithm, aimed at selecting suitable training environments for multi-disturbance scenarios in orbit games. By analyzing the causality generated by disturbance factors on the game policy, we filter the environment subsets that can maximize diversity for parallel interaction through multi-agent proximal policy optimization (MAPPO), thereby constructing a robust orbit game policy suitable for multiple disturbances environments. Without loss of generality, the satellite cluster cooperation pursuit of space debris is selected as the scenario. Atmospheric drag, solar radiation pressure, satellite mass, and solar illumination are used as multiple disturbances to numerically simulate and analyze the proposed method. Simulation experiments show that the proposed CDMMR method significantly improves robustness compared to the original MAPPO in multi-disturbance environments.