错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Offline Modeling Approach to Air Combat Maneuvering Policy

  • Zhu Ziqiang,
  • Fu Yupeng,
  • Deng Xiangyang,
  • Xu Tao

摘要

To address the issues of high environment exploration cost and low utilization of empirical data in the online Reinforcement Learning (RL) based air combat maneuver policy design, the offline modeling method of maneuver policy is investigated, and the policy-guided implicit Q-learning algorithm (PIQL) is proposed. According to the Minimax Theory, the policy model is designed as a guidance network and an execution network, the guidance network predicts the current situation and outputs the predicted disadvantageous situation, and the execution network selects the optimal maneuvering behaviors based on the predicted situation and self-state, which improves the model’s out-of-distribution generalizability through decoupling of the situations and behaviors. Updating the parameters based on the implicit Q-learning algorithm (IQL) during training improves the model performance by utilizing the in-distribution data. Simulations are based on the data sampled from suboptimal policy, and the results show that the algorithm improves the episode returns, and the generated policy model presents better intelligence.