<p>Air combat maneuver decision based on reinforcement learning (RL) for within-visual-range air combat has been a long-standing challenge due to the highly dynamic maneuvers and long decision process. In this paper, we propose a novel reward self-shaping PPO with guided prioritized experience replay (RS-PPO with GPER), which is a sample efficient RL paradigm that involves flexible reward for different scenarios and introduction of expert guidance in RL training. RS-PPO aims to release the difficulty in reward function design for different scenarios; practically, we incorporate five sub-goals into the action space of PPO algorithm to inject flexibility in reward. We develop guided prioritized experience replay mechanism to expedite training process and improve the efficiency in exploration by introducing expert guidance. Experimental results demonstrate that the RS-PPO with GPER can achieve better convergence speed in different scenarios than conventional methods, and shows superiority on sample efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Air Combat Maneuver Decision Based on Deep Reinforcement Learning with Expert Guidance

  • Dunwang Li,
  • Wenhan Dong,
  • Lei He,
  • Ming Cai,
  • Pin Zhang,
  • Xin Zhang

摘要

Air combat maneuver decision based on reinforcement learning (RL) for within-visual-range air combat has been a long-standing challenge due to the highly dynamic maneuvers and long decision process. In this paper, we propose a novel reward self-shaping PPO with guided prioritized experience replay (RS-PPO with GPER), which is a sample efficient RL paradigm that involves flexible reward for different scenarios and introduction of expert guidance in RL training. RS-PPO aims to release the difficulty in reward function design for different scenarios; practically, we incorporate five sub-goals into the action space of PPO algorithm to inject flexibility in reward. We develop guided prioritized experience replay mechanism to expedite training process and improve the efficiency in exploration by introducing expert guidance. Experimental results demonstrate that the RS-PPO with GPER can achieve better convergence speed in different scenarios than conventional methods, and shows superiority on sample efficiency.