Distributed Advantage-Based Weights Reshaping Algorithm with Sparse Reward
摘要
The air-combat is a dynamic game process with rapid situation switching, while conventional method for strategy modeling use the parametric approach as the reward, difficult to demonstrate the positive correlation between returns and win rate. In this work, the strategy model is trained by Reinforcement Learning (RL) with sparse reward function based on sparse events that are deep bonding with game results. To address the issue of slow convergence of sparse reward task, based on the Distributed Proximal Policy Optimization (DPPO) algorithm framework, this paper proposes an advantage-based weights reshaping algorithm, which optimizes value network and policy gradient weights, avoiding value overestimation issues in sparse reward environment. Simulation results show that the proposed algorithm improves both the convergence and returns of the algorithm, as well as the win rate.