<p>Due to the variable operational environments and high-dimensional state and action spaces employed in multiple unmanned surface vehicle (multi-USV) scenarios, it is challenging to formulate accurate, rapid control policies for making decisions in real time, as well as stable and flexible strategies for addressing the difficulty of fitting the desired behavior. To address these challenges, this study proposes a hierarchical decision-making framework that integrates a high-level collaboration policy and a low-level control policy to effectively reduce the complexity of decision-making and improve the efficiency of training. Initially, to enable the policies to be optimized by RL, a Markov decision process is designed for the control policy, and a Markov game model is established to capture interactions in multi-USV scenarios. To eliminate the steady-state error and improve the convergence speed of the model, the proximal policy optimization with integral–differential compensation (PPO-IDC) algorithm is presented for cooperatively controlling the velocities and headings of the USVs. In addition, the PPO-IDC algorithm combines a policy and values in a network structure to accelerate the performed calculations. Furthermore, the course-learning delayed self-play (CL-DSP) training algorithm is proposed with delayed update networks to train stable and flexible strategies. The experimental results demonstrate that the proposed hierarchical framework enables accurate control, fast and stable convergence, and flexible collaboration in complex competitive multi-USV scenarios, highlighting its potential for real-time maritime decision-making applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An integral-differential compensation-based RL control policy for hierarchical multi-USV decision-making

  • Luyu Jia,
  • Chengtao Cai,
  • Xingmei Wang

摘要

Due to the variable operational environments and high-dimensional state and action spaces employed in multiple unmanned surface vehicle (multi-USV) scenarios, it is challenging to formulate accurate, rapid control policies for making decisions in real time, as well as stable and flexible strategies for addressing the difficulty of fitting the desired behavior. To address these challenges, this study proposes a hierarchical decision-making framework that integrates a high-level collaboration policy and a low-level control policy to effectively reduce the complexity of decision-making and improve the efficiency of training. Initially, to enable the policies to be optimized by RL, a Markov decision process is designed for the control policy, and a Markov game model is established to capture interactions in multi-USV scenarios. To eliminate the steady-state error and improve the convergence speed of the model, the proximal policy optimization with integral–differential compensation (PPO-IDC) algorithm is presented for cooperatively controlling the velocities and headings of the USVs. In addition, the PPO-IDC algorithm combines a policy and values in a network structure to accelerate the performed calculations. Furthermore, the course-learning delayed self-play (CL-DSP) training algorithm is proposed with delayed update networks to train stable and flexible strategies. The experimental results demonstrate that the proposed hierarchical framework enables accurate control, fast and stable convergence, and flexible collaboration in complex competitive multi-USV scenarios, highlighting its potential for real-time maritime decision-making applications.