<p>To solve the problem of multi-ship collision avoidance in unpredictable maritime traffic environment, a multiple Marine Autonomous Surface ships (MASS) collision avoidance decision-making method based on the Deep Reinforcement Learning (DRL) algorithm is designed. The innovations of this study are as follows: (1) A clipping method based on the Penalized Point Policy Difference (P3D) is employed to improve the Proximal Policy Optimization (PPO) algorithm, which effectively mitigates excessive policy update amplitudes by constraining differences in point probabilities, rather than relying solely on measures of policy distribution disparities. (2) The Long Short-Term Memory (LSTM) network architecture facilitates the retention of historical state information, thereby enhancing the algorithm's predictive capability in modeling the state space at subsequent time steps. (3) Regarding the design of the reward function, we consider not only factors such as distance to target and stability of yaw angle, but also the COLREGs compliance and MASS maneuverability restrictions in ship domain. In the simulation, the proposed method exhibits excellent performance in terms of average reward, loss convergence, and success rate. It is noteworthy to emphasize that the proposed method demonstrates remarkable collision avoidance capability in complex scenes, and effectively responds to the emergency collision avoidance actions under immediate danger situation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A collision avoidance decision-making method for multiple marine autonomous surface ships based on P3DL-PPO algorithm

  • Zhewen Cui,
  • Wei Guan,
  • Xianku Zhang

摘要

To solve the problem of multi-ship collision avoidance in unpredictable maritime traffic environment, a multiple Marine Autonomous Surface ships (MASS) collision avoidance decision-making method based on the Deep Reinforcement Learning (DRL) algorithm is designed. The innovations of this study are as follows: (1) A clipping method based on the Penalized Point Policy Difference (P3D) is employed to improve the Proximal Policy Optimization (PPO) algorithm, which effectively mitigates excessive policy update amplitudes by constraining differences in point probabilities, rather than relying solely on measures of policy distribution disparities. (2) The Long Short-Term Memory (LSTM) network architecture facilitates the retention of historical state information, thereby enhancing the algorithm's predictive capability in modeling the state space at subsequent time steps. (3) Regarding the design of the reward function, we consider not only factors such as distance to target and stability of yaw angle, but also the COLREGs compliance and MASS maneuverability restrictions in ship domain. In the simulation, the proposed method exhibits excellent performance in terms of average reward, loss convergence, and success rate. It is noteworthy to emphasize that the proposed method demonstrates remarkable collision avoidance capability in complex scenes, and effectively responds to the emergency collision avoidance actions under immediate danger situation.