A control method based on proximal policy optimisation (PPO) algorithm is proposed to control autonomous berthing of unmanned vessels. Firstly, Bessel curves are used to generate the desired path, and the heading error, speed error, and lateral error are obtained by calculating with the current state of the ship, and then these errors are used as inputs to establish the reward function mechanism, and then the state of the ship and the reward obtained are inputted into the PPO algorithm so that the intelligent body can take behaviours according to the state and the reward, and finally, the actions of the intelligent body are inputted into the ship model to achieve bi-variable control of the ship’s rudder angle and thrust. Finally, the actions of the intelligent body are inputted into the ship model to achieve the bivariate control of the rudder angle and thrust. The experimental results show the effectiveness of PPO reinforcement learning in the automatic berthing process of unmanned vessels with low speed, interference and low rudder efficiency.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Based on PPO Reinforcement Learning for Ship Berthing and Departure Control

  • Ziyan Zhou,
  • Feng Ma,
  • Xuan Zhou

摘要

A control method based on proximal policy optimisation (PPO) algorithm is proposed to control autonomous berthing of unmanned vessels. Firstly, Bessel curves are used to generate the desired path, and the heading error, speed error, and lateral error are obtained by calculating with the current state of the ship, and then these errors are used as inputs to establish the reward function mechanism, and then the state of the ship and the reward obtained are inputted into the PPO algorithm so that the intelligent body can take behaviours according to the state and the reward, and finally, the actions of the intelligent body are inputted into the ship model to achieve bi-variable control of the ship’s rudder angle and thrust. Finally, the actions of the intelligent body are inputted into the ship model to achieve the bivariate control of the rudder angle and thrust. The experimental results show the effectiveness of PPO reinforcement learning in the automatic berthing process of unmanned vessels with low speed, interference and low rudder efficiency.