<p>This paper proposes a deep reinforcement learning (DRL) method for the course control of an underactuated autonomous underwater vehicle (AUV). The control task considers time-varying external disturbances (ED) and input constraints (IC) within an optimal control framework. Traditional DRL methods for AUV motion control suffer from poor generalization and limited control stability under ED. To address these issues, we simplify the Actor-Critic algorithm and integrate it with optimal control. This yields a nonlinear motion control scheme named op-AC. The op-AC method has several key improvements. First, all neural network (NN) training is completed offline before control tasks begin. The controller can therefore directly apply the learned policy without online optimization, and retraining is unnecessary even when the desired course changes. Second, the training process is simplified compared to traditional DRL methods. Third, we design a novel action and reward mechanism over a sufficiently long time step for offline optimal control. This mechanism better handles the effects of ED and IC. Simulation results demonstrate that the proposed controller accurately drives the underactuated AUV to follow the desired course despite ED and IC. It also shows clear advantages over other comparison methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Underactuated AUV course stability control based on optimal deep reinforcement learning with external disturbances and input constraints

  • Zhanyuan Wang,
  • Rongmin Chen,
  • Yuchen Liao,
  • Yanyun Wang,
  • Fuqiang Luo,
  • Jiaxi Wang,
  • Shuoshuo Ding,
  • Chengfu Liu,
  • Dapeng Jiang,
  • Kailin Xin

摘要

This paper proposes a deep reinforcement learning (DRL) method for the course control of an underactuated autonomous underwater vehicle (AUV). The control task considers time-varying external disturbances (ED) and input constraints (IC) within an optimal control framework. Traditional DRL methods for AUV motion control suffer from poor generalization and limited control stability under ED. To address these issues, we simplify the Actor-Critic algorithm and integrate it with optimal control. This yields a nonlinear motion control scheme named op-AC. The op-AC method has several key improvements. First, all neural network (NN) training is completed offline before control tasks begin. The controller can therefore directly apply the learned policy without online optimization, and retraining is unnecessary even when the desired course changes. Second, the training process is simplified compared to traditional DRL methods. Third, we design a novel action and reward mechanism over a sufficiently long time step for offline optimal control. This mechanism better handles the effects of ED and IC. Simulation results demonstrate that the proposed controller accurately drives the underactuated AUV to follow the desired course despite ED and IC. It also shows clear advantages over other comparison methods.