<p>This study presents a novel approach using continuous-time Q-learning for tracking control of a three-wheeled mobile robot (WMR). The method is designed based on the principles of zero-sum game theory, and it does not rely on any prior knowledge of the robot’s model. The algorithm proposed in this study is based on a single control loop, as opposed to the utilization of a complex mathematical model for the WMR with separate kinematic and dynamic control loops. The algorithm is developed using an actor-critic disturbance (ACD) structure. The ACD structure incorporates three neural networks that utilize synchronized update laws, which have been developed using the integral reinforcement learning (IRL) methodology. Lyapunov stability criteria ensure the convergence of the weight update rules for the networks inside this structure, as well as the stability of the closed-loop system. The efficacy of the method is demonstrated through the simulation results.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Model-Free Q-Learning-Based Adaptive Optimal Control for Wheeled Mobile Robot

  • Cuong Nguyen Duc,
  • Sen Huong Thi Pham,
  • Nga Thi-Thuy Vu

摘要

This study presents a novel approach using continuous-time Q-learning for tracking control of a three-wheeled mobile robot (WMR). The method is designed based on the principles of zero-sum game theory, and it does not rely on any prior knowledge of the robot’s model. The algorithm proposed in this study is based on a single control loop, as opposed to the utilization of a complex mathematical model for the WMR with separate kinematic and dynamic control loops. The algorithm is developed using an actor-critic disturbance (ACD) structure. The ACD structure incorporates three neural networks that utilize synchronized update laws, which have been developed using the integral reinforcement learning (IRL) methodology. Lyapunov stability criteria ensure the convergence of the weight update rules for the networks inside this structure, as well as the stability of the closed-loop system. The efficacy of the method is demonstrated through the simulation results.