<p>Quadruped robots’ nonlinear complexity makes traditional modeling challenging, while deep reinforcement learning (DRL) learns effectively through direct environmental interaction without explicit kinematic and dynamic models, becoming an efficient approach for quadruped locomotion across diverse terrains. Conventional reinforcement learning methods typically combine multiple reward criteria into a single scalar function, limiting information representation and complicating the balance between multiple control objectives. We propose a novel multi-head critic and dynamic policy gradient SAC (MHD-SAC) algorithm, innovatively combining a multi-head critic architecture that independently evaluates distinct reward components and a dynamic policy gradient method that adaptively adjusts weights based on current performance. Through simulations on both flat and uneven terrains comparing three approaches (Soft Actor-Critic (SAC), multi-head critic SAC (MH-SAC), and MHD-SAC), we demonstrate that the MHD-SAC algorithm achieves significantly faster learning convergence and higher cumulative rewards than conventional methods. Performance analysis across different reward components reveals MHD-SAC’s superior ability to balance multiple objectives. The results validate that our approach effectively addresses the challenges of multi-objective optimization in quadruped locomotion control, providing a promising foundation for developing more versatile and robust legged robots capable of traversing complex environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Quadruped robot locomotion via soft actor-critic with muti-head critic and dynamic policy gradient

  • Yanan Fan,
  • Zhongcai Pei,
  • Hongbing Shi,
  • Meng Li,
  • Tianyuan Guo,
  • Zhiyong Tang

摘要

Quadruped robots’ nonlinear complexity makes traditional modeling challenging, while deep reinforcement learning (DRL) learns effectively through direct environmental interaction without explicit kinematic and dynamic models, becoming an efficient approach for quadruped locomotion across diverse terrains. Conventional reinforcement learning methods typically combine multiple reward criteria into a single scalar function, limiting information representation and complicating the balance between multiple control objectives. We propose a novel multi-head critic and dynamic policy gradient SAC (MHD-SAC) algorithm, innovatively combining a multi-head critic architecture that independently evaluates distinct reward components and a dynamic policy gradient method that adaptively adjusts weights based on current performance. Through simulations on both flat and uneven terrains comparing three approaches (Soft Actor-Critic (SAC), multi-head critic SAC (MH-SAC), and MHD-SAC), we demonstrate that the MHD-SAC algorithm achieves significantly faster learning convergence and higher cumulative rewards than conventional methods. Performance analysis across different reward components reveals MHD-SAC’s superior ability to balance multiple objectives. The results validate that our approach effectively addresses the challenges of multi-objective optimization in quadruped locomotion control, providing a promising foundation for developing more versatile and robust legged robots capable of traversing complex environments.