<p>This paper proposes a Reinforcement Learning framework for bipedal robots that integrates Lyapunov stability principles with gait prediction based on the Linear Inverted Pendulum Model (LIPM). Unlike conventional approaches that rigidly enforce the policy to track model-planned footholds, the proposed method emphasizes the convergence trend whereby actual footholds gradually approach the desired ones. To achieve this, a neural Lyapunov critic is constructed to incorporate LIPM-predicted footholds and their errors as inputs, learning energy-decreasing patterns to softly constrain policy behavior. Furthermore, a Lyapunov-augmented advantage estimation mechanism is developed, enabling the policy to benefit from both task rewards and stability compensation, thereby unifying physical interpretability, stability guarantees, and the flexibility of Reinforcement Learning. This study conducts large-scale experiments in Isaac Gym with 4096 parallel environments and further validate the approach on the TRON1A bipedal robot. Results demonstrate that the proposed method achieves superior performance in terrain adaptability, velocity tracking capability, stair-climbing success rate, and highly dynamic locomotion tasks. Moreover, the approach exhibits strong sim-to-real transferability, confirming its robustness and generalization in diverse real-world environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Neural Lyapunov-Guided Reinforcement Learning with LIPM Foothold Convergence for Stable Bipedal Locomotion

  • Yan Luo,
  • Li Ping Chen,
  • Yu Hao Song,
  • Huang Chen,
  • Jian Wan Ding

摘要

This paper proposes a Reinforcement Learning framework for bipedal robots that integrates Lyapunov stability principles with gait prediction based on the Linear Inverted Pendulum Model (LIPM). Unlike conventional approaches that rigidly enforce the policy to track model-planned footholds, the proposed method emphasizes the convergence trend whereby actual footholds gradually approach the desired ones. To achieve this, a neural Lyapunov critic is constructed to incorporate LIPM-predicted footholds and their errors as inputs, learning energy-decreasing patterns to softly constrain policy behavior. Furthermore, a Lyapunov-augmented advantage estimation mechanism is developed, enabling the policy to benefit from both task rewards and stability compensation, thereby unifying physical interpretability, stability guarantees, and the flexibility of Reinforcement Learning. This study conducts large-scale experiments in Isaac Gym with 4096 parallel environments and further validate the approach on the TRON1A bipedal robot. Results demonstrate that the proposed method achieves superior performance in terrain adaptability, velocity tracking capability, stair-climbing success rate, and highly dynamic locomotion tasks. Moreover, the approach exhibits strong sim-to-real transferability, confirming its robustness and generalization in diverse real-world environments.