Obstacle Avoidance for Guided Quadruped Robots in Complex Environments
摘要
When guiding vision-impaired individuals, quadruped robots must promptly avoid obstacles and crowds and follow a safe path. Though reinforcement learning approaches are widely adopted for obstacle avoidance and navigation in complex and unknown environments, most of them suffer from low learning efficiency. This work proposes an interaction of hindsight experience replay (HER), intermediate waypoints, and a direction-aware reward function to tackle reward sparsity and learning inefficiency. Training results in the simulation environment demonstrate that the proposed algorithm significantly increases both success rate and reward compared to the Twin Delayed Deep Policy Gradient (TD3) algorithm. Furthermore, test results on the trained model indicate that the proposed algorithm performs better than TD3 in the same environment. This work enables quadruped robots to perform obstacle-avoidance navigation in complex and crowded environments, reinforcing their safety in guiding vision-impaired individuals.