Path planning in dynamic structured environments using transformer-enabled twin delayed deep deterministic policy gradient for mobile robots in simulation
摘要
Mobile robot path planning in partially observable environments remains a fundamental challenge in robotics applications. While deep reinforcement learning (DRL) approaches have shown promise, existing methods suffer from several limitations: slow training convergence, low success rates, and susceptibility to local optima, primarily due to insufficient environmental perception capabilities. To address these challenges, we propose a novel algorithm called Transformer-TD3, which integrates a Transformer network with the Twin Delayed Deep Deterministic Policy Gradient architecture. Our approach incorporates a Transformer-based state feature extraction network that significantly enhances the algorithm’s ability to perceive and adapt to unknown environments. This innovation substantially reduces unnecessary exploration, thereby accelerating network convergence. Furthermore, we employ heuristic search to determine dynamic sub-goals, which the Transformer-TD3 algorithm then utilizes for precise path planning. The robot efficiently navigates toward its target by intelligently transitioning between these sub-goals, resulting in improved path planning success rates and enhanced path efficiency. Extensive simulation experiments validate the effectiveness of our proposed algorithm in continuous action space mobile robot path planning. The results demonstrate notable improvements in minimizing redundant exploration, ensuring robust obstacle avoidance, and optimizing path generation strategies. Most significantly, our experimental findings show substantial enhancements in robot autonomy and adaptability within complex, partially observable environments. These results underscore the potential of our approach in advancing the field of autonomous mobile robot navigation.