Optimizing Twin Delayed Deep Deterministic Policy Gradient for 7-DOF Robotic Arm Grasping: Overcoming Suboptimality with Exploration-Enhanced Contrastive Learning
摘要
Twin delayed deep deterministic policy gradient (TD3) is a robust model-free reinforcement learning (RL) algorithm effective for robotic arm manipulation. However, TD3 has not been extensively researched, leading to limited adaptability in dynamic manufacturing environments, while traditional control methods face computational impracticality. We developed an exploration-enhanced contrastive learning (EECL) module for widespread adoption of TD3. The EECL module integrates with TD3, a state buffer, and a KDTree to efficiently identify robots’ states and provide additional and intrinsic rewards for their discovery. This targeted incentive mechanism accelerates exploration and improves TD3’s adaptability. The EECL-enhanced TD3 algorithm was evaluated on a simulated 7-DOF robotic arm manipulation task. Experimental results consistently demonstrated significant improvements over baseline TD3 across multiple random seeds. Specifically, the EECL module increased average cumulative rewards, achieved faster convergence speed, and enhanced exploration efficiency, facilitating diverse operations and policy optimization. This enhanced performance validates the EECL module’s ability to address inherent challenges in RL for complex robotic controls. The module’s generalizability is validated, showing its potential for integration with other RL algorithms and enabling efficient and adaptable reinforcement learning strategies.