A dual-arm cooperative control method based on improved proximal policy optimization
摘要
To address the challenges of high-dimensional control in dual-arm collaborative tasks, the complexity of multi-stage decision-making, and the limitations of traditional Proximal Policy Optimization (PPO) algorithms due to their single constraint mechanism, which results in policy bias and insufficient convergence efficiency, this paper proposes a dual-arm collaborative control method based on an improved Proximal Policy Optimization algorithm. Based on deep reinforcement learning, the state space and action space of the dual-arm system are first defined, and a perception-decision-update closed-loop interaction mechanism is constructed. Subsequently, a Hierarchical Constrained Hybrid Proximal Policy Optimization algorithm (HCH-PPO) is proposed, which designs a dual-timescale hierarchical policy, establishes dynamic hybrid constraints, and incorporates an adaptive parameter adjustment mechanism. While maintaining the efficiency of Proximal Policy Optimization (PPO), the algorithm introduces Trust Region Policy Optimization (TRPO) to enhance the stability of the optimization process and the policy exploration capability. This hierarchical optimization framework effectively enables efficient state-to-action mapping learning. Finally, experimental results demonstrate that, compared to traditional PPO, the proposed method achieves a 56.82% improvement in convergence speed and a 12% increase in task success rate in dual-arm collaborative grasping and placing tasks, indicating significant performance enhancement.