A Formation Control Deep Reinforcement Learning Model Based on MAPPO Algorithm for Unmanned Ground Vehicles
摘要
In order to address the problem of cooperative formation planning of multiple Unmanned Ground Vehicles (UGVs) in complex environments, a formation control method based on the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm is proposed, where firstly, the path planning problem is transformed into a partially observable Markov decision process, secondly, the MAPPO algorithm is extended to multi-UGVs with centralized training and decentralized execution, and finally, a simulation scenario is designed to verify the effect of the algorithm. To ensure the consistency with the real scene, the simulation scene contains multiple unmanned vehicles, multiple static obstacles, and multiple random dynamic obstacles. The experimental results show that in the complex environment where multiple dynamic and static obstacles exist, the method in this paper could control multiple unmanned vehicles to maneuver from the starting area to the destination area according to the perceived environmental information, with better obstacle avoidance and formation maintenance effects.