A Method for Security Traffic Patrolling Based on Structural Coordinated Proximal Policy Optimization
摘要
Multi-agent patrolling has significant implications for addressing real-world security concerns. In multi-agent systems, the actions of agent directly influence those with whom it interacts. Traditional reinforcement learning-based multi-agent security patrolling methods overlook the role of these localized interactions in coordination among agents, thus failing to enhance the efficiency. To address this issue, this paper introduces a security patrolling approach based on Structured Coordinated Proximal Policy Optimization (PPO). The multi-agent patrolling task is modeled as a finite-time-step distributed partially observable semi-Markov decision process. This method, grounded in the Shapley Value, designs a multi-agent credit allocation function. The efficiency of this function is amplified using the structure of localized interactions. By accurately evaluating the contributions of each agent’s selected actions, this function fosters enhanced coordination among agents. Extensive experiments in various scenarios were conducted, and the results demonstrate that our algorithm outperforms benchmark algorithms in terms of convergence speed and patrolling performance.