This study proposes a novel Multi-Head Attention-enhanced Proximal Policy Optimization (MHA-PPO) algorithm for collaborative truck-drone delivery systems in dynamic logistics scenarios. To address the limitations of traditional multi-agent reinforcement learning in handling complex spatiotemporal dependencies and real-time decision-making, we integrate a multi-head attention mechanism into the PPO framework, enabling adaptive feature fusion of heterogeneous agent observations (e.g., drone positions, truck routes, and package urgency levels). A key innovation lies in the design of a variable emergency mechanism that models a random urgency generation module and dynamically adjusts task priorities based on the urgency of customer demand for packages. The system is implemented in a high-fidelity Unity simulation environment incorporating 3D urban terrain, realistic drone aerodynamics, and traffic flow patterns. Comparative experiments against baseline algorithms (PPO, Multi-Agent Proximal Policy Optimization (MAPPO), Multi-Head Attention Deep Deterministic Policy Gradient (MHA-DDPG)) demonstrate that our MHA-PPO achieves superior performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Collaborative Optimization of Truck-Drone Based on Multi-agent Reinforcement Learning

  • Wenhao Zhang,
  • Guoyin Wang,
  • Qun Liu

摘要

This study proposes a novel Multi-Head Attention-enhanced Proximal Policy Optimization (MHA-PPO) algorithm for collaborative truck-drone delivery systems in dynamic logistics scenarios. To address the limitations of traditional multi-agent reinforcement learning in handling complex spatiotemporal dependencies and real-time decision-making, we integrate a multi-head attention mechanism into the PPO framework, enabling adaptive feature fusion of heterogeneous agent observations (e.g., drone positions, truck routes, and package urgency levels). A key innovation lies in the design of a variable emergency mechanism that models a random urgency generation module and dynamically adjusts task priorities based on the urgency of customer demand for packages. The system is implemented in a high-fidelity Unity simulation environment incorporating 3D urban terrain, realistic drone aerodynamics, and traffic flow patterns. Comparative experiments against baseline algorithms (PPO, Multi-Agent Proximal Policy Optimization (MAPPO), Multi-Head Attention Deep Deterministic Policy Gradient (MHA-DDPG)) demonstrate that our MHA-PPO achieves superior performance.