<p>In this paper, we investigate the mission planning problem for multiple Unmanned Aerial Vehicles (UAVs), where visit gains and the resource requirements for each mission are dependent on UAV types. This problem involves both mission allocation and path planning. Existing deep reinforcement learning (DRL) algorithms, particularly those utilizing graph attention networks (GAT), face challenges in extracting intricate state features. Additionally, these algorithms often lack practical applicability for mission planning problems where visit gains are highly dependent on UAV types. To address these issues, we propose an end-to-end DRL framework that employs a heterogeneous graph attention network (HAN) to derive the optimal scheduling strategy. Firstly, the problem is formulated as a Markov decision process (MDP). In this formulation, UAVs and missions are represented as a fully connected heterogeneous graph, which serves as the scheduling state across different time steps. We develop a heterogeneous graph neural network and integrate it with a multi-head attention mechanism to embed the latent relationships between various types of nodes and edges, ultimately yielding an overall graph embedding. The graph-level representation vector of the heterogeneous graph, along with the node vectors, is fed into a Transformer-based decoder to autoregressively generate the sequence of nodes to be visited. The reward function is designed to maximize the total visit gains while minimizing the total flight distance and flight time. Moreover, we propose a DRL training algorithm based on an Actor-Critic framework. The critic network estimates state-value functions by utilizing a fully connected HAN. Experimental results show that our proposed model exhibits a substantial advantage in solution efficiency, compared to CPLEX, several well-known heuristics, and a mainstream DRL framework. Specifically, when compared to the widely adopted mainstream GAT-based method (a representative architecture for neural combinatorial optimization), the average optimal gap for large-scale cases is reduced from 12.36% to 2.22%, while maintaining comparable solution times.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-UAV Mission Planning Based on Heterogeneous Graph Attention Network and Deep Reinforcement Learning

  • Jian-Jiang Wang,
  • Xiao Feng,
  • Rui Wang,
  • Xue-Jun Hu

摘要

In this paper, we investigate the mission planning problem for multiple Unmanned Aerial Vehicles (UAVs), where visit gains and the resource requirements for each mission are dependent on UAV types. This problem involves both mission allocation and path planning. Existing deep reinforcement learning (DRL) algorithms, particularly those utilizing graph attention networks (GAT), face challenges in extracting intricate state features. Additionally, these algorithms often lack practical applicability for mission planning problems where visit gains are highly dependent on UAV types. To address these issues, we propose an end-to-end DRL framework that employs a heterogeneous graph attention network (HAN) to derive the optimal scheduling strategy. Firstly, the problem is formulated as a Markov decision process (MDP). In this formulation, UAVs and missions are represented as a fully connected heterogeneous graph, which serves as the scheduling state across different time steps. We develop a heterogeneous graph neural network and integrate it with a multi-head attention mechanism to embed the latent relationships between various types of nodes and edges, ultimately yielding an overall graph embedding. The graph-level representation vector of the heterogeneous graph, along with the node vectors, is fed into a Transformer-based decoder to autoregressively generate the sequence of nodes to be visited. The reward function is designed to maximize the total visit gains while minimizing the total flight distance and flight time. Moreover, we propose a DRL training algorithm based on an Actor-Critic framework. The critic network estimates state-value functions by utilizing a fully connected HAN. Experimental results show that our proposed model exhibits a substantial advantage in solution efficiency, compared to CPLEX, several well-known heuristics, and a mainstream DRL framework. Specifically, when compared to the widely adopted mainstream GAT-based method (a representative architecture for neural combinatorial optimization), the average optimal gap for large-scale cases is reduced from 12.36% to 2.22%, while maintaining comparable solution times.