Post-disaster Emergency Supplies Distribution Optimization: A Deep Reinforcement Learning Approach
摘要
After the sudden disaster occurred, the damage to communication infrastructure and road capacity can lead to an explosive increase in demand for emergency supplies, and the demands for different types of emergency supplies can vary depending on the period. Due to the poor communication between the affected area and the external environment, the demand generation pattern will be unpredictable. To address this challenge, we propose a novel anticipatory routing, acceptance, and postponement policy for the Multi-period Dynamic Vehicle Routing Problem with Stochastic Requests (MDVRPSR). We first provide a detailed description of the unique features of MDVRPSR, followed by the formulation of a mathematical model using Markov Decision Process (MDP). Our model is designed to respond to stochastic requests within a multi-period dynamic framework, offering a comprehensive perspective on the decision state, reward setting, transition process, and objective function. Subsequently, we employ Deep Reinforcement Learning (DRL) to optimize the routing policy. Experiments demonstrate that the DRL-based policy improves the effectiveness of emergency supplies distribution in dynamic and uncertain scenarios.