<p>Vehicle scheduling and supply chain inventory management have long adopted a split optimization strategy, which fails to fully consider their dynamic coupling relationship, resulting in increased overall operating costs and reduced resource utilization. This paper addresses the inefficiency caused by the absence of collaborative optimization. A joint optimization algorithm integrating a graph attention mechanism and multi-agent deep reinforcement learning is proposed, with a structure-aware strategy learning framework designed to achieve efficient collaboration among agents through centralized training and distributed execution. The algorithm integrates inventory levels, vehicle status, demand forecasts, and traffic disturbances into a unified state representation, constructs an Actor-Critic structure based on Proximal Policy Optimization, jointly models and optimizes inventory replenishment and vehicle scheduling strategies, and enhances convergence efficiency using a time-difference error weighted experience replay mechanism. Experimental results show that when the supply chain scale expands from 10 to 50 nodes, the joint cost rises from 34,500 yuan to 70,500 yuan. Under extremely high demand disturbance, the out-of-stock rate is controlled at 8.1%, and over 83% of customer service scores exceed 0.85, demonstrating the algorithm’s collaborative stability and service assurance in a complex dynamic environment. The study demonstrates that the method coordinates inventory and transportation efficiently under dynamic uncertainty, providing a viable path and theoretical support for supply chain operation optimization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Joint optimization algorithm for vehicle scheduling and supply chain inventory management based on multi-agent deep reinforcement learning

  • Linlin Feng

摘要

Vehicle scheduling and supply chain inventory management have long adopted a split optimization strategy, which fails to fully consider their dynamic coupling relationship, resulting in increased overall operating costs and reduced resource utilization. This paper addresses the inefficiency caused by the absence of collaborative optimization. A joint optimization algorithm integrating a graph attention mechanism and multi-agent deep reinforcement learning is proposed, with a structure-aware strategy learning framework designed to achieve efficient collaboration among agents through centralized training and distributed execution. The algorithm integrates inventory levels, vehicle status, demand forecasts, and traffic disturbances into a unified state representation, constructs an Actor-Critic structure based on Proximal Policy Optimization, jointly models and optimizes inventory replenishment and vehicle scheduling strategies, and enhances convergence efficiency using a time-difference error weighted experience replay mechanism. Experimental results show that when the supply chain scale expands from 10 to 50 nodes, the joint cost rises from 34,500 yuan to 70,500 yuan. Under extremely high demand disturbance, the out-of-stock rate is controlled at 8.1%, and over 83% of customer service scores exceed 0.85, demonstrating the algorithm’s collaborative stability and service assurance in a complex dynamic environment. The study demonstrates that the method coordinates inventory and transportation efficiently under dynamic uncertainty, providing a viable path and theoretical support for supply chain operation optimization.