De-Supply: Deep Reinforcement Learning for Multivariate Supply Chain Optimization
摘要
This paper presents De-Supply, a novel approach utilizing deep reinforcement learning for multi-objective supply-chain optimization. It formulates the problem as a Markov Decision Process with complex action and observation spaces, addressing real-world supply chain challenges. De-Supply’s custom policy network enables agents to make informed procurement decisions, resulting in efficient stock management and cost control. Extensive experiments demonstrate its effectiveness, using historical data to outperform baselines and human-level performance. De-Supply offers robust solutions for warehouse businesses and holds promise for broader applications in the industry.