Research on Food Dynamic Pricing Algorithm Based on Deep Reinforcement Learning
摘要
In order to mitigate food waste resulting from irrational pricing and concurrently augment overall business revenue, this paper introduces a finite-shelf-life food dynamic pricing framework denoted as W-DQN, founded upon the principles of deep reinforcement learning theory. Initially, the dynamic pricing problem for perishable goods is formulated as a Markov decision process. Subsequently, a dynamic pricing algorithm model and corresponding reward function are devised to bolster business revenue while curtailing food waste. Experimental findings unequivocally demonstrate that, relative to tabular-based dynamic pricing algorithm models, W-DQN attains commendable returns. Furthermore, the proposed reward function effectively reduces waste. In comparison to conventional pricing approaches, W-DQN significantly diminishes food waste, thereby enhancing overall business revenue.