<p>Accurate environmental modeling is essential for model-based offline policy optimization. Traditional empirical risk minimization can yield biased models if data collection is subject to selection bias, potentially misguiding policy optimization. This issue is especially pertinent in real-world decision-making scenarios where data collection often depends on optimized, non-random behavior policies. This paper addresses such a practical challenge of offline policy optimization under selection bias, with a focus on delivery incentive policies for food delivery platforms. We propose a novel framework for offline optimization of these policies, based on a de-biased environmental model. Initially, the framework learns a de-biased order acceptance rate and delivery time prediction model from historical data through adversarial weighted empirical risk minimization, constituting the environment model. Subsequently, it employs operation research solvers to derive historic best actions based on the learned de-biased environment model, determining the optimal bonus amount and reasonable incentive time limit for each order under budget constraints. Finally, a policy neural network is trained to map environmental states to these optimized actions, enabling efficient and executable policies for real-time decision-making. To verify the effectiveness and efficiency of our framework, both offline experiments on a real-world dataset and online A/B tests on the Meituan food delivery platform are conducted. Results demonstrate that our framework outperforms baseline methods in both model accuracy and policy optimization performance in offline experiments and realizes a 9% reduction in the customer complaint rate in reality.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning de-biased environment models for delivery incentive policy optimization on food delivery platforms

  • Yu-Ren Liu,
  • Xiong-Hui Chen,
  • Siyuan Xiao,
  • Xinyu Yang,
  • Xintong Qi,
  • Linjun Zhou,
  • Yang Yu,
  • Fangsheng Huang

摘要

Accurate environmental modeling is essential for model-based offline policy optimization. Traditional empirical risk minimization can yield biased models if data collection is subject to selection bias, potentially misguiding policy optimization. This issue is especially pertinent in real-world decision-making scenarios where data collection often depends on optimized, non-random behavior policies. This paper addresses such a practical challenge of offline policy optimization under selection bias, with a focus on delivery incentive policies for food delivery platforms. We propose a novel framework for offline optimization of these policies, based on a de-biased environmental model. Initially, the framework learns a de-biased order acceptance rate and delivery time prediction model from historical data through adversarial weighted empirical risk minimization, constituting the environment model. Subsequently, it employs operation research solvers to derive historic best actions based on the learned de-biased environment model, determining the optimal bonus amount and reasonable incentive time limit for each order under budget constraints. Finally, a policy neural network is trained to map environmental states to these optimized actions, enabling efficient and executable policies for real-time decision-making. To verify the effectiveness and efficiency of our framework, both offline experiments on a real-world dataset and online A/B tests on the Meituan food delivery platform are conducted. Results demonstrate that our framework outperforms baseline methods in both model accuracy and policy optimization performance in offline experiments and realizes a 9% reduction in the customer complaint rate in reality.