The airline industry’s New Distribution Capability has spurred interest in continuous dynamic pricing, where fares aren’t restricted to fixed price points. Forecasting demand across a continuous price range is difficult, and no optimal solution algorithm exists for general demand functions. This study employs the Soft Actor-Critic (SAC) algorithm for continuous action spaces in reinforcement learning to tackle these challenges in continuous dynamic pricing. In the context of dynamic pricing, we found that due to its inherent model limitations, the Soft Actor-Critic algorithm requires a substantial number of training iterations to address the issue of local convergence to the boundary when learning optimal actions near the boundary. Therefore, utilizing the structure of the existing approximate solutions, we propose the Reward Shaping Soft Actor-Critic (RSAC) algorithm, which is based on a penalty for low prices. The experimental results indicate that RSAC significantly improved the convergence speed and performance of the SAC algorithm.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reinforcement Learning for Airline Continuous Dynamic Pricing

  • Zhicheng Yao,
  • Wenguo Yang

摘要

The airline industry’s New Distribution Capability has spurred interest in continuous dynamic pricing, where fares aren’t restricted to fixed price points. Forecasting demand across a continuous price range is difficult, and no optimal solution algorithm exists for general demand functions. This study employs the Soft Actor-Critic (SAC) algorithm for continuous action spaces in reinforcement learning to tackle these challenges in continuous dynamic pricing. In the context of dynamic pricing, we found that due to its inherent model limitations, the Soft Actor-Critic algorithm requires a substantial number of training iterations to address the issue of local convergence to the boundary when learning optimal actions near the boundary. Therefore, utilizing the structure of the existing approximate solutions, we propose the Reward Shaping Soft Actor-Critic (RSAC) algorithm, which is based on a penalty for low prices. The experimental results indicate that RSAC significantly improved the convergence speed and performance of the SAC algorithm.