Adaptive queue management in healthcare using supervised Q-learning with time-varying reward and cost structures
摘要
Efficient patient management in hospitals requires adaptive decision-making under time-varying demand and dynamic service environments. This study proposes a heterogeneous medical patient queueing model that integrates reinforcement learning with stochastic queue dynamics to minimize overall patient waiting time. The model distinguishes between two categories of service providers (SPs): those attending first-time patients and those serving returning patients. Each category may differ in service rate but not in medical specialty. Patient arrivals follow a non-homogeneous Poisson process (NHPP) to capture realistic time-dependent flow variations. A Q-learning framework with a supervised