A controlling estimation bias method: Max_Mix_Min estimator for Q-learning
摘要
Although Q-learning (QL) is widely used in reinforcement learning, it suffers from overestimation bias, which can lead to poor performance in stochastic environments due to its susceptibility to maximization bias. To address this problem, various bias correction mechanisms have been proposed. However, while these mechanisms may reduce overestimation bias, some of them introduce underestimation bias, which is undesirable in some environments. To leverage both overestimation and underestimation biases, we introduce an underestimation mechanism called the min estimator, followed by our proposed Max_Mix_Min Q-learning (M3QL) method, which incorporates a balance parameter