Exploring Advantage Estimation with Higher-Order Summation of State Values in Deep Reinforcement Learning: Experimental Insights for Online Temporal Difference Approach
摘要
This study investigates the potential advantages of redefining advantage calculation in actor-critic deep reinforcement learning. Departing from the conventional temporal difference approach, where the advantage is determined by the difference between the values of two consecutive states, we investigate an alternative method. Our approach involves computing the advantage as the difference between the value of the current state and the summation of multiple discounted values of subsequent states. Through a series of experiments, we explore the implications of this alternative advantage calculation on the training dynamics of actor-critic networks. Furthermore, we extend the actor’s loss function by incorporating the summation of even multiple discounted rewards with appropriate gamma discounts. The experiments aim to discern whether this extended loss function enhances the learning capabilities of the actor in complex reinforcement learning tasks. We find that the advantages are contingent upon the task’s reward horizon—indicating the frequency of rewards and the presence of negative rewards. Notably, the effectiveness of extended discounting is highly dependent on selecting an appropriate order of calculation, signifying the number of state values considered after initial. This order, akin to context, proves crucial in determining the efficacy of the extended discounting approach.