Examining Policy Entropy of Reinforcement Learning Agents for Personalization Tasks
摘要
This paper examines the behavior of reinforcement learning systems in personalization environments and details the differences in policy entropy associated with the type of learning algorithm utilized. We observe that as agents evolve towards the optimal policy, the trajectory of the learned policy is intricately linked to the chosen learning paradigm. Through a series of numerical experiments, we consistently observe differences in policy entropy values between Policy Optimization and Q-Learning agents during the training process. Our empirical findings are complimented by a theoretical analysis that sheds light on this phenomenon, which has not yet been explored in existing literature.