In the realm of Multi-Armed Bandits (MAB), the integration of habituation with intermittent breaks emerges as a compelling paradigm to enhance exploration-exploitation trade-offs. This study thoroughly investigates the application of habituation with breaks in two prominentstrategies: Softmax and Upper Confidence Bound (UCB). Empirical findings indicate that habituation with breaks can enhance long-term performance, mitigate reward stagnation, and support continuous adaptation. By elucidating the nuanced interplay between habituation, breaks, and MAB strategies, this work aims to inform future developments in decision-making algorithms designed for dynamic, real-world applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating Habituation Effects with UCB and Softmax Multi-Armed Bandit Algorithms for Optimized Digital Content Delivery

  • Kamil Bortko,
  • Kacper Fornalczyk,
  • Jarosław Jankowski

摘要

In the realm of Multi-Armed Bandits (MAB), the integration of habituation with intermittent breaks emerges as a compelling paradigm to enhance exploration-exploitation trade-offs. This study thoroughly investigates the application of habituation with breaks in two prominentstrategies: Softmax and Upper Confidence Bound (UCB). Empirical findings indicate that habituation with breaks can enhance long-term performance, mitigate reward stagnation, and support continuous adaptation. By elucidating the nuanced interplay between habituation, breaks, and MAB strategies, this work aims to inform future developments in decision-making algorithms designed for dynamic, real-world applications.