Differentiable Good Arm Identification
摘要
This paper focuses on a variant of the stochastic multi-armed bandit problem known as good arm identification (GAI). GAI is a pure-exploration bandit problem that aims to identify and output as many good arms as possible using the fewest number of samples. A good arm is defined as an arm whose expected reward is greater than a given threshold. We present our study in the context of a structured bandit setting and introduce DGAI, a novel, differentiable good arm identification algorithm. By leveraging a data-driven approach, DGAI significantly enhances the state-of-the-art HDoC algorithm empirically. Additionally, we demonstrate that DGAI can improve the cumulative reward maximization problem when a threshold is provided as prior knowledge for the arm set. Extensive experiments have been conducted to validate our algorithm’s performance. The results demonstrate that our algorithm significantly outperforms the competitors in both synthetic and real-world datasets for both the GAI and MAB tasks.