Existing Automated machine learning (AutoML) systems have achieved considerable success in offline machine learning. Nonetheless, they are not applicable to data streams that requires real-time training and prediction, not to mention class imbalanced data streams. This paper proposes an Online Automated framework designed for imbalanced data streams learning, called OAutoIDSL. Firstly, we adopt and improve Thompson Sampling (TS) with an imbalanced reward design for combined algorithm selection and hyperparameter tuning (CASH) that enables online learning and optimisation. Secondly, we introduce two mechanisms to further tackle data non-stationarity and class imbalance – discounting outdated knowledge and adaptive weight tuning in imbalanced rewards. The effectiveness of our approach is demonstrated through the empirical evaluation on a set of synthetic imbalanced data streams encompassing various stationary and non-stationary scenarios, along with four real-world data sets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Online Automated Imbalanced Learning via Adaptive Thompson Sampling

  • Zhaoyang Wang,
  • Shuo Wang

摘要

Existing Automated machine learning (AutoML) systems have achieved considerable success in offline machine learning. Nonetheless, they are not applicable to data streams that requires real-time training and prediction, not to mention class imbalanced data streams. This paper proposes an Online Automated framework designed for imbalanced data streams learning, called OAutoIDSL. Firstly, we adopt and improve Thompson Sampling (TS) with an imbalanced reward design for combined algorithm selection and hyperparameter tuning (CASH) that enables online learning and optimisation. Secondly, we introduce two mechanisms to further tackle data non-stationarity and class imbalance – discounting outdated knowledge and adaptive weight tuning in imbalanced rewards. The effectiveness of our approach is demonstrated through the empirical evaluation on a set of synthetic imbalanced data streams encompassing various stationary and non-stationary scenarios, along with four real-world data sets.