This paper presents a semi-supervised adaptive ensemble model designed to improve predictive performance in scenarios with limited labeled data. By integrating RandomForest, XGBoost, and an Attention-based Multi-Layer Perceptron (AttentionMLP), the model leverages both labeled and unlabeled data, using only 50% of the available labeled data alongside unlabeled data through an iterative pseudo-labeling process and an adaptive weighting scheme. The AttentionMLP incorporates a sample-wise attention mechanism to prioritize informative samples, enhancing robustness. The model’s performance is evaluated on three diabetes classification datasets: BRFSS2015, Pima Indian, and Diabetes Diagnosis. Results demonstrate that the proposed model achieves superior Area Under the Curve (AUC), F1 Score, and Accuracy on the Pima Indian and Diabetes Diagnosis datasets, with AUC improvements of up to 12.4% over baseline models such as LSTM, GRU, and BiLSTM. On the BRFSS2015 dataset, the model performs competitively, highlighting its effectiveness across diverse data distributions. The findings suggest that the ensemble’s combination of traditional and deep learning methods, augmented by attention and pseudo-labeling with limited labeled data, offers a powerful approach for classification tasks in data-scarce environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention-Driven Ensemble Learning: Enhancing Diabetes Prediction in Data-Scarce Environments

  • Md Shahadat Kabir,
  • Usman Gani Joy,
  • Tanvir Azhar

摘要

This paper presents a semi-supervised adaptive ensemble model designed to improve predictive performance in scenarios with limited labeled data. By integrating RandomForest, XGBoost, and an Attention-based Multi-Layer Perceptron (AttentionMLP), the model leverages both labeled and unlabeled data, using only 50% of the available labeled data alongside unlabeled data through an iterative pseudo-labeling process and an adaptive weighting scheme. The AttentionMLP incorporates a sample-wise attention mechanism to prioritize informative samples, enhancing robustness. The model’s performance is evaluated on three diabetes classification datasets: BRFSS2015, Pima Indian, and Diabetes Diagnosis. Results demonstrate that the proposed model achieves superior Area Under the Curve (AUC), F1 Score, and Accuracy on the Pima Indian and Diabetes Diagnosis datasets, with AUC improvements of up to 12.4% over baseline models such as LSTM, GRU, and BiLSTM. On the BRFSS2015 dataset, the model performs competitively, highlighting its effectiveness across diverse data distributions. The findings suggest that the ensemble’s combination of traditional and deep learning methods, augmented by attention and pseudo-labeling with limited labeled data, offers a powerful approach for classification tasks in data-scarce environments.