Comparative Study of Supervised Machine Learning Algorithms for Predicting Oversampled Imbalanced Medical Data
摘要
This study addresses the challenge of class imbalance in medical datasets, where negative cases significantly outnumber positive cases, hindering accurate health status predictions based on clinical characteristics. A comprehensive empirical comparison of supervised machine learning algorithms is conducted, incorporating models such as Logistic Regression, K-Nearest Neighbors, Naive Bayes, Decision Tree, Support Vector Machine, as well as Ensemble Learning models like Random Forest, Adaptive Boosting, and eXtreme Gradient Boosting. To mitigate the class imbalance issue, synthetic oversampling techniques such as Random Oversampling, Synthetic Minority Oversampling Technique (SMOTE), Adaptive Synthetic Sampling, and Borderline-SMOTE are employed. The study aims to enhance prediction performance, evaluating models based on metrics including Balanced Accuracy, Area Under the Receiver Operating Characteristic Curve, F1-score, and Matthews Correlation Coefficient. This research contributes valuable insights into algorithmic effectiveness and the impact of oversampling techniques, guiding the selection of robust models for improved health data classification.