This study addresses the challenge of class imbalance in medical datasets, where negative cases significantly outnumber positive cases, hindering accurate health status predictions based on clinical characteristics. A comprehensive empirical comparison of supervised machine learning algorithms is conducted, incorporating models such as Logistic Regression, K-Nearest Neighbors, Naive Bayes, Decision Tree, Support Vector Machine, as well as Ensemble Learning models like Random Forest, Adaptive Boosting, and eXtreme Gradient Boosting. To mitigate the class imbalance issue, synthetic oversampling techniques such as Random Oversampling, Synthetic Minority Oversampling Technique (SMOTE), Adaptive Synthetic Sampling, and Borderline-SMOTE are employed. The study aims to enhance prediction performance, evaluating models based on metrics including Balanced Accuracy, Area Under the Receiver Operating Characteristic Curve, F1-score, and Matthews Correlation Coefficient. This research contributes valuable insights into algorithmic effectiveness and the impact of oversampling techniques, guiding the selection of robust models for improved health data classification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Study of Supervised Machine Learning Algorithms for Predicting Oversampled Imbalanced Medical Data

  • Alvine Fandio,
  • O. Olawale Awe

摘要

This study addresses the challenge of class imbalance in medical datasets, where negative cases significantly outnumber positive cases, hindering accurate health status predictions based on clinical characteristics. A comprehensive empirical comparison of supervised machine learning algorithms is conducted, incorporating models such as Logistic Regression, K-Nearest Neighbors, Naive Bayes, Decision Tree, Support Vector Machine, as well as Ensemble Learning models like Random Forest, Adaptive Boosting, and eXtreme Gradient Boosting. To mitigate the class imbalance issue, synthetic oversampling techniques such as Random Oversampling, Synthetic Minority Oversampling Technique (SMOTE), Adaptive Synthetic Sampling, and Borderline-SMOTE are employed. The study aims to enhance prediction performance, evaluating models based on metrics including Balanced Accuracy, Area Under the Receiver Operating Characteristic Curve, F1-score, and Matthews Correlation Coefficient. This research contributes valuable insights into algorithmic effectiveness and the impact of oversampling techniques, guiding the selection of robust models for improved health data classification.