The class imbalance issue is evident in medical datasets, posing a hurdle for accurate predictive modeling. Maintaining the natural characteristics of medical data while mitigating this issue is paramount in ensuring the success of clinical decision support systems. Therefore, this study proposes a genetic algorithm-based data selection method (GA-DS) for imbalanced medical data reserving its distribution and characteristics. It optimizes the assignment of samples into train and test sets by improving the recognition of rare samples on test data. We presented three versions of the GA-DS, one unconstrained and two constrained (train-set \(>=\) 50%, train-set = 70%) to ensure the reliability of the results. Evaluation of the GA-DS methods on three commonly used imbalanced medical datasets and comparison with SMOTE, random oversampling, and stratified random sampling demonstrated the superior performance of GA-DS by improving the sensitivity score on test data. Future work will investigate its performance on large and highly imbalanced data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Proposing a Genetic Algorithms-Based Data Selection Method for Imbalanced Medical Datasets

  • Mabrouka Salmi,
  • Dalia Atif,
  • Ajith Abraham,
  • Sebastian Ventura

摘要

The class imbalance issue is evident in medical datasets, posing a hurdle for accurate predictive modeling. Maintaining the natural characteristics of medical data while mitigating this issue is paramount in ensuring the success of clinical decision support systems. Therefore, this study proposes a genetic algorithm-based data selection method (GA-DS) for imbalanced medical data reserving its distribution and characteristics. It optimizes the assignment of samples into train and test sets by improving the recognition of rare samples on test data. We presented three versions of the GA-DS, one unconstrained and two constrained (train-set \(>=\) 50%, train-set = 70%) to ensure the reliability of the results. Evaluation of the GA-DS methods on three commonly used imbalanced medical datasets and comparison with SMOTE, random oversampling, and stratified random sampling demonstrated the superior performance of GA-DS by improving the sensitivity score on test data. Future work will investigate its performance on large and highly imbalanced data.