<p>Accurate prediction in medical science is fraught with challenges. The presence of missing values, outliers, class imbalance, and suboptimal classifiers further complicates the task. While many researchers have addressed these issues, a comprehensive solution remains elusive. This research proposes a novel predictive model for diabetes, leveraging advanced techniques to enhance accuracy and robustness. The model incorporates IDBMI for missing value imputation, MFLOF for outlier detection and elimination, ASENN for class balancing, and a Multi-Model FusionNet Classifier to boost predictive capabilities. IDBMI imputed missing values by considering the density of surrounding data points, while MFLOF identified and removed outliers to prevent skewed results. The ASENN method balanced the classes to avoid bias towards one class, and the Multi-Model FusionNet Classifier combined multiple algorithms to enhance the accuracy, reliability and robustness of the predictions. The model’s efficacy was validated using two distinct diabetes datasets: NHANES and PIMA Indian Diabetes Dataset (PIDD). On the NHANES dataset, the model achieved an accuracy of 97.88%, precision of 0.976, recall of 0.982, F1-score of 0.979, and ROC of 0.979. Similarly, on PIDD, it attained an accuracy of 97.95%, precision of 0.971, recall of 0.988, F1-score of 0.980, and ROC of 0.971. The novel methods of IDBMI, MFLOF, ASENN, and Multi-Model FusionNet Classifier effectively addressed missing values, outliers, and class imbalance, making it a valuable tool for early diabetes detection. This research contributes novel solutions to address the challenges of diabetes detection, marking it as a noteworthy contribution to the health domain.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comprehensive Machine Learning Approach for Early Detection of Diabetes on Imbalanced Data with Missing and Outlier Values

  • Yogendra Singh,
  • Mahendra Tiwari

摘要

Accurate prediction in medical science is fraught with challenges. The presence of missing values, outliers, class imbalance, and suboptimal classifiers further complicates the task. While many researchers have addressed these issues, a comprehensive solution remains elusive. This research proposes a novel predictive model for diabetes, leveraging advanced techniques to enhance accuracy and robustness. The model incorporates IDBMI for missing value imputation, MFLOF for outlier detection and elimination, ASENN for class balancing, and a Multi-Model FusionNet Classifier to boost predictive capabilities. IDBMI imputed missing values by considering the density of surrounding data points, while MFLOF identified and removed outliers to prevent skewed results. The ASENN method balanced the classes to avoid bias towards one class, and the Multi-Model FusionNet Classifier combined multiple algorithms to enhance the accuracy, reliability and robustness of the predictions. The model’s efficacy was validated using two distinct diabetes datasets: NHANES and PIMA Indian Diabetes Dataset (PIDD). On the NHANES dataset, the model achieved an accuracy of 97.88%, precision of 0.976, recall of 0.982, F1-score of 0.979, and ROC of 0.979. Similarly, on PIDD, it attained an accuracy of 97.95%, precision of 0.971, recall of 0.988, F1-score of 0.980, and ROC of 0.971. The novel methods of IDBMI, MFLOF, ASENN, and Multi-Model FusionNet Classifier effectively addressed missing values, outliers, and class imbalance, making it a valuable tool for early diabetes detection. This research contributes novel solutions to address the challenges of diabetes detection, marking it as a noteworthy contribution to the health domain.