Diabetes continues to pose a significant global health challenge, necessitating innovations in its early diagnosis. This study uses machine learning to improve the prediction of diabetes by analyzing critical health indicators, such as glucose levels, BMI, and family medical history. Using the Pima Indian Diabetes Dataset, the research evaluated multiple classification algorithms, including Support Vector Machine (SVM), Random Forest, Decision Tree, Logistic Regression, and Gaussian Naive Bayes. Rigorous preprocessing methods ensured data quality through normalization, outlier removal, and missing value imputation. Model performance was assessed using metrics such as accuracy, precision, recall, and F1-score, with SVM, Random Forest, and Logistic Regression demonstrating the most balanced results. The findings underscore the transformative potential of machine learning in predictive healthcare, which offers pathways for timely interventions and personalized care. Future directions include expanding the diversity of the dataset and integrating real-time data for improved diagnostic accuracy and applicability in clinical settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing Early Diabetes Detection Through Machine Learning: A Comparative Analysis of Classification Models

  • Nandini Das,
  • Jhalak Dutta,
  • Atyasha Bhattacharyya,
  • Swati Raj

摘要

Diabetes continues to pose a significant global health challenge, necessitating innovations in its early diagnosis. This study uses machine learning to improve the prediction of diabetes by analyzing critical health indicators, such as glucose levels, BMI, and family medical history. Using the Pima Indian Diabetes Dataset, the research evaluated multiple classification algorithms, including Support Vector Machine (SVM), Random Forest, Decision Tree, Logistic Regression, and Gaussian Naive Bayes. Rigorous preprocessing methods ensured data quality through normalization, outlier removal, and missing value imputation. Model performance was assessed using metrics such as accuracy, precision, recall, and F1-score, with SVM, Random Forest, and Logistic Regression demonstrating the most balanced results. The findings underscore the transformative potential of machine learning in predictive healthcare, which offers pathways for timely interventions and personalized care. Future directions include expanding the diversity of the dataset and integrating real-time data for improved diagnostic accuracy and applicability in clinical settings.