This paper presents a thorough evaluation of machine learning algorithms for assessing the risk of having diabetes with the help of the Pima Indian Diabetes dataset. In view of the global diabetes epidemic, timely and precise risk assessment is imperative. Our study involves an in-depth exploration of the data, uncovering a robust correlation between glucose levels and the likelihood of diabetes. We deploy a diverse set of nine machine learning models, encompassing logistic regression, decision trees, random forests, AdaBoost, support vector machine, K-nearest neighbors, Naive Bayes, XGBoost, and an artificial neural network approach. The results illuminate the strengths and weaknesses of each model, providing valuable insights for potential clinical applications. The logistic regression approach, in particular, showcases its ability to capture intricate patterns within the dataset, underscoring its effectiveness in diabetes risk assessment. In addition to traditional classifiers, we introduce an ensemble model that combines the strengths of five best performing classifiers. These findings not only improve the accuracy of diabetes risk assessment but also establish a benchmark for future research in medical diagnosis. Identifying the most effective model assists healthcare practitioners in early intervention and tailored treatment strategies. This research advances the field of healthcare analytics, facilitating more informed decisions in diabetes prevention and management. The integration of a robust ensemble model approach broadens the scope of potential applications, marking a significant contribution to the field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance Comparison of Different Machine Learning Classifiers for Diabetes Prediction

  • Dipayan Ghosh,
  • Abhik Ganguly,
  • Rounak Chakraborty,
  • Pawan Kumar Singh,
  • Aimin Li

摘要

This paper presents a thorough evaluation of machine learning algorithms for assessing the risk of having diabetes with the help of the Pima Indian Diabetes dataset. In view of the global diabetes epidemic, timely and precise risk assessment is imperative. Our study involves an in-depth exploration of the data, uncovering a robust correlation between glucose levels and the likelihood of diabetes. We deploy a diverse set of nine machine learning models, encompassing logistic regression, decision trees, random forests, AdaBoost, support vector machine, K-nearest neighbors, Naive Bayes, XGBoost, and an artificial neural network approach. The results illuminate the strengths and weaknesses of each model, providing valuable insights for potential clinical applications. The logistic regression approach, in particular, showcases its ability to capture intricate patterns within the dataset, underscoring its effectiveness in diabetes risk assessment. In addition to traditional classifiers, we introduce an ensemble model that combines the strengths of five best performing classifiers. These findings not only improve the accuracy of diabetes risk assessment but also establish a benchmark for future research in medical diagnosis. Identifying the most effective model assists healthcare practitioners in early intervention and tailored treatment strategies. This research advances the field of healthcare analytics, facilitating more informed decisions in diabetes prevention and management. The integration of a robust ensemble model approach broadens the scope of potential applications, marking a significant contribution to the field.