Diabetes is a lifelong disease by which millions of people around the globe are affected and the number of patients is increasing annually. According to the International Diabetes Federation (IDF), 1 in 2 people with diabetes (240 million) are undiagnosed. It is predictable early on and can be treated more efficiently and effectively. Data mining algorithms are often used to predict diabetes in its early stages. In this study, we investigate widely used data mining methods for the aforementioned problem, and aim to find the most reliable and accurate method among them. Our results demonstrate a comparison between classification-based data mining methods KNN, Naïve-Bayes, decision trees, random forest, and Adaptive boosting (AdaBoost). Naive Bayes and Random Forest techniques provide the highest accuracies of 0.804 and 0.801, respectively, which can help clinicians make treatment decisions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Study on Classification Based-Data Mining Techniques in Early Diabetes Prediction

  • Yoshita Dahra,
  • Aman Jatain

摘要

Diabetes is a lifelong disease by which millions of people around the globe are affected and the number of patients is increasing annually. According to the International Diabetes Federation (IDF), 1 in 2 people with diabetes (240 million) are undiagnosed. It is predictable early on and can be treated more efficiently and effectively. Data mining algorithms are often used to predict diabetes in its early stages. In this study, we investigate widely used data mining methods for the aforementioned problem, and aim to find the most reliable and accurate method among them. Our results demonstrate a comparison between classification-based data mining methods KNN, Naïve-Bayes, decision trees, random forest, and Adaptive boosting (AdaBoost). Naive Bayes and Random Forest techniques provide the highest accuracies of 0.804 and 0.801, respectively, which can help clinicians make treatment decisions.