A Comparative Analysis of Machine Learning Algorithms for Early Prediction of Diabetes
摘要
Diabetes is a chronic disease with multifactorial etiologies such as age, obesity, and lack of exercise, complicating manual prediction. However, machine learning techniques have demonstrated high accuracy in diabetes prediction. This study evaluates the efficacy of several classification algorithms, including K-Nearest Neighbors (KNN), Decision Trees, Random Forests, Support Vector Machines (SVM), Neural Networks, Naive Bayes Classifiers, and Perceptron Learning Algorithms (PLA), in predicting diabetes. The research employs both brute force and ANNOY methods for KNN, genetic algorithms for SVM and Naive Bayes, and Neural Architecture Search (NAS) for optimizing neural network structures. The results indicate that KNN with brute force achieves the highest accuracy and stability, though at the cost of computational time. The ANNOY technique effectively reduces this time complexity. Neural networks, while highly accurate, show variability in performance. Naive Bayes classifiers modified by genetic algorithms exhibit competitive accuracy and stability comparable to Random Forests. These findings suggest that KNN, particularly with dataset normalization, is highly effective for early diabetes prediction.