Optimizing Early Diabetes Detection Through Machine Learning: A Comparative Analysis of Classification Models
摘要
Diabetes continues to pose a significant global health challenge, necessitating innovations in its early diagnosis. This study uses machine learning to improve the prediction of diabetes by analyzing critical health indicators, such as glucose levels, BMI, and family medical history. Using the Pima Indian Diabetes Dataset, the research evaluated multiple classification algorithms, including Support Vector Machine (SVM), Random Forest, Decision Tree, Logistic Regression, and Gaussian Naive Bayes. Rigorous preprocessing methods ensured data quality through normalization, outlier removal, and missing value imputation. Model performance was assessed using metrics such as accuracy, precision, recall, and F1-score, with SVM, Random Forest, and Logistic Regression demonstrating the most balanced results. The findings underscore the transformative potential of machine learning in predictive healthcare, which offers pathways for timely interventions and personalized care. Future directions include expanding the diversity of the dataset and integrating real-time data for improved diagnostic accuracy and applicability in clinical settings.