An Application of Cross Validation Strategies for Prediction of Diabetes Mellitus
摘要
Diabetes is a prevalent chronic disease affecting the world today, making prevention and detection paramount. The complications associated with diabetes are significantly exacerbated if an individual predisposed to the condition is not accurately diagnosed and treated. The side effects encompass renal failure, visual impairment, premature myocardial infarctions, and limb amputation. We have analysed diabetes utilising machine learning techniques. A total of 768 entries from the Pima diabetes dataset were utilised. This paper examines the effects of Stratified K-fold cross-validation, Train-Test Split, and K-fold methods utilising the Random Forest classifier. The maximum and minimum accuracies across various folds, as per the Stratified K-Fold cross-validation results, are 86.84% and 58.97%, respectively. K fold cross-validation demonstrated maximum and minimum accuracy rates of 92.10% and 64.10%, respectively, across various folds.