Application of Data Mining Methods and Techniques for Diabetes Prediction Using Pima Indian Dataset
摘要
Diabetes is a health condition that only occurs when the body either does not use insulin effectively or produces an insufficient amount of insulin from the pancreas. Insulin, a hormone that regulates blood sugar levels, is crucial. Untreated diabetes can lead to hyperglycaemia, a condition characterized by elevated levels of glucose in the bloodstream. This can have a profound impact on several physiological systems such as blood vessels and neurons. Computer vision is a crucial component in the realm of human health as it offers precise outcomes and eliminates any human biases. This research aims to improve diabetes classification. The study focused on analysing how data mining and ML techniques compare to traditional algorithms in predicting diabetes accurately. The “Pima Indians Diabetes Dataset” Standard was used, and feature selection was conducted to improve its potential. To assess the performance of the models, Naive Bayes, SVM, and KNN techniques were utilized, and evaluation criteria such as precision, recall, specificity, and mean absolute error were taken into account. Cross-validation was used to enhance accuracy across all models. The results indicated that the SVM method had an accuracy rate of 78%, making it the best classifier for the sample dataset. Thus, this analysis could help predict diabetes more effectively.