Investigating Role of SVM, Decision Tree, KNN, ANN in Classification of Diabetic Patient Dataset
摘要
Diabetes, which is a long-term ailment, is characterized primarily by high levels of sugar in the blood. It has been connected to a broad range of different types of complicated disorders, such as heart attack, renal failure, and stroke, among others. Almost 422 million people throughout the world were diagnosed with diabetes in 2014, and according to IDF Atlas 2021 report, 10.5% of the adult population (20–79 years) has diabetes, by 2045, IDF projections show that 1 in 8 adults approx. 783 million will be living with that disease an increment to 46% making it the most common metabolic. Logistic regression was used in traditional research to determine the characteristics that increase a person’s likelihood of developing diabetes based on probability value and odds ratio. The authors utilize many classifiers to make predictions about diabetes patients, including NB, DT, AB, and RF. Twenty separate tests were conducted, each using one of three partitioning strategies. These classifier's effectiveness is measured by their accuracy and area under the curve. The overall accuracy rate of ML systems was 90.62% with conventional research. The K10 procedure combined LR-based feature selection with an RF-based classifier to obtain an ACC of 94.25% and an AUC of 0.95. The major goal of this work is to compare the performance of SVM, Decision Tree, KNN, & ANN on a dataset of diabetes patient classifications. The study’s goals include improved accuracy and trustworthy results.