Investigating Sophisticated Machine Learning Paradigms for the Prognostication of Diabetes
摘要
The primary objective of this research is prediction of the probable presence of diabetes in people based on their health factors. In this paper, a dataset consisting of 100,000 records with 9 data points is used for conducting experiment. Pearson Correlation coefficient is calculated to check if input variables have any effect on output variable. Standard scalar is used to compute z-score of sample statistics. One-hot encoding is used for categorical data that has more than two values. Synthetic Minority Over-sampling Technique with the sampling strategy of 0.1 and Random Under Sampler with a sampling strategy of 0.5 is used to balance class. To find best hyper parameters for the algorithm, the execution of machine learning algorithms is done using Grid Search and cross validation class. Logistic Regression Classifier, K Nearest Neighbor Classifier, Support Vector Machine Classifier, Decision Tree Classifier, and Random Forest Classifier are used for prognostication of diabetes.