Analyzing the Impact of Oversampling on Classifier Performance for Cardiac Disease Classification
摘要
Cardiac disease remains the primary cause of mortality on a global scale, highlighting the urgent requirement for precise and dependable classification methods in the identification and management of cardiac conditions. The imbalanced distribution of data within cardiac disease datasets, on the other hand, poses a significant challenge to the accuracy and effectiveness of classification models. Oversampling techniques have emerged in recent years as a promising approach to addressing this issue and improving classification performance. This study focuses on analyzing the impact of one such oversampling method, SMOTE, on the classification of cardiac diseases. We investigate the effects of combining SMOTE with commonly used classifiers such as SVM, KNN, DT, RF, and NB. Further, we use the Cleveland (DCL), Hungarian (DHN), and combined datasets (DCM) from the UCI repository, to conduct our experiments. The experimental results show that using SMOTE improves classification accuracy significantly. The average performance accuracy improvement of SVM is 6.65%, for KNN is 2.88%, for DT is 7.86%, for NB is 6.78%, and for RF is 15.08% across the datasets studied. These findings highlight SMOTE's efficacy in addressing the class imbalance challenge in cardiac disease datasets, resulting in improved classification accuracy.