An Examination of the Effectiveness of SMOTE-Based Algorithms on Software Defect Prediction
摘要
Software engineering involves the design, development, testing, and maintenance of software products. The testing stage is crucial for retaining the reliability of said software product. Detection, verification, and resolution of software bugs that might have been generated at any point during the development life cycle is of critical significance and toughness. However, software bugs can be rare. Among the numerous software modules to be tested, only a handful might turn out to be buggy. This gives rise to a highly imbalanced dataset. To combat this, the usage of the Synthetic Minority Over-sampling Technique (SMOTE) concept from Machine Learning (ML) has seen a significant rise in the prediction of software bugs. ML, though an excellent prediction tool, has to be maneuvered carefully to handle skewed and imbalanced data to get optimum results. In this paper, a comparative analysis has been conducted, involving 87 existing variants of SMOTE to determine a variant that performs optimally across specific evaluation metrics and a wide range of classifiers. This comparative analysis was carried out on six different open-source software datasets, which consisted of only 8.9% of minority data on average, using K-nearest Neighbours (KNN) and RandomForest (RF) Classifiers. Through acquired results, the SMOTE_Cosine variant was determined to be the best-performing variant, appearing in the top quartile of the performing variants for all used datasets. The variants were taken through two rounds of selection, the first one involving F1 scores and frequency of appearance in the top quartiles of the list when ranked for F1 score. The second round took the Recall score of the classifier and variant combination into account and ranked them accordingly. The variant, when combined with RF Classifier, had an average F1 score of 50% and the highest F1 score achieved was 69.9%. During the second round of selection, the SMOTE_Cosine variant acquired the highest median Recall score of 61.81%. The selected SMOTE variant was observed to show impressive performance in terms of F1 score, Recall, Precision, and False Negative Rate (FNR), on the datasets.