Software Fault Prediction (SFP) processes find utility across various phases of the software development lifecycle, with a particularly valuable role in early identification of flawed or buggy components. Machine learning (ML) techniques are commonly harnessed in this domain. Predictive models rely on a set of 22 object-oriented metrics drawn from the CK and Martin metric collections, serving as input features. To evaluate the efficacy of these prediction models, experiments are conducted using ten publicly available datasets sourced from the PROMISE repository. Performance metrics, including Receiver Operating Characteristic (ROC) curve, accuracy (Acc), precision(Pre), F1-score(Fs), recall(Rc), and mean weighted error, are employed to gauge the predictive capabilities of the models. In this study, nine machine learning (ML) classifiers are employed to predict the accuracy rates of software fault proneness. We utilize ten of the most popular datasets in their latest versions from the PROMISE repository to evaluate the performance of algorithms, including Logistic Regression-LR, Decision Tree-DT, Support Vector Machine-SVM, Naïve Bayes-NB, Random Forest-RF, K-Nearest Neighbor-Knn, AdaBoost-AB, Gradient Boosting-GB, and Extreme Boosting-XB. The results highlight the superiority of the Random Forest over other examined algorithms in prediction accuracy. When using the RF classifiers, the achieved benchmarks include an Accuracy of 90%, Precision of 89%, Recall of 89.50%, F1-Score of 88%, and AUC-ROC of 75%. The comparative analysis indicates that the Random Forest delivers outstanding accuracy and reduced error rates. This research underscores the potential of the Random Forest, alongside other deep learning techniques, in advancing the field of software fault prediction precision.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fault Predictions Based on Base Learnings and Mean Weighted Score Using Machine Learning Techniques

  • Prachi Sasankar,
  • Gopal Sakarkar

摘要

Software Fault Prediction (SFP) processes find utility across various phases of the software development lifecycle, with a particularly valuable role in early identification of flawed or buggy components. Machine learning (ML) techniques are commonly harnessed in this domain. Predictive models rely on a set of 22 object-oriented metrics drawn from the CK and Martin metric collections, serving as input features. To evaluate the efficacy of these prediction models, experiments are conducted using ten publicly available datasets sourced from the PROMISE repository. Performance metrics, including Receiver Operating Characteristic (ROC) curve, accuracy (Acc), precision(Pre), F1-score(Fs), recall(Rc), and mean weighted error, are employed to gauge the predictive capabilities of the models. In this study, nine machine learning (ML) classifiers are employed to predict the accuracy rates of software fault proneness. We utilize ten of the most popular datasets in their latest versions from the PROMISE repository to evaluate the performance of algorithms, including Logistic Regression-LR, Decision Tree-DT, Support Vector Machine-SVM, Naïve Bayes-NB, Random Forest-RF, K-Nearest Neighbor-Knn, AdaBoost-AB, Gradient Boosting-GB, and Extreme Boosting-XB. The results highlight the superiority of the Random Forest over other examined algorithms in prediction accuracy. When using the RF classifiers, the achieved benchmarks include an Accuracy of 90%, Precision of 89%, Recall of 89.50%, F1-Score of 88%, and AUC-ROC of 75%. The comparative analysis indicates that the Random Forest delivers outstanding accuracy and reduced error rates. This research underscores the potential of the Random Forest, alongside other deep learning techniques, in advancing the field of software fault prediction precision.