<p>Software Defect Prediction (SDP) is an essential step in the software engineering community over the past few decades. SDP helps to deliver high-quality software products and reduce maintenance of software based on risk by identifying defective software. Various automatic prediction tools are employed to predict defects in software source codes. However, the existing SDP techniques have encountered difficulties in predicting data instances from an imbalanced database. Thus, an automatic ensemble-based feature selection technique with Ensemble_Machine Learning (Ensemble_ML) is introduced for SDP. Here, data transformation is executed using Yeo-Jhonson transformation, and the selection of relevant features is executed by subjecting the transformed data to the ensemble-based method. Following this, the synthetic software data points are generated by augmenting the selected data features using the Synthetic minority oversampling technique (SMOTE). After that, the SDP is performed using the designed Ensemble_ML technique, and majority voting is performed to determine the final predicted output. Moreover, the results show that Ensemble_ML outperforms other competing models with a Magnitude of relative error<b> (</b>MMRE) of 0.070, Root Mean Squared Error (RMSE) of 0.193, Mean absolute percentage error (MAPE) of 0.075, and accuracy of 0.965. Here the major limitation of the model is the increase in computational complexity as the feature size grows, which limits the scalability of the model. The high-performance results indicate that the model can be integrated with broader applications in software engineering, delivering more transparent and reliable decision support to the industry.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ensemble-based feature selection and machine learning models for software defect prediction

  • Gaurav Kishor Kanaujiya,
  • Prabhat Verma

摘要

Software Defect Prediction (SDP) is an essential step in the software engineering community over the past few decades. SDP helps to deliver high-quality software products and reduce maintenance of software based on risk by identifying defective software. Various automatic prediction tools are employed to predict defects in software source codes. However, the existing SDP techniques have encountered difficulties in predicting data instances from an imbalanced database. Thus, an automatic ensemble-based feature selection technique with Ensemble_Machine Learning (Ensemble_ML) is introduced for SDP. Here, data transformation is executed using Yeo-Jhonson transformation, and the selection of relevant features is executed by subjecting the transformed data to the ensemble-based method. Following this, the synthetic software data points are generated by augmenting the selected data features using the Synthetic minority oversampling technique (SMOTE). After that, the SDP is performed using the designed Ensemble_ML technique, and majority voting is performed to determine the final predicted output. Moreover, the results show that Ensemble_ML outperforms other competing models with a Magnitude of relative error (MMRE) of 0.070, Root Mean Squared Error (RMSE) of 0.193, Mean absolute percentage error (MAPE) of 0.075, and accuracy of 0.965. Here the major limitation of the model is the increase in computational complexity as the feature size grows, which limits the scalability of the model. The high-performance results indicate that the model can be integrated with broader applications in software engineering, delivering more transparent and reliable decision support to the industry.