<p>Software defect prediction (SDP) models are being developed to predict the post-release defects in a software module. This reduces the time and costs of testing the software module. To build an SDP model, it is essential to generate large volumes of defect data however, most literature on SDP focuses majorly on mitigating issues related to developing the supervised prediction models. With the rise in the development of large, complex software systems, a significant portion of data encountered is unlabelled. This emphasizes the need for efficient methods to assign accurate labels to extensive unlabelled datasets. For this objective, we propose a diverse ensemble of self-training semi-supervised learning frameworks that make use of bootstrapping and multi-inducer strategies to generate more accurate defect data for the SDP task. We conduct an empirical analysis on the 11 cross-software project data (that include a total of 40 released versions data) using various baselines. Experimental results show that our model significantly outperforms the recent state-of-the-art and baseline semi-supervised learning models in terms of the traditional and project-specific performance measures, indicating the use of a combination of various self-trainers in assigning the class labels to the unlabelled data points.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DEST: Diverse Ensemble of Self-Trainers for Software Defect Prediction

  • Bhutamapuram Umamaheswara Sharma,
  • Ravichandra Sadam,
  • Vinay Raj,
  • Sathish Jayabalan,
  • Sharan Krishnan,
  • Keerthana Saravanakumar,
  • Sanjana Maturi

摘要

Software defect prediction (SDP) models are being developed to predict the post-release defects in a software module. This reduces the time and costs of testing the software module. To build an SDP model, it is essential to generate large volumes of defect data however, most literature on SDP focuses majorly on mitigating issues related to developing the supervised prediction models. With the rise in the development of large, complex software systems, a significant portion of data encountered is unlabelled. This emphasizes the need for efficient methods to assign accurate labels to extensive unlabelled datasets. For this objective, we propose a diverse ensemble of self-training semi-supervised learning frameworks that make use of bootstrapping and multi-inducer strategies to generate more accurate defect data for the SDP task. We conduct an empirical analysis on the 11 cross-software project data (that include a total of 40 released versions data) using various baselines. Experimental results show that our model significantly outperforms the recent state-of-the-art and baseline semi-supervised learning models in terms of the traditional and project-specific performance measures, indicating the use of a combination of various self-trainers in assigning the class labels to the unlabelled data points.