Nowadays the trend of significant effort estimations is in demand. The software development industry requires an accurate software effort estimation to produce high-quality software that is delivered on time, on budget, and meets all the project’s requirements. However, missing data that is frequently encountered in the effort software estimation will influence the accuracy of the estimation. There are only a few studies that has been carried out that investigate the imputation of missing data in depth. Hence, this study proposed a modified MissForest Multiple Imputation (MFMI) approach, which integrates regression models for non-missing values within the MissForest algorithm. Five benchmark datasets including Cocomo, Dershanais, Albrecht, Kemerer, and ISBSG used in this study to evaluate the performance of the proposed model. The results show that MFMI outperforms original Random Forest and other existing imputation methods based on Root mean squared error (RMSE) and Mean Absolute Percentage Error (MAPE) performance measurements.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Data Quality with Miss Forest Imputation Techniques for Software Effort Estimation

  • Siti Hajar Binti Arbain,
  • Noorfa Haszlinna Mustaffa,
  • Dayang Norhayati Abang Jawawi,
  • Nor Azizah Ali

摘要

Nowadays the trend of significant effort estimations is in demand. The software development industry requires an accurate software effort estimation to produce high-quality software that is delivered on time, on budget, and meets all the project’s requirements. However, missing data that is frequently encountered in the effort software estimation will influence the accuracy of the estimation. There are only a few studies that has been carried out that investigate the imputation of missing data in depth. Hence, this study proposed a modified MissForest Multiple Imputation (MFMI) approach, which integrates regression models for non-missing values within the MissForest algorithm. Five benchmark datasets including Cocomo, Dershanais, Albrecht, Kemerer, and ISBSG used in this study to evaluate the performance of the proposed model. The results show that MFMI outperforms original Random Forest and other existing imputation methods based on Root mean squared error (RMSE) and Mean Absolute Percentage Error (MAPE) performance measurements.