Since a large number of credit scoring models are built on a known set of data (training data) collected in the past or from other regions/domains, a prerequisite for applying these models to new instances (test data) is that the test data is comparable to the training data. The comparability between the test data and the training data also has a strong impact on the performance of credit scoring models. However, most studies have focused on the methods or algorithms for model construction, there is a lack of research on the impact of the comparability between the training data and the test data on the accuracy of credit scoring models. To fill this gap, we have used the time lag (difference in years) to represent the comparability between the training and test data collected from different years, and investigated how this time lag affects the accuracy of credit scoring models. This paper aims to extend our previous research by collecting a larger number of samples and performing a more detailed analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Investigation of the Impact of the Time Lag Between Training and Test Data Sets on the Accuracy of Credit Scoring Models

  • Yanwen Dong,
  • Noriki Ogura

摘要

Since a large number of credit scoring models are built on a known set of data (training data) collected in the past or from other regions/domains, a prerequisite for applying these models to new instances (test data) is that the test data is comparable to the training data. The comparability between the test data and the training data also has a strong impact on the performance of credit scoring models. However, most studies have focused on the methods or algorithms for model construction, there is a lack of research on the impact of the comparability between the training data and the test data on the accuracy of credit scoring models. To fill this gap, we have used the time lag (difference in years) to represent the comparability between the training and test data collected from different years, and investigated how this time lag affects the accuracy of credit scoring models. This paper aims to extend our previous research by collecting a larger number of samples and performing a more detailed analysis.