Credit business is one of the primary operations in banking, and the control of credit default risk is critically important. A major challenge in the analysis of existing credit data is the highly imbalanced distribution between default and non-default cases. This imbalance can lead to biases in risk assessment. To address this issue, this study proposes an innovative method that self-attention CTGAN with an EasyEnsemble (SACTGAN-EE) mixed sampling method to handle data imbalance. This method boosts the data capturing capability of CTGAN through self-attention, enabling CTGAN to generate synthetic samples that are closer to the actual data distribution and significantly improving sample diversity and authenticity. Additionally, the use of EasyEnsemble technology integrates multiple data subsets to effectively balance majority and minority classes, thus reducing the bias caused by class imbalance. The final experimental results show that this integrated approach significantly outperforms traditional methods in dealing with the imbalance of credit data. It not only generates high-quality balanced datasets but also improves model capability in recognizing minority classes, effectively reducing overfitting risk.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SACTGAN-EE Imbalanced Data Processing Method for Credit Default Prediction

  • Shuxian Liu,
  • Guoqiang Wang,
  • Zhida Liu

摘要

Credit business is one of the primary operations in banking, and the control of credit default risk is critically important. A major challenge in the analysis of existing credit data is the highly imbalanced distribution between default and non-default cases. This imbalance can lead to biases in risk assessment. To address this issue, this study proposes an innovative method that self-attention CTGAN with an EasyEnsemble (SACTGAN-EE) mixed sampling method to handle data imbalance. This method boosts the data capturing capability of CTGAN through self-attention, enabling CTGAN to generate synthetic samples that are closer to the actual data distribution and significantly improving sample diversity and authenticity. Additionally, the use of EasyEnsemble technology integrates multiple data subsets to effectively balance majority and minority classes, thus reducing the bias caused by class imbalance. The final experimental results show that this integrated approach significantly outperforms traditional methods in dealing with the imbalance of credit data. It not only generates high-quality balanced datasets but also improves model capability in recognizing minority classes, effectively reducing overfitting risk.