SACTGAN-EE Imbalanced Data Processing Method for Credit Default Prediction
摘要
Credit business is one of the primary operations in banking, and the control of credit default risk is critically important. A major challenge in the analysis of existing credit data is the highly imbalanced distribution between default and non-default cases. This imbalance can lead to biases in risk assessment. To address this issue, this study proposes an innovative method that self-attention CTGAN with an EasyEnsemble (SACTGAN-EE) mixed sampling method to handle data imbalance. This method boosts the data capturing capability of CTGAN through self-attention, enabling CTGAN to generate synthetic samples that are closer to the actual data distribution and significantly improving sample diversity and authenticity. Additionally, the use of EasyEnsemble technology integrates multiple data subsets to effectively balance majority and minority classes, thus reducing the bias caused by class imbalance. The final experimental results show that this integrated approach significantly outperforms traditional methods in dealing with the imbalance of credit data. It not only generates high-quality balanced datasets but also improves model capability in recognizing minority classes, effectively reducing overfitting risk.