Credit risk is one of the primary risks that banks face, deserving special attention. Previously, modelling credit risk data using parametric models required significant labour, which was time-consuming. Despite the dominance of machine learning (ML) models, deep learning (DL) models for tabular data have emerged to address their drawbacks, including interpretability issues. We seek to determine whether the TabNet model is worth paying the price of its sophisticated computation and interpretability abilities. We used 37991 Italian manufacturing companies to determine their default likelihood. We adopted the Boruta method and sequential attention mechanism for feature selection, SMOTEENN for data balance, and SHAP values to quantify features’ contribution toward model output. A comparative analysis revealed that XGBoost remains a state-of-the-art model in balanced and imbalanced data cases. Thus, leveraging XGBoost can assist lenders in predicting and classifying potential defaulters. Data limitations and feature exclusions set the stage for further exploration of TabNet’s performance in default prediction tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Learning for Tabular Data: Application to Credit Risk Modeling

  • Steven Mphaya,
  • Marialuisa Restaino,
  • Michele La Rocca

摘要

Credit risk is one of the primary risks that banks face, deserving special attention. Previously, modelling credit risk data using parametric models required significant labour, which was time-consuming. Despite the dominance of machine learning (ML) models, deep learning (DL) models for tabular data have emerged to address their drawbacks, including interpretability issues. We seek to determine whether the TabNet model is worth paying the price of its sophisticated computation and interpretability abilities. We used 37991 Italian manufacturing companies to determine their default likelihood. We adopted the Boruta method and sequential attention mechanism for feature selection, SMOTEENN for data balance, and SHAP values to quantify features’ contribution toward model output. A comparative analysis revealed that XGBoost remains a state-of-the-art model in balanced and imbalanced data cases. Thus, leveraging XGBoost can assist lenders in predicting and classifying potential defaulters. Data limitations and feature exclusions set the stage for further exploration of TabNet’s performance in default prediction tasks.