Ensemble Machine Learning Models for Corporate Credit Risk Prediction: A Comparative Study
摘要
The objective of the study is to assess the predictive capabilities of four machine learning models such as XGBoost, CATBoost, LightGBM, and LSTM to identify the significant variables that determine whether a company is default or non-default in terms of credit risk. The study also applied two feature engineering techniques such as principal component analysis (PCA) and random forest to extract the important features and then applied the machine learning models. The analysis used a dataset from the Prowess database, which consists information on companies that are rated by the ACUITE and BRICKWORK credit rating agencies. The dataset comprised of 212 non-default-rated companies and 198 default-rated companies from 2016 to 2024. The study utilized 27 explanatory variables. The analysis showed that ensemble models, particularly the XGBoost, CATBoost, and LightGBM showed higher performance when the model was created using the features selected by the random forest algorithm compared to the models created with components created by PCA. However, LSTM showed higher performance with the components reduced using PCA when compared to the model created using feature selection with random forest. These models help both banking and non-banking financial institutions to make lending decisions, based on the borrower firm’s creditworthiness. This research also helps in the selection of most important features which would help to identify additional insights to assess a firm’s financial performance.