Understanding Bankruptcy Prediction Using Data Mining Algorithms—Evidence from Taiwan’s Economy
摘要
Bankrupt companies are entitled to pay back creditors a fraction of their debts during business failure resulting in substantial monetary losses for customers. Thus, companies seek bankruptcy protection to reassess capital and develop plans to support creditors. Hence, predicting bankruptcy is critical for stakeholders (e.g., banking industry, creditors). Researchers have predicted bankruptcy based on data (e.g., financial, non-financial, economic, and market variables) and have adopted a range of models (e.g., statistical, financial, machine learning, and deep learning). However, its full potential of selecting a subset of variables and using selected machine learning models on real datasets is yet to be explored. Here, we collected bankruptcy data from 1999 to 2009 on Taiwan’s economy. Here, we employ four methods (e.g., variance threshold, recursive feature elimination, logistic stepwise, and consideration of economic factors) to shortlist variables. Further, we deploy four ML models (e.g., logistic regression, random forest (RF), gradient boosting, and XGBoost). Key findings indicated that XGBoost (97% accuracy) and RF (96% accuracy) perform best when the SMOTE algorithm addresses data imbalance. As per the RF classifier, variables including net income to stockholder’s equity, net value growth rate, borrowing dependency, and persistent EPS in the last four seasons ranked higher for bankruptcy prediction. Our model is easy to adopt and performs better to present bankrupt and non-bankrupt firms’ discrepancies. The results discussed here should be useful to decision-makers (e.g., bankers, policy analysts, and economic and financial agencies) across the world and they would benefit by launching preventative measures to avoid financial loss.