An Empirical Assessment of Credit Risk of Indian Companies Using Machine Learning Algorithms
摘要
The study aims to assess the predictive capabilities of six machine learning models such as logistic regression, decision tree, support vector machine, artificial neural networks, random forest, and Adaboost to categorize the significant variables influencing the companies’ likelihood to default. The analysis utilized a dataset from the Prowess database, containing information on companies rated by the ACUITE and BRICKWORK credit rating agencies. The dataset comprised 198 high-security-rated companies and 176 default-rated companies from 2016 to 2023. The study incorporates 27 explanatory variables. The analysis revealed that ensemble models, particularly the Adaboost classifier and random forest classifier, exhibit superior predictive performance compared to traditional models. The random forest classifier showed 97% precision, 97% recall, 97% F1 score, 99% area under the curve, and 98% Gini coefficient, and the AdaBoost algorithm showed 99% precision, 99% recall, 99% F1 score, 100% area under the curve and 99% Gini coefficient. The findings suggested that banks and other non-banking financial institutions can apply these models to refine their lending strategies based on the borrower firms’ credit quality. Moreover, this study extends beyond traditional financial metrics by identifying key default drivers, providing insights beyond conventional financial statement analysis.