<p>In this paper, we investigate using machine learning models to predict credit defaults using financial ratios. We compare the performance and interpretability of two ensemble learning algorithms: Random Forest and XGBoost. To improve the models' capability to detect defaults, we exploit the inherent class imbalance of default prediction tasks with the ROSE (Random Over-Sampling Examples) technique to balance the dataset. Both models are trained on imbalanced and balanced datasets. We used Accuracy, Sensitivity, Specificity, F1 Score, and AUC (Area Under the ROC curve) to evaluate the models' performances. we validate model performance using Rank Graduation Accuracy (RGA) to assess ranking consistency, revealing superior predictive power on imbalanced data (RGA = 0.991–0.993) versus balanced distributions (RGA = 0.959–0.965). Contrary to oversampling orthodoxy, ROSE balancing degraded performance aligning with theoretical critiques of synthetic data in mature classifiers. We also interpret the models by calculating feature importance using Shapley-Lorenz values. Partial Dependence Plots (PDPs) help to visualize how key financial ratios impact the predicted probability of default. Results show non-linear relationships between key financial ratios, such as Return on Assets (R6), Debt to Equity Ratio (R8), and default risk. The key features shown are similar for Random Forest and XGBoost, though the interpretation of the feature importance differs slightly. To enhance the robustness and credibility of our feature effect analysis, we conducted ALE (Accumulated Local Effects) plots as they provide a more robust framwork that accounts for feature interractions. This study advances credit default prediction in Tunisia's banking sector by enhancing interpretability through Accumulated Local Effects analysis alongside Partial Dependence Plots, providing robust insights into feature effects, particularly for key financial.&#xa0;Results offer insights about important ratio thresholds and their impact on default probability prediction, such as the sharp drop in default risk when R6 becomes positive.&#xa0;These advancements provide regulators and financial institutions with more reliable tools for credit risk assessment in Tunisia's economic context, bridging the gap between sophisticated machine learning techniques and practical, interpretable financial decision-making.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Partial dependence analysis of financial ratios in predicting company defaults: random forest vs XGBoost models

  • Monia Antar,
  • Tahar Tayachi

摘要

In this paper, we investigate using machine learning models to predict credit defaults using financial ratios. We compare the performance and interpretability of two ensemble learning algorithms: Random Forest and XGBoost. To improve the models' capability to detect defaults, we exploit the inherent class imbalance of default prediction tasks with the ROSE (Random Over-Sampling Examples) technique to balance the dataset. Both models are trained on imbalanced and balanced datasets. We used Accuracy, Sensitivity, Specificity, F1 Score, and AUC (Area Under the ROC curve) to evaluate the models' performances. we validate model performance using Rank Graduation Accuracy (RGA) to assess ranking consistency, revealing superior predictive power on imbalanced data (RGA = 0.991–0.993) versus balanced distributions (RGA = 0.959–0.965). Contrary to oversampling orthodoxy, ROSE balancing degraded performance aligning with theoretical critiques of synthetic data in mature classifiers. We also interpret the models by calculating feature importance using Shapley-Lorenz values. Partial Dependence Plots (PDPs) help to visualize how key financial ratios impact the predicted probability of default. Results show non-linear relationships between key financial ratios, such as Return on Assets (R6), Debt to Equity Ratio (R8), and default risk. The key features shown are similar for Random Forest and XGBoost, though the interpretation of the feature importance differs slightly. To enhance the robustness and credibility of our feature effect analysis, we conducted ALE (Accumulated Local Effects) plots as they provide a more robust framwork that accounts for feature interractions. This study advances credit default prediction in Tunisia's banking sector by enhancing interpretability through Accumulated Local Effects analysis alongside Partial Dependence Plots, providing robust insights into feature effects, particularly for key financial. Results offer insights about important ratio thresholds and their impact on default probability prediction, such as the sharp drop in default risk when R6 becomes positive. These advancements provide regulators and financial institutions with more reliable tools for credit risk assessment in Tunisia's economic context, bridging the gap between sophisticated machine learning techniques and practical, interpretable financial decision-making.