This research explores the application of various machine learning models for credit risk assessment in the context of a highly imbalanced dataset. The study encompasses logistic regression, random forest, gradient boosting algorithms (LightGBM and XGBoost), neural networks, and ensemble methods. An extensive data processing pipeline, including categorical encoding, missing value imputation, and feature scaling, is implemented to enhance model performance. Evaluation metrics, primarily focusing on AUC-ROC scores, are used to assess model accuracy. The findings reveal that gradient boosting algorithms, specifically LightGBM, outperform other models, achieving the highest AUC-ROC score of 0.80123. Deep learning models, such as TabNet, demonstrate potential but require further development for competitive performance. Logistic regression and random forest models yield respectable results, showcasing the significance of feature engineering in enhancing their predictive capabilities. Ensemble methods, including stacking and weighted averaging, are explored to leverage the strengths of individual models. The study emphasizes the importance of model selection and parameter tuning in achieving optimal results. Logistic regression, despite its simplicity, performs well due to the incorporation of a feature-engineered pipeline. In conclusion, the research suggests that, while current deep learning models for tabular data show promise, gradient boosting algorithms remain superior for credit risk assessment. The study highlights the need for further exploration of advanced techniques and emphasizes the potential of regression algorithms for a comprehensive credit assessment score, complementing the classifier algorithms discussed.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Journey Through Multifaceted Data in Machine Learning Predictions on Financial Viability

  • R. Revathi

摘要

This research explores the application of various machine learning models for credit risk assessment in the context of a highly imbalanced dataset. The study encompasses logistic regression, random forest, gradient boosting algorithms (LightGBM and XGBoost), neural networks, and ensemble methods. An extensive data processing pipeline, including categorical encoding, missing value imputation, and feature scaling, is implemented to enhance model performance. Evaluation metrics, primarily focusing on AUC-ROC scores, are used to assess model accuracy. The findings reveal that gradient boosting algorithms, specifically LightGBM, outperform other models, achieving the highest AUC-ROC score of 0.80123. Deep learning models, such as TabNet, demonstrate potential but require further development for competitive performance. Logistic regression and random forest models yield respectable results, showcasing the significance of feature engineering in enhancing their predictive capabilities. Ensemble methods, including stacking and weighted averaging, are explored to leverage the strengths of individual models. The study emphasizes the importance of model selection and parameter tuning in achieving optimal results. Logistic regression, despite its simplicity, performs well due to the incorporation of a feature-engineered pipeline. In conclusion, the research suggests that, while current deep learning models for tabular data show promise, gradient boosting algorithms remain superior for credit risk assessment. The study highlights the need for further exploration of advanced techniques and emphasizes the potential of regression algorithms for a comprehensive credit assessment score, complementing the classifier algorithms discussed.