In this paper, we consider the application of machine learning methods such as gradient boosting and random forest to solve the problem of classifying financial data related to the prediction of defaults on credit obligations. The purpose of the study was to compare the effectiveness of these models and identify the most significant factors influencing the prediction of default. The models were evaluated based on accuracy, completeness, F1-score metrics and error matrices. The study showed that both methods demonstrated high accuracy, but revealed difficulties in classifying rare events such as default. The most important factors influencing the result were the annual income and the balance of the client's bank account. The work highlights the potential of using these methods to predict financial risks and make decisions in the banking sector.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prospects for Using Ensemble Machine Learning Methods for Predicting Credit Defaults

  • Vadim Tynchenko,
  • Ksenia Degtyareva,
  • Svetlana Kukartseva

摘要

In this paper, we consider the application of machine learning methods such as gradient boosting and random forest to solve the problem of classifying financial data related to the prediction of defaults on credit obligations. The purpose of the study was to compare the effectiveness of these models and identify the most significant factors influencing the prediction of default. The models were evaluated based on accuracy, completeness, F1-score metrics and error matrices. The study showed that both methods demonstrated high accuracy, but revealed difficulties in classifying rare events such as default. The most important factors influencing the result were the annual income and the balance of the client's bank account. The work highlights the potential of using these methods to predict financial risks and make decisions in the banking sector.