In this study, an effort is made to see if machine learning classification techniques can be used to evaluate the efficiency of loan reimbursement in creating jobs for young people and comprehending organizational business goals. Youth loan data is gathered, cleaned, processed, transformed, and then ready for model building using the classification algorithm to create a prediction model. The final dataset ready for model building included six attributes and 1505 records collected from 6 different Woredas of Addis Ababa Savings and Credit Association Nifas Silk Lafto Sub-city Branch, which are kept in various Excel files in the form of reports from 2017 to 2020. For the experiment, four classification algorithms, namely Logistic Regression, K Nearest Neighbor, Support Vector Machine, and Random Forest, are used for model building. Random under-sampling and SMOTE oversampling class imbalance handling techniques are implemented using Python programming. The preprocessed dataset was split into 90–10 for building and testing the models. The model trained with a support vector machine algorithm achieved the best performance, with an accuracy of 0.92.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Machine Learning for Assessing Youth Loan Reimbursement Impact on Job Creation in Ethiopia

  • Tarikwa Tesfa,
  • Isayas Feyera,
  • Yimer Amedie,
  • Sudhir Kumar Mohapatra,
  • Jukka-Pekka Skön,
  • Jukka Heikkonen,
  • Rajeev Kanth

摘要

In this study, an effort is made to see if machine learning classification techniques can be used to evaluate the efficiency of loan reimbursement in creating jobs for young people and comprehending organizational business goals. Youth loan data is gathered, cleaned, processed, transformed, and then ready for model building using the classification algorithm to create a prediction model. The final dataset ready for model building included six attributes and 1505 records collected from 6 different Woredas of Addis Ababa Savings and Credit Association Nifas Silk Lafto Sub-city Branch, which are kept in various Excel files in the form of reports from 2017 to 2020. For the experiment, four classification algorithms, namely Logistic Regression, K Nearest Neighbor, Support Vector Machine, and Random Forest, are used for model building. Random under-sampling and SMOTE oversampling class imbalance handling techniques are implemented using Python programming. The preprocessed dataset was split into 90–10 for building and testing the models. The model trained with a support vector machine algorithm achieved the best performance, with an accuracy of 0.92.