Academic dropout is one of the main challenges faced by higher education institutions, affecting both the efficiency of the educational system and the development of qualified professionals. These studies suggest that effective school inclusion, especially for students with special needs, requires understanding individual differences, collaborative and well-trained support among educators and families, and that simply adding technology or policies is not enough without addressing broader social and educational factors. This study proposes a machine learning-based predictive model to identify key factors associated with student dropout in the state of Pernambuco, Brazil, supporting data-driven decision-making in educational management. The dataset, sourced from the Faculdade Pernambucana de Saúde, includes academic, socioeconomic, and demographic variables. The modeling pipeline involved data normalization, feature selection using Boruta with Random Forest, class balancing via SMOTEENN, and hyperparameter tuning with GridSearchCV and stratified cross-validation. Multiple models were evaluated—Decision Tree, SVM, Logistic Regression, and XGBoost—using accuracy, F1-score, and ROC AUC. Additionally, SHAP was applied to interpret model outputs and assess feature contributions. Results highlighted the importance of academic performance, distance to campus, and financial status in predicting dropout, offering valuable insights for the development of more effective institutional retention policies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Explainable Approach to Predicting Academic Dropout: A Case Study at Faculdade Pernambucana de Saúde

  • Jefferson Rodrigues de Melo,
  • Maynara Donato de Souza,
  • Dalila Duraes,
  • Flávio A. O. Santos

摘要

Academic dropout is one of the main challenges faced by higher education institutions, affecting both the efficiency of the educational system and the development of qualified professionals. These studies suggest that effective school inclusion, especially for students with special needs, requires understanding individual differences, collaborative and well-trained support among educators and families, and that simply adding technology or policies is not enough without addressing broader social and educational factors. This study proposes a machine learning-based predictive model to identify key factors associated with student dropout in the state of Pernambuco, Brazil, supporting data-driven decision-making in educational management. The dataset, sourced from the Faculdade Pernambucana de Saúde, includes academic, socioeconomic, and demographic variables. The modeling pipeline involved data normalization, feature selection using Boruta with Random Forest, class balancing via SMOTEENN, and hyperparameter tuning with GridSearchCV and stratified cross-validation. Multiple models were evaluated—Decision Tree, SVM, Logistic Regression, and XGBoost—using accuracy, F1-score, and ROC AUC. Additionally, SHAP was applied to interpret model outputs and assess feature contributions. Results highlighted the importance of academic performance, distance to campus, and financial status in predicting dropout, offering valuable insights for the development of more effective institutional retention policies.