Identifying at-risk students for timely intervention to improve academic performance remains a significant challenge for schools. This study introduces an interpretable deep learning framework to identify third-year middle school students at risk of underperformance based on the overall average (OA) prediction five months before the final exam. Using a deep neural network (DNN) regression model enhanced by SHapley Additive exPlanations (SHAP). The framework predicts overall student performance and explains the impact of academic, demographic, and socioeconomic features. Using a dataset of 551,272 students’ records, the model achieves a Mean Squared Error (MSE) of 0.1278, a Root Mean Squared Error (RMSE) of 0.3532, and an R-squared (R2) of 0.86, demonstrating strong predictive accuracy and the ability to capture a substantial proportion of the variance in student performance. Additionally, our model categorizes students into high-risk, moderate-risk, and no-risk groups with 85% accuracy. SHAP analysis identifies academic performance, poverty, and class size as key predictors, providing actionable insights to support early interventions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Early Identification of Students at Risk of Underperformance: A Deep Learning Approach with SHAP-Based Interpretability

  • Mohamed El Jihaoui,
  • Oum El Kheir Abra,
  • Khalifa Mansouri

摘要

Identifying at-risk students for timely intervention to improve academic performance remains a significant challenge for schools. This study introduces an interpretable deep learning framework to identify third-year middle school students at risk of underperformance based on the overall average (OA) prediction five months before the final exam. Using a deep neural network (DNN) regression model enhanced by SHapley Additive exPlanations (SHAP). The framework predicts overall student performance and explains the impact of academic, demographic, and socioeconomic features. Using a dataset of 551,272 students’ records, the model achieves a Mean Squared Error (MSE) of 0.1278, a Root Mean Squared Error (RMSE) of 0.3532, and an R-squared (R2) of 0.86, demonstrating strong predictive accuracy and the ability to capture a substantial proportion of the variance in student performance. Additionally, our model categorizes students into high-risk, moderate-risk, and no-risk groups with 85% accuracy. SHAP analysis identifies academic performance, poverty, and class size as key predictors, providing actionable insights to support early interventions.