<p>Accurate prediction of student performance is essential for facilitating early interventions and enhancing academic results. However, interpretability is lacking in traditional machine learning approaches, which frequently struggle to capture implicit behavioural and learning dependencies within educational data. The study utilizes a publicly available dataset from the Kaggle repository, containing student records with features related to academic, demographic, and behavioural attributes. The dataset was standardized using z-score normalization. Four deep learning models were developed based on their effectiveness in processing sequential data: LSTM, CNN, RNN, and a hybrid model combining LSTM and Transformer architectures (LSTM-Transformer). To ensure model transparency, we integrated Explainable Artificial Intelligence (XAI) techniques such as SHapley Additive exPlanations (SHAP) and Transformer attention mechanisms. The hybrid LSTM–Transformer model consistently achieved the best performance, with the lowest MSE and MAE and highest R<sup>2</sup> across most runs. Statistical tests confirmed significant differences among models (Friedman p &lt; 0.0001), with the hybrid outperforming CNN (p = 0.0286) and RNN (p &lt; 0.0001). While LSTM and the hybrid showed comparable accuracy (p &gt; 0.05). The hybrid model’s interpretability and performance make it a strong candidate for future real-world deployment; however, further work is needed to assess integration into educational systems, computational feasibility, and stakeholder acceptance. According to SHAP and attention studies, absenteeism, study time, and parental support were identified as the most influential predictors of performance, while gender and parents' level of education had little bearing. This research presents a balanced framework that advances both accuracy and interpretability in student performance modelling.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainable artificial intelligence in LSTM transformer models for student performance analysis

  • Elijah Ofori,
  • Delali Kwasi Dake

摘要

Accurate prediction of student performance is essential for facilitating early interventions and enhancing academic results. However, interpretability is lacking in traditional machine learning approaches, which frequently struggle to capture implicit behavioural and learning dependencies within educational data. The study utilizes a publicly available dataset from the Kaggle repository, containing student records with features related to academic, demographic, and behavioural attributes. The dataset was standardized using z-score normalization. Four deep learning models were developed based on their effectiveness in processing sequential data: LSTM, CNN, RNN, and a hybrid model combining LSTM and Transformer architectures (LSTM-Transformer). To ensure model transparency, we integrated Explainable Artificial Intelligence (XAI) techniques such as SHapley Additive exPlanations (SHAP) and Transformer attention mechanisms. The hybrid LSTM–Transformer model consistently achieved the best performance, with the lowest MSE and MAE and highest R2 across most runs. Statistical tests confirmed significant differences among models (Friedman p < 0.0001), with the hybrid outperforming CNN (p = 0.0286) and RNN (p < 0.0001). While LSTM and the hybrid showed comparable accuracy (p > 0.05). The hybrid model’s interpretability and performance make it a strong candidate for future real-world deployment; however, further work is needed to assess integration into educational systems, computational feasibility, and stakeholder acceptance. According to SHAP and attention studies, absenteeism, study time, and parental support were identified as the most influential predictors of performance, while gender and parents' level of education had little bearing. This research presents a balanced framework that advances both accuracy and interpretability in student performance modelling.