Comparative Analysis of Classical Machine Learning Techniques for Predicting Students’ Exam Performance
摘要
The prediction of student performance is a critical area of research in educational data mining, aiming to identify factors that contribute to academic success or failure. Accurate prediction models can help educators and policymakers develop interventions to improve student outcomes. Despite the availability of various machine learning techniques, there remains a need for a comprehensive comparison of these methods applied to a single dataset. This study addresses this gap by applying several classical machine learning algorithms, including Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, K-Nearest Neighbors, Naive Bayes, and Gradient Boosting, to predict student performance using a publicly available dataset from Kaggle. The dataset includes demographic and educational attributes such as gender, race/ethnicity, parental level of education, lunch type, and test preparation course. The models were evaluated based on accuracy, precision, recall, F1 score, and ROC curves. Logistic Regression emerged as the top-performing model with an AUC of 0.72, accuracy of 0.71, and F1 score of 0.81, indicating robust predictive capability and high interpretability. Naive Bayes also performed competitively, highlighting the effectiveness of simple probabilistic models. These findings provide valuable insights for educators and policymakers, emphasizing the importance of machine learning in early identification of students at risk of underperforming, thereby enabling timely interventions. Future research should explore advanced techniques such as deep learning, feature engineering, and longitudinal data analysis to further enhance predictive accuracy. Additionally, the interpretability and ethical implications of deploying these models in educational settings must be considered to ensure responsible usage.