A Review on Understanding the Patterns of Student Dropout
摘要
Higher education institutions serve as crucial centers for the accumulation of vast quantities of student data, providing fertile ground for information gathering, knowledge development, and monitoring. This data reservoir, comprising demographic, socioeconomic, macroeconomic, and academic statistics, is pivotal in understanding and addressing issues such as early dropout rates and failure rates in higher education. These rates not only have immediate ramifications but also exert a profound influence on economic growth factors. Early dropout and failure rates impact not just individuals but also society, institutions, and families. The analysis utilized a dataset of approximately 4424 students enrolled in various undergraduate programs (agronomy, design, education, etc.) across first and second semesters. The data encompasses demographic, socioeconomic, and macroeconomic factors collected at the time of enrollment from sources like the General Directorate of Higher Education (DGES) and the Contemporary Portugal Database (PORDATA). F1 score was employed for model selection, while Logistic Regression, K-Nearest Neighbors, Random Forest (ML), and Artificial Neural Networks (DL) were used for prediction and evaluation. Additionally, quantitative metrics like accuracy, precision, and recall were employed for decision-making and model comparisons. By leveraging this system effectively, higher education institutions can play a pivotal role in steering the trajectory of students’ educational journeys and contributing to the broader development of society. The results that we achieved demonstrate that the DL approach, that is ANN, performed much better on student dropout data compared with other available methodologies.