Dropout Prediction for Higher Education: Data Sets and Methods, a Brief Overview
摘要
Dropout prediction represents one of the head educational objectives in higher education. One problem with making this prediction is generating or acquiring the data sets to train and validate the models that will make it. As a preliminary approach to these challenges, this study presents a comparison of five datasets used by researchers in projects related to dropout prediction, along with a summary of machine learning techniques such as decision trees, random forests, linear regression, artificial neural networks, support vector machine, random forests, among others. Finally, this investigation will provide to researchers and practitioners of educational data mining useful and important information to achieve dropout prediction goals effectively and efficiently, highlighting relevant features and particularities of datasets and what are the most popular techniques to find insights for dropout prediction.