Data-Driven Peer Tutoring: Machine Learning for Effective Group Composition
摘要
This study explores the use of machine learning techniques to optimize peer tutoring in education, identifying students best suited for the role of tutor and ensuring equitable distribution in study groups. Using data from the UK Open University, the research analyzes student demographic information, interactions with e-learning platforms and academic performance metrics to build predictive models. A rigorous methodology was adopted, including data preprocessing, feature engineering and model evaluation. Among the algorithms tested, CatBoost showed the best predictive performance with an F1 score of 0.719. By adjusting the classification thresholds, the study achieved a balance between precision and recall, ensuring reliable tutor selection minimizing false positives. Feature importance analysis showed that assiduity in e-learning resources has a significant influence on the results. These results underscore the potential of machine learning in enhancing collaborative learning environments and promoting academic success.