This research studies the viability of small samples of basic student demographical information as training data for machine learning models to perform multiclass classification of the specializations/majors of students enrolled in the Master of Applied Computing course at Taylor’s University. If successful, this approach can be used by educational institutions with sparse student data to recommend university majors/specializations. This research uses SMOTE-NC to augment the training dataset from 25 rows to 100 rows and trained three machine learning models on the augmented dataset—logistic regression, random forest, and support vector machine. The support vector machine classifier performed the best in terms of accuracy at 86%, but all three classifiers showed highly inconsistent performance when predicting different specializations. Further research is required to understand this outcome, but potential causes include weak correlations between the input features and the target feature and small sample size.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Viability of Using Data Augmentation with a Small Sample of Demographical Information to Predict Student Specialization in Master of Applied Computing Course

  • Lau Noel Kuan Kiat,
  • Humaira Ashraf,
  • Navid Ali Khan

摘要

This research studies the viability of small samples of basic student demographical information as training data for machine learning models to perform multiclass classification of the specializations/majors of students enrolled in the Master of Applied Computing course at Taylor’s University. If successful, this approach can be used by educational institutions with sparse student data to recommend university majors/specializations. This research uses SMOTE-NC to augment the training dataset from 25 rows to 100 rows and trained three machine learning models on the augmented dataset—logistic regression, random forest, and support vector machine. The support vector machine classifier performed the best in terms of accuracy at 86%, but all three classifiers showed highly inconsistent performance when predicting different specializations. Further research is required to understand this outcome, but potential causes include weak correlations between the input features and the target feature and small sample size.