<p>Classification of academic results of students using machine learning plays an important role in educational data mining (EDM) for uncovering hidden patterns in students’ academic behavior. This not only helps educators perform early remedial interventions to reduce chances of student failure but can also enable students, particularly, in universities to take charge of their learning process. As student results depend on numerous academic and non-academic factors and owing to the nature of university student data, usually, imbalanced datasets with comparatively higher dimensionality than the sample size, are used in classification of academic results at university level. This in turn leads to low accuracy in classification due to poor generalization of machine learning based classifiers. The study investigates effectiveness of dimensionality reduction of such unbalanced university student datasets in improving perfromance of contemporary machine learning. Dimensionality reduction of primary student datasets from two public universities is analyzed using various state-of-the-art dimensionality reduction techniques. A relative analysis of contemporary machine learning algorithms for classifying student results at the end of a semester is performed on original data with and without non-academic inputs and various dimensionality reduced versions of the dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Academic result classification of university students and the impact of dataset dimensionality reduction

  • Sharifa Rajab,
  • S. M. K. Quadri

摘要

Classification of academic results of students using machine learning plays an important role in educational data mining (EDM) for uncovering hidden patterns in students’ academic behavior. This not only helps educators perform early remedial interventions to reduce chances of student failure but can also enable students, particularly, in universities to take charge of their learning process. As student results depend on numerous academic and non-academic factors and owing to the nature of university student data, usually, imbalanced datasets with comparatively higher dimensionality than the sample size, are used in classification of academic results at university level. This in turn leads to low accuracy in classification due to poor generalization of machine learning based classifiers. The study investigates effectiveness of dimensionality reduction of such unbalanced university student datasets in improving perfromance of contemporary machine learning. Dimensionality reduction of primary student datasets from two public universities is analyzed using various state-of-the-art dimensionality reduction techniques. A relative analysis of contemporary machine learning algorithms for classifying student results at the end of a semester is performed on original data with and without non-academic inputs and various dimensionality reduced versions of the dataset.