Automated Essay Scoring (AES) has gained prominence in educational settings due to its potential to enhance grading efficiency and consistency. While most AES models output a single holistic score, fewer studies have focused on generating multiple elementary scores under rubric-based settings. To fill this gap, we develop a novel multi-view machine learning algorithm that learns a multi-output AES model, which combines BERT (Bidirectional Encoder Representations from Transformers), SentenceBERT, and manually extracted features. We train and evaluate our model on two real-world datasets prepared from formal high school assessments in New Zealand, comparing its performance against a single-output BERT model, a multi-output BERT model, and a widely used commercial IntelliMetric model for AES. Results show that our model consistently outperforms these methods, achieving the highest Quadratic Weighted Kappa (QWK) scores across all rubric elements. This suggests that the multi-view representation is an effective approach for multi-output AES under rubric settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Rubric-Based Automated Essay Scoring with Multi-view BERT: A Case Study in New Zealand

  • Jie Dong,
  • Xiaoying Gao,
  • Yi Mei

摘要

Automated Essay Scoring (AES) has gained prominence in educational settings due to its potential to enhance grading efficiency and consistency. While most AES models output a single holistic score, fewer studies have focused on generating multiple elementary scores under rubric-based settings. To fill this gap, we develop a novel multi-view machine learning algorithm that learns a multi-output AES model, which combines BERT (Bidirectional Encoder Representations from Transformers), SentenceBERT, and manually extracted features. We train and evaluate our model on two real-world datasets prepared from formal high school assessments in New Zealand, comparing its performance against a single-output BERT model, a multi-output BERT model, and a widely used commercial IntelliMetric model for AES. Results show that our model consistently outperforms these methods, achieving the highest Quadratic Weighted Kappa (QWK) scores across all rubric elements. This suggests that the multi-view representation is an effective approach for multi-output AES under rubric settings.