Optimizing Answer Representation Using Metric Learning for Efficient Short Answer Scoring
摘要
Automatic short answer scoring (ASAS) has received considerable attention in the field of education. However, existing methods typically treat ASAS as a standard text classification problem, following conventional pre-training or fine-tuning procedures. These approaches often generate embedding spaces that lack clear boundaries, resulting in overlapping representations for answers of different scores. To address this issue, we introduce a novel metric learning (MeL)-based pre-training method for answer representation optimization. This strategy encourages the clustering of similar representations while pushing dissimilar ones apart, thereby facilitating the formation of a more coherent same-score and distinct different-score answer embedding space. To fully exploit the potential of MeL, we define two types of answer similarities based on scores and rubrics, providing accurate supervised signals for improved training. Extensive experiments on thirteen short answer questions show that our method, even when paired with a simple linear model for downstream scoring, significantly outperforms prior ASAS methods in both scoring accuracy and efficiency.