Labeling Disagreements: Illuminating the Classification of Mathematics Teacher Questions
摘要
Text classification models are increasingly used to label teacher-talk and provide teachers with feedback about their classroom teaching. Following calls for increasing the transparency of models used to support learning, the performance of these models is typically reported using quantitative techniques that explain why models make classification decisions. However, the use of qualitative methods to better understand the decisions that models make have received little attention, despite their widespread use in building trustworthiness in other areas of educational measurement. This paper explores the use of the rater norming process, where disagreements between labelers are compared and discussed to better understand how classifications are made. It uses a dataset of text messages collected from chat conversations between pre-service teachers during simulated teaching activities. The labels used were drawn from research-based teacher question classifications used by teacher educators to provide feedback to pre-service teachers. After labeling the teacher text messages, we compared disagreements between two humans and then between a human and a trained model. Results describe the question ambiguities that impact labeling and demonstrate how qualitative approaches can provide additional model transparency. The paper also presents implications for future model development, for instructors of pre-service teacher instructors, and for learning tool designers.