Can Large Language Models Generate Middle School Mathematics Explanations Better Than Human Teachers?
摘要
The rapid development of large language models has led to increased interest in evaluating the efficacy of models for education technology. In this paper, we are particularly interested in GPT created by OpenAI. Specifically, we are interested in comparing GPT-3.5 and GPT-4 in their ability to generate explanations for math problems. Computer-generated explanations could be readily utilized in online tutoring systems as a supplement to existing teacher-created explanations. Conclusions from prior work regarding GPT-3.5 generated explanations found a high mathematical error rate of 90% and 57% respectively for two methods tested and a poor quality rating when assessed by teachers. In this work, we present improved methods using GPT-4 that significantly reduced errors in generated explanations to 6%. Additionally, a preregistered study with evaluators rated the GPT-4 explanations as higher in quality compared to human explanations. Our main interest is looking at the quality of generated explanations, so other AIED systems can begin using GPT-4 to create feedback to better foster student learning.