Can GPT4 Answer Educational Tests? Empirical Analysis of Answer Quality Based on Question Complexity and Difficulty
摘要
While recent advancements in Large Language Models (LLMs) suggest their potential to tackle these challenges, limited research exists on how well LLMs respond to open-ended questions with varying difficulty and complexity. This paper addresses this gap by comparing GPT4’s performance with human counterparts, considering question difficulty (assessed through Item Response Theory – IRT) and complexity (categorized based on Bloom’s taxonomy levels) using a dataset of 7,380 open-ended questions related to high school topics. Overall, the results indicate that GPT4 surpasses non-native speakers and demonstrates comparable performance to native speakers. Moreover, despite facing challenges in tasks involving basic recall or creative thinking, GPT4’s performance notably improves with increasing question difficulty. Therefore, this paper contributes empirical evidence on GPT4’s effectiveness in addressing open-ended questions, enhancing our understanding of its potential and limitations in educational settings. The findings offer valuable insights for practitioners and researchers seeking to incorporate LLMs into educational practices, such as assessment, virtual assistant and feedback.