<p>This study explores ChatGPT’s ability to solve reading skill questions from the PISA exams, an internationally recognized assessment designed to evaluate 15-year-old students’ reading, mathematics, and science competencies. 90 questions from the 2000 and 2009 PISA reading assessments were administered to ChatGPT twice within one week. Questions were analyzed based on their text type (continuous, non-continuous, or multiple), question format (open-ended, multiple-choice, or short answer), and difficulty level defined by the PISA framework. The accuracy of ChatGPT’s responses was evaluated using the official answer keys provided by the OECD. Results indicate that ChatGPT achieved a 91% accuracy rate in both rounds. However, there were discrepancies in the questions answered incorrectly between the two rounds. ChatGPT performed better on continuous texts than non-continuous texts and achieved its highest accuracy rates on questions of lower difficulty. The analysis also highlights inconsistencies in ChatGPT’s responses, particularly in handling non-continuous texts and higher-level questions based on Bloom’s taxonomy. While ChatGPT has limitations, such as occasional incorrect or inconsistent answers, it remains a valuable tool with the potential for educational integration, particularly in automating assessment processes. The study underscores the importance of understanding ChatGPT’s strengths and limitations to enhance its application in educational settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing AI in Educational Evaluation: A Comprehensive Analysis of ChatGPT’s Performance on PISA Reading Skills

  • Mehmet Başaran,
  • Ömer Faruk Vural,
  • Cennet Tandırcı

摘要

This study explores ChatGPT’s ability to solve reading skill questions from the PISA exams, an internationally recognized assessment designed to evaluate 15-year-old students’ reading, mathematics, and science competencies. 90 questions from the 2000 and 2009 PISA reading assessments were administered to ChatGPT twice within one week. Questions were analyzed based on their text type (continuous, non-continuous, or multiple), question format (open-ended, multiple-choice, or short answer), and difficulty level defined by the PISA framework. The accuracy of ChatGPT’s responses was evaluated using the official answer keys provided by the OECD. Results indicate that ChatGPT achieved a 91% accuracy rate in both rounds. However, there were discrepancies in the questions answered incorrectly between the two rounds. ChatGPT performed better on continuous texts than non-continuous texts and achieved its highest accuracy rates on questions of lower difficulty. The analysis also highlights inconsistencies in ChatGPT’s responses, particularly in handling non-continuous texts and higher-level questions based on Bloom’s taxonomy. While ChatGPT has limitations, such as occasional incorrect or inconsistent answers, it remains a valuable tool with the potential for educational integration, particularly in automating assessment processes. The study underscores the importance of understanding ChatGPT’s strengths and limitations to enhance its application in educational settings.