Massive Open Online Courses offer scalable learning opportunities, yet ensuring robust assessment remains challenging. Advances in Artificial Intelligence, especially large language models like GPT-4, present new possibilities for benchmarking and refining course materials. This paper investigates how GPT-4 compares to human learners in ten introductory programming MOOCs—six Python and four Java courses—on multiple-choice and coding exercises. We analyze the model’s performance at an exercise level compared to average student scores. Results show GPT-4 frequently outperforms average learners, particularly in Python courses, with statistically significant advantages observed. However, the model’s advantages are less pronounced or insignificant in Java courses. Crucially, we highlight scenarios where GPT-4 struggles: exercises involving specialized libraries, multimedia content, or external contextual knowledge. These insights assist course designers in identifying ambiguous instructions and context-dependent tasks, enabling targeted course design improvements. Such refinements ensure materials better meet human learners’ needs while anticipating AI advancements, thus sustaining MOOC quality and fairness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

When GPT-4 Goes to Class: Benchmarking MOOCs and Enhancing Course Design with AI

  • Mohamed Elhayany,
  • Christoph Meinel

摘要

Massive Open Online Courses offer scalable learning opportunities, yet ensuring robust assessment remains challenging. Advances in Artificial Intelligence, especially large language models like GPT-4, present new possibilities for benchmarking and refining course materials. This paper investigates how GPT-4 compares to human learners in ten introductory programming MOOCs—six Python and four Java courses—on multiple-choice and coding exercises. We analyze the model’s performance at an exercise level compared to average student scores. Results show GPT-4 frequently outperforms average learners, particularly in Python courses, with statistically significant advantages observed. However, the model’s advantages are less pronounced or insignificant in Java courses. Crucially, we highlight scenarios where GPT-4 struggles: exercises involving specialized libraries, multimedia content, or external contextual knowledge. These insights assist course designers in identifying ambiguous instructions and context-dependent tasks, enabling targeted course design improvements. Such refinements ensure materials better meet human learners’ needs while anticipating AI advancements, thus sustaining MOOC quality and fairness.