The expense and logistics of organizing experiments to train and evaluate teaching policies, as well as the potential negative impacts of these policies on the very first students, are significant challenges in the field of education. In this paper, we explore the feasibility of using offline reinforcement learning (RL) to learn an adaptive feedback policy that improves student progress in Pyrates, a programming platform. Leveraging an existing dataset of teacher-provided feedback and student interactions and codes, we have developed an offline RL model capable of learning an optimal feedback policy without direct interaction with students. The trained policy is then evaluated to assess its effectiveness in maximizing student progress and its suitability for online deployment in real-world educational settings. Our evaluation yields promising results, demonstrating that offline-trained policies can significantly enhance student progress, highlighting their potential for scalable deployment in programming platforms. However, challenges remain in ensuring robustness and adaptability when transitioning to online use.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning Feedback Policy from Historical Data: An Offline Approach Within Pyrates

  • Badmavasan Kirouchenassamy,
  • Amel Yessad,
  • Sébastien Jolivet,
  • Matthieu Branthôme,
  • Sébastien Lallé,
  • Vanda Luengo

摘要

The expense and logistics of organizing experiments to train and evaluate teaching policies, as well as the potential negative impacts of these policies on the very first students, are significant challenges in the field of education. In this paper, we explore the feasibility of using offline reinforcement learning (RL) to learn an adaptive feedback policy that improves student progress in Pyrates, a programming platform. Leveraging an existing dataset of teacher-provided feedback and student interactions and codes, we have developed an offline RL model capable of learning an optimal feedback policy without direct interaction with students. The trained policy is then evaluated to assess its effectiveness in maximizing student progress and its suitability for online deployment in real-world educational settings. Our evaluation yields promising results, demonstrating that offline-trained policies can significantly enhance student progress, highlighting their potential for scalable deployment in programming platforms. However, challenges remain in ensuring robustness and adaptability when transitioning to online use.