The application of GPT-4 in grading design university students’ assignment: an exploratory study
摘要
This study investigates whether GPT-4 can reliably grade assignments in a design education context, where tasks are typically open-ended and lack a single correct answer. Such subjectivity often leads to inconsistent grading between human raters. Using an iterative research process, we developed a customized GPT-4 and tested its ability to deliver consistent assessments. The findings show that, after several rounds of refinement, the inter-rater reliability between GPT-4 and human assessors reached a level generally accepted in educational settings. This suggests that, with well-crafted prompts and customization, GPT-4 can serve as a reliable complement to human raters. We also observed moderate consistency in GPT-4’s grading over time, with intra-rater reliability scores ranging from 0.65 to 0.78. As consistency and comparability are key principles of reliable assessment, this study explores both whether a Custom GPT can meet these criteria and how iterative refinement supports this goal. The preliminary results suggest that, with appropriate prompting and iteration, GPT-4 may support more consistent grading in design education, offering a possible complement to human assessment.