ChatGPT, Doubao, and human teachers in evaluating secondary school students’ handwritten EFL essays: evidence from the Chinese high-stakes exam context
摘要
This study investigates the scoring performance of ChatGPT and Doubao in evaluating handwritten English essays produced by Grade-9 students in Shanghai, China’s high-stakes exam context. Although large language models (LLMs) have shown promise in EFL writing evaluation in higher education contexts, little is known about their effectiveness in K–12 high-stakes exam settings or when applied to handwritten exam essays. Using an explanatory mixed-methods design, we compared the scores assigned by ChatGPT and Doubao in October 2025 to those of human teachers across three official criteria: Content, Grammar & Vocabulary, and Creativity (N = 32). Quantitative analyses, including Correlation Coefficients, Intraclass Correlation Coefficient, Mean Absolute Differences, and repeated-measures ANOVA, revealed that Doubao aligned more closely with human teachers in scoring the more objective criteria of Content and Grammar & Vocabulary. Both AI models showed noticeable divergence from human teachers in evaluating Creativity, a criterion heavily shaped by socio-cultural values, implicit preferences, and exam-oriented expectations. Follow-up interviews with EFL teachers further explained these discrepancies and highlighted the influence of emotional tone and positively oriented value expressions on human scoring. The findings suggest that Doubao holds potential as a supplementary scoring tool for high-stakes EFL writing evaluations in mainland China and other regions with limited access to ChatGPT. This study also underscores the limitations of LLMs in capturing subjective, culturally embedded evaluative criteria in handwritten essay scoring.