Evaluating the performance of ChatGPT and Claude in automated writing scoring: Insights from the Many-facet Rasch model
摘要
The use of Large Language Models (LLMs) in automated writing scoring (AWS) systems presents a compelling alternative to traditional hand-scoring. However, few studies have compared the performance of different LLMs across distinct prompting strategies. Additionally, commonly used correlation and agreement-based indices, while critical for evaluating the alignment between LLMs and human raters, often overlook important estimators such as rater severity and potential gender bias. To address these issues, this study examines the grading performance of two prominent LLMs—ChatGPT and Claude—across two prompting strategies (zero-shot and few-shot prompting) using the Many-Facet Rasch Model (MFRM). Our sample included 117 English as a Foreign Language (EFL) students in China, who submitted essays that were evaluated by four human raters and four LLM-prompting combinations. The results reveal that: (1) LLM-based raters were more severe in scoring than human raters. ChatGPT with few-shot prompting (0.10 logits) was the closest to human raters, followed by ChatGPT with zero-shot (0.31 logits), Claude with few-shot (0.38 logits), and Claude with zero-shot (0.46 logits); (2) Few-shot prompting improved the consensus of LLM-based raters with human raters compared to zero-shot prompting; (3) LLM-based raters demonstrated greater scoring consistency compared to human raters; and (4) none of the LLM-based raters exhibited evidence of gender bias. This study highlights MFRM as an effective framework for evaluating LLMs performance in AWS and underscores the potential of LLMs as supplementary tools in educational assessment. Also, it emphasizes the need for continued optimization of LLM-based raters to ensure accurate and fair grading.