<p>The use of Large Language Models (LLMs) in automated writing scoring (AWS) systems presents a compelling alternative to traditional hand-scoring. However, few studies have compared the performance of different LLMs across distinct prompting strategies. Additionally, commonly used correlation and agreement-based indices, while critical for evaluating the alignment between LLMs and human raters, often overlook important estimators such as rater severity and potential gender bias. To address these issues, this study examines the grading performance of two prominent LLMs—ChatGPT and Claude—across two prompting strategies (zero-shot and few-shot prompting) using the Many-Facet Rasch Model (MFRM). Our sample included 117 English as a Foreign Language (EFL) students in China, who submitted essays that were evaluated by four human raters and four LLM-prompting combinations. The results reveal that: (1) LLM-based raters were more severe in scoring than human raters. ChatGPT with few-shot prompting (0.10 logits) was the closest to human raters, followed by ChatGPT with zero-shot (0.31 logits), Claude with few-shot (0.38 logits), and Claude with zero-shot (0.46 logits); (2) Few-shot prompting improved the consensus of LLM-based raters with human raters compared to zero-shot prompting; (3) LLM-based raters demonstrated greater scoring consistency compared to human raters; and (4) none of the LLM-based raters exhibited evidence of gender bias. This study highlights MFRM as an effective framework for evaluating LLMs performance in AWS and underscores the potential of LLMs as supplementary tools in educational assessment. Also, it emphasizes the need for continued optimization of LLM-based raters to ensure accurate and fair grading.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the performance of ChatGPT and Claude in automated writing scoring: Insights from the Many-facet Rasch model

  • Rui Jin,
  • Mingren Zhao,
  • Chunling Niu,
  • Yuyan Xia,
  • Hao Zhou,
  • Na Liu

摘要

The use of Large Language Models (LLMs) in automated writing scoring (AWS) systems presents a compelling alternative to traditional hand-scoring. However, few studies have compared the performance of different LLMs across distinct prompting strategies. Additionally, commonly used correlation and agreement-based indices, while critical for evaluating the alignment between LLMs and human raters, often overlook important estimators such as rater severity and potential gender bias. To address these issues, this study examines the grading performance of two prominent LLMs—ChatGPT and Claude—across two prompting strategies (zero-shot and few-shot prompting) using the Many-Facet Rasch Model (MFRM). Our sample included 117 English as a Foreign Language (EFL) students in China, who submitted essays that were evaluated by four human raters and four LLM-prompting combinations. The results reveal that: (1) LLM-based raters were more severe in scoring than human raters. ChatGPT with few-shot prompting (0.10 logits) was the closest to human raters, followed by ChatGPT with zero-shot (0.31 logits), Claude with few-shot (0.38 logits), and Claude with zero-shot (0.46 logits); (2) Few-shot prompting improved the consensus of LLM-based raters with human raters compared to zero-shot prompting; (3) LLM-based raters demonstrated greater scoring consistency compared to human raters; and (4) none of the LLM-based raters exhibited evidence of gender bias. This study highlights MFRM as an effective framework for evaluating LLMs performance in AWS and underscores the potential of LLMs as supplementary tools in educational assessment. Also, it emphasizes the need for continued optimization of LLM-based raters to ensure accurate and fair grading.