Comparative Analysis of LLM-Based Writing Tools for Error Correction and Feedback: A Study on ChatGPT-3.5, ChatGPT-4, Gemini, and Claude 3
摘要
Second language writing instruction has increasingly focused on the writing process and idea generation, while less emphasis has been placed on helping students produce error-free sentences. However, grammatical accuracy remains challenging and essential for L2 learners to express their ideas effectively. Recently, large language models (LLMs), such as ChatGPT, have gained attention for their advanced ability to generate and refine text, offering essential support for writing. This study evaluates the performance of four prominent LLM-based writing tools—ChatGPT-4, ChatGPT-3.5, Gemini, and Claude 3—on 560 ungrammatical sentences representing 20 error types from Common Mistakes in English (Longman). The tools were assessed on their accuracy in detecting and correcting grammatical errors, with Grammarly as the benchmark. Results show that all LLM tools outperformed Grammarly, with ChatGPT-4 leading in both error detection and correction, followed by ChatGPT-3.5, Claude, and Gemini. The analysis also identified strengths and weaknesses in the correction of specific error types, including misused forms, incorrect omissions, unnecessary words, misplaced words, and confused words. Additionally, the tools’ explanations for corrections were evaluated for readability, comprehensibility, detail, and accuracy. Overall, the present study sheds light on the capabilities and limitations of LLM-based writing tools, offering valuable insights into their effectiveness in language education.