This study explores the applicability of Large Language Models (LLMs) in the assessment of student translations, focusing on their stability, reliability, and adaptability to evaluation criteria. Based on the Human Translation Quality Evaluation (HTQE) criteria, the study selects culturally themed expository texts from English-Chinese bidirectional translation corpora to assess the consistency and bias of LLM-generated scores using Intraclass Correlation Coefficients (ICC) and Pearson correlation analysis. Results indicate that LLMs exhibit stability and alignment with human scoring in structured dimensions such as grammatical accuracy and terminology consistency. However, reliability fluctuates significantly in higher-order semantic dimensions, such as intent reproduction and cultural adaptation. Furthermore, LLM-generated scores for English-to-Chinese translation show higher consistency than those for Chinese-to-English translation, reflecting the influence of training data distribution. Current LLMs remain constrained by the transparency of their evaluation mechanisms and their ability to comprehend context, making it challenging to replicate the holistic judgment of human evaluators. To enhance the applicability of LLMs in translation pedagogy, this study proposes optimizing human–machine collaborative evaluation, refining translation evaluation criteria, and fostering students’ translation evaluation skills. By expanding the application of LLMs in translation quality evaluation, this study provides empirical support for the enhancement of intelligent evaluation systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Empirical Study on AI-Driven Evaluation of Student Translations Using Large Language Models

  • Sirui Peng,
  • Jing Zhang

摘要

This study explores the applicability of Large Language Models (LLMs) in the assessment of student translations, focusing on their stability, reliability, and adaptability to evaluation criteria. Based on the Human Translation Quality Evaluation (HTQE) criteria, the study selects culturally themed expository texts from English-Chinese bidirectional translation corpora to assess the consistency and bias of LLM-generated scores using Intraclass Correlation Coefficients (ICC) and Pearson correlation analysis. Results indicate that LLMs exhibit stability and alignment with human scoring in structured dimensions such as grammatical accuracy and terminology consistency. However, reliability fluctuates significantly in higher-order semantic dimensions, such as intent reproduction and cultural adaptation. Furthermore, LLM-generated scores for English-to-Chinese translation show higher consistency than those for Chinese-to-English translation, reflecting the influence of training data distribution. Current LLMs remain constrained by the transparency of their evaluation mechanisms and their ability to comprehend context, making it challenging to replicate the holistic judgment of human evaluators. To enhance the applicability of LLMs in translation pedagogy, this study proposes optimizing human–machine collaborative evaluation, refining translation evaluation criteria, and fostering students’ translation evaluation skills. By expanding the application of LLMs in translation quality evaluation, this study provides empirical support for the enhancement of intelligent evaluation systems.