LLM-Based Machine Interpreting Evaluation: Implications for Interpreter Education in the Age of AI
摘要
This study investigates the potential of Large Language Model (LLM)-based approaches for evaluating machine interpreting (MI) output, with a focus on implications for interpreter education. First, two experiments using GPT-4o were conducted to assess the quality of Automatic Speech Recognition (ASR) and Machine Translation (MT) outputs. The first experiment applied an adapted GEMBA-SQM (Scalar Quality Metric) prompting strategy to generate overall scores across three interpreting modes: ASR-MT, ST-ASR-MT, and ST-ASR-MT-RT. Although the ST-ASR-MT mode showed relatively better alignment with human judgments, the overall performance was limited. The GEMBA approach produced rough scores without pinpointing specific errors or their severity. To address these limitations, a more structured prompting method—MIEval—was developed. This framework incorporates Word Error Rate (WER) for evaluating ASR and the Multidimensional Quality Metric (MQM) for assessing MT. Applied to the ST-ASR-MT and ST-ASR-MT-RT modes, MIEval significantly improved alignment with human evaluations, particularly in the ST-ASR-MT-RT mode, where reference texts supported more accurate error identification and severity analysis. The findings highlight that prompts with detailed evaluation criteria yield more reliable and informative results. Beyond evaluating technical performance, this study also demonstrates how educators can adapt the experimental approach for pedagogical purposes, offering strategies to train the next generation of interpreters in the age of AI.